Industrial sensor data — vibration, temperature, pressure, throughput readings from field equipment — arrives out of order, at uneven rates, and from devices whose clocks aren't perfectly synchronized. A pipeline designed around batch-ETL assumptions (data arrives in order, a "day" is a clean boundary, late data is rare enough to ignore) will silently produce wrong aggregates. Three decisions determine whether the pipeline holds up under real field conditions: how you handle event-time versus processing-time, how you define late-arrival tolerance, and how you absorb burst load without losing data or falling permanently behind.
Event Time vs. Processing Time Is the First Decision, Not a Detail
Every reading from a sensor has (at least) two timestamps that matter: event time — when the reading actually happened on the device — and processing time — when your pipeline received and processed it. In a batch pipeline these are close enough to treat as the same thing. In a streaming pipeline over industrial sensors, they diverge routinely: a device buffers readings during a network outage and sends a burst once connectivity returns; two devices on the same physical process disagree by a few seconds because their clocks weren't synced; a gateway aggregates readings from several sensors and forwards them with variable latency.
Aggregating by processing time when the analysis needs event-time semantics — "what was the average vibration reading between 2:00 and 2:05" — produces numbers that are wrong in a way that's hard to detect after the fact, because the pipeline ran without error; it just windowed the data by the wrong clock. Streaming frameworks that support explicit event-time windowing (assigning each record to a window based on its embedded timestamp, not its arrival time) exist specifically for this. The design decision has to be made explicitly and early — retrofitting event-time semantics onto a pipeline built around processing-time windows usually means rebuilding the windowing logic, not patching it.
Watermarks: Defining How Late Is Too Late
Once you commit to event-time windowing, you need a policy for a question batch ETL never has to answer: when is a window "done," given that a record from that window might still be in flight? A watermark is the mechanism for this — a declared bound stating "we don't expect event-time data older than X to arrive after this point," which lets the pipeline emit a window's aggregate once the watermark passes the window's end, while still holding a defined grace period open for stragglers.
The mistake in both directions is common. Too tight a watermark, and legitimately late data — a sensor's buffered burst after reconnecting — arrives after the window has already closed and gets dropped or routed to a late-data path that downstream consumers don't check. Too loose a watermark, and windows stay open far longer than the data justifies, which delays every downstream aggregate and increases the pipeline's state-retention requirements, since it has to keep window state around for however long the watermark tolerates lateness.
Getting this right requires actually characterizing the lateness distribution of the sensor fleet — how long a device typically buffers during an outage, how skewed device clocks actually are — rather than picking a watermark value that feels reasonable. A pipeline tuned against assumed lateness rather than measured lateness tends to either drop real data or lag unnecessarily, and both failure modes look identical from the outside: numbers that don't match what the field equipment actually reported.
Backpressure and Burst Absorption
Industrial sensor volume is not steady-state. A process upset condition, a batch of equipment coming online simultaneously, or a network partition resolving and flushing buffered readings all produce burst load that can be an order of magnitude above steady-state throughput. A pipeline sized only for average or even peak-observed-so-far volume will fall behind during a burst it hasn't seen yet, and "falling behind" in a streaming context isn't graceful degradation — it's accumulating backlog that has to be worked off later, during which downstream consumers are seeing stale or delayed data without any signal that anything is wrong.
Designing for this means an explicit backpressure strategy at each stage: a message broker (Kafka or equivalent) sized with retention long enough to absorb bursts without data loss, consumers that can scale horizontally rather than being pinned to a fixed partition count that caps throughput, and — critically — monitoring on consumer lag itself, not just on whether the pipeline is "running." A pipeline that's technically up but steadily accumulating lag looks identical to a healthy one on a simple liveness check.
Schema Drift From the Field, Not From a Deploy
Batch ETL schema changes are usually deliberate — someone changes a table definition and deploys it. Sensor fleets drift for reasons that have nothing to do with your deploy cycle: a firmware update on a subset of devices adds a field or changes units, a new sensor model gets deployed alongside older ones and reports a slightly different payload shape, a device sends a malformed reading during a fault condition. A pipeline that assumes a fixed schema will either crash on the first unexpected payload or, worse, silently coerce it into something that parses without error but represents the wrong value.
The practical mitigation is treating schema validation as an explicit pipeline stage — reject or quarantine records that don't match the expected contract, with visibility into what's being quarantined and why — rather than assuming upstream data always conforms. For a fleet that gets firmware updates and hardware refreshes on a rolling basis, "the schema is fixed" is an assumption that will be violated during the pipeline's normal operating lifetime, not an edge case.
Where Exactly-Once Actually Matters, and Where It's Overkill
Streaming systems get a lot of design attention paid to exactly-once processing guarantees, and it's worth being deliberate about where that guarantee is actually load-bearing for industrial sensor analytics. A raw vibration or temperature reading feeding a rolling average is often tolerant of occasional duplication or loss — the statistical aggregate barely moves. A discrete event — an alarm trip, a threshold breach that triggers a downstream action — is not tolerant of duplication (a double-triggered shutdown) or loss (a missed alarm). Applying uniform exactly-once machinery across the whole pipeline when only the event/alarm path needs it adds latency and operational complexity to the high-volume, loss-tolerant path for no benefit. Separating the two — at-least-once or best-effort for continuous telemetry, exactly-once (via idempotent writes keyed on a stable event ID) for discrete triggered events — is usually the better-calibrated design.
The Same Root Cause, Five Different Symptoms
The failure pattern across all of the above is the same: assumptions that were safe in batch ETL — ordered arrival, a stable schema, roughly steady volume, processing time as a proxy for event time — don't hold for a live industrial sensor fleet, and a pipeline built without explicitly re-examining each of them will pass every test against clean synthetic data and then produce quietly wrong aggregates the first time a device buffers a burst, a firmware update ships, or a network partition resolves at an inconvenient moment.