Skip to content
Corentin GS

Out-of-Order Traces: Redis Timestamp Coordination Model

№55 · · ·1073 words ·5 min read
In this piece

A child span can reach the aggregate processor before its root. If the root’s timestamp falls in the previous month, publishing both rows under their own timestamps splits the trace across monthly partitions.

A trace is one logical event with many delivery times

OpenTelemetry traces are trees. A root span records the operation; child spans record database calls, model calls, tool calls, and internal work. Spans with one trace ID can arrive in any order. The streaming OTLP parser decodes spans one at a time, so downstream code receives separate span records.

Spans from the same trace can reach different workers. Kafka preserves order within a partition, but a child can still arrive before its root.

The previous ClickHouse schema article defines an aggregate projection for trace-level queries. ClickHouse orders those rows by dimensions including date and timestamp. If a trace’s children retain their individual timestamps while the root crosses a month boundary, one trace can be spread across partitions.

Queries must scan every relevant partition, and retention jobs must account for rows from one trace split across partition boundaries.

Use each span’s timestamp for span-level events. A trace-level projection needs one stable timestamp, so the pipeline needs an explicit policy for choosing it.

The small state model

Two Redis keys coordinate each trace:

KeyValueLifetimeJob
trace:timestamp:{traceID}Root timestamp in Unix milliseconds24 hoursMakes the root’s time available to later children.
trace:pending:{traceID}JSON-serialized child aggregate rows1 hourHolds children received before the root.

The client must set both timeouts explicitly; Redis supplies neither default.

  1. T0 · Child

    Child · T0Child span finishes
  2. T1 · Child

    Child · T1Child reaches aggregator
  3. T2 · Redis

    Redis · T2Root timestamp missing
  4. T3 · Redis

    Redis · T3Buffer child for up to 1 hour
  5. T4 · Root

    Root · T4Root reaches aggregator
  6. T5 · Redis

    Redis · T5Store root timestamp for 24 hours
  7. T6 · Redis

    Redis · T6Rewrite and publish child
Timeline: the child reaches storage before its root

The child cannot adopt the root timestamp because the root has not yet reached the processor, so it waits in Redis until the root arrives.

When a root arrives

In this model, the root handler stores the root’s Unix-millisecond timestamp in Redis with a 24-hour TTL. It publishes the root to the aggregate Kafka topic, retrieves buffered children, and publishes them with the root timestamp.

A later child looks up the root timestamp. If the key exists, the processor parses its value, replaces the child’s timestamp, and publishes it.

Child first: buffer until the root arrives

When a child arrives first, the timestamp lookup misses. The processor serializes the aggregate row, pushes it onto trace:pending:{traceID}, then sets the list TTL to one hour.

text
child: Get trace:timestamp:{id} → missing
child: LPush trace:pending:{id} = serialized child
child: Expire trace:pending:{id}, 1h

Each successful Expire after LPush resets the pending-list TTL. LPush alone leaves any existing TTL unchanged; a failure between the two calls can leave a newly created list without an expiry.

When the root arrives, the processor reads the list with LRange, rewrites each child’s timestamp, and publishes it.

The pending-list flush is non-transactional

The pending-list flush has three separate operations:

  1. Read the list with LRange.
  2. Publish each rewritten child to Kafka.
  3. Delete the list.

No transaction spans the Redis list operations and Kafka publications. The gaps between those operations create duplication and loss windows.

If the process crashes after publishing some children and before deleting the list, a later retry can publish them again. Downstream consumers need an idempotency or deduplication policy.

If publication fails during the loop and the handler still deletes the list, that child is lost. If the handler leaves the list for retry after publishing some children, those children can be duplicated.

A child can miss the timestamp, then LPush after the root calls LRange but before Delete. The root’s LRange does not include that child, and Delete removes it.

  1. PROCESSRoot processor
    1. LRangeto Pending list
    2. publish buffered childrento Aggregate topic
    3. Deleteto Pending list
  2. PROCESSLate child
    1. LPush after readto Pending list
  3. KAFKAAggregate topic
  4. REDIS LISTPending list
The pending-list flush has a gap

Redis reads and deletes the list separately from Kafka publication; a late child or partial failure falls between those operations.

This two-key model covers root-first and child-first delivery, but it leaves duplication and loss windows around partial flushes.

TTLs determine failure retention

An orphaned pending list expires without publishing its children. A child that arrives after the root timestamp expires takes the child-first path, even if the root was already processed.

Monitor expiry failures, Redis-side pending-list cardinality, and the oldest pending span’s age. An in-process count resets on restart and cannot show how many lists Redis holds. Without a size limit, repeated children can exhaust Redis memory; choose an overflow policy.

Test paths and their interleavings

Test root-first publication, child-first buffering, flush-on-root, and later children adopting the stored timestamp. A broker that always accepts publishes cannot expose a failed second publish.

Inject failures after one child has published, between LRange and Delete, and when Expire runs; restart after a partial flush. For each case, decide whether work is retried, dropped, or dead-lettered.

Alternative coordination strategies

Accept partition skew. Keep each span’s original timestamp and pay for cross-partition trace assembly in queries and maintenance.

Reconcile at query time. Normalize presentation after lookup, while aggregate tables retain the original spread.

Use a compacted Kafka state topic. Keep root timestamps as durable, replayable state, at the cost of restoration, compaction lag, and another ordering problem.

Redis requires less operational machinery while coordination state is short-lived and its failure windows are acceptable.

Start with the invariant, then choose the cache

Write the invariant before choosing the cache:

All aggregate rows for one trace use the root span’s timestamp when the root is known within the coordination window.

Then test root-first and child-first delivery, a missing root, Redis failures, failed flush publication, the LRange/delete race, and a restart after partial publication.

Choose the pending-list failure policy before choosing Redis: lost children on expiry, duplicate rows on retry, or a durable write path with different operational cost.

Sources

Explore this subject

More on Devlog