Question-led guide · decision
Should an OpenTelemetry pipeline use persistent queues or a broker?
Choose collection buffers by failure domain, outage duration, replay ownership, and acknowledgement meaning instead of treating every queue as durable delivery.
Direct answer
Use a persistent Collector queue for a bounded interruption when the queue storage survives the failure you intend to cover. Add a broker when independent replay, longer retention, fan-out, or a wider failure boundary justifies its operating cost. For either design, identify what acceptance acknowledges and who recovers work after capacity or storage is lost.
Start with the interruption, not the queue product
Suppose a fictional team loses telemetry whenever the database pauses for maintenance. Its first question is how long the database may be unavailable and whether the producers continue running. A local persistent queue may bridge that interruption. If the same requirement includes losing the entire producer node, a disk attached only to that node does not cover the larger failure.
Write down the acceptance boundary
Draw the receiver, queue, exporter, database adapter, and query service. Above every successful response, write the state it proves. Accepted into a local queue, committed to a broker, stored by a database, and visible to a reader are separate milestones. A dashboard that reports only successful export requests cannot tell an operator how far the records actually traveled.
Make capacity a time budget
Queue capacity becomes meaningful when paired with a traffic distribution. A request-count limit hides differences between small batches and large ones. Estimate the bytes or records generated during the required outage, then test with realistic batching. Decide what happens when the queue fills: rejection, producer blocking, dropping, or spill to another bounded store. Each choice moves a consequence somewhere.
Pay for a broker only when its contract is needed
A broker can separate producers from several consumers and retain work for controlled replay. It also adds partition management, lag monitoring, storage, security, and recovery duties. Identify the specific obligation it owns before adopting it. A broker that acknowledges data cannot by that fact prove that a downstream materialized view is current or that an investigator can see a complete trace.
| Requirement | Candidate boundary | Remaining question |
|---|---|---|
| Collector process restart | Persistent sending queue | Does its volume survive? |
| Extended database outage | Sized queue or broker | How much backlog accumulates? |
| Independent consumers | Broker with replay policy | Who owns offsets and duplicates? |
| Entire node loss | Storage beyond that node | Which copies survive together? |
Rehearse recovery with known records
Send an identifiable sequence, interrupt the selected component, restore it, and reconcile expected versus observed records. Exercise queue exhaustion and a lost response after acceptance. Test explicit rejection separately from an ambiguous timeout. For OTLP partial success, follow the protocol’s response semantics rather than blindly resending every accepted record in the batch.
Monitor the backlog as part of the product
Track oldest queued age, disk reserve, export failures, dropped records, and the time needed to return to normal freshness. Show investigators when admitted data is still waiting upstream. The deployment is ready when an owner can explain both its covered outage and its loss boundary, then demonstrate recovery under those conditions. The mere presence of a queue configuration is not that evidence.
Evidence and scope
- OpenTelemetry Collector resiliency: Collector sending queues can use persistent storage, while queue capacity and storage survival bound recovery.
- OTLP specification: OTLP defines response, retry, and partial-success behavior at the receiving endpoint.
The proposed checks are teaching tools; validate their behavior in the actual environment.
Evidence
Collector sending queues can use persistent storage, while queue capacity and storage survival bound recovery.
Collector sending queues can use persistent storage, while queue capacity and storage survival bound recovery.
Primary source · official-doc · checked Sep 11, 2026
Limit: This source supports the named mechanism, not the outcome or thresholds of the illustrative workflow.
OTLP defines response, retry, and partial-success behavior at the receiving endpoint.
OTLP defines response, retry, and partial-success behavior at the receiving endpoint.
Primary source · official-doc · checked Sep 11, 2026
Limit: This source supports the named mechanism, not the outcome or thresholds of the illustrative workflow.
Limitations
The scenarios and decision worksheets are original teaching examples. They are not measured deployments or guarantees; adapt the checks to the actual system and its documented behavior.
FAQ
- Is a persistent queue equivalent to a replicated broker?
- No. Storage survival, replay, retention, consumer isolation, and operating responsibilities differ.
- Can a successful export prove that a trace is queryable?
- Only if the documented endpoint contract includes that condition and the implementation verifies it; otherwise measure visibility separately.
Related guides
Continue within ClickHouse observability, or use one of these adjacent diagnostics:
Editorial QA: automated native-English, structure, source-presence, and link checks completed . This record is not an independent expert endorsement. Review boundary.
