Question-led guide · decision
How should I compare observability platform data architectures?
Compare collection, normalization, storage, query, and retention boundaries through the investigations and failure modes the platform must support.
Direct answer
Compare architectures from required queries and service obligations, not from a single ingestion-rate claim. Map each signal from producer through collection, transformation, storage, index, and query path. Record where data can be lost, delayed, sampled, reinterpreted, or made too expensive to retrieve. Test mixed ingestion and investigation load with representative records. The appropriate architecture depends on those measured workload and failure conditions.
Begin at the query that must work
A trace search, long-range metric trend, log reconstruction, and service inventory lookup place different demands on storage and indexing. Write the operator’s query shape, freshness need, retention period, and expected concurrency. Without that contract, a fast ingest benchmark may optimize the wrong side of the system. Describe which result is acceptable when a signal is missing or delayed.
Mark every acknowledgement boundary
A producer may report export success when a collector accepted a batch, while the query store remains unavailable. A queue can absorb an outage only within its configured capacity and persistence behavior. Storage may acknowledge writes before derived views catch up. Draw these promises separately, including retry and replay effects, so buyers know which component owns loss, duplication, freshness, and recovery.
A million writes do not answer one incident
In a constructed test, a pipeline sustains 500,000 spans per second with no query traffic. During a simulated outage, replay adds 200,000 spans per second and investigators open twenty concurrent trace searches. P99 query time jumps from two seconds to forty. The ingestion headline remains true for its narrow test, but the mixed workload fails the incident journey the platform was bought to support.
Draw a boundary map before selecting technology
A compact map forces explicit ownership for the data path.
| Boundary | Promise to test | Failure example |
|---|---|---|
| Producer to collector | Accepted signal identity | Dropped export |
| Queue to store | Bounded retry and replay | Full buffer |
| Store to query | Freshness and availability | Merge contention |
| Projection to view | Preserved meaning | Double-counted derived row |
| User to result | Supported investigation | Missing context |
Test signal meaning alongside throughput
A platform that changes metric units, trace attributes, or sampling populations can return fast but misleading answers. Use representative payloads and compare stored records with source semantics. Test retention boundaries, high-cardinality filters, access isolation, and failure recovery under the same load. The architecture decision should include the cost of preserving the fields investigators actually need.
State the evidence behind the recommendation
Record versions, data volume, compression, query mix, resource limits, outage duration, and observed recovery. Separate vendor documentation from your measured result. A small laboratory can reveal a design flaw without proving production scale; a reference customer story may not represent your workload. Choose the simplest architecture that meets the named promises with a rehearsed failure path.
Evidence boundary for platform architectures
- OpenTelemetry Collector resiliency: OpenTelemetry documents sending queues, persistence, and retry behavior. It does not measure any candidate platform in this example.
- ClickHouse observability schema design: ClickHouse discusses query-driven schema choices for observability. Those design considerations do not prove a vendor-neutral performance result.
All workload rates and latencies are invented. Architecture acceptance requires representative local mixed-load tests.
Evidence
Collector queues and retries have explicit capacity and recovery limits.
OpenTelemetry documents sending queues, persistence, and retry behavior.
Primary source · official-doc · checked Oct 7, 2026
Limit: It does not measure any candidate platform in this example.
Observability storage design should follow actual query patterns.
ClickHouse discusses query-driven schema choices for observability.
Primary source · official-doc · checked Oct 7, 2026
Limit: Those design considerations do not prove a vendor-neutral performance result.
Limitations
The map identifies tests but cannot predict hardware sizing, vendor implementation details, or commercial cost without measurements.
FAQ
- Is ingest throughput enough to compare platforms?
- No. Test query latency, freshness, replay, meaning preservation, and recovery under the workload users actually need.
- Can one queue guarantee no telemetry loss?
- No. Its capacity, persistence, upstream behavior, and failure domain determine the protection it can provide.
Related guides
Continue within Enterprise observability platform selection, or use one of these adjacent diagnostics:
Editorial QA: automated native-English, structure, source-presence, and link checks completed . This record is not an independent expert endorsement. Review boundary.
