Question-led guide · how-to

How should an operational graph represent change and time?

Keep observed topology, declared dependencies, deployments, and validity intervals distinct in an operational graph used for incident review.

Direct answer

Represent an operational relationship with its source, observation time, valid interval, and confidence rather than storing a timeless edge. Keep declared dependencies separate from runtime observations and inferred links. During an incident, query the graph as it was known at the relevant time and show gaps. A graph can organize candidates for investigation; it cannot turn proximity or a shared path into a causal finding by itself.

A topology model connects source and target services with valid interval, observation time, evidence source, and confidence.
Temporal edge: This schema is an author-designed operational representation. An edge supports investigation context, not a causal verdict. This is an author-created explanatory model, not measured system evidence.

A timeless edge rewrites the past

A graph is often displayed as a current set of nodes and lines. That view cannot answer what an operator knew during yesterday’s incident. A dependency may be added, removed, or simply discovered later. Preserve when the relationship was believed valid and when the system learned about it. These times are different when inventories lag behind deployments or telemetry arrives late.

Separate three kinds of relationship

A declared service dependency comes from configuration or ownership records. An observed call comes from runtime telemetry. An inferred relation comes from a rule or model comparing observations. Store their sources and confidence separately so a query cannot silently promote a possible link into an operating contract. Stable entity identity also matters: a reused service name should not combine two deployments into one enduring node.

A payment edge appears after the outage

Consider a fictional incident at 10:00. A new payment gateway was deployed at 13:00, and its trace edge was first seen at 13:04. The present-day graph includes that gateway. A retrospective query that ignores valid time incorrectly puts it in the morning failure path. The investigator filters by the 10:00 state, checks the older gateway, and records the data gap for any missing observations.

Give each edge an inspection record

Use a small schema before selecting a graph database.

Field Meaning
Source and target identity Scoped, versioned entities
Relation kind Declared, observed, or inferred
Valid interval When the link applied to the system
Observation time When evidence was recorded
Provenance Configuration, trace, or review reference
Status Supported, contradicted, or unknown

Reconcile without erasing disagreement

Compare inventory, deployment records, and trace observations for a selected incident window. A missing trace is not proof that a dependency did not exist: sampling and instrumentation gaps may explain absence. A declared edge with no observed traffic may be dormant or stale. Keep both claims, attach their timestamps, and request a targeted check when the difference could change the incident decision.

Ask a graph question it can answer

Use the graph to identify candidate affected services, owners, and evidence paths. Then test a causal hypothesis with chronology, contrasting healthy and failing cohorts, and mechanism-specific observations. If the relevant edge is inferred or its time interval is unknown, lower the strength of the conclusion. A plausible path is an investigation lead, not authorization to change production.

Evidence boundary for operational graph time

  • OpenTelemetry resource conventions: OpenTelemetry specifies resource attributes and identity conventions. Resource attributes do not encode historical dependency validity or cause.
  • OpenTelemetry signals: OpenTelemetry describes the signal types used to gather runtime observations. Combining signals in a graph requires local identity and time policies.

The gateway timeline is fictional. Implementations need actual deployment and telemetry timestamps and identity reconciliation.

Evidence

  1. Telemetry resources describe the entity producing observations through defined attributes.

    OpenTelemetry specifies resource attributes and identity conventions.

    Primary source · official-doc · checked Oct 7, 2026

    Limit: Resource attributes do not encode historical dependency validity or cause.

  2. Traces, metrics, and logs are different observation forms with distinct scopes.

    OpenTelemetry describes the signal types used to gather runtime observations.

    Primary source · official-doc · checked Oct 7, 2026

    Limit: Combining signals in a graph requires local identity and time policies.

Limitations

The schema is a design worksheet. It does not solve distributed clock uncertainty, missing telemetry, or identity collisions without system-specific controls.

FAQ

Can a trace edge replace a dependency inventory?
It can add observed evidence, but a missing edge may reflect sampling or inactivity and does not erase a declared dependency.
Does a path through the graph prove root cause?
No. It identifies a candidate relationship that still needs mechanism and counterevidence checks.

Continue within AIOps alerting and operations, or use one of these adjacent diagnostics:

Editorial QA: automated native-English, structure, source-presence, and link checks completed . This record is not an independent expert endorsement. Review boundary.