Question-led guide · decision
When should an agent trajectory propose a semantic change?
Distinguish a candidate gap in semantic definitions from retrieval, query, data-quality, and user-intent failures in agent traces.
Direct answer
Use a failed or corrected agent run as a lead, not an automatic rewrite request. Reconstruct the user decision, selected definition, source records, query plan, result, and correction. Classify whether the failure came from missing meaning, stale retrieval, bad mapping, data quality, or ambiguous intent. Propose the smallest semantic change only when independent evidence supports it, with an owner, expected benefit, affected consumers, and tests.
Treat correction as evidence with an uncertain cause
A user saying “that number is wrong” does not identify which layer failed. The agent may have misunderstood the question, selected the wrong existing metric, joined on an unstable key, read delayed data, or used a valid definition that the user disputes. Preserve the original request, selected semantic revision, query, source state, and correction before assigning a cause.
Compare the run with a known-good path
Ask whether a governed definition already exists and whether the agent retrieved it. If it did, inspect whether the query obeyed its grain and filters. If it did not, test the retrieval index before editing the ontology. A candidate definition change is warranted only when the current approved rule cannot express the real decision and independent domain evidence supports a replacement.
The missing metric was already present
In a fictional revenue analysis, an agent uses “gross order value” for a net-revenue question and returns a larger total. The user corrects it. An automatic editor proposes adding a new net-revenue class, but the semantic catalog already contains a governed net metric; its synonym was absent from retrieval. The team fixes the lookup and adds a test before considering any semantic model change.
Use a triage card for candidate changes
The card keeps failure evidence separate from proposed authority.
| Finding | Evidence to keep | Disposition |
|---|---|---|
| Intent ambiguity | User question and clarified goal | Ask or document |
| Retrieval miss | Candidate rankings and index revision | Repair retrieval |
| Query violation | SQL and definition revision | Repair generation |
| Source defect | Data-quality record | Correct source |
| Semantic gap | Owner examples and counterexamples | Review candidate |
Bound the proposal and its blast radius
If a real gap remains, name the exact rule to add or revise, the questions it should improve, and the existing consumers that may change. Keep the old definition available for comparison. The agent may draft a change packet, but a domain owner must inspect source support and approve the release boundary. Do not infer broad transfer from one corrected run.
Test the classification itself
Review a sample of failed tasks with human experts and measure how often triage labels match the underlying repair. Include cases where no semantic update is needed. A system that proposes changes for every failure can create costly drift. Preserve rejected proposals as evidence for improving the diagnostic process rather than silently retrying a different ontology edit.
Evidence boundary for semantic change proposals
- OpenLineage object model: OpenLineage records jobs, runs, datasets, and associated metadata. It does not classify business-semantic failures in agent runs.
- NIST AI Risk Management Framework: NIST connects governance, measurement, and management of AI risks. It does not authorize autonomous ontology edits.
The revenue failure is constructed. A real change needs the task trace, source data, and accountable domain review.
Evidence
Lineage distinguishes dataset and job events that help reconstruct a data task.
OpenLineage records jobs, runs, datasets, and associated metadata.
Primary source · official-doc · checked Oct 7, 2026
Limit: It does not classify business-semantic failures in agent runs.
AI changes require governed and measured risk decisions.
NIST connects governance, measurement, and management of AI risks.
Primary source · official-doc · checked Oct 7, 2026
Limit: It does not authorize autonomous ontology edits.
Limitations
Trajectory evidence can be incomplete or biased by user feedback. The card does not replace domain authority or regression testing.
FAQ
- Can user correction automatically update the ontology?
- No. It identifies a candidate problem; classify the failure and require independent source and owner review.
- What if the right definition already exists?
- Repair retrieval or query construction, then add a regression case so the agent selects it next time.
Related guides
Continue within Evolving semantic layers for data agents, or use one of these adjacent diagnostics:
Editorial QA: automated native-English, structure, source-presence, and link checks completed . This record is not an independent expert endorsement. Review boundary.
