Question-led guide · how-to
How should a data agent retrieve definitions while working?
Resolve task-relevant semantic definitions by identity, time, scope, and provenance before an agent builds a data query.
Direct answer
Retrieve a governed definition at the moment the agent needs to decide a metric, entity, or join. Match the user’s question to a scoped semantic identifier and effective time, then return the rule, source, owner, and conflicts. Keep retrieved text as evidence rather than authority to rewrite the definition. If several candidates remain, ask for clarification or expose uncertainty before generating SQL or a business conclusion.
Retrieve when a decision needs meaning
Loading an entire business glossary into every prompt wastes space and still leaves the relevant definition ambiguous. Detect the metric, entity, or join choice that the task requires. Use a stable identifier when available; otherwise collect candidate definitions with their domains and owners. The agent should not generate an answer from the first semantic hit merely because the text sounds familiar.
Scope identity by domain and effective time
Two groups can use “active customer” with different inclusion rules. A definition is selected by more than a name: domain, owner, validity interval, permitted audience, and source revision all matter. Retrieve these fields together. If the user asks about last year’s cohort, a current definition may be the wrong one even when it is the newest and most popular result.
Two teams share a metric label
In a constructed company, the sales team counts an account active after a signed contract, while the product team requires recent usage. A user asks for active accounts in an adoption report. A naive retriever returns the sales definition because its name is an exact match. The agent compares task domain and report purpose, then selects the product definition or asks which population the user intends.
Run a retrieval checklist before query construction
The checklist is an interface between search and decision.
| Check | Required evidence |
|---|---|
| Identity | Metric or entity ID, not only label |
| Scope | Domain, population, and permitted user |
| Time | Effective interval for the question |
| Authority | Owner and approved revision |
| Conflict | Competing definitions and reason |
| Use | Query field mapping and citation |
Return conflicts as data
When two governed definitions fit, provide their distinguishing rules and ask the user which decision they are making. When a source is missing, say that the selection cannot be verified. A retrieval score ranks candidate text, not business authority. Keep the result linked to its original record so corrections and version changes can be traced into later agent runs.
Check the answer against the chosen rule
Inspect the generated query’s grain, filters, time zone, and joins against the retrieved definition. Execute a boundary case where competing definitions yield different results. Log the selected semantic revision with the task, while protecting sensitive query data. A successful query execution establishes syntax and access, not that the chosen meaning matches the user’s decision.
Evidence boundary for semantic retrieval
- dbt semantic models: dbt documents semantic-model elements that can be retrieved and applied. It does not resolve this fictional organization’s conflicting labels.
- W3C PROV-O: PROV-O defines relations for entities, activities, and attribution. Provenance records do not determine which definition the user intended.
The competing active-customer definitions are invented. Retrieval ranking and authorization need local implementation evidence.
Evidence
Semantic models define entities, dimensions, and measures for consistent use.
dbt documents semantic-model elements that can be retrieved and applied.
Primary source · official-doc · checked Oct 7, 2026
Limit: It does not resolve this fictional organization’s conflicting labels.
Provenance can link a retrieved definition to its source and responsible agent.
PROV-O defines relations for entities, activities, and attribution.
Primary source · standard · checked Oct 7, 2026
Limit: Provenance records do not determine which definition the user intended.
Limitations
The checklist presumes governed semantic records. It cannot compensate for absent ownership, stale indexes, or wrong source mappings.
FAQ
- Should I always choose the newest definition?
- No. The question may refer to a historical period or a different domain; use effective time and scope.
- Can a retrieval score break a business-policy tie?
- No. It measures text relevance, not policy authority; present the conflict or ask for clarification.
Related guides
Continue within Evolving semantic layers for data agents, or use one of these adjacent diagnostics:
Editorial QA: automated native-English, structure, source-presence, and link checks completed . This record is not an independent expert endorsement. Review boundary.
