Topic map
Navigate by system boundary.
These hubs organize questions by the part of the production system that must change—not by whichever tool is currently popular.
AIOps alerting and operations
Evaluate alert quality, grouping, forecasting, diagnosis, and bounded recovery around real operations decisions.
02LLM judge evaluation
Design task-specific rubrics, expert reference sets, error measures, stress tests, and calibration records for model judges.
03Evolving semantic layers for data agents
Govern metric definitions, provenance, retrieval, semantic changes, and drift in data-agent workflows.
04Enterprise observability platform selection
Compare platform architecture, investigation coverage, operating cost, proof-of-concept evidence, and migration readiness.
05ClickHouse observability
Plan ClickHouse-backed telemetry around ingestion durability, signal meaning, investigative queries, capacity, and platform health.
06Ontology-driven enterprise delivery
Connect business decisions to governed meaning, source mappings, operational authority, and reusable product capability.
07Build your first app with AI
Follow a beginner path from a readable web page to a small app with checkable changes, recoverable data, and a deliberate release.
08MCP in production
Operating Model Context Protocol services through accountable boundaries, catalog identity, service objectives, compatibility, admission, and remote dependency control.
09Agent skills in production
Reusable agent capabilities with focused discovery, deliberate authoring, immutable releases, coexistence evidence, and governed improvement.
10AI agent platform engineering
Production deployment, release compatibility, control-plane resilience, worker lifecycle, tenant scheduling, and admission for agent services.
11AI agents for SRE and root cause analysis
Evidence-aware incident investigation across traces, logs, metrics, topology, and operational change.
12AI cost engineering
Outcome units, request ledgers, model routing, caching, capacity, and the cost of failed attempts.
13Context engineering, memory, and state
Production context assembly, memory governance, durable workflow state, provenance, and side-effect safety.
14Engineering judgment with AI
Problem framing, constraints, failure design, verification, and deliberate learning with AI-assisted development.
15AI agent evaluation
Task portfolios, trace grading, simulation, red teaming, stochastic reliability, and production quality gates.
16SEO and GEO for technical sites
Accurate discovery, retrieval, understanding, evidence absorption, citation, and measurement across search and answer engines.
17AI agent observability
Tracing model calls, retrieval, tools, memory, delegation, state transitions, evaluations, and human review.
18AI agent security
Prompt injection containment, tool permissions, identity, sandboxing, memory poisoning, and audit trails.
19Forward deployed engineering
Turning urgent field problems into production outcomes, reusable learning, and scalable product capability.
20Ontology and operational semantics
Identity, relationships, state, provenance, and lifecycle semantics across CMDB, telemetry, and operational systems.
