Topic map

Navigate by system boundary.

These hubs organize questions by the part of the production system that must change—not by whichever tool is currently popular.

01

AIOps alerting and operations

Evaluate alert quality, grouping, forecasting, diagnosis, and bounded recovery around real operations decisions.

02

LLM judge evaluation

Design task-specific rubrics, expert reference sets, error measures, stress tests, and calibration records for model judges.

03

Evolving semantic layers for data agents

Govern metric definitions, provenance, retrieval, semantic changes, and drift in data-agent workflows.

04

Enterprise observability platform selection

Compare platform architecture, investigation coverage, operating cost, proof-of-concept evidence, and migration readiness.

05

ClickHouse observability

Plan ClickHouse-backed telemetry around ingestion durability, signal meaning, investigative queries, capacity, and platform health.

06

Ontology-driven enterprise delivery

Connect business decisions to governed meaning, source mappings, operational authority, and reusable product capability.

07

Build your first app with AI

Follow a beginner path from a readable web page to a small app with checkable changes, recoverable data, and a deliberate release.

08

MCP in production

Operating Model Context Protocol services through accountable boundaries, catalog identity, service objectives, compatibility, admission, and remote dependency control.

09

Agent skills in production

Reusable agent capabilities with focused discovery, deliberate authoring, immutable releases, coexistence evidence, and governed improvement.

10

AI agent platform engineering

Production deployment, release compatibility, control-plane resilience, worker lifecycle, tenant scheduling, and admission for agent services.

11

AI agents for SRE and root cause analysis

Evidence-aware incident investigation across traces, logs, metrics, topology, and operational change.

12

AI cost engineering

Outcome units, request ledgers, model routing, caching, capacity, and the cost of failed attempts.

13

Context engineering, memory, and state

Production context assembly, memory governance, durable workflow state, provenance, and side-effect safety.

14

Engineering judgment with AI

Problem framing, constraints, failure design, verification, and deliberate learning with AI-assisted development.

15

AI agent evaluation

Task portfolios, trace grading, simulation, red teaming, stochastic reliability, and production quality gates.

16

SEO and GEO for technical sites

Accurate discovery, retrieval, understanding, evidence absorption, citation, and measurement across search and answer engines.

17

AI agent observability

Tracing model calls, retrieval, tools, memory, delegation, state transitions, evaluations, and human review.

18

AI agent security

Prompt injection containment, tool permissions, identity, sandboxing, memory poisoning, and audit trails.

19

Forward deployed engineering

Turning urgent field problems into production outcomes, reusable learning, and scalable product capability.

20

Ontology and operational semantics

Identity, relationships, state, provenance, and lifecycle semantics across CMDB, telemetry, and operational systems.