Question-led guidance

Start with the engineering question.

Every guide states a direct answer, defines its scope, shows the diagnostic path, and names its limitations.

Browse by problem domain. Every reviewed guide owns a separate engineering decision or failure mode, evidence set, and working artifact.

Technical domain

AIOps alerting and operations

Evaluate alert quality, grouping, forecasting, diagnosis, and bounded recovery around real operations decisions.

Technical domain

LLM judge evaluation

Design task-specific rubrics, expert reference sets, error measures, stress tests, and calibration records for model judges.

Technical domain

Evolving semantic layers for data agents

Govern metric definitions, provenance, retrieval, semantic changes, and drift in data-agent workflows.

Technical domain

Enterprise observability platform selection

Compare platform architecture, investigation coverage, operating cost, proof-of-concept evidence, and migration readiness.

Technical domain

ClickHouse observability

Plan ClickHouse-backed telemetry around ingestion durability, signal meaning, investigative queries, capacity, and platform health.

Technical domain

Ontology-driven enterprise delivery

Connect business decisions to governed meaning, source mappings, operational authority, and reusable product capability.

planning · for enterprise architect and forward deployed engineer

What should a thin ontology production slice prove?

Choose one decision-shaped delivery slice and test evidence, mappings, authority, failure handling, user value, and operational handoff before expanding.

Technical domain

Build your first app with AI

Follow a beginner path from a readable web page to a small app with checkable changes, recoverable data, and a deliberate release.

Technical domain

MCP in production

Operating Model Context Protocol services through accountable boundaries, catalog identity, service objectives, compatibility, admission, and remote dependency control.

Technical domain

Agent skills in production

Reusable agent capabilities with focused discovery, deliberate authoring, immutable releases, coexistence evidence, and governed improvement.

comparison · for agent capability owner

When should a workflow become an agent skill?

Decide whether recurring work needs a reusable agent skill, a prompt, a deterministic program, or a tool by writing the capability contract first.

Technical domain

AI agent platform engineering

Production deployment, release compatibility, control-plane resilience, worker lifecycle, tenant scheduling, and admission for agent services.

Technical domain

AI agents for SRE and root cause analysis

Evidence-aware incident investigation across traces, logs, metrics, topology, and operational change.

Technical domain

AI cost engineering

Outcome units, request ledgers, model routing, caching, capacity, and the cost of failed attempts.

how to · for AI FinOps engineer

How do I calculate cost per AI outcome?

A fully loaded outcome-ledger method that includes accepted, rejected, reversed, and reviewed AI work instead of dividing one provider invoice by request count.

decision · for AI application architect

What should an LLM application cache?

A decision framework for prompt prefixes, retrieval results, tool reads, embeddings, responses, and derived artifacts under staleness and privacy constraints.

governance · for AI product owner

What belongs in an AI cost release gate?

A release record that ties AI unit cost to quality, reliability, sensitive-data handling, capacity, fallbacks, owners, and rollback thresholds.

Technical domain

Context engineering, memory, and state

Production context assembly, memory governance, durable workflow state, provenance, and side-effect safety.

how to · for agent platform engineer

How do I allocate a context-window budget?

A context-frame budget for instructions, authority, state, evidence, tools, history, uncertainty, and output reserve in production agents.

how to · for RAG platform engineer

How should retrieval preserve provenance?

An evidence supply chain connecting source records, versions, chunks, retrieval runs, context placement, claims, citations, and downstream effects.

Technical domain

Engineering judgment with AI

Problem framing, constraints, failure design, verification, and deliberate learning with AI-assisted development.

diagnostic · for early-career developer

AI helps me code faster—why am I learning less?

A judgment-loop diagnosis for developers who are shipping faster with AI but retaining less understanding, transfer, debugging skill, and design confidence.

diagnostic · for backend developer

Why can retries make an incident worse?

A failure exercise for retry amplification, duplicate effects, synchronized load, stale work, budget exhaustion, and unsafe compensation.

Technical domain

AI agent evaluation

Task portfolios, trace grading, simulation, red teaming, stochastic reliability, and production quality gates.

evaluation · for AI evaluation lead

How can I simulate users without fooling myself?

A validation card for agent-evaluation user simulators covering persona distributions, hidden goals, behavior calibration, leakage, and transfer to real users.

Technical domain

SEO and GEO for technical sites

Accurate discovery, retrieval, understanding, evidence absorption, citation, and measurement across search and answer engines.

comparison · for technical content strategist

GEO vs SEO: what actually changes?

A responsibility matrix separating durable SEO foundations from the evidence design and answer-level observations added by generative search experiences.

decision · for technical publisher

Is llms.txt worth adding to a technical site?

A bounded experiment for llms.txt that preserves accessible HTML, sitemaps, robots controls, analytics, maintenance, and honest success criteria.

governance · for technical publisher

Which AI crawlers should a publisher allow?

A crawler-purpose matrix separating search discovery, user-requested fetching, training, and other automated access while preserving security boundaries.

planning · for GEO program owner

What should a 90-day GEO program deliver?

A practical 90-day GEO program that ships technical access, evidence-backed content, measurement baselines, distribution experiments, and an executive decision.

Technical domain

AI agent observability

Tracing model calls, retrieval, tools, memory, delegation, state transitions, evaluations, and human review.

how to · for observability engineer

What belongs in an AI agent trace schema?

A field-level design for tracing agent runs across models, retrieval, tools, policy decisions, outcomes, cost, and privacy boundaries.

Technical domain

AI agent security

Prompt injection containment, tool permissions, identity, sandboxing, memory poisoning, and audit trails.

diagnostic · for agent platform engineer

Can a system prompt stop prompt injection?

Why instruction wording cannot be the only prompt-injection control, and how to build boundaries around evidence, tools, authority, and effects.

Technical domain

Forward deployed engineering

Turning urgent field problems into production outcomes, reusable learning, and scalable product capability.

definition · for engineering leader

What is a forward-deployed engineer, really?

A decision-oriented definition of forward-deployed engineering, including ownership, customer proximity, product feedback, and boundaries with consulting.

Technical domain

Ontology and operational semantics

Identity, relationships, state, provenance, and lifecycle semantics across CMDB, telemetry, and operational systems.