Working notes from the frontier of agent deployment.
Need an agent that still works after launch?
We build production AI agents with the evals, routing, tools, review loops, and runbooks that keep them useful in the real world.
Agent Instruction Files Are Source Code: Versioning AGENTS.md and CLAUDE.md Like Config
AGENTS.md and CLAUDE.md are production config with no failing tests. How to add ownership, review, drift checks, and regression tasks before they rot.
Thinking Budgets Are a Product Decision: Tuning Reasoning Effort per Task, Not per Model
Reasoning effort settings decide cost, latency, and accuracy per request. How to tune thinking budgets by workflow step instead of one global default.
Read more ->Durable Agents: Checkpointing, Resumption, and Why Long Tasks Need a Workflow Engine
Durable execution for AI agents explained: what to checkpoint, why idempotency comes first, and how Temporal, Restate, and Inngest fit under agent loops.
Read more ->Agent Simulation Testing: Put a Simulated User in Front of Your Agent Before Real Traffic Does
Multi-turn agents fail on trajectory, not single responses. How to build a user simulator, seed it from real transcripts, and version it like a test fixture.
Read more ->News
No posts match that search yet.