Reliable agents
Recover cleanly when tools, providers, and state fail.
Co-founder & CTO · Systems builder
I design AI systems that stay reliable under load, predictable under failure, and efficient under constraint.
Recover cleanly when tools, providers, and state fail.
Control spend across a workflow, not one request at a time.
Treat agent behavior as software that can be verified.
Building production-ready agent systems.
Led Kubernetes migration at company scale.
2018 Amazon Alexa Prize winner.
Selected work / 01
The interesting problems begin after the demo: traffic spikes, provider failures, runaway costs, and workflows that need to resume cleanly.
Operating principles / 02
Design recovery, fallback, and observability before the happy path ships.
Optimize whole workflows—not isolated requests—for quality, latency, and cost.
Treat agents like applications: reproducible, idempotent, and accountable.
Field notes / 03
Why per-request routing misses the point—and what workflow-level cost control looks like.
Why autonomous agents need deterministic testing, and how idempotency failures surface in production.
The case for treating agent systems as software—with integration tests, not just evals.
Toolbox / 04