Notes from the work of making autonomy reliable.
Practical writing for teams building agents and devices that must remember, perceive, respond, and improve in a world that will not hold still.
All articles
8 articles
The agent passed the demo. Then it lost the thread.
Model capability has moved quickly. Operational proof has not. The difficult work now lives between one impressive response and a reliable outcome.
Simulation infrastructure is a painkiller, not a research accessory
Production is the most expensive place to discover how an autonomous system behaves under pressure. Simulation moves that learning earlier.
Regression is not just a code problem
A model swap, tool change, or context policy can alter behavior without touching the business logic. Compare systems in the same world, not against different anecdotes.
The telemetry you need is not a single score
Scores decide whether to look. Trajectories explain what changed, who felt it, and which engineering lever can fix it.
Physical AI needs more than a photorealistic world
Better world models and vision-language-action systems raise the ceiling for robotics. They also make behavioral simulation more urgent.
Context engineering needs a world, not a longer prompt
The agent's context is an internal view. Reliability depends on whether that view remains aligned with a changing external world.
The industry is building scoreboards. Autonomous systems need worlds.
Static datasets tell you what a model knows. Operating worlds reveal what an autonomous system does after its action changes the situation.
Synthetic users are useful only when they disagree
Volume is easy. The value comes from credible variation in incentives, memory, constraints, and reactions across a customer population.