- Inside the Sandbox: Unconstrained Shell Execution
Exploring why allowing agents arbitrary host shell access leads to non-deterministic failures and why strict Docker isolation is mandatory for production.
2 min
- A Plan Can Validate and Still Be Unsafe to Implement
A controlled indexed-planner study shows why structural validation, path grounding, and semantic review must be reported as separate gates.
5 min
- Designing Bounded Repair Loops for Agent Plans
A bounded repair action makes plan validation observable and safe — but repair telemetry must be tied to the final outcome, not counted as success itself.
3 min
- What a Two-for-Two Agent Result Actually Proves
A 2/2 completion result is useful evidence of a cohort outcome; it is not automatically proof of a mechanism or a production default.
4 min
Back