An agent that reads is a toy. An agent that writes is a risk.
No side effect without explicit confirmation.
"Just put it in the prompt" isn't enough
Js
- Confirmation is a required field. The model can't call the tool without stating, explicitly, whether the user agreed.
- The tool refuses without it, and the refusal message tells the model what to do instead. The model reads the error, asks the caller "Shall I move it to Thursday at 10?", and tries again once they say yes.
The other three guardrails
1. Fall back instead of failing
2. Hand off when unsure
3. Log every decision
A checklist for any AI agent you're about to deploy
- Which tools change something? List them.
- Does each of those tools refuse to run without confirmation, in code?
- What happens when the model or an API is down?
- Where does the agent hand off to a human?
- Can you see what it did and why, after the fact?
Planning an agent that will act on your real systems? Let's make it safe first.