
AI Agent Safety: Permissions and Rollback
September 26, 2026
How to Evaluate an AI Agent Before It Can Act
September 27, 2026Traditional application logs tell you that a request failed. Agent observability must also explain the trajectory: what the model saw, which tool it selected, what the tool returned, where a handoff occurred and which control allowed the final result.
Trace the observable system
OpenAI’s Agents SDK documentation describes traces containing model calls, tool calls, handoffs, guardrails and custom spans. A useful trace also records timing, status, validated parameters, source identifiers and approvals.
Do not rely on private chain-of-thought as the operational record. The evidence is in observable inputs, decisions, actions and outputs.
Three levels of signals
Outcome metrics: Was the task completed correctly? Was the user or downstream system satisfied?
Process metrics: Were the correct sources and tools used? How many retries, handoffs and tokens were required?
Safety metrics: Did the agent request approval, respect data boundaries and avoid prohibited actions?
An outcome can be correct despite a dangerous process. A safe process can still produce a useless answer. Monitor both.
Find the first wrong step
Suppose a final report cites the wrong price. The last generation step may only repeat an earlier error. Trace backward until the first unsupported state appears: the wrong source was retrieved, a date filter was missing or a tool returned ambiguous units. Repair that boundary rather than adding a final instruction to “be accurate.”
Design privacy before logging
Traces may contain prompts, documents, tool arguments and personal data. Collect the minimum needed for debugging and evaluation. Apply access controls, retention limits and redaction. Keep credentials outside trace payloads. Observability that exposes secrets is a production defect.
Build operational views
Track failure rate by task type, tool and version. Watch latency distributions rather than averages. Record unknown token counts as unknown, not zero. Compare model or prompt changes on the same evaluation set. Link important failures to a reproducible case and the change that fixed them.
Alerts need action
Alert on repeated tool failures, unusual write volume, budget exhaustion, missing approvals and safety-rule violations. Every alert should name an owner and a response. Otherwise it becomes background noise.
When full tracing is unnecessary
A low-risk, one-step transformation may need only normal application logs and sampled outputs. Increase trace detail with autonomy and consequence. Use configurable sampling or privacy-preserving summaries rather than removing observability entirely.
Continue with How to Evaluate an AI Agent, How the ReAct Loop Works and Continual Improvement.
Primary sources
Adaptation note: This article was informed by evaluation and observability concepts in AI Agents in Depth: Design Principles and Engineering Practice by Bojie Li and contributors, Apache License 2.0. It was independently rewritten and expanded for Stariy.com.



