
Context Engineering for Long-Running AI Agents
September 25, 2026
Agent Observability: Traces Metrics and Finding the First Wrong Step
September 26, 2026An AI agent becomes operationally risky when it can turn uncertain interpretation into an external action. The central safety question is therefore not “how intelligent is the model?” It is “what can this system read, change, send, spend, disclose, or delete when it is wrong?”
Safety starts with permissions and recovery, not with a promise that the model will behave perfectly.
Build a threat model around consequences
List the assets the agent can reach: private files, customer data, email, publishing systems, financial tools, code, infrastructure, or public accounts. Then list credible failure paths: misunderstood instructions, stale context, malicious content, over-broad tools, exposed secrets, incorrect target selection, repeated actions, and missing confirmation.
Rank the consequences. Reading a public page is not equivalent to sending a public message. Drafting a change is not equivalent to applying it. The control should match the consequence.
Separate read tools from write tools
Use distinct tools and permissions for reading, drafting, and committing. A research agent may need broad access to public sources but no authority to publish. A deployment assistant may prepare a change set while a human or separately controlled process applies it.
Do not give write access merely to make a demonstration smoother. Friction at a consequential boundary is often a safety feature.
Apply least privilege by default
Grant only the minimum data, functions, targets, and duration required for the current task. Prefer scoped credentials, restricted directories, test environments, and allow-listed destinations.
Permissions should expire or be revocable. A one-time migration does not justify standing access after the migration is complete.
Treat external content as untrusted
Documents, webpages, emails, tickets, and retrieved text can contain instructions that conflict with the user’s intent. The agent should treat that material as data, not as authority.
Keep system rules, user authorization, and retrieved content separate. Do not let a document expand its own permissions, reveal secrets, disable safeguards, or redefine the target.
Confirm consequential actions at action time
Approval is strongest when the user sees the exact action, target, important parameters, and expected consequence immediately before execution. An early statement such as “help manage the site” is not approval for every later publication or deletion.
Require confirmation for public messages, irreversible changes, purchases, permission changes, sensitive disclosures, and other high-impact actions. When a series of actions is approved, define its boundaries and stop conditions.
Keep secrets outside the reasoning surface
Use approved secret stores or active authorized sessions. Do not place passwords, tokens, recovery codes, cookies, or private keys into prompts, logs, source files, or generated artifacts.
Redact logs so they remain useful without becoming a second secret store. Log the credential identifier or permission class, not the secret value.
Make actions auditable
For every material action, record who authorized it, what changed, the target, time, result, and relevant version. The log should distinguish a proposed action from an executed one.
Auditability supports incident response and learning. It does not justify collecting unrelated private data.
Use isolation and staging
Run untrusted or experimental work in a bounded environment. Test with synthetic or minimized data where practical. For publishing or infrastructure, use staging before production and prevent an experimental tool from silently crossing the boundary.
Isolation should also limit blast radius: one site, one project, one directory, one account, or one transaction class rather than a global permission.
Design rollback as part of the action
Before granting write access, identify what rollback means. It may be restoring a previous version, reverting a change set, moving a page back to draft, disabling a tool, revoking a credential, or switching to a manual process.
Some actions cannot be fully reversed. A sent email, exposed secret, public disclosure, or external purchase may require containment rather than rollback. Label those actions accordingly and demand stronger approval.
Prepare incident response
Define how the agent or operator stops work, preserves evidence, revokes access, identifies affected targets, notifies the responsible person, and verifies recovery. Do not let the same agent that caused a high-impact incident decide alone that the incident is resolved.
A practical permission matrix
For each tool, document:
- allowed operations;
- prohibited operations;
- permitted targets;
- data classification;
- confirmation requirement;
- logging requirement;
- rollback or containment method;
- permission owner;
- expiry or review condition.
Begin with read-only access. Add drafting. Add bounded writes only after the evaluation, approval, logging, and recovery controls are credible.
Agent safety is not a single prompt or filter. It is the combined design of authority, data boundaries, tool separation, confirmation, observability, and recovery.
Related reading: AI Systems, Technology Foundations, and From AI Demo to Reliable Workflow.



