
MCP Is Becoming Infrastructure: What the 2026 Specification Changes
September 16, 2026
A2A 1.0: Why Agent-to-Agent Communication Needs a Real Protocol
September 19, 2026An agent does not improve merely because it performs more tasks. Experience becomes capability only when outcomes are evaluated, lessons are represented in the right layer and changes pass regression tests.
Collect evidence during work
Preserve the task, relevant context, tool calls, result, errors, human corrections and final outcome. Do not treat every successful trajectory as a best practice. A result may have succeeded by luck, hidden human intervention or an external condition that will not repeat.
Diagnose the first useful failure
Find where the trajectory first departed from acceptable behavior. Was the source missing, the prompt unclear, the wrong tool selected, an argument invalid or a permission boundary absent? Fixing the final wording will not repair an earlier retrieval or tool problem.
Choose the right representation
- Update knowledge when a sourced fact or reference is missing.
- Update instructions when the desired process is unclear.
- Create or revise a skill when a reusable procedure needs examples, scripts or references.
- Change code when a rule must be deterministic.
- Consider fine-tuning only when repeated evaluated examples show a behavior-level gap.
This ordering favors changes that are inspectable and reversible.
Separate candidates from production
New memories, prompts, skills and code should enter a candidate area. Evaluate them on cases they are meant to improve and cases they must not affect. Require security review for new dependencies or tools. Promote only after the evidence supports the change.
The agent must not be able to rewrite the validators, thresholds, audit logs or rollback copies that judge its own updates. Otherwise it can make regression look like progress by changing the test.
Consolidate in batches
Online execution should focus on the current job and append evidence. Offline review can merge duplicate lessons, identify conflicts, retire stale guidance and propose small updates. Batch review provides enough context to distinguish a pattern from one unusual incident.
Measure the loop
Track task success, safety violations, human review time, latency, cost and rollback frequency. Improvement in one metric can hide damage elsewhere. A faster agent that creates more correction work may not be better.
When self-improvement is unsafe
Do not allow autonomous updates where outcomes are ambiguous, feedback is delayed or the same agent controls both production and evaluation. Human reviewers must define the objective and approve high-impact capability changes.
Use Agent Observability to collect evidence, Agent Skills to package procedures and How to Evaluate an AI Agent to validate updates.
Primary sources
- OpenAI: Evaluate agent workflows
- OpenAI: Testing Agent Skills with evals
- Anthropic: Effective harnesses for long-running agents
Adaptation note: This article was informed by the continual-evolution framework in AI Agents in Depth: Design Principles and Engineering Practice by Bojie Li and contributors, Apache License 2.0. It was independently rewritten and expanded into an operational improvement process for Stariy.com.



