Traditional application logs tell you that a request failed. Agent observability must also explain the trajectory: what the model saw, which tool it selected, what the tool returned, where a handoff occurred and which control allowed the final result.
Long-running agents fail when their working context becomes noisy, stale or incomplete. Context engineering treats attention as a managed system resource.
An agent does not improve merely because it performs more tasks. Experience becomes capability only when outcomes are evaluated, lessons are represented in the right layer and changes pass regression tests.
An agent does not gain useful memory merely because a conversation is long. Useful memory is selected, stored, retrieved and corrected outside the momentary model call.