
AI Accelerators Beyond TOPS: Memory, Bandwidth and Real Workloads
September 21, 2026
Multi-Agent Systems: Collaboration Patterns Costs and Failure Modes
September 22, 2026An AI demo answers a narrow question: can this tool produce a convincing result once? A workflow must answer a harder question: can people use it repeatedly, with known inputs, measurable quality, controlled access, visible failures, and a recovery path?
The gap between those two questions is where many promising experiments fail. The model may be capable, but the surrounding process is undefined. Inputs arrive in inconsistent forms. Success is judged by impression. Permissions are broader than necessary. Edge cases appear only after someone depends on the output. When something goes wrong, no one knows whether to retry, correct, escalate, or roll back.
The way forward is not to add more automation immediately. First, define the work.
1. Name the user and the decision
Write one sentence that identifies who uses the workflow, what job they are doing, and what decision or action the output supports. “Summarize documents” is too vague. “Prepare a source-linked briefing that an editor reviews before publication” is bounded enough to evaluate.
Then list what is outside scope. A system that drafts a briefing is not automatically authorized to publish it. A system that recommends an action is not automatically authorized to execute it.
2. Define the input and output contract
Specify the allowed input types, required fields, privacy classification, source rules, and size limits. For the output, define structure, required citations, tone only where it matters, and the conditions that make the result unusable.
A contract makes variation visible. Without one, the same workflow may appear reliable simply because each person expects something different.
3. Establish a non-AI baseline
Describe how the work is done without the new system. Record the time, steps, review burden, common errors, and quality standard only when real evidence is available. The baseline does not need to be elegant; it needs to be comparable.
If there is no baseline, the first implementation goal should be observability. Do not claim improvement merely because the new output looks more polished.
4. Build a representative evaluation set
Create a small set of cases that reflects normal work, difficult work, incomplete inputs, conflicting sources, and inputs that should be rejected. Remove private data or use clearly labeled synthetic examples.
For each case, write acceptance criteria before reviewing the AI output. Criteria might include factual support, required fields, prohibited content, correct escalation, formatting, or a human decision. Avoid changing the criteria after seeing the result unless the change is documented.
5. Separate capability from permission
A model’s ability to call a tool does not mean it should receive that authority. Begin with read-only access and the smallest dataset required for the job. Separate drafting from sending, recommending from buying, and preparing a change from applying it.
Use explicit human confirmation for actions that are public, costly, destructive, difficult to reverse, or likely to affect another person. AI Agent Safety: Permissions and Rollback provides a fuller control model.
6. Design the human checkpoints
State who reviews the output, what they check, and what options they have. “Human in the loop” is not a control unless the human receives the evidence and has enough time and authority to reject the result.
Good checkpoints appear before consequences. A final review after an email has already been sent or a record overwritten is not meaningful approval.
7. Catalogue failure modes
List failures in observable terms: missing source, unsupported statement, wrong record, stale information, malformed output, unsafe instruction, unavailable dependency, excessive latency, or unexpected cost. For each one, define detection, containment, owner, and next action.
Do not use “the model made a mistake” as the entire diagnosis. Ask which boundary failed: input validation, retrieval, instruction design, permission scope, evaluation, review, or recovery.
8. Measure cost and latency in context
The relevant cost includes more than model usage. Count preparation, review, retries, exception handling, storage, tools, and maintenance when real measurements exist. The relevant latency is the time until an acceptable decision or output, not merely the model response time.
This draft does not assert benchmark values. Record dated measurements for the chosen implementation before making performance or savings claims.
9. Add monitoring and change control
Track the signals that reveal drift: rejection rate, unsupported claims, missing citations, escalation volume, failure by input type, cost per accepted result, and reviewer disagreement. Select only metrics tied to the workflow’s purpose.
Record model, prompt, tool, source, and policy changes. Re-run representative evaluations when a material dependency changes.
10. Plan rollback before expansion
Rollback may mean disabling write tools, returning to manual processing, restoring a previous configuration, reversing a staged change, or stopping a workflow while preserving its audit trail. Name the trigger and the person authorized to act.
Only expand scope after the current version has evidence of acceptable behavior, bounded failures, usable monitoring, and credible recovery.
A practical go/no-go review
Before relying on the workflow, ask:
- Is the user, job, and decision explicit?
- Are input and output contracts documented?
- Is there a representative evaluation set with prewritten criteria?
- Are permissions limited to what the current job requires?
- Do human approvals occur before consequential actions?
- Are failures detectable and assigned to an owner?
- Are cost and latency measured through accepted output?
- Can changes be traced and important evaluations repeated?
- Is rollback credible and tested at the appropriate level?
- Is the remaining risk acceptable to the responsible human?
A “no” does not always end the project. It tells you what must remain manual, what evidence is missing, or where the scope should shrink.
Related reading: AI Systems, Workflows, and Editorial Policy.



