
A2A 1.0: Why Agent-to-Agent Communication Needs a Real Protocol
September 19, 2026
From AI Demo to Reliable Workflow
September 22, 2026AI accelerator comparisons often begin with a single large number: TOPS, FLOPS or a vendor's claimed speedup. These figures can be useful, but only after the workload and measurement conditions are known. A system that performs enormous amounts of low-precision arithmetic may still be the wrong choice if the model does not fit in memory, data moves too slowly or the required software path is immature.
The practical question is not “Which chip has the biggest number?” It is “What prevents this workload from producing the required result at the required latency, cost and power?”
Capacity decides whether the model fits
Model weights, runtime state and the key-value cache all consume accelerator memory. Longer contexts and more concurrent requests increase the working set. If that set does not fit, the system may need to split the model across devices or move data through a slower path. Either choice adds communication and operational complexity.
Current data-center products illustrate why capacity has become a headline specification. NVIDIA documents 141 GB of HBM3e for H200, while its enterprise material lists higher capacities for later Blackwell systems. AMD documents up to 288 GB of HBM3E for its MI350 series. Those figures do not prove which device wins a workload; they show that keeping larger working sets close to compute is now a central design goal.
Bandwidth decides how quickly data reaches compute
Large language model inference repeatedly moves weights and state through memory. During some phases, arithmetic units can wait for data rather than for more theoretical compute. High-bandwidth memory is designed to reduce that bottleneck.
NVIDIA lists 4.8 TB/s for H200 and up to 8 TB/s for B200/B300 configurations. AMD lists up to 8 TB/s for MI350. These are peak specifications under defined product conditions. They are best used to identify a possible constraint, not to predict application throughput on their own.
Precision changes both speed and meaning
Lower-precision formats can reduce memory use and increase throughput, but results depend on model support, quantization method, software kernels and acceptable quality loss. A quoted FP4 or INT8 figure should not be compared directly with an FP16 requirement as if the workloads were identical.
Ask which precision the model actually uses, whether accuracy was re-evaluated after conversion and which parts of the workflow remain at higher precision. Peak low-precision compute is valuable only when the application can use it without breaking its quality target.
The benchmark must resemble the job
MLCommons describes MLPerf Inference as an open, peer-reviewed and architecture-neutral benchmark suite. Its value is not that one table permanently ranks all hardware. Its value is that systems are tested against specified models, scenarios and rules.
Interactive serving, offline batch inference, image generation and speech recognition stress different parts of a system. Even within language models, input length, output length, batch size and latency constraints can change the result. Compare submissions only when the workload, accuracy target and scenario align with your intended use.
A five-part hardware decision
- Fit: Can the model, cache and concurrency target fit in available memory?
- Movement: Is bandwidth or interconnect likely to constrain the workload?
- Precision: Which data types are supported end to end without unacceptable quality loss?
- Software: Are the framework, kernels, drivers and deployment tools mature for this model?
- System result: What are measured latency, throughput, power and cost under your own traffic pattern?
Hardware specifications narrow the search. A representative evaluation makes the decision. For smaller on-device workloads, continue with Local AI Hardware: NPU, GPU or Cloud?. For the broader evaluation method, read How to Evaluate an AI Agent.
Primary sources
- NVIDIA H200 product specifications
- NVIDIA HGX enterprise reference architecture
- AMD CDNA architecture and MI350 specifications
- MLPerf Inference v5.1
Source and adaptation note: This article is an original Stariy.com hardware explainer. General system-design framing was informed by AI Agents in Depth: Design Principles and Engineering Practice by Bojie Li and contributors, distributed under Apache License 2.0. The text and analysis were independently developed.



