If the key phrase for AI infrastructure over the previous few years was merely extra compute, the large-scale deployment of Agentic AI is forcing the business to confront a extra sensible query: How effectively can all that compute truly be used?
On the 2026 Apsara Convention, H3C put the query by way of Tokens. As AI brokers constantly name on basis fashions, trillion-Token-scale providers are rising as a brand new infrastructure requirement.
Merely including extra GPUs doesn’t essentially ship proportional efficiency good points. Idle compute, community congestion, inadequate knowledge provide, and the operational and failure challenges of large-scale clusters can all erode the worth of extra compute.
That’s the reason H3C’s product lineup on the occasion, whereas spanning a variety of applied sciences, follows a comparatively clear logic: from compute to networking and storage, and from software program to operations, AI infrastructure is transferring past {hardware} stacking towards system-level optimization.

Extra GPUs don’t essentially imply extra efficiency
Zhu Shiyin, normal supervisor of H3C’s Superior Know-how Analysis Division, described AI clusters as extremely built-in techniques that require shut hardware-software coordination.
As cluster sizes develop from a whole lot of GPUs to 1000’s and even tens of 1000’s, GPUs themselves turn into only one a part of the equation. Chip-to-chip communication, knowledge switch, process scheduling, energy provide, and cooling can all immediately have an effect on general compute utilization.
H3C’s UniPoD S80000 Collection SuperPod, showcased on the occasion, displays this strategy. It helps configurations starting from 32 to 1,024 GPUs and might scale additional to 16,384 GPUs, whereas supporting heterogeneous computing assets together with CPUs, GPUs, NPUs, and DPUs and mainstream frameworks.
The main focus will not be merely on placing extra chips right into a single system, however on coordinating compute, networking, storage, cloud, safety, and operations in order that various kinds of computing assets might be managed as a unified system.
In different phrases, the competitors in AI clusters is shifting from what number of GPUs a system has to what number of helpful Tokens every GPU can truly produce.

As GPUs get quicker, the community can turn into the bottleneck
Zhu additionally highlighted high-speed interconnects, an more and more vital a part of AI infrastructure. The reason being simple: as GPU efficiency improves, the quantity of information exchanged between GPUs additionally grows. If the community can not sustain, GPUs are compelled to attend.
That is why H3C showcased three interconnect situations on the occasion: Scale-Up, Scale-Out, and Scale-Throughout.
Inside a node, H3C launched the S9828-128EO, a 102.4T NPO silicon-photonics clever computing swap that makes use of co-packaged optics to scale back end-to-end latency by 15%. Between nodes, the corporate’s 1.6T clever computing swap S9828-64FP, adopts 224G SerDes applied sciences, it cuts energy consumption by greater than 20% in high-density environments to assist large-scale cluster enlargement.
Throughout knowledge facilities, the 800G clever computing DCI swap S12500R-64EP is designed to assist intra-city cross-data-center compute scheduling and gradient synchronization.
The logic behind these three layers comes all the way down to a core query: As AI clusters develop, how can interconnects keep away from turning into the brakes on compute efficiency?
Interconnects are additionally not merely a matter of bandwidth. As clusters increase, protocols, congestion management, hyperlink reliability, and fault prognosis turn into equally vital. If distributors develop closed technical ecosystems independently, they might create new know-how silos. Open protocols and business requirements will subsequently turn into more and more vital.

Past compute, there’s one other continuously underestimated problem: knowledge
GPUs can not run with out knowledge. In large-model coaching, knowledge preparation, mannequin coaching, and parameter change all contain huge quantities of information motion. Throughout inference, KV Cache provides additional strain on storage and caching assets.
H3C subsequently additionally positioned high-performance storage as a part of the broader Token manufacturing pipeline. Its UniStor X20000 collection X20836 can ship as much as 200GB/s of bandwidth and three million IOPS per node, whereas supporting interoperability throughout block, file, object, and HDFS protocols.
In keeping with H3C, its full-speed engine can cut back GPU ready time by 30%, whereas XCache inference acceleration can reduce time-to-first-token latency by as much as 90% in inference situations.
The takeaway is pretty sensible: AI infrastructure is evolving from a compute heart right into a complete knowledge processing system. From knowledge preparation and coaching to inference and Token technology, each interval of ready finally interprets into value.

Software program stands out as the subsequent battleground
As soon as {hardware} reaches a sure scale, software program optimization turns into more and more vital. Zhu famous that all the things from the working system and communication libraries to useful resource administration and scheduling platforms must work intently with the underlying {hardware}.
Communication optimization, overlapping computation with communication, process scheduling, dynamic load balancing, and unified useful resource administration can all assist cut back GPU idle time and communication conflicts.
That is additionally why fashionable AI infrastructure more and more resembles a super-system. Chips present compute, networks join assets, storage provides knowledge, and energy and cooling preserve the system operating. Software program, in the meantime, is chargeable for orchestrating these assets right into a functioning complete.
If any considered one of these parts turns into a bottleneck, the impression is not restricted to a single {hardware} metric. It finally impacts what number of helpful Tokens the system can generate per second and the way a lot every Token prices to provide.
The following section of AI infrastructure is a contest over compute utilization
From H3C’s exhibition sales space to its discussion board classes on the Apsara Convention, the corporate was primarily making the identical level from completely different angles.
Supernodes deal with how compute assets might be organized at scale. Excessive-speed networking addresses how these assets can talk effectively. Excessive-performance storage ensures that knowledge might be delivered in time. Software program, scheduling, operations, and unified administration then flip these assets into AI capabilities that can be utilized constantly and effectively.
The following section of competitors in AI infrastructure, subsequently, is probably not decided just by who has extra GPUs. As mannequin sizes attain the trillion-parameter scale and AI brokers more and more name basis fashions, metrics resembling compute utilization, Token throughput, latency, vitality consumption, and price per Token have gotten more and more vital.
That’s the deeper that means behind the pursuit of maximum Token value effectivity: it’s not about placing extra GPUs collectively, however about making each GPU, each community hyperlink, and each unit of storage within the system spend much less time ready and extra time working.
As AI enters the stage of large-scale manufacturing, compute is not only a chip downside. It’s turning into a contest over the effectivity of the complete infrastructure system.