AI Clusters and the New Economics of Data Center Optics: A Conversation with Cisco's Bill Gartner
Key Highlights
- Optical transceivers are becoming more economically significant as network speeds increase, with optics sometimes costing more than the switch ports they support at 400G and above.
- AI infrastructure is divided into three tiers—scale-up, scale-out, and scale-across—each with distinct networking requirements and increasing reliance on high-capacity optical links.
- Link reliability is crucial in AI networks; errors can cause significant performance drops, making optical stability a key measure of compute efficiency.
- The industry is moving toward 1.6T optics, with early deployments expected, driven by hyperscalers, to improve cost and power efficiency per bit.
- Emerging architectures like co-packaged optics aim to reduce power and increase port density but introduce operational challenges related to component failure and interoperability.
For much of the modern data center era, optics occupied a supporting role in the network. Optical transceivers connected switches, routers and facilities, but were generally treated as accessories to the more consequential infrastructure around them.
Artificial intelligence is overturning that calculation.
The enormous bandwidth requirements of AI training clusters have made optics one of the most economically significant—and operationally sensitive—components in the infrastructure stack. As networks progress from 400-gigabit and 800-gigabit connections toward 1.6 terabits per second, the cost of the optics can now exceed the cost of the switch port they support.
At the same time, a faulty connection can leave some of the world’s most expensive computing infrastructure waiting for a network to recover.
“AI has really changed the relevance of optics in the industry,” said Bill Gartner, vice president and general manager of Cisco’s Optical Systems and Optics business, during a recent appearance on the Data Center Frontier Show podcast.
At 10 gigabits per second, Gartner said, optics accounted for roughly 10% of the bill of materials associated with a network port. At 100G, that share increased to approximately 50%.
“At 400 gig and above, optics can actually be more expensive than the cost of the port itself,” he said.
That economic inversion is unfolding as data center operators face an equally consequential power equation. Every watt consumed by the network is a watt unavailable to the GPUs producing the actual training or inference output.
Interconnect is essential, Gartner noted, but its value lies in keeping compute infrastructure operating rather than performing the computation itself. That gives network designers a strong incentive to reduce the power and cost of moving each bit.
“We’d like to drive that to as close to zero as we can,” he said. “Obviously, we’re not going to get to zero, but reducing the power penalty in that interconnect drives to much better solutions for GPUs and CPUs.”
Three Networks Inside the AI Factory
Understanding the optical challenge begins with recognizing that an AI cluster contains several distinct networking environments.
Gartner divided AI infrastructure into three broad tiers: scale-up, scale-out and scale-across. Each operates over a different distance, carries a different level of traffic and creates a different set of requirements for the interconnect.
Scale-up describes the connections within a rack, where operators place as many GPUs as possible inside servers and then pack those servers into the available rack footprint. Gartner estimated that the bandwidth within this environment can be approximately 500 times that of a traditional wide-area network application.
Scale-up connections still rely heavily on electrical interfaces because electrical interconnects remain relatively inexpensive and power efficient over short distances.
Once the compute capacity of a rack has been exhausted, the cluster must expand into additional racks. This is the scale-out network, where 400G and 800G pluggable optics connect large numbers of GPU systems operating in parallel.
Gartner characterized scale-out bandwidth as roughly 50 times the capacity associated with a conventional WAN environment.
The third tier, scale-across, emerges when a data center reaches its practical power limit and the AI infrastructure must extend into another facility. Those data centers may be separated by tens or hundreds of kilometers, requiring coherent optical technology capable of carrying extremely high-capacity signals over longer distances.
Scale-across networks can represent approximately 14 times traditional WAN bandwidth, according to Gartner.
Taken together, the three tiers illustrate why optics has become inseparable from the AI infrastructure discussion. Network capacity must expand inside the rack, across rows of racks and increasingly between separate data centers—all without consuming an untenable share of the power budget or introducing failures that leave GPUs idle.
When One Link Slows the Whole Cluster
The reliability requirement for AI networks differs sharply from that of conventional enterprise or internet infrastructure.
In a traditional TCP/IP network, an error may trigger a packet retransmission. The application or user may never realize that the error occurred.
AI training clusters operate differently. GPUs work on large problems in parallel and must remain synchronized. When an error occurs on one link, the impact can extend across the cluster.
“All of the GPUs need to stop and back up to a checkpoint and restart,” Gartner said. “Because everything is running in parallel in a GPU infrastructure, a link error on one link actually impacts neighboring links.”
Those failures can be caused by a bad connection, a dirty optical connector, noise on the link or other physical and electronic problems. The resulting instability is often described as a link flap—a connection repeatedly dropping and recovering.
Gartner said information Cisco has received from hyperscale customers indicates that link flaps can reduce the efficiency of GPU infrastructure by as much as 40%.
That figure changes the practical meaning of optical reliability. A transceiver is no longer simply one component among thousands in the network. An unstable link can undermine the utilization of the far more expensive accelerators attached to it.
“We have every incentive to make sure that the optics are super reliable,” Gartner said.
The economics are stark. Operators are investing billions of dollars in AI campuses designed to keep accelerators running at the highest possible utilization. A network issue that forces thousands of GPUs to stop, return to a checkpoint and restart can erode the productivity of the entire deployment.
In that environment, optical reliability becomes a compute-efficiency metric.
1.6T Moves Toward Deployment
Most large AI training networks are currently being built around 400G and 800G connectivity. Gartner described those technologies as ubiquitous within hyperscale training infrastructure, with each serving different portions of the network.
The next step is already approaching.
Cisco is sampling 1.6-terabit optics with customers, and Gartner said the first volume deployments are expected to begin later this year. As with previous transitions from 10G to 100G, 100G to 400G and 400G to 800G, the appeal lies in improving cost and power efficiency per bit.
“There’s no question that there’s a demand for 1.6T,” he said. “We’re looking beyond 1.6T as well, as to what happens next.”
Adoption will remain uneven across customer segments.
Hyperscalers are driving most of the demand and volume for 800G and 1.6T technology. Traditional service providers remain earlier in their 400G and 800G deployments, while 400G is still in the early stages within many enterprise networks.
That distinction matters because hyperscale purchasing volume will help determine the technology’s cost curve. Gartner advised service providers and enterprises to closely watch the architectures being adopted by the largest cloud and AI infrastructure operators.
“If you want to be on the low-cost curve, you should be on the hyperscaler cost curve,” he said.
The Coming Electrical-to-Optical Transition
The progression from 400G to 800G and then 1.6T is significant, but Gartner views it as a familiar generational evolution. The more disruptive transition will occur when electrical connections can no longer carry the bandwidth required within the scale-up network.
The constraint comes down to physics.
As the bitrate of a signal increases, the distance it can travel generally decreases. Electrical interconnects remain attractive inside the rack because they are inexpensive and efficient, but the industry will eventually reach a point where they cannot support the required capacity over even short distances.
“My bet is ultimately on an optical solution,” Gartner said.
Several potential architectures are under development, including co-packaged optics, near-packaged optics, wide electrical buses using slower-rate signals and radio-frequency technologies.
Co-packaged optics and near-packaged optics move optical components away from the front panel of the switch and closer to the switching silicon. Instead of transmitting a high-speed electrical signal several inches across a line card to a pluggable transceiver, the system converts the signal into light nearer to the chip.
That shorter electrical path can reduce the need to repeatedly retime the signal, lowering power consumption. Removing pluggable modules from the faceplate could also increase port density because optical connectors require less space than complete transceivers.
Cisco expects co-packaged solutions to enter the market this year, giving operators an opportunity to gain practical experience with the architecture.
But the transition introduces operational tradeoffs.
Pluggable optics are replaceable, interoperable and available from multiple vendors. When one fails, a technician can replace the affected module. In a co-packaged design, the optical components may be mounted directly to a line card or integrated with the switch silicon.
That raises a more difficult question: What happens when a single optical channel fails?
Operators may need to rely on software to route around the failed link, replace a larger portion of the system or reconsider how they stock and service network equipment. Questions involving multi-vendor sourcing and component interoperability will also have to be resolved.
Gartner expects those operational issues to be worked through over the next several years before co-packaged or near-packaged optics become commonplace in scale-up networks.
Routed Optics Reduce the Scale-Across Penalty
Optical architecture is already changing how operators connect data centers.
Traditional long-distance optical networks use transponder line cards housed in dedicated chassis. The transponder receives an optical signal from a switch or router, converts it into a form capable of traveling over a long distance and combines it with other wavelengths on the same fiber.
Through its acquisition of Acacia, Cisco developed an architecture known as routed optical networking. The approach replaces the standalone transponder line card with a coherent pluggable optic inserted directly into the router.
Gartner said the architecture can consume approximately 90% less power than the conventional transponder-based approach while requiring less space and reducing cost.
Hyperscalers were among the earliest adopters, using coherent pluggables to connect AI infrastructure across data centers. The technology is now moving into service-provider networks and enterprise campus environments.
For AI operators, routed optical networking makes scale-across architectures more practical. When the power ceiling at one location prevents a cluster from expanding further, coherent optical links can extend the training environment into another facility.
Rather than acting as a bottleneck, Gartner said, optics is helping AI networks continue to scale beyond the physical and electrical limits of a single building.
“We’re going to see increasing use of multiple data centers involved in a single training infrastructure,” he said.
Coherent Optics Moves Inside the Data Center
The boundaries between scale-out and scale-across may also become less distinct.
Conventional data center optics commonly use intensity-modulation direct-detection technology, or IMDD, for connections rated over distances of approximately two kilometers. As bitrates continue increasing, maintaining that reach becomes more difficult.
The physical distance between switches does not become shorter simply because the network has moved to a faster generation. Operators expect each new technology to support the cabling and layouts already in place.
At some point, Gartner said, conventional data center optics may no longer be able to maintain the required reach. The industry may then need to adopt techniques from coherent optics, which have historically been used for metro, regional and long-haul networks.
“The likelihood is that we’ll have to borrow some of the coherent techniques that are used outside the data center and bring them in the data center,” he said.
That evolution could further blur the distinction between a data center campus, a metro cluster and geographically distributed AI infrastructure. Optical systems capable of supporting connections beyond 1,000 kilometers are already extending the practical reach of data center interconnection.
Bringing coherent techniques inside the facility would create new power and cost considerations, but it could allow operators to continue increasing bitrates without redesigning every physical network around shorter reaches.
Inference Changes the Optimization Target
Training clusters currently command much of the industry’s attention because of their enormous scale and capital requirements. Gartner expects the next major deployment wave to arrive as trained models move into inference.
Inference environments generally require smaller systems than frontier training clusters, but they will be deployed much more broadly across hyperscalers, service providers and enterprises.
The optimization target will therefore change.
Training places a premium on scale and the ability to connect enormous numbers of GPUs. Inference will place greater pressure on cost efficiency, power efficiency and the ability to support specialized models closer to users and applications.
That will keep optics at the center of the infrastructure equation, even as the type and size of the deployment changes.
For data center operators, the underlying lesson is that networking can no longer be treated as a secondary layer added after decisions about compute, power and facilities have been made. Optical cost, power consumption, reach and reliability are becoming design inputs for the AI factory itself.
As Gartner put it, optics is not limiting the growth of AI networks today. It is one of the technologies allowing them to move beyond the power and geographic boundaries of a single data center.
At Data Center Frontier, we talk the industry talk and walk the industry walk. In that spirit, DCF Staff members may occasionally use AI tools to assist with content. Elements of this article were created with help from OpenAI's GPT5.
Keep pace with the fast-moving world of data centers and cloud computing by connecting with Data Center Frontier on LinkedIn, following us on X/Twitter and Facebook, as well as on BlueSky, and signing up for our weekly newsletters using the form below.
About the Author
Matt Vincent
Matt Vincent is Editor in Chief of Data Center Frontier, where he leads editorial strategy and coverage focused on the infrastructure powering cloud computing, artificial intelligence, and the digital economy. A veteran B2B technology journalist with more than two decades of experience, Vincent specializes in the intersection of data centers, power, cooling, and emerging AI-era infrastructure. Since assuming the EIC role in 2023, he has helped guide Data Center Frontier’s coverage of the industry’s transition into the gigawatt-scale AI era, with a focus on hyperscale development, behind-the-meter power strategies, liquid cooling architectures, and the evolving energy demands of high-density compute, while working closely with the Digital Infrastructure Group at Endeavor Business Media to expand the brand’s analytical and multimedia footprint. Vincent also hosts The Data Center Frontier Show podcast, where he interviews industry leaders across hyperscale, colocation, utilities, and the data center supply chain to examine the technologies and business models reshaping digital infrastructure. Since its inception he serves as Head of Content for the Data Center Frontier Trends Summit. Before becoming Editor in Chief, he served in multiple senior editorial roles across Endeavor Business Media’s digital infrastructure portfolio, with coverage spanning data centers and hyperscale infrastructure, structured cabling and networking, telecom and datacom, IP physical security, and wireless and Pro AV markets. He began his career in 2005 within PennWell’s Advanced Technology Division and later held senior editorial positions supporting brands such as Cabling Installation & Maintenance, Lightwave Online, Broadband Technology Report, and Smart Buildings Technology. Vincent is a frequent moderator, interviewer, and keynote speaker at industry events including the HPC Forum, where he delivers forward-looking analysis on how AI and high-performance computing are reshaping digital infrastructure. He graduated with honors from Indiana University Bloomington with a B.A. in English Literature and Creative Writing and lives in southern New Hampshire with his family, remaining an active musician in his spare time.




