top of page

Scale-Across Networking Confronts the AI Data Center Cost Wall

Aug 11
12 min read

Google News surfaced a critical conflict on August 11: AI clusters need more space and power, but connecting multiple sites can erase their economic gains.

The underlying discussion came from Data Center Dynamics and an industry panel on data center interconnect, or DCI, which carries traffic between facilities. Google, Meta, Nokia, Lumen Technologies, and Zayo participated in the March 2026 session. Their shared problem was more consequential than another increase in headline bandwidth.

AI developers increasingly want multiple buildings to behave like one computing system. That approach, called scale-across networking, can overcome a single facility’s physical limits. However, it also exposes valuable accelerators to fiber failures, congestion, distance, and synchronization delays.

The core contest is therefore cost versus performance. Operators can place computing capacity near available power and distribute it across several sites. They must then build an optical network fast and dependable enough to keep expensive processors productive.

Google and Meta are already describing architectures that cross building boundaries. Equipment suppliers are promoting faster coherent optics, which encode data onto light for transmission over longer distances. None of that guarantees that distributed training will remain economical at production scale.

Google News Highlights a Network Architecture Moving Into Production

Scale-across networking is shifting from an industry proposal into a design requirement for the largest AI clusters.

The OFC panel program described DCI as the fastest-growing network segment, with annual growth exceeding 50 percent. It also identified scale-across as a distinct back-end use case rather than ordinary inter-site traffic.

Traditional DCI transports application, storage, replication, and customer traffic between facilities. Those flows can demand high capacity, yet they usually tolerate more variation than synchronized AI training.

A back-end AI network connects accelerators while they work on the same training or inference job. Each processor repeatedly exchanges model parameters, intermediate results, and synchronization messages. Delayed traffic can leave processors waiting even when local computing resources remain available.

Scale-across extends that back-end fabric beyond one building. The participating facilities form a larger computing domain, even though fiber and geographic distance separate them.

That distinction matters because operators are no longer connecting independent data centers only for resilience or data movement. They are attempting to coordinate accelerators across sites as components of one machine.

Nokia’s account of the panel separated AI connectivity into three dimensions. Scale-up connects processors within a rack, scale-out links racks inside a facility, and scale-across joins clusters beyond one facility.

These categories carry different engineering requirements. Short copper connections remain common inside racks, while optical links dominate longer paths. Scale-across adds coherent transmission, line systems, fiber routes, and operational controls to the AI fabric.

Google engineer Kevin Croussore joined representatives from Meta, Zayo, Lumen Technologies, and Nokia on the panel. Their participation showed that the problem spans cloud operators, network carriers, and equipment suppliers.

The timing is also important. In April 2026, Google introduced Virgo Network, its new scale-out fabric for the AI Hypercomputer. Google said training requirements now exceed the power and space available in individual data centers.

The Virgo architecture uses a flat, two-layer topology and independent control domains. Google designed it to reduce network tiers while providing predictable latency and fault isolation.

Virgo does not make every distant data center equivalent to a local rack. It does show that Google now treats a campus as one computing environment rather than a collection of isolated buildings.

That architectural change creates the article’s central tension. Breaking a cluster into multiple facilities can unlock power and floor space. It also introduces a network whose cost and behavior directly affect useful computing output.

Operators can no longer judge DCI by raw capacity alone. They must ask how much accelerator time the network preserves, how quickly it recovers, and how much power each transported bit consumes.

Power Scarcity Is Forcing AI Clusters Across Buildings

The pressure comes from a mismatch between concentrated computing demand and the slower expansion of power-ready data center capacity.

Large AI systems concentrate thousands of accelerators, memory devices, switches, cooling components, and power conversion systems. Adding processors also increases the supporting infrastructure needed to feed, cool, and connect them.

A single building eventually reaches limits imposed by utility delivery, substations, cooling systems, construction schedules, and local permitting. Operators can design denser racks, but density does not create additional grid capacity.

Moving data is often more practical than moving electrical power across long distances. Fiber therefore becomes a tool for combining separately powered facilities into a larger resource pool.

Google described this principle directly in May 2026. The company said it locates data centers near energy resources, then distributes AI workloads across campuses through its network.

Its global infrastructure reflects the scale of that investment. Google says its wide-area network bandwidth increased sevenfold between 2020 and 2025. That network supports consumer services, cloud traffic, training data, and AI computing resources.

Meta faces the same physical constraint at a different scale. Its Prometheus cluster is designed for one gigawatt of capacity across at least five data center buildings.

Prometheus also incorporates adjacent colocation space and temporary weatherproof structures, according to Meta’s infrastructure account. The arrangement makes the network part of the cluster’s basic construction plan.

Meta’s larger Hyperion project is expected to begin coming online in 2028. The company says the completed cluster will support up to five gigawatts.

Those figures demonstrate why scale-across is gaining attention now. Facilities of that size cannot depend on one conventional data hall or one local switching fabric.

Meta’s back-end aggregation system, known as BAG, connects multiple AI fabrics across facilities. BAG is a centralized Ethernet super-spine layer, meaning it links the upper levels of separate network fabrics.

The company says back-end aggregation will help interconnect tens of thousands of GPUs used by Prometheus. It must also bridge two different Meta fabric designs.

This approach pressures more than hyperscalers. Cloud providers need enough inter-site capacity to offer large accelerator pools. Carriers need routes with predictable latency, physical diversity, and faster installation.

Optical suppliers must improve capacity without letting power consumption, module cost, or failure rates grow at the same pace. Data center developers must also plan fiber paths before buildings open.

Enterprise buyers face an indirect consequence. Their AI service costs depend partly on how efficiently providers use accelerators. A distributed cluster with poor network utilization can waste expensive computing capacity.

The forced response is long term. Operators need to coordinate site selection, energy procurement, fiber routes, computing architecture, and workload software years before a large cluster enters service.

Scale-across networking gives them another design option, but it does not eliminate scarcity. It exchanges a concentrated power constraint for a distributed systems problem.

Cost-Effective Scale-Across Networking Depends on Useful GPU Time

The cheapest network component is not necessarily the network that produces the lowest training cost.

AI training proceeds through repeated computation and communication. Accelerators calculate results locally, exchange information, and wait for other participants before continuing.

This behavior makes tail latency unusually important. Tail latency measures the slowest responses within a group, not the average response time across all traffic.

One delayed link, congested path, or struggling node can hold back a much larger job. The cost appears as idle accelerator time rather than a simple networking charge.

That is why raw bandwidth offers an incomplete measure of scale-across economics. A high-capacity link can still deliver inconsistent completion times, weak failure recovery, or excessive power consumption.

Google designed Virgo around deterministic latency, high bisection bandwidth, and hardware fault isolation. Bisection bandwidth measures how much traffic can cross between two halves of a network simultaneously.

Those features aim to keep distributed tasks moving predictably. They also illustrate how AI networking differs from best-effort application traffic.

Meta is addressing the problem through network and software changes. Its Prometheus work includes Twine and MAST, systems intended to support training across geographically distributed data centers.

Scheduling becomes crucial because not every workload should cross every distance. An operator can place tightly synchronized tasks within one locality while distributing less sensitive stages across a wider region.

Coherent optics make longer high-capacity connections possible. The technology uses sophisticated modulation and signal processing to preserve data across fiber links extending beyond ordinary data center distances.

Modern coherent modules can fit into switch-facing pluggable formats. This reduces reliance on separate transport shelves for some applications and can simplify deployments.

However, the module’s purchase cost is only one part of network economics. Operators must consider fiber availability, line systems, amplification, monitoring, replacement labor, security, and spare capacity.

The network also consumes power that could otherwise support computing. A design that increases bandwidth while requiring more optical and switching equipment can intensify the site’s original power constraint.

Ciena advertises 400G, 800G, and 1.6T systems across its DCI portfolio. Marvell has similarly described 1.6T coherent modules intended for emerging scale-across deployments.

These roadmaps signal substantial capacity growth. They do not establish that every AI workload can use that capacity efficiently across multiple facilities.

Software must coordinate placement, routing, congestion control, checkpointing, and recovery. Operators also need visibility across computing and transport layers that were previously managed as separate systems.

A failed optical path should not force a large training job to restart from its beginning. The network and workload scheduler need coordinated recovery strategies that limit lost work.

That requirement changes how buyers should compare designs. Cost per transported bit remains useful, but cost per completed training step is closer to the business outcome.

Job completion time, accelerator utilization, optical power, and recovery duration belong in the same analysis. Treating them as separate procurement metrics can hide the true expense.

Scale-across becomes cost-effective only when the network preserves enough productive processor time to offset its added infrastructure. Faster optics support that goal, but architecture and operations determine the result.

Open Ethernet Faces Proprietary Performance Pressure

Scale-across networking intensifies the contest between interoperable Ethernet systems and tightly controlled accelerator fabrics.

Ethernet has an enormous installed base, a broad supplier market, and mature operational tools. Those advantages can reduce vendor concentration and give operators more component choices.

Meta’s BAG design uses Ethernet to aggregate different back-end fabrics. That decision shows how a hyperscaler can preserve a common interconnection layer while developing specialized systems underneath it.

Google also builds vertically across chips, systems, network fabrics, software, and cloud services. Vertical control lets the company tune several layers for its workloads, even when individual interfaces use established standards.

Nvidia offers a different reference point through NVLink, InfiniBand, Spectrum-X Ethernet, and its rack-scale systems. Its portfolio links accelerators, switches, interface hardware, and software under a coordinated architecture.

Tighter integration can improve predictable performance and reduce deployment uncertainty. It can also limit component substitution and concentrate roadmap decisions with one supplier.

Open systems promise flexibility, but interoperability does not guarantee efficient distributed training. Devices from different vendors can support the same nominal link rate while behaving differently under congestion or failure.

Operators therefore face a difficult purchasing decision. A more integrated architecture can reduce engineering effort, while an open design can improve sourcing flexibility and long-term negotiating leverage.

The debate extends into optical interfaces. Pluggable modules allow replacement and supplier choice, while co-packaged optics place optical engines closer to switching silicon.

Co-packaged optics can reduce the electrical distance between a switch chip and its optical connection. Shorter electrical paths can improve signal integrity and energy efficiency at high data rates.

Serviceability becomes harder when optics and switching hardware are more tightly integrated. Replacing a small failed module is simpler than servicing an optical engine attached to an expensive switch package.

Linear pluggable optics remove a digital signal processor from the module. This approach can reduce power and latency, but it places more signal-processing responsibility on host equipment.

Data Center Dynamics identified optics reliability and tail latency as emerging constraints in its network complexity analysis. It also noted that faster refresh cycles increase pressure on design and installation teams.

The physical layer remains easy to overlook. More fiber creates larger cable bundles, denser panels, tighter loss budgets, and more demanding testing procedures.

A logical network diagram may show two sites connected by redundant paths. In practice, both paths can share ducts, bridges, power systems, or construction risks unless operators verify physical diversity.

Standards can lower costs by creating larger component markets. They can also mature more slowly than proprietary systems developed around one vendor’s schedule.

This is why the primary contest remains cost versus performance, not Ethernet versus one named competitor. Supplier strategy supports that contest without replacing it.

An open network that leaves processors idle is expensive. A proprietary network that locks buyers into unfavorable upgrades can also become expensive.

The winning approach will deliver predictable job performance while preserving manageable power, maintenance, and sourcing requirements. No single interface specification proves that outcome by itself.

The Optical Roadmap Still Has a Reliability Test

The largest uncertainty is whether emerging optical systems can deliver advertised capacity under continuous AI workloads without creating new operational costs.

Optical links already carry traffic within and between large data centers. Scale-across changes the traffic pattern, failure sensitivity, and expected utilization of those links.

AI back-end traffic can be sustained and synchronized. It may fill links for long periods, leaving less room for maintenance events or unpredictable performance.

Higher signaling rates also shrink engineering margins. Components must preserve signal quality while operating in dense, hot, power-constrained environments.

The industry is moving from 400G toward 800G coherent pluggables, while vendors prepare 1.6T products. Each transition affects switches, modules, test equipment, fiber designs, and operational practices.

IEEE Spectrum reported that optical suppliers are also developing dense wavelength division multiplexing for shorter AI connections. DWDM sends several optical wavelengths through one fiber.

One 2026 design spreads traffic across eight lower-rate channels and targets aggregate speeds reaching 1.6 terabits per second. Its developers plan broader production before expected deployments in 2028.

The optical scaling roadmap is promising, but supplier targets are not production evidence. Manufacturing yield, thermal behavior, error rates, and long-duration reliability still require field validation.

The same caution applies to vendor performance claims. Lower module power does not automatically reduce total system power if the architecture requires more links or spare capacity.

Greater bandwidth does not guarantee lower job completion time. Congestion control, workload placement, and software recovery can dominate results during failures.

Distance creates another fixed constraint. Light travels quickly, but propagation delay cannot be removed through software or better modulation.

Additional switches, optical conversions, queues, and error correction add more delay. Operators can reduce these penalties, yet they cannot make a distant facility behave exactly like an adjacent rack.

This limits the workloads suitable for scale-across. Some training jobs can tolerate wider distribution, especially with software designed around locality and checkpointing.

Other tasks require extremely frequent synchronization. Their economics can deteriorate when network distance repeatedly stalls processors.

Reliability also has a different meaning at cluster scale. A component with a low individual failure rate can still generate frequent incidents when deployed across an enormous link population.

Repair time then matters as much as failure frequency. Technicians may need to locate damaged fiber, replace modules, clean connectors, and retest paths across several facilities.

Operational skills can become a bottleneck. Dense optical plants demand specialized installation, documentation, measurement, and troubleshooting practices.

Supply chains introduce further uncertainty. Accelerated adoption of coherent modules, lasers, connectors, and fiber equipment can create shortages or uneven quality.

The skeptical conclusion is not that scale-across will fail. Google, Meta, and their suppliers have clear reasons to invest in it.

The unresolved question is whether distributed clusters can maintain high utilization after accounting for real failures, maintenance, and software overhead. Public capacity specifications do not answer that question.

Buyers should treat cost-reduction claims as architecture-specific. A result from one hyperscaler’s custom network may not transfer to a smaller operator using different facilities, fiber routes, and software.

What Google News Readers Should Watch Next

Three signals will show whether scale-across networking is becoming an economical production model rather than an expensive workaround.

The first signal is measured utilization from operating multi-building clusters. Meta’s Prometheus deployment offers a particularly relevant test because its planned one-gigawatt capacity spans several structures.

Meta has explained its network topology, but the decisive evidence will come from sustained workload performance. Useful disclosures would include job completion time, accelerator utilization, recovery duration, and network-related interruptions.

Strong utilization across multiple buildings would support the argument that fiber can unlock stranded power without wasting computing resources. Frequent synchronization stalls would weaken it.

The second signal is the production arrival of 1.6T coherent modules. Vendors have announced products and sampling schedules, but operators need volume availability and field reliability.

Watch for qualification by hyperscalers, carrier deployments, interoperability testing, and power measurements at the complete system level. Module specifications alone reveal little about deployment cost.

Successful qualification from several suppliers would improve capacity and sourcing options. Delays, weak yields, or high failure rates would preserve the economics of 800G systems longer.

The third signal is how Google, Meta, and Nvidia divide workloads across network domains. Their software choices will reveal where scale-across adds value and where physical locality remains necessary.

Google says its network spans the fabric inside the AI Hypercomputer, the fabric across it, and its global backbone. Meta is connecting distinct training fabrics through BAG.

Nvidia continues developing integrated systems for scale-up, scale-out, and scale-across environments. The boundaries among these domains will matter as much as their maximum speeds.

A move toward locality-aware scheduling would confirm that distance remains a hard design constraint. Broad placement of synchronized jobs across facilities would indicate that network and software improvements are reducing that penalty.

Google News is useful here as a discovery channel, not as technical evidence. Readers should follow the underlying engineering publications, deployment reports, and independent testing.

For developers, the shift affects available accelerator capacity and the behavior of distributed training platforms. For enterprise buyers, it can influence service reliability, regional availability, and the cost of AI computing.

Network architects should evaluate complete workload outcomes instead of comparing link rates in isolation. Procurement teams should also request failure recovery, power, and interoperability evidence.

The most important question is practical: does another connected building produce proportional computing value after network costs and delays are included?

Track those three signals before accepting claims that scale-across has solved the AI data center cost problem. The network must prove that distributed power becomes productive computing, not stranded capacity joined by expensive fiber.

Give every agent the context to do better work

Connect your agents to the knowledge, decisions, and history already organized in remio.

remio currently supports Windows 10+ (x64) and Macs with Apple silicon.

Your AI Partner at Work
Get more done with remio

Plan. Create. Deliver.
All in one place.

bottom of page