Inder's Desk
Inder's Desk Podcast
Podcast: Wiring the AI Factory
0:00
-28:37

Podcast: Wiring the AI Factory

Welcome to Inder’s Desk. I’m Inder. Today, we’re mapping the network that wires the AI factory together.

A GPU that can’t talk to the other ninety-nine thousand, nine hundred and ninety-nine GPUs is a space heater.

That’s the entire thesis in one sentence.

For the past three years, investors have argued about who makes the best AI chips.

NVIDIA.

AMD.

Google’s TPUs.

Amazon’s Trainium.

But powerful chips sitting alone are like brilliant musicians who can’t hear the rest of the orchestra.

The performance comes from coordination.

And that coordination depends on the network.

As AI clusters get larger, the connections between the chips are becoming more valuable, more complicated, and more essential.

This is the other AI trade.

The companies that build the roads, intersections, bridges, and express lanes carrying data through the AI factory.

They can get paid regardless of which model wins, and sometimes regardless of which accelerator wins.

Before we begin, this discussion is educational. It is not financial advice, and I’m not predicting stock prices. The goal is to understand the technology, the competitive landscape, and why it matters.


There is one framework I want you to remember.

Copper inside the rack.

Optics between racks.

Coherent optics between buildings.

Three boundaries.

Think of an AI data center as a city.

Copper handles the short streets within a neighborhood.

Optics runs the highways connecting neighborhoods.

And coherent optics operates the high-speed rail linking separate cities.

Almost every company in AI networking sits somewhere along those three routes.

And much of the industry’s competitive struggle comes down to where each boundary falls, how quickly it moves, and who collects the toll.

Copper inside the rack.

Optics between racks.

Coherent optics between buildings.

Keep that framework in mind, and the rest becomes much easier to understand.


The first thing most people get wrong is imagining an AI data center as one enormous network.

It is actually three separate networks, each designed for a different job.

The first is called scale-up.

Scale-up is the network inside a single rack, where GPUs communicate directly with other GPUs.

Imagine seventy-two chefs trying to prepare one enormous meal.

It is not enough for every chef to be individually talented.

They need to exchange ingredients, coordinate timing, and avoid getting in one another’s way.

If communication is slow, the whole kitchen slows down.

Scale-up networking is the communication system inside that kitchen.

Its goal is to connect a group of accelerators so closely that software can treat them as one enormous computing engine.

This is the highest-bandwidth and most tightly controlled layer in the entire data center.

NVIDIA’s technology here is called NVLink.

The current generation moves roughly one point eight terabytes of data per second, per GPU.

A flagship seventy-two-GPU rack can move around one hundred and thirty terabytes per second across the full system.

The next generation is expected to increase that substantially.

That one-point-eight-terabyte number matters because the connection inside the rack is roughly ten times faster than the network connecting one rack to another.

It is the difference between handing a document to the person sitting beside you and shipping it to another office across town.

And surprisingly, the connection inside the rack runs primarily on copper.

Actual copper wire.

We’ll come back to why.

The second network is called scale-out.

This is also known as the back-end network.

It connects one rack to another, turning individual systems into clusters containing ten thousand, one hundred thousand, or eventually even more accelerators.

If scale-up turns one rack into a single machine, scale-out turns an entire warehouse into a single computer.

This is the central battleground in AI networking.

It is where NVIDIA’s proprietary technology competes against the open merchant ecosystem.

The third network is the front end.

That handles storage, data ingest, system management, and ordinary enterprise traffic.

Think of it as the loading dock and administrative office.

It matters operationally, but it is not the most differentiated or strategically contested part of the AI network.

So our focus is scale-up inside the rack and scale-out between racks.


The central scale-out battle is InfiniBand versus Ethernet.

In one corner is InfiniBand.

InfiniBand is NVIDIA’s proprietary networking fabric.

It is purpose-built, lossless, extremely low-latency, and supplied by a single vendor.

NVIDIA became the only major commercial supplier after acquiring Mellanox in 2019.

Think of InfiniBand as a private railway.

One company owns the tracks, the trains, the signaling system, and the stations.

Because everything is designed together, the system can run with extraordinary precision.

InfiniBand won the first phase of the AI build-out for a straightforward reason.

For tightly coupled training workloads, it worked reliably and delivered exceptional performance.

In the other corner is Ethernet.

Ethernet is the public highway system.

Many companies can build the vehicles.

Many vendors can supply the roads and traffic-control equipment.

Customers are not locked into one operator.

But ordinary office Ethernet was not originally designed for tens of thousands of GPUs trying to communicate simultaneously.

That would be like putting Formula One cars onto suburban streets and wondering why traffic backs up.

So the industry began rebuilding Ethernet for AI.

The Ultra Ethernet Consortium brings together much of the non-NVIDIA ecosystem, including AMD, Broadcom, Arista, Cisco, Meta, Microsoft, and Oracle.

Its purpose is to make Ethernet behave more like a purpose-built AI fabric while preserving the benefits of an open, multi-vendor standard.

A newer approach called M R C was also introduced by a group including OpenAI, Microsoft, Broadcom, AMD, and, notably, NVIDIA itself.

Its goal is to create much larger and more efficient switch configurations, scale beyond one hundred and thirty thousand computing engines, and reduce the number of switches required by roughly sixty percent.

Imagine replacing a maze of connecting flights with one enormous airport hub.

Fewer stops.

Fewer handoffs.

Less equipment.

Lower cost.

Now here is the data point that captures the direction of the market.

In the first quarter of 2026, data-center Ethernet switch revenue grew sixty-one percent year over year, surpassing ten billion dollars.

And the number-one vendor in data-center Ethernet was NVIDIA.

Its Ethernet revenue reached roughly two point one billion dollars, nearly three times the prior-year level, placing it ahead of Arista and Cisco.

Think about what that means.

NVIDIA has the strongest economic interest in preserving its proprietary InfiniBand ecosystem.

Yet one of its fastest-growing networking businesses is Ethernet.

It is like the owner of the private railway becoming the biggest supplier of trucks for the public highway.

That does not mean the railway is disappearing.

But it tells you NVIDIA has no intention of watching the open market grow without participating.

The current AI back-end market is approximately two-thirds Ethernet and one-third InfiniBand.

But this is not a clean victory.

InfiniBand revenue also rebounded sharply during the same period.

NVIDIA is not abandoning its proprietary fabric.

It is playing both sides of the board.

This is a long competitive grind, not an overnight displacement.

Ethernet is gaining ground for three main reasons.

First, merchant silicon reduces dependence on a single vendor.

Hyperscalers do not want one company controlling the engine, the transmission, the roads, and the tollbooths.

Second, Ethernet network designs can be more efficient.

Some can reach full cluster scale in three switching tiers, while comparable InfiniBand architectures may require four.

Think of each tier as another connection at an airport.

Every additional connection requires more gates, more baggage transfers, more time, and more opportunities for delay.

Removing one tier can reduce the number of optical transceivers by roughly one-third.

And those transceivers cost real money.

Third, every hyperscaler wants negotiating leverage against NVIDIA.

Even customers that depend heavily on NVIDIA GPUs do not necessarily want NVIDIA controlling every surrounding layer.


But inside the rack, NVIDIA remains in a much stronger position.

This is the scale-up layer.

And so far, the open ecosystem has not cracked it.

NVLink is deployed, mature, and approximately twice as fast as the emerging alternatives.

NVIDIA has also made a strategically clever move called NVLink Fusion.

Instead of reserving NVLink only for NVIDIA-designed systems, the company will license portions of the interconnect so that third-party processors and custom chips can connect to NVIDIA’s fabric.

Imagine a country realizing it cannot stop neighboring countries from building their own cars.

So instead, it invites all those cars onto its roads and charges them to use the highway.

Rather than simply losing customers who develop custom silicon, NVIDIA is trying to pull those chips into its own networking ecosystem.

The open alternative is called U A Link.

The second version of the standard was published in April 2026.

It has broad industry support, including AMD, Broadcom, Google, Intel, Meta, Microsoft, Apple, and Amazon.

The architecture is designed to connect as many as one thousand and twenty-four accelerators within a pod.

But the current competitive position remains clear.

U A Link is a blueprint and early construction.

NVLink is a finished bridge already carrying traffic.

And NVIDIA keeps extending that bridge while competitors are still completing theirs.

The merchant ecosystem is making progress in the open scale-out layer while remaining behind in the proprietary scale-up layer.


Now we move between racks.

This is where copper runs out of road and optics takes over.

Inside a rack, the distances are short enough for copper.

But once data needs to travel from one rack to another, copper begins losing signal quality and consuming too much power.

Fiber solves that problem by carrying the data as light.

Think of copper as a delivery van.

It works extremely well for short trips around the neighborhood.

Optical fiber is the freight train.

It costs more to load, but once the distance and traffic increase, it can carry vastly more data much more efficiently.

This is why optics has become one of the highest-growth parts of AI infrastructure.

Every AI accelerator needs to communicate with accelerators in other racks.

That requires optical connections.

And the number of those connections rises as AI clusters get larger.

At the same time, the industry is upgrading from eight-hundred-gigabit connections to one-point-six-terabit connections.

In simple terms, the road is becoming twice as wide while the number of vehicles using it is also increasing.

That gives optical suppliers two growth drivers at once.

More connections.

And more expensive connections.

To understand the companies, it helps to think of an optical module as having three basic jobs.

First, someone has to create the light.

That is the laser.

Second, someone has to translate the electrical signal from the chip into a clean optical signal, and then translate it back at the other end.

Third, someone has to assemble all those pieces into the finished transceiver that plugs into the network switch.

The laser is the hardest part to manufacture.

Many companies can assemble a module.

Far fewer can reliably produce the advanced lasers required for the newest one-point-six-terabit connections.

That makes the laser one of the key bottlenecks in the optical supply chain.

Coherent and Lumentum are two of the important Western suppliers here.

Coherent has an advantage because it manufactures both lasers and finished optical products.

It is like a restaurant that owns the farm supplying its most important ingredient.

When the ingredient is scarce, controlling your own supply matters.

Lumentum offers more concentrated exposure to lasers and optical components.

That can give it greater upside when optical demand accelerates, but it also means the business is more exposed when demand slows.

Then comes the signal-processing chip inside many optical modules.

This chip is called an optical D S P.

Think of it as a translator combined with noise cancellation.

At very high speeds, the signal becomes distorted as it travels.

The D S P cleans it up so the receiving equipment can understand it.

Marvell is the leading merchant supplier of these chips.

Broadcom is the other major player.

Together, they control most of this part of the market.

NVIDIA’s strategic investment in Marvell shows how important this technology has become.

Even NVIDIA, with its enormous internal engineering resources, wants a close relationship with one of the leading optical-signal specialists.

Finally, someone has to assemble the finished module.

Fabrinet is one of the major contract manufacturers doing that work.

But assembly usually has lower margins and fewer technological barriers than producing the laser or the signal-processing chip.

Think of the optical module as a premium smartphone.

The assembler matters.

But the companies supplying the advanced processor and camera sensor often capture more of the economics.

There are also major Chinese transceiver suppliers, including Innolight and Eoptolink, that lead the market by shipment volume.

For an investor, the simple map is this.

Coherent and Lumentum help make the light.

Marvell and Broadcom help translate and clean the signal.

Fabrinet and other manufacturers assemble the finished product.

The most durable bottleneck appears to be the laser.

The D S P is another attractive layer because two companies dominate it.

Assembly benefits from booming demand, but it is generally the most competitive part of the chain.

There is one potential disruption worth understanding.

It is called linear-drive optics, or L P O.

L P O tries to remove the D S P from the optical module.

The attraction is lower power consumption and lower cost.

But removing the D S P is like removing noise-canceling technology from a phone call.

It works when the connection is short and the environment is controlled.

It becomes much harder as speed, distance, and signal noise increase.

That is why L P O may win some shorter-distance connections without replacing D S Ps everywhere.

The larger point is straightforward.

AI clusters need more optical connections.

Those connections are moving to faster and more expensive technology.

And within that supply chain, the greatest value is likely to sit with the companies controlling the hardest components to manufacture, not necessarily the companies assembling the final box.


Now let’s return to copper.

Why does the fastest network inside the data center still rely on copper rather than optics?

The answer is physics.

Every time signaling speed doubles, copper’s practical reach is cut roughly in half.

It is similar to trying to shout a message through a crowded room.

The faster you speak, the harder it becomes for someone far away to understand you.

At current speeds, passive copper can carry a signal for about one meter.

Active copper, which includes a small signal-conditioning chip inside the connector, can extend that range to roughly two and a half or three meters.

That one-to-three-meter range corresponds almost perfectly to the dimensions inside a rack.

Beyond those distances, optics becomes necessary.

The reach limitation is not a minor technical detail.

It defines the addressable market for an entire category of companies.

NVIDIA’s flagship seventy-two-GPU rack illustrates the point.

It contains approximately five thousand copper cables, totaling around two miles of copper, with no optical connections inside the rack.

That was a deliberate engineering decision.

Using one-point-six-terabit optics inside the rack would have increased power consumption, reduced reliability, and added tens of kilowatts to the system.

Copper remains inside the rack for the same reason people take an elevator inside a building but use a train to travel across a city.

The right transportation technology depends on the distance.

Credo is one of the cleanest public-market exposures to active electrical cables.

Its revenue increased more than two hundred percent during its most recent fiscal year.

Astera Labs sells retimer chips and fabric switches.

A retimer is like a relay runner receiving a tired signal halfway through the race, refreshing it, and sending it onward at full speed.

Astera has one of the richest valuations in the networking ecosystem, reflecting extremely high expectations.

Amphenol is the diversified incumbent.

It sells connectors, cables, and backplane systems across many markets, while its data-center business has been growing exceptionally quickly.

The trade-off is customer concentration.

Credo’s largest customer represented roughly two-thirds of annual revenue, while its top ten customers represented around ninety percent.

That pattern appears throughout the highest-growth parts of AI networking.

The more direct the AI exposure, the more fragile the customer base can become.

It is like owning the only food truck outside one enormous factory.

Business is fantastic while the factory is running three shifts.

But if that customer changes its schedule, your revenue changes overnight.

A single delayed hyperscaler program can materially alter a supplier’s growth trajectory.


Now let’s look at two technologies on the horizon.

The first is co-packaged optics, usually shortened to C P O.

Today, optical transceivers generally plug into the front of a network switch.

Co-packaged optics moves the optical engine much closer to the switch chip itself.

Think of today’s design as placing an airport several miles outside the city.

Passengers must first travel across town before they can board the plane.

Co-packaged optics moves the airport next door.

The electrical signal travels a much shorter distance before becoming light.

That can reduce power consumption and increase bandwidth density.

Both NVIDIA and Broadcom have announced co-packaged optical platforms.

This has led some investors to assume that pluggable transceivers will soon be displaced.

The timing matters.

The most defensible current view is that co-packaged optics becomes a meaningful volume technology in 2027 and beyond.

The first deployments may occur earlier, but pluggable modules are expected to remain above ninety percent of switch ports in the near term.

Co-packaged optics is likely to complement the pluggable market before it begins materially displacing it.

It is similar to electric vehicles.

The long-term direction may be clear, but the existing installed base does not disappear the moment the new technology arrives.

That makes C P O an important future risk for transceiver companies, but not necessarily an immediate threat to the current upgrade cycle.

A broader technology land grab is already underway.

Marvell is acquiring Celestial AI for approximately three and a quarter billion dollars, aiming to bring optical connectivity into the scale-up domain.

Private companies including Ayar Labs and Lightmatter are also developing optical input-output technologies.

That may become the next major architectural contest.

The second forward-looking opportunity is data-center interconnect.

Individual data-center sites are beginning to reach physical power limits.

In many locations, operators cannot obtain enough additional electricity to continue expanding one building or one campus.

Imagine an airport that has run out of land for new runways.

Instead of continuing to expand one location, the operator links several nearby airports and manages them as one system.

AI operators are beginning to do something similar.

They connect multiple data centers across a metropolitan area and operate them as one logical training fabric.

The industry often calls this scale-across.

This is the third boundary from our opening framework.

Coherent optics between buildings.

High-end router demand is becoming increasingly influenced by data-center connectivity rather than traditional telecommunications.

That represents a major structural shift for the optical-transport industry.

Ciena is one of the clearest public-market exposures to this trend.

Its revenue recently grew around forty percent year over year, while demand for coherent optical products has reportedly exceeded available supply.


So how should investors organize the competitive map?

The cleanest combination of AI exposure and durability may be in switch silicon and network systems.

Broadcom represents the silicon side.

Arista represents the branded-systems side.

Think of them as selling the traffic lights and highway interchanges.

They benefit whether the vehicles are powered by one kind of engine or another.

They can participate across multiple optical architectures.

They do not need to predict the exact timing of co-packaged optics, the winning transceiver design, or which laser supplier gains share.

They sell the switching layer at the center of the network.

The optical laser layer may offer the largest absolute-dollar growth combined with a legitimate manufacturing bottleneck.

Coherent and Lumentum are two of the clearest examples.

The module itself may become more standardized.

The laser remains harder to produce.

The highest-beta portion of the market is copper cables and retimers.

Credo and Astera Labs have some of the fastest growth rates, but they also carry extreme customer concentration and demanding valuations.

They are speedboats.

They move quickly when the water is calm and demand is strong.

But they are less stable when conditions change.

The value-oriented corner includes companies such as Cisco, where AI networking is an additional source of growth rather than the entire investment thesis.

Cisco is more like a cargo ship.

It may not accelerate as quickly, but one AI program is less likely to determine the entire voyage.

Data-center interconnect suppliers may also offer a later-cycle and more diversified way to participate.

The area requiring the most caution is thin-margin contract manufacturing that has re-rated primarily because of an AI narrative without a comparable improvement in its underlying margin structure.

A company assembling the shovel does not necessarily earn the same economics as the company that owns the scarce metal used to make it.

Highly valued pure plays can also become vulnerable when market prices move beyond even the optimistic assumptions of the analysts covering them.


Now let’s state the bear cases clearly.

The first is that NVIDIA’s full-stack strategy may continue winning.

NVIDIA offers the GPU, NVLink, InfiniBand, Ethernet switches, network-interface cards, and software.

It owns the engine, transmission, road network, and navigation system.

That is the most integrated and highest-attach networking stack in the market.

The merchant ecosystem offers openness and multi-vendor flexibility.

Which approach ultimately controls the AI back end remains unsettled.

The second risk is customer concentration.

Credo, Arista, Fabrinet, Astera Labs, Celestica, and several other suppliers depend heavily on a small number of hyperscalers.

The greatest advantage of a pure play is also its greatest weakness.

When one customer is the wind filling your sails, a change in direction can become a problem very quickly.

The third risk is the white-box model.

Hyperscalers can combine Broadcom switch silicon with open-source networking software and generic hardware.

That puts pressure on branded network-system vendors, especially among the largest and most sophisticated customers.

It is the networking equivalent of buying high-quality ingredients and cooking the meal yourself instead of paying a restaurant.

Arista’s primary defense is the quality and consistency of its software platform.

The fourth and most important risk is capex digestion.

Valuations across this group assume that hyperscaler capital spending continues rising.

It does not require a collapse to create problems.

A pause can be enough.

Think of an escalator that stops moving.

Nobody has fallen through the floor.

The building is still standing.

But everyone who assumed they would keep moving upward now has to walk.

When growth expectations are extreme, merely flattening capital expenditures can produce falling earnings estimates and multiple compression at the same time.


So let’s bring the map back to its simplest form.

Copper inside the rack.

Optics between racks.

Coherent optics between buildings.

Networking is the other AI trade.

It is the infrastructure layer that gets paid regardless of which model wins, and in many cases, regardless of which accelerator wins.

The cleanest and most durable exposure may be in switch silicon and network systems because those companies can benefit across multiple architectural paths.

The highest operating leverage sits in optics, particularly around the laser bottleneck, where the transition from eight hundred gigabits to one point six terabits combines higher prices with rising unit demand.

The highest-beta companies also carry the richest valuations and the greatest customer concentration.

Maximum torque and maximum fragility often arrive together.

From here, there are three developments worth watching.

First, the share shift between Ethernet and InfiniBand in the scale-out back end.

Second, the timing of co-packaged optics, which currently appears to be a 2027-and-beyond volume event, leaving more runway for the pluggable upgrade cycle.

And third, hyperscaler capital-spending guidance.

Because every valuation across this ecosystem depends, directly or indirectly, on that line continuing to rise.

The memorable takeaway is simple.

The AI factory is not defined only by how many chips it contains.

It is defined by how effectively those chips can communicate.

A room full of geniuses who cannot exchange ideas is not a team.

It is just a crowded room.

And a GPU that cannot talk to the rest of the cluster is not an AI accelerator.

It is a space heater.


Thanks for listening to Inder’s Desk.

If you found this useful, subscribe on Substack for more deep dives into the businesses, technologies, and market forces shaping tomorrow’s winners.

Until next time, keep learning, keep questioning, and keep investing with conviction.

This episode is for educational and informational purposes only. It is not financial advice. Do your own research.

Discussion about this episode

User's avatar

Ready for more?