Alibaba's newest AI chip, the Zhenwu V900, is a useful sign of where the AI industry is heading. The company says the processor delivers three times the performance of its previous generation and can be clustered at enormous scale. But Alibaba is far from alone. Google, Amazon, Meta and OpenAI are all developing custom silicon for AI.

That doesn't mean Nvidia-style GPUs have suddenly become obsolete. It means the biggest AI operators have reached a scale where designing hardware around their own workloads can make economic and technical sense.

The short version: a general-purpose AI accelerator has to be good at many jobs. A custom chip can be built around the jobs one company expects to run millions or billions of times. The chip becomes less like an off-the-shelf engine and more like a component shaped around one machine.

The first reason is efficiency

AI models do two broad kinds of heavy computational work: training and inference.

Training is the expensive process of building or updating a model from huge amounts of data. Inference is what happens after training, when the model actually responds to a prompt or performs another task for a user.

Those jobs don't stress hardware in exactly the same way. Google, for example, has split its latest TPU generation into two designs. TPU 8t is aimed at compute-intensive training, while TPU 8i puts more emphasis on memory bandwidth and low-latency inference.

Meta is taking a similar approach with MTIA. It says its custom chips are designed around its own ranking and recommendation workloads as well as generative AI, with upcoming generations primarily targeting inference.

Here's what that actually means: if a company knows precisely what its servers spend their days doing, it can stop paying for some flexibility it doesn't need. The hardware can instead be optimised around the work it does need.

At enormous scale, even modest efficiency gains can reduce the amount of hardware or power needed to do the same work.

Power has become part of the problem

The AI chip race isn't only about making processors faster. Data centres also have to power and cool them. Power is increasingly a ceiling, not just another line on the electricity bill.

OpenAI's first custom inference chip, Jalapeño, illustrates the point. OpenAI says it designed the processor alongside the memory and networking around language-model inference. In tests published in August, the company reported higher performance per watt and lower latency than the comparison systems it used.

Those are OpenAI's own benchmark results, not an independent verdict on every AI accelerator. But they show what the company is trying to optimise: more useful AI work from a fixed amount of power while keeping responses fast.

Google makes the same argument for its TPUs, saying customisation across the system lets it improve energy efficiency.

When a company is planning data-centre capacity in gigawatts, power efficiency becomes a business constraint as well as an engineering metric.

Owning the chip means controlling more of the stack

A custom processor isn't useful by itself. It has to work with the surrounding memory and networking, plus the software and models running above it.

That makes custom silicon part of a bigger strategy often described as owning more of the "stack" — the layers of technology between the physical chip and the AI product a customer actually uses. More ownership gives an operator more knobs to turn when one layer becomes a bottleneck.

Alibaba's September announcement is unusually clear about this. Alongside the Zhenwu V900, Alibaba described a strategy that reaches from its AI models into the hardware and infrastructure used to run them. OpenAI makes the same broader point with Jalapeño, which it says was designed around the needs of its own inference workloads.

The attraction is control. A company that understands both its models and its hardware can change one with the other in mind.

There is a trade-off, though. Custom chips demand deep expertise and specialised partners before they become useful at scale. Most companies will never have enough AI demand to justify that investment.

Cost matters more as AI usage grows

The cost of a chip is only part of the bill. What matters to a large operator is how much useful work an entire system can deliver for the money and power it consumes.

Amazon has spent years building its Trainium processors for this reason. AWS markets Trainium around price-performance for AI training and inference, and Amazon said in 2026 that its broader custom-chip business had reached an annual revenue run rate above $25 billion.

For a cloud provider, there is another incentive: a successful in-house accelerator becomes something it can sell to customers as computing capacity.

That turns custom silicon from a cost-saving project into a product. At hyperscale, fractions on a spreadsheet can become buildings full of hardware.

This is not simply an escape from Nvidia

It is tempting to read every custom AI chip announcement as evidence that the big technology companies want to stop buying Nvidia GPUs. Their own plans suggest a more complicated picture.

Meta explicitly describes its infrastructure strategy as a portfolio approach, combining its own silicon with chips from industry suppliers. OpenAI says it plans to continue widely deploying Nvidia accelerators and hardware from other partners even as it rolls out Jalapeño.

Different processors can also be better suited to different stages of AI. A company may use one system for frontier-model training and another for high-volume inference.

So the shift isn't from one universal chip supplier to another. The hardware market is starting to look more like a toolbox than a single race for one winning processor.

Why this is happening now

Scale changes the calculation because small savings become meaningful when they are repeated across huge numbers of requests.

A smaller AI company can rent computing capacity and let a cloud provider worry about the machinery underneath. A company serving AI at enormous scale has a different problem: tiny improvements in cost or latency can be multiplied across vast numbers of requests, while energy savings accumulate across entire data centres.

At the same time, inference is becoming more important as AI moves from model development into everyday products. Interactive agents make latency especially noticeable because a single task can require many model calls in sequence. One slow step is manageable; repeated slow steps turn into waiting.

That helps explain why so much custom-chip work now focuses on inference rather than simply chasing the biggest possible training processor.

Alibaba's V900 is the latest headline, but the broader story is bigger than one chip. The largest AI companies increasingly see hardware as something they can shape around their models, rather than a fixed layer they simply buy.

They are not all trying to build the same processor. That's the point.