The Mainframes of 2030? (Continued)

Audio reading

Audio reading by Polly on Amazon Web Services

Artificial Intelligence · Data Centers · Energy Consumption · Computing Infrastructure · Cloud Computing · tech

AI may be heading in much the same direction.

It has two pretty different jobs. One is creating intelligence. The other is using it.

Creating a frontier model is still heavy industry. The processors used to train one are no longer just faster versions of the chips in our laptops. NVIDIA, Google and others build processors specifically for AI, then connect thousands of them with huge amounts of high-speed memory and networking. At that scale, the data center starts behaving less like a building full of computers and more like one enormous machine.⁴ That isn’t moving into my spare bedroom anytime soon.

At the same time, though, the machines themselves are getting much better.

Stanford reports that the performance of leading machine-learning hardware has been increasing about 43 percent a year, while price-performance has improved about 30 percent annually and energy efficiency about 40 percent. Epoch AI estimates that leading AI hardware has roughly doubled its energy efficiency every two years.⁵ This is a cousin of Moore’s Law, although the gains now come from a lot more than squeezing additional transistors onto a chip. The processors themselves are being redesigned around AI.

The models are getting more efficient too. Put those two things together and the cost of actually using AI comes down very quickly. Stanford found that processing a million tokens at roughly GPT-3.5 performance fell from about $20 in November 2022 to seven cents in October 2024, a decline of more than 280-fold in less than two years.⁵

I wouldn’t bet on that exact rate continuing. Technology curves eventually slow down. But I wouldn’t bet against the general direction either.

And this is where I think we may be asking the wrong question. We keep asking whether tomorrow’s laptop will be powerful enough to run the giant frontier models being trained today. Why should it have to?

Most of what I ask AI to do is pretty mundane stuff: turn a recipe photo into text and change it for smaller portions; find the best product to restore a mahogany deck—and what aisle it’s in at Walmart; find an obscure, family-owned B&B that’s too small for the travel services. I’m not asking it to discover a new cancer drug or solve some impossible mathematics problem.

For most of us, a lot of AI is going to be stuff like that.

Apple is already designing its AI around that idea. Its current on-device model contains 20 billion parameters, but its sparse architecture activates only one to four billion for a particular request. Harder jobs can move to larger cloud models. Microsoft, with more than 1.6 billion active Windows devices, has described the PC as a place for “unmetered intelligence at the edge.”⁶,⁷

That sounds about right to me. Let the laptop do what it can. A business may have a somewhat bigger private system of its own. If the job really needs more horsepower, send it up to the cloud. There is no particular reason every request has to make the whole trip.

That changes the economics of these giant data centers too. Why spend billions developing a frontier model if much of the work eventually gets done on somebody else’s machine?

Because the giant model may not be what most people end up using. It may be the master version from which a lot of smaller ones are made.

Meta demonstrated the idea with Llama 3.1. Along with a 405-billion-parameter model, it released versions with 70 billion and eight billion parameters and specifically described the large model as useful for generating synthetic training data and distilling capability into smaller models.⁸ So the cloud doesn’t just become a warehouse full of training data. It becomes something closer to a model foundry.

← BackThe Mainframes of 2030? · Page 2Continue →