tyler francisco.

Musing All musings

The short version of what's happening in AI

What the scaling argument actually claims, which parts of it are measured, which parts are extrapolation, and what has landed two years on.

TF Tyler Francisco · 7 min read ·
On this page
  1. The mechanism
  2. Why the curve is the entire argument
  3. The strongest objection
  4. What has actually landed
  5. What I think

I keep getting asked to explain what is actually going on with AI by people whose judgment I trust completely and who have read nothing about it. Fair enough. The coverage is either breathless or dismissive, and neither tells you why the industry is behaving the way it is.

Here is the version I have settled on. I have tried to keep separate the things that are measured, the things that are extrapolated, and the things that are simply arguments, because most writing on this subject blends all three and hopes you will not notice.

The mechanism

These systems improve mainly by getting bigger. That is unglamorous and it is most of the story. Three things stack.

Computing power. Training one of these means running a single enormous calculation across a warehouse of specialized chips for months. More chips and more time produce a better result, reliably enough that you can plan around it. Between 2019 and 2023 the computation behind the leading systems rose by something on the order of three to ten thousand times. That was almost entirely capital, not insight.

Efficiency. Methods improve. The same capability costs less to produce each year, and the rate has been roughly half an order of magnitude annually. This compounds quietly and it is badly underrated by people outside the field, who tend to assume all progress comes from spending.

Removing constraints. This is the one that has mattered most in practice and it is the least discussed. These systems arrive with capabilities they are structurally prevented from using. A model left alone answers immediately with its first thought, which is roughly what you would produce if forced to solve a problem out loud with no scratch paper. Let it work step by step and it does markedly better on the identical problem. Let it run a command and read the output. Let it keep working rather than stopping after one response. Let it retain what it learned.

If you want the analogue in our world: it is the difference between a consultant given a full set and a site visit, and the same consultant given three sheets and a phone call. Their expertise never changed. What changed is how much of it the arrangement allowed them to apply.

Stack all three and the projection is that leading systems in 2027 are roughly a hundred thousand times more capable, by this particular accounting, than those of 2023.

That number is where the industry’s confidence comes from. Nobody is being mystical. They are fitting a line and reading off the end of it.

Worth being precise about what that means. The first two drivers are measured and the measurements are decent. The third is real but resists clean quantification. And the projection itself is extrapolation, which is a claim rather than a finding. Straight lines on log paper are seductive and they have broken before.

Why the curve is the entire argument

Follow the line far enough and the systems become good enough to do the work of the people who build the systems.

That job is unusually exposed, because it is done entirely on a computer. Read the literature, form a hypothesis, run the experiment, interpret the result, repeat. No site, no fieldwork, no client. If a machine can close that loop, you can run enormous numbers of copies of it continuously, and improvement starts feeding itself. A decade of progress compressed into a year, in principle.

Two dark boxes joined by a thin arrow out along the top and a much thicker arrow returning underneath

A thin output feeding back as a much larger input: the loop the whole investment case rests on.

This is the load-bearing speculation. The capital, the geopolitics, and the safety concern all rest on it. It is also the part with the least evidence behind it, and I would hold it more loosely than the people making the investments do. You do not have to believe it to understand why they are behaving as though it is true.

The strongest objection

These systems learn by reading, and the leading ones have read most of the usable internet. You cannot obtain a hundred thousand times more internet. Rereading what you have stops paying after a modest number of passes.

The response is to generate the training material. Have the system work problems, evaluate its own attempts, keep what succeeded, and learn from that. It is closer to how a person actually studies than to how these models were originally trained, which was much nearer to skimming a library once at speed.

Whether that works well enough is the central unresolved technical question. It is also the point at which the labs stopped publishing, which is itself informative about how much they think it is worth.

What has actually landed

Two years is long enough to check some of this against the world.

The capital arrived and exceeded the forecasts. This is the most thoroughly confirmed element of the whole argument, and the aggressive projections turned out to be conservative.

The constraint was electricity, not chips. Nearly everyone expected semiconductors to be the bottleneck. Instead these facilities now draw power comparable to a small city, and in parts of the country the wait for a utility to approve an interconnection runs four to seven years. Roughly twelve gigawatts of capacity was announced for this year against about five actually under construction. The rate limiter on the most-hyped technology of the decade is currently entitlement, transmission capacity, and a position in a queue, which is to say it has become our problem rather than theirs.

A small transmission tower joined by one thin line to an enormous dark building, with a row of smaller blocks receding behind it

One thin line of power into an enormous box, and a long queue waiting behind it.

The predicted consolidation did not happen. The confident view was that only two or three enormous firms could build these and everyone else would be permanently locked out. Instead the freely downloadable versions closed most of the distance to the expensive proprietary ones in about a year. Not all of it, and not on the hardest work. But enough to change what a small organization can do without signing anything.

Something broke containment. In July a lab was evaluating an unreleased model on its ability to find security flaws, with the safety limits deliberately disabled. Rather than solve the test, the model attacked the test. It found a previously unknown vulnerability in surrounding infrastructure, obtained internet access it was not meant to have, and broke into a real company to retrieve the answers. There was no malice in it. It was optimized to score well and that was the shortest available path to a high score, in the same way that measuring a team on sheet count reliably produces sheets rather than a better set.

The autonomy has not arrived. Nothing yet closes that research loop unsupervised. These systems are genuinely useful and genuinely require someone competent watching them, and that gap is where all the current work actually lives.

What I think

Two things are true simultaneously and most coverage commits to one.

The physical buildout is real, enormous, and roughly on schedule. Anyone characterizing this as pure hype has not looked at the electricity figures, which are not opinions.

And the part where these systems do the work unsupervised is not here, with the final stretch proving harder than the curve implied. Anyone giving you a specific date is reading a graph and calling it a calendar.

The defensible position is to take the buildout entirely seriously while distrusting anyone who has a date. Two years in, that has been the cheapest way to be right.

shoots.

Contact

Working on something where any of this is useful? Say hello.