In August 2025, Epoch AI published a finding that ought to have settled the argument about who gets to own artificial intelligence. Using a single top-of-the-line gaming card, it reported, anyone can locally run models matching the absolute frontier of performance from six to twelve months ago. The card it named was Nvidia's RTX 5090. The price it put on that card, in parentheses, was under $2,500.
Every technical word of that sentence is still true. The arithmetic has not changed, the models have improved, and the measured lag has if anything narrowed.
Only the parenthesis stopped being true.
On 14 September 2026, Tom's Hardware reported that the RTX 5090 had vanished from first-party stock at the major American online retailers, with third-party sellers asking between $6,500 and $9,500 and the cheapest actual listing at $6,395. In June the median had been $4,299, and in early September the same publication's price tracker had logged a low of $5,199. On 14 September that tracker stood at $6,389, against a list price of $1,999. Micro Center, alone among large retailers, was still selling at $4,299 to whoever reached the shelf first.
So a research organisation described a route to owning frontier-adjacent intelligence, priced it at under $2,500, and thirteen months later the route is intact while the toll has more than doubled against that figure and more than tripled against the card's list price. Nobody legislated that. No regulator decided that open weights should cost more to execute. The argument about whether AI will be permitted to remain in private hands is being conducted in legislatures, and the question is being settled on a wafer line in Icheon.
The constraint almost entirely missing from the ownership debate is the price of memory. That price has a documented cause, a concentrated set of buyers, and a published timetable for when supply might ease.
Four Months Behind, and Forty Billion Parameters Short
Start with what actually works, because the capability picture is better than most people assume and it is the part of this story nobody is fighting over.
Epoch AI tracks the distance between the strongest open-weight models and the strongest closed ones. In October 2025 it put that gap at about three months, an average of roughly seven points on its capability index, a distance it compared to that between two consecutive commercial releases. By its May 2026 assessment the gap had averaged about four months since January, or eight index points. OpenRouter, which routes real traffic across providers rather than running benchmarks, described the same gap in June 2026 as three to six months and stable for eighteen months.
The benchmark positions agree, and they agree more strongly than the headline numbers suggest. On SWE-bench Verified, the test of whether a model can resolve real software issues, the leading open-weight entry on Epoch's board is Z.ai's GLM-5.2 at 78.7 percent, against 83.5 percent for Anthropic's Claude Opus 4.7 at the top of Epoch's board. Just under five points separate the best thing you can download from the best thing that exists. On Artificial Analysis's composite index in April 2026, open leaders sat at 54 against 60 for the frontier, where a year earlier the best open model had scored 22. That index has since been rebased, so the numbers on the live page no longer match, which is its own small lesson about how fast this field invalidates its own measurements.
Architecture is why the gap closed. Alibaba's Qwen3.5-122B-A10B, released in February 2026, carries 122 billion parameters and activates about 10 billion for any given token, which lets a very large model run at the computational cost of a small one. Its context window is 262,000 tokens natively. It is free to download under Apache 2.0.
Now the part that decides everything downstream, and that most writing on this subject skips.
Running a model has two separate costs. The first is computation, the arithmetic performed on every token, and that cost has been falling for three years through better architectures, sparse activation and quantisation. The second is residency, the plain requirement that the weights sit in memory the processor can reach at speed. Quantisation helps here too, and it helps enormously: at four bits, the format in which the overwhelming majority of local models are actually distributed, a parameter costs roughly half a byte instead of one. Epoch's own threshold assumes exactly this, and puts a 32-gigabyte card at models of about 40 billion parameters.
But quantisation is a discount rather than an exemption. Qwen's 122-billion-parameter model at four bits still needs something in the region of sixty-five to seventy gigabytes before you have loaded a single token of context. The engineers made the model cheap to run and the market made the machine expensive to hold. What arithmetic gave away, procurement took back, and what procurement withholds, no arithmetic supplies.
One more word needs settling before the argument leans on it. Ownership here is narrower than sovereignty over every layer of the stack, since an open-weight model still runs on somebody's silicon, somebody's firmware and somebody's toolchain, and that dependency is the subject of other work in this archive. What it means is the ability to execute a model without asking a provider for permission each time, and without a meter running while you do it.
The part of this problem that looked hardest in 2023 is the part that got solved.
$1,999 to $6,389
The hardware curve moved while almost nobody was describing it, and it moved twice.
Take the tier Epoch used. The RTX 5090 carries 32 gigabytes and a list price of $1,999. In early August 2026, Tom's Hardware put its street price at $4,288, an increase of 114.5 percent. Six weeks later it reported that first-party stock had disappeared from the major American online retailers, with resellers asking up to $9,500, and its running price tracker stood at a best available figure of $6,389 on 14 September. The RTX 5080 and 5070 Ti, both with 16 gigabytes, listed at $1,256 and $949 in that August survey and at $1,579 and $1,149 on the tracker in September.
Now take the tier above, the one that holds a 122-billion-parameter model. Nvidia's RTX PRO 6000 Blackwell carries 96 gigabytes of GDDR7. It opened pre-orders at $7,673 in March 2025 and listed at $8,565 at launch. In June 2026 Nvidia raised the list price to $13,250, and in August 2026 to $16,000. Tom's Hardware calculated the whole move, from pre-order to current list, at almost 90 percent over seventeen months.
Read that again, because it is the strangest fact in this piece. The same SKU nearly doubled in quoted price without gaining a single gigabyte of memory or a single feature. Somebody who placed a pre-order in the spring of 2025 and simply did not cancel it has, by doing nothing at all, outperformed most investments available over the same period.
The effect is an inversion that would have sounded absurd two years ago. A Mac Studio with the M5 Ultra, announced in August and shipping from 22 September, starts at $5,499 and carries 96 gigabytes of unified memory, exactly what the sixteen-thousand-dollar Nvidia card carries, at a third of the price. Apple also announced 256 and 512-gigabyte configurations; the step from 96 to 256 alone adds four thousand dollars, and the 512-gigabyte machine does not ship until late October with no price published. The M5 Max starts at $2,499 and configures upward to 128 gigabytes. Framework's desktop with 128 gigabytes of unified memory lists at $3,449 and is out of stock, though so is every other configuration in that line, and the company advertises a 192-gigabyte version as coming soon.
These products share no manufacturer, no supply chain and no market position. Each is a way of buying a large quantity of fast memory with a processor attached, and all of them ultimately draw on the same concentrated DRAM manufacturing base.
The card did not gain a gigabyte. What changed was the market competing for the memory inside it.
Nine Hundred Thousand Wafers
Here is the sentence that converts this from a coincidence into a mechanism, and it has been publicly available since the first day of October 2025.
On 1 October 2025, OpenAI announced that Samsung Electronics and SK hynix had joined the Stargate project, following discussions at the Presidential Office in Seoul attended by President Lee Jae-myung, Sam Altman, Samsung's Jay Y. Lee and SK's Chey Tae-won. The stated production target, in OpenAI's own language, is 900,000 DRAM wafer starts per month at an accelerated capacity rollout.
Not 900,000 chips. Wafer starts, the unit by which fabrication capacity itself is measured.
Set that against the size of the industry. TechInsights projected global 300-millimetre fab capacity at roughly 10 million wafer starts per month in 2025, of which DRAM accounted for approximately 2.25 million. Tom's Hardware ran the arithmetic on publication day: the stated target is equivalent, on that industry-capacity comparison, to roughly forty percent of monthly DRAM wafer starts. Two cautions belong with that number. It is a production target that the partnerships are scaling toward, not a monthly reservation anyone has already claimed. And wafer starts are not bits, since different nodes, dies and yields produce very different quantities of memory per wafer, so the comparison is against the industry's manufacturing throughput rather than its finished output. Throughput is the more useful measure anyway, because it is the thing that cannot be conjured.
Alongside the memory pacts, OpenAI signed a memorandum of understanding with the Korean Ministry of Science and ICT to evaluate data-centre sites outside the Seoul metropolitan area, and a partnership with SK Telecom to explore capacity. Hold on to those two items. They come back later in a form that changes what they mean.
One distinction has to be made precisely here, because everything downstream depends on getting it right.
The memory in an AI accelerator and the memory on a graphics card are different products. An accelerator uses high-bandwidth memory, stacks of DRAM dies bonded vertically and wired through the silicon, which requires advanced packaging that only a handful of lines in the world can perform. A graphics card uses GDDR7, a flat and far cheaper part. The RTX PRO 6000 carries 96 gigabytes of GDDR7, not HBM. Nobody at OpenAI is bidding against a workstation buyer for the same chip, and anyone who tells you otherwise has skipped a step.
What the two products share is everything upstream of the chip. The same three companies make both. Both come off advanced-node DRAM wafer lines, and both draw on the same capital budgets, the same fab floor, the same engineers and the same expansion plans. And high-bandwidth memory consumes, on Micron's own figure, roughly three times the wafer capacity per gigabyte that ordinary DDR5 does, because the stacking, the through-silicon vias and the yield losses in packaging all cost output. A line converted toward HBM lowers how much memory the industry produces in total, while earning far more per wafer than what it replaced.
So this is an allocation mechanism rather than a competition for the same parts, and allocation is the harder kind to see. Memory fabrication capacity is not elastic on any timescale shorter than years. Samsung was reported in September 2025 to be targeting around 60,000 wafers a month of 1c DRAM, the node that feeds HBM4, by the end of that year, and by November the reported plan had risen to 200,000 a month by the end of 2026, roughly a third of its total DRAM output. Every wafer that moves toward HBM is a rational decision by a manufacturer serving the customer who pays the most. Every wafer that moves is also a wafer no longer making something else. Every one of those decisions is invisible from outside the industry, and the sum of them tightens everything that needs DRAM.
How much of any one product's retail price that tightening explains varies enormously from product to product. What can be said without qualification is what happened to the inputs. Tom's Hardware, citing Taiwanese trade reporting, put DRAM contract prices up roughly 172 percent year on year as of the third quarter of 2025, and NAND flash wafer contract prices rose between twenty and more than sixty percent in November alone.
The bill for making intelligence cheap to rent is itemised, public, and denominated in wafer starts.
San Jose, March
In the middle of March 2026, Chey Tae-won stood on the floor of Nvidia's GTC conference in San Jose and told reporters that the global memory shortage was likely to persist for another four to five years.
Consider who was speaking and where. Chey is chairman of SK Group, which controls SK hynix, which at that point held around fifty-seven percent of the global high-bandwidth memory market and thirty-two percent of DRAM overall, a lead Samsung has since been narrowing. He was standing at the annual conference of the company whose accelerators channel much of that demand. Five months earlier he had been at the Presidential Office in Seoul for the meeting that produced the largest memory supply commitment in the industry's history.
The number he gave described physical supply rather than price. Industry-wide, he said, wafer supply lags demand by more than twenty percent.
He also said, in the same remarks, the thing that explains this entire piece in eight words. Nearly all new capital expenditure is going toward HBM lines, where margins are highest.
That is the incentive layer, stated by the man whose company collects the margin. No manufacturer is choosing to make workstation cards expensive. Each is choosing to point its scarcest asset at its most profitable customer, which is what a manufacturer is for. The consequence for anyone trying to buy ninety-six gigabytes of memory is identical to what a deliberate policy would have produced, and it arrives without one.
Then he set out the timetable, which is the part almost nobody reported. SK hynix would break ground the following month, on 22 April, on a 19-trillion-won HBM packaging and test facility at Cheongju, close to thirteen billion dollars, with its wafer-test line due from late 2027 and its packaging line in 2028. That is advanced packaging rather than new wafer capacity, which matters for a bottleneck measured in wafer starts. Samsung's P5 facility at Pyeongtaek is expected online by 2028. Micron's HBM plant in Hiroshima, announced at around 9.6 billion dollars, produces nothing before 2028.
There is a counterweight in the same file and it belongs here rather than in a footnote. SK hynix has separately committed to doubling its memory wafer capacity within five years. If that lands, the constraint this piece describes has an end date inside the decade.
The first supply answer has a published date and it begins in 2028. The industry's largest supplier says the shortage itself may run to 2030.
Thirty-Five Percent
The consequences reached ordinary buyers before they reached the AI conversation, and they arrived first in a setting where people are obliged to be accurate about their costs.
On 24 February 2026, on HP's first-quarter earnings call, chief financial officer Karen Parkhill told analysts that memory and storage had accounted for roughly fifteen to eighteen percent of the company's PC bill of materials in the prior quarter, and that the company now estimated roughly thirty-five percent for the year ahead. The share of a computer's cost made up of the components that remember things was on course to double inside a single fiscal year. Parkhill added that the company was pulling every lever available to offset what she called unprecedented headwinds.
Bruce Broussard, three weeks into the job as interim chief executive after his predecessor stepped down, listed the countermeasures: new suppliers qualified, strategic inventory built, the time to qualify new material cut in half, long-term agreements secured to cover memory requirements for fiscal 2026, and targeted pricing actions for the rest. The company maintained full-year guidance while saying results would land at the lower end of the range.
HP buys memory at a scale almost no other company matches, did all of that, and still expects to finish the year at the bottom of its own guidance. Dell's chief operating officer, Jeff Clarke, had told analysts in November 2025 that the company had never witnessed costs escalating at the current pace, and Lenovo, speaking the same week, called the squeeze unprecedented while carrying memory inventory around fifty percent above normal levels. Ranjit Atwal of Gartner expects the sub-five-hundred-dollar entry-level PC segment to disappear by 2028.
Now transpose that to the machine this piece is about. If memory and storage are on course to double their share of an ordinary computer, consider a workstation card built around ninety-six gigabytes of the fastest memory made. Such a card is unusually exposed to a memory shock, precisely because memory is not incidental to the product. It is the reason the product exists.
I want to be careful about the limit of what I can establish. I have not found a study that decomposes the increase in consumer and workstation GPU street prices into a memory component, a tariff component, a packaging component and a pricing-strategy component. Those other factors are real and I cannot assign them a share. What is documented rather than inferred is the middle of the chain: the partnerships behind one AI project target production capacity equivalent, on that industry comparison, to roughly forty percent of monthly DRAM wafer starts, the manufacturers pointed new capital at the highest-margin product and said so on the record, prices rose across every category drawing on the same lines, and the firms buying the most memory in the world have been reporting it in earnings calls for a year.
Three of the largest computer makers on earth have now used the same word to analysts, and the word is unprecedented.
Where the Money Went
Stargate is the largest single line item and it is not the only one.
In 2025, four American companies, Microsoft, Amazon, Meta and Google, committed at least $350 billion in artificial intelligence capital expenditure, according to Bloomberg Intelligence estimates collected in a U.S.-China Economic and Security Review Commission staff report. The same four are projected to exceed $400 billion in 2026. China's major cloud providers committed under $40 billion and are expected to stay flat.
That money buys physical objects. Stargate Abilene alone is reported to be targeting more than 450,000 accelerators. Each of those carries a stack of high-bandwidth memory produced by three companies on production lines that cannot be expanded in a quarter.
The asymmetry between American and Chinese capital expenditure produces an irony worth naming. The open models that closed the capability gap are overwhelmingly Chinese, released by laboratories operating at roughly a tenth of the American spending level. Much of the demand shock tightening the memory market originates in the extraordinary AI capital expenditure of American hyperscalers. The country spending the most on computing has helped make its competitors' software free and the hardware to run that software scarce.
So the circle closes. The capital expenditure helping make intelligence cheaper to rent is the same capital expenditure tightening the inputs required to own the means to produce it. The cheapness of the service and the expense of the alternative share a cause, and it is not a policy.
8.9 Million Developers, and Eleven Percent
Two adoption figures sit in tension, and holding both is the only way to describe what is happening.
Ollama, the tool most people use to run models on their own machines, reported 8.9 million monthly active developers in July 2026 and says it is present in eighty-five percent of the Fortune 500, on fourteen employees. Both figures are its own, unaudited, released alongside a $65 million round. Hugging Face hosts 2.96 million public model repositories, and in a single mid-2026 month Qwen's quantised files were downloaded 39.6 million times. By any reasonable measure, running your own model is a mass activity.
Then set that against enterprise spending. Menlo Ventures found in December 2025 that open-source models accounted for eleven percent of enterprise spending on model APIs, down from nineteen percent a year earlier.
That comparison needs a warning label, and the warning matters more than the figure. Menlo's number measures money paid to inference providers. A company self-hosting Qwen on its own hardware contributes nothing to that denominator, which means the metric is structurally blind to exactly the population Ollama's 8.9 million describes. Some of the apparent tension between those two numbers is a measurement artefact and I should say so before a reader finds it. Menlo's own explanation for the decline is Llama's stagnation rather than hardware cost, a competing causal story this piece does not resolve.
What is not an artefact is the download distribution. On Hugging Face, models above 100 billion parameters account for one percent of all-time downloads. Models below one billion account for eighty-three percent. Of 2026 downloads, three percent were of models above 70 billion parameters.
So the shape of this is breadth without depth. Millions run small models on ordinary hardware for ordinary tasks, which is valuable and unthreatened. Almost nobody runs the large ones.
Notice which way the causation runs there, because it is easy to read backwards. A one-percent share for very large models could mean hardly anyone wants them. It could equally mean hardly anyone can hold them. The distribution alone cannot separate a preference from a budget, and the same curve would appear under either explanation. What tilts the reading is the price series, which moved in the same window and in the direction consistent with affordability constraining that distribution.
Ownership did not narrow at the bottom. It narrowed at the size where it would have mattered.
What the Rules Actually Say
This is the point where the argument I expected to write fell apart, and the falsification is worth showing rather than hiding.
The expectation was straightforward and widely held: that AI regulation, whatever its stated purpose, would function as a moat, because compliance is a fixed cost and fixed costs favour scale. A capability rule with no threshold catches the person fine-tuning a model in a spare room as surely as the laboratory running a hundred thousand accelerators, and only one of them has a general counsel.
The European Union's AI Act is the most developed AI regulation in existence, and it goes the other way at every point where that thesis predicts it should not.
Article 2(12) provides that the Regulation does not apply to AI systems released under free and open-source licences, unless they are placed on the market or put into service as high-risk systems, or fall under the prohibited-practice or transparency articles. That exemption covers systems rather than general-purpose models, which are governed separately, and which is why Article 53(2) exists at all. Article 53(2) exempts open-source model providers from the technical documentation and downstream-information obligations where weights, architecture and usage information are public. It does not exempt them from the copyright-policy or training-data-summary duties, and it falls away entirely once a model is designated as carrying systemic risk. Article 51(2) sets the threshold for that designation at 10^25 floating-point operations of training compute.
The small-firm provisions are equally specific. Article 11(1) permits simplified technical documentation. Article 58(2)(d) makes regulatory sandboxes free of charge for small firms. Article 62(1) gives them priority access, and Article 62(2) requires their size and market share to be reflected in conformity-assessment fees. Article 63(1) allows a simplified quality management system, originally for microenterprises alone and widened to all small firms by the 2026 amendments. Article 99(6) caps their fines at whichever of the percentage or the absolute figure is lower.
Two provisions that get quoted as statute are not statute, and the distinction cuts against my own argument rather than for it. The indicative figure of 10^23, for whether a model counts as general-purpose at all, does not appear in the Regulation. Neither does the rule that a developer fine-tuning somebody else's model becomes a provider in their own right only above one third of the original model's training compute. Both sit in the Commission's guidelines on the scope of general-purpose AI obligations, published on 18 July 2025, which bind nobody. They are the operative administrative position and they are not law. A piece arguing that legislators wrote the exemptions in should be honest that on these two the legislators did not, and the Commission did.
Nor did any of this tighten. The Digital Omnibus on AI, Regulation (EU) 2026/1744, was adopted on 8 July 2026, published on the 24th and entered into force on the 27th. It postponed high-risk obligations to December 2027 and August 2028, extended small-firm reliefs to small mid-caps, and created a Union-level sandbox with priority access. It did not alter the 10^25 threshold and it did not alter the open-source exemptions. In March 2026 the Parliament's internal market and civil liberties committees adopted a joint position in favour of the postponement, 101 votes to 9 with 8 abstentions.
There is a further observation here and it is not flattering to the legislation. The compute threshold at which the Act's systemic-risk obligations bite is a training threshold. It says nothing about what it costs to run a model once trained, which is the constraint this piece is about. The Act governs the making of intelligence and says nothing about the holding of it. What it governs, almost nobody is fighting over. What people are fighting over, it does not govern.
The most developed AI regulation in the world distinguishes by scale in six separate places, and says nothing whatsoever about the constraint that is actually binding.
The Argument That Did Not Survive
The moat thesis can be tested one more way. If frontier laboratories were building a regulatory moat, the behavioural record would show them opposing exemptions for small developers and for open weights. Lobbying is a documented activity. Disclosure filings, consultation responses and formal submissions are public.
On the size question the record is empty, and emphatically so. Meta's response to the American open-weights consultation argued against restrictions on the grounds that limiting open-sourcing would undermine national interests. OpenAI's own comment to the same consultation warned that catastrophic-risk assessments may cost a substantial fraction of the budget of small training runs, with a chilling effect on innovation. Anthropic's published position on open-weights models, in July 2026, endorses safety testing while exempting less capable models, such as those from startups and academia, entirely. I found no filing, from any major laboratory, opposing a size-based exemption.
On the open-source question the record is not empty, and I had it wrong. In that same July 2026 document, published seven weeks before this piece, Anthropic argues that all sufficiently capable models, open and closed, should go through mandatory safety testing, on the grounds that open weights are harder to guard and cannot be withdrawn once released. That is explicit opposition to an open-source carve-out. It is a carve-out from testing rather than from the size thresholds, but it exists, it is recent, and a claim that the record showed nothing would have been wrong.
The surviving distinction is sharper than the absence I went looking for. Laboratories do contest where the open-closed line falls. They do not contest where the size line falls. And the moat thesis predicted the second.
What survives on the cost side is narrower and better supported. Compliance costs are fixed, and fixed costs do fall harder on small firms. A cross-country study of the GDPR by Carl Benedikt Frey and Giorgio Presidente, published in Economic Inquiry in 2024, found an average fall of around eight percent in profits among exposed firms, roughly double that among small IT firms, and no statistically significant effect on Facebook, Google or Apple. That is a real asymmetry. An asymmetry in who bears a cost is something short of a moat built on purpose, and in the AI case the legislators wrote the exemptions in anyway.
It would be easier to write that someone arranged all of this. I am not going to, and not because the evidence is merely thin. The moat thesis is attractive because it explains an outcome by assigning it a motive, and motive is satisfying. The mechanism actually doing the work explains the same outcome by assigning it an order book. Nobody in the memory story wants private AI ownership to be difficult. Samsung has no view on it. OpenAI would presumably be indifferent if the RTX PRO 6000 cost four hundred dollars. The effect is produced by ordinary purchasing at extraordinary scale, which is how most structural outcomes get produced, and which is exactly the kind of cause a search for intent walks straight past.
A mechanism that predicts a lobbying position and finds the opposite one has told you something, and what it told me was that I was looking in the wrong place.
Seoul, 28 August
One government treated the problem as what it is, and the shape of its answer is the clearest evidence that the constraint is physical rather than legal.
On 28 August 2026, South Korea's Ministry of Science and ICT announced the outcome of a national procurement. Three consortia, led by SK Telecom, KT and Kakao and selected from six applicants, will provide generative AI to the public free and without usage limits. Beta service was scheduled for September, with full launch in December. At a briefing on 4 September the three winners set their targets: five million monthly active users at launch for Kakao and KT, ten million by the end of 2027 for SK Telecom.
Look at what the state actually supplied. Not a right, not a rule, not a licence. In 2026 the government is providing 512 Nvidia B200 accelerators in kind, with part of the running cost carried by the national budget from 2027. It bought hardware.
There is a condition attached, and it is industrial rather than protective. At least half of each service must run on South Korean sovereign foundation models, with at least thirty percent from other domestic providers. Foreign models are permitted only where minimally necessary, and are not subsidised.
Now bring back the items from the Stargate announcement, because the overlap is nearly total.
The Ministry of Science and ICT, which ran this procurement, is the same ministry that signed a memorandum of understanding with OpenAI in October 2025 to evaluate Korean data-centre sites. SK Telecom, which leads one of the three winning consortia, is the same SK Telecom that signed the Stargate capacity partnership, and belongs to the same SK Group whose chairman committed to the memory supply and later gave the shortage four to five years. Samsung and SK hynix, the two firms scaling toward 900,000 wafer starts a month, are Korean firms.
So one country occupies every position in this story at once. It manufactures much of the memory at the centre of the scarcity described here. It hosts and partners the project whose production target accounts for a plurality of that throughput. And it offers one of the clearest examples of a state answering an AI-access problem by buying accelerators and distributing the capacity.
Hypocrisy and conspiracy both overexplain this. It is what a country does when it can see the whole supply chain from the inside, because it owns most of it. Seoul could read the order book without needing a theory about AI freedom, and the order book said access would be settled by whoever held accelerators. So it went and got 512 of them.
Set that beside the argument happening elsewhere. In much of the American public debate the visible question has been how capability should be governed. In this Korean procurement the operative question was how much compute had to be physically secured, and the answer was 512 accelerators.
The country that treated this most explicitly as a procurement problem is also the country sitting closest to the memory supply chain that defines it.
The Strongest Objection
The strongest counterargument to this reading comes from Epoch AI rather than from industry, from the same organisation whose capability measurements open this piece, and it is the objection I found hardest to answer.
In August 2025 Epoch published exactly the finding that undercuts a hardware-scarcity thesis. A single RTX 5090, one 32-gigabyte card, runs models of around 40 billion parameters at four bits, and those models match the absolute frontier of six to twelve months earlier. The measured lags were 7.4 months on GPQA-Diamond, 7.3 on MMLU-Pro, 6.3 on the Artificial Analysis index and 12.4 on LM Arena Elo. On this reading the hundred-gigabyte tier is a hobbyist concern, the band that matters is already available on one consumer card, and a piece that measures freedom by the price of a workstation card has picked the wrong instrument.
The industry version of the objection reinforces it. Hardware shortages are cyclical and have always ended. Three new plants are under construction and SK hynix has committed to doubling its wafer capacity within five years. Eighty-three percent of downloads are models under one billion parameters, which looks like revealed preference rather than exclusion. Distillation keeps lifting small-model quality, and a distilled 27-billion model in 2026 does work that needed a frontier system in 2024.
Those objections are serious and the download distribution genuinely supports them. It is why this piece does not claim ownership is ending.
One part of the objection does not survive contact with the release record, and it is the part I expected to be strongest. The open frontier is diverging in both directions at once rather than converging on sizes that fit a consumer card. In mid-August 2026 Alibaba released Qwen3.8 twice: as a dense 27-billion model under Apache 2.0, and two days earlier as a 2.4-trillion-parameter mixture with roughly 95 billion active. Moonshot's Kimi K3, at 2.8 trillion, had gone out a month before that. The small model runs on the card Epoch named. The large ones need well over a terabyte of memory even at four bits, which is not a hobbyist tier, a workstation tier or a Mac Studio tier.
And the licences moved with the sizes, which is the part that belongs in a piece about ownership. The 27-billion model is Apache 2.0, with no conditions worth reading. The 2.4-trillion model is not. It ships under a bespoke licence requiring a separate commercial agreement once a licensee's model-serving revenue passes fifty million dollars in twelve months, and Kimi K3 carries its own custom terms. Hugging Face flagged the pattern in its August survey: at the very top of the open range, weights are now arriving with revenue gates and attribution requirements attached. So the models moving furthest beyond ordinary private hardware are also the models that stopped being unconditionally free, and both movements began in the same year.
Here is what survives. Epoch put a price on its own conclusion, and the price was under $2,500. That card is now $6,389 where it can be found at all, and gone from first-party stock at the large American retailers. Not one word of Epoch's technical argument has become false. The whole of its accessibility argument has. A conclusion that depends on a consumer card being a consumer purchase does not survive the consumer card tripling, and the tripling is the subject of this piece rather than an exception to it.
The claim is therefore narrower than a threshold and broader than a tier. The entire curve moved, at every capacity from sixteen gigabytes to ninety-six, in the same window and in the same direction. Whether the band that matters sits at 32 gigabytes or at 96 is a question I cannot settle and do not need to, because both ends of it repriced together.
But the objection does change what this piece is finally about, and it is worth naming the change rather than burying it. This began as a story about two curves running in opposite directions: the capability gap between open and closed models, which is narrowing, and the price of memory, which is rising. There is a third curve underneath both, and on a long enough view it is the one that decides the outcome. Call it capability per gigabyte: how much useful intelligence fits in a fixed quantity of memory, which improves through distillation, sparsity, better quantisation and architectures nobody has published yet.
If engineers can compress frontier-adjacent performance into thirty gigabytes faster than the memory industry can make a hundred affordable again, ownership wins without the shortage ever ending, and this piece will have measured a real constraint that turned out not to bind. If capability per gigabyte improves slowly while the memory constraint runs to 2030, the band that matters drifts further out of reach every quarter. The race is therefore not only between memory demand and new fab capacity. It is also between the price of a gigabyte and the amount of intelligence engineers can fit inside one, and that second race is a better description of the question than anything in the regulatory debate.
What I can say is that the instrument for reading this is a catalogue rather than a statute.
What Runs Out, and When
Consider what the past twelve months contained. The gap between what you can download and what you can rent narrowed to about four months, and to under five points on the benchmark that measures real work. The consumer card that Epoch priced at under $2,500 reached $6,389 and left the shelves. The workstation card that holds a 122-billion-parameter model doubled to $16,000. The partnerships behind a single AI project set a production target equivalent, on that comparison, to roughly forty percent of monthly DRAM wafer starts. Memory went from a sixth to a third of the cost of an HP laptop in ninety days. Four American companies committed more than $350 billion to buying the components those cards are made from, and the chairman of the largest memory supplier said in public that nearly all new capital is going where margins are highest. Nine million people installed software to run models on their own machines, and almost all of them ran small ones. A European regulation exempted open-source models, set a training-compute threshold, and postponed most of its own obligations. A Korean ministry bought 512 accelerators. Senator Bernie Sanders and Representative Greg Casar announced the Ban Artificial Superintelligence Act on 3 September, carrying penalties of up to twenty years in prison for persons who attempt to circumvent its prohibitions, and as of 17 September it had still not been formally introduced in either chamber.
Every one of those facts was published. None was hidden. What is missing is the sentence that puts them in the same paragraph, and it is missing because the debate and the mechanism are being conducted in different vocabularies. The people arguing about AI freedom are arguing about permission. The thing determining AI freedom is availability. Permission is decided in public and availability is decided in procurement, and only one of those has a press office.
It is also the rare structural question with a real timetable, and the timetable runs on two clocks that do not keep time together. New fabrication capacity begins arriving in 2028, with Samsung's P5 and Micron's Hiroshima plant, while SK hynix's Cheongju build is packaging rather than wafers and adds no starts. The shortage actually clearing is a different event, and the chairman of the group that makes more than half the world's high-bandwidth memory put that at four to five years from March 2026. Anyone who wants to know whether private ownership of frontier-adjacent AI survives should watch two construction sites, one capacity commitment, and a single line in a quarterly earnings report.
The question people think they are asking is whether they will be allowed to own an artificial intelligence. The question that is actually being answered is whether they will be able to afford one, and it is being answered by people who are not thinking about them at all.
A right you cannot exercise is not taken away. It is outbid.
Evidence Map
Facts, interpretations, forecasts, and disconfirming signals.
Core claim. The capability gap between downloadable and rentable AI models has narrowed to roughly four months, while the hardware required to run open models repriced upward by a factor of two to three at every capacity tier during the same period. The binding constraint on private ownership of frontier-adjacent AI is the cost and availability of advanced DRAM rather than regulation, and that constraint traces to manufacturing capacity and capital reallocated toward the AI buildout by the same firms whose services are the alternative to ownership. The mechanism is allocation of shared wafer capacity and margin, not direct competition for the same component.
Evidence level. Facts: high. Capability-gap measurements are Epoch AI's, corroborated by OpenRouter; benchmark positions are from Epoch's published dataset and from Artificial Analysis in April 2026, on an index since rebased. GPU prices are from Tom's Hardware, 3 August and 14 September 2026, and from manufacturer listings; Mac Studio and Framework prices are from the manufacturers' own configurators, with Apple's 512-gigabyte configuration unpriced and unshipped as of publication. The 900,000 wafer-starts-per-month target is from OpenAI's announcement of 1 October 2025; the roughly forty percent figure is Tom's Hardware's arithmetic against TechInsights wafer-start capacity projections and refers to manufacturing capacity, not finished bit output. Chey Tae-won's remarks were made at Nvidia GTC in March 2026 and reported by Bloomberg and the Korea Times. HP figures are from the Q1 2026 earnings call of 24 February 2026 and the company's own transcript, in which the thirty-five percent is a full-year estimate rather than a realised quarter. Hyperscaler capital expenditure is Bloomberg Intelligence via a U.S.-China Economic and Security Review Commission staff report of March 2026, and covers those four companies rather than American spending generally. Adoption figures are Ollama's own, unaudited, and Hugging Face's August 2026 review; the enterprise share figure is Menlo Ventures, December 2025, and measures API spending only. EU AI Act articles are quoted from the consolidated Regulation, with the 10^23 figure and the one-third fine-tuning rule identified as non-binding Commission guidance of 18 July 2025. South Korean details are from the Ministry of Science and ICT announcement of 28 August 2026 and Korean trade press. Interpretation: medium. The capacity-reallocation mechanism is documented at industry level; the link from it to specific street prices is inference from scale, timing and category-wide movement. HBM and GDDR7 are distinct products sharing upstream capacity and capital, so the chain runs through allocation rather than substitution. No study decomposing GPU price increases into memory, tariff, packaging and pricing-strategy components was found. Forecast: speculative. Whether and when prices ease depends on fab completions that have announced dates and no track record.
Weakest link. The chain runs from AI capital expenditure to memory demand to capacity reallocation to device pricing to local-AI economics. The weakest established link is not the first one. AI demand driving memory scarcity is documented at industry level, by the buyer's own announcement and the supplier's own remarks. The weak link is the last step before the headline: how much of any specific GPU price increase memory scarcity actually accounts for, against tariffs, packaging constraints, distribution and Nvidia's own pricing power. If somebody decomposes the RTX PRO 6000 increase and memory turns out to be a minor share, the broader argument survives and the hook does not.
What would confirm this. Memory and GPU prices staying elevated or rising while hyperscaler capital expenditure continues to grow. The share of Hugging Face downloads above 70 billion parameters remaining at or below three percent. Further device makers reporting memory as a rising share of bill of materials. Open-weight releases clustering at sizes that fit whatever capacity is cheapest rather than whatever capability is best.
What would disprove this. Street prices returning toward list while frontier capital expenditure keeps rising, which would sever the link between the two. A collapse in advanced DRAM prices ahead of the 2028 fab completions. SK hynix's commitment to double wafer capacity within five years landing early, or the reported cooling of the memory price surge as consumers reach an affordability limit turning into a sustained reversal. Small and distilled models continuing to improve fast enough that the capacity question stops mattering, which would make the instrument used here the wrong one. Or evidence that tariffs, packaging constraints and pricing strategy, rather than memory allocation, account for most of the increase.
Watchlist. Street prices for the RTX 5090 and RTX PRO 6000 through Q4 2026, against a dated baseline: on 14 September 2026 Tom's Hardware reported first-party stock gone from major US online retailers, third-party listings at $6,500 to $9,500, a cheapest actual listing of $6,395, a June 2026 median of $4,299 and an early-September tracker low of $5,199. Its continuously updated GPU price tracker stood at $6,389 as of its 14 September refresh; that page carries no fixed date, so it should be re-read rather than cited. A return toward $4,299 with hyperscaler capital expenditure still rising is the cleanest disconfirming signal available. Apple's price for the 512-gigabyte Mac Studio when it ships in late October. Samsung P5 and Micron Hiroshima completions in 2028, and SK hynix's capacity-doubling commitment. Hyperscaler capital expenditure guidance for 2027. Memory as a share of bill of materials in HP, Dell and Lenovo earnings calls. The Epoch open-closed capability gap, and any update to its consumer-GPU analysis, which is the cleanest published proxy for capability per gigabyte. Watch whether the parameter count that fits on one consumer card buys a shorter lag each year. Hugging Face's next model-size download distribution. The Korean beta and whether the domestic-model requirement holds. Whether the Stargate Korea data-centre plans clear their reported site and power constraints.
Related from The Manifest Archive
The ownership illusion: licence versus ownership in the digital age
Everyone is watching the AI boom. Nobody is watching the transformer that can't be built in time.
Palantir locked federal data. Now federal agencies cannot leave.
Jerry van der Laan writes The Manifest Archive, a long-form investigation into how institutions, language and infrastructure decide what counts as reality. He works from primary documents, market data and the public record. He traces the structures beneath them.