Real Technology, Phantom Profits: Who Is Likely to Actually Get Paid for the AI Buildout?: MONDAY MAMLMS
This week brought me news increasing my confidence with respect to one thing it is not: it is unlikely to be a hyperprofits engine for research labs with companies attached making money serving up calculations made in datacenters by simulated neural networks:
What is now called “AI”—what I insist on calling MAMLMs, Modern Advanced Machine-Learning Models; but perhaps CIP, Complex Information-Processing, would be better—these days is so many things:
First of all, it is wishful mnemonics: the opposite of the rectification of names, in that the very idea is to call it something that it is not, and all because Marvin Minsky was primarily a showman seeking a bigger budget.
Most important, it is natural-language interfaces to information systems, especially, so far at least, databases of example code from successful software projects.
But it is also:
A potential disruptor and destroyer of platform monopolies’ existing megaprofits flows.
Hence a spur to platform monopolists to spend fortunes to try to protect those existing platform-monopoly profits from this disruption and destruction.
The next normal general-purpose technology driving the next generational doubling of human productive potential while leveling and rebuilding a fifth of the economy.
And, on the other side, the favored post-crypto way to grift the gullible separating gullible, separating VCs and wannabes from their money.
A bubble that knows it is a bubble—boosted by those confusing the long-run benefits to society from technology-driven investment from the actual profits to be earned by the investors.
The Rapture of the Nerds: trying to build a Digital God as the latest and strangest of the many American religious-cult millennarian outbreaks since the First Great Awakening.
Extraordinary plentiful tools for megasuperscale, hyperdimensional, flexible and extensible functional classification analyses at new levels.
A better internet search-and-summarization engine that has not yet been polluted by SEO.
A boilerplate, ritual, and slop text and image and video generator.
Highly verbal software pets and coaches.
Clever Hans at lightning scale and speed: trying thousands of things in the time it would take you to try one, and when and if it can be well and properly harnessed, tremendously productive.
Quite possibly, the savior of authoritarian states that want to live East Germany’s STASI surveillance dream.
Autocomplete on super-steroids: “that’s just kernel smoothing” vs. “you can do that with just kernel smoothing?!”; “that’s just a Markov chain” vs. “you can do that with just a Markov chain?!?!”
A jagged frontier produced by a stochastic parrot that is being harnessed to build a bicycle for the mind.
This week brought me news increasing my confidence with respect to one thing it is not: it is unlikely to be a hyperprofits engine for research labs with companies attached making money serving up calculations made in datacenters by simulated neural networks.
Today we have a review article for an $11,750 machine, the M5Ultra Mac Studio with 256GB of unified memory and a 2TB hard disk( cf.: the original 128K Macintosh’s price back in 1984 was $8000 of today-value dollars, with a Motorola 68000 processor, 128KB of RAM, and 400KB of floppy-disk storage per disk. 5 x 10^9 storage, 2 x 10^6 memory capacity, 5 x 10^5 memory bandwidth, 1 x 10^10 scalar and matrix addition and multiplication computrons, and with respect to the kinds of computations MAMLMs do “the honest ratio is division-by-zero: the 1984 chip couldn’t do it at all”):
Federico Viticci: M5 Ultra Mac Studio Review: The Dream Mac for Local AI Agents <https://www.macstories.net/stories/m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents/>: ‘The M5 Ultra Mac Studio is a dream machine for local AI agents… powered by local models with great performance and no additional cloud costs…. The personal assistants I use the most are now entirely powered by a model running locally on a Mac Studio…. A team of agents… ran 24/7, for 99 days, to… transcribe my favorite WWDC sessions (using
summarizeplus LLM processing)… extract features of iOS and iPadOS 27… cross-reference features across different sources, and keep track… extract features and bugs from screenshots… work with the Notion API to organize…. Relying on the OpenAI or Anthropic APIs for this kind of always-on, persistent background task would be… cost-prohibitive, to say the least. So I pivoted to local AI…. The result is the iOS and iPadOS 27 review you can read on MacStories….Having a model such as Qwen3.8-Flash-Next clear 100 tokens/second on short prompts and still write at 60 to 85 with 64K to 256K of context behind it… enables the kind of agentic back-and-forth between you and the model that feels great to use…. For my taste, 5-bit quantization hits the sweet spot on this version of the Ultra with a balance of intelligence, performance, and memory consumption…
The natural local on-device comparison:
NVIDIA’s RTX 5090 is still faster than Apple’s M5 Ultra…. Prompt processing speeds are dictated by… one giant matrix multiplication, which is exactly the job NVIDIA’s Tensor Cores were built for…. The M5 Ultra read at ~1,700 tok/s; the 5090 delivered a staggering ~3,000…. Token generation… is bandwidth… so the 5090’s 1.79 TB/s against the M5 Ultra’s 1.2 TB/s gives it a steady ~25% lead…. [But] the 5090 doesn’t have… memory…. The moment I want to run anything exceeding 32 GB… the 5090 must offload model layers over PCIe to (much slower) system RAM, and that’s no way to live…
Cf.: the NVIDIA RTX 5090 launched eighteen months ago at a MSRP of $1999 costs $7,500 today, and that is for the graphics card alone:
The bottom line:
I have a simple, tangible result: the M5 Ultra lets me run local agents with incredible performance, with less time spent staring at a blank screen and everything happening on a single, compact, cool, and quiet machine on my desk. This would have seemed impossible a couple of years ago. But here we are…
Running Federico’s agent team as claude-code-sonnet for three months at current API-prices would, I guess, have cost him at least twice the price of his machine.
Now, on the one hand, this kind of always-on, high-input-token workload and sustained agentic use is exactly what the hyperscalers are counting on to monetize the buildout.
And, the other hand, this workload means that $10K+ of Mac (or NVIDIA!) hardware pays for itself in a month. Profitable datacenter demand thus runs on frictions: people who need the technology’s Clever Hans trick—try a million paths, let the harness keep the winners with nine-nines reliability—but who can’t be bothered to hire a work-study IT tech to spin up what the open-weight community already gives away for free.
Thus my Visualization of the Cosmic All with respect to the future of MAMLMs is becoming clearer:
Will anyone other than NVIDIA ever show us real money from the AI datacenter buildout? Probably not.
