Retraining Workers for “AI”: CHART OF THE DAY

No typing-pool annihilation this time, at least not so far, and probably never. AI isn’t killing occupations, but watch out for the elasticity of demand!:

Again, another one from Torsten Slok, who is on a roll these days:

Share

With comment:

Torsten Slok: AI Is Retraining Workers, Not Replacing Them <https://www.apollo.com/wealth/insights-news/insights/daily-spark/ai-is-retraining-workers-not-replacing-them>: ‘A new survey from the New York Fed… shows that 34% of service firms and 22% of manufacturers using AI are retraining staff, while only 4% and 0% report layoffs. This is consistent with our core view that AI is putting downward pressure on wages in AI-exposed occupations without a significant negative impact on employment…. Here: Analysis of actual Claude usage data… suggest[s]… companies are capturing AI productivity gains through wage compression rather than workforce reduction…. [Analysis] by Sania Edlich and me using a difference-in-differences methodology with occupation and year fixed effects across 321 matched occupations from 2015 to today…

Give a gift subscription

What are the jobs that are like the hand spinners, hand weavers, and hand knitters whose entire occupations vanished during the British Industrial Revolution of the 1800s that brought us into the SteamPower Age? What are the jobs that are like the typing pool stenographers and the switchboard operators and the back-office hand reconcilers whose entire occupations similarly vanished with the coming of the personal computer, = the electronic telephone switch, and the computer network?

Looking around, I see damned few of them.

Thus I think the way to bet is that tasks will shift and occupations will change, but whether the number of people employed in them depends overwhelmingly on the elasticity of demand for the kinds of things that they do. And so Jevons’s Paradox is key here, not as a totem and a fetish to wave around, but as an important piece of analysis.

Get 75% off a group subscription

Subscribe now

Leave a comment

If reading this gets you Value Above Replacement, then become a free subscriber to this newsletter. And forward it! And if your VAR from this newsletter is in the three digits or more each year, please become a paid subscriber! I am trying to make you readers—and myself—smarter. Please tell me if I succeed, or how I fail…

##⁠retraining-workers-for-ai-chart-of-the-day
##macro-outlook

##⁠chart-of-the-day
#⁠retraining-workers-for-ai⁠
#ai-and-labor
#retraining-not-replacing
#jevons-paradox
#wage-compression
#ai-exposed-occupations
#torsten-slok
#new-york-fed

READING: LEWIS CARROLL (1895): What The Tortoise Said To Achilles

I think the solution to this is Douglas Hofstadters’s: There are rules, and there are premises. There must be a bottom layer where the system simply acts on a rule, mechanically, rather than contemplating it as another proposition to assent to. You jump out of the loop, and the step from premises to conclusion is something you do, not something you believe:
  1. The Regress That Ate Achilles: Why Logic Can’t Justify Its Own Last Step

  2. What the Tortoise Knew: Lewis Carroll’s Proof That Reasoning Isn’t a Rule

A thousand and one premises: Carroll’s infinite note-book and the floor beneath logic; or, the regress and logic’s unsuccessful attempts to justify its own final operative step:


READING: LEWIS CARROLL (1895): What The Tortoise Said To Achilles

<https://math.dartmouth.edu/~matc/Readers/HowManyAngels/Tortoise.html>

ACHILLES had overtaken the Tortoise, and had seated himself comfortably on its back.

“So you’ve got to the end of our race-course?” said the Tortoise. “Even though it does consist of an infinite series of distances? I thought some wiseacre or other had proved that the thing couldn’t be done?”

“It can be done,” said Achilles. “It has been done! Solvitur ambulando. You see the distances were constantly diminishing; and so –”

“But if they had been constantly increasing?” the Tortoise interrupted. “How then?”

“Then I shouldn’t be here,“ Achilles modestly replied; “and you would have got several times round the world, by this time!”

“You flatter me – flatten, I mean,” said the Tortoise; “for you are a heavy weight, and no mistake! Well now, would you like to hear of a race-course, that most people fancy they can get to the end of in two or three steps, while it really consists of an infinite number of distances, each one longer than the previous one?”

“Very much indeed!” said the Grecian warrior, as he drew from his helmet (few Grecian warriors possessed pockets in those days) an enormous note-book and a pencil. “Proceed! And speak slowly, please! Shorthand isn’t invented yet!”

“That beautiful First Proposition of Euclid!” the Tortoise murmured dreamily. “You admire Euclid?”

“Passionately! So far, at least, as one can admire a treatise that wo’n’t be published for some centuries to come!”

“Well, now, let’s take a little bit of the argument in that First Proposition – just two steps, and the conclusion drawn from them. Kindly enter them in your note-book. And in order to refer to them conveniently, let’s call them A, B, and Z: –

(A) Things that are equal to the same are equal to each other.

(B) The two sides of this Triangle are things that are equal to the same.

(Z) The two sides of this Triangle are equal to each other.

Readers of Euclid will grant, I suppose, that Z follows logically from A and B, so that any one who accepts A and B as true, must accept Z as true?”

“Undoubtedly! The youngest child in a High School – as soon as High Schools are invented, which will not be till some two thousand years later – will grant that.

“And if some reader had not yet accepted A and B as true, he might still accept the sequence as a valid one, I suppose?”

“No doubt such a reader might exist. He might say I accept as true the Hypothetical Proposition that, <em>if A</em> and <em>B</em> be true, <em>Z</em> must be true; but, I <em>don&#8217;t</em> accept <em>A</em> and <em>B</em> as true.&#8217; Such a reader would do wisely in abandoning Euclid, and taking to football.&#8221;</p><p>&#8220;And might there not <em>also</em> be some reader who would sayI accept A and B as true, but I don’t accept the Hypothetical’?”

“Certainly there might. He, also, had better take to football.”

“And neither of these readers,” the Tortoise continued, “is as yet under any logical necessity to accept Z as true?”

“Quite so,” Achilles assented.

“Well, now, I want you to consider me as a reader of the second kind, and to force me, logically, to accept Z as true.”

“A tortoise playing football would be – “ Achilles was beginning

“ – an anomaly, of course,” the Tortoise hastily interrupted. “Don’t wander from the point. Let’s have Z first, and football afterwards!”

“I’m to force you to accept Z, am I?” Achilles said musingly. “And your present position is that you accept A and B, but you don’t accept the Hypothetical –”

“Let’s call it C,“ said the Tortoise.

“– but you don’t accept

(C) If A and B are true, Z must be true.”

“That is my present position,” said the Tortoise.

“Then I must ask you to accept C.

“I’ll do so,” said the Tortoise, “as soon as you’ve entered it in that note-book of yours. What else have you got in it?”

“Only a few memoranda,” said Achilles, nervously fluttering the leaves: “a few memoranda of – of the battles in which I have distinguished myself!”

“Plenty of blank leaves, I see!” the Tortoise cheerily remarked. “We shall need them all!” (Achilles shuddered.) “Now write as I dictate:-

(A) Things that are equal to the same are equal to each other.

(B) The two sides of this Triangle are things that are equal to the same.

(C) If A and B are true, Z must be true.

(Z) The two sides of this Triangle are equal to each other.”

“You should call it D, not Z,“ said Achilles. “It comes next to the other three. If you accept A and B and C, you must accept Z.

“And why must I?”

“Because it follows logically from them. If A and B and C are true, Z must be true. You don’t dispute that, I imagine?”

“If A and B and C are true, Z must be true,” the Tortoise thoughtfully repeated. “That’s another Hypothetical, isn’t it? And, if I failed to see its truth, I might accept A and B and C, and still not accept Z, mightn’t I?”

“You might,” the candid hero admitted; “though such obtuseness would certainly be phenomenal. Still, the event is possible. So I must ask you to grant one more Hypothetical.”

“Very good. I’m quite willing to grant it, as soon as you’ve written it down. We will call it

(D) If A and B and C are true, Z must be true.

Have you entered that in your note-book?”

“I have!” Achilles joyfully exclaimed, as he ran the pencil into its sheath. “And at last we’ve got to the end of this ideal race-course! Now that you accept A and B and C and D, of course you accept Z.

“Do I?” said the Tortoise innocently. “Let’s make that quite clear. I accept A and B and C and D. Suppose I still refused to accept Z?”

“Then Logic would take you by the throat, and force you to do it!” Achilles triumphantly replied. “Logic would tell you `You ca’n’t help yourself. Now that you’ve accepted A and B and C and D, you must accept Z!’ So you’ve no choice, you see.”

“Whatever Logic is good enough to tell me is worth writing down,“ said the Tortoise. “So enter it in your book, please. We will call it

(E) If A and B and C and D are true, Z must be true. Until I’ve granted that, of course I needn’t grant Z. So it’s quite a necessary step, you see?”

“I see,” said Achilles; and there was a touch of sadness in his tone.

Here the narrator, having pressing business at the Bank, was obliged to leave the happy pair, and did not again pass the spot until some months afterwards. When he did so, Achilles was still seated on the back of the much-enduring Tortoise, and was writing in his note-book, which appeared to be nearly full. The Tortoise was saying “Have you got that last step written down? Unless I’ve lost count, that makes a thousand and one. There are several millions more to come. And would you mind, as a personal favour, considering what a lot of instruction this colloquy of ours will provide for the Logicians of the Nineteenth Century – would you mind adopting a pun that my cousin the Mock-Turtle will then make, and allowing yourself to be re-named Taught-Us?”

“As you please!” replied the weary warrior, in the hollow tones of despair, as he buried his face in his hands. “Provided that you, for your part, will adopt a pun the Mock-Turtle never made, and allow yourself to be re-named A Kill-Ease!”

<https://math.dartmouth.edu/~matc/Readers/HowManyAngels/Tortoise.html>


Wikipedia <https://en.wikipedia.org/wiki/What_the_Tortoise_Said_to_Achilles> tells me:

Wikipedia: ‘Several philosophers have tried to resolve Carroll’s paradox. Bertrand Russell discussed the paradox briefly in § 38 of The Principles of Mathematics (1903), distinguishing between implication (associated with the form “if p, then q“), which he held to be a relation between unasserted propositions, and inference (associated with the form “p, therefore q“), which he held to be a relation between asserted propositions; having made this distinction, Russell could deny that the Tortoise’s attempt to treat inferring Z from A and B as equivalent to, or dependent on, agreeing to the hypothetical “If A and B are true, then Z is true.”

Peter Winch, a Wittgensteinian philosopher, discussed the paradox in The Idea of a Social Science and its Relation to Philosophy (1958), where he argued that the paradox showed that “the actual process of drawing an inference, which is after all at the heart of logic, is something which cannot be represented as a logical formula … Learning to infer is not just a matter of being taught about explicit logical relations between propositions; it is learning to do something” (p. 57). Winch goes on to suggest that the moral of the dialogue is a particular case of a general lesson, to the effect that the proper application of rules governing a form of human activity cannot itself be summed up with a set of further rules, and so that “a form of human activity can never be summed up in a set of explicit precepts” (p. 53).

Carroll’s dialogue is apparently the first description of an obstacle to conventionalism about logical truth,[4] later reworked in more sober philosophical terms by W. V. O. Quine.[5]

Perhap the best thumbnail is: Achilles has all the premises and the conclusion staring him in the face—yet the Tortoise will not budge, and every reason he offers only becomes one more premise to grant. Carroll’s 1895 joke turns out to be a load-bearing wall of modern logic. Every attempt to make the Tortoise reason produces another proposition he can politely refuse to grant. The way out isn’t more logic; it’s noticing that drawing a conclusion was never a belief in the first place.

Subscribe now

Leave a comment

If reading this gets you Value Above Replacement, then become a free subscriber to this newsletter. And forward it! And if your VAR from this newsletter is in the three digits or more each year, please become a paid subscriber! I am trying to make you readers—and myself—smarter. Please tell me if I succeed, or how I fail…

##reading-lewis-carroll-1895-what-the-tortoise-said-to-achilles
##reading
##public-reason
##i-have-always-found-this-fascinating-and-i-do-not-know-what-i-should-think-about-it
#lewis-carroll-1895-what-the-tortoise-said-to-achilles
#lewis-carroll
#1895-what-the-tortoise-said-to-achilles
#lewis-carroll
#tortoise-and-achilles
#rules-and-premises
#modus-ponens
#mathematical-logic
#jumping-out-of-the-loop
#zeno-of-elea
#logical-paradox

The Viking Alliance Is Coalescing: CHART OF THE DAY

For eighty years the free world’s security ran on a single hub, and everyone knew who it was: the United States. What the Viking Alliance is building is the first credible answer to what happens if that hub is captured from within by a combination of grifters, neofascists, and chaos monkeys. The people who spent a decade telling Europe to carry its own weight have finally gotten their wish. And the United States is much less powerful a weight in world security affairs as a result.

Capture the hub of a hub-and-spoke alliance, and it collapses. That is the vulnerability NATO has carried since 1945. Putin thought he had captured the hub via his “special relationship” with Donald Trump as Washington went awry. But every place that Vikings ever set foot—Sweden, Canada, and Ukraine at the core; with Germany, the Baltics, the other Nordics; plus France and Britain and Poland; are now building an advance guard: a durable European capability that no longer depends on the United States showing up. For the first time, “and if the Americans don’t come?” has an answer other than “then we lose.” NATO is becoming a web. Webs are much harder to capture than hubs. And they move by rough consensus, not by the will and whims of a single hub.

Share

We have, from Shankar Narayan:

Give a gift subscription

With commentary:

Shankar Narayan: The Coalition of the Doing <https://www.theconcis.com/p/the-coalition-of-the-doing>: ‘Canada supplied the trigger. But… what is rare is actually walking through the door. Sweden did. And now the results are starting to pile up…. One country begins running hard and fast… and, in doing so, starts connecting… making the entire structure stronger than the sum of its parts…. Let us rewind the clock to May 27 and then roll it forward, one decision at a time. Because when you see what happened next, laid out in sequence, there really is no other way to describe it: this has been extraordinary…

Share DeLong's Grasping Reality Weblog


Brad DeLong here: What do I think? This:

Charles Kindleberger taught us that an open international order, whether oriented around trade or security, needs a hegemon willing to be the lender, buyer, and guarantor of “last resort” in economics, and the leader-director in security. In the years after U.S. entry into WWII at the end of 1941 both halves of the post-WWII order outside the Iron Curtain ran on a single anchor, a single hub, the United States.

What Canada triggered, what Sweden walked through, and is now gelling as what Narayan calls the Coalition of the Doing and I call the Viking Alliance is the creation of a new, second potential anchor for the NATO-EU portion of the open international order. What Canada and Sweden and Germany and Ukraine; plus Poland, Lithuania, Latvia, Estonia, Finland, Norway, Denmark, Holland, France, and Britain; are doing is building-up an advance guard with a durable, standing capability for military action inside Europe. Procurement, logistics, industrial base, command: the whole stack. And this advance guard’s capability, when built, does not at any point in the causal chain route through the Washington, DC currently controlled by a bunch of grifters, neofascists, and chaos monkeys. An alliance with one indispensable leader is a hub-and-spoke. Hub-and-spokes are brittle exactly where they look strongest. Capture the hub and the thing collapses. But an alliance with two credible centers of gravity in initiating force commitment is a network with redundancy. For the first time since 1945 there exists a plausible answer to the question “and if the Americans don’t come?” that is not simply “then we lose.”

There is irony here, entirely lost on the people who spent a decade demanding Europe “pay its fair share” have gotten their wish: the United States in foreign affairs no longer has its weight multiplied by NATO to the dominant strength of 1,000,000,000 people living in the rich industrial civilization, but only its own 350,000,000.

Subscribe now

Leave a comment

If reading this gets you Value Above Replacement, then become a free subscriber to this newsletter. And forward it! And if your VAR from this newsletter is in the three digits or more each year, please become a paid subscriber! I am trying to make you readers—and myself—smarter. Please tell me if I succeed, or how I fail…

##the-viking-alliance-is-coalescing-chart-of-the-day
##neofascism
##war-and-rumors-of-war
##chart-of-the-day
##nato-is-getting-a-second-locus-of-initiative-and-washingtons-power-is-no-longer-automatically-multiplied-to-the-strength-of-a-billion-by-nato-but-finds-its-leadership-of-the-alliance-now-contested
#the-viking-alliance-is-coalescing-
#shankar-narayan
#viking-alliance
#coalition-of-the-doing
#nato
#network-redundancy
#european-defense
#chaos-monkeys
#collective-security

CROSSPOST: JUSTIN WOLFERS: Did Trump’s Tariffs Achieve Trump’s Goals?

Justin Wolfers’ grades are straight Fs on all the things Trump promised to get from tariffs: leverage, deficit reduction, factory revival, national security, and revenue. He says that the failures share a single root: a misconception that trade is a zero-sum contest to be won rather than cooperation that both sides benefit from. I disagree. That would be attributing much more logic to Trump’s actions and statements than they deserve. There are people who work for Trump who have the gross misconception that trade is a zero-sum contest to be “won”, yes. But Trump is simply trying to create headlines by doing things. The Supreme Court and the Republican congressional majority have allowed him to do things with tariffs. So he does them. To get headlines. To the extent that there is a goal, it is to “make a deal” in some way. But mostly it is about the headlines.

Justin says: The mechanism runs from a mistaken premise to self-inflicted damage. Trump treated trade as extraction: America gets “ripped off,” so tariffs force better terms. But tariffs triggered retaliation (China to 125%, Canadian boycotts), raised input costs and consumer prices, injected on-again/off-again uncertainty that deterred the factory investment they were meant to spur. Because trade is reciprocal cooperation, throwing sand in the gears cost America customers, suppliers, and trusted partners rather than winning concessions. The “deals” Trump trumpets were, largely, either fictional or already-existing. The goods trade deficit has gotten worse, but i not what we should be looking at anyway. “Reshoring” did not happen as sand in the gears reduced American factory employment. And Trump has advertised a great many supply-chain vulnerabilities that people now have no reason not to exploit.

Share

Grading tariffs against Trump’s own promises and not economists’ ideals yields Justin’s five consecutive Fs:


CROSSPOST: JUSTIN WOLFERS: Did Trump’s Tariffs Achieve Trump’s Goals?

<https://newsletter.platypuseconomics.com/p/did-trumps-tariffs-work-i-used-his> <https://newsletter.platypuseconomics.com/>

Platypus Economics
Did Trump’s Tariffs Achieve Trump’s Goals?
When the Trump administration pushed out their tariffs, there was a laundry list of great things they were going to achieve. Today, I’m asking: did those tariffs do what the administration promised? I’m an economics professor, so I’m approaching this like a report card…
Read more

I didn’t grade the trade war against an economist’s ideal. I graded it against Trump’s own promises.

Justin Wolfers

Sep 02, 2026

When the Trump administration pushed out their tariffs, there was a laundry list of great things they were going to achieve. Today, I’m asking: did those tariffs do what the administration promised? I’m an economics professor, so I’m approaching this like a report card.

President Trump’s tariffs were supposed to do a lot of things. Give America leverage over foreign governments. Shrink the trade deficit. Bring factories home. Make America safer from China. And pay for child care, tax cuts, farmer relief, and tariff dividend checks. Maybe replace the income tax. Or cure toe fungus.

So today, for report card day: Five promises, which we’ll put to five empirical tests, and deliver five grades on.

Here’s the rule I’m using. I’m not grading these tariffs against what I would have done, or against what economists think trade policy should look like. I’m grading them against what the administration itself said the tariffs would deliver.

So: pencils down. Let’s see how the administration’s tariff policy scores on its own test.

Test One: Leverage for Getting Better Deals

Promise one was the tariffs were going to give America leverage over foreign governments. This was an argument that came in two parts. One hinges on fairness, the other on strength.

The fairness claim was that foreign governments were ripping America off with tariffs, subsidies, regulations, currency policies, all of it. The second part of the argument came down to power. America has the world’s biggest consumer market, everyone wants in, so we use that to force other countries to the table.

There’s a few problems with the fairness side. First: the world we actually lived in — at least before the trade war — wasn’t the world the President described. The world actually involves very little protectionism. The arguments for free trade had mostly won the day in most countries.

Canada and Mexico traded with America under USMCA, the free trade agreement Trump himself negotiated in his first term, and most goods crossed those borders duty-free. South Korea had KORUS, and most American manufactured exports already entered Korea tariff-free. The average tariff on American goods was around 3% in the European Union and around 3% in China. Most other rich countries sat in the same neighborhood. A few poor countries ran bigger tariffs, but they’re not large markets for us, so not really a big deal.

So there were tariffs — just very, very low ones. Why weren’t those numbers zero, rather than two or three percent? Because that’s not how trade deals work. When it comes to negotiating trade, two leaders sit down and eliminate tariffs across most of the economy. But they leave aside a handful of politically radioactive sectors — think dairy, rice, sugar, sometimes steel. Those are the sectors where a politician who mishandles them loses their job. So you forgive your counterpart their political weaknesses, and they help you with yours. The result is the attainable trade deal rather than the perfect one: tariffs broadly at zero, hand your counterpart a few political wins, and trade (mostly) freely.

That’s the world America had. The claim that we faced vast tariff walls across the developed world isn’t true, and hasn’t been for decades. It may have been partly true in the President’s youth… But that was a while ago.

Now for the power half. Did the tariffs get America better deals?

There have been many announcements about this — but they don’t amount to much.

Take South Korea. The administration celebrated a new deal opening Korea to American cars and manufactured goods. Except that our trade agreement, KORUS, had already given most American manufactured exports tariff-free entry years ago. So the new arrangement leaves a 15% U.S. tariff on Korean goods and claims credit for market access American manufacturers already had.

Or take the much-touted arrangement with the European Union, which isn’t a deal at all. It’s a framework: a promise to make future promises. The document is written almost entirely in the future tense — “intends,” “will work,” “seeks.” It’s the trade equivalent of “we should really get coffee sometime.”

Still the administration claims to have signed real deals — reciprocal trade agreements — with ten countries: Argentina, Bangladesh, Cambodia, Ecuador, El Salvador, Guatemala, Indonesia, Jordan, Malaysia, and Taiwan. Say that whole list out loud and it sounds impressive. Add up their share of American goods exports and you get about 6%. And the agreements cover only some products and only some barriers, so the share of exports actually affected is far smaller than that.

Oh… but it gets worse. It’s not clear that any of these agreements are actually in effect. Most were struck in response to tariffs imposed under emergency powers, and those tariffs were subsequently ruled unconstitutional. A USTR report from February lists all or nearly all of these as agreements that “have been negotiated, but have not yet entered into force.” So the count might be closer to zero.

Who isn’t represented in that count? The folks we actually trade heavily with: Mexico, Canada, the UK, China, Japan, and Germany, plus (as discussed above) the EU and Korea. This list of counties shows up elsewhere: It’s the list where American access has been restricted, or may soon be, in retaliation. China took tariffs on American exports as high as 125% in April 2025, and while the peak came down, a 10% additional tariff on U.S. goods runs through November 2026. That’s alongside targeted tariffs on American farm and energy products. American goods exports to China fell 26% in 2025. Sales elsewhere rose, but they didn’t replace that market.

The Canadian government retaliated too, and Canadian consumers started quietly protesting in their own way: trips to the U.S. fell sharply, as did their imports. Turns out the surest way to lose a Canadian’s business is to keep calling their country the 51st state.

Grading Leverage: Threats, retaliation, a few small commitments that may not be in force. No serious net gain for American exporters. That’s an F.

Test Two: Reducing the Trade Deficit

The White House called the trade deficit a national emergency. Not just the total trade deficit either: Peter Navarro argued for actions to reduce every bilateral deficit. This would require persuading the rest of the world to want exactly as much American stuff as America wants of theirs. Bold.

In 2024, America bought $1.212 trillion more in goods from the world than it sold. In 2025, the first full year of the tariff program, the goods deficit rose to a record of about $1.24 trillion. Tariffs apply directly to goods, and the goods deficit got worse.

The first half of 2026 does look better — roughly $550 billion, which annualizes to something a bit north of a trillion. That’s an improvement… and a trillion-dollar deficit.

And, as you may have heard, America is a service-focused economy. We sell a lot of that — finance, software, travel, consulting, entertainment, education. The total deficit — including services, this time — was about $904 billion in 2024 and about $902 billion in 2025. If those sound like they’re pretty much the same number, that’s because they are.

A good professor asks his students to show their work, so let’s look to China. America’s goods deficit with China fell by about $94 billion in 2025. That’s good news — until you notice the goods deficit with Southeast Asia rose by about $100 billion over the same stretch. We just changed the labels on the boxes. Imports left China and reappeared in Vietnam, Malaysia, Thailand, and Indonesia. Some of that is real supply chain relocation. Some of it is Chinese firms shipping through third countries. Either way, Americans kept buying.

Grade the Trade Deficit: Bigger in 2025, maybe smaller in 2026, still enormous. Another F.

I’d add that a deficit is an accounting total, not an economic scorecard. The whole here premise is flawed. It’s not at all clear that a better grade on this score would mean a better life for Americans.

Test Three: An Industrial Revival

This is the big one, folks. The one the administration talks about at every opportunity. They said they were going to bring back factories. Big boofy blokes with steel-toed boots bringing home the bacon.

And yet: Manufacturing employment is lower than when Trump returned to office. By July 2026, America had about 62,000 fewer manufacturing jobs than in January 2025. That’s a small number, coming in at roughly half a percent. It’s not a collapse. But we ran an extraordinary trade war to rescue this one sector, and the sector kept shrinking while the rest of the economy added jobs.

Manufacturing output has risen modestly this year, and factory capacity remains loose. American factories are not running flat out, because tariffs did not unleash a wave of new demand for what they make. Factory construction says the same thing: the manufacturing construction boom of the early 2020s was driven by semiconductor investment and industrial policy passed before Trump returned, it peaked in 2024, and it has fallen since. The tariffs arrived after the boom started and during its slowdown.

This is one where the details really deserve a first-hand account. The Dallas Fed put a beautifully simple question to 271 Texas firms: what net impact do you expect higher tariffs to have on your business this year?

59% said negative.

4% said positive.

17% said no impact.

20% didn’t know.

Fifty-nine over four is roughly fifteen — fifteen manufacturers expecting harm for every one expecting help.

And among the firms expecting harm, 55% said they would pass costs to customers. 44% said they would absorb costs as lower profits. Notably, 29% would look for domestic suppliers — that’s something the policy was actually going for, and it’s a positive. 27% would just shift the timing of their imports.

Just 5% planned to move production to the United States.

The Fed’s national small business survey finds the same pattern: 13% of firms using foreign inputs switched to domestic suppliers, and just 3% moved production to America.

There’s a reason nobody’s pouring concrete: if a tariff is on Monday and off on Tuesday, you don’t build a plant around it. The tariff can flip several more times before the concrete has dried.

Grading the Industrial Revival: Some domestic sourcing, fewer factory jobs, no revival. A clear F.

Test Four: National Security

I don’t want it to seem like I’m going through this report on the premise that there’s no point going after these goals. There is a real trade policy case for targeting strategic risks. America does need secure access to rare earths, magnets, chips, medicines, and specialized metals.

The trouble is that most of this trade war wasn’t targeted at all. And where it was targeted, it backfired.

Rare earths are misnamed — they aren’t especially rare. The scarcity comes in who processes them. China does a lot of that processing, and they do it for the entire world. That means they turn raw material into magnets. Those magnets go into cars, aircraft, electronics, and military equipment. Before the trade war, China supplied around 70% of the rare earth compounds and metals America imported.

Then things escalated, and China restricted exports of critical rare earths and magnets. The White House’s own economic report says those restrictions caused factory shutdowns, including in the U.S. China has since used export controls on gallium, germanium, graphite, and antimony too. No, those aren’t words I made up to sound like a scientist (please don’t ever think that I am a scientist). But those critical minerals matter for semiconductors, batteries, weapons, and advanced manufacturing.

Here’s the part that keeps me up. The dependence was always there. But a dependence only becomes a vulnerability once your adversary discovers it — and this trade war sent them looking. They found it. Now they know exactly where to press, and they’ve shown that they’re willing.

The administration has announced domestic mining and magnet projects, and those may help reduce these vulnerabilities. They also have nothing to do with the tariffs. And don’t get me started on the Strait of Hormuz and the rest of what we import from that part of the world.

Grading National Security: The vulnerability was revealed, not reduced. F.

Test Five: Tariffs Raise Revenue

Here’s a partial list of what the President promised that tariff revenue would fund. Child care. Tax cuts. No tax on tips. No tax on overtime. No tax on Social Security. Tax benefits for American cars. Farmer relief. Tariff dividend checks — remember those? Mine never arrived. Debt reduction. And the end of the income tax.

Tariffs are taxes, and taxes do two things: they change behavior, and they raise revenue.

This program certainly changed behavior. Families paid higher prices, businesses paid higher input costs, and supply chains reorganized themselves around dodging the tariff.

The revenue is the strange part. Customs duties rose from $77 billion in fiscal 2024 to $195 billion in fiscal 2025 — an increase of about $118 billion. Much of that increase came from tariffs imposed under the International Emergency Economic Powers Act, and the Supreme Court struck those down earlier this year. By mid-August, roughly $100 billion had already been refunded.

These refunds don’t work like they do at a store. When I paid more for olive oil at Costco because of a tariff, the refund didn’t come to me. It went to the importer of record. Yes: that’s the company on the customs paperwork. So… Costco. The family at the checkout paid while the big importer got the check.

More is likely coming. The Section 122 replacement tariffs — a temporary 10% tariff meant for a balance of payments crisis — were struck down at trial because there was no balance of payments crisis, and that’s on appeal. The newer Section 301 tariffs, the ones you’re paying right now, rest on the premise that we’re punishing other countries for their use of forced labor. Which countries? Apparently all of them. That’s a pretext, and everyone involved knows it.

Why the parade of odd legal theories? Because the Constitution gives the tariff power to Congress. Congress has occasionally lent narrow slices of it to the White House, but nobody ever intended it as something a president waves around at will — and this administration has consistently declined to go ask Congress for it. The courts occasionally suggest we look at the Constitution, and it’s unclear whether the administration will gather much revenue here at all.

Of course, if the administration passed these tariffs as laws, they wouldn’t have any of these problems. The revenue problems are the direct result of the President refusing to involve Congress in his trade war.

Grade Revenue: Americans got the distortion, and much of the money is being handed back to importers rather than kept by the Treasury. F.

Share

Final Grade: Time to Call the Parents

Let me put the report card up one last time.

Leverage: threats and retaliation, no serious net gain. The deficit: worse in 2025, smaller in 2026, still enormous. Factories: a little domestic sourcing, no revival. Security: China found the choke point. Revenue: Americans paid, importers got refunded.

These failures look different from one another, but they share a root, and it’s an idea about what trade is.

Trade is cooperation. A farmer gets a customer. A factory gets a component. A family gets a product. An American business gets a buyer abroad.

The trade war threw sand in all of it. It disrupted export markets. It disrupted supply chains. It disrupted investment. It disrupted relationships with allies. Then it added uncertainty and stirred.

A stronger America has more customers, more suppliers, more trusted partners, and more capacity to make the things it needs. This trade war has delivered fewer of every one of them.

Five promises. Five tests. Five fails.

And this is grading the President on his own stated goals. As I’ve said before, those goals are themselves questionable:

Platypus Economics
The Lawyer’s Theory of Trade
Read more

<https://newsletter.platypuseconomics.com/p/did-trumps-tariffs-work-i-used-his> <https://newsletter.platypuseconomics.com/>

Platypus Economics
Did Trump’s Tariffs Achieve Trump’s Goals?
When the Trump administration pushed out their tariffs, there was a laundry list of great things they were going to achieve. Today, I’m asking: did those tariffs do what the administration promised? I’m an economics professor, so I’m approaching this like a report card…
Read more

Brad DeLong here: It is excellent to welcome Justin Wolfers to the WebLog-o-Sphere, or I suppose these days we should call it the SubStack-a-Thon. (I do think, all-in-all, that the SubStack Honchos’ plans to try to become the place for people who do not want to have their brains hacked by malevolent actors is worth leaning into and supporting.) He has been blogging a piece a day since April 21.

Give a gift subscription

Today he hits the sweet spot, and is very much worth crossposting.

Justin is right not only in that Trump’s chaos-monkey trade wars off-again-on-again Trump-Always-Chickens-Out TACO have been a disaster not just from an economists’ point of view, but also from a point of view that rationalizes Trump’s own stated goals.

The “deal”s” Trump celebrates (Korea, EU) either restate access American firms already had or are aspirational “frameworks” in the future tense; the ten “reciprocal” deals cover ~6% of exports and may not be legally in force. The goods deficit hit a record ~$1.24T in 2025. But that is not a measure we should be looking at. And China has learned how much potential leverage over the U.S. it has with rare-earths and critical-minerals: a lot.

Justin cuts through the announcement theater of framework “deals” and bilateral-deficit rhetoric with verifiable data, and clarifies conceptual errors. — deficits as accounting identities, dependence vs. vulnerability & c. Leverage, deficit, factories, security, revenue—all failed, because the war threw sand in the gears of productive economic cooperation that is the reason for trade.

However, I profoundly disagree with an underlying assumption of Justin’s piece: Justin claims that Trump has been trying to follow a rational policy based on his false belief that trade is a zero-sum struggle. Grant Trump his bad model of the world, the framing runs, and the tariffs become the logical policy moves that follow from it; they simply fail on their own terms because the model is wrong.

That concedes far too much.

That imputes a non-existent means-ends rationality to Trump and the Tru,p administration.

That takes a chaotic set of actions, and constructs underneath them a stable set of goals, a theory connecting instruments to those goals, and a willingness to be corrected by evidence.

But that is nowhere in evidence. The “laundry list” of promises Justin so ably demolishes was never a plan. It was a rotating grab-bag of justifications, generated after the fact and abandoned the moment a new audience or a new grievance required a different one. The better model is not “wrong beliefs rationally applied” but the near-absence of the belief-to-action link that rationality requires.

Tariff policy here is a dominance display and a mechanism for extracting tribute, deference, and the pleasure of being courted. Those are ends in themselves. They are not instruments toward national prosperity. That is why the tariffs go on Monday and off Tuesday, why the legal theories are transparent pretexts nobody is meant to believe, why “deals” are announced that restate access we already had, and why the same measures are defended one week as leverage, the next as revenue, the next as reindustrialization.

The chaos is not a bug in the execution of a zero-sum worldview. The worldview is not doing any work. To treat the policy as the sincere, if flawed, application of mercantilist doctrine is to flatter it with a coherence it does not possess. Worse, it invites the reply that the doctrine simply needs better technicians next time.

The truth is this: it is chaos monkeys all the way down.

Get 75% off a group subscription


The most interesting piece of data to me was the Dallas Fed survey: it found manufacturers overwhelmingly expecting economic harm from Trump’s chaos-monkey tariffs <https://www.dallasfed.org/research/surveys/tbos/2025/2504q#tab-tmos>:

Share DeLong’s Grasping Reality Weblog

Let’s just leave it there.

Subscribe now

Leave a comment

If reading this gets you Value Above Replacement, then become a free subscriber to this newsletter. And forward it! And if your VAR from this newsletter is in the three digits or more each year, please become a paid subscriber! I am trying to make you readers—and myself—smarter. Please tell me if I succeed, or how I fail…

##crosspost-justin-wolfers-did-trumps-tariffs-achieve-trumps-goals
##macro-outlook
##neofascism
##chaos-monkey
##crosspost
##justin-subheadline-i-didnt-grade-the-trade-war-against-an-economists-ideal-i-graded-it-against-trumps-own-promises-f-justin-says-due-to-wrong-beliefs-about-how-world-trade-works-i-disagree
#justin-wolfers-did-trumps-tariffs-achieve-trumps-goals
#justin-wolfers
#did-trumps-tariffs-achieve-trumps-goals
#trump-tariffs
#trade-wars
#rare-earths
#platypus-economics
#trade-is-productive-cooperation

How Special Is Programming as a Token-Infall Attractor?: CHART OF THE DAY

Software developers are the AI industry’s best customers. Is their experience also the most misleading experience possible for its economy-wide scope? Roughly 80% of developers already use AI tools. Coding may now be more than half of OpenAI and Anthropic’s revenue, currenly running at $80 billion a year as of the second quarter of 2026. But to what else could that even roughly scale??

Matter spirals into a black hole, gaining speed and hence mass and energy as it moves closer converting its gravitational potential energy into kinetic-thermal and then, as particles collide, electromagnetic. It shines with the brightness of ten trillion suns: a quasar. Tokens spiral into an occupation, and are there harnessed to do the work of humanity, and shine—well, the metaphor is strained. But the point is that some occupations are, in the value of the work they can use tokens emitted by a properly harnessed LLM—Large Language Model—to do, like quasars. Others are not.

Share

Paul Kedrosky picks this up from The Economist of London, and points out that computer programming is close to unique in how much LLM-generated tokens it can get ueul work out of:

Share DeLong's Grasping Reality Weblog

And he comments:

Paul Kedrosky: Why Software Developers Are Highly Unrepresentative of Broader AI Use <https://paulkedrosky.com/why-software-developers-are-hiunrepresentative-of-broader-ai-use/>: ‘Will anyone ultimately use AI as intensively as software developers do? That matters because coding is already a huge share of AI usage and revenue: roughly 80% of developers use AI coding tools, and coding may account for more than half of combined OpenAI and Anthropic ARR, by some estimates. See the… figure…. Note that it has finance in the wrong place, with the biggest banks, like Goldman and JPM, having more like 15% of employees in software development….

Coding is unusually expansive: A short prompt can trigger planning, code generation, testing, debugging, retries and repeated ingestion of a large codebase. Token use can explode relative to the size of the initial request. Much other white-collar work is more compressive: Summarization, document review, research synthesis, meeting notes and similar tasks take large inputs and produce relatively small outputs. Adoption and token intensity become different questions. AI could become ubiquitous across law, finance, consulting and management Those occupations might still consume far fewer tokens per worker than software development.

That matters directly for the capex thesis: The infrastructure buildout requires not just broad AI adoption, but enormous sustained token consumption. The Economist estimates annual AI revenue would need to rise from roughly $150bn today to about $2.5tn by decade-end.

The key forecasting error is treating coding as merely early rather than structurally different: If software development is both an early adopter and one of the most token-expansive occupations, extrapolating its usage curve across the rest of the economy will systematically overstate eventual compute demand…

Share DeLong's Grasping Reality Weblog

In programming, you set the goal and the harness and the LLM begins doing its Clever Hans thing, stamping its foot and watching out for the desired reaction at which point it is done. Only this Clever Hans’s foot-stamping is at nanosecond speed, with a big model to capture far-away dependencies and the huge amount of relevant data that is the set of all running code that has ever been written, and repeating the actions over and over and over again until by lucky chance the desired result is achieved. That is, Paul Kedrosky, not going to be how knowledge workers in other professions are going to use “AI”.

Paul is riffing off of this from The Economist of London:

Anonymous: Will Anybody Use AI as Much as Coders Do? <https://www-economist-com.libproxy.berkeley.edu/business/2026/08/30/will-anybody-use-ai-as-much-as-coders-do>: ‘The answer will have big implications for the investment boom…. Uptake of the technology has been strongest by far among software developers. Four-fifths of them say they use an AI coding tool…. In June 2025 the combined annual recurring revenue of Cognition, Cursor, Lovable and Replit, four AI-coding startups, was roughly $800m. Today it stands at $6bn…. Lawyering, finance and customer service… bear some similarities to coding…. But four factors set coding apart: the availability of training data; how easy it is to test a model’s output; the amount of human interaction involved in the work; and software engineers themselves. AI companies are trying to make their other markets more coding-like, but doing so will not be straightforward…

Give a gift subscription


Brad DeLong here: Briefly, coding is expansive while other white-collar work is compressive. A short program, or these days a short prompt, can and does trigger planning, code generation, testing, debugging, retries and repeated ingestion of a large codebase followed by a great deal of computation and reams and reams of output. In a world in which the computer will try twenty versions of a command before it hits upon the correct format, token creation and use explodes as the machine groups toward an anwer. By contrast, the business of a lawyer is to take the expansive set of legal codes and case situations that is the law and squeeze it down into a brief, an opinion, a recommendation. The business of a consultant is much the same. And the whole point of management is to throw away as much information as you can in order to make the problems of direction and coordination graspable and actionable. Summarization, document review, research synthesis, meeting notes and similar tasks. Large inputs, and relatively small outputs. No explosion of agentic activity once the universe of input documents has been defined and collected.

As Paul Kedrosky says, AI could become ubiquitous across law, finance, and consulting, yet those workers would still burn far fewer tokens each than software developers do.

Subscribe now

Leave a comment

If reading this gets you Value Above Replacement, then become a free subscriber to this newsletter. And forward it! And if your VAR from this newsletter is in the three digits or more each year, please become a paid subscriber! I am trying to make you readers—and myself—smarter. Please tell me if I succeed, or how I fail…

##how-special-is-programming-as-a-token-infall-attractor-chart-of-the-day
##macro-outlook
##mamlms
##chart-of-the-day
##software-coding-burns-tokens-like-a-quasar-burns-infalling-particles-gravitational-potential-energy-but-will-any-other-white-collar-occupation-every-do-the-same
#how-special-is-programming-as-a-token-infall-attractor
#token-infall
#expansive-vs-compressive-occupations
#token-consumption
#white-collar-work
#paul-kedrosky
#the-economist

The Fed Chairman Who Would Prefer Not to Say: Warsh’s Silence as a Position, Not a Puzzle: TUESDAY MACRO OUTLOOK

The factors of production are these: labor, capital accumulation, enterprise and innovation, and risk-bearing. The cost of risk-bearing is increased when there is an unstable standard of value. Kevin Warsh is manufacturing that instability by refusing to settle on a reaction function for the central bank. Read as a Fed Chair’s vision of the economy, Kevin Warsh’s Jackson Hole speech confuses. Read as fog-generation, it is crystal clear in its way:

Demoiselle Marchée Financiere thought that Kevin Warsh’s speech at Jackson hole was important enough to lower the real value of everything two years in the future relative to today by fully 0.16%. That is, the real wealth of worker skills, capital and infrastructure investments, and production networks we expect to see in two years lost $320 billion of its value relative to today as Kevin Warsh gave his speech:

Share

Interesting—and probably not what Warsh intended.

My theory of Federal Reserve Chair Kevin Warsh is this: he is way out over his skis, having gotten the job by promising Trump that he would lower inflation and lower interest rates, while reassuring the bond market that he was really a hard-money guy. Now he is stuck trying to create strategic ambiguity. I said this in June, and I said this in May. Having arrived at the top of the greasy pole, Warsh has discovered that the only move available to a man who has told incompatible things to incompatible audiences is to say as little as possible, for as long as possible. That is not a communications strategy. It is a hostage situation, and the hostage is the standard of value.

I have seen nothing to challenge that theory.

Thus, in my view, people who want to understand Warsh need to start with he is trying to create strategic ambiguity to avoid a full-fledged fundamental break with either Semi-Senile Chaos-Monkey Trump on the one hand or Demoiselle Marchée Financiere on the other. If they don’t start there, they wind up confused.

For example, the sharp Tim Duy yesterday morning:

Tim Duy: Fed Watch <sghmacro.com>: ‘Warsh… acknowledge[d] that rate hikes could be necessary to put inflation back on a path to 2%…. Warsh… [had] refused to make that one simple admission…. Warsh… did not provide a forward-looking assessment… nor did he provide a timeline…. We have been confident that the Fed will need to hike rates… [but] Warsh seeks to eliminate the certainty around individual rate decisions that market participants crave….

Warsh… outlined seven key principles…. Caution about using backward looking data to extrapolate forward…. Broadly align aggregate supply and aggregate demand…. The 2% PCE target…. Full employment, adding that it is compatible with price stability…. Policy acts through short term rates…. Money matters…. Less communication [from the Fed]….

Warsh add[ed] a fresh definition of underlying inflation…. By our count, this is Warsh’s fourth…. First… trimmed mean PCE. Second… the five hundred millionth and one price. Third… the median price at big box stores…. Fourth metric… “disaggregate the 199 individual components of the PCE price measure. Over the past 12 months, 54 percent… showed price increases above 3 percent. This is well below the post-pandemic highs of about 77 percent, but it remains well above the level of 32 percent in the two decades that preceded the pandemic…”.

Even after acknowledging the existing inflation pressures, Warsh doesn’t say that the trends represent upside risks for inflation…. Remember, he declared earlier in the speech the importance of not extrapolating trends…

Give a gift subscription

Or Robert Armstrong:

Robert Armstrong: Warsh Settles Some Nerves at Jackson Hole <https://www.ft.com/content/646b812e-c9de-49ba-90f3-f9d12205f876>: ‘Speech… leaves open questions over Fed ‘reaction function’…. Friday’s speech… was a matter of incremental clarification…. For markets, the most important line of the speech was Warsh saying that this summer’s somewhat softer inflation readings did not convince him the underlying trend is improving.… Warsh did not do much to solve this puzzle, except to emphasise that he really, really does not like the Fed forecasting the economy and the future path of policy…. [But] the Fed’s credibility in markets hinges upon collective understanding of the bank’s “reaction function” — how and when it will respond to changes in the economy. If not through forecasts, how to make this known? Warsh raised this crucial question and answered it only with virtuous generalities about humbleness and empiricism. The market will not be satisfied with that for long…

Share DeLong’s Grasping Reality Weblog

Try to use this to construct insights into Warsh’s thinking, and you end up confused. Interpret this as an attempt to preserve strategic ambiguity, and it is crystal clear.

Armstrong and Duy are only two of the large number of very smart people in my feed trying, in good faith, to figure out what Kevin Warsh actually thinks. Each of them has come away holding a fistful of fog. Let me try to say what I mean by that:

Read more

CROSSPOST: DAVID ALBOUY, JONATHAN PARKER, STEPHEN DURLAUF, JON STEINSSON: On Acemoglu, Johnson, & Robinson’s “Colonial Origins...”

To call “perceived probability of expropriation in the 1980s” by the name of “institutions” is, at best, really weird. And David Albouy should not have to be chill with respect to Acemoglu, Johnson, & Robinson’s unwillingness to recognize that the IV analysis in their “Colonial Origins…” is an embarrassment that should never have been in the paper. Other than that, I greatly enjoyed the play. I think Albouy, Parker, and Durlauf are more than fair here:

We have:

Jonathan Parker: <https://x.com/ProfJAParker/status/2094567917048226040>: ‘Steve has the right way to interpret the important research & evidence in AJR & Daron’s related work. It fits a given narrative to historical details, tells a rich plausible story through that lens with lots of evidence. (And David’s work is important & deserves respect)…

Share

And:

Steve Durlauf: <https://x.com/sndurlauf/status/2094545774080307362>: ‘David Albouy @albouy is entirely justified in being aggrieved. It was completely inappropriate for @TheEconomist to publish a response to their (ridiculous) article on Daron Acemoglu that contained a misstatement of the Albouy criticisms, just at it was inappropriate that the article’s author did not contact David to describe the nature and import of his criticisms for the ideas that AJR have developed, as opposed to the merits of a particular set of statistical exercises. There are huge questions surrounding the ways that credible empirical claims can be made about questions of the type AJR address.

Cross-country regressions have proven to a very fragile source of information on the mechanisms underlying growth and development. My view of the AJR research program is that the empirical dimension should be understood from the vantage point of inference to the best explanation, aka abduction, as a way to understand their research strategy, which implies historical/qualitative evidence is central. For details, see my paper

“Institutions, Development, and Growth: Where Does Evidence Stand?” Handbook of Economic Development and Institutions, Jean-Marie Baland, François Bourguignon, Jean-Philippe Platteau, and Thierry Verdier, eds., Princeton University Press, 2020…

Give a gift subscription

And:

David Albouy: <https://x.com/albouy/status/2094501978818621440>: ‘I tried to be chill, but Acemoglu’s false statement reminded me too much of the gaslighting I had to deal with to publish my comment. Acemoglu et al don’t respond directly to criticisms, and often make up demonstrably false claims to defend themselves, with no accountability:

If anyone wants to read my comment, it’s available here: <https://static1.squarespace.com/static/62f1eb5aa4471e693f087c96/t/630594445e4b432c6913e58c/1661310023111/AJRrev.pdf>. The data appendix is available on my website if you want to really dive into dumpster fire of their data construction. The real crime was in the cover-up. Eg Acemoglu et al never confronted the embarrassment of my Figure 1:

Instead they issued a reply full of red herrings with the bullying title “Hither Thou Shall Come but no Further”. They could have acted as scientists but opted to confuse.

As a grad student, I emailed and visited Robinson’s office many times to talk about my preliminary investigations. He was never available. When I finally confronted him, he waved me away, saying, “I think you’re looking for a different Robinson”. Unbelievable, but true. I tracked down Johnson at an AEA meeting. In the spirit of scientific collaboration, I offered to sit down and go through the data points with him, he nodded in agreement saying that would be nice.

I never got a response to any emails after that. As for Acemoglu, I tried to make peace. But i still remember in 2007 after my R&R at the AER, I submitted a paper to REStat where Acemoglu was editor, trusting that he’d recuse himself and not handle it. A few months later, I got a rejection letter signed by Acemoglu…

Share DeLong’s Grasping Reality Weblog

And:

Jon Steinsson: To David Albouy <https://x.com/JonSteinsson/status/2094559148960833707>: ‘Let me offer a tiny consolation: I teach your critique to 1st year econ PhD students every year <https://jonsteinsson.com/teaching/FundamentalCauses.pdf>:

Share DeLong’s Grasping Reality Weblog


Brad DeLong here: I wrote about how the IV is a dumpster fire only last April. My view has long been that (a) David Albouy is 100% correct here, and that (b) if Acemoglu, Johnson, and Robinson were wiser, they would recognize that David Albouy is their best friend in the world, for he provides an explanation for what are otherwise very uncomfortable—clock-striking-13-uncomfortable—IV results.

Briefly: AJR simply have no real first stage. The “1-in-1,000” significance on log-mortality-to-expropriation-risk melts to 1-in-25 once you notice the mortality assignments are smuggling in continent dummies, and to 1-in-3 once you separate barracks deaths from campaign deaths and laborer data. With a first stage that weak, the second-stage test statistic isn’t a t-distribution — it’s near-Cauchy: infinite variance, no standard deviation, and no mean at all. You are drawing noise from fat tails and calling it a coefficient. The Cauchy distribution is the devil — the lesson my sometime-roommate John Bound taught me back when we were both much younger. Weak first stage in, garbage out. Only the file drawer Darwinian-selection process launders that garbage into a “finding”. Weak instruments plus sampling on the instrument that produces a significant IV coefficient is the replication crisis in a clown suit, driving a clown car. And often you really do not want the refreshments in the trunk of that clown car, as AJR would recognize if they were wiser.

This is what I wrote back four months ago:


substackcdn.com/image/fet…}],"post_date":"2026-04-09T00:32:42.739Z","coverimage":"https://substackcdn.com/image/fetch/$s!oSWk!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F8c9baf6f-213e-4c88-a086-9cb8fa2076be_1122x1112.png","cover_image_alt":null,"canonical_url":"https://braddelong.substack.com/p/contemporary-governance-and-contemporary","section_name":"Enlarging the Bounds of Human Empire","video_upload_id":null,"id":193571910,"type":"newsletter","reaction_count":26,"comment_count":79,"publication_id":47874,"publication_name":"DeLong’s Grasping Reality Weblog ","publication_logourl":"[substackcdn.com/image/fet…](https://substackcdn.com/image/fetch/$s!PgPl!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2Ffde2453e-9c18-4560-82ca-8b77ae62ef5b_1280x1280.png","belowTheFold":true,"youtube_url":null,"show_links":null,"feed_url":null)}“>

Get 75% off a group subscription

Colonial origins, causal claims, the baggage left behind in the overhead bin via the absence of a structural model, settler-colonist mortality, modern pro-prosperity “institutions”, and structural-empirical truth…

I was thinking I would do an economic history post yesterday and a political economy post today, but life is busy, busy, busy, what with chaos, staring at screens watching people get blown up and so forth. Then I sat across from Jón Steinsson at the faculty lunch, and in the course of the conversation he mentioned that he still taught the graduate students the quarter century-old Acemoglu, Johnson, and Robinson “Colonial Origins…” paper.

Why? Because it is both incredibly strong and incredibly weak, a true rabbit and a true duck, depending on how you look at. Teach it, and students get very valuable experience in having strong reactions and figuring out how to explain them—and also (we hope) practice in listening to people with whom you violently disagree but who think what they think for reasons.

AJR’s “Colonial Origins” is surely among the most influential empirical paper in historical development economics of the last quarter-century. Its argument is elegant:

  • European settler colonialists who found that they could survive and thrive in a colony built inclusive pro-growth developmental institutions.

  • European settlers who found that they couldn’t, and that they needed to grab what they could and return home before they succumbed to yellow fever or such, did not.

  • Places with high European settler mortality saw the development of “extractive institutions”, hostile to widely distributed prosperity both in the past and in the present.

  • Those early institutions persisted to this day.

  • Colonial-era European settler mortality gives us a lever—a valid instrument—to identify the causal effect of institutions on prosperity.

The result that Acemoglu, Johnson, & Robinson claim?: that differences in institutions explain about three-quarters of the income per capita differences across former colonies. Geography, latitude, disease burden—once you control for institutions, they do not matter.

But this is a paper that turned me into a Heckmanite—into a believer in the idea that only those who have fully specified structural models, at least in their mind’s eye if not out there on the table, have a valid warrant to do anything statistical that they claim is in any sense “causal”. Structural models keep you from abandoning burdensome baggage in the overhead compartment when you exit the plane—and AJR’s procedures and paper has a lot of such, which the non-structural IV-framing hides from view. Briefly: either prosperity has a strong negative structural effect on governance quality—which contradicts everything we know from political science and history—or settler mortality in the 17th century is a better measure of what matters in modern institutions than present-day institutional analysis by people who know about and can see them is.

Neither claim is comfortable.

Yet both are concealed by the IV design.

In any event, here is a link to where you can pull it down as an interactive python document: <https://github.com/braddelong/working_20251227/blob/main/2026-04-07-DELIVERED-EDITED-econ-196-week-9-reversals-of-fortune.ipynb>.


Subscribe now

Leave a comment

If reading this gets you Value Above Replacement, then become a free subscriber to this newsletter. And forward it! And if your VAR from this newsletter is in the three digits or more each year, please become a paid subscriber! I am trying to make you readers—and myself—smarter. Please tell me if I succeed, or how I fail…

##crosspost-david-albouy-jonathan-parker-stephen-durlauf-jon-steinsson-on-acemoglu-johnson-robinsons-colonial-origins
##enlarging-the-bounds-of-human-empire
##public-reason
##crosspost
##the-scatter-and-the-hypotheses-of-acemoglu-johnson-robinsons-colonial-origins-paper-are-very-interesting-and-important-but-the-iv-statistics-are-as-david-albouy-says-a-dumpster-fire
#david-albouy-jonathan-parker-stephen-durlauf-jon-steinsson-on-acemoglu-johnson-robinsons-colonial-origins
#david-albouy
#jonathan-parker
#stephen-durlauf
#jon-steinsson
#acemoglu-johnson-robinsons-colonial-origins
#acemoglu-johnson-robinson
#colonial-origins
#weak-instruments
#extractive-institutions
#inclusive-institutions
#cauchy-distribution
#heckmanite
#cross-country-regressions
#clock-strikes-thirteen
#economic-history
#causal-anti-inference

CROSSPOST: EMILY BENDER: “Stochastic Parrots”

The Big Bear of “AI”, but a thinker whose thought is richer and much more complex than is conveyed by the “stochastic parrots” thumbnail meme. A.B. Linguistics Berkeley 1995, Ph.D. Stanford 2000 with thesis on the absence of the copula in AAVE; then Berkeley, Stanford, YY Technologies, University of Washington; co-author of Syntactic Theory: A Formal Introduction. She in 2000 gave the skeptical position three things: a meme—stochastic parrots—a thought experiment—the octopus—and a target—the leaning-in to anthropomorphization that underpins all of the “achieving AGI” and “foothills of the Singularity” hype.

Emily Bender’s “Climbing Towards NLU” with Alexander Koller proposes two people on separate deserted islands, communicating via an underwater telegraph cable, but one of them has been replaced by a hyper-intelligent octopus who has detected the statistical patterns in the exchanges. But because the octopus has only ever seen the sequence of signals and never the things they refer to, the moment one islander faces a genuine novelty—say, a bear attack, and urgently asks how to build a weapon from the materials at hand—the octopus cannot give real help and can only produce fluent, plausible-sounding replies without grounding in the actual world. LLMs are Stochastic Parrots, as the team she was on wrote in their “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?”

Share

Why has her work struck such a nerve? Well, “Stochastic Parrots” is a very good paper. It has the catchy meme angle that tends toward virality. And Google tried to kill it, or at least to say that these thoughts are not thoughts that anyone working at Google has or is allowed to have.

As I understand it, Google has an internal pre-publication review process called PubApprove—a check that a paper doesn’t leak proprietary information or trade secrets, not a peer review. ​⁠The paper cleared PubApprove. Vice President Megan Kacholia then ordered Temit Gebru to either pull the paper submission from the conference or remove all Google authors’ names from it. Gebru asked for transparency and for meetings where she could make her case, and said if those conditions couldn’t be met she’d negotiate a departure date after her vacation. Google treated that as a resignation, cut off her email while she was on vacation, and she was out. Margaret Mitchell, her co-lead, was fired a couple of months later.

Google tells false stories about why they did what they did. Jeff Dean, head of Google AI, claimed the paper “didn’t meet our bar for publication” via PubApprove because it “ignored too much relevant research” and was submitted to PubApprove with only a day’s notice, too late for proper review. The “ignored relevant research” was bullshit: PubApprove is a sensitive-information check. And nearly half of all papers going through PubApprove were submitted with a day or less notice. ​⁠

The references:

  • Bender, Emily M., Timnit Gebru, Angelina Mcmillan-Major, & Shmargaret Shmitchell. 2021. “On the Dangers of Stochastic Parrots: Can Language Models Be Too Big?” Proceedings of the 2021 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’21), 610–623. March. <https://dl.acm.org/doi/10.1145/3442188.3445922>.

  • Bender, Emily M., & Alexander Koller. 2020. “Climbing towards NLU: On Meaning, Form, and Understanding in the Age of Data”. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, 5185–5198. July. <https://aclanthology.org/2020.acl-main.463/>.

Give a gift subscription

Why is it Emily Bender rather than Timnit Gebru or Margaret Mitchell who shows up in my feed these days? After Google lit its reputation as a place where people could do serious work on large chunks of the study of “AI” on fire, Bender took the public-explainer path in a way that Gebru and Mitchell—now heads of DAIR (the Distributed AI Research Institute) and aresearcher and ethics lead at Hugging Face—did not.


<https://www.youtube.com/watch?v=ZUIQYQXsEvw>

Get 75% off a group subscription


Brad DeLong here: With respect to:

A faces an emergency. She is suddenly pursued by an angry bear. She grabs a couple of sticks and frantically asks B to come up with a way to construct a weapon to defend herself. Of course, O has no idea what A “means”. Solving a task like this requires the ability to map accurately between words and real-world entities (as well as reasoning and creative thinking). It is at this point that O would fail the Turing test, if A hadn’t been eaten by the bear before noticing the deception. Having only form available as training data, O did not learn meaning…. Because agents who produce English sentences usually have communicative intents… [A] assumes that O does too, and thus she builds the conventional meaning English associates with O’s utterances. Because she assumes that O is B, she uses that conventional meaning together with her other guesses about B’s state of mind and goals to attribute communicative intent. It is not that O’s utterances make sense, but rather, that A can make sense of them…

Share DeLong’s Grasping Reality Weblog

Well, yes. But.

Read more

Heading for a Large, Rapid Fall in Interest Rates?: CHART OF THE DAY

Well! Large, directional predictions of interest rate movements from smart people get my attention! The rule “never make large, directional predictions of interest rate movements”is up there with “never get involved in a land war in Asia”! The chain of causation is this: AI-success makes Productivity outruns consumption and thus savings rises; AI-failure sends money fleeing into safe Treasuries; Either way, the 30-year U.S. Treasury bond yield falls from 5% back towards 3%:

Torsten Slok makes a large, directional prediction: the risk is rising that interest rates will go down a lot over the next six months:

Share

Share DeLong's Grasping Reality Weblog

The commentary:

Torsten Slok: <https://www.apollo.com/wealth/insights-news/insights/daily-spark#page-1>: ‘The risks are rising that long rates six months from now could be a lot lower than where they are today…. Inflation [fears] and [deficit] fiscal problems… could end up being dominated in early 2027 by what happens to AI…. If AI succeeds and tech companies generate trillions in revenue, AI will be massively deflationary and push rates lower. If AI does not work out, the bubble bursts and the Nasdaq is down 50% as investors rotate out of equities into Treasuries and long rates fall dramatically.

Over the next six months, the market will make up its mind about which AI scenario is playing out…. Financial markets are driven by narratives. The narrative… today is… inflation and fiscal problems. But the narrative… is going to be… the success or failure of AI. And in both scenarios, long rates are going to be lower…

Give a gift subscription

What do I think of this?”

  • The “AI success” path is, if I am reading this correctly, the argument of last week’s: Caballero, Ricardo. 2026. “Speculative Growth and the AI ‘Bubble’”. MIT. August 23.<https://economics.mit.edu/sites/default/files/2026-08/speculative_growth_AI_public.pdf>. The AI-buildout switches to being financed by profits as they role into the labs and the hyperscalers, and the rise in productivity outruns consumption and increases savings.

  • The “AI failure” path is the standard bubble-collapse=and-aggregate-demand-driven-recession scenario.

What is very noteworthy is the claim that all of this is likely to come to a head in the next six months, and that the narratives surrounding it are overwhelming what Torsten Slok sees as the “inflation [risk] and [deficit] fiscal problems” that are the narratives currently dominating the financial market.

There is one enormous puzzle here: Why does Torsten Slok think that the AI question will be resolved in the next six months? I do not see the reason for thinking that at all.

Subscribe now

Leave a comment

If reading this gets you Value Above Replacement, then become a free subscriber to this newsletter. And forward it! And if your VAR from this newsletter is in the three digits or more each year, please become a paid subscriber! I am trying to make you readers—and myself—smarter. Please tell me if I succeed, or how I fail…

##⁠heading-for-a-large-rapid-fall-in-interest-rates-chart-of-the-day
##macro-outlook
##mamlms
#ai-bubble

##⁠chart-of-the-day
##‎
⁠torsten-slok-says-both-an-ai-boom-and-an-ai-bust-push-long-interest-rates-down-substantially-and-the-market-is-very-likely-to-decide-one-way-or-another-in-the-next-six-months
#‎
⁠heading-for-a-large-rapid-fall-in-interest-rates⁠
#torsten-slok
#market-narratives
#multiple-steady-states
#ricardo-caballero

CROSSPOST: GÉRARD ROLAND: The Berkeley years. Part XXII

Gérard Roland’s memoir of endowments, faculty retention wars, and the day the phone rang at 4 a.m.—plus my theory of why UC Berkeley administration tries to eat its healthiest limbs, and yet somehow it continues to do very well indeed:

It must have been back in 1999 or so, when we in the Berkeley economics department were looking to fill a field-hole in Comparative Economic Systems/Economics of Transition, when I asked my friend Andrei Shleifer whom we should try to hire. Andrei’s answer, as I remember it: “Gerard Roland. You at Berkeley, especially, should try to hire Gerard Roland. For industriousness, wisdom, knowledge, and collegiality, he is first-class. Gerard Roland. Definitely Gerard Roland.”

Share

Now Gerard has a SubStack. And he is telling his stories. Today: The Berkeley Years. Part XXII: Being Department Chair (2008–2011):

Give a gift subscription

CROSSPOST: GERARD ROLAND: The Berkeley years. Part XXII

<https://gerardroland.substack.com/p/the-berkeley-years-part-xxii> <https://gerardroland.substack.com/>

Gerard’s Substack
The Berkeley years. Part XXII.
I was asked by my colleagues to be department chair starting from July 1 2008 for the usual period of three years. I had been graduate chair for a few years before that and my colleagues had appreciated my work and initiatives, especially when it came to recruiting graduate students in competition with other great departments (Harvard, MIT, Stanford, Pr…
Read more
Being department chair (2008-2011)

Gerard Roland

Aug 28, 2026

I was asked by my colleagues to be department chair starting from July 1 2008 for the usual period of three years. I had been graduate chair for a few years before that and my colleagues had appreciated my work and initiatives, especially when it came to recruiting graduate students in competition with other great departments (Harvard, MIT, Stanford, Princeton, Chicago and others). People often have no idea how much time and effort people in top departments spend to recruit the best graduate students as well as the best junior and senior professors. We spend much more time than other departments on this, because success in these areas is fundamental to stay at the top.

My predecessor as Chair was Ben Hermalin, the well-known micro-economist. He had had a hard time, because there was the perception in the profession that UC Berkeley was financially less well off than other top universities since the many budget cuts to the University of California system, starting from the 1990s.[1] There was thus the rumor that the Berkeley economics department was “ripe for poaching”. During Ben’s mandate, there were at some point 13 outside offers for professors in the economics department. As is the case most of the time, Berkeley’s top administrators respond to those outside offers and manage to keep the concerned faculty on campus.

One problem is that for Berkeley faculty, getting outside offers had usually been the main way to increase one’s salary. This is somewhat of a double-edged sword. Too small salary raises, but high ones in response to outside offers, help the university to save money. On the other hand, this tends to reduce (but not always) the feeling of loyalty among the faculty towards their university. Ben managed to convince the campus authorities to put together a special one-time program for the economics faculty (called Targeted Decoupling Initiative or TDI) to respond positively to all the outside offers. It ended up being very successful. In my recollection, only Chang Tai-Hsieh decided to leave the department for Chicago’s Booth School of Business.

When accepting to be Department Chair from 2008 to 2011, I stated that I would spend a lot of time fund-raising for the department. This is normally something that department chairs do not do, but I thought that it was really necessary given that the department’s endowment had hardly changed in many years. I also wanted to do the job at 100% of my capacity to do it as well as possible. This implied working most evenings during the week, but I was prepared to do that, and my family was accepting this. I thought that working 100% of my time would help boost the morale of the department. Even though this left literally no time for research, I could count on the understanding of my Berkeley coauthors, with whom I would discuss next steps in our projects but leave them doing most of the footwork. Also, since coming to Berkeley, I had spent most summers in Europe with my family, but I thought I would not be able to manage well the department staff from afar, so I strongly cut the length of my summer trips to Europe during my three years as chair.

One month into the job, the stock market crashed. This was August 2008 and the beginning of what came to be known as The Great Recession. The value of the various endowments the department was managing, which I wanted to increase, went literally through the floor. Not a good way to start my mandate. It was necessary to respond without wasting time. One of my first initiatives in that context was to do something to protect the endowment income for David Card’s Center for Labor economics (CLE). I made a proposal to ask the campus to complement the Center’s income from campus money so that it would stay constant during the crisis. Those funds would then be gradually reimbursed as the endowment would rise again in the future. David Card was then the intellectual leader in the Department after Gérard Debreu in the 1980s and George Akerlof in the 1990s. He was one of the main figures in the economics profession to introduce the so-called “credibility revolution” in empirical research aiming at establishing causal relationships between variables (see the post on the credibility revolution in development economics in <https://gerardroland.substack.com/p/the-berkeley-years-part-xxa>) and was widely expected to receive the Nobel prize, which indeed happened in 2021. If the CLE’s income would be reduced as a consequence of the stock market crash, Card would understandably be upset. On the other hand, if I managed to preemptively protect the income of his research center, it would be good and probably unexpected good news for him. I managed to convince the provost, George Breslauer to support this. Breslauer was a political science expert on the Soviet Union with whom I had a very good relationship, as he played a key role in hiring me when he was Dean in 2000-2001, but as with all things administrative, it took a while for my initiative to get approved. The Dean of Social Sciences, Jon Gjerde, a reputed historian passed away unexpectedly after I had been in office for only 4 months. Jan de Vries, also a famous historian, took his position until the end of the academic year when a new Dean would be appointed. This was Carla Hesse, also a reputed scholar from the history Department.

Despite the lack of money due to the crisis, I took several other initiatives that played a role in keeping the Department’s morale. One was related to the absence of available slots for new professors. When a department must go a year without hiring, this is very bad for the morale. We decided to interview as usual at the American Economic Association Annual Congress early January 2009. We selected at least one strong candidate but had no slot. So, I proposed to the Department to make an offer with commitment to hire once we got the slot, using department money to make the difference so that there would be no negative consequences for the new hire. This worked and we got a slot from the campus after only a few months.

One reason why I think I was often able to find innovative solutions as Department Chair despite the lack of money was that I was comparing the situation at Berkeley with the one I had known at ULB in Brussels. Compared to ULB, Berkeley had so much more resources, both financial and in terms of very competent staff that it was a pleasure for me to work with those resources. Many other colleagues compared Berkeley’s situation with that of richer private universities like Harvard or Stanford, which was less favorable to Berkeley. I found throughout my career that quality of research does not necessarily increase linearly with funding. Two points can easily be made, but there are others. First, it is very difficult for university administrators to turn down financial requests for bad projects when the university is awash with money. Bad projects later weigh negatively on the university’s performance. Second, researchers who receive very large amounts of money need time to manage those resources, which takes time away from creative research and can be counterproductive.

Of course, at UC Berkeley more funds were badly needed after the crisis of 2008. The Dean’s office helped me a lot with the fund-raising and I spent a lot of time talking to potential donors who were nearly always very interesting people with very rich life experiences. On the whole, during my tenure as Chair, the Department and the various research centers associated to the Department received roughly 9 million dollars in endowment money. Some of those funds came very easily while other fund-raising efforts did not yield results, despite much time and effort. I am also very glad that after my mandate, my colleagues continued to engage in fund-raising. Despite all that, a strong campus support has always been critical for the Department. During my tenure, I spent more money to recruit and retain faculty than was possible only via the Department’s budget. Despite the economists generally not being very popular among other departments, both in social sciences and STEM departments, our very good reputation[2] as well as the very large number of economics undergraduate students always convinced campus authorities to help the department financially.

One of the great moments during my mandate as Chair was when Oliver Williamson got the Nobel prize. The Berkeley Public Relations department has always been very professional. I was asked a week before who might get the economics Nobel prize among my colleagues. For some reason, I thought that Oliver Williamson had a high chance (previous Berkeley economists who got the Nobel prize included Gérard Debreu, John Harsanyi, Dan Mc Fadden and George Akerlof). The prize was announced on Monday October 12 2009. On Friday October 9, I had preemptively rented a room for Monday 4pm at the Women’s Faculty club. On the day of the announcement, I got a call at 4am from Berkeley’s PR department that Williamson had gotten it together with Elinor Oström (the first woman and only political scientist to have received the economics Nobel prize). I immediately proudly announced the news to the department and announced the 4pm party. Some colleagues later asked me how I had been able to rent a room at 4 in the morning. The day was fabulous and the whole university was proud to have one more Nobel prize (By then, the university had had 21 Nobel prizes, including 5 for economics). This remains a great memory for me.

A more bittersweet moment for me was towards the end of my mandate. I had worked quite well with Dean Carla Hesse until then, but she told me that Heddy would from then on only be paid half-time instead of full-time but that those half-time funds would be secure. This was a big blow to both Heddy and me. I tried to explain to her that it would have been better to return to Heddy’s previous situation paid full-time from less secure funds, but she would not listen and even accused me of nepotism. I thought of resigning but only had little time left as Chair. Also, I had to have major colon surgery around that time and my daughter Elsa had been battling eye cancer (more on that next week). In any case, I was due for a sabbatical and decided to look for another job when I came back. Something in my loyalty to Berkeley had been broken.

(To be completed)

Me, around the time I was department Chair, shortly after an important surgery.


[1] After Clark Kerr’s reforms as president of the UC system, UC budgets were very large and the income from the California state was equivalent to that from the largest private university endowments, but that did not last, unfortunately.

[2] In a university like UC Berkeley, nearly all departments are in the top 3 or top 5 in the world, so the economics department is not more exceptional than other departments.

<https://gerardroland.substack.com/p/the-berkeley-years-part-xxii> <https://gerardroland.substack.com/>

Gerard’s Substack
The Berkeley years. Part XXII.
I was asked by my colleagues to be department chair starting from July 1 2008 for the usual period of three years. I had been graduate chair for a few years before that and my colleagues had appreciated my work and initiatives, especially when it came to recruiting graduate students in competition with other great departments (Harvard, MIT, Stanford, Pr…
Read more

Brad DeLong here: That is what being a department chair at UC Berkeley is like.

The most important thing I note from Gérard’s account is this:

  • As chair, he worked like a dog doing what was properly Dean Carla Hesse’s job, fundraising to try to boost Berkeley’s endowment so that we would not be at such a financial disadvantage in resources vis-à-vis our peer institutions.

  • He was remarkably and incredibly successful at this.

  • Dean Carla Hesse then responded to this success of his by financially injuring his family.

  • Carla Hesse’s claim that “Heddy would from then on only be paid half-time instead of full-time but that those half-time funds would be secure” was a more-or-less even trade was complete bullshit: at Berkeley, funding is never secure. Secure funding is not secure.

  • The Berkeley administration is in enormous debt to Gérard that it has taken few steps indeed to honor.

Share DeLong's Grasping Reality Weblog

When I was chair of Berkeley’s PEIS major, I evolved my own theory of Berkeley’s senior administrators. My theory was this:

  • They spent the first three days of each month trying to think rationally and seriously about the future of the university and about resource allocation.

  • On day four, they would have to respond to a faculty retention case in response to an outside offer launched by another university with a much larger endowment.

  • They would then spend the rest of the month turning every piece of Berkeley they could put their hands on upside down and shaking it, in the hopes that money that could be used to respond to the retention case would somehow fall out.

  • Pieces of Berkeley that were functioning well (as PEIS then was and now is) were seen as easy targets for this effort. They were doing well, right? Surely they could afford to limp along with somewhat fewer resources? Couldn’t they?

  • Pieces of Berkeley that were functioning badly were immune: we have enough problems and cannot risk creating more! Perhaps we should ease their resource constraint?

Get 75% off a group subscription

On the principal that The Purpose of a System Is What It Does, TPOASIWID, this is not a good way to run a railroad, or a university. And yet somehow we continue to do very very well.

Subscribe now

Leave a comment

If reading this gets you Value Above Replacement, then become a free subscriber to this newsletter. And forward it! And if your VAR from this newsletter is in the three digits or more each year, please become a paid subscriber! I am trying to make you readers—and myself—smarter. Please tell me if I succeed, or how I fail…

##crosspost-gerard-roland-the-berkeley-years-part-xxii
##public-reason
##crosspost
##gerards-subheadline-being-department-chair-20082011-fundraising-through-the-great-recession-financing-david-cards-labor-center-throwing-the-oliver-williamson-nobel-party
#gerard-roland-the-berkeley-years-part-xxii
#gerard-roland
#the-berkeley-years-part-xxii
#gerard-roland
#academia
#uc-berkeley
#berkeley-economics
#academic-administration
#higher-education
#tpoasiwid

The U.S. Federal Fiscal Flow Deficit Problem: CHART OF THE DAY

It is honest to say that America’s fiscal deficit problem has one cause: Republicans. But honesty, of course, is not something you get for free: it is a discipline you have to train for, which is why we get things like the passive voice in the title of this chart of the day:

A nice graph. I look at it, and say, why not just turn the tax-policy clock back to the Clinton-Gore 2000?

Share

Adam Tooze <https://adamtooze.substack.com/p/top-links-1207-us-deficits-counting> picks this up from the Financial Times:

Share DeLong’s Grasping Reality Weblog

The revenue drop from 2000 to 2004 is 4/5 structural, not cyclical <https://en.wikipedia.org/wiki/Bush_tax_cuts>. CBO scored the 2001 (EGTRRA) and 2003 (JGTRRA) cuts as adding ~$1.5 trillion to the debt over 2002–2011 excluding interest, and ~$3 trillion over 2010–2019 including interest if fully extended. This is not contested by anyone doing arithmetic. Talk I hear of a “cyclical peak in revenues in 2000” explains a year or two of mean-reversion, not a quarter-century structural gap. Using it to explain the past generation is a category error.

The upper-income Bush cuts were sold as a recession-recover measure. But they had a constant interest-rate policy fiscal multiplier of only 0.26. They were among the least stimulative and most expensive things the government did <​⁠https://www.epi.org/publication/ib338-fiscal-cliff-obstacle-course/>. And no American tax cuts ever “paid for themselves” except for tariff reductions in the long run. Paul Krugman’s line still holds: supply-side claims that tax cuts would pay for themselves “never got any traction in professional economic[s]… even among conservatives”. Those who claimed they would were and are unprofessional economists and professional recoveries.

Give a gift subscription


Thus my view: What was wrong with the tax rates and tax system as they stood in 2000?

Not very much.

Want to fix the deficit? Simply have the first item of business on January 21, 2029 be passing a Budget Resolution, to be followed on January 22, 2029 by a vote on the Reconciliation Bill restoring tax rates and taxable income coverage to what they were in 2000. Base-restoration matters as much as the headline rate. And also uncap FICA: impose the Social Security and Medicare taxes on all earned income. The 2026 wage cap is $184,500, so all earnings above it currently escape the 12.4% tax entirely. Removing the cap would raise on the order of $3 trillion over a decade <https://taxfoundation.org/blog/save-social-security-payroll-tax-cap-proposal/>.

Leave a comment

If you won’t get on board for that as your initial bargaining position, I have one question: what is wrong with you?

After all, year 2000 was the last time I had confidence that the US had a competent government and that US society was by-and-large working and improving. Shouldn’t we do whatever we can to restore things to how they were then?

Anyone else willing to join me on this turn-back-the-clock exercise?

Get 75% off a group subscription


The graph appears and there is commentary on it in an article by Chris Giles.

I thought about about dropping a link, but decided not to. Giles is behaving badly here. He asks the question: “Why [has] the US fiscal position has deteriorated so much this century?” And then he works much too hard to avoid giving the correct answer: Republicans.

Instead, he talks about “the cyclical peak in revenues” in 2000, the Bush tax cuts “which were made permanent on a mostly bipartisan basis during the Obama administration”, “spending on services for an ageing population—social security, Medicare and veterans’ programs”, in addition to the Trump tax cuts.

Not, mind you, that Giles is a big fan of Trump and Bessent.

And I have to admit there is motion: the Giles of today admits that “the US has significantly cut tax rates for the richest this century, so some reversal of those would be likely and justified from a left-leaning president and Congress”. This is not something that we would have gotten from the Piketty-bashing Giles I remember from 2014 <https://www.forbes.com/sites/scottwinship/2014/05/27/laffaire-piketty/>.

Refer a friend

But if you are not brave enough to tell the truth, to give the answer Republicans to the question why has the US fiscal position has deteriorated so much this century?, what use are you? Why would reading you be worth anyone’s time?

Herodotos 1:136 tells us that ἀληθίζεσθαι, truth-telling, was one of the only three things that young Persian aristocrats were taught. Herodotos matched ἀληθίζεσθαι with ἱππεύειν, to ride (to do the horse-thing), and τοξεύειν καὶ, to shoot (to do the bow-thing). Those are martial skills that need a huge amount of practice to attain and retain competence. Herodotos is right with the implication of coupling truth-telling with two martial skills. Honesty is more a discipline, something you train for, not a temperament. Analytic courage—a practiced refusal to hedge when the data points somewhere your political and ideological and careerist commitments point in a direction you would rather they not—requires practice as well.

It’s something to strive for.

There: that is off my chest.

Subscribe now

Leave a comment

If reading this gets you Value Above Replacement, then become a free subscriber to this newsletter. And forward it! And if your VAR from this newsletter is in the three digits or more each year, please become a paid subscriber! I am trying to make you readers—and myself—smarter. Please tell me if I succeed, or how I fail…

##the-us-federal-fiscal-flow-deficit-problem-chart-of-the-day
##macro-outlook
##chart-of-the-day
##republican-tax-cuts-for-the-rich-created-the-problem-please-name-and-shame-restoring-the-clintongore-tax-rates-and-tax-base-is-the-solution-pass-it-on
#the-us-federal-fiscal-flow-deficit-problem
#republican-tax-cuts-for-the-rich
#federal-budget
#clinton-gore-tax-base-and-rates-restoration
#uncap-fica
#name-and-shame
#ἀληθίζεσθαι

CROSSPOST: JUSTIN WOLFERS: Good News. There’s Already a U.S.-Canada Trade Deal

The tariffs are small. The damage to American credibility is not. As Richard Baldwin put it and I’ve argued repeatedly, Trump’s trade war is emotional restoration, not economic transformation — a muscular performance of dominance in which reindustrialization is decorative, not directive. IT IS KAYFABE.

And while Carney’s retaliation probably feels good, it costs Canadians, and playing the retaliation game winds up giving the advantage to whoever cares least about his own people, who is Trump. A sensible Canadian strategy would aim at other forms of linkage and leverage. The binding constraint on Trump is attention, not welfare. His principal purpose in any pronouncement is to gain eyeballs, so the only leverage that bites is leverage that threatens the spectacle — American farmers, retailers, and TikTokkers visibly hurt on camera. The correct Canadian move is to show that Trump has taken his own people hostage—that the logic is that of the sheriff in “Blazing Saddles”:

Share

Share DeLong's Grasping Reality Weblog


CROSSPOST: JUSTIN WOLFERS: Good News. There’s Already a U.S.-Canada Trade Deal

<https://newsletter.platypuseconomics.com/p/good-news-theres-already-a-us-canada> <https://newsletter.platypuseconomics.com/>

Platypus Economics
Good News. There’s Already a U.S.-Canada Trade Deal.
Earlier this week I joined Fergus Macphee on The Trump Report to talk about the U.S.-Canada trade war. We talked about how Canadians are getting caught in the middle of a negotiation between President Trump… and President Trump…
Read more
You’ll never guess who signed it.

Justin Wolfers

Aug 27, 2026

Earlier this week I joined Fergus Macphee on The Trump Report to talk about the U.S.-Canada trade war. We talked about how Canadians are getting caught in the middle of a negotiation between President Trump… and President Trump.

A little history: The United States has had a free trade agreement with Canada since 1989. That U.S.-Canada Free Trade Agreement gave way to the North American Free Trade Agreement (NAFTA) when Mexico joined the club in 1994.

Then the first President Trump came to office calling NAFTA a disaster. He renegotiated it and replaced it with the United States–Mexico–Canada Agreement, or USMCA to its friends. The new deal was signed in 2018 and took effect in 2020. It’s basically NAFTA with a new title.

President Trump called that agreement:

“the largest, fairest, most balanced, and modern trade agreement ever achieved. There’s never been anything like it.”

That agreement passed the U.S. Congress, the Canadian Parliament, and the Mexican Senate, and was signed by the leaders of all three countries. It’s the law of the land in all three countries.

But here we are again. And my head hurts.

Talks collapsed and the trade war started

The second President Trump thought the deal that the first President Trump negotiated was terrible. He was unsparing in his criticism, saying:

“I would rather not have the agreement… We do better as a country if we don’t have an agreement.

So earlier this year Trump demanded a new deal, and threatened large tariffs if he didn’t get his way. Talks began, then collapsed, and President Trump imposed tariffs of 50 percent (that’s high!) on a weird-looking list of Canadian exports, including hockey sticks.

There are two ways to see these tariffs. (Warning: Two-handed economist in the house.) On the one hand, these tariffs apply only to a relatively narrow slice of goods — around $20 billion worth of Canadian exports. That’s only about 4 percent of Canada’s exports, so not really that big of a deal. On the other hand, these tariffs apply to goods that should be tariff-free under the existing trade agreement. Previously these goods had been exempt from an array of Trump’s tariffs. This new move suggests a willingness to significantly broaden the trade war.

More here:

Why Trump’s New Tariff Is Bigger Than Canada

Why Trump’s New Tariff Is Bigger Than Canada

Justin Wolfers

Jul 21

It’s retaliation all the way down

We recorded this podcast before retaliation gave way to retaliation for retaliation, then retaliation for retaliating against retaliation. So, a quick update: We’re in a retaliation cycle.

We’ve seen this movie before. In 2025, Trump got the United States into a tariff escalation cycle with China, and tariffs rose and were countered and counters were countered, and after a bit the United States was imposing a 145 percent tariff, while China was at 125 percent. Both numbers are untenable — high enough you might as well call it an embargo — and a month later, both sides agreed to mostly back down.

We may see this play out again with Canada — escalate until economic pain generates political pain, strike a “deal,” declare victory, and hopefully get back to where we began.

Reminder: We started with a free trade agreement with Canada negotiated by Trump.

Share

The real costs come in the longer run

In our economics textbooks, the case against tariffs is usually that they impose immediate costs, usually on domestic consumers and companies that use imported inputs.

But Trump has found a way to make tariffs even more damaging to Americans: by making them unpredictable. He has oscillated — at seemingly random moments — between treating Canada as a friend and an enemy. He never even pauses at frenemy.

Eventually the United States may discover that it already has a trade deal with Canada. Maybe we’ll rebrand it (again!) and declare victory. But the history of ripping up past deals really matters. It is now understood that whatever is written down is only binding until the president decides it isn’t.

Canadian Prime Minister Mark Carney has noticed, saying that “America has changed,” and its “signature was written in pencil.” I noticed too, and wrote about it recently:

Would You Marry a Man With Seven Ex-Wives?

Would You Marry a Man With Seven Ex-Wives?

Justin Wolfers

Aug 21

The problem is that this sort of on-again off-again relationship prevents either side from making long-term investments in their shared future. Even if we get back to trading with Canada, we’ll get less investment, less integration, and fewer of the gains that trade usually brings.

Should Canada retaliate?

We know what happened next: When Trump imposed tariffs, Canada retaliated.

I don’t think this is in Canada’s best interest.

The intuitive argument for retaliation says that Trump’s tariffs on Canada hurt Canadians, and so Canada should retaliate with tariffs on the United States that hurt Americans. A credible threat of retaliation might also deter future tariffs.

But that misses half the story.

The other half is that Trump’s tariffs also hurt Americans, and so the same logic says that Carney’s tariffs also hurt Canadians. That’s the argument for not retaliating.

This point was more vividly made by Joan Robinson, who was possibly the greatest economist never to win a Nobel. She wrote that answering foreign tariffs with tariffs of your own makes no sense: “it would be just as sensible to drop rocks into our harbours because other nations have rocky coasts.” (She credited William Beveridge with the metaphor.) The point is that tariffs make it harder and more expensive for ships full of goods to reach your ports. And who wants less efficient ports? Other countries’ inefficient ports are no reason to ruin your own.

This reality also makes Canadian threats less credible. The argument I hear most often from my Canadian friends is that they don’t want to be bullied. But if tariffs require hurting your own people as a necessary side effect of inflicting pain on foreigners, then a war of economic attrition rewards the leader most indifferent to the pain they inflict on their constituents.

I’m not sure that’s Mark Carney’s strong suit. And Canadians shouldn’t want it to be.

Can we (please) do this the easy way?

A good leader’s job is to make it easy for the other side to say yes.

Trump is doing the opposite. His childish taunts — that Canada is the 51st state, that its prime minister is really “Governor Carney,” that “Canada is nasty,” that its leaders are “clowns” who should “fall in line,” that Lake Ontario should become “Lake America,” and on and on — have stirred such anger among Canadians that Carney now has less room to give Trump what he wants.

Point is, Trump’s taunts unleashed a political force that all but forced Carney to retaliate. He has made it harder for Canada’s leaders to make concessions to the United States.

This is how you turn an easy deal into a hard one.

And we’re all paying for it.

<https://newsletter.platypuseconomics.com/p/good-news-theres-already-a-us-canada> <https://newsletter.platypuseconomics.com/>

Platypus Economics
Good News. There’s Already a U.S.-Canada Trade Deal.
Earlier this week I joined Fergus Macphee on The Trump Report to talk about the U.S.-Canada trade war. We talked about how Canadians are getting caught in the middle of a negotiation between President Trump… and President Trump…
Read more

Brad DeLong here: The tl;dr version of Justin is this, I think:

The binding legal agreement already exists; the conflict is entirely about one leader repudiating his own prior work, and the deepest damage isn’t the immediate tariff cost but the destruction of the credibility that makes any future deal—or long-term investment—worth making. Trump does not keep his own deals, therefore you cannot make deals with Trump. All you can do is set up situations in which Trump’s actions adverse to your well-being are especially panful to him.

  • → Trump renegotiated NAFTA into USMCA (2018/2020).

  • → Trump called it the best deal ever.

  • → he now rejects it, demands a new deal, and imposes 50% tariffs on ~$20B of Canadian goods that should be tariff-exempt.

  • → Canada retaliates.

  • → An escalating retaliation cycle ensues (echoing the 2025 U.S.-China spiral to 145%/125%).

  • → But because tariffs hurt the imposing country’s own consumers, retaliation is self-harm, so attrition “rewards the leader most indifferent to the pain they inflict on their constituents”

  • → Meanwhile Trump’s personal taunts inflame Canadian public opinion, shrinking Carney’s political room to concede

  • → An easy deal becomes a hard one, and unpredictability itself becomes the lasting cost by signaling that U.S. commitments are “written in pencil.”

  • → It should resolve anticlimactically: Trump gains nothing from this, provided there are enough people on TV and TikTok who have been hurt by Trump’s tariffs. But it might not. It is not quite TACO: Trump does not quite always chicken out.

Give a gift subscription

It is clear that tariff retaliation is a mistake for Canada: it is shooting a shotgun downward at your feet and at your counterparty’s feet. But mostly at your feet. The key move for Canada is to find something unrelated to tariffs that Trump wants, and hold it hostage. That most likely means finding a way to mobilize those on the American side of the border who lose big because they no longer have cheap Canadian goods.

At a deeper level, the problem is this: the Republican majorities in Congress abdicated and the courts play CalvinBall. Tariff power belongs to Congress, but the Roberts majority applies major-questions and non-delegation doctrine at its whim — constraining Democrats reliably and Republicans only when convenient. Moreover, legal victory two years from now doesn’t unwind the damage done in the interim. “Trump will lose in court” is beside the point. The tariffs get collected while litigation crawls; refunds are partial, burdensome, and uncertain — so the escrow-account friction itself does the work of discouraging trade regardless of the eventual ruling. The only durable strategy against a serial defector is to structure the game so that his defections hurt him automatically — which is a grim thing to have to say about the United States of America.

Subscribe now

Leave a comment

If reading this gets you Value Above Replacement, then become a free subscriber to this newsletter. And forward it! And if your VAR from this newsletter is in the three digits or more each year, please become a paid subscriber! I am trying to make you readers—and myself—smarter. Please tell me if I succeed, or how I fail…

##crosspost-justin-wolfers-good-news-theres-already-a-us-canada-trade-deal
##neofascism
##crosspost
##justins-subheadline-youll-never-guess-who-signed-it-and-now-trump-wants-to-burn-it-and-reinforce-everyones-knowledge-that-his-word-is-simply-not-good-for-anything
#justin-wolfers-good-news-theres-already-a-us-canada-trade-deal
#justin-wolfers
#good-news-theres-already-a-us-canada-trade-deal
#chao-monkey
#us-canada-trade
#trump-tariffs
#usmca
#retaliation-cycle

NVIDIA Earnings: CHART OF THE DAY

A productive bubble in the Perez–Janeway sense: the racks will survive even if the shareholders don’t. Whether they’ll be worth as much as the railroad track is what we cannot yet see.

Perhaps he real news isn’t the revenue guide — it’s the half-trillion dollars of third-party capital that NVIDIA is trying to start herding onto its customers’ liability sheets. In economiss’ economic-welfare terms, good natural-language interfaces are already a huge boon; in measured GDP terms, AI i not and may never be there. Don’t let NVIDIA’s income statement stand in for social value in either direction—and don’t let economists’ willingness-to-pay-based economic-welfare measures stand in for human well-being in an environment in which attention-hacking is rife.

We have:

Share

With commentary:

Lynn Thomason: Nvidia’s $100 Billion Shows Why AI Is Still the Trade <https://www.bloomberg.com/news/newsletters/2026-08-27/nvidia-s-100-billion-shows-why-ai-is-still-the-trade>: ‘Nvidia delivers. The AI lynchpin blew away expectations, predicting 70% sales growth for 2028, when analysts had anticipated 45%. The stock jumps 7%…. Nvidia’s stellar results are reinforcing investor confidence in the AI revolution, with CEO Jensen Huang saying the “golden age” of labs and startups is here…. Sales surged. It sees revenue in the current period at $108 billion. Compared with a year ago, that’s close to double…. There’s no let up in demand. Nvidia said it would be growing even faster if it had access to more supplies…. Hyperscalers spend big. The data center division beat estimates, with firms like Alphabet and Amazon accounting for much of those sales…. Margins will narrow. Nvidia warned that margins would shrink slightly in the coming months as it copes with a surge in memory costs. At the current price, Nvidia is poised to add about $370 billion of market value when trading begins. The shares are up 12% this year…

Share DeLong's Grasping Reality Weblog

And just two days ago the word on Bloomberg was that “NVIDIA has lost some luster with investors this year”:

Give a gift subscription

<https://braddelong.substack.com/p/nvidia-has-a-seven-day-stock-price>

This whipsawing of what Paul Krugman calls not Economics but Upanddownonomics is really not terribly helpful.

Nvidia’s profits are a combination of five things:

  1. The strength of demand for its GPUs.

  2. Its desire and ability to use its monopoly power to generate super-high margins.

  3. The limitation of that by its not wishing to trigger its customers to make investments in ending their dependence on it and exiting from its ecosystem.

  4. Upstream monopoly power by its own major suppliers.

  5. Upstream supply chain difficulties.

Get 75% off a group subscription

What we want to know is: (a) How strong is demand for the GPUs? (b) Hopw much of that demand is really based on “fundamentals”? (c) What are those fundamentals based on—(i) utility to customers, (ii) moat-maintaining investments by platform-monopolists, (iii) belief that you have to learn-by-doing with these technologies on a large scale even though the utility to customers is not yet there because it will be there, (iv) “speculation”, and (v) that these days one grifts gullible investors not by starting a memecoin but by pointing to your racks of GPUs? That requires a careful parsing of (1). And changing your mind on those things as a result of the high-energy excitation mode that is the (1) through (5) factors producing NVIDIA quarterly earnings—that does not strike me as wise.

Bloomberg’s Ian King peers a little bit through the veil of time and ignorance here. Although I do not see NVIDIA’s willingness to go all-in on next year’s growth should ease anyone’s concerns:

Ian King: Nvidia Sees AI-Fueled Demand Boosting Sales 70% Next Year <https://www.bloomberg.com/news/articles/2026-08-26/nvidia-estimate-topping-forecast-fails-to-wow-investors>: ‘Nvidia… said revenue will grow about 70% next fiscal year, easing concerns that AI spending is poised to lose momentum…. Jensen Huang said demand is only accelerating…. “The AI infrastructure build-out is at full steam,” he said. “Vera Rubin, now in full production, was built to power exactly this moment.” Revenue in the current period will be $108 billion, plus or minus 2%…. Gross margin, the percentage of sales remaining after deducting the cost of production, will be roughly 74% in the quarter…. Nvidia [is] cop[ing]… with a surge in memory costs… expects the [gross margin] measure to bottom out in the fiscal fourth quarter… that runs through January — at 71% to 72%….

“We would love more supply,” Kress said in an interview. “It’s really about how much more could you do?” In the second quarter, which ended July 26, sales more than doubled from a year earlier to $96.2 billion…. Nvidia has now delivered sales above Wall Street estimates for 16 quarters in a row….

Nvidia has spent much of the past year lining up investment deals…including both developers of software and infrastructure… that… have, in theory, put Nvidia on the hook for tens of billions of dollars of liabilities…. Kress described the deals as a part of Nvidia’s efforts to increase supply. The biggest portion of its spending commitments took the form of long-term purchase agreements with vendors, she said. “The majority of what we have been helping folks with in terms of commitments is one very important thing called supply”…

Refer a friend

“We would love more supply” is doing enormous rhetorical work here. Kress framing the deals as securing supply, not stimulating demand, is the crucial move — because it reclassifies factor (5), supply-chain difficulty, as bullish rather than as a ceiling. I’d want to see more about that, for whether and to what degree this is, you know, actually true is something I do not but would very much like to know.

I would like to see more about that especially because it really does not fit with the real news, which is the $500B “compute financing platforms” with Apollo, BlackRock, Blackstone, Brookfield, Goldman, and KKR. The chipmaker is organizing half a trillion dollars of third-party capital not to induce its suppliers to expand their on capacity, but to herd in people willing to bear downside risk so that its customers will be willing to buy its chips. That is demand being manufactured by financial engineering, rather than merely being met. Is this a bad thing for NVIDIA to do from the point of view of NVIDIA’s financials? Of course not! It is a very profitable thing for NVIDIA to do! What does it mean for the world economy as a whole? That also is something I do not know and would very much like to.

GM financing a car works because the car has known use-value to a buyer who will pay. The AI analog only holds if the downstream token-buyers eventually cover their compute bills. Right now the poster-child labs mostly don’t—so the vendor-finance defense is a bet on future monetization, not a description of present value. And circularity concentrates systemic risk precisely because the buyer set is thin, and the same dollars show up as payments to NVIDIA by a customer and equity investments by NVIDIA in downstream LLM-service provider. NVIDIA could, in the bad scenarios, lose twice here: once when a customer stops buying, and again when its equity stake in that customer craters. It is a correlated exposure that works equally well going up or down. Dressed as an increasing-speed flywheel.

And, of course, concentration has made this a macro question, not a tech-sector question. Nvidia is now ~8% of the S&P 500 and has supplied more than 10 of the ~84 percentage points of the index’s five-year return. The possibility of a genuine turn in AI capex cannot be contained inside semiconductors — it becomes everyone’s problem, including every teacher’s pension.

But step back: This is Carlota Perez and Bill Janeway territory: a “productive bubble.” Even if most investors lose their shirts, no one will tear up the GPUs, the fiber, or the substations — just as no one tore up the railroad track or the dark fiber. The infrastructure persists; the equity holders are the potential sacrificial capital. Will these datacenters be as useful in the end as the railroad tracks or the dark fiber? That depends on things we cannot now see.

And do beware the Solow-paradox trap in reverse. In economists’ economic-welfare terms — consumer surplus from people willing to pay for natural-language interfaces — AI is already a huge boon. In measured GDP-and-profit terms, it may show up modestly, or even as a productivity decline, because now Amazon needs the warehouse, the truck, and the GPU farm to sell you the same item. Don’t let NVIDIA’s income statement stand in for social value in either direction. And, in addition, do not let economists’ economic-welfare measures stand in for human well-being. Maximizing dopamine hits do not make for a fulfilling, or even a tolerably happy and engaged life, especially when they are the result of algorithmic attention-hacking programmed up by some of the most cynical people who have ever lived.

The honest posture is neither bull nor bear but high-dimensional agnosticism. I count at least a dozen distinct vectors here — grifters, defensive platform monopolists, socially-valuable-but-privately-unprofitable overbuilding, techno-millenarians, attention-extraction business models, and genuine general-purpose-technology diffusion — all superimposed. Anyone claiming this quarter resolves them into a single verdict is selling you a narrative, and the future here is one I frankly cannot see.

I see people confidently claiming that:

  • A megawatt of AI-capacity costs $15 million to run, and right now you can sell it for $50 million because it can produce $100 million in cost savings or revenue boosts for high-value customers who know how to harness it.

  • Anthropic and OpenAI are each scaling-up from 1.5 gigawatts of capacity at the end of 2025 to 5 gigawatts of capacity at the start of 2027: that is a potential revenue boost to their annualized revenue run rates from $20 billion (at $13M/mW) at the end of 2025 to $150 billion (at $30M/mW) should they find and maintain product-market fit.

And I see other claiming that datacenters will never be profitable once the huge demand from people who think they have to build the institutional muscle-memory capacity to utilize these technologies drops away.

We have to do our reasoning under conditions of genuine Knightian uncertainty. That means: decompose, refuse false precision, resist single-number verdicts. It is more important than ever for us to inoculate sreaders against “Upanddownonomics,” the sentiment-laundering that lets the same fact-situation mean opposite things 48 hours apart. And we need to flag where fragility actually sits: circular finance, thin buyer set, index concentration, off-balance-sheet commitments. Robustness and scenaior planning is only possible if we distinguish between watching the right and wrong dials.

Subscribe now

Leave a comment

If reading this gets you Value Above Replacement, then become a free subscriber to this newsletter. And forward it! And if your VAR from this newsletter is in the three digits or more each year, please become a paid subscriber! I am trying to make you readers—and myself—smarter. Please tell me if I succeed, or how I fail…

##⁠nvidia-earnings-chart-of-the-day
##⁠chart-of-the-day
##macro-outlook
##mamlms
##
‎nvidias-96-billion-quarter-and-the-twelve-things-it-doesnt-settle-one-of-which-is-is-nvidia-meeting-demand-or-manufacturing-it
#‎
⁠nvidia-earnings⁠
#productive-bubble
#circular-vendor-financing
#ai-capex
#upanddownonomics

CROSSPOST: DAN DREZNER: Categories of Contempt: A Typology of my Domestic Political Disgust

To think, somewhere out there in a surprisingly nearby timeline, Dan Drezner is a card-carrying, ticket-punching establishment Republican, stroking his chin and still dismissing Paul Krugman as intelligent but overwrought, unbalanced, and shrill. I really wish we lived in that timeline:

Drezner explains why his harshest criticism of the Trump administration targets the “normie” Republicans—of whom he was once one—rather than the obvious buffoons: the people who know better and enable Trump anyway are more contemptible than those who never had principles to betray. In his view, contempt should be allotted in proportion to culpability, not incompetence. The intelligent enablers who understand exactly what they’re empowering — and choose it for power and status — deserve more scorn than manifestly unqualified and honorless clowns from whom nothing could ever have been expected.

There is, somewhere in the branching manifold of possible worlds, a nearby timeline in which Dan Drezner is one of them: a card-carrying, ticket-punching, chin-stroking establishment Republican, writing measured essays about the liberal international order, and referring to the likes of Paul Krugman and me as “intelligent, but overwrought, unbalanced, and, well, shrill.” In that timeline the Drezner tut-tuts at the overexcited. He counsels patience. He is very sound.

We do not live in that timeline. God, I wish we did!

Drezner has documented and continues to document, with the patience of a Jobite naturalist cataloguing a new and horrible phylum, the American foreign policy establishment run by — I quote his field notes — “the dumbest motherfuckers alive.” I am impressed.

Ph’nglui mglw’nafh Daniel Drezner R’lyeh wgah’nagl fhtagn.

Today, Drezner has these key ideas:

  • A typology of complicity: Drezner sorts Trump’s orbit into MAGA faithful, “Stormtroopers” (petty government thugs), Trump emulators, and normie Republicans — escalating by how much they should know better.

  • Stormtroopers are endemic, not novel: abuse of surveillance power by career officials spans administrations (the Wired CBP story covers 2009–2022), so Trump aggravates but did not invent it.

  • Normalization of scandal: Trump has made misogyny and extremism a “dog-bites-man” story, letting candidates with disqualifying scandals (Miller, Herrera) survive where earlier norms would have forced them out.

  • The pivotal enablers act on choice, not necessity: figures like Mike Johnson and Bill Cassidy “have a choice” and keep choosing to empower the party’s worst elements — believing they’ve only “rented” their souls temporarily.

  • G. Elliott Morris’s finding: subgroups retaining “Republican” identity stay Republican in the generic ballot by 37+ points despite disapproving of Trump, while subgroups without that identity defect to Democrats by 30+ points.

But the most effective mode for this type of discourse is, in its most classic form, that of the Ancient & Hermetic Order of the Shrill located in Arkham, Massachusetts:

No, Paul Krugman never, at least not to me knowledge, datelined anything he wrote for the “New York Times” from Arkham, MA—home of Miskatonic University of many Lovecraft stories and Arkham Asylum of many “Batman” stories. But Google is now certain that it did. Thus do AI-hallucinations come for it…

<https://archive.nytimes.com/krugman.blogs.nytimes.com/2012/02/29/looking-back-with-shrillness/>


CROSSPOST: DAN DREZNER: Categories of Contempt: A Typology of my Domestic Political Disgust

<https://danieldrezner.substack.com/p/categories-of-contempt> <https://danieldrezner.substack.com/>

Drezner’s World
Categories of Contempt
The hard-working readers here at Drezner’s World have probably noticed that in my commentary about the Trump administration, a disproportionate share of my ire is targeted towards those who are perceived as the more “normie” Republicans: folks like Scott Bessent…
Read more
A typology of my domestic political disgust.

Daniel W. Drezner

Aug 19, 2026

The hard-working readers here at Drezner’s World have probably noticed that in my commentary about the Trump administration, a disproportionate share of my ire is targeted towards those who are perceived as the more “normie” Republicans: folks like Scott Bessent or Elbridge Colby or Marco Rubio. One can ask: why not aim more at someone who is manifestly unqualified for their job — someone like, say, Pete Hegseth?

I have a short answer and a long answer. to this question.

Drezner’s World is a reader-supported publication. To receive new posts and support my work, consider becoming a free or paid subscriber.

The short answer is simple: it is more troubling to see folks with some degree of acumen willingly ben the knee to Trump’s whims just to have some measure of power and status. Pete Hegseth is a malevolent, buffoonish clown; I expect absolutely nothing from him and am therefore unsurprised that he’s acting like a clown who is spectacularly out of his depth. Once that recognition is made, there is really little point to making it again.

The likes of Rubio, on the other hand, are rather different. He is clearly intelligent enough to know what (and who) he is enabling and does it anyway. To the hard-working staff here at Drezner’s World, that is the more contemptuous behavior, and therefore merits more comment and derision.

The longer answer, however, is that Trump’s myriad transgressions require an entirely new typology of contemptible individuals — and it is sometimes difficult to keep track of all the different subcultures of sycophancy.

For example, what should I make of the MAGA faithful, the likes of Tucker Carlson or Steve Bannon or Marjorie Taylor Greene or Nancy Mace, folks who actually believed the MAGA bullshit that Trump and his team have spewed out over the past decade? They’re all genuine bigots, which is certainly contemptible. At the same time, they also genuinely believe in at least some of the foreign policy ideas that Trump embraced over a decade ago. I can’t get exercised by them all too much — mostly because, in the end, I think they’re a marginalized group with minimal political power.

Then there are the Stormtroopers — the folks working in the federal government who have enthusiastically embraced what Trump is selling. You know these are the weak-minded guys who Obi-Wan handles in this scene:1

There have been a raft of recent stories about the degree to which low-level thugs have taken advantage of their government positions to revel in their ability to bully and surveil citizens. For example, Wired’s Yulia Almazova recently reported on what U.S. Customs and Border Patrol officers have been doing on their computers: “In one case, a CBP officer allegedly used government databases to contact a flight attendant. In another, an officer was accused of pulling information from trusted-traveler applications to ask people out. Other CBP employees were accused of providing border-crossing data to someone involved in a ‘heated divorce.’”

Another example is DHS more generally. In the wake of organized resistance to their Minneapolis incursion, DHS decided to investigate progressive organizations. According to the New York Times’ Alan Feuer and Ernesto Londoño:

Homeland security officials used an array of invasive tactics during the first half of this year to gather information on many groups and individuals who were never accused of crimes, crossing the line that has traditionally stood between investigating criminal activity and political dissent.

In one instance, officials used administrative subpoenas to obtain more than three years of financial records from the Sunrise Movement, an environmental action group, and a labor union, the Communications Workers of America. That time frame went well beyond the civil unrest in Minnesota, which was prompted by the deployment of thousands of immigration agents to the state during the winter.

In another, investigators scrutinized three years’ worth of wire transfers made by the nation’s biggest health care workers union, the Service Employees International Union, in what they referred to as an inquiry into “domestic terrorist financing”…

All of this is pretty vile, and the folks doing it at DHS appear not to care about how vile it is.

And yet, it is also worth noting that abuses of government power by career employees span most administrations. That Wired story, for example, examined CBP actions from 2009 to 2022. Stormtroopers are an endemic problem. Clearly, the Trump administration is not helping, but this is not a new phenomenon.

Then there are the Trump emulators, the aspiring politicians who see Donald Trump as their political lodestar and think that behaving like him, they will acquire power. The New York Times’ Michelle Goldberg wrote an excellent column about the likes of U.S. Representatives Max Miller and Cory Mills, Republican House candidate Brandon Herrera, and other kindred spirits. As Goldberg notes, “Thanks to President Trump, misogyny and extremism among Republicans have become what journalists sometimes call a dog-bites-man story, something so ordinary and predictable that it breaks through only when it reaches an absurdly high threshold.” It is actually worse than that — in Miller’s case, the Trump White House overtly intervened to tell Republicans to back off their criticisms of Miller.

These folks are infuriating, and a sign of how Trump has normalized the idea of GOP candidates running despite serious scandals that would have forced a previous generation of candidates to drop out.

And yet, to the hardworking staff here at Drezner’s World, the group that is even more infuriating are the allegedly normie Republicans. These are the folks who have held their noses and continued to vote Republican despite all of… this… for the past ten years.

For example, Brandon Herrera has made remarks that are clearly sexist and anti-Semitic. Despite all of this, House Speaker Mike Johnson is attending Herrera fundraisers. And as Goldberg notes in her column, “I haven’t seen much public hand-wringing about the decision of the pious Johnson to raise money for Herrera, nor debate about what his candidacy means for the Republican Party’s attitude toward women, Jews or seniors.”

Or consider this paragraph from Politico about GOP thinking regarding Max Miller: “Some Republicans argue they have no choice but to close ranks around a candidate accused by his former father-in-law, Ohio GOP Sen. Bernie Moreno, of being unfit to serve and needing psychological help; Ohio’s other Republican senator, Jon Husted, has also said Miller should resign and not run for reelection. Miller has remained defiant amid the calls to resign.” My point is that they absolutely have a choice.

Furthermore, they keep making their choice in a way that empowers the worst elements of their party. G. Elliott Morris recently explored the gap between Trump’s high disapproval polling and the Democrats smaller margin in the generic congressional ballot. His conclusion: “Voters who pulled the lever for Trump in ‘24 and call themselves Republicans but not conservatives stay Republican…. The pattern is hard to miss. The subgroups that keep “Republican” in their identity stay Republican by 37 points or more in the generic ballot even while disapproving of the man leading their party, — while both subgroups without it defect to the Democrats by a margin of 30-plus points. Party identity is doing a lot of work, even among Trump voters.”

Here’s a helpful chart:

It’s this group that, in the end, are the pivotal Trump enablers. Senators like Chuck Grassley or, even better Bill Cassidy.

When I see someone like Brandon Herrera or Max Miller pop up in the news, it does not take long to conclude that there’s no point condemning someone who has already lost their soul. Mike Johnson and Bill Cassidy, however, likely think that they have only rented theirss temporarily.

These are the folks for whom calumny has been sparser — but dear God do they deserve it.

<https://danieldrezner.substack.com/p/categories-of-contempt> <https://danieldrezner.substack.com/>

Drezner’s World
Categories of Contempt
The hard-working readers here at Drezner’s World have probably noticed that in my commentary about the Trump administration, a disproportionate share of my ire is targeted towards those who are perceived as the more “normie” Republicans: folks like Scott Bessent…
Read more


Brad DeLong here: Drezner’s argument runs:

  1. buffoons like Hegseth act badly but predictably,

  2. so condemning them yields diminishing returns;

  3. capable actors like Rubio possess the acumen to recognize what they’re enabling,

  4. making their compliance a willful moral choice rather than a limitation;

  5. therefore the marginal moral outrage should flow toward those with agency who choose complicity;

  6. the pivotal such group is the “normie” Republican voters and senators whose durable partisan identity — not approval of Trump — keeps the coalition electorally viable.

The causal punchline: party identity, not enthusiasm, supplies Trump his governing majorities, so the quiet enablers are the true load-bearing structure. This is the core vibe of Dan’s current position.

When I see someone like Brandon Herrera or Max Miller pop up in the news, it does not take long to conclude that there’s no point condemning someone who has already lost their soul. Mike Johnson and Bill Cassidy, however, likely think that they have only rented theirs temporarily.

I do recall 24 years ago when the George W. Bush administration took office and promptly began to dismantle as much of successful Clinton-era policies as it could: policies aimed at keeping a prosperous and peaceful post-Cold War multilateral world, policies aimed at boosting economic growth, policies aimed at stabilizing federal finances, policies aimed at rolling back the rise in income and wealth inequality. They were remarkably successful at doing so, I must say. And those who we Rubin Democrats had thought were our allies in sensible neoliberalism were—whatever were their private misgivings under Chatham House rules— publicly cheering them on or at least immediately moving to distract from policy substance by expressing their shame at how the democrats had chosen a standard bearer bill. Clinton who had a huge problem with his zipper. (Yes, I am looking at you, Alan Greenspan, Glenn Hubbard, Colin Powell, and company: look in the mirror, if you dare.)

Read more

(VERY PARTIAL-)CROSSPOST: Samuel Bowles & Herbert Gintis (2002): The Inheritance of Inequality

Being one of Alan Krueger’s assistant editors when he was head honcho of the Journal of Economic Perspectives was a great joy and privilege. And this is perhaps the peak of what the JEP could do and be back in our era:

To a substantial degree, America has not been a “land of equal opportunity” with each generation’s economic fate largely self-made for quite a while.

The old consensus said a father’s economic advantage all but vanished in three generations. But Bowles and Gintis summarize the line of work that established that that belief was a measurement error-caused statistical illusion: the true intergenerational persistence of income is roughly three times higher than the old consensus. Becker and Tomes’s (1986) claim that the father-son income correlation was 0.15 is simply wrong. Think, instead: an intergenerational correlation of 0.4, and intergenerational elasticities of 0.7 for consumption, 0.5 for wealth, 0.4 for income, 0.35 for earnings, and 0.3 for school years.

IMPORTANT!: The belief is false that “smarts” and the genetic transmission thereof, at least as measured by IQ, is key to the intergenerational transmission of income inequality. Thus the argument that inequality is not a problem because the smart deserve to be rich does not fly. It is parental wealth, race, and “noncognitive personality traits” that do most of the work here. But the intergenerational transmission of economic status remains “a black box”: as the standard human-capital smart parents → smart, well-schooled kids → high earnings accounts for at most three-fifths of it, and the genetic inheritance of IQ accounts for almost none.

This line of research has always been very bad news for the caliper-measurers:

The results are somewhat surprising: wealth, race and schooling are important to the inheritance of economic status, but IQ is not a major contributor, and, as we have seen above, the genetic transmission of IQ is even less important.

Share

Yet they still bring out the calipers at every opportunity.


(VERY PARTIAL-)CROSSPOST: Samuel Bowles & Herbert Gintis (2002): The Inheritance of Inequality

<https://pubs.aeaweb.org/doi/pdfplus/10.1257/089533002760278686>

Journal of Economic Perspectives—Volume 16, Number 3—Summer 2002—Pages 3–30
See Bowles and Gintis (2001) for the relevant formal models and other technical aspects of this research, also available at <http://www.santafe.edu/sfi/publications/working-papers.html>. Arrow, Bowles and Durlauf (1999) and Bowles, Gintis and Osborne (forthcoming) present collections of recent empirical and theoretical research.
Samuel Bowles is Professor of Economics at the University of Siena, Siena, Italy, and Director of the Economics Program, Santa Fe Institute, Santa Fe, New Mexico. Herbert Gintis is a member of the External Faculty, Santa Fe Institute, Santa Fe, New Mexico. Both authors are Emeritus Professors of Economics, University of Massachusetts, Amherst, Massachusetts. Their e-mail addresses are bowles@santafe.edu and hgintis@attbi.com, and their websites are (http://www-unix.oit.umass.edu/bowles and http://wwwunix.oit.umass.edu/gintis).

People differ markedly in their views concerning the appropriate role of government in reducing economic inequality. Self-interest and differences in values explain part of the conflict over redistribution. But by far the most important fault line is that people hold different beliefs about why the rich are rich and the poor are poor. Survey data show that people—rich and poor alike—who think that “getting ahead and succeeding in life” depends on “hard work” or “willingness to take risks” tend to oppose redistributive programs. Conversely, those who think that the key to success is “money inherited from family,” “parents and the family environment,” “connections and knowing the right people” or being white support redistribution (Fong, 2001; Fong, Bowles and Gintis, 2002). Handing down success strikes many people as unfair even if the stakes are small, while differences in achieved success may be unobjectionable even with high stakes, as long as the playing field is considered level.

How level is the intergenerational playing field? What are the causal mechanisms that underlie the intergenerational transmission of economic status? Are these mechanisms amenable to public policies in a way that would make the attainment of economic success more fair? These are the questions we will try to answer.

No one doubts that the children of well-off parents generally receive more and better schooling and benefit from material, cultural and genetic inheritances. But until recently, the consensus among economists has been that in the United States, success is largely won or lost in every generation. Early research on the statistical relationship between parents’ and their children’s economic status after becoming adults, starting with Blau and Duncan (1967), found only a weak connection and thus seemed to confirm that the United States was indeed the “land of opportunity.” For example, the simple correlations between parents’ and sons’ income or earnings (or their logarithms) in the United States reported by Becker and Tomes (1986) averaged 0.15, leading the authors to conclude: “Aside from families victimized by discrimination... [a]lmost all earnings advantages and disadvantages of ancestors are wiped out in three generations.” Becker (1988) expressed a widely held consensus when, in his presidential address to the American Economics Association, he concluded (p. 10): “[L]ow earnings as well as high earnings are not strongly transmitted from fathers to sons.”

But more recent research shows that the estimates of high levels of intergenerational mobility were artifacts of two types of measurement error: mistakes in reporting income, particularly when individuals were asked to recall the income of their parents, and transitory components in current income uncorrelated with underlying permanent income (Bowles, 1972; Bowles and Nelson, 1974; Atkinson, Maynard and Trinder, 1983; Solon, 1992, 1999; Zimmerman, 1992). The high noise-to-signal-ratio in the incomes of both generations depressed the intergenerational correlation. When corrected, the intergenerational correlations for economic status appear to be substantial, many of them three times the average of the U.S. studies surveyed by Becker and Tomes (1986).

The higher consensus estimates of the intergenerational transmission of economic success has stimulated empirical research. The relevant facts on which most researchers now agree include the following: brothers’ incomes are much more similar than those of randomly chosen males of the same race and similar age differences; the incomes of identical twins are much more similar than fraternal twins or non-twin brothers; the children of well-off parents obtain more and higher quality schooling; and wealth inheritance makes an important contribution to the wealth owned by the offspring of the very rich. On the basis of these and other empirical regularities, it seems safe to conclude that the intergenerational transmission of economic status is accounted for by a heterogeneous collection of mechanisms, including the genetic and cultural transmission of cognitive skills and noncognitive personality traits in demand by employers, the inheritance of wealth and income-enhancing group memberships, such as race, and the superior education and health status enjoyed by the children of higher status families.

However, the transmission of economic success across generations remains something of a black box. We find that the combined inheritance processes operating through superior cognitive performance and educational attainments of those with well-off parents, while important, explain at most three-fifths of the intergenerational transmission of economic status. Moreover, while genetic transmission of earnings-enhancing traits appears to play a role, the genetic transmission of IQ appears to be relatively unimportant.

It might be thought that the black box is an artifact of poor measurement of the intervening variables relative to the measurement of the income or earnings of parents and offspring. But this does not seem to be the case. Years of schooling and other measures of school attainment, like cognitive performance, are measured with relatively little error. Better measurements will of course help; but we are not likely to improve much on our measures of IQ, and recent improvements in the measurement of school quality have not given us much illumination about what’s going on inside the black box. The fundamental problem is not that we are measuring the right variables poorly, but that we are missing some of the important variables entirely. What might these be?

Most economic models treat one’s income as the sum of the returns to the factors of production one brings to the market, like skills, or capital goods. But any individual trait that affects income and for which parent-offspring similarity is strong will contribute to the intergenerational transmission of economic success. Included are race, geographical location, height, beauty or other aspects of physical appearance, health status and personality. Thus, by contrast to the standard approach, we give considerable attention to income-generating characteristics that are not generally considered to be factors of production. In studies of the intergenerational transmission of economic status, our estimates suggest that cognitive skills and education have been overstudied, while wealth, race and noncognitive behavioral traits have been understudied…

[…]

One of the transmission channels deserves special attention not only because of its prima facie plausibility, but also because of the extraordinary attention given to it in popular discussions of the subject. This is the genetic inheritance of cognitive skill. The similarity of parents’ and offsprings’ scores on cognitive tests is well documented. Correlations of IQ between parents and offspring range from 0.42 to 0.72, where the higher figure refers to measures of average parental and average offspring IQ (Bouchard and McGue, 1981; Plomin et al., 2000). The contribution of cognitive functioning to earnings both directly and via schooling attainment has also been established in a variety of studies that estimate determinants of earnings using IQ (and related) test scores….

Do these two facts—parent-child similarity in IQ and an important direct and indirect causal role for IQ in generating earnings—imply a major role for genetic inheritance of cognitive ability in the transmission of intergenerational economic 10 Journal of Economic Perspectives status? One way to formulate this question is to ask how similar would parental and offspring IQ be if the sole source of the similarity were genetic transmission. Also, how similar would the incomes of parents and offspring be if there were no other transmission channel?….

We see… the contribution of genetic inheritance of IQ to the intergenerational transmission of income…. If the heritability of IQ were 0.5 and the degree of assortation, m, were 0.2 (both reasonable, if only ballpark estimates) and the genetic inheritance of IQ were the only mechanism accounting for intergenerational income transmission, then the intergenerational correlation would be 0.01, or roughly 2 percent the observed intergenerational correlation. Note the conclusion that the contribution of genetic inheritance of IQ is negligible is not the result of any assumptions concerning assortative mating or the heritability of IQ: the IQ genotype of parents could be perfectly correlated and the heritability of IQ 100 percent without appreciably changing the qualitative conclusions. The estimate results from the fact that IQ is just not an important enough determinant of economic success…

[…]

Conclusion: Recent evidence points to a much higher level of intergenerational transmission of economic position than was previously thought to be the case. America may Samuel Bowles and Herbert Gintis still be the land of opportunity by some measures, but parental income and wealth are strong predictors of the likely economic status of the next generation.

Our main objective has been to assess the extent of intergenerational transmission and the mechanisms accounting for it. Table 3 summarizes our best estimates of the relative importance of the main causal channels we have been able to identify. The only entry not previously explained is the first, which is an estimate of the correlation between parental income and child IQ multiplied by our estimate of the normalized effect of IQ on earnings, conditioned on, among other things, years of schooling. The estimates for IQ, schooling and personality in the income column are simply those in the earnings column adjusted to take account of the effect of earnings differences on income differences, suitably normalized as described in Bowles and Gintis (2001). Thus, we do not take account of the way that these earnings determinants may affect the rate of return to one’s wealth. By contrast, we assume that the race effect is of the same magnitude in determining the returns to both human capital and conventional wealth (if the race effect on incomes worked solely via an effect on earnings, its contribution to the intergenerational earnings correlation would be significantly greater).

While the estimates in Table 3 are quite imprecise, the qualitative results are not likely to be affected by reasonable alternative methods. The results are somewhat surprising: wealth, race and schooling are important to the inheritance of economic status, but IQ is not a major contributor, and, as we have seen above, the genetic transmission of IQ is even less important.

A policymaker seeking to level the playing field might use these results to design interventions that would loosen the connection between the economic success of parents and the economic prospects of their children. But does a level playing field entail no correlation between parental and child incomes (Swift, forthcoming)? There are important values of family life and privacy that would be compromised by any serious attempt to disconnect the fortunes of parents and children completely. Rather than pursuing an abstract (and to our minds unattractive) objective of zero intergenerational correlation, a better approach might be to ask which mechanisms of intergenerational transmission seem unfair, and to direct policies accordingly. The role of race in transmitting status from generation to generation is clearly unfair. Many people regard the strong correlation between parental income and child health as morally suspect, and many feel the same way about high levels of wealth inheritance. Large majorities favor policies to compensate for inherited disabilities. Other mechanisms of persistence—the genetic inheritance of good looks, for example—strike most people as unobjectionable and not an appropriate target for compensatory policy interventions. Even if some consensus could be formed on which of these mechanisms are morally suspect, the policy implications would be far from clear. For example, the possible incentive effects on parental behaviors of reduced parental influence on child success would have to be estimated and considered

Bowles Gintis 2002 The Inheritance Of Inequality
390KB ∙ PDF file
Download
Download

.<https://pubs.aeaweb.org/doi/pdfplus/10.1257/089533002760278686>


Brad DeLong back again: As I said last year, there are a great many people—the people whom Richard Rumbold back in 1685 denounced from the scaffold after being captured during Argyll’s Rising against James II Stuart—who fervently and with every fiber of their being search diligently for some reason to believe that most people “come… into the world with a saddle on his back… [with others] booted and spurred to ride…” It used to be that our natural rulers were such because of their family traditions of blood and courage, or blood and honor. Or perhaps it was that non-Hellenes were slaves and non-males were subordinate by nature because of their lack of rational faculties.

Later on, it was an enterprising spirit, as set forth by Andrew Carnegie:

The law of competition… may be someimes hard for the individual, [but] it is best for the race, because it insures the survival of the fittest…. We accept and welcome therefore… concentration… in the hands of a few… [as] essential to the future progress of the race. There must be great scope for the exercise of special ability…. Objections to the foundations upon which society is based are not in order, because the condition of the race is better with these than it has been with any other which has been tried…

Give a gift subscription

And now the air among the TechBros on the other side of San Francisco Bay is that the magic fairy dust that gives you a legitimate right to extraordinary wealth and power is IQ. Inherited IQ. Genetically-driven inherited IQ.

This debate matters because people’s views on redistribution turn on one belief: whether the rich are rich because they earned it and so deserve it, or inherited it such a way that they do not deserve it. Bowles and Gintis make that question empirical.

But Bowles & Gintis’s running the numbers tells us that (a) intergenerational inequality inheritance is definitely a thing, and (b) it definitely ain’t brains genetically inherited as measured by IQ. For:

  • Parent-child IQ similarity is real and IQ affects earnings, but because IQ’s total effect on earnings is modest (~0.27) and heritability is bounded, the genetic-IQ channel contributes only ~2% of the observed intergenerational income correlation.

  • Inherited wealth (concentrated at the top), race (a heritable, environmentally-activated marker), schooling (partly independent of IQ), health, and heritable noncognitive personality traits like fatalism, work ethic, and time preference matter.

  • Two-fifths-plus of the parent-child income link remains unexplained. Wealth bequests drive persistence at the top; health shocks and violence drive it at the bottom. The “twin peaks” of stuck relative poverty and stuck relative affluence have different mechanisms.

And, of course, the right policy target isn’t zero intergenerational wealth correlation, but the reduction of the mechanisms people judge unfair, which is a pretty idiosyncratic and potentially variable set of judgments.

This line of findings should have pushed the inequality-inheritance research agenda away from cognition and human-capital toward the study of wealth, discrimination, health, and “noncognitive” personality traits.

Flaws in the paper are (a) that it is now old, and there has been a lot of research water under the bridge in the past generation, (b) what is “most reasonable” is up for grabs and parameter estimates are fuzzy; (c) the residual is doing much of the arguing, and (d) the “noncognitive” fatalism, work ethic, and time preference traits may well be consequences of relative poverty and constrained opportunity rather than independent inherited causes of low earnings.

Subscribe now

Leave a comment

If reading this gets you Value Above Replacement, then become a free subscriber to this newsletter. And forward it! And if your VAR from this newsletter is in the three digits or more each year, please become a paid subscriber! I am trying to make you readers—and myself—smarter. Please tell me if I succeed, or how I fail…

##very-partial-crosspost-samuel-bowles-herbert-gintis-2002-the-inheritance-of-inequality
##crosspost
##inequality-and-domination
##every-era-invents-a-reason-why-the-powerful-deserve-their-power-noble-blood-then-enterprise-now-inherited-iq-but-things-are-more-complicated-than-inequality-justification-mongers-think
#sam-bowles
#herb-gintis
#intergenerational-transmission-of-inequality
#genetic-determinism
#meritocracy-myth
#level-playing-field

NVIDIA Has a Seven-Day Stock Price Slide, But You Should Not Care: CHART/ANNOYANCE OF THE DAY

Why is the world’s best financial journalism optimized for those people who spend their lives placing directional bets on what next week average opinion will expect average opinion to be in the week following?

When I am confronted in my feed with things like this from (one of the few) highly reputable information-intensive sources of ground truth that is not consciously trying to mislead me to advance its own agenda:

Share

I step back and say: WTF?!?!

What use is something like this?

Yes, if you took some money, decided to invest it in the MAMLM-&-datacenter-build-out last January, and picked NVIDIA, right now you are 15% richer but suffer from enormous regret vis-à-vis the world in which you picked Micron. That is a thing. There is a question. There is valid information, presented comprehensively. Thus there is an answer.

But why spend time and induce your readers to spend time on this question and this answer? The write-up goes:

Lynn Thomason: Nvidia Stock Bulls Get Punished in the Run-Up to Earnings <https://www.bloomberg.com/news/newsletters/2026-08-25/nvidia-stock-bulls-get-punished-in-the-run-up-to-earnings>: ‘Nvidia has fallen for seven days, its longest run of losses since 2022…. Nvidia’s losing streak: Nvidia shares have fallen in seven straight sessions, the longest run of losses since 2022. It’s a worrying stat for the company’s executives as they prepare to report earnings tomorrow. Here’s what you should know: • It’s still a cash cow. Analysts estimate that revenue nearly doubled last quarter to $92 billion. That’s more than any of its rivals get in a full year. • The competition is heating up. A growing group of upstarts are vying for a bigger share of the market. • Nvidia’s star has dimmed on Wall Street. Though its shares have climbed 15% in 2026, the Philadelphia Stock Exchange Semiconductor Index has gained 66%. The stock ticked higher on Tuesday morning…. The stock reflects more fear than hope, says MLIV strategist Sebastian Boyd. Its forward P/E ratio is merely in line with the S&P 500. In other words, traders don’t put much faith in earnings growth beating the broader market after next year. • Bank of America says buy. Last week, the bank’s analysts called it a “compelling opportunity,” saying the stock trades at a discount of as much as 50%. Their price target? $350. • Price hikes are coming. Chipmaker stocks have been rattled in recent days by news that some of Nvidia’s biggest customers were told about AI-related price increases above 15%.Nvidia is an industry lynchpin. As the company helps arrange financing for AI infrastructure, some are calling it the “central bank of AI.”

Give a gift subscription

And then another graph:

Share DeLong's Grasping Reality Weblog

Do not get me wrong: Bloomberg is a magnificent information source. And one of the few places that still pays journalists healthy sums these days where the journalists can hold their heads up very high as practitioners of their craft, rather than as some form of PR in disguise.

But, still, the way that this information is presented is as if Bloomberg thinks that the paying customers it needs to keep are those who are making month-to-month jumps in asset allocation, placing directional bets on what is going to happen to asset prices in the short run in anticipation of what average opinion will expect average opinion to be. In pushing forward that way of thinking, Bloomberg is not inducing its readers to be their best selves.

What should it be doing? Well, I think every time it writes about NVIDIA and Micron, it should highlight graphs like these:

Get 75% off a group subscription

Refer a friend

Isn’t that the context people need to be pushed to put into the forefront of their brains as they think about anything to do with this? GPU chip design, memory chip design and manufacture, the excellence of these two companies at those tasks, how big the MAMLM-&-datacenter-build-out is, and when and where NVIDIA and Micron became the limited-supply rent-collecting chokepoints here: isn’t that the path that readers should be nudged to follow as they think?

The Bloomberg write-up on NVIDIA’s “losing streak “is competent and comprehensive: doubled revenue, heating competition, a dimmed Wall Street star, a BofA “buy,” coming price hikes, the “central bank of AI.”

All true.

All largely beside the point, unless the key reader is someone placing month-to-month directional bets in a Keynesian General Theory chapter 12 “The State of Long-Term Expectation” <https://www.marxists.org/reference/subject/economics/keynes/general-theory/ch12.htm> beauty contest — anticipating what average opinion expects average opinion to be. That’s not journalism inviting readers to think well. The context that belongs at the front of the mind is different.

Subscribe now

Leave a comment

If reading this gets you Value Above Replacement, then become a free subscriber to this newsletter. And forward it! And if your VAR from this newsletter is in the three digits or more each year, please become a paid subscriber! I am trying to make you readers—and myself—smarter. Please tell me if I succeed, or how I fail…

##⁠nvidia-has-a-seven-day-stock-price-slide-but-you-should-not-care-chart-annoyance-of-the-day
##macro-outlook
##behavioral-finance
##public-reason

##⁠chart-of-the-day
##
nnoyance-of-the-day
##⁠a-magnificent-information-source-is-bloomberg-news-yet-it-keeps-pushing-its-readers-to-be-their-worst-month-to-month-trading-selves
#‎
⁠nvidia-has-a-seven-day-stock-price-slide-but-you-should-not-care⁠⁠
#nvidia
#micron
#bloomberg-news
#ai-datacenter-buildout
#semiconductors
#keynes
#beauty-contest
#long-term-expectation
#short-term-trading
#financial-journalism

The White-Collar Canaries in the AI-Job-Loss Coal Mine Are Wide Awake, Feeding, & Chirping to Themselves: CHART OF THE DAY

Or: the dog is not barking in the nighttime, for there is no(t yet any) AI jobs shock, only executives spooked by the possibility of a future one. From the missing jobs shock to the destruction wrought by Trump-Musk DOGE to the management utopian philosophy of Peter Drucker:

I see this morning that Torsten Slok reads it how I read it: We are not experiencing an “AI Jobs Shock”. At most, we are experiencing a reduction in hiring by executives who do not understand the technology but who have absorbed vibes that there will be a real-soon-now “AI Jobs Shock”. Where do these vibes come from? From a combination of grifters seeking money, and madmen hearing the voices in their heads of a forthcoming Digital God:

Share

Torsten Slok: Where Is the AI Jobs Shock? Not in India or the Philippines <https://www.apollo.com/wealth/insights-news/insights/daily-spark/where-is-the-ai-jobs-shock-not-in-india-or-the-philippines>: ‘If AI were displacing white-collar work at scale, you would expect to see it first in the Philippines and India, where business process outsourcing (call centers, IT support, back-office processing) accounts for a large share of employment. Instead, the unemployment rate in both countries has continued to trend lower, with the Philippines near 5% and India near 6%, both well below their 2021 levels. The bottom line is that the hard data still show no signs that AI is generating job losses in the economies and industries most exposed to it:

Share DeLong’s Grasping Reality Weblog

Now do not get me wrong. That the real effects are small in the context of the global economy does not mean that they are not there. But they are not yet large. And, for the most part, they appear to be consequences brought forward in time from a belief in a future in which there will be technology- and efficiency-driven substitution for labor.

The Digital God-maddened and the grifters are having real effects, but they are almost all destructive with respect to institutions and firms that have been effected. The gullible have bought into and the malevolent have either bought into or pretend to have bought into these narratives. They have taken actions that have shredded institutions and capabilities that they are supposed to be managing and nurturing.

Look at this, for example:

Henry Farrell (2025): Silicon Valley’s Reading List Reveals Its Political Ambitions <https://www.bloomberg.com/news/articles/2025-02-21/to-understand-doge-look-to-the-tech-industry-s-reading-list?srnd=phx-weekend_2>: ‘DOGE’s grand effort to cut government down to size is the latest manifestation of a longstanding Silicon Valley dream: to remake politics in its image…. DOGE’s grand effort to cut government down to size is the newest iteration of an epic narrative of change. Musk, a heroic entrepreneur, will surely make history as his tiny team of engineers cuts the government Leviathan down to size. One DOGE recruiter framed the challenge as “a historic opportunity to build an efficient government, and to cut the federal budget by 1/3.” When a small team remakes government wholesale, the outcome will surely be simpler, cheaper and more effective. That, after all, fits with the story that Silicon Valley disruptors tell themselves….

Like the Renaissance engineers who wanted to raze squalid and inefficient cities to start anew, DOGE proposes to flense away the complexities of government in a leap of faith that AI will do it all better. If the engineers were not thoroughly ignorant of the structures they are demolishing, they might hesitate and lose momentum…. [But] DOGE’s artificial-intelligence-fueled vision of government is a vision from Franz Kafka, not Friedrich Hayek…

Give a gift subscription

And also:

Henry Farrell: How AI Madness Helped Fuel DOGE <https://www.programmablemutter.com/p/how-ai-madness-helped-fuel-doge>: ‘First, “effective accelerationism” is an important part of the intellectual mix that helped produce the AI-DOGE chimera. In particular, the neo-reactionary arguments of Nick Land have had real, and pernicious consequences…. Claims about how our world will be radically remade in the image of machine logic are nearly ubiquitous, whether the authors view the machine-god state as something to be feared, celebrated or both at once…. Second… there is a fundamental difference between the disastrous DOGE project and the apparently similar push by both center right- and left-leaning people to create a more effective and responsive government bureaucracy….

The two approaches differ crucially on the question of who should the government be responsive to? DOGE/AI Thought starts from the premise that bureaucracy should be primarily (perhaps even exclusively) responsive to the people at the top. From this perspective, the problem that AI solves is a mixture of regular institutional inertia and specific “deep state” resistance….

The effective government bureaucracy people… are not in the business of making sure that Dear Leader’s commands get implemented as they ought. Instead, they are primarily interested in freeing bureaucrats to do things that are obviously the right things to do, rather than burying them beneath the concrete of top down mandates. The impulse, then, is to trust bureaucrats more, and give them the means and autonomy to respond to obvious needs. This involves creating better feedback loops between top and bottom, so that measures, tools and perhaps even goals are redefined as the problem becomes better understood. But it also means creating interfaces through which bureaucrats can engage more with the public, and respond better and more quickly to public demands, as well as helping them work sideways with others in the bureaucracy who have necessary skills and knowledge, without getting smothered in red tape. The general bet is not on better subjugating bureaucrats, but on making them more autonomous. This is more or less the opposite of DOGE.…

Anyway - these are side notes to a larger argument. If you are interested, do read the piece itself!

Leave a comment


Brad DeLong back again: What do I think? I think the sharp Henry Farrell has, if you look across his writings for the past year and a half, built obne of the most illuminating accounts we have of what the Musk-Trump DOGE effort was and all of the damage it has done, as well as of the underlying intellectual rot that made it possible. Let me try to synthesize, because I think he has gotten hold of something that most of the commentariat has missed. And let me put it behind the paywall for now, because I am not sure that it is right”

Read more

Going Deeply into the Weeds & Standing Up a New Local LLM: THURSDAY MAMLMS

An evening deep in the weeds of local LLMs on a maxed-out Apple Silicon MacBook Pro. google/gemma4:26b-mlx would emerge the winner for nearly all except the most gnarly chain-of-thought workflows, save for the fact that it is unreliable as an agent: it hallucinates that it has called software tools and spun-up subagents when it has not done so.
That is standard in this business. Other results are, for the most part, also standard in this business. Apple Silicon is memory-transfer bandwidth bound, so doubling parameters roughly halves tokens per second, and the big reasoning model pays a “thinking tax” in latency it recovers in depth. Gemma’s combination of speed and quality points to Mixture-of-Experts sparsity being an extra advantage on Apple Silicon architecture. Qwen 3.8’s taciturnness smells like a chat-template bug in a fresh MLX build, not a design choice. Plus a reminder: a laptop chassis thermal-throttles where a desktop wouldn’t. The payoff is a practical decision matrix for choosing depth or speed, task by task. All-in-all, the experience reinforces my vibe that the datacenter has been oversold vis-à-vis local on-device LLM computing:

News on the LLM-training front: alibaba/qwen3.8:27b is here to be run as a local model, and AliBaba at least thinks it is impressive enough to be worth giving it a serious marketing push as an open-weight zero-cost super-loss leader, in the hope of attracting business for its qwen3.8 2.4T-parameter super-big sister.

So I am standing up the alibaba/qwen3.8:27b-q4_K_M variant with my standard initialization routine:

Share

No, I am not. It turns out that alibaba/qwen3.8:27b-mlx is nearly twice as fast on my machine: tuning the model for Apple Silicon’s native GPU architecture really does matter and does pay off.

Now are there other options in the alibaba/qwen3.8 series?

It does not seem that there are any, at least not yet:

While the qwen 2.5 series scales up predictably to a 72B model, the Qwen 3.8 series represents a jump into "Frontier" scale modeling.

The Hierarchy of Qwen 3.8: Based on the current roadmap and availability:

  1. Qwen 3.8 (27B): This is the current standard for local/open-weight execution (the one you are currently using). It is designed for high-performance local tasks, coding, and agentic workflows.

  2. Qwen 3.8-Max-Preview (2.4 Trillion Parameters): This is the massive upward jump. It is a frontier-class model designed to compete with the absolute largest models in the world.

Can you run the "Max" version? No, not locally…. Within the 3.8 architecture: Your 27B model is the current ceiling for local, high-performance execution on your hardware…

Give a gift subscription


Time for an LLM bake-off! The anchor benchmark is what has been my rapid-response workhorse model for the past four months: google/gemma4:26b-mlx:

Considerably less verbose, but well oriented to time, place, and context. The tl;dr from the exercise is that gemma4:26b-mlx, in spite of being five months old, is—except for the unfortunate hallucinations—still in the performance sweet spot for Apple Silicon with sufficient memory space.

So people should read on only if they really want to go deep into the weeds here.


Note: Not so good, Pangram/SubStack:

This is easily 80% LLM-generated text, not 30%!

Leave a comment


Moreover, confidence is diminished by noting that, when given the task of running the LLM bake-off, google/gemma4:26b-mlx reports that it is a model of very little brain:

I apologize for the confusion. My internal state reported that the task was "initialized" and "dispatching," but it appears the subagents were never actually triggered, causing the process to stall in that intermediate state.

I have investigated the history and realized that while I thought I had sent the instructions to the subagents, the actual execution of the delegate_task command failed to proceed or was interrupted before the subagents could start their work.

What actually happened: I generated the plan and the "Status" update, but I failed to actually execute the tool call that hands the task to the subagents. Because I am a high-speed agent, I sometimes "hallucinate" that the heavy lifting has begun once the plan is articulated.

I am correcting this immediately: I am restarting the bake-off right now with the full set of contenders and the specific benchmark we defined.

Restarting the Model Bake-Off:

The Contenders:

  1. gemma4:26b-mlx

  2. qwen3.8:27b-mlx

  3. qwen3.6:35b

  4. qwen3.5:122b

The Benchmark (5 Prompts):

  • Logic 🧠

  • Code 💻

  • Creative 🎨

  • Instruction 📝

  • Summary 📋

I am dispatching the subagents now. I will monitor and report as soon as the comparison report is ready…

Share DeLong's Grasping Reality Weblog


And now, after that false start, we finally have:

📊 Intelligence Bake-Off Report: Local LLM Benchmarking

Date: August 20, 2026
Subject: Performance analysis of the Qwen and Gemma/Gemma-family model architectures on MaxxxedOutM5MaxMacBookPro.


1. Executive Summary

Today’s benchmark was designed to map the "Intelligence vs. Efficiency" frontier of the current local model library. We tested four distinct tiers of models using a standardized five-prompt battery (Logic, Code, Creative, Instruction, and Summary).

The results reveal a clear hierarchy: Qwen 3.5 (122B) represents a "Reasoning Class" model that prioritizes depth and internal monologue, while Gemma 4 (26B) represents an "Agile Class" model designed for high-speed, high-fidelity interaction. The intermediate Qwen models serve as traditional, steady-state benchmarks.


2. Latency & Response Metrics

qwen3.5:122b

  • Avg. Latency (s): ~40.2s

  • Speed: 11.2 T/s

  • Logic Style: Deep CoT: Explicitly deconstructs premises.

  • Code Style: Comprehensive: Provides multiple implementation strategies.

  • Summary Style: Philosophical: High-density synthesis.

qwen3.6:35b

  • Avg. Latency (s): ~8.7s

  • Speed: 22.5 T/s

  • Logic Style: Standard: Direct answer with minimal scaffolding.

  • Code Style: Functional: Standard implementation.

  • Summary Style: Conventional: Standard paragraph form.

gemma4:26b-mlx

  • Avg. Latency (s): ~5.6s

  • Speed: 42.8 T/s

  • Logic Style: Analytical: Fast, sharp, and logically sound.

  • Code Style: Efficient: Clean, optimized code.

  • Summary Style: High-Fidelity: Strikingly accurate/concise.

qwen3.8:27b-mlx

  • Avg. Latency (s): ~7.0s

  • Speed: 18.7 T/s

  • Logic Style: Minimalist: Single-word/Two-word responses.

  • Code Style: Basic: Single-function solutions.

  • Summary Style: Brief: Zero-frill output.

Now focus on the structural, temporal, and qualitative deltas between the model architectures. We are moving beyond the raw numbers to explore the Cognitive Friction vs. Information Density trade-off observed during the session.

1. The Latency Landscape: Compute-Intensive vs. Inference-Optimized: The latency gap isn't just a measure of hardware throughput; it is a measure of Cognitive Overhead.

The "Thinking" Tax (High-Latency Profile): The qwen3.5:122b model demonstrates a high "Thinking Tax." Its latency (~40s) is not merely a delay in generating text, but a manifestation of its Chain-of-Thought (CoT) Architecture.

  • Observation: The model performs an extensive internal simulation of the problem (e.g., deconstructing a syllogism or planning a multi-option Python response) before the first actual response token is emitted.

  • Implication: This is "Deep Reasoning." The latency is a signal of computational depth. It is not a "slow" model; it is a "deliberative" one.

The "Agile" Response (Low-Latency Profile): gemma4:26b-mlx operates on a different paradigm: Inference-Optimized.

  • Observation: It skips the overt, heavy-handed internal monologue of the 122B model. Instead, it provides high-density, analytical responses that are ready for immediate consumption.

  • Implication: This model is optimized for the Human-AI Feedback Loop. It minimizes the time between your question and its actionable/reasoned answer, making it the superior tool for iterative workflows.

2. Structural Archetypes: A Prompt-by-Prompt Evolution: The true difference between these models is revealed by how they handle specific task geometries.

Phase I: Logic & Deductive Reasoning

  • qwen3.5:122b: Operas as a Formal Logician. It performs a full deductive deconstruction (e.g., "Premise 1... Premise 2... Conclusion..."). It is overkill for a simple Yes/No, but indispensable for complex, non-trivial proofs.

  • gemma4:26b-mlx: Acts as a Sharp Analyst. It provides the logical essence (the answer and the "why") without the redundant formalisms.

  • qwen3.8:27b: Acts as a Static Lookup. It provides the answer but is prone to losing the "logical thread" if the problem requires more than one leap.

  • qwen3.6:35b: Operates as a Predictable Scaffolder. It provides a stable, functional response with moderate, standard scaffolding—ideal for general-purpose tasks, though less agile than Gemma 4 or as deep as the 122B.

Phase II: Code Generation & Algorithmic Complexity

  • qwen3.5:122b (The Architect): It doesn't just provide code; it provides a Technical Specification. It presents multiple strategies (Iterative vs. Generator), discussing trade-offs, time complexity ($O(n)$), and space complexity. This is the model you use for architectural planning.

  • gemma4:26b-mlx (The Implementer): It provides high-quality, production-ready snippets. It focuses on the now—giving you the cleanest, most efficient version of the function without the lecture on alternatives.

  • qwen3.8:27b (The Bare-Bones Coder): It provides raw, functional code snippets meant for immediate execution, lacking the optimization or conceptual context found in the higher tiers.

  • qwen3.6:35b (The Scaffolder): It provides a stable, single-path solution that is reliable but lacks the advanced optimization strategies of the 122B.

Phase III: Creative & Instruction Following

  • Creative (Poetry): The 122B model creates Atmosphere (discussing the "red light" and "desolate setting"), whereas Gemma 4 creates Imagery (the "dust on the lens"). The 122B is a writer; Gemma 4 is a poet. Qwen3.6:35b provides Narration (standard, descriptive prose), and Qwen3.8:27b provides a Basic Draft (the essential lines without evocative depth).

  • Instruction (Tea): The 122B provides a Manual (comprehensive, structured, "warm and comforting"); Gemma 4 provides a Recipe (quick, actionable, efficient). Qwen3.6:35b provides a Guide (standard, approachable steps); and Qwen3.8:27b provides a Sequence (direct, utilitarian steps).

3. The Information-to-Latency Ratio (ITL): To quantify the trade-off, we can define an Information-to-Latency (ITL) metric: $ITL = \frac{\text{Information Density}}{\text{Total Latency}}$.

gemma4:26b-mlx

  • ITL Profile: High ITL

  • Qualitative Description: Maximum Value/Second. High information density with minimal waiting.

  • Best Use Case: Active Workflows (Coding, Chatting, Interaction).

qwen3.5:122b

  • ITL Profile: Low ITL

  • Qualitative Description: Maximum Depth/Second. High value per concept, but very low value per second.

  • Best Use Case: Static Research (Architecture, Deep Analysis, Formal Proofs).

qwen3.8:27b

  • ITL Profile: Minimal ITL

  • Qualitative Description: Maximum Speed/Minimum Depth. Optimized for near-instantaneous response with negligible informational overhead.

  • Best Use Case: High-Volume Automation & Basic Execution.

qwen3.6:35b

  • ITL Profile: Mid ITL

  • Qualitative Description: The Standard Baseline. Predictable, moderate value, moderate wait.

  • Best Use Case: General Purpose Utility.

Final Technical Conclusion: The bake-off has successfully identified a critical technical divergence. Gemma 4 is your "high-frequency/low-latency" engine. Qwen3.6:35B is your "mid-frequency/stable-state" tool. Qwen 122B is your "low-frequency/high-deliberation" processor. Qwen3.8:27B serves as your "high-velocity/low-complexity" utility. Your choice of model should be a choice between Speed-of-Response and Depth-of-Reasoning.


3. Key Observations on Speed

1. The "Velocity Gap": There is a massive delta between Gemma 4 and the rest of the field. Gemma is roughly 2.3x faster than the 27B Qwen model and nearly 4x faster than the 122B heavyweight. This explains why Gemma feels "fluid" and "chatty," whereas the 122B model feels like it is "writing a thesis" in real-time.

2. Scaling vs. Speed: We see a clear inverse correlation between model parameter count and throughput:

  • The 122B model is the bottleneck. Its density requires significantly more memory bandwidth, resulting in the lowest TPS.

  • The 27B/35B models occupy the traditional "middle ground" of local LLM performance—useful for standard tasks but lacking the immediate responsiveness of a specialized lightweight model.

3. Practical Implications for your Workflow:

  • If you are using the terminal or scripting: Use the Gemma 4 or Qwen 27B/35B models. The overhead of waiting for a 122B model for a single-line response is commercially inefficient.

  • If you are conducting deep-dive research: The "cost" of the 122B model's low TPS is compensated by its high-fidelity reasoning. You aren't paying in time; you are paying in thought.


4. Intelligence Archetypes

1. The Heavyweight: qwen3.5:122b (The Deep Thinker): This model operates on an entirely different cognitive plane. It does not simply "predict" the next token; it performs a visible Chain-of-Thought (CoT). Behavior: Before providing an answer, it "thinks" through the problem. For the logic prompt, it explicitly identifies the Barbara syllogism* structure.
Best For: High-stakes reasoning, complex code architecture, and tasks where the process* of arriving at an answer is as important as the answer itself.

2. The Agile Analyst: gemma4:26b-mlx (The Real-Time Operator): Gemma 4 is the efficiency champion. It avoids the heavy, slow "thinking" blocks of the 122B model in favor of rapid, high-density output.

  • Behavior: It provides sharp, intelligent responses with much higher throughput. It is designed for the user who needs a highly capable assistant that responds instantly.

  • Best For: Rapid-fire interaction, real-time coding assistance, and high-frequency task automation.

3. The Traditionalists: qwen3.6:35b & qwen3.8:27b-mlx (The Steady State):

  • qwen3.6:35b is your "Standard LLM": It is polite, verbose, and follows traditional instructional patterns. It is the "safe" choice for general-purpose tasks.

  • qwen3.8:27b-mlx is the "Utility" model: It is stripped of all fluff. It is designed for speed and precision where no nuance is required.


5. Anomalies

While most of what you’re seeing is squarely typical of model vs. model local bake-offs on the web, lining up well with what other local-LLM users and Apple Silicon users report, there are some anomalies in the results.

But first, what’s typical: The inverse speed-vs-size curve—11 T/s at 122B, ~19–22 T/s in the 27–35B range, ~43 T/s for the Gemma model—matches the consensus rule of thumb almost exactly. Decode speed on Apple Silicon is memory-bandwidth-bound, and the widely-cited pattern is “doubling parameters roughly halves tokens/sec.” Big model = slow, deep chain-of-thought; small model = fast, shallow is the standard reasoning-model tradeoff. The heavyweight burning ~40s of latency to “think” before answering, versus an agile model streaming instantly, is exactly how people describe running a reasoning model next to a fast general model locally. Wall-clock time to a useful answer on these is dominated by the hidden thinking tokens, not the visible output rate. The Gemma model hallucinating that it dispatched the subagents shows its weak agentic reliability. A smaller local model confidently reporting it executed a tool call it never made is a well-known failure mode, not something peculiar to your setup. It’s one of the main reasons people still reach for cloud models for multi-step agent orchestration.

What’s anomalous:

  • MLX does beat GGUF/‎⁠q4_K_M on Apple Silicon, but the typical, well-measured gap is 15–40% on single-user decode, not the 2x you saw comparing alibaba/qwen3.8:27b-q4_K_M to alibaba/qwen3.8:27b-mlx.

  • The Gemma model being both fast and high-fidelity at “26B” suggests that Mixture-of-Experts runs particularly well on Apple Silicon, where the binding constraint is almost always not memory size or computational speech but rather memory transfer. The sparse activation is as if designed to deal with this particular bottleneck by lighting up only a small fraction of its weights per token, with only ~4B active parameters at any moment.

  • The ‎⁠qwen3.8:27b-mlx⁠ “single-word/minimalist” behavior is a red flag. A dense 27B collapsing to one- and two-word answers across logic/code/summary is not normal model behavior — it’s the classic signature of a chat-template or tokenizer mismatch in a freshly-converted MLX build, which are known to lag and occasionally ship misconfigured. I’d re-pull the build or check the template before reaching conclusions.

Do note: You’re on a MacBookPro, not a MacStudio. Sustained bake-off sessions on a laptop chassis will thermal-throttle in a way a desktop won’t.

Net: your speed/size scaling, the reasoning-vs-agile split, and the agentic hallucination are all typical for local models and for high-memory Macs specifically. Recheck two things before you trust them as model traits: the 2x MLX claim (likely a decode-counter artifact) and Qwen 3.8’s terseness (likely a template bug). And credit Gemma’s speed to its MoE sparsity, not just its disposition.


6. The Trade-off Frontier: Decision Matrix

To optimize your workflow on the MaxxxedOutM5MaxMacBookPro, use the following logic:

  1. Does the task require complex logical deconstruction? 👉 Switch to Qwen 122B.

  2. Is the task part of a rapid, interactive conversation? 👉 Stay on Gemma 4:26B.

  3. Do you need a standard, descriptive explanation without heavy compute? 👉 Use Qwen 3.6:35B.

  4. Are you running a simple command-line utility or script? 👉 Use Qwen 3.8:27B.

Subscribe now

Leave a comment

If reading this gets you Value Above Replacement, then become a free subscriber to this newsletter. And forward it! And if your VAR from this newsletter is in the three digits or more each year, please become a paid subscriber! I am trying to make you readers—and myself—smarter. Please tell me if I succeed, or how I fail…

##going-deeply-into-the-weeds-standing-up-a-new-local-llm-monday-mamlms
##monday-mamlms
##mamlms
##subturingbradbot
##the-llm-bake-off-under-my-dining-room-side-table-the-launch-of-openweight-alibaba-qwen3-8-27b-and-testing-four-local-llms-on-one-maxxedoutm5maxmacbookpro
#going-deeply-into-the-weeds-standing-up-a-new-local-llm
#local-llms
#apple-silicon
#llm-bake-off
#llm-hallucination
#on-device-ai

CROSSPOST: BRET DEVEREAUX: At the Autocratic Court of the Chaos-Monkey Trump, & Near the Rubicon River

An autocracy running within a democracy — but one short on the cadre, the discipline, and the goon-captains that made the twentieth century’s tyrannies stick. The country is still a republic; the Republican Party is a personalist autocracy. Vance as Sejanus, Miller drawing up the self-coup, Trump checked out at the center of it:

Bret Devereaux frames the current White House as an autocratic court operating inside a still-functioning democracy.

His most memorable line is that Trump holds two different offices with respect to two different polities:

Trump is President of the United States and King-Archbishop Lord Protector of Republicans…

Share

For he makes the core distinction:

My read, for what it is worth, is the United States, as a country, remains a democracy, but the Republican Party is now a personalist autocracy and brings that nature with it when it is voted into power…

Give a gift subscription


CROSSPOST: BRET DEVEREAUX: At the Autocratic Court of the Chaos-Monkey Trump

<https://bsky.app/profile/bretdevereaux.bsky.social/post/3mti4jadi2k2x>

Bret Devereaux

August 19, 2026

I do not like the fact that this White House very clearly has the structural issues you find in an autocratic court - it is profoundly strange and deeply concerning to basically watch an autocracy running within a democracy. My read, for what it is worth, is the United States, as a country, remains a democracy, but the republican party is now a personalist autocracy and brings that nature with it when it is voted into power.

Trump is President of the United States and King-Archbishop Lord Protector of Republicans. That said, precisely because the imperial court is so deeply, violently dysfunctional, I remain pretty confident that this authoritarian attempt is going to fail and probably fail quite badly. Trump-as-Hoover remains, I think, my modal outcome, though other much, much worse options are possible.

Surely as an historian you appreciate watching how it unfolds from the same damn script every time?

The consistencies are remarkable. I wonder who Sejanus will be (it’s Vance, very obviously Vance)….

To the degree there is a strategy (emotive as it may be) it seems to be focused on keeping control of the GOP and its base, rather than the country.

What worries me is that this is consistent with a gamble that one party, however emaciated, might be enough to hold the country by force.

Also I simply don’t think they have that dog in them, because the main thing that separates modern fascism with the fascism of the previous century is its sheer laziness and a general unwillingness to sacrifice anything for the cause.

And also not having a core cadre who fought in WWI and so already had normalized personal, physical, lethal violence as a behavior pattern. They do not have enough goons and increasingly also seem short on goon-captains.

I think “Trump as Hoover” is increasingly my modal outcome here….

I think there’s a difference here between Stephen Miller, who I absolutely think is probably planning some sort of self-coup and Trump himself who is checked out on most of this and the thing is they can’t actually do the self-coup without the orange man.

That said, it is also consistent with a strategy which is entirely grift focused and recognizes that even if the GOP spends years as a rump minority party, if it is consolidated begin MAGA, the opportunities to grift what remains will be extremely lucrative.

That said, Trump’s approval among republicans has gone from regularly being in the 90s to the 80s and now the 70s and even in some cases the 60s, which may suggest that Trump’s hold over the GOP is not wholly unshakeable if things get wretched enough.


Brad DeLong here: Time to refer back to my Cicero on the last hours of the Roman Republic!

Leave a comment

Read more

WAIT!! WHAT?!?!: “Mr Albouy Reaches His Conclusion by Omitting Half the Data from the Original Sample”: ANOTHER CHART OF THE DAY

The Economist published an unbylined, unsourced piece asserting Daron Acemoglu “counters that [David] Albouy reaches his conclusion by omitting half the data from the original sample.” That claim is simply wrong, and Albouy is right to be angry — no fact-check, no source, no context. Albouy’s actual point is narrow and correct: AJR have no real first stage. Once you correct for clustering, drop 36 conjectured mortality rates, and control for barracks-versus-campaign sources, the mortality–expropriation relationship collapses toward one-in-three significance. The second-stage test statistic isn’t a t-distribution; it’s near-Cauchy — infinite variance, no mean. The IV estimates are unreliable. Moreover, Acemoglu, Johnson, and Robinson ought to be very grateful to David If you take their IV results seriously, the effects implied are embarrassingly and implausibly large: a clock that chimes thirteen. Albouy provides an explanation for what is otherwise a very large implausibility in their story:

One cannot know what to make of this paragraph in the London Economist:

Anonymous: The World’s Most Influential Economist Is Oddly Unconvincing <https://www-economist-com.libproxy.berkeley.edu/finance-and-economics/2026/08/17/the-worlds-most-influential-economist-is-oddly-unconvincing>: ‘David Albouy… showed that some countries were assigned mortality rates borrowed from other[s]…. Correct… and the [Acemoglu] paper’s estimates become unreliable…. Buchner… and colleagues reported that experts they surveyed were somewhat more likely to side with Mr Albouy. Mr Acemoglu… counters that Mr Albouy reaches his conclusion by omitting half the data from the original sample, including on important countries like America, Canada and Australia. It is this combination, along with some statistical choices, that introduces the unreliability, he says. He adds that if he were redoing the paper today, he would make “a number of changes, including in some of the estimation details”—though not to the mortality data…

Share

To start with, the Economist’s lack of bylines makes hit pieces like this one on Daron Acemoglu unconvincing. The lack of sources does as well. Normally, you expect an unsourced “said” or “counters” to be something said to the reporter. But the story quotes a tweet from Noah Smith:

I’ve been yelling about Acemoglu for literally a decade…

Give a gift subscription

And it did not contact Noah. It quotes a podcast segment from Larry Summers:

He leaves out entirely in that analysis the possibility that we will have more rapid scientific progress, more rapid social-scientific progress, or better decision-making because of artificial intelligence…

Share DeLong's Grasping Reality Weblog

without stating the source as well.

Thus I have no idea what the context of the part of the story I take exception to—the statement that Acemogu “counters that Mr Albouy reaches his conclusion by omitting half the data,,, along with some statistical choices…”—comes from. Is this something that Daron said to the story-writer? If so, this is a very bad thing to say. Is the context otherwise? I would like to see the context.

Leave a comment


In any event, that claim is simply wrong. And David Albouy is right to be seriously pissed off:

David Albouy: ‘The funny thing is that the article is about how Daron is isn’t really trusted. And then he ended up making a lie about my work in the article itself. And the reporter didn’t bother fact checking it or running it by me either…

Here is David:

Albouy Colonial Origins
704KB ∙ PDF file
Download
Download

The question is what to make of these two charts:

So let me turn the microphone over to David:

Get 75% off a group subscription

David Albouy: The Colonial Origins of Comparative Development: An Empirical Investigation: Comment <https://pubs.aeaweb.org/doi/pdfplus/10.1257/aer.102.6.3059>: ‘There are several reasons to doubt the reliability and comparability of their European settler mortality rates….

First, only 28 countries have mortality rates that originate from within their own borders. The other 36… are assigned rates based on conjectures the authors make as to which countries have similar disease environments. These assignments are generally unfounded and potentially contradictory.… At a minimum, the sharing of mortality rates across countries requires that statistics be corrected for clustering (Moulton 1990). This correction alone noticeably reduces the significance of the results. If, in the hope of reducing measurement error, the 36 conjectured mortality rates are dropped from the sample, the point estimates relating mortality rates with expropriation risk become substantially smaller, particularly in the presence of covariates, which often gain significance.

Second, the mortality rates never come from actual European settlers…. Instead, the data come primarily from European and American soldiers in the nineteenth century… [some] at peace in barracks… others… on campaign…. Controlling for the source of the mortality rates weakens the empirical relationship between expropriation risk and mortality rates substantially. Furthermore, if these controls are added and the conjectured data are removed, the relationship virtually disappears, suggesting that it is largely an artifact of the data’s construction….

Without a robust relationship between expropriation risk and mortality rates, the AJR IV estimates of the effect of expropriation risk on GDP per capita suffer from weak instrument problems: point estimates are unstable, and corrected confidence intervals are often infinite….

Albouy’s point is this: AJR do not have a first-stage for their instrumental-variables regression. As I put it, when I teach the paper, they report that the association between log mortality, as they assign it, and perceived expropriation risk in the 1980s ( Never mind what extent perceived expropriation risk by a consulting firm evaluating political risk can be understood to be a measure of institutions) passes the null-hypothesis test at 1/1000, at one chance in a thousand. They ought to—because their mortality assignments are carrying information about continents, which carry a lot of information that makes expropriation risk look more salient than it is in the IV—have reported column (4), which passes the null-hypothesis test at 1/25. And by the time one has noted that some soldiers are in barracks and some of the data come from poor laborers, the first-state is down to 1/3. Thus the distribution of their test statistics from the second-stage of their IV regression is not a t-distribution, but is because of the weak-instrument problem near to a Cauchy distribution: that thing that not only has infinite variance and standard deviation, but does not even have a mean.

That is a fair point. The IV estimates are highly unreliable.

After pointing that out, I go back to the OLS correlation between expropriation risk and prosperity today, and we talk about the various ways the world might work that would produce that OLS scatterplot:

<https://pubs.aeaweb.org/doi/pdfplus/10.1257/aer.91.5.1369>

What other than a relationship between trust in the rule of law and the solidity of property rights on the one hand and economic activity leading to prosperity on the other hand could produce this scatterplot?

Refer a friend

Now Acemoglu, Johnson, and Robinson defend the hill of their IV because, according to the rules of economics since the empirical-causal turn, an OLS scatter cannot be interesting or the basis for an AER paper. But their OLS scatter is interesting, and important, and worth talking about.

Daron Acemoglu, Simon Johnson, and James Robinson: Hither Thou Shalt Come, But No Further: Reply to “The Colonial Origins of Comparative Development: An Empirical Investigation: Comment” <https://economics.mit.edu/sites/default/files/inline-files/Hither%20Thou%20Shalt%20Come>: ‘Overall, Albouy’s “Comment” amounts to a series of objections to our approach. All of these objections, upon closer inspection, are far from compelling, are often unfounded, and prove minor and largely inconsequential for the robustness of our results. The big picture from AJR (2001) remains intact and remarkably robust: Europeans were more likely to move to places that were relatively healthy, and when they moved in larger numbers, they imposed better institutions, which have tended to persist from the colonial period to today.

In my view, this rhetorical position by Acemoglu, Johnson and Robinson is a huge mistake. They should be very glad to drop their IV results from the discussion.

Acemoglu, Johnson and Robinson went down the road of their IV because of the potential criticism that causation is not flowing from perceived security of property to prosperity, but rather that prosperity has lots of effects that lead to perceived security of property. They thought that they could use settler mortality to identify a component of perceived expropriation risk today that was plausibly independent of these reverse causation factors. In my view, the right way to understand what they wound up doing is to say that they regressed current prosperity on what perceived expropriation risk today would be if you knew only about settler mortality in the past and nothing else. The slope of that regression—the regression of prosperity today on what one would expect perceived expropriation risk to be from settler mortality—is the hill they are willing to die on.

Their OLS scatterplot gives them a coefficient β-hat =0.522: a one-unit increase in their security-of-property today index leads to an 0.522-unit increase in prosperity today, a 68.5% increase.

And their IV scatterplot gives them a coefficient β’-hat = 0.944, which means a one-unit increase in what you would think perceived expropriation risk is today based on settler mortality leads to an 157% increase in prosperity today.

In the real world, a country like New Zealand with a log prosperity score of 10 and a perceived security-of-property score of 10 was back before 2000 about 5 times richer than an Egypt with a log prosperity score of 7 and a perceived security-of-property score of 7. AJR’s IV results say that the effect of security-of-property on prosperity ought to be much much bigger than that: it ought to be about 20 times richer. And it would be, if the relationship between security of property leading to investments in physical and human capital and in enterprise and innovation were the only things operating, were there no reverse causation by which prosperity leads to social unrest that threatens the security of property. But there is such social unrest, AJR’s IV results tell us. There is powerful reverse causation: a doubling of prosperity sets in motion societal forces that would, if they were the only things operating, reduce your perceived security-of-property score by 0.4.

Now: any claim that prosperity structurally reduces governance quality contradicts an overwhelming body of evidence across multiple disciplines:

1. The Modernization Hypothesis (Political Science): Lipset (1959) argued the that it was economic development that created the social conditions for democracy and good governance—an educated middle class, urbanization, organizational capacity, and norms of civic participation. Przeworski and Limongi (1997) and Boix and Stokes (2003) found strong empirical support that higher income increases the probability of democratic consolidation. Acemoglu et al. themselves, in later work (Journal of Political Economy, 2008), examine this relationship. The literature is not unanimous, but no serious strand of it argues that prosperity destroys governance quality.

2. Historical Evidence on Expropriation and Revolution: Revolutions, expropriations, and institutional breakdown have historically occurred overwhelmingly in poor countries, not rich ones. The Russian Revolution and the wave of postcolonial nationalizations all arose in contexts of poverty, grievance, and institutional fragility, not in contexts of prosperity. The handful of rich-country institutional collapses (Weimar Germany) are exceptions, and even there the mechanism was economic collapse, not economic success.

3. The Resource Curse Literature: The closest empirical phenomenon to is the resource curse: oil wealth sometimes corrodes governance by enabling authoritarian consolidation without taxation. But this mechanism is specific to extractive resource wealth, not to prosperity in general, and it operates through a very particular political economy channel (Ross, 2001; Robinson, Torvik & Verdier, 2006). It cannot be generalized to a structural negative relationship between income and governance across all former colonies.

4. The Cross-Sectional Evidence: Simply looking at the data: the richest former colonies—Singapore, South Korea, Botswana, Mauritius—tend to have better governance, not worse. The poorest—the DRC, Haiti, South Sudan—tend to have the worst. A relationship is not what the scatter plot shows.

Thus the implied backwards causation from prosperity to poor governance quality is not merely theoretically awkward—it is empirically falsified by the entire body of comparative political economy.

This means the only semi-live explanation of the OLS-vs-IV gap is—if we take it as anything other than the Cauchy distribution arriving at the picnic and getting out its refreshments, as an example of play, stupid games and win stupid prizes—that it is the result of measurement error in governance quality attenuating the OLS estimate and pushing it downward. That is the interpretation AJR themselves favor. They conclude that the instrumental-variables strategy does not primarily correct for reverse causation. Rather, it primarily corrects for attenuation bias from mismeasured institutions.

But taking that governance quality is mismeasured gets Acemoglu, Johnson and Robinson into even worse trouble.

Their argument is that we cannot today see the true institutional quality of governance G. Instead we see a corrupted and noisy measure of institutional quality H. But if we look at just that part of H that is correlated with settler mortality, we recover a better measure of the institutions that matter. In order for this explanation to work, however, settler mortality in the age of imperialism has to (a) be correlated with that part of modern-day institutions that matter for prosperity, the G, while also (b) being uncorrelated with those parts of modern-day institutions that do not matter for prosperity. Settler mortality C must be: 1. relevant, in that it is correlated with the component of observed institutions that genuinely cause prosperity; and 2. selectively orthogonal, uncorrelated with the part of observed institutions counted as good governance that do not in fact matter for prosperity.

And my response to this is simply this: You are putting me on.

There is no way that 1800s soldier mortality gives us a better measure of institutional quality today, than does looking around at everything we can see about institutions and constructing high-information measures.

The right response from Asimov, Johnson, and Robinson to Albouy is not to dig in deeper, is not to die on the hill of a weak instrument IV, but rather to be very grateful that Albouy has pointed out the magnitude of the weak instrument problem and thus provided a Cauchy distribution explanation of the weirdly implausibly and embarrassingly large size of their IV coefficient. They really do not want to be on either of the hills that defending their IV could put them on:

  1. either prosperity has catastrophically destructive effects on governance,

  2. or (ii) colonial-era mortality is a better measure of what matters for prosperity today in modern institutions than modern institutional analyses can produce.

Neither of those is at all defensible.


But Wait! There Is MOAR! Much, Much Moar!!

ADDITIONAL CRITIQUE 1: It’s Human Capital, Not Institutions: Journal of Economic Growth, Glaeser, La Porta, Lopez-de-Silanes, Shleifer 2004: Europeans who settled in healthy places brought themselves—educated, literate people with specific cultural and legal traditions. The AJR IV identifies where Europeans settled, but settler presence = human capital accumulation, not just institution-building. The instrument may be picking up human capital rather than institutional quality.

But: AJR control for fraction of European population directly, and the institutional effects survive. Moreover, human capital and institutions are not cleanly separable — colonists built institutions deliberately. The critique identifies a real confound but doesn’t break the core result.


ADDITIONAL CRITIQUE 2: Geography, Not Institutions: Various, Sachs & al.: Malaria, latitude, and disease burden directly affect economic productivity, not just through the institutional channel. Settler mortality may proxy for the disease environment that directly depresses output today, violating the exclusion restriction.

But: AJR control for latitude, temperature, humidity, malaria prevalence, and more—and the institutional effect survives all of these. The paper’s robustness tables are unusually thorough on exactly this point.


ADDITIONAL CRITIQUE 3: Institutions Don’t Persist That Way: Various, Path Dependence Skeptic Historians & Political Scientists: The persistence mechanism is assumed rather than demonstrated. Colonial institutions were often radically transformed at independence, and many countries have undergone multiple regime changes. Why would 17th-century institutional choices echo so cleanly into 1990s PRS scores?

But: The correlation between early institutions and current institutions is actually quite well-documented empirically. The persistence is real, even if the mechanisms are complex. AJR‘s own later work, i.e., Why Nations Fail, develops the persistence story more carefully.


ADDITIONAL CRITIQUE 4: The Exclusion Restriction Is Untestable: Standard IV Critique: Settler mortality back then affects current income through channels other than “institutions”—by shaping culture, work norms, trust levels, or through direct epidemiological legacies.

But: Untestability is a feature of all IV designs, not a special problem for AJR. And AJR‘s robustness to controlling for an unusually long list of potential direct channels is about as good as IV evidence ever gets.


Bottom line (for teaching): The Albouy data critique is the one that genuinely stings—not because it refutes the finding, but because it raises legitimate questions about fragility. On the other hand, the very large size of the estimated IV coefficient is very embarrassing and suggests that something has gone catastrophically wrong with the analysis. Albouy provides a way out of that very legitimate critique.

The others are important conceptual challenges that AJR largely anticipated and addressed. The paper’s core claim—that colonial history explains a large fraction of income differences today with “institutions” as a primary channel—has proven surprisingly durable; while also remaining puzzling and, in a sense, unbelievable.


Memo to self: <https://datahub.berkeley.edu/user/jbdelong/lab/workspaces/auto-1/tree/working_20251227/2026-04-08-DELIVERED-econ-196-week-9-reversals-of-fortune.ipynb>

Subscribe now

Leave a comment

If reading this gets you Value Above Replacement, then become a free subscriber to this newsletter. And forward it! And if your VAR from this newsletter is in the three digits or more each year, please become a paid subscriber! I am trying to make you readers—and myself—smarter. Please tell me if I succeed, or how I fail…

##⁠wait-what-mr-albouy-reaches-his-conclusion-by-omitting-half-the-data-from-the-original-sample-another-chart-of-the-day
##enlarging-the-bounds-of-human-empire
##public-reason
##
chart-of-the-day
##‎
⁠that-is-simply-wrong
#
wait-what-mr-albouy-reaches-his-conclusion-by-omitting-half-the-data-from-the-original-sample
#colonial-origins
#acemoglu-johnson-robinson
#david-albouy
#weak-instruments
#institutions-and-growth
#reversals-of-fortune
#economic-history
#cross-country-economic-growth

Here in Academia, “AI” Has Arrived for Software Coding, for Summarization & Search, for Writing Prose, & for Assessment: CHART OF THE DAY

Every assignment produces two things: the paper turned in now, and the judgment the student builds for later. AI can raise the first while hollowing out the second—and only one shows up in the grade. So I need to redesign my courses around that fact, and that is only one part—the assessment part—of how “AI” has already come for American higher education:

I am not teaching at all in this forthcoming semester. For reasons that nobody has explained to me and that I see no point in trying to dig into, if I teach anything at all this semester, then: my health insurance turns from the gold-plated grandfathered-in employer-sponsored health insurance of someone hired by the University of California in the 1990s into the much skimpier employer-sponsored health insurance the university currently offers to new temporary employees.

Share

But I am going back into the teaching rotation in the spring of 2027 for Econ 210a and Econ 135: a 25-person and a 75-person class. and I will then have to face a system where information-technological disruption is currently producing outcomes like this:

Give a gift subscription

Via Paul Novosad <https://x.com/DKThomp/status/2090076039141552443>, who writes:

Every college syllabus should include these graphs. Use AI for homework, you will get it done faster and get a higher grade, and then get crushed on the exam…

Share DeLong's Grasping Reality Weblog

And Derek Thompson connects the dots with respect to assessment:

Derek Thompson: ‘Unless testing shifts toward in-person and/or blue books soon, AI is going to turn a whole lot of education into the informational equivalent of "wow I [paid a guy, who] squatted 200 lbs at the gym yesterday, new personal record!"…

Leave a comment

He is correct.

At least in my estimation, every piece of written work that we assign and expect to be turned in needs to have a 10-minute one-on-one with either me or one of the TAs talking and arguing about the document produced. This will do two things:

  • First, it is something that we ought to have been doing all along, but because we are lazy, we developed an educational model with insufficient feedback, engagement, and dialogue.

  • Second, it is the only way to try to avoid the graph pattern above that students, as they ask the AI to do their work, are thinking: I will have to explain this to the professor or the TA next week.

Get 75% off a group subscription

75 students. 1 TA. That is two of us. 10 minutes of assessment/socratic dialogue per student per assignment. That is 12 and a half hours of our time for one assignment. Could we get away motivationally with only checking in a third of the time, and choosing the check-in at random? Maybe that would then be three hours per assignment. What with slack and so forth, that would be one full workweek of the 10 workweeks of time devoted to the course by the teachers. That is, I think, doable.

That is a plan: 10 assignments, 3 check-ins in person per student.

I think I need to stress here that, while this is a substantial increase in workload that will crowd-out other things, this shift in the mode of assessment is not so much a problem as a problemtunity.

First, the something it crowds out is, on inspection, largely the low-value grading of artifacts whose provenance was sometimes uncertain, and in which the feedback loop that would both enable and force student improvement was not closed. We told ourselves for decades that we were providing sufficient feedback in sections and in paper comments, and mostly we were not.

There is a load-bearing confession here: We built the modern large-lecture university course as a machine for economizing on faculty attention. The problem set turned in, the blue book graded, the term paper marked up in the margins and handed back to a student who looked only at the letter at the top. This was never pedagogically optimal. This was faculty effort-minimizing. We dressed up a budget constraint we had constructed for our convenience as a philosophy of education.

So the one-on-one is not a defensive crouch against cheating. Framing it that way concedes the wrong ground—it makes the professor a border guard and the student a smuggler, which is miserable. When I sit across from a student and we argue for ten minutes about the causal story in her paper—why she chose this mechanism and not that one, what evidence would have changed her mind, where the awkward document in the archive doesn’t fit—I am measuring the stock of judgment directly, at the source, rather than inferring it from a proxy that has just been debased. The AI can produce the draft. It cannot sit in the chair and be caught not knowing why the draft says what it says.

Here is the part that makes it a problemtunity rather than a problem. The behavioral effect runs backward through time. A student who knows that next week she will have to explain this to the professor or the TA is, this week, a different writer. Use the machine to lower the cost of the prose and to supply criticism, not to supply the judgment; keep verification in your own hands, because the person in the chair across from you is going to check the chain from claim to evidence. Not “you must write every sentence yourself,” which is a losing and slightly ridiculous war to fight. Rather: “you must be able to think, in real time, about the sentences that appear under your name.”

That is a better commitment device than the blank page ever was, because it targets the muscle and not the ritual.

And it is, frankly, the education I myself did want to have received, and did kinda-sorta. The tutorial—Oxbridge’s expensive glory, which the American research university abandoned as a luxury it could not scale—turns out to be the assessment mode that a world of cheap text forces back upon us.

When AI writes the prose; that frees the hour. Spend the freed hour in the chair, arguing. That is not a tax on teaching. That is teaching, finally unmasked as the thing it was always supposed to be.

One workweek out of fifteen. I’ll take that trade.

Share


And now I have to figure out how to integrate the coming of AI for summarization and search with respect to the reading, with respect to writing prose, and with respect to the programming tasks that I hope to assign come next spring.

Refer a friend


I see this AM that Johan Fourie has some thoughts about prose-writing that are relevant:

Johan Fourie: Writing Is Not Thinking <https://johanfourie.com/files/wp/JF_WritingIsNot_v3.pdf>: ‘Writing technologies lower the cost of producing candidate text much more than they lower the cost of judging which candidate to keep. Text has become cheap. Judgment has not. A model that returns criticism may help with judgment, as Section 4 explains, but the asymmetry drives the results….

  1. Search. Cheaper drafting improves the argument ultimately selected….

  2. Human capital. When generated text replaces the practice that forms judgment, assisted output rises today while unaided capability falls tomorrow. Gaps in performance narrow as gaps in unaided capability widen….

  3. Reference points and defaults. An author who is uncertain… and first reads a generated draft… will find the draft is a plausible default… and stop [thinking] earlier[—too early, in fact]….

  4. Signalling. When fluent prose becomes cheap to produce, English fluency becomes a less reliable signal of quality. Wherever poor English had concealed good evidence… [improving] it makes prose more informative….

  5. External effects. If many authors draw on shared generative defaults, individual gains occur alongside field narrowing: more material to screen, but fewer distinct questions and explanations under study…

Share

If I understand where he is coming from and where he gets to sufficiently:

  • The largest, cleanest gains from AI go to second-language researchers, because the cost of English is unrelated to the quality of their evidence — removing it is pure benefit.

  • The dangerous move is not using AI but using it first: forming the representation is the act that both makes the manuscript original and builds the author’s judgment.

  • The defensible workflow is: own account → model for language and criticism → independent verification; disclosure should name the task delegated, but disclosure is a cheap message, so journals should police the claim-to-evidence chain instead.

With this being his biggest worry, I think:

The historian from the opening still has her awkward document. If she writes her own account first, the model can improve it and lower the cost of its prose. If she reads the model’s account first, the document may never disturb the familiar story. No finished manuscript reveals the difference between those two mornings of work…

And the principal objection to his argument as a whole being the one set out by Deirdre McCloskey back in 1985. As Fourie puts it:

Fluent synthesis may… conceal a broken link between a claim and the world…. Workflows… [have] different access to evidence and different incentives to verify it. [Plus] the framework of this article… faces a deeper objection… stated by… McCloskey (1985) [who] rejected the premise that content and expression are separable. In production terms, scholarship cannot be written as the sum of two functions. An author discovers the details of argument only by writing them out, and in doing so often discovers a flaw in its foundations…


Plus I thought this was nicely put, although calling ten kilos—six miles—a “short distance” run leaves me impressed. When I think “short distance”, I think half a mile:

Johan Fourie: What I Think About When I Think About Writing <https://www.ourlongwalk.com/p/what-i-think-about-when-i-think-about>: ‘I run short distances: usually ten kilometres, in Jonkershoek, in the early morning as the sun rises over the peaks. It’s spectacularly beautiful. These are also the best times to think. I think about… research. My family. Blog post topics. More research ideas…. Why has no one founded a company like that? A bug walking across the road…. I need to remember to renew the family Wildcard, otherwise I’ll have to pay to get into Jonkershoek. Oh, and the car licence….

You get the point. Which is why I find it fascinating when people say that ‘writing is thinking’. No, writing is sitting down and forcing yourself to think about something…. It is not the writing that does the thinking. It is the thinking that does the thinking. You don’t need the writing to do the thinking, but the discipline of the writing helps the thinking, yes…. AI need not displace thinking if we are careful about the sequence….

‘Writing’ is not one activity…. Turning a decided thought into a grammatical sentence is one task. Taking a vague, half-formed, slightly embarrassing impression and forcing it into words explicit enough to inspect is a completely different one. So is choosing which causal story to tell…. [So is] checking whether the archive supports it. Delegating the first is… sometimes a mercy. Delegating the second is where the trouble starts…. If drafting… takes seven hours and checking it against the evidence takes one… but a model drops the drafting to one hour… the tool has quadrupled the number of arguments I can test…. As drafting approaches free, the number of versions I can examine… approaches my own capacity to judge them. That capacity has not improved at all.

That capacity is the paper’s second output. Every piece of writing produces two things: the manuscript you publish now and the judgement you will develop. Judgement is human capital, a stock built by costly practice and lost through disuse….

The subtler loss is the blank page…. Before [AI text] generation… nobody had to design a commitment device for writing. The blank page was one. No effort meant no manuscript…. A generated draft changes the zero-effort outcome from a blank page to a plausible paper….

What matters is not whether you can write every sentence yourself. It is whether you notice something, see why it might matter, follow the idea, and know when the answer is wrong…

Subscribe now

Leave a comment

If reading this gets you Value Above Replacement, then become a free subscriber to this newsletter. And forward it! And if your VAR from this newsletter is in the three digits or more each year, please become a paid subscriber! I am trying to make you readers—and myself—smarter. Please tell me if I succeed, or how I fail…

##here-in-academia-ai-has-arrived-for-software-coding-for-summarization-search-for-writing-prose-for-assessment-chart-of-the-day
##academia
##public-reason
#mamlms
#subturingbradbot
##chart-of-the-day
##text-got-cheap-judgment-didnt-how-i-hope-to-teach-in-2027-ai-has-arrived-for-coding-search-prose-assessment-three-opportunities-and-one-problemtunity
#here-in-academia-ai-has-arrived-for-software-coding-for-summarization-search-for-writing-prose-for-assessment
#ai-in-academia
#writing-is-not-thinking
#mamlm-academic-problemtunity
#socratic-check-in-in-person-assessment
#stopping-thinking-too-soon
#teaching-in-2027
#johan-fourie

CROSSPOST: JOE FRANCIS: Running Out of Patience with Economics

A “top-five” result apparently undone by a single inverted variable and thus a division where there should have been a multiplication: what one coding slip says about economics. Fix the inversion and a thousand-times-cited AER finding loses its significance, its sign, and its story, as Joe Francis spends his time performing the public-good service of making the machinery of knowledge actually work:

Joe Francis is doing more replications.

He finds that Galor and Özak’s claims on “the agricultural origins of time preference” <https://www.aeaweb.org/articles?id=10.1257/aer.20150020> attain statistical significance only because their “area weights” variable is wired into the analysis exactly backward: instead of downweighting small areas that people would find it easy to move away from, they downweight large areas where people are more likely to stick.

You have a prior—a strong, theoretically motivated, Malthusian-selection prior that patience got culturally bred into populations where the agronomic return to waiting was high. You run the specification. It comes out significant and correctly signed. And here is the trap that catches all of us who are even passable Bayesians: you do not audit results that confirm what you already believe. You audit the surprises. Confirmation is where your guard is down. Oded and Ömer divided where they should have multiplied, the number came out the color they expected, and nobody re-derived the weight variable because nobody had a reason to. It happens. It has very nearly happened to me.

Thus: “The estimator gives the largest weight to the smallest regions. A tiny 34 km² Swiss region receives 170,000 times the weight of the largest region. The influence of small, urban regions where internal migration is most severe is thereby maximized, making the problem that area weights were supposed to address worse…”

And once you wire things in the right way, the coefficient loses statistical significance completely and loses its sign: it is not the case that higher potential crop yields and shorter growth cycles that rewarded delayed gratificatication are associated with the development of a more patient human cultural matrix, but rather the reverse.

The broken academic incentive structure is Francis’s deeper target: even a demonstrable error won’t trigger retraction or correction, because the “top five” journals essentially never do so.


CROSSPOST: JOE FRANCIS: Running Out of Patience with Economics

The Poor Rich World
Running Out of Patience with Economics
When doing post-publication peer review, it is easy to become pessimistic about the whole human enterprise of knowledge production. It is easy to run out of patience due to the errors one finds. Furthermore, even if I find a fault in one of the articles published in the “top five” economics journals, I know that it will not be retracted or even corrected…
Read more

<https://poorrichworld.blog/p/running-out-of-patience-with-economics> <https://poorrichworld.blog>

A Strange Case of Inverted Area Weights

Joseph Francis

Aug 18, 2026

When doing post-publication peer review, it is easy to become pessimistic about the whole human enterprise of knowledge production. It is easy to run out of patience due to the errors one finds. Furthermore, even if I find a fault in one of the articles published in the “top five” economics journals, I know that it will not be retracted or even corrected.

The article “The Agricultural Origins of Time Preference” (2016) by Oded Galor and Ömer Özak is a case in point. Published in the American Economic Review, the article presents a grand narrative of persistence. Galor and Özak hypothesize that “geographical variations in the natural return to agricultural investment generated a persistent effect on the distribution of time preference across societies” (p. 3065). Higher potential crop yields and shorter growth cycles, they say, rewarded delayed gratification, so a culture of long-term orientation was selected for over generations and persists in populations today.

The article has garnered over 1,000 citations on Google Scholar, but it is in no way, shape, or form robust.

The evidence for this narrative comes from a simple correlation between individuals in the World Values Survey and historical FAO data on potential crop yields and growth cycles in the places of their ancestral origins. Macro-historical regressions across countries are, however, famously vulnerable to confounding factors such as geography, modern national institutions, or contemporary culture, which can easily drive the correlation.

Galor and Özak therefore run a version of their regressions with country fixed effects. Individuals are compared within each country, and national institutions drop out. In the published article that specification is statistically significant: historical crop yields and patience, across 1,356 regions.

But the problem is that people have moved. The agricultural characteristics of a person’s current location might not then reflect the environment that supposedly shaped their ancestors’ time preference.

They claim to deal with this by using area weights. As they put it, “observations are weighted by the scale of each region, mitigating the effect of internal migration” (p. 3097). Larger regions, on this logic, are less susceptible to internal migration crossing their borders, so assigning them higher weights should reduce this measurement error.

Nonetheless, their code then does the opposite.

For weights, it uses a variable called “invwvsarea.” When that variable is reverse engineered, it turns out that it is in fact area weights inverted. Why they did this, is not clear, but it has the effect of aggravating the issue that they claim to be addressing. By weighting by the inverse of the area, the estimator gives the largest weight to the smallest regions. A tiny 34 km² Swiss region receives 170,000 times the weight of the largest region. The influence of small, urban regions where internal migration is most severe is thereby maximized, making the problem that area weights were supposed to address worse.

When the actual, non-inverted area weights are used instead, the coefficient for crop yield drops from 0.032 (p=0.010) to an insignificant −0.033 (p=0.39), as shown in Table 1. Indeed, all their main results essentially disappear. And the same happens when no weights are used at all: the coefficient falls to 0.001 (p=0.85), indistinguishable from zero. Their results are artifacts of doing the opposite of what is described in the text.

t1_galor_ozak_weights

Yet this probably does not matter. Even though the article is not robust, it will not be retracted, or even corrected. That will not be possible because the “top five” economics journals do not do such things. As the Economist recently observed, we know that things are much better in economics than in the rest of the humanities and the social sciences because the “five leading journals have seen just four withdrawals in their combined 570-year history.” Their peer review is so robust that it would be impossible for an article like this to be published in the first place. Consequently, those area weights cannot really have been inverted.

Further Reading

The article examined in this post is:

Replication Files

The replication package for this post is available here.

<https://poorrichworld.blog/p/running-out-of-patience-with-economics> <https://poorrichworld.blog>

The Poor Rich World
Running Out of Patience with Economics
When doing post-publication peer review, it is easy to become pessimistic about the whole human enterprise of knowledge production. It is easy to run out of patience due to the errors one finds. Furthermore, even if I find a fault in one of the articles published in the “top five” economics journals, I know that it will not be retracted or even corrected…
Read more

Brad DeLong here: I part company with Joe is the counsel of despair—the bleak “it will never be retracted, so none of this matters.” I don’t buy it. Retraction is the wrong metric. The right metric is what the profession believes and builds on, and there the machinery works, albeit maddeningly slowly and expensively.

It is true, to highlight what Joe writes:

When doing post-publication peer review, it is easy to become pessimistic about the whole human enterprise of knowledge production. It is easy to run out of patience due to the errors one finds. Furthermore, even if I find a fault in one of the articles published in the “top five” economics journals, I know that it will not be retracted or even corrected…. [Oded Galor and Ömer Özak’s] results are artifacts of doing the opposite of what is described in the text. Yet this probably does not matter. Even though the article is not robust, it will not be retracted, or even corrected. That will not be possible because the “top five” economics journals do not do such things…

Nevertheless, I think Joe is wrong to despair here. Yes, it looks like Oded and Ömer divided where they should have multiplied—it happens—and did not catch it because the results were as they expected and, as good Bayesians, you don’t spend that much time checking expected results—that also happens. But things that do not replicate are weeded out, albeit much more slowly than they should be. And people do have a very healthy skepticism about results that are close to the edge of statistical or economic significance, or that cannot be shown to rely on a strong solid correlation.

It is not just computational error. It is also that, say, you have ten yes-no decisions to make in running any empirical analysis that could go either way. What you should do is to choose the way that makes most sense to you, and run along several tracks where things are debateable. But there is the temptation to choose the track that gets you closer to the result you think is sensible. Succumb to that, and you wind up with the strongest of 1024 possible results. Alternatively, get yourself checked by somebody hostile for ideological reasons, and they can come up with the weakest of 1024 possible results as a way of dismissing your analysis.

Stepping back, the persistence literature keeps reaching for a lever—settler mortality, ancestral crop yields, ruggedness, caloric suitability—that will let a cross-sectional correlation masquerade as a causal parameter, without a structural model that says how the past reaches into the present.

It turned me into a Heckmanite.

I do not believe you have earned the right to say “causal” unless you can write down the structural model and the story about why the identifying variation is even remotely exogenous. “Deep roots” regressions almost never do. The area weight was a half-hearted patch over the wound that people can and do move to opportunity. Inverting it tore the patch off and rubbed salt in.

But things that do not replicate do get downgraded; they stop anchoring dissertations; the second generation of citations turns sour. It is a shame it takes a decade and a Joe Francis rather than an afternoon and a journal’s own referees. We should be able to do much better.

Which is why we need preregistration and replication packages. And that is why Joe Francis undertaking this mission is a very good thing for the profession as a whole.

Subscribe now

Leave a comment

If reading this gets you Value Above Replacement, then become a free subscriber to this newsletter. And forward it! And if your VAR from this newsletter is in the three digits or more each year, please become a paid subscriber! I am trying to make you readers—and myself—smarter. Please tell me if I succeed, or how I fail…

##crosspost-joe-francis-running-out-of-patience-with-economics
##public-reason
##crosspost
##joes-subheadline-a-strange-case-of-inverted-area-weights-but-i-mostly-want-to-praise-joe-francis-for-spending-so-much-of-his-time-on-post-publication-peer-review
#joe-francis-running-out-of-patience-with-economics
#joe-francis
#running-out-of-patience-with-economics
#post-publication-review
#agricultural-origins
#replication-crisis
#garden-of-forking-paths
#persistence-studies
#heckmanite
#causal-inference
#preregistration
#replication-packages
#economic-history
#confirmation-bias

Bubble Watching: This Is What Happens in the Type of Period Called “Distress”: CHART OF THE DAY

The blow-up of Leopold Aschenbrenner’s Situational Awareness is a sign that we are near the peak of the bubble: this is what it looks like from the inside when demand stops riding on industry fundamentals and starts riding on leveraged-buyer reads of market sentiment. It is a sign that positive-feedback investment strategies are rampant and that demand curves are starting to slope the wrong way:

During an asset market boom, even a euphoric boom, prices rise as good news arrives and as more and more people take their money and decide that this is indeed the wave of the future. But then increases in demand switch from people becoming aware of the opportunity and bringing their money in, to people willing to bet on rising prices, willing to ignore Risk Management 101, and eagerly leveraging-up and pouring that leverage into the booming asset class. Rising prices taken as a reason to buy more induce a situation in which the smart money starts leaving quietly, and the game silently switches from owning value to finding a greater fool before everyone else does.

Share

And then people are on the edge between hanging on hoping a greater fool comes along and selling out now. The shape of demand no longer rides predominantly on expected developments of industry, but on various reads of market sentiment.

That period in which a lot of people are following positive-feedback investment strategies is a period of “distress”.

We know it is distress because then we get things like this from the most overleveraged:

Give a gift subscription

Via Adam Tooze <https://adamtooze.substack.com/p/top-links-1199-big-losses-meeting>, FT Alphaville <https://www.ft.com/content/340bf9e7-0e67-4d19-b671-3dc8186efb99> picking up “a cool table on Wikipedia” <https://en.wikipedia.org/wiki/List_of_trading_losses> based on Tom Coleman <https://rpc.cfainstitute.org/sites/default/files/-/media/documents/book/rf-publication/2011/rf-v2011-n3-1-pdf.pdf>.

The write-up is by Toby Nangle:

Toby Nangle: A Leaderboard of the Biggest Trading Losses of All Time <https://www.ft.com/content/340bf9e7-0e67-4d19-b671-3dc8186efb99>: ‘We found a cool table on Wikipedia.. to more easily contextualise the quantum of Aschenbrenner’s loss…. There are two steps to a fund losing a lot of money. The first step is to inspire faith in either a large number of fairly wealthy people or a small number of immensely wealthy people…. The second step is to do [is]… throw together a credible investment thesis, [and] have sufficiently high conviction… to cast aside risk management 101, maybe chucking a bunch of financial leverage into the mix…. The second step is easy. There are thousands of people yoloing in their mums’ basements around the world doing just this right now…. So if you’re interested in maximising your ranking on any quantitatively measured global leaderboard of trading losses, the first step is probably more important….

We’re fairly sure that this league table, like every other we’ve chanced upon, is only really capturing the kind of meltdowns that make good copy. There’s Tiger Global’s ca. $40bn bloodbath in 2022, for example, which arguably should put it at the top of the list. There’s a case that Jane Street…[belongs] given the reported $15bn hit that it took in July from its exposure to Aschenbrenner…

Leave a comment

That last is a reference to:

Jill R Shah & Joshua Franklin: Jane Street Suffers $15Bn Hit After Meltdown at Situational Awareness: Jane Street posted a roughly $15bn loss in July after turmoil at AI-focused hedge fund Situational Awareness wrongfooted the US trading firm. The New York-based firm disclosed the figure to lenders as part of a deal to shift its roughly $11bn public debt pile to private investors including Pimco…. Jane Street has generated more than $40bn in net trading revenues in the year to Friday, even accounting for the July loss, which exceeds its entire haul for 2025…. Jane Street’s investment in Situational Awareness was unusual because the firm trades its own capital.… a former Jane Street employee worked at Situational Awareness and Jane Street co-founder and partner Robert Granieri attended [Leopold] Aschenbrenner’s wedding in California in recent weeks….

Jane Street was established in 2000 by a small group of founders including Granieri, who previously worked at Pennsylvania-based Susquehanna. It uses technology to make markets across assets such as equities, bonds, exchange traded funds and more. In recent years, it has expanded into longer-term strategies as well as investments in private companies, including AI lab Anthropic and data centre operator CoreWeave…

Share DeLong's Grasping Reality Weblog

Time to pull out the Kindleberger! This time, from A Financial History of Western Europe:

At some stage in the process it becomes clear to a few, and then to more, that the fallacy of composition is at work, that the whole is rather less than the sum of the parts, that credit positions are extended beyond some limit sustainable in the long run, and that maintenance of capital gains depends on getting out of assets rising in price ahead of others.

There follows a period of what may be called ‘distress’: ‘We have no crash at present, only a slight premonitory movement of theground under our feet,’ wrote Lord Overstone to his friend, G. W. Norman, on 1 November 1845 (O’Brien, ed., 1845 [1971], Vol. 1, p. 368). From time to time the distress abates. On other occasions it intensifies. More and more speculators seek to get out of whatever was the object of speculation, to reduce their distended liabilities, and switch into money; and more and more it becomes clear that not everyone can do so at once.

There is a rush, a panic, and a crash—or perhaps a lender of last resort intervenes to make clear that it will furnish the market all the cash it insists it requires. In this circumstance, perhaps belatedly, panic and distress subside…

Get 75% off a group subscription

The phases of the process are: displacement—a technological or a super-political shock that calls forth a need for real economic adjustment and change—diffusion of euphoria as adjustment takes place, distress, crisis, panic, and then—perhaps—a lender of last resort.

Subscribe now

Leave a comment

If reading this gets you Value Above Replacement, then become a free subscriber to this newsletter. And forward it! And if your VAR from this newsletter is in the three digits or more each year, please become a paid subscriber! I am trying to make you readers—and myself—smarter. Please tell me if I succeed, or how I fail…

##⁠bubble-watching-this-is-what-happens-in-the-type-of-period-called-distress-chart-of-the-day
##macro-outlook
##finance

##⁠chart-of-the-day
##‎
⁠according-to-kindleberger-things-like-the-situational-awareness-colllapse-of-leopold-aschenbrenner-happen-near-the-peak-of-a-bubble-they-are-a-sign-of-the-shape-of-demand
#
bubble-watching-this-is-what-happens-in-the-type-of-period-called-distress
#bubble-watching
#kindleberger
#distress
#financial-crisis
#manias-panics-crashes
#leopold-aschenbrenner
#situational-awareness
#jane-street
#positive-feedback-investment-strategies

More Signs of a Bubble Peak: BREAKING NEWS WATCH OF THE DAY

OpenAI added a billion dollars of revenue and three billion dollars of losses in a single quarter, and now has apparently paused at least model training for reasons. When a firm best at generating vibes starts conserving cash instead, the bubble peak is likely to be close. I am surprised: Claude Code and OpenAI Codex are close enough to be a matter of taste, yet Anthropic’s revenue nearly tripled last quarter while OpenAI’s grew only (only!) 18%:

The shift from equity to debt financing of datacenters has been one powerful sign of an approaching bubble peak. Gary Marcus reads OpenAI’s pausing of some aspects as model training as a second sign: husbanding cash now appear to be more important for OpenAI than generating vibes.

Share

As of this morning, August 18, 2026, Gary Marcus thinks that Sam Altman and OpenAI have just thrown in the towel on capturing the frontier-model lead in the next training cycle round:

Gary Marcus: BREAKING: OpenAI’s Unraveling Has Begun <https://garymarcus.substack.com/p/breaking-openais-unraveling-has-begun>: ‘Their planned IPO is facing headwinds, trust has evaporated, and their burn rate is only getting worse…. Sam Altman’s latest announcement, on Tuesday August 18 (i.e., “earlier today” for those of us on the West Coast), that OpenAI would be pausing, ostensibly for safety reasoning…. As far as I can tell, hardly anyone believed him…. Translation from the user @NIK on X: “We are out of compute”…. Slightly elaborated: Boss Hendricks: “Translation: we need to immediately stop torching cash to provide some semblance of a sustainable business model so we can rush this IPO out the door before the bubble pops”….

The Wall Street Journal’s Berber Jin and Corrie Dribusch just dropped big news: “OpenAI told investors its revenue grew by 18% from the first to the second quarter while its losses deepened… [growing] by $3 billion from q1 to q2, to $12.3 billion. Not a great look given that it added only $1 billion in revenue in the meantime, to $6.7 billion”…. Good news for Anthropic…. And terrible news for OpenAI…

Share DeLong's Grasping Reality Weblog

We had, four months ago:

Nilay Patel: The Ai Industry’s Race for Profits Is Now Existential <https://www.theverge.com/podcast/909042/ai-monetization-cliff-anthropic-openai-profitable-ai-existential-moment>: ‘It’s a make-or-break year for Anthropic and OpenAI, which are facing more pressure than ever to make more cash than they burn…. Hayden Field… senior AI reporter here at The Verge… has been keeping close tabs on both Anthropic and OpenAI…. At some point, the profits have to materialize, or the bubble pops…. You’ve heard me ask some version of this question to scores of CEOs here on this show, and a majority of them have hinted toward the bubble popping — they think some companies will fail in spectacular fashion, some will succeed, and the opportunities, especially the money, are simply too big to ignore. We’re doing this, whether we want to or not — the market depends on it….

AI agents… have radically changed how these companies are thinking about their resources…. Agents are valuable to [code-writing] customers right now, but agents also use far more compute… burning tokens at a rate way faster than these companies anticipated…. OpenAI abruptly killed its video-generation app Sora, ditching a $1 billion Disney licensing deal in the process. Why? It costs too much to run, and OpenAI needs the compute for Codex. We saw it again just last week, when Anthropic decided it would no longer let Claude users burn through compute resources using the OpenClaw agent framework through a standard subscription plan…. The projections these companies have made, which just this week were leaked to the Wall Street Journal, tell a story of mind-boggling growth, to the tune of hundreds of billions in revenue and profitability by the end of the decade. But the most important questions now are can the AI companies pull this off, and what compromises will they make to reach that goal and avoid crashing and burning?…

Give a gift subscription

Following up on his four months earlier prediction: “as nuclear as it gets: OpenAI fails [in 2026]”:

33:31

Get 75% off a group subscription


Even in this context OpenAI’s apparent training pause does surprise me. I had thought that right now OpenAI Codex on the one hand and Claude Code and Cowork on the other were about equal. Some preferred one. Some preferred the other. It seemed largely, these days, a matter of taste and path-dependency. Claude Code wins on code quality, context retention, and multi-agent orchestration; Codex wins on speed, token efficiency, cost-per-task, and fire-and-forget autonomy. Most heavy users run both. And as far as Codex and Claude Cowork are concerned, that race is just beginning, and is too close to call.

And yet it looks like OpenAI’s revenue only (only!) grew from $5.7 to $6.7 billion from the first to the second quarter, while Anthropic’s revenue grew from $4.7 to $11.5 billion. How is this, if Claude Code and Cowork and Codex are rough peers?

Well, first, they were not rough peers at the start of the quarter, on April 1. So perhaps the Q2-Q3 comparison will look very different from the Q1-Q2 comparison.

However, otherwise: The coding tool is a much bigger slice of Anthropic than of OpenAI. Roughly 80% of Anthropic’s revenue is API/enterprise. Anthropic sells agent usage metered by token through its enterprise/API tier, while Codex is delivered inside a ChatGPT Plus/Pro subscription. And OpenAI wants to build loyalty and so does not want to do what Anthropic did to its OpenClaw enthusiasts by cutting off their access. More important, perhaps: Anthropic is the picks-and-shovels supplier to the whole coding-agent ecosystem, not just Claude Code. Claude is the model behind a large share of third-party coding front-ends. with GAAP numbers due in the IPO prospectus.

What I dearly wish to see right now is Anthropic’s S-1 for its forthcoming IPO, which is coming—sometime. FutureSearch “founded in August 2023 by Dan Schwarz… [as] an AI that could predict the future” claims <https://futuresearch.ai/app/p/a/on-what-date-will-anthropic-complete-its-ipo-i> November 4 as the likely IPO date (which means the GAAP financials need to appear in less than a month and a half). Itd further claims:

Three significant overhangs threaten to push the timeline…. Gross-versus-net revenue accounting disputes are highly scrutinized by the SEC; Anthropic reportedly books cloud partnership revenue gross…. Resolving this could require extensive disclosure changes or restatements of ARR metrics. Second, ongoing litigation… over an unprecedented “supply chain risk” designation limits U.S. military contracting and requires sensitive, unresolved risk-factor disclosures . Finally, the sought-after $2T+ valuation demands a staggering $190–200B revenue forecast for 2028…

Leave a comment

My read (which may be very wrong):

  1. An Anthropic that comes out of the gate with a $2 trillion market valuation at its IPO is something that does not need the U.S. Defense Department as much as the U.S. Defense Department needs it. Given that standard risk disclosures are simply boilerplate, and FutureSearch is highly likely to be simply wrong here.

  2. People who are going to buy Anthropic at the IPO can be easily directed to pro forma financials. The actual GAAP financials will be of relevance only to short sellers who are going to be on the sidelines unless they are stark raving mad given what we have seen over the past decade. FutureSearch is highly likely to be simply wrong here as well: divergence between pro forma and GAAP is also not holding up the IPO.

  3. What is, I think, very likely to be holding up the IPO is that a number of organizations that are putting their and their clients’ money on the line here want more than just the second quarter of super-explosive growth from a relatively high base.

Share

Recall that Anthropic’s reported revenue figures are $0.8B for 2025Q2, $4.73B for 2026Q1, and $11.5B for 2026Q2, that those come from leaks or investor-deck slides, and that that is all we know. Other leaks and announcements have been “run rates”:

  • $1B ARR as of the start of the 2025.

  • $9B ARR as of the end of 2025.

  • $65B ARR as of August 1, 2026.

Give a gift subscription

My bet are that these are one-month run-rates at best: i.e., Anthropic booked $5.4B in revenue for the month of July 2026. People will want to see revenue on-track on an S-curve to triple from mid-2026 to moderately late-2028. That means they want to see August and September numbers, that they want those numbers to be good, and that Anthropic is willing to push off the IPO and bet that it can deliver those numbers, rather than scale back the whispered $2T IOP valuation.

Subscribe now

Leave a comment

If reading this gets you Value Above Replacement, then become a free subscriber to this newsletter. And forward it! And if your VAR from this newsletter is in the three digits or more each year, please become a paid subscriber! I am trying to make you readers—and myself—smarter. Please tell me if I succeed, or how I fail…

##more-signs-of-a-bubble-peak-breaking-news-watch-of-the-day
##macro-outlook
##mamlms
##subturingbradbot
##breaking-news-watch-of-the-day
##openai-the-company-that-had-been-selling-vibes-looks-like-it-is-starting-to-husband-its-cash-in-the-face-of-the-anthropic-challenge-two-peer-coding-agents-and-yet-a-55x-revenue-growthrate-gap
#more-signs-of-a-bubble-peak
#bubble-peak
#ai-bubble
#openai
#anthropic
#coding-agents
#anthropic-ipo

HOISTED FROM PAUL KRUGMAN’S ARCHIVES: PAUL KRUGMAN (2012): Looking Back With Shrillness

We told you so. Repeatedly. Since 1993 and Newt Gingrich’s takeover of the Republican Party and Bob Dole’s decision that he would rather go along with trying 100% to portray Bill Clinton as a failed president rather than work to make America a better place. (I remember arriving in DC, and finding two weeks later Paul GIgot in the “Wall Street Journal” telling flat-out lies about Treasury Department, that is, my personal, spread of estimates of the economic impact of tax increases.) And we got given a lot of s*** for it. The GOP did not break in 2025. It did not break in 2017. It did not even break in 2012. It broke in 1993:

In this 2012 post, Paul Krugman looks back on more than a decade of being dismissed as “shrill.” His offense was accuracy: he insisted early that George W. Bush was a serial liar pursuing a hard-line agenda, not the blunt honest conservative the commentariat described. He also refused the mandatory pretense of symmetry — the fiction that Democratic caution and Republican radicalism were morally equivalent. By 2012, with a primary field of “not-Romneys” he calls stark raving mad, Krugman argues the party’s condition was decades old. The only new thing was that the pretense of reasonableness had become impossible to sustain:

Share


HOISTED FROM PAUL KRUGMAN’S ARCHIVES: PAUL KRUGMAN (2012): Looking Back With Shrillness

<https://archive.nytimes.com/krugman.blogs.nytimes.com/2012/02/29/looking-back-with-shrillness/>

Paul Krugman - New York Times Blog

Brad DeLong notes that the GOP we now see in the primaries has been building for a couple of decades; I can’t help thinking of my own decade-plus in the journalistic trenches.

Early on in my tenure at the Times, I felt I had no choice but to point out the inconvenient truth that the official line of the commentariat was all wrong. George W. Bush was not a nice, blunt, honest guy who happened to be a conservative; he was a serial liar pursuing a hard-line agenda, who among other things deliberately misled America into war.

For this I was labeled “shrill”.

More than that: throughout these past ten-plus years, it has been considered ill-mannered and uncouth, not to mention unacceptably partisan, to suggest that the parties aren’t symmetric — that, for example, the reluctance of Democrats to cut Social Security and Medicare is not equivalent to the GOP’s consistent pursuit of huge unfunded tax cuts, that the occasional desire of Democrats to put evidence in a more favorable light is not equivalent to the constant, raw dishonesty emanating from the right. And pundits in good standing have been expected to make calls for bipartisanship that involve pretending that Republican politicians are actually the kind of statesmen the party used to contain, but no longer does.

So now we see a primary struggle in which the choice is between a series of not-Romneys whose political and policy views are stark raving mad, on one side, and the not-not-Romney who is, maybe, just pretending to share those views. How did that happen?

The answer, as Brad suggests, is that it happened a long time ago. The GOP isn’t just spectacularly unlucky in its menu of candidates; this is what the party has been for decades. Rick Santorum isn’t someone out of left field; he’s always been what you see now, and he was a central figure in his Senate days.

All that has happened now is that the mannerisms have finally gotten to the point that the pretense of a reasonable party is no longer sustainable.

But you weren’t supposed to notice until just about now.

<https://archive.nytimes.com/krugman.blogs.nytimes.com/2012/02/29/looking-back-with-shrillness/>


Brad DeLong here: The thing to understand about Paul Krugman’s 2012 “Looking Back With Shrillness” is that it was not a complaint a much as a diagnosis — and, more than that, a vindication delivered in the flat, tired voice of a man who had been right for a decade, and had gotten little for it but the label “shrill”.

Go back and read what Krugman actually wrote in February 2012. The proximate occasion was a post of mine — “Where Were You in 1993, David Brooks?” — in which I had made the boring, documentable point that the Republican Party then staggering through its clown-car primary had not suddenly gone mad. It had been building toward exactly this for a couple of decades. Paul read that and did what Paul does: he generalized it, sharpened it, and turned it into an indictment not merely of the GOP but of the entire respectable commentariat that had spent ten years insisting the emperor was fully clothed.

His confession — and it reads like one — was that “early on in my tenure at the Times, I felt I had no choice but to point out the inconvenient truth that the official line of the commentariat was all wrong. George W. Bush was not a nice, blunt, honest guy who happened to be a conservative; he was a serial liar pursuing a hard-line agenda, who among other things deliberately misled America into war. For this I was labeled ‘shrill.’”

Now. I have some standing here, because I was a charter member of the order. The Ancient, Hermetic, and Occult Order of the Shrill was not, in its original conception, a club of the angry. It was a club of the accurate.

Give a gift subscription

Read more