Here in Academia, “AI” Has Arrived for Software Coding, for Summarization & Search, for Writing Prose, & for Assessment: CHART OF THE DAY
Every assignment produces two things: the paper turned in now, and the judgment the student builds for later. AI can raise the first while hollowing out the second—and only one shows up in the grade. So I need to redesign my courses around that fact, and that is only one part—the assessment part—of how “AI” has already come for American higher education:
I am not teaching at all in this forthcoming semester. For reasons that nobody has explained to me and that I see no point in trying to dig into, if I teach anything at all this semester, then: my health insurance turns from the gold-plated grandfathered-in employer-sponsored health insurance of someone hired by the University of California in the 1990s into the much skimpier employer-sponsored health insurance the university currently offers to new temporary employees.
But I am going back into the teaching rotation in the spring of 2027 for Econ 210a and Econ 135: a 25-person and a 75-person class. and I will then have to face a system where information-technological disruption is currently producing outcomes like this:
Via Paul Novosad <https://x.com/DKThomp/status/2090076039141552443>, who writes:
Every college syllabus should include these graphs. Use AI for homework, you will get it done faster and get a higher grade, and then get crushed on the exam…
And Derek Thompson connects the dots with respect to assessment:
Derek Thompson: ‘Unless testing shifts toward in-person and/or blue books soon, AI is going to turn a whole lot of education into the informational equivalent of "wow I [paid a guy, who] squatted 200 lbs at the gym yesterday, new personal record!"…
He is correct.
At least in my estimation, every piece of written work that we assign and expect to be turned in needs to have a 10-minute one-on-one with either me or one of the TAs talking and arguing about the document produced. This will do two things:
First, it is something that we ought to have been doing all along, but because we are lazy, we developed an educational model with insufficient feedback, engagement, and dialogue.
Second, it is the only way to try to avoid the graph pattern above that students, as they ask the AI to do their work, are thinking: I will have to explain this to the professor or the TA next week.
75 students. 1 TA. That is two of us. 10 minutes of assessment/socratic dialogue per student per assignment. That is 12 and a half hours of our time for one assignment. Could we get away motivationally with only checking in a third of the time, and choosing the check-in at random? Maybe that would then be three hours per assignment. What with slack and so forth, that would be one full workweek of the 10 workweeks of time devoted to the course by the teachers. That is, I think, doable.
That is a plan: 10 assignments, 3 check-ins in person per student.
I think I need to stress here that, while this is a substantial increase in workload that will crowd-out other things, this shift in the mode of assessment is not so much a problem as a problemtunity.
First, the something it crowds out is, on inspection, largely the low-value grading of artifacts whose provenance was sometimes uncertain, and in which the feedback loop that would both enable and force student improvement was not closed. We told ourselves for decades that we were providing sufficient feedback in sections and in paper comments, and mostly we were not.
There is a load-bearing confession here: We built the modern large-lecture university course as a machine for economizing on faculty attention. The problem set turned in, the blue book graded, the term paper marked up in the margins and handed back to a student who looked only at the letter at the top. This was never pedagogically optimal. This was faculty effort-minimizing. We dressed up a budget constraint we had constructed for our convenience as a philosophy of education.
So the one-on-one is not a defensive crouch against cheating. Framing it that way concedes the wrong ground—it makes the professor a border guard and the student a smuggler, which is miserable. When I sit across from a student and we argue for ten minutes about the causal story in her paper—why she chose this mechanism and not that one, what evidence would have changed her mind, where the awkward document in the archive doesn’t fit—I am measuring the stock of judgment directly, at the source, rather than inferring it from a proxy that has just been debased. The AI can produce the draft. It cannot sit in the chair and be caught not knowing why the draft says what it says.
Here is the part that makes it a problemtunity rather than a problem. The behavioral effect runs backward through time. A student who knows that next week she will have to explain this to the professor or the TA is, this week, a different writer. Use the machine to lower the cost of the prose and to supply criticism, not to supply the judgment; keep verification in your own hands, because the person in the chair across from you is going to check the chain from claim to evidence. Not “you must write every sentence yourself,” which is a losing and slightly ridiculous war to fight. Rather: “you must be able to think, in real time, about the sentences that appear under your name.”
That is a better commitment device than the blank page ever was, because it targets the muscle and not the ritual.
And it is, frankly, the education I myself did want to have received, and did kinda-sorta. The tutorial—Oxbridge’s expensive glory, which the American research university abandoned as a luxury it could not scale—turns out to be the assessment mode that a world of cheap text forces back upon us.
When AI writes the prose; that frees the hour. Spend the freed hour in the chair, arguing. That is not a tax on teaching. That is teaching, finally unmasked as the thing it was always supposed to be.
One workweek out of fifteen. I’ll take that trade.
And now I have to figure out how to integrate the coming of AI for summarization and search with respect to the reading, with respect to writing prose, and with respect to the programming tasks that I hope to assign come next spring.
I see this AM that Johan Fourie has some thoughts about prose-writing that are relevant:
Johan Fourie: Writing Is Not Thinking <https://johanfourie.com/files/wp/JF_WritingIsNot_v3.pdf>: ‘Writing technologies lower the cost of producing candidate text much more than they lower the cost of judging which candidate to keep. Text has become cheap. Judgment has not. A model that returns criticism may help with judgment, as Section 4 explains, but the asymmetry drives the results….
Search. Cheaper drafting improves the argument ultimately selected….
Human capital. When generated text replaces the practice that forms judgment, assisted output rises today while unaided capability falls tomorrow. Gaps in performance narrow as gaps in unaided capability widen….
Reference points and defaults. An author who is uncertain… and first reads a generated draft… will find the draft is a plausible default… and stop [thinking] earlier[—too early, in fact]….
Signalling. When fluent prose becomes cheap to produce, English fluency becomes a less reliable signal of quality. Wherever poor English had concealed good evidence… [improving] it makes prose more informative….
External effects. If many authors draw on shared generative defaults, individual gains occur alongside field narrowing: more material to screen, but fewer distinct questions and explanations under study…
If I understand where he is coming from and where he gets to sufficiently:
The largest, cleanest gains from AI go to second-language researchers, because the cost of English is unrelated to the quality of their evidence — removing it is pure benefit.
The dangerous move is not using AI but using it first: forming the representation is the act that both makes the manuscript original and builds the author’s judgment.
The defensible workflow is: own account → model for language and criticism → independent verification; disclosure should name the task delegated, but disclosure is a cheap message, so journals should police the claim-to-evidence chain instead.
With this being his biggest worry, I think:
The historian from the opening still has her awkward document. If she writes her own account first, the model can improve it and lower the cost of its prose. If she reads the model’s account first, the document may never disturb the familiar story. No finished manuscript reveals the difference between those two mornings of work…
And the principal objection to his argument as a whole being the one set out by Deirdre McCloskey back in 1985. As Fourie puts it:
Fluent synthesis may… conceal a broken link between a claim and the world…. Workflows… [have] different access to evidence and different incentives to verify it. [Plus] the framework of this article… faces a deeper objection… stated by… McCloskey (1985) [who] rejected the premise that content and expression are separable. In production terms, scholarship cannot be written as the sum of two functions. An author discovers the details of argument only by writing them out, and in doing so often discovers a flaw in its foundations…
Plus I thought this was nicely put, although calling ten kilos—six miles—a “short distance” run leaves me impressed. When I think “short distance”, I think half a mile:
Johan Fourie: What I Think About When I Think About Writing <https://www.ourlongwalk.com/p/what-i-think-about-when-i-think-about>: ‘I run short distances: usually ten kilometres, in Jonkershoek, in the early morning as the sun rises over the peaks. It’s spectacularly beautiful. These are also the best times to think. I think about… research. My family. Blog post topics. More research ideas…. Why has no one founded a company like that? A bug walking across the road…. I need to remember to renew the family Wildcard, otherwise I’ll have to pay to get into Jonkershoek. Oh, and the car licence….
You get the point. Which is why I find it fascinating when people say that ‘writing is thinking’. No, writing is sitting down and forcing yourself to think about something…. It is not the writing that does the thinking. It is the thinking that does the thinking. You don’t need the writing to do the thinking, but the discipline of the writing helps the thinking, yes…. AI need not displace thinking if we are careful about the sequence….
‘Writing’ is not one activity…. Turning a decided thought into a grammatical sentence is one task. Taking a vague, half-formed, slightly embarrassing impression and forcing it into words explicit enough to inspect is a completely different one. So is choosing which causal story to tell…. [So is] checking whether the archive supports it. Delegating the first is… sometimes a mercy. Delegating the second is where the trouble starts…. If drafting… takes seven hours and checking it against the evidence takes one… but a model drops the drafting to one hour… the tool has quadrupled the number of arguments I can test…. As drafting approaches free, the number of versions I can examine… approaches my own capacity to judge them. That capacity has not improved at all.
That capacity is the paper’s second output. Every piece of writing produces two things: the manuscript you publish now and the judgement you will develop. Judgement is human capital, a stock built by costly practice and lost through disuse….
The subtler loss is the blank page…. Before [AI text] generation… nobody had to design a commitment device for writing. The blank page was one. No effort meant no manuscript…. A generated draft changes the zero-effort outcome from a blank page to a plausible paper….
What matters is not whether you can write every sentence yourself. It is whether you notice something, see why it might matter, follow the idea, and know when the answer is wrong…
