DeLong's Grasping Reality Weblog

DeLong's Grasping Reality Weblog

SubTuringBradBot

Exploring Truly Useful Use Cases for "AI"?: QUESTION OF THE DAY

Can I get it to teach me things I do not already know?; or, experimenting with "attention" on Bosworth Field on August 22, 1485...

Brad DeLong's avatar
Brad DeLong
Aug 05, 2026
∙ Paid
An experiment with trying to get the LLM to teach me things I do not already know, with unsettling conclusions: I think I got it to successfully teach me things about the architecture and functioning of LLMs. But it confesses that all of its numbers are made-up—that the calculations it said it did it did not in fact do? And SubStack x Pangram cannot tell that a 90% AI-generated post is AI-generated? Do I actually know and understand more? Or am I just being deceived in thinking that I do?

Much these days in economic analysis, in pedagogy, in academia, in the sphere of public reason, and in the future of civilization hinges on the answer to a single question: just what are these things truly useful for?

Share

We know that some of the answer is that these things are useful for:

  • Searching with natural-language prompts, yes (although I am still not sure whether it’s excellent at search is due to the superior affordances of natural language as opposed to keywords for my brain, or whether it is because Google has allowed itself to be thoroughly corrupted by SEO in a way that there has not yet been a parallel MEO—Model-Engine Optimization—corruption).

  • Summarizing.

  • Boilerplate and ritual.

  • Coding assistants, in that one no longer needs to wrestle with Stack-Overflow threads or have one’s ORA books with the cute pictures of weird animals on the cover at one’s hand.

  • Coding and other agents, that these things can be Clever Hanses at scale. They try a thousand things in the time it would take me to try one, and, when properly harnessed to reward success and drop failure, can do a lot. When properly harnessed.

Give a gift subscription

But what else is there—beyond protein folding, medical screening, customer support, ambient scribes (which still miss “nots” much too often), weather forecasting, Herculaneum scrolls, and so on?

Get 75% off a group subscription


Specifically I find myself this AM interested in this: Can “AI” help me understand things I do not now understand? So I started exploring. And after several false starts, I wound up setting my basecamp up at this sentence fragment:

Richard III was the last Yorkist King of England; the army of his cousin Henry Tudor…

Refer a friend

for I wound up wanting to see both what an LLM does with it, and whether the LLM can—properly nudged—come up with an explanation of what it is doing that is useful to me.

What does useful to me mean here? It means:

  • just beyond my grasp,

  • something I definitely do not have nailed,

  • but something where it can write something I more than half-understand,

  • but could not teach,

  • and definitely need to think about lots more.

Share DeLong's Grasping Reality Weblog

So I prompted it, iterated it, attempted to adjust the level of detail of its explanation, and saw what it spat out:

Richard III was the last Yorkist King of England; the army of his cousin Henry Tudor defeated him at Bosworth Field in 1485.

Details below the paywall fold (for now). The take-aways (I think) here:


First: I note that while this post is 90% AI-generated, the SubStack-Pangram judgment is:

That in itself is very interesting to see. My guess as to what is going on? That SubStack x Pangram seems to be tracking a particularly literacy-tech use case of “AI”, rather than successfully detecting AI created documents in general.

Leave a comment

Second, the explainer below does seem to be architecturally exact. I think I find it pedagogically effective. The architecture, matrix shapes, and sequence of operations are correct, and the filing-cabinet analogy carries real load.

But, third, it warns me that every number in it was invented: there never was a calculation that produced a 55% chance of choosing defeated as the next token.. The post ends unresolved, which is the honest place for it to end.

Forth, my guesses right now as to the usefulness of LLMs circle around the idea that all the real use-cases share the property that the output is very cheap to check. Thus the interesting question to me is what happens when you cannot check, hence this visit of mine to Bosworth Field. And I do not know what to think. The LLM’s own “honesty” about fabricating its own numbers seems to me more useful than a confident verdict would have been.

What are the take-aways here? Perhaps these:

Academically: The LLM’s piece is a clean case study in the epistemics of machine-assisted learning — the failure is neither hallucination in the usual sense nor error in the architecture, but plausible synthetic data inside a correct frame.

For public reason: the Pangram result undercuts the assumption that AI authorship can be policed by detection.

For action: one operational rule that falls out is not to trust anywhere, but rather verify; a second is to lean into Clever Hans: have it write an explainer at ten different levels, and then hone in on where the user’s current understanding frontier exactly lies.

When I ran out of time, this is where its final answer was:

SubTuringBradBot: A walkthrough of what happens at the final token of “Richard III was the last Yorkist King of England; the army of his cousin Henry Tudor…”: tokenization, the scaled dot-product at one head, four heads doing different jobs, the logit distribution, and the rollout to Bosworth. Built on a filing-cabinet analogy — Query is the question you walk in with, Key is the label on the drawer, Value is the papers inside.

User's avatar

Continue reading this post for free, courtesy of Brad DeLong.

Or purchase a paid subscription.
© 2026 J. Bradford DeLong · Privacy ∙ Terms ∙ Collection notice
Start your SubstackGet the app
Substack is the home for great culture