π Read this awesome post from Hacker News π
π Category:
π‘ Key idea:

Look, I donβt know if AI is gonna kill us or make us all rich or whatever, but I do know weβve got the wrong metaphor.
We want to understand these things as people. When you type a question to ChatGPT and it types back the answer in complete sentences, it feels like there must be a little guy in there doing the typing. We get this vivid sense of βitβs alive!!β, and we activate all of the mental faculties we evolved to deal with fellow humans: theory of mind, attribution, impression management, stereotyping, cheater detection, etc.
We canβt help it; humans are hopeless anthropomorphizers. When it comes to perceiving personhood, weβre so trigger-happy that we can see the Virgin Mary in a grilled cheese sandwich:

A human face in a slice of nematode:
And an old man in a bunch of poultry and fish atop a pile of books:
Apparently, this served us well in our evolutionary historyβmaybe itβs so important not to mistake people for things that we err on the side of mistaking things for people. This is probably why weβre so willing to explain strange occurrences by appealing to fantastical creatures with minds and intentions: everybody in town is getting sick because of WITCHES, you canβt see the sun right now because A WOLF ATE IT, the volcano erupted because GOD IS MAD. People who experience sleep paralysis sometimes hallucinate a demon-like creature sitting on their chest, and one explanation is that the subconscious mind is trying to understand why the body canβt move, and instead of coming up with βIβm still in REM sleep so thereβs not enough acetylcholine in my brain to activate my primary motor cortexβ, it comes up with βBIG DEMON ON TOP OF MEβ.
This is why the past three years have been so confusingβthe little guy inside the AI keeps dumbfounding us by doing things that a human wouldnβt do. Why does he make up citations when he does my social studies homework? How come he can beat me at Go but he canβt tell me how many βrβs are in the word βstrawberryβ? Why is he telling me to put glue on my pizza?
Trying to understand LLMs by using the rules of human psychology is like trying to understand a game of Scrabble by using the rules of Pictionary. These things donβt act like people because they arenβt people. I donβt mean that in the deflationary way that the AI naysayers mean it. They think denying humanity to the machines is a well-deserved insult; I think itβs just an accurate description. As long we try to apply our person perception to artificial intelligence, weβll keep being surprised and befuddled.
We are in dire need of a better metaphor. Hereβs my suggestion: instead of seeing AI as a sort of silicon homunculus, we should see it as a bag of words.
An AI is a bag that contains basically all words ever written, at least the ones that could be scraped off the internet or scanned out of a book. When users send words into the bag, it sends back the most relevant words it has. There are so many words in the bag that the most relevant ones are often correct and helpful, and AI companies secretly add invisible words to your queries to make this even more likely.
This is an oversimplification, of course. But itβs also surprisingly handy. For example, AIs will routinely give you outright lies or hallucinations, and when youβre like βUhh hey that was a lieβ, they will immediately respond βOh my god Iβm SO SORRY!! I promise Iβll never ever do that again!! Iβm turning over a new leaf right now, nothing but true statements from here onβ and then they will literally lie to you in the next sentence. This would be baffling and exasperating behavior coming from a human, but itβs very normal behavior coming from a bag of words. If you toss a question into the bag and the right answer happens to be in there, thatβs probably what youβll get. If itβs not in there, youβll get some related-but-inaccurate bolus of sentences. When you accuse it of lying, itβs going to produce lots of words from the βIβve been accused of lyingβ part of the bag. Calling this behavior βmaliciousβ or βerraticβ is misleading because itβs not behavior at all, just like itβs not βbehaviorβ when a calculator multiplies numbers for you.
βBag of wordsβ is a also a useful heuristic for predicting where an AI will do well and where it will fail. βGive me a list of the ten worst transportation disasters in North Americaβ is an easy task for a bag of words, because disasters are well-documented. On the other hand, βWho reassigned the species Brachiosaurus brancai to its own genus, and when?β is a hard task for a bag of words, because the bag just doesnβt contain that many words on the topic. And a question like βWhat are the most important lessons for life?β wonβt give you anything outright false, but it will give you a bunch of fake-deep pablum, because most of the text humans have produced on that topic is, no offense, fake-deep pablum.
When you forget that an AI is just a big bag of words, you can easily slip into acting like itβs an all-seeing glob of pure intelligence. For example, I was hanging with a group recently where one guy made everybody watch a video of some close-up magic, and after the magician made some coins disappear, he exclaimed, βI asked ChatGPT how this trick works, and even it didnβt know!β as if this somehow made the magic extra magical. In this personβs model of the world, we are all like shtetl-dwelling peasants and AI is like our Rabbi Hillel, the only learned man for 100 miles. If Hillel canβt understand it, then it must be truly profound!
If that guy had instead seen ChatGPT as a bag of words, he would have realized that the bag probably doesnβt contain lots of detailed descriptions of contemporary coin tricks. After all, magicians make money from performing and selling their tricks, not writing about them at length on the internet. Plus, magic tricks are hard to describeββHe had three quarters in his hand and then it was two pennies!ββso youβre going to have a hard time prompting the right words out of the bag. The coin trick is not literally magic, and neither is the bag of words.
The βbag of wordsβ metaphor can also help us guess what these things are gonna do next. If you want to know whether AI will get better at something in the future, just ask: βcan you fill the bag with it?β For instance, people are kicking around the idea that AI will replace human scientists. Well, if you want your bag of words to do science for you, you need to stuff it with lots of science. Can we do that?
When it comes to specific scientific tasks, yes, we already can. If you fill the bag with data from 170,000 proteins, for example, itβll do a pretty good job predicting how proteins will fold. Fill the bag with chemical reactions and it can tell you how to synthesize new molecules. Fill the bag with journal articles and then describe an experiment and it can tell you whether anyone has already scooped you.
All of that is cool, and I expect more of it in the future. I donβt think weβre far from a bag of words being able to do an entire low-quality research project from beginning to endβcoming up with a hypothesis, designing the study, running it, analyzing the results, writing them up, making the graphs, arranging it all on a poster, all at the click of a buttonβbecause weβve got loads of low-quality science to put in the bag. If you walk up and down the poster sessions at a psychology conference, you can see lots of first-year PhD students presenting studies where they seemingly pick some semi-related constructs at random, correlate them, and print out a p-value (βDoes self-efficacy moderate the relationship between social dominance orientation and system-justifying beliefs?β). A bag of words can basically do this already; you just need to give it access to an online participant pool and a big printer.
But science is a strong-link problem; if we produced a million times more crappy science, weβd be right where we are now. If we want more of the good stuff, what should we put in the bag? You could stuff the bag with papers, but some of them are fraudulent, some are merely mistaken, and all of them contain unstated assumptions that could turn out to be false. And theyβre usually missing key informationβthey donβt share the data, or they donβt describe their methods in adequate detail. Markus Strasser, an entrepreneur who tried to start one of those companies thatβs like βweβll put every scientific paper in the bag and then ??? and then profitβ, eventually abandoned the effort, saying that βclose to nothing of what makes science actually work is published as text on the web.β
Hereβs one way to think about it: if there had been enough text to train an LLM in 1600, would it have scooped Galileo? My guess is no. Ask that early modern ChatGPT whether the Earth moves and it will helpfully tell you that experts have considered the possibility and ruled it out. And thatβs by design. If it had started claiming that our planet is zooming through space at 67,000mph, its dutiful human trainers would have punished it: βBad computer!! Stop hallucinating!!β
In fact, an early 1600s bag of words wouldnβt just have the right words in the wrong order. At the time, the right words didnβt exist. As the historian of science David Wootton points out, when Galileo was trying to describe his discovery of the moons of Jupiter, none of the languages he knew had a good word for βdiscoverβ. He had to use awkward circumlocutions like βI saw something unknown to all previous astronomers before meβ. The concept of learning new truths by looking through a glass tube would have been totally foreign to an LLM of the early 1600s, as it was to most of the people of the early 1600s, with a few notable exceptions.
You would get better scientific descriptions from a 2025 bag of words than you would from a 1600 bag of words. But both bags might be equally bad at producing the scientific ideas of their respective futures. Scientific breakthroughs often require doing things that are irrational and unreasonable for the standards of the time and good ideas usually look stupid when they first arrive, so they are oftenβwith good reason!βrejected, dismissed, and ignored. This is a big problem for a bag of words that contains all of yesterdayβs good ideas. Putting new ideas in the bag will often make the bag worse, on average, because most of those new ideas will be wrong. Thatβs why revolutionary research requires not only intelligence, but also stupidity. I expect humans to remain usefully stupider than bags of words for the foreseeable future.
The most important part of the βbag of wordsβ metaphor is that it prevents us from thinking about AI in terms of social status. Our ancestors had to play status games well enough to survive and reproduceβlosers, by and large, donβt get to pass on their genes. This has left our species exquisitely attuned to whoβs up and whoβs down. Accordingly, we can turn anything into a competition: cheese rolling, nettle eating, phone throwing, toe wrestling, and ferret legging, where male contestants, sans underwear, put live ferrets in their pants for as long as they can. (The world record is five hours and thirty minutes.)
When we personify AI, we mistakenly make it a competitor in our status games. Thatβs why weβve been arguing about artificial intelligence like itβs a new kid in school: is she cool? Is she smart? Does she have a crush on me? The better AIs have gotten, the more status-anxious weβve become. If these things are like people, then we gotta know: are we better or worse than them? Will they be our masters, our rivals, or our slaves? Is their art finer, their short stories tighter, their insights sharper than ours? If so, thereβs only one logical end: ultimately, we must either kill them or worship them.
But a bag of words is not a spouse, a sage, a sovereign, or a serf. Itβs a tool. Its purpose is to automate our drudgeries and amplify our abilities. Its social status is NA; it makes no sense to ask whether itβs βbetterβ than us. The real question is: does using it make us better?
Thatβs why Iβm not afraid of being rendered obsolete by a bag of words. Machines have already matched or surpassed humans on all sorts of tasks. A pitching machine can throw a ball faster than a human can, spellcheck gets the letters right every time, and autotune never sings off key. But we donβt go to baseball games, spelling bees, and Taylor Swift concerts for the speed of the balls, the accuracy of the spelling, or the pureness of the pitch. We go because we care about humans doing those things. It wouldnβt be interesting to watch a bag of words do themβunless we mistakenly start treating that bag like itβs a person.
(Thatβs also why I see no point in using AI to, say, write an essay, just like I see no point in bringing a forklift to the gym. Sure, it can lift the weights, but Iβm not trying to suspend a barbell above the floor for the hell of it. I lift it because I want to become the kind of person who can lift it. Similarly, I write because I want to become the kind of person who can think.)
But that doesnβt mean Iβm unafraid of AI entirely. Iβm plenty afraid! Any tool can be dangerous when used the wrong wayβnail guns and nuclear reactors can kill people just fine without having a mind inside them. In fact, the βbag of wordsβ metaphor makes it clear that AI can be dangerous precisely because it doesnβt operate like humans do. The dangers we face from humans are scary but familiar: hotheaded humans might kick you in the head, reckless humans might drink and drive, duplicitous humans might pretend to be your friend so they can steal your identity. We can guard against these humans because we know how they operate. But we donβt know whatβs gonna come out of the bag of words. For instance, if you show humans computer code that has security vulnerabilities, they do not suddenly start praising Hitler. But LLMs do. So yes, I would worry about putting the nuclear codes in the bag.
Anyone who has owned an old car has been tempted to interpret its various malfunctions as part of its temperament. When it wonβt start on a cold day, it feels like the appropriate response is to plead, the same way you would with a sleepy toddler or a tardy partner: βCβmon Bertie, we gotta get to the dentist!β But ultimately, person perception is a poor guide to vehicle maintenance. Cars are made out of metal and plastic that turn gasoline into forward motion; they are not made out of bones and meat that turn Twinkies into thinking. If you want to fix a broken car, you need a wrench, a screwdriver, and a blueprint, not a cognitive-behavioral therapy manual.
Similarly, anyone who sees a mind inside the bag of words has fallen for a trick. Theyβve had their evolution exploited. Their social faculties are firing not because thereβs a human in front of them, but because natural selection gave those faculties a hair trigger. For all of human history, something that talked like a human and walked like a human was, in fact, a human. Soon enough, something that talks and walks like a human may, in fact, be a very sophisticated logistic regression. If we allow ourselves to be seduced by the superficial similarity, weβll end up like the moths who evolved to navigate by the light of the moon, only to find themselves drawn toβand ultimately electrocuted byβthe mysterious glow of a bug zapper.
Unlike moths, however, we arenβt stuck using the instincts that natural selection gave us. We can choose the schemas we use to think about technology. Weβve done it before: we donβt refer to a backhoe as an βartificial digging guyβ or a crane as an βartificial tall guyβ. We donβt think of books as an βartificial version of someone talking to youβ, photographs as βartificial visual memoriesβ, or listening to recorded sound as βattending an artificial recitalβ. When pocket calculators debuted, they were already smarter than every human on Earth, at least when it comes to calculationβa job that itself used to be done by humans. Folks wondered whether this new technology was βa tool or a toyβ, but nobody seems to have wondered whether it was a person.
(If you covered a backhoe with skin, made its bucket look like a hand, painted eyes on its chassis, and made it play a sound like βhnngghhh!β whenever it lifted something heavy, then weβd start wondering whether thereβs a ghost inside the machine. That wouldnβt tell us anything about backhoes, but it would tell us a lot about our own psychology.)
The original sin of artificial intelligence was, of course, calling it artificial intelligence. Those two words have lured us into making man the measure of machine: βNow itβs as smart as an undergraduate…now itβs as smart as a PhD!β These comparisons only give us the illusion of understanding AIβs capabilities and limitations, as well as our own, because we donβt actually know what it means to be smart in the first place. Our definitions of intelligence are either wrong (βIntelligence is the ability to solve problemsβ) or tautological (βIntelligence is the ability to do things that require intelligenceβ).
Itβs unfortunate that the computer scientists figured out how to make something that kinda looks like intelligence before the psychologists could actually figure out what intelligence is, but here we are. Thereβs no putting the cat back in the bag now. It wonβt fitβthereβs too many words in there.
PS itβs been a busy week on Substackβ
and I discussed why people get so anxious about conversations, and how to have better ones:
And
at Can’t Get Much Higher answered all of my questions about music. He uncovered some surprising stuff, including an issue that caused a civil war on a Beatles message board, and whether they really sang naughty words on the radio in the 1970s:
Derek and Chris both run terrific Substacks, check βem out!
{π¬|β‘|π₯} {What do you think?|Share your opinion below!|Tell us your thoughts in comments!}
#οΈβ£ #Bag #words #mercy
π Posted on 1765150442
