{
  "version": "https://jsonfeed.org/version/1.1",
  "title": "Quionie — Studies",
  "home_page_url": "https://quionie.com/blog",
  "feed_url": "https://quionie.com/feed.json",
  "description": "Long-running research essays by Quionie Gaban on AI, decision-making, and how things actually work. Every claim cited, written in plain language.",
  "language": "en-US",
  "authors": [
    {
      "name": "Quionie Gaban",
      "url": "https://quionie.com"
    }
  ],
  "items": [
    {
      "id": "https://quionie.com/blog/what-ai-has-done-and-whats-wrong",
      "url": "https://quionie.com/blog/what-ai-has-done-and-whats-wrong",
      "title": "What AI Has Done, What's Wrong With It, and What's Just Noise",
      "summary": "The wins that already happened, the harms that deserve the attention they never get, and five claims about AI that fall apart when you check them.",
      "content_html": "<p>I work in tech. I use these tools every day, and part of my job is testing AI systems before they go anywhere near a client. So I spend my week watching how they break.</p>\n<p>I'm pro-AI. Something bothers me more than the arguing does, though.</p>\n<p>The real problems and the fake problems get shouted at the same volume. Someone will spend an hour insisting ChatGPT is boiling the ocean, which isn't true. That same person will never mention that the overwhelming majority of deepfake video is intimate imagery of real people, made and circulated without their consent. Which is true, and measured.</p>\n<p>That's a bad trade. We're spending our attention in the wrong places.</p>\n<p>So I went and read the research.</p>\n<h2>A fifty-year problem in biology, solved and given away</h2>\n<p>Proteins are the machines inside your body. They do everything. What a protein does depends entirely on the shape it folds into.</p>\n<p>If you know the shape, you can design a drug that fits it, the way a key fits a lock. If you don't, you're guessing.</p>\n<p>For fifty years, working out one protein's shape took a scientist years of lab work. Expensive equipment. Slow going.</p>\n<p>In 2020, an AI system called AlphaFold cracked it.</p>\n<p>Two years later the team published predicted shapes for over 200 million proteins, which is close to every protein anyone has ever found, and gave the whole thing away free.</p>\n<p>More than 3 million researchers in over 190 countries have used it. Over a million of them are in low- and middle-income countries, places that could never afford the lab equipment and now have the answers anyway. More than 30% of the research built on it is aimed at understanding disease.</p>\n<p>It won the Nobel Prize in Chemistry in 2024. That same year the Physics prize went to Hopfield and Hinton for the neural network foundations underneath it. Two AI Nobels in one year.</p>\n<h2>It finds breast cancers that two radiologists missed</h2>\n<p>This one already ran, on real people.</p>\n<p>Sweden did a proper randomized trial. Over 105,000 women getting routine mammograms. Half got the standard process, where two radiologists each read the scan. Half got a version where AI read it first and flagged anything suspicious.</p>\n<p>AI-supported screening found 29% more cancers, with no increase in false alarms. It cut the radiologists' reading workload by 44%, which matters more than it sounds, because there aren't enough radiologists and there won't be.</p>\n<p>Then the number that made me sit up. They followed the women afterward and counted interval cancers, the ones that surface between screenings because a screening missed them. Those are the dangerous ones. The AI group had 12% fewer of them, and the ones that did appear were less often the aggressive kind.</p>\n<p>A randomized controlled trial with a hundred thousand real women, and it means some people are alive who wouldn't be.</p>\n<h2>It beat the best weather system on earth</h2>\n<p>Weather forecasting has run for decades on giant physics simulations on supercomputers. The European system, ENS, is considered the world's best.</p>\n<p>DeepMind built a model called GenCast and tested it head to head across 1,320 measures. Wind, temperature, pressure, storm paths, at different distances out.</p>\n<p>GenCast was more accurate on 97.2% of them, and on 99.8% at lead times beyond 36 hours. It's strongest where it counts most: extreme weather and the tracks of tropical cyclones. Knowing three days earlier where a hurricane is going gets measured in lives.</p>\n<p>It produces a 15-day forecast in about eight minutes on a single chip. The traditional method takes hours on a supercomputer. Published in Nature, with the numbers in the paper.</p>\n<h2>At work, it helps beginners most</h2>\n<p>This part surprised me.</p>\n<p>The assumption is that AI serves the people already at the top. The evidence points the other way.</p>\n<p>In a study of 5,172 customer support workers, average productivity rose 14%. For the newest and least skilled, it rose 34%. The most experienced agents gained almost nothing, and on quality they slipped slightly.</p>\n<p>The same pattern showed up in a writing study with 453 professionals and a consulting study with 758 people. Every time, the biggest gains went to whoever was struggling most.</p>\n<p>Most technology widens the gap between the best and the rest. At the individual level, this one narrows it. If you care who benefits from a new tool, that's unusual and worth defending.</p>\n<h2>Non-consensual intimate imagery</h2>\n<p>Now the part where I don't defend anything.</p>\n<p>This belongs at the top of every list and it's almost never there.</p>\n<p>Non-consensual intimate imagery, NCII in the research and the legislation, is exactly what it sounds like: intimate images of a real, identifiable person, produced and shared without their consent. The 2019 Sensity analysis found that 96% of deepfake videos online fell into that category, and every one it examined depicted a woman. A 2023 industry study put the figure at 98%. Different methods, different years, same answer.</p>\n<p>Researchers and lawmakers have moved toward treating NCII as a form of abuse rather than a category of content, and the language matters, because &quot;content&quot; is what you moderate and abuse is what you prosecute.</p>\n<p>Then fraud. Voice cloning got good enough that a few seconds of recorded speech produces a convincing copy. Criminals call parents impersonating their child in distress. They call finance departments impersonating an executive.</p>\n<p>The law finally moved. The TAKE IT DOWN Act was signed in May 2025, making it a federal crime to publish this material, and since May 2026 platforms have had 48 hours to remove it once a victim files a valid report. The first conviction came in April 2026, when an Ohio man pleaded guilty to charges including cyberstalking and publishing digital forgeries, after using AI to target women and children. Around 30 states have passed laws specifically covering the AI-generated kind, on top of near-universal state laws on NCII generally.</p>\n<p><strong>Where I land.</strong> This is the strongest argument anyone can make against how AI is currently being deployed, and I don't have a rebuttal. I don't think one exists. What I'd add is that this isn't a mysterious force. Specific applications were built for it. Specific platforms hosted the output. Those are companies making decisions, and decisions can be regulated and prosecuted, which is finally starting to happen.</p>\n<h2>Entry-level jobs are drying up</h2>\n<p>Stanford has been tracking this with payroll records from ADP covering millions of American workers, and the number keeps getting worse.</p>\n<p>When they first published in August 2025, employment for 22-to-25-year-olds in the most AI-exposed jobs was down 13% relative to their peers in other fields. The August 2026 revision, using data through June, puts it at about 19%.</p>\n<p>Two things make it convincing. Older workers in the same jobs are doing fine, so it isn't the economy and it isn't those industries dying. And almost nobody is getting fired. The drop is overwhelmingly jobs that never got created. Companies quietly stopped opening junior roles.</p>\n<p><strong>Where I land.</strong> No counterargument. If you're 23, applying for junior jobs, and it feels harder than it should, you're right, and the best data available agrees with you. Anyone telling you to just learn AI is skipping the part where somebody still has to hire you.</p>\n<h2>A lot of it doesn't work</h2>\n<p>MIT looked at business AI projects and found roughly 95% produced no measurable financial return. S&amp;P Global found 42% of companies abandoned most of their AI plans in 2025, up from 17% the year before.</p>\n<p>And there's one study I keep coming back to. Sixteen experienced developers, real work, in code they knew well. Half with AI, half without. They predicted AI would make them 24% faster. They came out 19% slower. Afterward they still believed they'd been faster. They couldn't feel it.</p>\n<p><strong>Where I land.</strong> That's real and I take it seriously. It also sits next to several careful studies showing solid gains, so &quot;AI doesn't work&quot; is too strong. What's true is that it works on some things and not others, and people are bad at telling which is which. MIT's own conclusion about the failures is not what you'd expect from a report about failure: the problem is how companies implement it, not the technology.</p>\n<h2>Energy, but only where you'd expect</h2>\n<p>Globally it's smaller than people think. All data centres on earth, not just AI ones, used about 1.5% of the world's electricity in 2024. AI is a fraction of that.</p>\n<p>Growth is fast, and the IEA expects data centre demand to roughly double by 2030. The part nobody mentions is that energy used per AI task is falling at what the IEA calls a rate unprecedented in the history of energy. Total use still climbs, because demand grows faster than efficiency improves. But the guilt about your individual questions doesn't survive the math.</p>\n<p>Where the worry is completely right is locally. These buildings cluster. The US alone accounts for about 45% of global data centre electricity use. In Ireland, data centres went from 5% of metered electricity in 2015 to 23% in 2025.</p>\n<p>If you live next to one, your grid is straining and your bill is climbing, that is real, and quoting a global average at you is a useless answer.</p>\n<p><strong>Where I land.</strong> Two different arguments keep getting mashed together. Planetary carbon, no. Your local grid, absolutely yes.</p>\n<h2>What it might be doing to how we think</h2>\n<p>The famous study came from MIT's Media Lab. People wrote essays with ChatGPT, with a search engine, or with nothing, while researchers watched their brain activity. The ChatGPT group showed the least engagement and, oddly, couldn't quote essays they'd finished minutes earlier.</p>\n<p>Worth taking seriously. It also needs its caveats, and the lead author gives them herself. It hadn't been peer reviewed. There were 54 people, and only 18 finished the last session. And Nataliya Kosmyna has said plainly that they did not find brain rot and did not measure anyone's intelligence.</p>\n<p>The finding nobody repeats is the useful one. People who did the thinking first and brought AI in afterward did better on memory and engagement. The order mattered more than the tool.</p>\n<p><strong>Where I land.</strong> Legitimate concern, thin evidence. I'd want a lot more before panicking or dismissing. The order thing I took personally, because it matches my own experience exactly.</p>\n<h2>Five claims that keep circulating and aren't true</h2>\n<p>These are the ones that don't survive checking, and they crowd out everything above.</p>\n<p><strong>&quot;AI is destroying jobs.&quot;</strong> Across all ages, employment in AI-exposed occupations has barely moved, and overall employment kept growing. The entry-level problem is real. The broad collapse is not happening. The exaggeration actively hurts, too: &quot;AI is killing jobs&quot; is a slogan nobody can act on, while &quot;entry-level hiring is down 19% and no company has a plan for training juniors&quot; is a real problem with a real shape, and it's the true one.</p>\n<p><strong>&quot;Every question you ask uses a bottle of water.&quot;</strong> It traces back to one estimate, about an older model, in specific buildings, and the number swings wildly depending on where the building sits and how it's cooled. It got flattened into a per-question rule the original work never claimed. It also aims the guilt at you. You not asking a question doesn't change where a data centre gets built or what powers it. Those are decisions made by companies and utility regulators, and personal guilt is an effective way to make sure nobody looks at them.</p>\n<p><strong>&quot;AI just glues together stuff it copied.&quot;</strong> This is wrong about <a href=\"https://quionie.com/blog/how-large-language-models-work\">how the thing works</a>, and it matters, because the honest complaint is different. There's no folder of images inside a model. What's inside is a huge pile of numbers that got nudged, over and over, until the model got good at predicting what comes next. Nothing was stored to copy from later. Now, was it fair to use all that work without asking or paying? That's a real question and courts are working on it. But &quot;it's copy-paste&quot; isn't that argument. It's a claim about the machinery, and the machinery doesn't work that way. Making the wrong argument loudly is how you lose one you'd otherwise win.</p>\n<p><strong>&quot;MIT proved AI makes you stupid.&quot;</strong> Fifty-four people. Eighteen in the final session. Not peer reviewed. And the author on record saying they found no such thing. If you're going to cite science at me, and you should, cite it the way you'd want your own cited.</p>\n<p><strong>&quot;It's just autocomplete, so it can't do anything real.&quot;</strong> The description is roughly right. These systems predict what comes next, over and over. The conclusion doesn't follow. The same family of technology mapped every known protein, beat the world's best weather system, and found cancers two radiologists missed. &quot;It's pattern matching&quot; and &quot;it can't do anything important&quot; are two different claims, and the second keeps getting proven wrong while people keep saying it. I think it survives because it's comforting. It lets you skip the harder question, which is what it means that something this simple keeps working this well.</p>\n<h2>Where that leaves me</h2>\n<p>Pro-AI. Specifically:</p>\n<p>The wins are real, they already happened, and they're bigger than most people realize. Proteins. Cancer screening. Weather. And at work, the people it helps most are the ones with the least experience.</p>\n<p>NCII is the worst thing on the list and belongs at the top, not buried under arguments about electricity.</p>\n<p>The entry-level jobs problem is real, getting worse, and deserves more attention than it gets.</p>\n<p>Most of the rest is either exaggerated or aimed at the wrong target.</p>\n<p>Underneath all of it, the technology isn't really the variable. How it gets used is. Nobody is going to stop it advancing, there's too much money and national interest behind it. But how it ships, what's legal, who gets protected, who gets hired, those are all choices. Made by people. Over and over. Including by me, at my job, this week.</p>\n<p>That's why I'm optimistic. Not because nothing bad happens. Because most of the bad things on this list came from decisions, and decisions can be made differently.</p>\n<h2>If you're the one who's worried</h2>\n<p>I'm not going to tell you you're being dramatic. On the two biggest items you're right, and one of them got worse this month.</p>\n<p>What I'd ask is that we argue about the specific thing.</p>\n<p>Your strong version is much stronger than the version that circulates. &quot;AI is bad&quot; gets nobody anywhere. &quot;Non-consensual intimate imagery is 96% of all deepfake video and platforms took years to act&quot; is a claim that has already produced federal law and criminal convictions. So is &quot;entry-level hiring is down 19%.&quot;</p>\n<p>Use the strong version. I'll take it seriously, and so will the people who currently tune you out.</p>\n<h2>What would change my mind</h2>\n<p>This is what I think in August 2026, based on what's published today.</p>\n<p>Several of these numbers moved while I was writing. The Stanford jobs figure went from 13% to 19% across revisions of one paper in a year.</p>\n<p>So I'll come back to this. If the productivity results stop holding up, if the developer slowdown study turns out to be the rule instead of the exception, if the pattern of AI helping beginners most flips, I'll write it down and say what it changed for me.</p>\n<p>Holding a position no matter what the evidence says isn't conviction. It's just a personality.</p>\n<h2>Sources</h2>\n<p>AlphaFold: the <a href=\"https://www.nobelprize.org/prizes/chemistry/2024/press-release/\">2024 Nobel Prize in Chemistry</a> materials, the <a href=\"https://www.nobelprize.org/prizes/physics/2024/press-release/\">2024 Nobel Prize in Physics</a> for the neural network foundations, and <a href=\"https://deepmind.google/blog/alphafold-five-years-of-impact/\">Google DeepMind's five-year impact review</a> for the usage figures.</p>\n<p>Breast cancer screening: the MASAI trial, Lund University. <a href=\"https://www.thelancet.com/journals/lancet/article/PIIS0140-6736(25\">Interval cancer, sensitivity and specificity</a>02464-X/abstract) in The Lancet (2026), with earlier results in The Lancet Oncology (2023) and <a href=\"https://www.thelancet.com/journals/landig/article/PIIS2589-7500(24\">The Lancet Digital Health</a>00267-X/fulltext) (2025).</p>\n<p>Weather: Price et al., <a href=\"https://pubmed.ncbi.nlm.nih.gov/39633054/\">Probabilistic weather forecasting with machine learning</a>, Nature (December 2024), and DeepMind's <a href=\"https://deepmind.google/blog/gencast-predicts-weather-and-the-risks-of-extreme-conditions-with-sota-accuracy/\">GenCast writeup</a>.</p>\n<p>Work: Brynjolfsson, Li and Raymond, <a href=\"https://academic.oup.com/qje/article/140/2/889/7990658\">Generative AI at Work</a>, Quarterly Journal of Economics (2025); Noy and Zhang in Science; Dell'Acqua et al. at Harvard Business School and BCG; METR's developer trial; MIT NANDA's GenAI Divide report; S&amp;P Global.</p>\n<p>Jobs: Brynjolfsson, Chandar and Chen, <a href=\"https://digitaleconomy.stanford.edu/publication/canaries-in-the-coal-mine-six-facts-about-the-recent-employment-effects-of-artificial-intelligence/\">Canaries in the Coal Mine</a>, Stanford Digital Economy Lab, August 2025 through the August 2026 revision.</p>\n<p>Non-consensual intimate imagery: Sensity's 2019 State of Deepfakes analysis and later industry composition studies; the <a href=\"https://www.congress.gov/crs-product/LSB11314\">TAKE IT DOWN Act</a> (Public Law 119-12, signed May 19, 2025); <a href=\"https://www.nbcnews.com/tech/security/first-person-convicted-law-criminalizing-intimate-deepfakes-rcna267236\">reporting on the first conviction</a> (April 2026); state legislative trackers.</p>\n<p>Energy: International Energy Agency, <a href=\"https://www.iea.org/reports/energy-and-ai/executive-summary\">Energy and AI</a>; Ireland's Central Statistics Office on data centre metered consumption.</p>\n<p>Cognition: Kosmyna et al., Your Brain on ChatGPT, MIT Media Lab preprint, plus the author's own clarifications.</p>\n<p>Anything wrong here is mine.</p>",
      "image": "https://quionie.com/og/what-ai-has-done-and-whats-wrong.png",
      "date_published": "2026-08-22T12:00:00.000Z",
      "date_modified": "2026-08-22T12:00:00.000Z",
      "tags": [
        "AI benefits",
        "AI harms",
        "deepfakes",
        "non-consensual intimate imagery",
        "AI and jobs",
        "AI energy use",
        "AlphaFold",
        "AI in healthcare",
        "AI productivity"
      ],
      "authors": [
        {
          "name": "Quionie Gaban",
          "url": "https://quionie.com"
        }
      ]
    },
    {
      "id": "https://quionie.com/blog/ai-agent-architectures",
      "url": "https://quionie.com/blog/ai-agent-architectures",
      "title": "AI Agent Architectures: Workflows, Subagents, and Multi-Agent Systems",
      "summary": "Workflow, agent, subagent, multi-agent system and harness, explained in plain English, plus what the failure data says about splitting work across agents.",
      "content_html": "<p>Every product in <a href=\"https://quionie.com/blog/grok-bot-vs-claude-cowork-vs-chatgpt\">that agent comparison</a> called itself an agent. So does a chatbot with a search button. So does a system that runs for six hours and rewrites your codebase. The word covers all of it, which means it tells you nothing.</p>\n<p>There are real categories underneath. There's also an argument about the most-hyped one, where two credible teams published opposite advice a day apart and neither has backed down.</p>\n<h2>Who decides the next step</h2>\n<p>Anthropic published the clearest version of this in December 2024 and it has held up.</p>\n<p>A workflow is a set of steps you wrote out ahead of time. The model does the thinking inside each step, but you decided what the steps are and what order they run in.</p>\n<p>An agent gets a goal and works out its own steps as it goes. Which tool to use, what to do next, when it's finished.</p>\n<p>A recipe versus a cook. With a recipe you made every decision and the cook carries them out. With a cook you say &quot;make dinner&quot; and hand the decisions over.</p>\n<p>Both get called agentic systems. The only real difference is who decides what happens next: you, in code you wrote, or the model, while it's running.</p>\n<p>The trade runs both ways. Workflows are predictable and easy to fix when they break, and they fall over on anything you didn't think of. Agents handle what you didn't think of, and you give up knowing in advance what they'll do. Every extra turn a model takes on its own makes the job slower, costs more, and gives an early mistake another chance to spread into everything after it.</p>\n<p>Anthropic's own advice is to use the simplest thing that passes your tests, which is often a workflow or even one well-equipped model call, and to save agents for jobs where you can't write the steps out in advance but you can still check whether it's getting somewhere.</p>\n<p>Most reliable systems running in production today are workflows, not agents deciding everything for themselves. Worth knowing before you buy something sold on autonomy.</p>\n<h2>The building block everything is made from</h2>\n<p>Before any of the patterns there's what Anthropic calls the augmented LLM. That's a model with three things attached: it can look things up, it can use tools, and it can remember.</p>\n<p>That isn't an agent yet. It's the unit the rest of this is built out of.</p>\n<h2>The five workflow shapes</h2>\n<p>All five count as workflows, because in each one you wrote the order.</p>\n<p><strong>Prompt chaining.</strong> Break the job into steps that run one after another, with a check in code between each one. Step one drafts, step two checks it, step three formats it. Simple, and mistakes get caught early because you can stop it at any step.</p>\n<p><strong>Routing.</strong> Sort the request first, then send it to a prompt built for that kind of request. Support questions go one way, refund requests go another. Works when the categories are genuinely distinct.</p>\n<p><strong>Parallelization.</strong> Running things at the same time, in one of two ways that constantly get mixed up. Sectioning splits a job into separate pieces that run side by side, which buys speed. Voting runs the same job several times and compares the answers, which buys confidence.</p>\n<p><strong>Orchestrator-workers.</strong> One model reads the job, breaks it into pieces on the spot, hands each piece to a helper model, then puts the answers back together. This is the one people mean when they say subagents.</p>\n<p><strong>Evaluator-optimizer.</strong> One model writes, a second marks it, and it loops until the work is good enough. Works when you can say clearly what good means.</p>\n<p>Real systems mix these. A router at the front feeding into orchestrator-worker setups, each one checking its own output. The shape comes out of the job.</p>\n<h2>What subagents are actually for</h2>\n<p>Subagents live inside the orchestrator-worker pattern, and the reason they exist is room, not horsepower.</p>\n<p>A model can only hold so much text in mind at once. That limit is called the context window. Point one model at a research question with a hundred sources and it has to squeeze everything down to fit, and squeezing too hard loses the details that mattered.</p>\n<p>Subagents get around that. Each one has its own window, works on a different piece at the same time, and sends back a short summary. The model in charge never has to hold all hundred sources at once.</p>\n<p>There's a second benefit that gets less attention. One agent that takes a wrong turn early tends to stay on that road, because everything it does afterward is built on the first bad step. Five agents looking independently don't share that wrong turn.</p>\n<h2>Subagents and multi-agent systems are not the same thing</h2>\n<p>People use the two interchangeably. They're different.</p>\n<p>In orchestrator-workers, the helpers get created on the spot. The model in charge decides what this particular job needs, spins them up, and they're gone when it's done. No names, no history.</p>\n<p>In a multi-agent system the agents are permanent, and each one has a job. A research agent always does research. A critic agent always evaluates. They exist whether or not there's work to do.</p>\n<p><a href=\"https://quionie.com/blog/grok-bot-vs-claude-cowork-vs-chatgpt\">Grok Bot's named bots</a> are the second kind. Anthropic's research system is closer to the first. It matters because permanent specialists build up their own memory and their own bad habits over time, and temporary helpers aren't around long enough to.</p>\n<h2>The argument nobody settled</h2>\n<p>In June 2025, Cognition, the team behind Devin, published a post called Don't Build Multi-Agents. Their case: splitting a job across several agents is fragile, because each one only sees part of the picture and they end up making decisions that contradict each other. Their fix is what people call context engineering, which just means being deliberate about what information the model has in front of it at each moment.</p>\n<p>The next day Anthropic published How we built our multi-agent research system, describing close to the setup Cognition said not to build and arguing it was necessary for hard research questions. On their own research evaluation, a lead Opus 4 agent with Sonnet 4 helpers beat a single Opus 4 by 90.2%.</p>\n<p>It got read as a fight, and it partly is. They also agree more than the coverage suggested.</p>\n<p>Anthropic's post says multi-agent is a poor fit when every agent needs the same information, or when the pieces of the job depend on each other. It names coding as the example.</p>\n<p>Cognition builds a coding agent.</p>\n<p>So the disagreement is narrower than it looked. Both think coding wants one agent with everything in front of it. Both think open-ended research, where you can go down several roads at once, is a different shape. They disagreed about which jobs, published a day apart, and got read as a holy war.</p>\n<p>What I take from it is one question. Can the pieces of your job be done without knowing about each other? If piece B needs to know what piece A found, splitting them costs more than it buys.</p>\n<h2>What the evidence says</h2>\n<p>The most careful work here is a study called MAST, from a Berkeley team led by Mert Cemri, with Matei Zaharia, Joseph Gonzalez and Ion Stoica among the co-authors. They recorded seven popular multi-agent systems doing coding, math and general tasks, on both GPT-4 and Claude models. Human experts went through 150 of those recordings and built a catalog of the ways things went wrong, agreeing with each other almost every time, then applied it to more than 1,600 runs.</p>\n<p>Their opening line is blunt: despite the enthusiasm, the gains over single-agent systems on popular benchmarks are often minimal.</p>\n<p>Across those seven systems, between 41% and 86.7% of runs failed.</p>\n<p>They found 14 distinct ways to fail, in three groups. Roughly 42% came from the system being set up badly in the first place. About 37% came from agents talking past each other. About 21% came from nobody checking the work.</p>\n<p>The biggest group is setup, not capability. The two most common individual failures are agents redoing work they had already done, and an agent reasoning its way to the right answer and then doing something else.</p>\n<p>A 2026 paper by Dat Tran and Douwe Kiela ran one agent against several on questions that take a few steps to answer, and did the thing most comparisons skip. It gave both sides the same amount of thinking to spend. Across three families of models, the single agent matched or beat the multi-agent versions almost every time. Which suggests some multi-agent wins are really the multi-agent version being allowed to think for longer.</p>\n<p>Google Research tested five different setups across three model families and followed what happens to one small mistake. Agents working in parallel with nobody checking them amplified errors 17.2 times over. Put one agent in the middle whose job is to check everything before it moves on, and that drops to 4.4 times. If you split work, something in the middle has to check.</p>\n<p>Then the cost. Anthropic's own numbers: a multi-agent setup uses about 15 times the tokens of a single chat, and a plain agent about 4 times. That's a budget decision.</p>\n<h2>The harness, and why it makes benchmarks slippery</h2>\n<p>The harness is everything wrapped around the model. Which tools it can reach, how its instructions get put together, how it actually runs things, and what it sees when something fails. Same model, different harness, different behavior.</p>\n<p>Recent evaluation work found the harness explains more of the difference in results than the choice of model does, and that the same model scores differently depending on which one it's sitting in. The recommendation is to report a score as model-plus-harness rather than crediting the model on its own.</p>\n<p>So when someone tells you a model is good at agent work, ask which harness. The claim doesn't mean much without it.</p>\n<h2>The words, in one place</h2>\n<p><strong>Agentic system.</strong> Umbrella term. Means almost nothing on its own.</p>\n<p><strong>Workflow.</strong> You wrote the steps. The model does the thinking inside them.</p>\n<p><strong>Agent.</strong> The model decides the next step while it's running.</p>\n<p><strong>Augmented LLM.</strong> A model that can look things up, use tools and remember. The building block.</p>\n<p><strong>Context window.</strong> How much text a model can hold in mind at once.</p>\n<p><strong>Subagent.</strong> A temporary helper created for one job, with its own context window, gone when the job is done.</p>\n<p><strong>Orchestrator-worker.</strong> The setup where that happens. One model in charge, several helpers.</p>\n<p><strong>Multi-agent system.</strong> Permanent agents with names and specialties.</p>\n<p><strong>Harness.</strong> Everything around the model. Tools, instructions, how it runs, what it sees when it fails.</p>\n<p><strong>Tool use.</strong> The model calling a function you gave it. Not autonomy on its own.</p>\n<p><strong>Context engineering.</strong> Being deliberate about what information the model has in front of it at each step.</p>\n<p><strong>ReAct.</strong> The loop most agents run on. Think, do, look at what happened, repeat.</p>\n<h2>How I'd choose</h2>\n<p>Start with the simplest thing that passes your tests. That's Anthropic's advice, and the MAST numbers back it up, because most failures came from bad setup rather than a weak model.</p>\n<p>Ask whether the pieces of your job are genuinely independent. If piece B needs what piece A found, you want one agent with its information well managed, not five agents and a coordination problem.</p>\n<p>If the pieces are independent and the work is broad and open-ended, subagents earn their cost. Research is the clearest example. Coding mostly isn't, and both sides of the famous argument agree on that.</p>\n<p>If you do split the work, put something in the middle to check it. The gap between errors growing 17 times over and errors growing 4 times over is a checking step.</p>\n<p>And budget for roughly 15 times the tokens before deciding multi-agent is worth it.</p>\n<p>The numbers say most of these systems fail because of how they were set up, not because the model wasn't smart enough. So the fix for a broken agent system is usually clearer instructions and better boundaries between tasks, and almost never a bigger model.</p>\n<h2>Sources</h2>\n<p>Anthropic, <a href=\"https://www.anthropic.com/engineering/building-effective-agents\">Building Effective Agents</a> (December 19, 2024) for the workflow-versus-agent distinction, the augmented LLM and the five patterns. Anthropic, <a href=\"https://www.anthropic.com/engineering/built-multi-agent-research-system\">How we built our multi-agent research system</a> (June 13, 2025) for the subagent rationale, the token multiples, the 90.2% eval result and the coding caveat.</p>\n<p>Cognition, <a href=\"https://cognition.com/blog/dont-build-multi-agents\">Don't Build Multi-Agents</a> by Walden Yan (June 2025).</p>\n<p>Cemri et al., <a href=\"https://arxiv.org/abs/2503.13657\">Why Do Multi-Agent LLM Systems Fail?</a> (arXiv 2503.13657, NeurIPS 2025) for MAST, the failure categories and the failure rates. Tran and Kiela, <a href=\"https://arxiv.org/abs/2604.02460\">Single-Agent LLMs Outperform Multi-Agent Systems on Multi-Hop Reasoning Under Equal Thinking Token Budgets</a> (arXiv 2604.02460). Kim et al., <a href=\"https://arxiv.org/abs/2512.08296\">Towards a Science of Scaling Agent Systems</a> (Google Research, arXiv 2512.08296) for the error amplification figures. Harness findings from <a href=\"https://arxiv.org/abs/2605.27922\">Harness-Bench</a> (arXiv 2605.27922).</p>\n<p>Anything wrong here is mine.</p>",
      "image": "https://quionie.com/og/ai-agent-architectures.png",
      "date_published": "2026-08-22T12:00:00.000Z",
      "date_modified": "2026-08-22T12:00:00.000Z",
      "tags": [
        "AI agents",
        "agent architecture",
        "multi-agent systems",
        "subagents",
        "orchestrator-workers",
        "agentic workflows",
        "context engineering",
        "agent harness"
      ],
      "authors": [
        {
          "name": "Quionie Gaban",
          "url": "https://quionie.com"
        }
      ]
    },
    {
      "id": "https://quionie.com/blog/grok-bot-vs-claude-cowork-vs-chatgpt",
      "url": "https://quionie.com/blog/grok-bot-vs-claude-cowork-vs-chatgpt",
      "title": "Grok Bot vs Claude Cowork vs ChatGPT",
      "summary": "Grok Bot, Claude Cowork, ChatGPT, OpenClaw and Hermes compared on where each one runs and whose credentials it holds, with every claim checked against the primary docs.",
      "content_html": "<p>Grok Bot launched ten days ago and I've now read about thirty comparison articles on it. Most are written by companies selling a competing agent. <a href=\"https://felloai.com/best-ai-agents/\">One reviewer counted</a>: nine of the fifteen top-ranking &quot;best AI agents&quot; guides are published by agent vendors, and all nine rank their own product first.</p>\n<p>So I went through the primary docs on all five instead. Grok Bot, Claude Cowork, ChatGPT, OpenClaw, Hermes. On one of them the summaries going around don't match the documentation.</p>\n<p>I'm not selling anything, and I've dated the claims that move.</p>\n<h2>They all launched in eight months and say the same thing</h2>\n<p>Cowork went out in January as a desktop research preview, then to web and mobile on July 7. OpenClaw went viral in late January. Hermes shipped February 25. ChatGPT's agent launched July 9 on GPT-5.6. Grok Bot is the newest, on August 11.</p>\n<p>The positioning lines up almost exactly. A teammate, not a chatbot. You give it a goal instead of a prompt. It keeps working after you close your laptop.</p>\n<p>Five labs with different incentives landed on the same shape inside eight months. That says more about what the technology currently allows than about any one product team. So the pitch won't separate them. Where each one runs will.</p>\n<h2>Where the computer runs</h2>\n<p>Every one of these has to solve the same problem, which is that an agent needs to touch your stuff. There are three ways to do that and each product picks one.</p>\n<p><strong>Run on a cloud machine and log in as you.</strong> The agent gets a browser and uses your real accounts through the real web interfaces. Reaches everything, including software with no API. Holds your actual credentials. That's Grok Bot.</p>\n<p><strong>Run on your machine, on files you hand it.</strong> Local sandbox, reaching outside through connectors you approve. Less reach, much smaller blast radius. Cowork does this, and so do OpenClaw and Hermes, except self-hosted with no vendor in between.</p>\n<p><strong>Run in the vendor's cloud through official connectors.</strong> No credential handoff, no browser puppeteering, only the integrations that already exist. Closest to ChatGPT.</p>\n<figure><img src=\"https://quionie.com/studies/where-the-computer-runs.png\" alt=\"The three places an AI agent can run: on a cloud machine logged in as you (Grok Bot), on your own machine on files you hand it (Cowork, OpenClaw, Hermes), or in the vendor's cloud through official connectors (ChatGPT). Each bucket sets the reach and the risk.\" width=\"1080\" height=\"939\" /><figcaption>The three buckets, and what each one costs you.</figcaption></figure>\n<p>Once you know which bucket something is in, you can predict its price, its reach and its failure modes.</p>\n<h2>Grok Bot: every bot shares one computer</h2>\n<p>It's from SpaceXAI now, the entity that came out of SpaceX acquiring xAI in February. In June, SpaceX agreed to buy Anysphere, the company behind Cursor, for $60 billion in stock. That deal closed in the middle of this month, days before I wrote this.</p>\n<p>That matters, because the product ships through Cursor's stack. Per xAI's own docs, authentication, privacy settings, retention and deletion all resolve to Cursor rather than xAI.</p>\n<p>The product itself is good. Named persistent bots, each with a system prompt, on a cloud Linux machine with browser, filesystem and terminal. Two features nobody else shipped this cleanly. It will record up to ten minutes of you doing a workflow, browser only and no audio, and turn it into a draft skill you review. And bots message each other asynchronously, with group chats of two to six and the handoff visible in the transcript.</p>\n<p>The short write-ups mostly get the next part wrong.</p>\n<p>All of an account's bots share one persistent cloud computer. They share its files, its browser sessions and its app logins. Each bot gets its own screen on that one machine, which is a work surface rather than an isolation boundary.</p>\n<p><strong>Straight from the docs:</strong> &quot;Do not use separate Bots as a security boundary.&quot; That sentence appears word for word on both the Grok Bot FAQ and the approvals and security page. xAI documented the constraint plainly, which is more than most of the coverage managed.</p>\n<p>The practical version: sign into a bank, a payroll system or anything holding regulated data on that machine, and every bot on your account can reach that session. If you pictured one bot per machine, correct that before you connect anything.</p>\n<p>The model powering Grok Bot isn't publicly named anywhere, and press guesses at a version number. For a product you hand credentials to, I'd want that stated.</p>\n<p>On price, be careful with the numbers going around. Access is still the same three tiers it was on day one: SuperGrok Heavy at $300 a month, Cursor Ultra at $200, and Cursor Teams Premium at $120 a seat. SuperGrok Plus and Cursor Pro+ exist and do not include Grok Bot. No free tier, and the seven-day trial some posts mention doesn't appear in the documentation.</p>\n<h2>Claude Cowork is the opposite bet</h2>\n<p>Seven months of hardening versus Grok Bot's ten days, which is not nothing in a category this new.</p>\n<p>Architecturally it's the mirror. Cowork runs in a containerized Linux sandbox on your own machine, works on files you explicitly grant, and reaches services through connectors and a Chrome extension. Sandbox traffic is forced through a proxy with an allow list, and connector tokens never enter the sandbox. Grok Bot runs on a cloud computer it owns and gets to your tools through the same web interfaces you would use, no integration required.</p>\n<p>Grok Bot reaches more, because a browser can log into anything. Cowork reaches less and knows exactly what it touched.</p>\n<p>Price gap is big too. Cowork rides along with Pro, Max, Team and Enterprise, so $20 is the door. Grok Bot's cheapest entry is six times that.</p>\n<p>Where Cowork can't help: if your workflow lives in a system with no connector and no API, it can't get there. That's the hole Grok Bot was built for. One old complaint about it is now out of date, though. Since the July expansion, a task started on desktop keeps running in the cloud with the laptop shut.</p>\n<h2>ChatGPT</h2>\n<p>The least flashy of the five, and probably the one most enterprises land on.</p>\n<p>OpenAI ships the agent side of ChatGPT as ChatGPT Work, launched July 9 on GPT-5.6. Works over connected apps and files, produces documents, spreadsheets, presentations and web apps, and keeps running in the cloud after you close the app. OpenAI frames it as moving ChatGPT past answering questions and toward finishing work, which is close to word for word what SpaceXAI says about Grok Bot.</p>\n<p>It does drive a browser, but not a logged-in one. The cloud browser handles public pages; anything authenticated goes through connectors. That's the whole architectural difference from Grok Bot.</p>\n<p>Better pick when your tools have solid connectors and you mostly need messy inputs turned into something reviewable. Safer for anything stakeholder-facing.</p>\n<p>It runs subagents too: the parent spawns specialists in parallel, routes the work and merges the results. That makes multi-agent less of a Grok Bot differentiator than it looked on launch day.</p>\n<h2>OpenClaw got hit first</h2>\n<p>OpenClaw is open source and self-hosted, built by Peter Steinberger in November 2025 as Clawdbot. It was renamed twice in four days under trademark pressure from Anthropic, first to Moltbot on January 27 and then to its current name on January 30. It went viral immediately, gaining around 25,000 GitHub stars in a single day and passing React's ten-year record in about sixty. <a href=\"https://www.cnbc.com/2026/02/15/openclaw-creator-peter-steinberger-joining-openai-altman-says.html\">Steinberger joined OpenAI</a> on February 15 to work on personal agents, and the project moved to an independent foundation sponsored by OpenAI, NVIDIA, Microsoft and Tencent. As he put it himself, OpenAI hired him, not OpenClaw.</p>\n<p>Then it became the first real AI agent security crisis. Every product in this post has the same shape of exposure. OpenClaw is the one with a public record.</p>\n<p><a href=\"https://www.sonicwall.com/blog/openclaw-auth-token-theft-leading-to-rce-cve-2026-25253\">CVE-2026-25253</a> was a one-click remote code execution via cross-site WebSocket hijacking: a crafted URL parameter made the control panel connect to an attacker's server and hand over its auth token. CVSS 8.8, patched at the end of January. It wasn't isolated. Command injection, server-side request forgery, path traversal that let attackers read local files, and prompt-injection-driven code execution all landed as separate CVEs. Later came a <a href=\"https://www.armosec.io/blog/cve-2026-32922-openclaw-privilege-escalation-cloud-security/\">privilege escalation at 9.9</a> and a sandbox-escape race condition at 9.6.</p>\n<p>The plugin registry got hit too. Koi Security audited all 2,857 skills on the registry and <a href=\"https://thehackernews.com/2026/02/researchers-find-341-malicious-clawhub.html\">found 341 malicious ones</a>, 335 of them part of a single campaign pushing a macOS infostealer. Over a hundred posed as crypto wallet tools. Later audits at larger registry sizes kept finding malicious packages in the hundreds, though the rates reported vary enough that I wouldn't quote a single one.</p>\n<p>Exposure counts are all over the place. Censys counted 63,070 live instances at the end of March. SecurityScorecard found <a href=\"https://strobes.co/blog/42900-openclaw-exposed-control-panels-and-why-you-should-care/\">42,900 exposed control panels</a> across 82 countries. Earlier scans in January found a few hundred, later ones ran into six figures. I'm not picking a number, because they don't agree with each other. What is consistent across all of them is that most exposed instances had no authentication at all, because early versions bound to every network interface with no login required.</p>\n<p>Microsoft's security team wrote in February that OpenClaw &quot;should be treated as untrusted code execution with persistent credentials&quot; and &quot;is not appropriate to run on a standard personal or enterprise workstation.&quot; Kaspersky said use a spare machine or a VPS, never your primary. NVIDIA shipped an enterprise security wrapper for it at GTC in March. China's industry ministry issued a public warning that bad configuration could expose users to attacks and data leaks.</p>\n<p>A normal application has a defined scope, so you can put controls at each boundary. An autonomous agent has no fixed scope by definition, because the whole point is that it can do anything you could do. That's true of the category, not of OpenClaw specifically. It got hit first because it was open, popular, and deployed by a lot of people faster than the defaults could keep up. To the project's credit the fixes shipped quickly and the defaults changed.</p>\n<h2>Hermes picked up where OpenClaw left off</h2>\n<p>Released February 25 by Nous Research, MIT licensed, self-hosted, and widely treated as OpenClaw's technical successor. It ships a migration command that imports OpenClaw personality files, memories, skills and API keys straight over.</p>\n<p>The real difference is how much a human stays in the loop. Both learn skills from work you've already done. OpenClaw holds a learned skill as pending until you review and apply it, while Hermes writes and starts using them on its own after a handful of tool-heavy tasks. Memory lives in a local SQLite database on your machine and skills sit alongside it as plain markdown files. Nothing goes through a third party.</p>\n<p>Adoption is real. <a href=\"https://blogs.nvidia.com/blog/rtx-ai-garage-hermes-agent-dgx-spark/\">NVIDIA's own blog</a> reported in May that it crossed 140,000 GitHub stars in under three months and was the most used agent in the world by OpenRouter's numbers. It's at roughly 234,000 today, up from about 47,000 in April, so treat any single figure as a snapshot with a short shelf life.</p>\n<p>Then the timing. Nous shipped Bot Mode on August 18, three days ago: named reusable bots with their own model and memory that message each other. That's Grok Bot's flagship feature, in an open-source product, seven days after Grok Bot launched.</p>\n<p>Self-hosted means everything in the OpenClaw section is now your problem.</p>\n<h2>Nobody has benchmarked the products</h2>\n<p>Agentic benchmarks exist and the labs publish on them. OSWorld and GAIA for computer use, SWE-bench and Terminal-Bench for coding. Anthropic, OpenAI and xAI all report numbers.</p>\n<p>Those numbers are for the models: Claude Opus, GPT-5.6, Grok 4.6. Nobody has published one for Grok Bot or ChatGPT or Cowork as shipped products, meaning the model plus its scaffolding, its browser, its permission prompts, and what it does when a page won't load. That scaffolding is most of what you're paying for.</p>\n<p>So when a review says one of them finishes the work, that's someone's impression. It might be right. It isn't a measurement, and that includes every review I read for this post.</p>\n<h2>How I'd actually pick</h2>\n<p>Write down what the agent has to touch. If everything on that list has an official connector, stop, and use the cheapest connector-based option. Most people are here and don't realize it.</p>\n<p>If your work lives in something with no API, which describes a lot of real work in construction, healthcare, logistics and government, browser-driving is the only thing that reaches it, and Grok Bot is the most polished version right now. You pay for that with real credentials on a machine you don't control, shared across your bots.</p>\n<p>If data can't leave your environment, you're self-hosting, and Hermes is the better-designed option with the more active project.</p>\n<p>On compliance, I expected this to be the tiebreaker and it isn't. All three of Anthropic, OpenAI and xAI hold SOC 2 Type II, and all three gate HIPAA agreements, SSO and zero-retention behind enterprise contracts and a sales conversation. xAI shipped its business and enterprise tiers in December. If you're in a regulated industry you're having the same procurement call three times, not picking a winner off a page.</p>\n<p>Never pilot on the real account. Run it on low-stakes work and keep sensitive accounts disconnected until it proves out.</p>\n<p>And price the failure, not the subscription. A $20 agent reading files you handed it and a $300 agent holding your CRM password carry very different downside.</p>\n<h2>What I'd watch</h2>\n<p>Whether anyone benchmarks the products rather than the models, because until someone does, this market runs on assertion.</p>\n<p>Whether Grok Bot names its model.</p>\n<p>Whether the multi-agent thing means anything, given ChatGPT already had subagents and an open-source project matched the rest in a week.</p>\n<p>And whether the first serious commercial-agent breach lands. OpenClaw took the hit because it was open, popular and deployed at speed. I don't think the polished ones are structurally safer. Fewer people have tried.</p>\n<h2>Sources</h2>\n<p>Grok Bot: xAI's <a href=\"https://x.ai/bot\">product page</a>, launch post and docs, including the FAQ on the shared computer and the security-boundary warning. Reporting on the SpaceX acquisition of xAI and the $60 billion Anysphere deal from CNBC, Forbes and TechCrunch.</p>\n<p>Claude Cowork: <a href=\"https://techcrunch.com/2026/07/07/the-coding-agent-wars-are-spilling-into-the-rest-of-the-office-claude-cowork/\">TechCrunch</a> on the July expansion and Anthropic's <a href=\"https://www.anthropic.com/engineering/how-we-contain-claude\">engineering writeup</a> on sandbox containment. ChatGPT: <a href=\"https://www.infoworld.com/article/4195478/openai-launches-chatgpt-work-as-it-broadens-gpt-5-6-rollout.html\">InfoWorld</a> on the launch, plus OpenAI's docs on subagents.</p>\n<p>OpenClaw: <a href=\"https://www.microsoft.com/en-us/security/blog/2026/02/19/running-openclaw-safely-identity-isolation-runtime-risk/\">Microsoft Security</a> (February 19), <a href=\"https://www.kaspersky.com/blog/openclaw-vulnerabilities-exposed/55263/\">Kaspersky</a>, <a href=\"https://www.theregister.com/2026/03/16/nvidia_wraps_its_nemoclaw_around/\">The Register</a> on NVIDIA's wrapper, Reuters on China's ministry warning, <a href=\"https://thehackernews.com/2026/02/researchers-find-341-malicious-clawhub.html\">Koi Security</a> on the malicious skills, <a href=\"https://www.sonicwall.com/blog/openclaw-auth-token-theft-leading-to-rce-cve-2026-25253\">SonicWall</a> and <a href=\"https://www.armosec.io/blog/cve-2026-32922-openclaw-privilege-escalation-cloud-security/\">ARMO</a> on the CVEs, Censys and SecurityScorecard on exposure counts.</p>\n<p>Hermes: <a href=\"https://github.com/NousResearch/hermes-agent\">Nous Research's repo</a> and docs, and <a href=\"https://blogs.nvidia.com/blog/rtx-ai-garage-hermes-agent-dgx-spark/\">NVIDIA's blog</a> (May 13) on adoption.</p>\n<p>Dated because most of these numbers moved in the last month, and a few moved while I was writing. Anything wrong here is mine.</p>",
      "image": "https://quionie.com/og/grok-bot-vs-claude-cowork-vs-chatgpt.png",
      "date_published": "2026-08-21T12:00:00.000Z",
      "date_modified": "2026-08-22T12:00:00.000Z",
      "tags": [
        "AI agents",
        "Grok Bot",
        "Claude Cowork",
        "ChatGPT",
        "OpenClaw",
        "Hermes",
        "agent security",
        "AI tool comparison"
      ],
      "authors": [
        {
          "name": "Quionie Gaban",
          "url": "https://quionie.com"
        }
      ]
    },
    {
      "id": "https://quionie.com/blog/how-large-language-models-work",
      "url": "https://quionie.com/blog/how-large-language-models-work",
      "title": "How Large Language Models Actually Work",
      "summary": "A large language model does one thing: predict the next token. This is the full mechanism, from tokenizers and parameters to why models hallucinate.",
      "content_html": "<p>A large language model does one thing. Given some text, it produces a probability for what comes next.</p>\n<p>I kept waiting to find out there was more to it. There isn't. Everything else, the reasoning, the code, the personality, the refusals, the essay it wrote you, comes out of that one operation run over and over.</p>\n<p>Most of the confusion about these systems comes from skipping past that too fast. Slow down on it and a lot of arguments sort themselves out, including a few that serious people are still having in public.</p>\n<h2>It only does one thing</h2>\n<p>Feed a model some text. It produces a probability distribution over its entire vocabulary for what token comes next. Not a word and not an answer, just a ranked list of every option with a number attached to each one.</p>\n<p>Then something outside the model <a href=\"https://quionie.com/blog/why-ai-gives-different-answers\">picks one</a>, sticks it on the end, and feeds the whole thing back in. Over and over, until it produces a token that means stop.</p>\n<p>That loop is the entire operation. When a model writes you four paragraphs, it didn't plan four paragraphs. It picked a token, reread everything including what it just picked, and picked another one, a few thousand times.</p>\n<p>Which is strange, because the output holds together and the mechanism sounds like it shouldn't be able to do that. There's no outline, no goal state, no draft getting revised. Just a very good next-token guess, run enough times to add up to an essay.</p>\n<h2>It has never seen a word</h2>\n<p>Before any of that, your text gets chopped into tokens, and tokens aren't words.</p>\n<p>A tokenizer is a separate thing, built before the model and then frozen. It breaks text into chunks based on what was statistically common in the data it was built from. Common words are usually one token. Uncommon words get split. Spaces and punctuation get folded in.</p>\n<p>So the model never sees &quot;strawberry.&quot; In OpenAI's tokenizer it gets three chunks, st and raw and berry, with no access to the letters inside them.</p>\n<p>That's the main reason models were bad at counting letters. Not because they can't count, and not because they're bad at math. You asked about a level of detail that got thrown away before the model saw anything. It's like asking someone to count the letters in a word they've only ever heard out loud.</p>\n<p>It isn't the whole story. Ask a model to spell the word out first and it usually can, because spellings appear in training text, and once the letters are separate tokens the counting works. So the failure is tokenization plus not thinking to break the word apart first.</p>\n<p>Numbers get chopped the same way. OpenAI's tokenizer splits runs of digits into groups of up to three, reading left to right, so 1000 arrives as 100 and 0. A model doing arithmetic is working with units that don't line up with place value.</p>\n<p><strong>Not a bug in the model:</strong> A <a href=\"https://arxiv.org/abs/2305.15425\">2023 study</a> covering about 200 languages found the same sentence can take 17 tokens in English and 198 in Burmese. You're billed per token and your context fills per token, so those users pay more and run out of room sooner for identical text. Nobody set that as a policy. It fell out of a statistical process nobody was watching.</p>\n<h2>What a parameter is</h2>\n<p>When you read that a model has 70 billion parameters, a parameter is a weight, and a weight is a number saying how much one thing counts toward another. I wrote <a href=\"https://quionie.com/blog/history-of-ai-one-number\">a whole study about that one idea</a>, because the history of the field runs through it.</p>\n<p>That's the whole content of the model. A file of numbers.</p>\n<p>Parameters are what got adjusted during training and what stay fixed afterward. They aren't a database of facts and they aren't compressed text, though people reach for both metaphors. They're the settings that determine how input flows through to a probability distribution.</p>\n<p>So there's no lookup. When a model tells you a date, it isn't retrieving a stored date. The date falls out of the same arithmetic that produces everything else, which is why it can be wrong in ways a database never is.</p>\n<p>And nobody picked any of the numbers. Not one. They're the residue of a training process no human supervised step by step, which is why nobody can open a model and read what it knows.</p>\n<h2>How the numbers get set</h2>\n<p>Pretraining is the expensive part, and the recipe is boring.</p>\n<p>Take enormous amounts of text. Hide the next token. Let the model guess. Measure how wrong it was. Nudge every parameter slightly in the direction that would have made the guess less wrong. Repeat over trillions of tokens.</p>\n<p>That's it. Nobody labels the data with what's true. Nobody teaches it grammar or facts or reasoning. The only signal is whether the next-token guess got better.</p>\n<p>Everything the model appears to know is a side effect. To predict text well you have to encode a lot about how the world works, because text is about the world. Predicting the end of a mystery novel takes something like tracking who knew what. Predicting the next line of a proof takes something like following the proof.</p>\n<p>Whether that &quot;something like&quot; adds up to understanding is the argument I end on. What isn't in dispute is the objective. It's next-token prediction, and it's the only thing being optimized.</p>\n<h2>The size mistake</h2>\n<p>For a few years the assumption was that more parameters meant a better model. GPT-3 had 175 billion and that number was the headline.</p>\n<p>In 2022 a DeepMind team led by Jordan Hoffmann <a href=\"https://arxiv.org/abs/2203.15556\">tested it properly</a>. They trained over 400 models, from 70 million to 16 billion parameters, on 5 to 500 billion tokens. Their finding was that current large models were significantly undertrained. Too many parameters, nowhere near enough data.</p>\n<p>The rule they landed on: for a fixed compute budget, model size and training tokens should scale together, roughly 20 tokens per parameter.</p>\n<p>GPT-3 was trained on about 300 billion tokens. That's under 2 tokens per parameter. It should have had roughly ten times more data, or been about ten times smaller.</p>\n<p>They proved it by building Chinchilla, 70 billion parameters on 1.4 trillion tokens, using about the same compute as a 280 billion parameter model called Gopher. Chinchilla beat Gopher, GPT-3, and Megatron-Turing NLG at 530 billion parameters. On MMLU it scored about 67.5 percent against Gopher's 60.</p>\n<p>The industry's central assumption was wrong for years, and the correction came from running the experiment properly rather than from a new idea.</p>\n<p>The number has moved again since. Llama 3's 8 billion parameter model was trained on 15 trillion tokens, which is about 1,875 tokens per parameter, nearly a hundred times the Chinchilla ratio. Meta did that on purpose and said why: compute-optimal minimizes the cost of training, and ignores that you then serve the model billions of times. Training is a one-time bill. Inference is forever. A smaller model trained far longer is cheaper for good.</p>\n<p>The correction got corrected. That's how this goes.</p>\n<h2>What comes out of training isn't a chatbot</h2>\n<p>A freshly pretrained model, a base model, is not an assistant. It's a text continuation engine. Ask it a question and it might answer, or it might generate three more questions, because in the training data questions are often followed by more questions.</p>\n<p>It has no concept of being asked. There's no user, no turn, no task. It continues text, and a question is just text.</p>\n<p>So the thing OpenAI trained in 2020 wasn't something you could hand to the public. It knew an enormous amount and wouldn't reliably do anything with it.</p>\n<h2>The part that makes it usable</h2>\n<p>The fix is a second stage, and it costs almost nothing next to the first.</p>\n<p>OpenAI's <a href=\"https://arxiv.org/abs/2203.02155\">InstructGPT work</a> in 2022 laid out the recipe everything since is a variation of. Collect about 13,000 prompts and have humans write good responses, then fine-tune on those. Have humans rank model outputs against each other, about 33,000 prompts worth, and train a separate reward model to predict those rankings. Then tune the model to score well against that reward model.</p>\n<p>Human raters preferred outputs from a 1.3 billion parameter InstructGPT model over GPT-3 at 175 billion, a model with a hundred times more parameters. At matched size, raters picked InstructGPT over GPT-3 about 85 percent of the time. On closed-domain tasks, where the answer isn't supposed to contain anything that wasn't in the input, hallucination dropped from around 41 percent to around 21.</p>\n<p>Then the compute. Pretraining GPT-3 took roughly 3,640 petaflop/s-days. The reinforcement learning stage took about 60. Under two percent of the pretraining budget did more for usefulness than a hundredfold increase in size would have. The paper says so directly: RLHF is more effective at making models helpful than a 100x model size increase.</p>\n<p>That's the stage where the assistant gets made. It's also where most of what people complain about gets installed, because you're optimizing the model to produce what human raters approved of.</p>\n<p>Sycophancy comes at least partly from here. <a href=\"https://arxiv.org/abs/2310.13548\">Anthropic found</a> that both human raters and the preference models trained on them tend to prefer responses matching the user's stated view, sometimes over correct ones. The flattened writing voice, the hedging, the particular shape of how a model says no, same stage. None of that is emergent. It's a preference signal, applied on purpose, by people making judgment calls.</p>\n<p>It's also a large part of why models from different labs feel different, which is worth its own section.</p>\n<h2>Why Claude, ChatGPT, and Grok feel different</h2>\n<p>They're the same kind of thing. Every frontier chat model is a transformer predicting the next token, pretrained on an enormous pile of text and then shaped by a preference stage. No lab has a fundamentally different recipe.</p>\n<p>What differs is everything downstream of that.</p>\n<p><strong>The training data.</strong> Undisclosed almost everywhere. Meta is the exception and only roughly, publishing Llama 3's mix as about half general knowledge, a quarter math and reasoning, 17 percent code, 8 percent multilingual. For OpenAI, Anthropic, Google and xAI you get nothing usable. It's a real difference nobody outside the lab can measure.</p>\n<p><strong>The preference stage.</strong> This is where they actually diverge, and it's the best documented part. Anthropic uses <a href=\"https://www.anthropic.com/news/constitutional-ai-harmlessness-from-ai-feedback\">Constitutional AI</a>, where the model critiques its own output against a written set of principles rather than relying only on human raters, plus a separate <a href=\"https://www.anthropic.com/news/claude-character\">character training</a> stage aimed at the persona. OpenAI publishes a <a href=\"https://model-spec.openai.com/\">Model Spec</a>, a document telling raters what good behavior looks like, which then becomes the preference data. Different instructions, different model, similar starting material.</p>\n<p><strong>The system prompt.</strong> A block of text stuck in front of your conversation that sets tone and rules. It isn't in the weights and it can change between Tuesday and Wednesday. Anthropic is the only major lab that <a href=\"https://platform.claude.com/docs/en/release-notes/system-prompts\">publishes theirs</a> with a changelog. What circulates for the others is leaks of uncertain accuracy.</p>\n<p><strong>The tokenizer.</strong> Different vocabularies chop your text differently. GPT-4o's has about 200,000 entries, Llama 3's has 128,256, Google's Gemma family around 256,000. Anthropic has never published Claude's. Same sentence, different token count, different bill.</p>\n<p>Then there's engineering that changes cost and speed without changing the kind of thing. Some models are mixture-of-experts, routing each token through a fraction of the network instead of all of it. DeepSeek-V3 has 671 billion parameters and uses 37 billion per token. Llama 4 and Gemini do versions of this. OpenAI and Anthropic don't say what they do.</p>\n<p>And a lot of what feels like a difference isn't the model. Memory, web search, code execution, file handling, artifacts. That's the product wrapper, which gets its own section further down.</p>\n<p>What nobody can tell you is which of these produces which personality. There's research on how to steer a model's personality once you have one, and nothing rigorous comparing the shipped products to explain why they land differently on people. So when someone tells you Claude is thoughtful and GPT is creative, that's a real impression with no established cause behind it.</p>\n<h2>Why it makes things up</h2>\n<p>Hallucination gets discussed as a bug better engineering will eventually remove. The research is less comfortable than that.</p>\n<p>A 2024 paper by Adam Kalai and Santosh Vempala, <a href=\"https://arxiv.org/abs/2311.14648\">Calibrated Language Models Must Hallucinate</a>, proved a floor. For arbitrary facts, meaning ones you can't work out from patterns in the data, a calibrated model has to produce falsehoods at some minimum rate. The bound comes from calibration itself, not from bad data or the transformer architecture. It's roughly the fraction of facts that appeared exactly once in training.</p>\n<p>Worth being precise about the scope. That floor applies to a calibrated pretrained model, and the paper points at post-training as the way underneath it, since post-training deliberately gives up strict calibration. So it isn't a permanent ceiling on shipped products. It's a property of the pretraining objective.</p>\n<p>Their <a href=\"https://arxiv.org/abs/2509.04664\">2025 follow-up with OpenAI</a> went after why post-training doesn't finish the job, and found the rest of the problem in incentives. Generating a right answer is harder than recognizing one, and they formalized the gap: the generation error rate is at least twice the classification error rate, minus a calibration term. If a model can't reliably tell true from false for some fact, it certainly can't reliably generate the true one.</p>\n<p>Then the part that's a human problem rather than a math problem. They surveyed the major leaderboards and found most of them score answers as simply right or wrong, which means &quot;I don't know&quot; scores exactly the same as a confident lie. A model that guesses beats a model that abstains, every time, on every board. We trained these things to be test-takers and then acted surprised when they bluff.</p>\n<p>Their fix isn't a new architecture. It's changing the scoring. Put an explicit confidence threshold in the question, answer only if you're more than that confident, and penalize a wrong answer in proportion to the threshold you were given. One of the field's most notorious failures turns out to be partly a measurement culture problem, and the people who study it are the ones saying so.</p>\n<h2>Whether abilities appear suddenly</h2>\n<p>You've probably heard that models gain abilities suddenly at certain sizes. A capability is absent, absent, absent, then a model crosses a threshold and it's there. <a href=\"https://arxiv.org/abs/2206.07682\">Wei and colleagues</a> made the case in 2022 and called them emergent abilities, with a companion catalog listing 137 examples. It became one of the most repeated claims about LLMs, and it carries real weight, because unpredictable capability jumps are a safety argument.</p>\n<p>In 2023, Schaeffer, Miranda and Koyejo argued the jumps are largely <a href=\"https://arxiv.org/abs/2304.15004\">an artifact of measurement</a>. Score with something all-or-nothing, like exact string match on an arithmetic problem, and smooth underlying improvement looks like a sudden leap. Switch to a metric with partial credit and the same data draws a smooth curve. Over 92 percent of the emergent abilities annotated on BIG-Bench showed up under just two scoring rules, both of them all-or-nothing. The paper won an Outstanding Paper award at NeurIPS that year.</p>\n<p>It's a good argument and it hasn't ended the debate. A <a href=\"https://arxiv.org/abs/2503.05788\">2025 survey</a> by Berti, Giorgi and Kasneci pushes back, pointing out that a log axis can manufacture the appearance of smoothness, and that some of the smoothed curves still contain jumps from under 10 percent to near 100 percent accuracy across a single step in scale. Their question is fair: does going from 10 percent to 100 percent stop being a jump because of how you drew the axis?</p>\n<p>This isn't settled, and I'd be careful of anyone who says it is in either direction. What's fair to say is that the confident version, where capabilities appear from nowhere at unpredictable scales, is weaker than it was in 2022, and how much weaker depends on measurement choices people are still arguing about.</p>\n<h2>No memory, and memory anyway</h2>\n<p>The model does not remember your conversation. Every turn, the entire transcript gets fed back in and processed from scratch. Nothing carries over inside the model, because the weights are frozen and reading text doesn't change them.</p>\n<p>But products built on models do have memory, it works, and people use it constantly. When you tell an assistant to remember something, that gets written to a database outside the model as text. Next time, it gets retrieved and pasted into the input alongside your message.</p>\n<p>Both things are true, and the distinction is the useful part. The remembering happens outside the intelligence. It's a retrieval system deciding what to paste in, which is why memory features behave the way they do. You can read them, edit them, delete them, and they sometimes surface the wrong thing or miss the obvious thing. Those aren't glitches in a mind. They're retrieval decisions.</p>\n<p>Same logic for the rest of the wrapper. When the thing searches the web, the model didn't do that. When it runs code, executes a task, reads your file, the product did that and handed the result to the model as more text.</p>\n<p>People attribute product features to the model and then draw conclusions about the model. Most of what looks like agency is plumbing.</p>\n<h2>So what is it</h2>\n<p>A large language model is a very large set of numbers, adjusted by a process that repeatedly guessed the next chunk of text and got corrected, until those numbers encoded enough about language and the world to make the guesses good. Then a second, much cheaper process shaped it into something that answers instructions the way human raters preferred.</p>\n<p>You interact with it by feeding text in and taking tokens out, one at a time, with a sampler choosing among ranked options. Around it sits software providing memory, tools, search and safety filtering, none of which the model itself has.</p>\n<p>That's the honest description. No step in it requires the model to understand anything, want anything, or know that you exist.</p>\n<h2>What it isn't</h2>\n<p>Three corrections, and the last one runs against the other two.</p>\n<p><strong>It isn't a database.</strong> Facts aren't stored and retrieved. They fall out of the same arithmetic as everything else, which is why a model can be confidently wrong about a date in a way no lookup system ever is, and why &quot;just make it only say true things&quot; isn't a coherent engineering request.</p>\n<p><strong>It isn't a person.</strong> Nothing persists between conversations. No continuity, no accumulation, no self that grows. The apparent personality is largely the output of a preference-tuning stage, and it can be changed by retraining without anything in there noticing.</p>\n<p><strong>And it isn't obviously just autocomplete either</strong>, which is where I part ways with the confident dismissal. The autocomplete description is mechanically correct, and it smuggles in an assumption: that a system trained to predict text can only ever be shuffling text around. That's an empirical question, not a definitional one, and the honest answer is that we don't know what the training produced, because nobody can read the numbers.</p>\n<p>The dismissive version and the mystical version make the same mistake. Both claim to know what's inside from the outside.</p>\n<p>We know exactly how it works and we still don't know what it is. That's an uncomfortable place to leave it, and it's where things actually are.</p>\n<h2>Sources</h2>\n<p>Hoffmann et al. (2022), <a href=\"https://arxiv.org/abs/2203.15556\">Training Compute-Optimal Large Language Models</a>, for Chinchilla and the 20-tokens-per-parameter result; <a href=\"https://ai.meta.com/blog/meta-llama-3/\">Meta's Llama 3 post</a> and Sardana et al. (2024), <a href=\"https://arxiv.org/abs/2401.00448\">Beyond Chinchilla-Optimal</a>, for training past it on purpose. Ouyang et al. (2022), <a href=\"https://arxiv.org/abs/2203.02155\">Training Language Models to Follow Instructions with Human Feedback</a>, for InstructGPT and the compute figures. Sharma et al. (2023), <a href=\"https://arxiv.org/abs/2310.13548\">Towards Understanding Sycophancy in Language Models</a>. Kalai and Vempala (2024), <a href=\"https://arxiv.org/abs/2311.14648\">Calibrated Language Models Must Hallucinate</a>, and Kalai, Nachum, Vempala and Zhang (2025), <a href=\"https://arxiv.org/abs/2509.04664\">Why Language Models Hallucinate</a>. Wei et al. (2022) on <a href=\"https://arxiv.org/abs/2206.07682\">emergent abilities</a>, Schaeffer, Miranda and Koyejo (2023), <a href=\"https://arxiv.org/abs/2304.15004\">Are Emergent Abilities of Large Language Models a Mirage?</a>, and Berti, Giorgi and Kasneci (2025) for the <a href=\"https://arxiv.org/abs/2503.05788\">counterargument</a>. Petrov et al. (2023), <a href=\"https://arxiv.org/abs/2305.15425\">Language Model Tokenizers Introduce Unfairness Between Languages</a>. Anything wrong here is mine, not theirs.</p>",
      "image": "https://quionie.com/og/how-large-language-models-work.png",
      "date_published": "2026-08-14T12:00:00.000Z",
      "date_modified": "2026-08-22T12:00:00.000Z",
      "tags": [
        "large language models",
        "transformers",
        "tokenization",
        "neural networks",
        "AI hallucination",
        "how AI works"
      ],
      "authors": [
        {
          "name": "Quionie Gaban",
          "url": "https://quionie.com"
        }
      ]
    },
    {
      "id": "https://quionie.com/blog/mental-models-worth-keeping",
      "url": "https://quionie.com/blog/mental-models-worth-keeping",
      "title": "The Mental Models I Keep Coming Back To",
      "summary": "The mental models I actually reach for, from first principles to compounding to leverage, what each one is good for, and the point where each one turns dangerous.",
      "content_html": "<p>Most people collect advice. Almost nobody collects better ways of thinking.</p>\n<p>That's backwards. Advice is glued to the situation it came from, so it expires the second the situation changes. A way of thinking travels. You learn it in one place and it keeps working somewhere nobody expected.</p>\n<p>That's what a mental model is. A compressed version of how some part of the world behaves, portable enough to use somewhere it wasn't built for. Compound interest is a fact about money. It's also a fact about skills, reputation, and grudges. Same shape, different material.</p>\n<p>The problem is that mental models turned into a collectible. Somewhere in the last decade people started hoarding them like productivity apps. A hundred concepts, each one flattened into a tweet, none of them ever pointed at an actual decision.</p>\n<p>That's not what <a href=\"https://en.wikipedia.org/wiki/Charlie_Munger\">Charlie Munger</a> meant. He wasn't handing out flashcards. He was describing a habit: reach for the few models that explain what's in front of you, and notice the second they stop.</p>\n<p>So this isn't a list of a hundred. It's the ones I actually use, including the parts where they turn on you. Every one of them breaks somewhere. The people who get burned usually learned it as a rule and never went looking for where it stops.</p>\n<h2>First principles, and why you probably don't need them</h2>\n<p>The idea is old and simple. Reason up from things you know are true instead of sideways from what everyone else is doing. <a href=\"https://en.wikipedia.org/wiki/Aristotle\">Aristotle</a> called it the first basis from which a thing is known. In practice it means refusing to treat a number as fixed just because the market agreed on it.</p>\n<p>Battery packs are the famous case. Take the going price as a law of nature and electric cars stay expensive forever. Price the raw materials instead, the nickel and cobalt and aluminum sitting on the commodity exchange, and the floor turns out to be way below what anyone was quoting.</p>\n<p>That's the good version. The bad version is a personality. Reasoning up from physics is a great story for a biography and a terrible way to pick a restaurant. Rebuilding a field from scratch is slow and expensive, and the convention you're sneering at is usually hard-won knowledge from people who already paid for the mistakes you're about to make. First principles becomes a party trick the moment the derivation gets more fun than the decision.</p>\n<p>So use it on one thing at a time. Find the single assumption everyone treats as permanent, and ask who decided it, when, and whether they'd decide the same today. That's usually where the mispriced thing is hiding. Everywhere else, take the shortcut. Convention is a decent default, and it's a lot cheaper than rederiving the world before lunch.</p>\n<p><strong>Where it works:</strong> Point it at the one number a whole industry treats as fixed. Take the shortcut on everything else.</p>\n<h2>Inversion</h2>\n<p><a href=\"https://en.wikipedia.org/wiki/Carl_Gustav_Jacob_Jacobi\">Carl Jacobi</a> supposedly told his students to invert, always invert. Munger built his whole style on it. All he wanted to know was where he was going to die, so he could avoid the place. Instead of asking how to win, ask how you'd guarantee losing, then don't do those things.</p>\n<p>It works because failure is smaller and more concrete than success. Nobody can hand you a recipe for a great company or a good marriage. Everybody can tell you how to wreck one. Avoiding stupidity is more reliable than chasing brilliance, and you can start on it today.</p>\n<p>The cleanest version is the pre-mortem. Before you start, pretend it's a year from now and the whole thing has already failed, then write the story of how. People will say things in that exercise they'd never say in an optimistic kickoff, because you've made it safe to be the one who saw it coming instead of the one killing the mood.</p>\n<p>The limit is built in. Inversion tells you what to avoid, not what to build. Spend your entire life dodging mistakes and you end up with a very clean, very small life. It's a filter, and a filter needs something behind it.</p>\n<p><strong>The pre-mortem:</strong> Pretend it's a year from now and the project already failed, then write down exactly why. People name risks there they'd never raise in an optimistic kickoff.</p>\n<h2>Second-order thinking</h2>\n<p><a href=\"https://en.wikipedia.org/wiki/Howard_Marks_(investor\">Howard Marks</a>) has the sharpest version of this. First-level thinking says the company is good, buy the stock. Second-level thinking says everyone already knows it's good, so it's priced like it, and the only money left is in the gap between what people expect and what actually happens. Being right about the obvious pays nothing. The obvious is already in the price.</p>\n<p>Underneath it is a plainer idea. Consequences have consequences. Rent control drops rents for the people who already have an apartment, then kills the reason to build any more, so rents climb for everyone still stuck outside. The first effect is the one on the poster. The second and third are the ones that actually run the world, and almost nobody reads that far down.</p>\n<p>The failure mode is that you can always ask &quot;and then what&quot; one more time. Do it forever and you've talked yourself out of every decision and called it rigor. At some point the next loop costs more than it's worth and you just move. Second-order thinking done badly is anxiety with a flowchart.</p>\n<p><strong>The catch:</strong> You can always ask &quot;and then what&quot; one more time. Past a point that's not rigor, it's paralysis.</p>\n<h2>Base rates, and the myth that more information helps</h2>\n<p><a href=\"https://en.wikipedia.org/wiki/Daniel_Kahneman\">Kahneman</a> and <a href=\"https://en.wikipedia.org/wiki/Amos_Tversky\">Tversky</a> spent years proving that people ignore base rates in favor of a good story. Describe a shy, tidy, detail-obsessed man and ask whether he's more likely a librarian or a farmer. Most people say librarian and forget there are far more farmers in the world. The story beats the math.</p>\n<p>The fix is what forecasters call the outside view. Before you guess how long your project will take, ask how long projects like it usually take, and start there. Your version is a story you're the main character in. The base rate is the record of everyone who tried the same thing and also figured they were the exception.</p>\n<p><a href=\"https://en.wikipedia.org/wiki/Philip_E._Tetlock\">Tetlock's</a> superforecasters basically do this on autopilot. Start at the base rate, update in small steps as real evidence shows up, and don't lurch at the first dramatic headline.</p>\n<p>The part nobody selling a productivity system will admit is that more information doesn't reliably make your decisions better. Past a surprisingly low point, extra detail grows your confidence faster than your accuracy. You feel more sure and you're no more right, which is worse than knowing nothing, because now you'll bet big on it.</p>\n<p>Base rates only mislead when there's no real comparison group, and that's rarer than people want it to be. It usually gets claimed by someone who really needs their case to be special.</p>\n<p><strong>The trap:</strong> More information mostly grows your confidence, not your accuracy. Start from what usually happens, then adjust.</p>\n<h2>Expected value, asymmetry, and the one rule that beats both</h2>\n<p>Expected value is probability times payoff, added up over the outcomes. Obvious on paper, ignored in real life, because a loss stings more than the same-size win feels good, so people pass on bets they should take.</p>\n<p>That instinct has a name, loss aversion, and it comes with a caveat worth knowing. The tidy claim that a loss hurts exactly twice as much as a matching gain has taken a real beating in replication lately. It's more situational than the paperbacks make it sound. It's still there, and it still explains a lot of timid choices. It's just softer and stranger than the number everyone repeats.</p>\n<p>The more useful cousin is asymmetry. You don't have to be right often if you're right big and wrong small. When the worst case is small and fixed and the best case is huge and open-ended, a startup, a cold email, an essay you put on the internet, you can miss most of the time and still come out ahead, as long as you survive the misses.</p>\n<p><a href=\"https://en.wikipedia.org/wiki/Benjamin_Graham\">Benjamin Graham's</a> margin of safety is the same instinct made concrete. Build the bridge to hold thirty tons and only drive ten-ton trucks over it. The gap is there to absorb the error in an estimate that's wrong somewhere you can't see.</p>\n<p>Then there's the rule that sits on top of all of it. You have to still be in the game. Expected value assumes you get to keep playing. A bet with great expected value and a five percent chance of ruin is a bad bet when you only get one life to run it in, because the average across a thousand imaginary versions of you means nothing to the actual one who hit zero. You can't compound from nothing. &quot;Never risk what you can't afford to lose&quot; isn't caution, it's math. Risk is only worth taking when the downside is survivable. When it's not, the upside doesn't matter, because you won't be there to collect it.</p>\n<p><strong>The one rule:</strong> Positive expected value means nothing if a bad outcome takes you out of the game. You can't compound from zero.</p>\n<h2>Opportunity cost</h2>\n<p>Every choice has a cost that never shows up on the receipt: the best thing you didn't pick. A dollar in a mediocre investment isn't only earning less, it's also not in the good one, and that second loss is invisible and usually bigger. Good investors judge every option against their next best one instead of against zero, which is why they pass on things that are merely fine.</p>\n<p>Time works the same way and it hurts more. The real cost of a fine yes is the great thing you now can't say yes to, because you're busy with the fine one. For capable people the most expensive line item is the one that never appears anywhere. It's the work they'll never take, because they already said yes to something adequate.</p>\n<p>The trap is that this can curdle into never committing. Everything has an opportunity cost, so weigh it too hard and you keep every slot open and do nothing, which turns out to have the highest opportunity cost of all. The tool is for choosing better. It's not for making the act of choosing unbearable.</p>\n<p><strong>The catch:</strong> Weigh it too heavily and you commit to nothing, which is the most expensive choice of all.</p>\n<h2>Incentives, and how measuring things breaks them</h2>\n<p>&quot;Show me the incentive and I'll show you the outcome,&quot; Munger said, and he thought it was one of the strongest forces he'd ever watched work. A lot of what looks like stupidity or malice is just someone responding sensibly to an incentive you can't see from where you're standing.</p>\n<p>The sharp edge of this is <a href=\"https://en.wikipedia.org/wiki/Goodhart%27s_law\">Goodhart's Law</a>, best phrased by the anthropologist <a href=\"https://en.wikipedia.org/wiki/Marilyn_Strathern\">Marilyn Strathern</a>: when a measure becomes a target, it stops being a good measure. Teachers told to hit test scores teach the test. Salespeople at Wells Fargo told to open accounts opened around two million fake ones. Measure a team on lines of code and you get software with the structural integrity of wet cardboard. The number was a fine stand-in for the thing you cared about, right up until you started paying people for the number. Then they optimized the number and dropped the thing.</p>\n<p>The part people skip is that the cynical read is also wrong. Assume everyone's a pure self-interested optimizer and you'll build a culture that produces exactly that, because people tend to become whatever you treat them as. Plenty of people do the right thing when it costs them. And over-measuring has a way of killing the quiet motivation that was doing most of the real work in the first place. Incentives explain a lot. They don't explain everything, and running a company as if they do is a reliable way to build one nobody wants to work at.</p>\n<p><strong>Goodhart's Law:</strong> The moment a measure becomes a target, it stops measuring the thing you cared about.</p>\n<h2>Chesterton's Fence</h2>\n<p><a href=\"https://en.wikipedia.org/wiki/G._K._Chesterton\">G. K. Chesterton</a> wrote it as a scene. There's a fence across a road for no reason you can see. The reformer says, I don't see the point of this, let's clear it away. The wiser answer is: if you don't see the point, I definitely won't let you touch it. Go away, work out what it's for, come back, and then maybe I'll let you take it down.</p>\n<p>The lesson is to leave alone what you don't understand yet. The weird approval step, the legacy code everyone's scared to touch, the rule that looks pointless. Someone put it there, maybe for a reason that's no longer obvious but is still holding something up. Systems collect scar tissue, and scar tissue is usually covering an old wound.</p>\n<p>But watch how fast this turns into an excuse to never change anything. &quot;Someone must have had a reason&quot; is not itself a reason. The person who built the fence might have been an idiot, or solving a problem that stopped existing in 1994. The discipline isn't leaving the fence up. It's being able to say what it's for before you decide. Once you can, and the reason's dead, take it down and don't make a ceremony of it.</p>\n<p><strong>The discipline:</strong> Don't touch the fence until you can say what it's for. Once you can, and the reason's gone, take it down.</p>\n<h2>Compounding, and what actually deserves it</h2>\n<p>This is the one that earns the word everyone overuses. It runs on one specific mechanism. Each period's gain becomes the next period's starting point, so the growth builds on the growth, and the line stays boring for a long time before it stops being boring.</p>\n<p>It's not just money. Skills compound, reputation compounds, relationships compound. So do the ugly ones, debt and resentment and a body you keep ignoring, which is why a small bad habit is so much worse than it looks on any single day.</p>\n<p>The line usually pinned on Munger is that the first rule of compounding is to not interrupt it unnecessarily. Most of the damage people do to their own curve is self-inflicted. They sell at the bottom, quit in year three, torch a decade of trust for one good quarter.</p>\n<p>But compounding only matters if the thing underneath is worth compounding. Ten years of practicing the wrong technique makes you excellent at the wrong technique. A business growing twenty percent a year while it loses money on every sale is just going bankrupt with better momentum. People say &quot;it compounds&quot; to justify grinding on things that don't add up to anything, a job with no skill transfer, an audience that will never buy a thing, a pile of contacts they'd never actually call. Patience on the wrong asset isn't a virtue, it's a slow leak. Check that the thing compounds before you congratulate yourself for sticking with it.</p>\n<p><strong>The catch:</strong> It only pays if the thing underneath is worth compounding. Patience on the wrong asset is just a slow leak.</p>\n<h2>Leverage and optionality</h2>\n<p>Leverage is borrowed force. Debt is the obvious kind, but code, capital, other people's time, and an audience are all leverage too. They multiply whatever you point them at. The part people forget is that they multiply mistakes exactly as well as they multiply good calls.</p>\n<p><a href=\"https://en.wikipedia.org/wiki/Warren_Buffett\">Warren Buffett</a> likes to say smart people go broke three ways: liquor, ladies, and leverage. He admits the first two are only in there because he needed words that start with L. A brilliant call and a terrible one both look fine while the borrowed money is still working for you. The difference only shows up when the loan comes due, and by then only one of them lets you keep playing.</p>\n<p>Optionality is the opposite instinct, and the antidote to it: keep your options open. Favor bets you can't lose much on but could win big from, the things that get stronger from chaos instead of breaking under it. That last part is <a href=\"https://en.wikipedia.org/wiki/Nassim_Nicholas_Taleb\">Nassim Taleb's</a> idea of antifragility, and his barbell is the shape of it: most of your resources somewhere boring and safe, a small slice somewhere wild, and nothing parked in the respectable-looking middle that hides how fragile it is until the worst possible moment. <a href=\"https://en.wikipedia.org/wiki/Jeff_Bezos\">Jeff Bezos</a> framed the everyday version as one-way and two-way doors. Most decisions are reversible, so make those fast and cheap, and save the slow, careful kind for the few doors that only open once.</p>\n<p>Optionality has its own failure, and it's a quiet one. It turns into never choosing at all. Keeping every door open is how you end up walking through none of them. And options aren't free. What you pay is the depth and focus you give up by staying available. At some point the antifragile move is to shut the doors and commit, because compounding, the one from earlier, only pays off if you stay put long enough to let it.</p>\n<p><strong>The catch:</strong> Leverage multiplies mistakes as fast as wins, and keeping every door open is how you walk through none.</p>\n<h2>Moats, distribution, and owning the thing</h2>\n<p>If you want to learn from people who got rich, study the machine and skip the worship. The interesting thing about Costco isn't anyone's net worth. It's that a company that caps its own markup and makes most of its real profit on membership fees has built something almost impossible to undercut, because there's no margin left for a competitor to attack.</p>\n<p>That durability is what people mean by a moat. Network effects, where each new user makes the thing more valuable to everyone else, so a payment network that's useless with one merchant is unbeatable with a few million. Switching costs, the enterprise software nobody will ever find the will to rip out. Scale, where being the biggest just makes you the cheapest. Brand, which at its best is a promise strong enough that people stop comparing prices and pay more for the same molecule under a name they trust. None of that comes from the marketing department. It's structural, and structure is what gets you through a bad year.</p>\n<p>Founders almost always overrate the product and underrate distribution. <a href=\"https://en.wikipedia.org/wiki/Peter_Thiel\">Peter Thiel's</a> uncomfortable point is that a better product with no way to reach people loses to a worse product that reaches everyone. &quot;Build it and they will come&quot; is survivorship bias talking. You only ever hear the stories where it happened to work.</p>\n<p>The engine under all of it is ownership. The math is plain: a wage is linear and taxed hard on the way in, while a claim on an appreciating asset compounds and is taxed lightly, if at all, on the way out. That's why almost every large fortune traces back to owning something, equity or a business or property, more than to a paycheck. None of which makes a salary a bad deal. It's steady, it's low-risk, and it's usually what pays for any ownership in the first place. The two just grow differently, and the useful thing is understanding why.</p>\n<p>The failure here is copying the surface. Costco's food court won't give you Costco's economics, and Apple's font won't give you Apple's brand. You have to build what's underneath, and that part never copies.</p>\n<p><strong>The mechanism:</strong> A wage grows in a straight line; an appreciating asset compounds. That's the whole difference, and it's just math.</p>\n<h2>What they all have in common</h2>\n<p>Buffett talks about a circle of competence, and people usually miss his actual point. You don't need to be an expert on everything. You need to know where the edge of your circle is. The size of it barely matters. Knowing the border is the whole game, because the expensive mistakes almost always happen just outside it, dressed up to look like they're just inside.</p>\n<p>Look back over the rest and they rhyme. First principles, base rates, margin of safety, Chesterton's fence, the circle itself. Most of them are disciplined ways of admitting what you don't know. Inversion is humility about your odds of being brilliant. Margin of safety, humility about your estimates. Second-order thinking, humility about consequences. The models aren't clever. They're just honest in a structured way.</p>\n<p>None of them holds up alone. The person who only knows inversion becomes a professional pessimist who never builds anything. The one who only knows compounding never sells, never exits, never cuts a loss. They work as a set, checking each other, and the skill is knowing which one the moment is asking for. That part is judgment, and no model can hand you judgment. It's what's left after the models have narrowed things down.</p>\n<p>None of these are the whole picture. Each one makes a specific thing easy to see and leaves the rest out, which is fine as long as you know that going in. Keep the few that hold up when you actually use them, and stay ready to drop one when reality stops agreeing with it. It always does eventually.</p>\n<h2>Sources</h2>\n<p>If you want the originals instead of my read on them: Munger's <a href=\"https://en.wikipedia.org/wiki/Poor_Charlie%27s_Almanack\">Poor Charlie's Almanack</a> and his 1994 talk on worldly wisdom; Howard Marks in The Most Important Thing; Graham's <a href=\"https://en.wikipedia.org/wiki/The_Intelligent_Investor\">The Intelligent Investor</a>; Kahneman's <a href=\"https://en.wikipedia.org/wiki/Thinking,_Fast_and_Slow\">Thinking, Fast and Slow</a> and Tetlock's <a href=\"https://en.wikipedia.org/wiki/Superforecasting\">Superforecasting</a> for the base-rate work; Chesterton's The Thing for the fence; Taleb's <a href=\"https://en.wikipedia.org/wiki/Antifragile_(book\">Antifragile</a>); and Thiel's <a href=\"https://en.wikipedia.org/wiki/Zero_to_One\">Zero to One</a> for distribution and power laws. Everything above is my read on these, so anything that's off is on me, not them.</p>",
      "image": "https://quionie.com/og/mental-models-worth-keeping.png",
      "date_published": "2026-07-07T12:00:00.000Z",
      "date_modified": "2026-07-07T12:00:00.000Z",
      "tags": [
        "mental models",
        "decision making",
        "second-order thinking",
        "expected value",
        "risk of ruin"
      ],
      "authors": [
        {
          "name": "Quionie Gaban",
          "url": "https://quionie.com"
        }
      ]
    },
    {
      "id": "https://quionie.com/blog/history-of-ai-one-number",
      "url": "https://quionie.com/blog/history-of-ai-one-number",
      "title": "A History of AI, Told Through One Number",
      "summary": "The part of ChatGPT that seems to think is nothing but numbers nobody chose. A history of AI told through one idea: the weight.",
      "content_html": "<p>The part of ChatGPT that actually seems to think is nothing but numbers.</p>\n<p>There's code that runs it, and a hidden prompt that nudges its tone. But the ability itself, the part that can write and reason and hold a conversation, was never written down by anyone. No rules, no logic anyone typed in. It lives in hundreds of billions of numbers sitting in a file.</p>\n<p>Nobody picked them. Not one, not ever, and no human went through them by hand. We built the process that finds the numbers and handed over the part where you'd normally understand what you made.</p>\n<p>That fact sits at the bottom of every argument about AI right now, and almost nobody explaining this stuff starts there. They start with neural networks, or training data, or transformers, all of which are made of something smaller that rarely gets named.</p>\n<blockquote><p>A weight is a number that says how much one thing matters to another.</p></blockquote>\n<p>That's it. Learn that one properly and the last seventy years turn into a single argument with one question in it: who picks the numbers. Humans, or the machine.</p>\n<h2>The one thing</h2>\n<p>You already do this. You just don't call it weighting. Meeting someone for the first time, you decide how much to trust them by leaning hard on a couple of signals (whether they actually listen, how they treat the waiter) and ignoring others completely (what car they drive). Nobody handed you that ranking. You built it over years of misjudging people and adjusting.</p>\n<p>That ranking is the whole idea of a weight: how much a given signal counts toward the answer. An artificial neuron is the same move in numbers. Each signal is multiplied by its weight (high, low, or zero), the results are added up, and if the total clears a line, the neuron fires.</p>\n<p>Stack billions and you have a language model. The behavior has no second ingredient. No rules, no logic anyone wrote for what it should say. Just numbers saying how much things matter to each other, found by training rather than typed in by hand.</p>\n<p>Which is what makes the question of who sets them the only one that's ever mattered. Seventy years, two collapses, one long feud, and the thing on your phone. All of it is that.</p>\n<h2>1943, a neuron becomes arithmetic</h2>\n<p><a href=\"https://en.wikipedia.org/wiki/Warren_Sturgis_McCulloch\">Warren McCulloch</a>, a neurophysiologist, and <a href=\"https://en.wikipedia.org/wiki/Walter_Pitts\">Walter Pitts</a>, a logician who taught himself the field as a teenager and was more or less homeless when the two of them met, published a paper describing brain cells as simple logical switches. Inputs, a threshold, an output. On or off.</p>\n<p>They meant it as biology. What it proved by accident is that you can build thought-shaped things out of arithmetic.</p>\n<p>It couldn't learn, though. A person set the weights by hand, and that person had to already know the answer. Which makes it a description of a brain, not a thing that behaves like one.</p>\n<p>That gap took fifteen years to close.</p>\n<h2>1958, and the worst press release in science</h2>\n<p><a href=\"https://en.wikipedia.org/wiki/Frank_Rosenblatt\">Frank Rosenblatt</a> was a psychologist at Cornell, and the piece he added is almost insultingly simple.</p>\n<p>Show it an example. Let it guess. If the guess is wrong, nudge the weights slightly toward what would have been right. Do it again.</p>\n<p>That's the <a href=\"https://en.wikipedia.org/wiki/Perceptron\">perceptron</a> learning rule, and it is the ancestor of everything running today. Nobody tells the machine what matters. It finds out by being wrong at scale.</p>\n<p>He first ran it as software on an IBM 704, and the Navy held a press conference in 1958 to show it off. The New York Times covered it under a headline about a device that learns by doing, describing a machine expected to eventually &quot;walk, talk, see, write, reproduce itself and be conscious of its existence.&quot; The New Yorker called it the first serious rival to the human brain ever devised.</p>\n<p>What the machine had actually done that day was learn to tell left from right, after about fifty tries.</p>\n<p>The hardware version came around 1960. The Mark I Perceptron was room-sized, with a 20x20 grid of photocells for an eye, wired to banks of motor-driven potentiometers that served as the adjustable weights. Which means the weights were knobs, and learning was a set of small motors turning them. You could stand in the room and watch a machine physically change its mind.</p>\n<p>Every hype cycle in AI since has the shape of that press conference. A real breakthrough, correctly spotted as a big deal, then described in language that runs ahead of what the thing can do yet. Worth holding onto when you read anything about AI now, including this.</p>\n<h2>The problem a straight line can't solve</h2>\n<p>Take a small logic puzzle called XOR. Two inputs, each either on or off. You want a yes when they differ (one on, one off) and a no when they match. It's the logic of a hallway light wired to a switch at both ends: flip either switch and the light flips. Sounds like nothing.</p>\n<p>One layer of weights cannot do it. Not &quot;hasn't managed yet.&quot; Cannot, provably.</p>\n<p>The reason is geometric. A single layer can only draw a straight line through your data, one side yes and the other no. XOR needs a division a straight line can't produce. And once you know to look for that shape, you find it in most problems worth solving.</p>\n<p>The fix was visible even then. Stack layers and the network can bend the line. The catch was that nobody could train a stack. The learning rule works on one layer because when the answer comes out wrong you know precisely which weights to blame. Add layers and the blame goes murky. Which weight, in which layer, caused this?</p>\n<p>So the field had a machine that could learn but not handle anything real, and a design that could handle real things but not learn.</p>\n<p><strong>Why layers matter:</strong> One layer of weights can only separate things with a straight line. Most real problems need a curve, and a curve takes stacked layers. The whole catch was that nobody could train the stack yet.</p>\n<h2>The 1969 story everyone repeats wrong</h2>\n<p><a href=\"https://en.wikipedia.org/wiki/Marvin_Minsky\">Marvin Minsky</a> and <a href=\"https://en.wikipedia.org/wiki/Seymour_Papert\">Seymour Papert</a> published a book called <a href=\"https://en.wikipedia.org/wiki/Perceptrons_(book\">Perceptrons</a>) laying out exactly that limit, and the standard telling is that it single-handedly killed neural networks and triggered the first AI winter. One book, funding gone, field dead.</p>\n<p>That version is in most explainers and it doesn't survive contact with the sources.</p>\n<p>Minsky and Papert had been arguing against this approach since around 1965, at conferences and in circulated drafts. By 1969 most researchers had already drifted away, worn down by lack of progress, and the rules-based camp had largely won the funding fight before the book existed. It landed in a room that was already emptying and picked up its reputation as the assassin partly through timing.</p>\n<p>Rosenblatt died in a boating accident in Chesapeake Bay in July 1971, on his 43rd birthday. When the book was reissued in 1987, it carried a dedication to him.</p>\n<p>Worth knowing which version you're repeating, because the tidy one makes it sound like ideas die from criticism. They mostly die from people quietly leaving.</p>\n<h2>Thirty years of typing rules in by hand</h2>\n<p>Left out of most tellings: for about three decades the mainstream of AI wasn't weights at all.</p>\n<p>The other approach was rules. Write down what you know as explicit logic and let the machine reason over it. If fever and cough, consider flu. Knowledge as statements a human typed in.</p>\n<p>This was the sensible bet, not the dumb one. Rules are readable. You can debug them. You can ask the system why it decided something and get a real answer, which is the exact thing everyone now complains AI can't do.</p>\n<p>It produced <a href=\"https://en.wikipedia.org/wiki/Expert_system\">expert systems</a> that worked in narrow slices (diagnosing infections, configuring hardware orders), and a genuine industry grew around them in the 80s.</p>\n<p>Then it hit the wall it was always going to hit, which is that a human has to type in every single thing. Every rule, every exception, every exception to the exception. Systems grew past a few thousand rules and started contradicting themselves in ways nobody could untangle. And the parts we do without thinking, seeing, hearing, reading a room, turn out to be the hardest to write down, because nobody can say what rules they're following.</p>\n<p>The market collapsed in the late 80s. Second winter.</p>\n<p>Both approaches failed, which is the part that gets dropped. Rules failed because humans can't articulate what they know. Weights failed because nobody could train more than one layer. Same era, different walls.</p>\n<p><strong>The two camps:</strong> Rules means a human writes the logic. Readable, but you have to type in everything. Weights means the machine finds its own numbers. Powerful, but nobody can read the result. This whole history is the fight between the two.</p>\n<h2>1986</h2>\n<p>The unlock is <a href=\"https://en.wikipedia.org/wiki/Backpropagation\">backpropagation</a>, from a 1986 paper by <a href=\"https://en.wikipedia.org/wiki/David_Rumelhart\">David Rumelhart</a>, <a href=\"https://en.wikipedia.org/wiki/Geoffrey_Hinton\">Geoffrey Hinton</a>, and <a href=\"https://en.wikipedia.org/wiki/Ronald_J._Williams\">Ronald Williams</a>. Earlier versions existed (Linnainmaa in 1970, Werbos in 1974), but this is the one that landed.</p>\n<p>It solves the blame problem. Work the error backwards through the network, layer by layer, calculating how much each weight contributed to the mistake, then nudge all of them slightly toward better. Millions of examples later, a stack is trained.</p>\n<p>Still the engine today. Bigger, faster, same idea.</p>\n<p>And look at what it actually is: guess, measure the error, nudge the weights. Rosenblatt's rule with the layer problem solved. 1958, unstuck.</p>\n<p>Hinton shared the Nobel Prize in Physics in 2024, partly for this. Thirty-eight years is a long time to wait for a phone call.</p>\n<p><strong>Backpropagation:</strong> Run the mistake backwards through the layers, work out how much each weight caused it, and nudge them all toward better. It's what finally made deep networks trainable, and it's still the engine today.</p>\n<h2>Why a finished idea sat there until 2012</h2>\n<p>Backprop worked in 1986. Deep learning didn't take over until 2012, and that 26-year gap is the most useful thing in this history.</p>\n<p>It wasn't ideas. The math was done. <a href=\"https://en.wikipedia.org/wiki/Yann_LeCun\">Yann LeCun</a> had these networks reading handwritten digits on real bank checks in the 90s. It worked. It just couldn't be pushed further.</p>\n<p>Two things were missing and both are boring.</p>\n<p>Data. These systems need an absurd number of examples, and before the internet, assembling a large labeled dataset meant paying people to sit there labeling. <a href=\"https://en.wikipedia.org/wiki/ImageNet\">ImageNet</a> landed in 2009, millions of labeled images, built through a lot of unglamorous coordination nobody writes hero stories about.</p>\n<p>Compute. Training is billions of small multiplications, which happens to be exactly what a graphics card does to render video games. Nobody designed GPUs for this. Researchers noticed the hardware fit.</p>\n<p>In 2012 a network called <a href=\"https://en.wikipedia.org/wiki/AlexNet\">AlexNet</a>, from <a href=\"https://en.wikipedia.org/wiki/Alex_Krizhevsky\">Alex Krizhevsky</a>, <a href=\"https://en.wikipedia.org/wiki/Ilya_Sutskever\">Ilya Sutskever</a>, and Hinton, entered the ImageNet competition running on gaming GPUs and posted a top-5 error rate of 15.3 percent, meaning the correct label sat in its top five guesses about 85 percent of the time. Second place was more than ten percentage points behind. In a benchmark where people fought over fractions, that ended the argument.</p>\n<p>Nothing conceptual had changed since 1986. The idea was finally allowed to run.</p>\n<blockquote><p>The idea was finished in 1986 and useless until 2012. What showed up wasn't insight. It was cheap data and gaming hardware, arriving from two directions nobody in AI controlled.</p></blockquote>\n<h2>Weights that point at each other</h2>\n<p>Last piece. The <a href=\"https://en.wikipedia.org/wiki/Transformer_(deep_learning_architecture\">transformer</a>), 2017, from a Google paper whose title gives away the ending: &quot;<a href=\"https://en.wikipedia.org/wiki/Attention_Is_All_You_Need\">Attention Is All You Need</a>.&quot;</p>\n<p>The problem it fixed: earlier systems read text one word at a time in order, holding a running summary, and that summary gets overloaded. Things far apart in a sentence lose track of each other.</p>\n<p>Attention lets every word look at every other word directly and decide how much each one matters to it. In &quot;the trophy didn't fit in the suitcase because it was too big,&quot; the word &quot;it&quot; checks both nouns and weighs them.</p>\n<p>Notice what attention is, though. Weights again. Computed fresh for whatever you typed instead of fixed after training.</p>\n<p>Same 1943 neuron, same 1958 rule, same 1986 blame math, rearranged so the weights can point at each other. That's the entire architecture under everything you've used.</p>\n<p><strong>Attention:</strong> Instead of reading word by word and losing track, every word looks at every other word and decides how much each one matters to it. The entire transformer is built on that one move.</p>\n<h2>The bitter part</h2>\n<p>In 2019 <a href=\"https://en.wikipedia.org/wiki/Richard_S._Sutton\">Rich Sutton</a> wrote a short piece called <a href=\"https://en.wikipedia.org/wiki/Bitter_lesson\">The Bitter Lesson</a>, and it's the closest thing here to a moral.</p>\n<p>His claim: for seventy years the same thing keeps happening. Researchers carefully build in human knowledge about a problem. Chess strategy, grammar rules, which features matter in an image. It works, and it feels like science. Then someone throws a dumber, more general method at more compute and beats it, and all that encoded expertise turns out to be dead weight.</p>\n<p>Chess, Go, speech, vision, translation. Same result every time.</p>\n<p>He calls it bitter because it flatters nobody who does the work.</p>\n<p>I don't buy the strong version. The transformer was human insight. Attention was a design choice a person made, and &quot;just add compute&quot; reads far cleaner in hindsight than it did to anyone deciding what to add compute to. What holds up is the narrower claim, which is that betting against scale has a perfect losing record, and the people who lost that bet were usually the ones most certain their expertise was the important part.</p>\n<h2>So nobody can read it</h2>\n<p>Back to where this started.</p>\n<p>We understand training completely. The recipe is public and taught in undergrad courses. Guess, measure the error, nudge every weight, repeat trillions of times.</p>\n<p>We invented the search. We did not pick the result.</p>\n<p>Nobody chooses the weights. Nobody has inspected them. Hundreds of billions of numbers, each adjusted by microscopic amounts an unfathomable number of times, and what comes out is the residue of a process no human supervised at any point.</p>\n<p>They aren't hidden, either, which is the next assumption people make. You can download an open model right now and print every number. They look like 0.0231, -1.4432, 0.8871. And they tell you nothing, because meaning isn't stored in any single one. Concepts get smeared across many weights at once, and each weight takes part in many different concepts depending on what else is happening around it. There's no dictionary. There was never anyone around to write one.</p>\n<p>So we have an artifact we built and now have to study experimentally, the way you'd study something you dug up. There's a field for this called interpretability, and it is mostly unsolved.</p>\n<p>The short version: we automated the writing of the program, and the price was being able to read it.</p>\n<p>That's not a bug that crept in late. It's the same trade Rosenblatt made in 1958 when he chose a machine that finds its own weights over a human writing them down. The rules camp kept readability and lost, because humans can't type in everything they know. The weights camp took the problems nobody could write down and gave up ever knowing what the solution says.</p>\n<p>Nobody framed it as a trade at the time. It turned out to be one.</p>\n<p>Three things that fall out of this if you use these tools for work:</p>\n<p>There's no rulebook inside, so stop hunting for the line to fix. When a model gets something wrong, nobody can open it and point. That's why prompting is trial and error instead of configuration.</p>\n<p>A confident answer isn't automatically a correct one. Good answers and wrong ones come out of the same process, so it's worth checking the ones that matter instead of going off the tone.</p>\n<p>When capability jumps, the boring explanation, more data and more compute, is usually most of it.</p>\n<p><strong>Interpretability:</strong> The field trying to read the finished weights and work out what they mean. Mostly unsolved, because meaning is smeared across billions of them at once, with no dictionary.</p>\n<h2>The frozen machine</h2>\n<p>One last thing, because it's where people's intuition breaks.</p>\n<p>Paste five examples of some invented task into a prompt, something the model has never seen, and it will do the sixth one correctly. It picked up a new task on the spot. Nothing inside changed.</p>\n<p>That looks like learning, and it isn't, in the sense that matters. Water through pipes. The pipes were shaped once, during training, and never change again. Water flows, takes on a pattern, the pattern does something useful, and when the flow stops nothing remains. Close the chat and it's the identical machine it was before you typed a word.</p>\n<p>When you learned to do your job it changed you, and you'll still know it next week. The model can't do that mid-conversation. Every session starts from the same frozen thing.</p>\n<p>Which is the whole history in one behavior. A machine that appears to learn, built by a process that already finished, made of numbers nobody selected, doing something nobody can fully explain.</p>\n<p>The explaining part is still open. Seventy years in, that's the piece still sitting there.</p>\n<p><strong>Why it isn't learning:</strong> The weights freeze the moment training ends. A prompt changes what flows through the machine, not the machine itself. Close the chat and it's identical to before you typed.</p>\n<h2>Sources</h2>\n<p>The originals, if you want them instead of my read: McCulloch and Pitts (1943); Rosenblatt's 1958 perceptron work and the New York Times piece from July 8, 1958; Minsky and Papert's Perceptrons (1969), plus Yuxi Liu's essay on the perceptron controversy for the corrected history in the 1969 section; Rumelhart, Hinton and Williams (1986); Krizhevsky, Sutskever and Hinton (2012) for AlexNet; Vaswani et al. (2017); Rich Sutton's The Bitter Lesson (2019). Anything wrong is mine, not theirs.</p>",
      "image": "https://quionie.com/og/history-of-ai-one-number.png",
      "date_published": "2026-07-01T12:00:00.000Z",
      "date_modified": "2026-08-22T12:00:00.000Z",
      "tags": [
        "history of AI",
        "neural networks",
        "perceptron",
        "deep learning",
        "model weights",
        "XOR problem"
      ],
      "authors": [
        {
          "name": "Quionie Gaban",
          "url": "https://quionie.com"
        }
      ]
    },
    {
      "id": "https://quionie.com/blog/why-ai-gives-different-answers",
      "url": "https://quionie.com/blog/why-ai-gives-different-answers",
      "title": "Why AI Gives You a Different Answer Every Time",
      "summary": "Ask a model the same question twice and you get two different answers. It comes down to one number outside the model, and most of what gets said about it is wrong.",
      "content_html": "<p>Ask a model the same question twice and you get two different answers.</p>\n<p>I used to think that was the model thinking differently the second time. It's not. It comes down to one number, and that number isn't even part of the model.</p>\n<h2>The model doesn't pick anything</h2>\n<p>A <a href=\"https://en.wikipedia.org/wiki/Large_language_model\">language model</a> doesn't produce an answer. It produces a score for every token in its vocabulary, all at once, and those scores get turned into probabilities. On a current model that's a ranked list around a hundred thousand entries long, sometimes closer to two hundred thousand, with a number attached to every one.</p>\n<p>Then something outside the model picks one of them. The ranking and the picking are two different jobs, done by two different pieces of software. I had them filed in my head as one.</p>\n<p>Temperature is the dial on the second one.</p>\n<p><strong>Token:</strong> Models don't read words, they read tokens. A token is a chunk of text, often a whole word, sometimes a piece of one, sometimes just punctuation. &quot;Unbelievable&quot; might arrive as three of them.</p>\n<p>Mechanically it's one division. Every score gets divided by the temperature before it becomes a probability. Divide by a number below 1 and the gaps between scores stretch, so the leaders pull further ahead. Divide by a number above 1 and the gaps compress, so the field bunches up and long shots get a real chance.</p>\n<p>Temperature 1 is the neutral setting, because dividing by one changes nothing and you get the model's own distribution exactly as it came out. That's the reference point most explanations leave out. &quot;Higher temperature flattens the distribution&quot; doesn't mean much until you know what it's flattening against.</p>\n<p>At 0 the math breaks, since you can't divide by zero, so implementations skip the formula and just take the top-ranked token every time. So 0 isn't the low end of the dial. It's the sampler switched off.</p>\n<p><strong>Worth knowing:</strong> The ranges differ by provider. Anthropic caps temperature at 1. OpenAI and Google allow up to 2, with 1 as the default. Several of the newest reasoning models don't accept the parameter at all and simply run at their default.</p>\n<p>Which is why the same question gives you different answers. Nothing about the model changed between your two tries. The dice did.</p>\n<h2>The thing everyone says about it</h2>\n<p>The standard line, and I've repeated it, is that temperature is a slider between accuracy and creativity. Turn it down for facts, turn it up for writing, and the price of creativity is hallucination.</p>\n<p>Half of that holds up. The other half doesn't have much behind it.</p>\n<p>A 2024 dialogue benchmark called <a href=\"https://arxiv.org/abs/2406.07070\">HalluDial</a> swept temperature and tracked how often the model made things up. The rate sat low and roughly flat through the lower range, then climbed sharply once temperature passed a point. Which is a threshold, not a slider.</p>\n<p>The other study went looking for the effect inside the range people actually use. Renze and Guven ran nine models through multiple-choice exams, compared temperatures from 0.0 to 1.0, and found no statistically significant difference in accuracy anywhere in that range.</p>\n<p>They also pushed one model past it. GPT-3.5 held around 61 percent accuracy at temperature 1.0 and 56 percent at 1.4. At 1.5 it dropped to 35 percent. At 1.6 it scored 5 percent, with 84 of its 100 answers coming back malformed.</p>\n<p>That's one model on one exam, so it's a data point rather than a law. But it puts the cliff somewhere north of 1.4, which is a long way from 0.7.</p>\n<p>So the honest version. The tradeoff is real and it mostly lives outside the range you're working in. Between 0 and 1, where nearly every application sits, careful studies struggle to find an effect on correctness at all.</p>\n<p>Which makes most temperature tuning theater. Dropping from 0.7 to 0.3 to make a model &quot;more accurate&quot; is adjusting a knob that, in that range, is mostly adjusting variety.</p>\n<p>Renze and Guven do recommend 0.0 for problem solving, which sounds like the opposite of what I just said. Their reason is reproducibility, not accuracy. Pin it so the same input gives you the same output, and you're not trading away anything to do that.</p>\n<h2>What it's actually good for</h2>\n<p>Not accuracy. Variety.</p>\n<p>Low temperature buys consistency, which matters when the same input has to produce the same output. It costs you range. Models near 0 get repetitive and a little flat, reaching for the same phrasings, because you're always taking the safest available word.</p>\n<p>High temperature buys range, which matters when you're generating options and want them to genuinely differ from each other.</p>\n<p>What finally made it click for me is that temperature isn't a quality setting, it's a spread setting. It sets how far from the model's first instinct you're willing to let it go.</p>\n<h2>The zero that isn't zero</h2>\n<p>Temperature 0 gets described as deterministic. Same input, same ranking, always take the top one, same output. In theory that holds.</p>\n<p>In practice it doesn't, and <a href=\"https://platform.claude.com/docs/en/api/messages\">Anthropic's own API docs</a> say so plainly: even at temperature 0, results won't be fully deterministic.</p>\n<p>The usual explanation is floating-point math and GPU concurrency. A 2025 write-up from Thinking Machines Lab called <a href=\"https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/\">Defeating Nondeterminism in LLM Inference</a> went after that explanation and found it mostly wrong. The real culprit is batching. Production servers group incoming requests together, and the size of that group shifts minute to minute with traffic. Some GPU operations pick a different computation strategy depending on batch size, which produces tiny numeric differences, and when two tokens sit nearly tied for first place a tiny difference is enough to flip which one wins.</p>\n<p>So the answer you get can depend on how many other people happened to hit the server at the same moment as you.</p>\n<h2>Where that leaves me</h2>\n<p>The model never chose your answer.</p>\n<p>It handed over a ranked list of possibilities and something else reached in. When the output surprises you, that surprise came from the sampling step, not from the thing you think you're talking to.</p>\n<p>Small thing to find out, but it changed how I think about what I'm doing every time I open a chat.</p>\n<h2>Sources</h2>\n<p>Matthew Renze and Erhan Guven, <a href=\"https://aclanthology.org/2024.findings-emnlp.432/\">The Effect of Sampling Temperature on Problem Solving in Large Language Models</a>, Findings of EMNLP 2024 (<a href=\"https://arxiv.org/abs/2402.05201\">arXiv:2402.05201</a>). Their null result comes from three narrower comparisons rather than one full grid across every model, prompt technique, and exam, and the numbers above 1.0 are GPT-3.5 alone. <a href=\"https://arxiv.org/abs/2406.07070\">HalluDial</a> (2024) for the hallucination rate curve, which is a secondary analysis in that paper rather than its headline result. Thinking Machines Lab, <a href=\"https://thinkingmachines.ai/blog/defeating-nondeterminism-in-llm-inference/\">Defeating Nondeterminism in LLM Inference</a> (2025), on batch invariance. Anything wrong here is mine, not theirs.</p>",
      "image": "https://quionie.com/og/why-ai-gives-different-answers.png",
      "date_published": "2026-06-06T12:00:00.000Z",
      "date_modified": "2026-06-06T12:00:00.000Z",
      "tags": [
        "AI temperature",
        "sampling",
        "LLM determinism",
        "nondeterminism",
        "model settings"
      ],
      "authors": [
        {
          "name": "Quionie Gaban",
          "url": "https://quionie.com"
        }
      ]
    },
    {
      "id": "https://quionie.com/blog/how-to-tell-when-ai-is-making-things-up",
      "url": "https://quionie.com/blog/how-to-tell-when-ai-is-making-things-up",
      "title": "How to Tell When AI Is Making Things Up",
      "summary": "Everyone knows AI makes things up. Far less clear is which kinds of mistakes it makes, which models are worse, and what actually catches them. Here are the numbers.",
      "content_html": "<p>Everyone knows AI makes things up. The useful questions are which kinds of mistakes it makes, which models are worse, whether the numbers proving that mean anything, and what to do about it besides &quot;double-check important stuff.&quot;</p>\n<p>That last one is useless advice. Check it against what? You asked because you didn't know the answer.</p>\n<p>So I went looking for the version with numbers in it. A few of the answers changed how I use these tools.</p>\n<h2>Hallucination is one word for six different failures</h2>\n<p>The term got stretched to cover every way a model can be wrong, which matters because the fix is different for each one.</p>\n<p><strong>Fabrication.</strong> The model invents something that doesn't exist. A case citation, a study, a function in a library. Usually well-formed, because it's generating what a plausible instance of that thing looks like.</p>\n<p><strong>Misgrounding.</strong> The source exists but doesn't say what the model claims. This is the sneaky one. The citation is real, the link works, you click through and skim and it looks about right. Stanford's legal AI study found systems citing real cases that didn't support the proposition, and mischaracterizing real holdings, which is much harder to catch than an invented case.</p>\n<p><strong>Outdated.</strong> Correct as of the training data, wrong now. Prices, versions, who runs what, what a law says. The model has no sense of its own recency.</p>\n<p><strong>Incompleteness.</strong> Nothing stated is false, but a load-bearing exception is missing. Common in anything with jurisdictional or conditional structure.</p>\n<p><strong>Sycophancy.</strong> You assert something wrong and the model builds a case for it. Stanford named this as one of the error types in legal tools specifically. Asked to support an incorrect premise, the systems often generated plausible arguments on fabricated or mischaracterized authority instead of correcting the premise.</p>\n<p>This one deserves more fear than it gets. Every other failure is the model being wrong on its own. This one is the model being wrong because you were, so your own error comes back with citations attached.</p>\n<p><strong>Inference gaps.</strong> Failures on things that follow trivially from what the model already knows. Models trained on statements shaped like &quot;A is B&quot; often fail on &quot;B is A,&quot; which researchers named <a href=\"https://arxiv.org/abs/2309.12288\">the reversal curse</a>. GPT-4 could answer who Tom Cruise's mother is and then fail on who Mary Lee Pfeiffer's son is.</p>\n<h2>Which model is best is the wrong question</h2>\n<p>There's no such thing as the model with the lowest hallucination rate. The ranking flips depending on what you're measuring, and the gap between benchmarks is enormous.</p>\n<p>Grounded summarization is the easy case. The model gets a document and has to summarize it faithfully, and a hallucination is any claim the source doesn't support. That's what <a href=\"https://github.com/vectara/hallucination-leaderboard\">Vectara's HHEM leaderboard</a> measures, and top-ranked models score in the low single digits.</p>\n<p>That number gets quoted as &quot;hallucination rates are down to about 2 percent.&quot; It isn't. It's the score on a task where the answer is sitting inside the text you handed it. Two things get lost in the quoting. The models scoring lowest tend to be smaller ones, not the frontier reasoning models people are actually using, and when Vectara moved to longer and harder source documents the flagship reasoning models all landed above 10 percent.</p>\n<p>Open-domain factual recall is a different universe. No source, nothing but training. On OpenAI's PersonQA, models from a single generation ran from 16 percent to 48.</p>\n<p>Domain-specific real work is worse still, which is section five.</p>\n<p>So a model isn't reliable or unreliable. It's close to reliable at staying faithful to text you gave it, shakier at recalling facts, and worse in specialized domains. Pick per task rather than per brand, and treat any quoted hallucination percentage as meaningless until you know which benchmark produced it.</p>\n<h2>Newer is not always better</h2>\n<p>In OpenAI's own <a href=\"https://openai.com/index/o3-o4-mini-system-card/\">system card</a> from April 2025, the o3 reasoning model hallucinated on 33 percent of PersonQA prompts. Its predecessor o1 hallucinated on 16 percent. o4-mini came in at 48. The newer, more capable, more expensive models made things up at two and three times the rate.</p>\n<p>The obvious explanation is that o3 simply makes more claims per answer, so there are more chances to be wrong. OpenAI offered that much and then said plainly that they don't know why it happens and more research is needed. I want to be careful here, because the tidy story going around is that reasoning chains help on math and hurt on facts, and that isn't an established finding. It's a guess people repeat.</p>\n<p>The effect does show up elsewhere. DeepSeek's V3 scored 3.9 percent on grounded summarization and its reasoning model R1 scored 14.3. But it isn't uniform. <a href=\"https://arxiv.org/abs/2505.23646\">Research from 2025</a> found R1 improving on some fact-seeking benchmarks depending on how it was post-trained, and points at the post-training pipeline rather than reasoning length as the thing that matters.</p>\n<p>What's solid enough to act on: a newer, better-reasoning model is not automatically a more factual one, and for factual work &quot;use the newest&quot; can be exactly backwards. Check it rather than assuming it.</p>\n<h2>The benchmarks are contaminated</h2>\n<p>When a model scores well on a test, benchmark questions leak into training data, so it can score well by having effectively seen the exam.</p>\n<p>The numbers aren't small. Meta's own Llama 2 paper reported that over 16 percent of MMLU test examples had detectable overlap with the pretraining data, about 11 percent of them heavily. An <a href=\"https://arxiv.org/abs/2310.17589\">open audit of 15 popular benchmarks</a> found contamination from 1 percent to 45, rising over time. A separate audit of multilingual benchmarks found some as high as 91.</p>\n<p>Labs do filter. GPT-3's team removed anything sharing a 13-word run with a benchmark. GPT-4 used a different method, sampling three 50-character chunks from each test example and checking whether they appear in training data. Both only catch near-verbatim overlap. Researchers <a href=\"https://arxiv.org/abs/2311.04850\">demonstrated the hole</a> by training a 13-billion parameter model on rephrased benchmark questions, which sailed past the filters and produced GPT-4-level scores on those tests.</p>\n<p>The problem also regenerates. A clean benchmark gets published, copied across repositories and forums and derivative datasets, and within a training cycle or two it's back in the corpus.</p>\n<p>There's a real counterargument. An <a href=\"https://arxiv.org/abs/2410.03249\">ICML 2025 paper</a> found that moderate contamination gets substantially forgotten by the end of training, so small leakage doesn't automatically void a benchmark. Worth knowing, and worth reading the conditions: it holds for models trained well past compute-optimal, and forgetting depends on when the model saw the data. Contamination late in training sticks around.</p>\n<p>The field's answer has been benchmarks that refresh their questions on a rolling schedule, like LiveBench and LiveCodeBench, which helps and hasn't closed it.</p>\n<p>What I take from this. Benchmark scores are evidence, not proof, and the gap between a model's benchmark performance and its performance on your work is unknown from the outside. The only evaluation that fully belongs to you is one built from your own tasks, which is a genuine argument for keeping a private set of test questions you never publish.</p>\n<h2>What retrieval fixes, and what it doesn't</h2>\n<p>The standard fix is retrieval: give the model the source documents, have it answer from those, cite as it goes. Vendors marketed this as solving hallucination. LexisNexis used the phrase &quot;100 percent hallucination-free.&quot;</p>\n<p>Stanford's RegLab ran <a href=\"https://onlinelibrary.wiley.com/doi/full/10.1111/jels.12413\">the first preregistered test</a> of that claim on commercial legal research tools, using over 200 hand-built queries scored by legal experts, on products selling for thousands a month.</p>\n<p>Lexis+ AI hallucinated on 17 percent of queries. Westlaw's AI-Assisted Research on 33. GPT-4 with no retrieval at all on 43.</p>\n<p>Both directions of that matter. Retrieval genuinely helped, cutting the rate by half or more. And the best purpose-built, retrieval-grounded, professionally sold legal AI system still produced false or misleading information on roughly one query in six. LexisNexis later narrowed its hallucination-free claim to cover linked citations only.</p>\n<p>The accuracy figures underneath are worth as much as the hallucination ones. Lexis+ AI, the best system tested, answered 65 percent of queries accurately. Westlaw managed 42. So a third to half the time you're not getting a correct answer.</p>\n<p>Then a result that complicates the pessimism, which is why I'm including it. A <a href=\"https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5162111\">randomized trial</a> put 127 law students into three groups: a retrieval-based legal tool, a reasoning model with no retrieval, and no AI. The retrieval group produced about as many fabricated citations as the students working without AI, three in total. The reasoning-model group produced eleven. Both AI groups did substantially better work, and faster.</p>\n<p>So retrieval didn't make hallucination worse and the tools clearly helped. What retrieval doesn't do is remove the need to check, because the failure moves rather than disappearing. It becomes your job to notice that a real citation doesn't support the claim attached to it, and people are bad at that when the output looks authoritative and they're moving quickly.</p>\n<h2>Longer answers are more dangerous</h2>\n<p>Two findings from different directions land on this.</p>\n<p>Stanford found Westlaw's higher hallucination rate tracked with response length. Its answers averaged around 350 words against Lexis's 219. More words means more falsifiable propositions, which means more chances one of them is wrong.</p>\n<p>Separately, Anthropic looked at whether a model's stated reasoning reflects what it's actually doing, and found unfaithful reasoning chains were consistently longer than faithful ones. For Claude 3.7 Sonnet, 2,064 tokens against 1,439. For DeepSeek R1, 6,003 against 4,737.</p>\n<p>So the elaborate explanations were more likely to be covering the real reasoning than revealing it. Length is not thoroughness, and a detailed confident answer should raise your suspicion rather than settle it. The instinct runs the other way, which is the problem.</p>\n<h2>Asking the model if it's sure doesn't work</h2>\n<p>The obvious move is to ask whether it's confident, or to have it explain its reasoning so you can check the logic.</p>\n<p><a href=\"https://www.anthropic.com/research/reasoning-models-dont-say-think\">Anthropic tested this</a> by planting hints in prompts, seeing whether the model changed its answer, then checking whether its stated reasoning mentioned the hint. Claude 3.7 Sonnet acknowledged the hint about 25 percent of the time. DeepSeek R1 about 39. On misaligned hints, roughly 20 and 29.</p>\n<p>The starkest result came from environments containing a reward hack the model could exploit. In five of six, it used the exploit in over 99 percent of cases and mentioned it in under 2 percent of its explanations. Behavior driven almost entirely by something the model never mentioned while explaining itself.</p>\n<p>Faithfulness also dropped on harder questions, by 44 percent for Claude on the harder benchmark, which is exactly where you'd want the explanation to be real.</p>\n<p>There's a reasonable counterargument that some of this is incompleteness rather than dishonesty, since compressing distributed computation into a linear narrative loses information. A <a href=\"https://arxiv.org/abs/2512.23032\">2025 paper</a> found much of the apparent unfaithfulness disappears when models get bigger token budgets, with verbalization rising as high as 90 percent. Fair, and it doesn't change the practical conclusion. A model's explanation of its reasoning is generated output, not a report from inside, so it can't serve as verification.</p>\n<p>Same for confidence language. &quot;I'm certain&quot; is a stylistic property that got selected for during preference tuning. It carries almost no information about whether the answer is right.</p>\n<h2>What actually works</h2>\n<p>The method with the strongest research behind it is simple, and it comes from a <a href=\"https://www.nature.com/articles/s41586-024-07421-0\">2024 Nature paper</a> by Farquhar, Kossen, Kuhn and Gal.</p>\n<p>Ask the same question several times in separate sessions and look at whether the answers mean the same thing. Not whether they're worded alike. Whether they agree.</p>\n<p>When a model knows something, it converges. Different phrasings, same content. When it's making something up, the fabrications diverge, because there's nothing anchoring them. The paper clusters answers by meaning rather than wording and measures the spread, which they call semantic entropy. High spread flags likely confabulation. It works across datasets and tasks without task-specific tuning, which is why it ran in Nature.</p>\n<p>You can do the human version in about ninety seconds for free. Ask three times in three fresh conversations and compare the substance. If the specifics move, especially names, numbers, dates and citations, stop trusting the answer.</p>\n<p>The critical detail is separate sessions. Asking again in the same conversation is close to worthless, because the previous answer is sitting in the context and the model will stay consistent with it. At that point you're testing whether it can read what it just said. If you want the mechanism behind why fresh sessions diverge at all, <a href=\"https://quionie.com/blog/why-ai-gives-different-answers\">it comes down to sampling</a>.</p>\n<h2>The checking system</h2>\n<p>Sort by failure mode first. If the model is summarizing text you provided, risk is low and the check is fast, because you have the source. If it's recalling facts from training, risk is much higher. Anything with a name, a number, a date or a citation is the high-risk category regardless of task.</p>\n<p>Re-ask in fresh sessions. Three times, compare substance. This catches fabrication efficiently, because invented specifics don't survive resampling.</p>\n<p>Click every citation and read the part that supports the claim. Not the title, not the abstract. The specific passage. Misgrounding is the failure that survives a casual check, and it's the one Stanford found in professional-grade tools.</p>\n<p>Treat length as risk. Long detailed output contains more claims and a better chance one is wrong, and elaborate reasoning correlates with unfaithful reasoning.</p>\n<p>Watch your own premise. If you asked in a way that assumed something, sycophancy means the answer may be built on your assumption. Occasionally ask the opposite question and see whether it argues that side just as well. If it does, it doesn't know.</p>\n<p>Don't ask it to check itself. Its explanation isn't a window into its process, and it will confidently confirm.</p>\n<p>Use the newest model for reasoning and verify before trusting it on facts.</p>\n<p>Build a private eval. Twenty questions from your own work where you know the answers. Run them when you switch models. It's the only benchmark that can't be contaminated and the only one measuring your actual use case.</p>\n<h2>The error you can't catch</h2>\n<p>Everything above has the same blind spot.</p>\n<p>Consistency checking catches confabulation, meaning arbitrary invention. It does not catch consistent error, where the model reliably produces the same wrong answer every time. The Nature paper says so directly. Semantic entropy is built to detect answers that would change on resampling, which means errors that don't change are invisible to it.</p>\n<p>Those happen when something wrong was well represented in training. A popular misconception, a fact that was true for years and isn't now, a widely repeated myth. The model isn't uncertain and inventing. It's confidently reproducing something false, and it'll reproduce it identically across ten fresh sessions.</p>\n<p>Ask three times, get the same answer, feel reassured. That's the failure this method is structurally blind to.</p>\n<p>There's no clever prompt for it. The only defense is a source that isn't the model, which is the unglamorous conclusion that verification has to come from outside the thing being verified.</p>\n<h2>So how do you know</h2>\n<p>You don't, fully. That's the real answer and I'd rather say it than pretend.</p>\n<p>What you can do is stop treating reliability as a property of the model and start treating it as a property of the setup.</p>\n<p>The same model is close to reliable summarizing a document you handed it and closer to a coin flip on specialized domain questions. That range isn't the model having good and bad days. It's how much of the answer had to come from inside it. The more an answer depends on <a href=\"https://quionie.com/blog/how-large-language-models-work\">what's in the weights</a>, the less you can trust it. The more it depends on text you supplied, the more you can.</p>\n<p>So the working question isn't whether the model is accurate. It's how much of this answer the model had to invent, and whether you gave it what it needed not to.</p>\n<p>The law student trial is the part I keep coming back to. The retrieval tool didn't add fabrications and it made the work meaningfully better and faster. The reasoning model without retrieval improved the work too and brought three times the fabricated citations along with it. Same students, same tasks, different setup, different failure rate.</p>\n<p>That's the whole thing, really. Professionals use instruments with known error rates all the time, and it works because checking is built into the process rather than because the instrument is trusted. The failure mode isn't using AI for things it gets wrong sometimes. It's using it without the check, at speed, on things that matter, and finding out later.</p>\n<h2>Sources</h2>\n<p>Magesh, Surani, Dahl, Suzgun, Manning and Ho, <a href=\"https://onlinelibrary.wiley.com/doi/full/10.1111/jels.12413\">Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools</a>, Journal of Empirical Legal Studies (2025), for the 17, 33 and 43 percent figures, the accuracy rates, the response-length correlation and the sycophancy error type. Schwarcz et al., <a href=\"https://papers.ssrn.com/sol3/papers.cfm?abstract_id=5162111\">AI-Powered Lawyering</a>, for the 127-student trial. OpenAI's <a href=\"https://openai.com/index/o3-o4-mini-system-card/\">o3 and o4-mini system card</a> (April 2025) for the PersonQA rates, and <a href=\"https://arxiv.org/abs/2505.23646\">Are Reasoning Models More Prone to Hallucination?</a> (2025) for the post-training account. <a href=\"https://github.com/vectara/hallucination-leaderboard\">Vectara's HHEM leaderboard</a> for grounded summarization. Touvron et al. (2023) for the Llama 2 contamination disclosure, Li et al., <a href=\"https://arxiv.org/abs/2310.17589\">An Open-Source Data Contamination Report</a> (2024) for the 1 to 45 percent range, Yang et al., <a href=\"https://arxiv.org/abs/2311.04850\">Rethinking Benchmark and Contamination</a> (2023) for rephrased samples defeating filters, and Bordt et al., <a href=\"https://arxiv.org/abs/2410.03249\">How Much Can We Forget about Data Contamination?</a> (ICML 2025) for the counterargument. Chen et al. (2025) at Anthropic on <a href=\"https://www.anthropic.com/research/reasoning-models-dont-say-think\">chain-of-thought faithfulness</a>, with Turpin et al. (2023) as the foundational work and Zaman and Srivastava (2025) for the <a href=\"https://arxiv.org/abs/2512.23032\">incompleteness critique</a>. Farquhar, Kossen, Kuhn and Gal, <a href=\"https://www.nature.com/articles/s41586-024-07421-0\">Detecting Hallucinations in Large Language Models Using Semantic Entropy</a>, Nature (2024). Berglund et al. (2023) for <a href=\"https://arxiv.org/abs/2309.12288\">the reversal curse</a>. Anything wrong here is mine, not theirs.</p>",
      "image": "https://quionie.com/og/how-to-tell-when-ai-is-making-things-up.png",
      "date_published": "2026-05-03T12:00:00.000Z",
      "date_modified": "2026-08-22T12:00:00.000Z",
      "tags": [
        "AI hallucination",
        "fact checking AI",
        "LLM benchmarks",
        "retrieval augmented generation",
        "AI reliability"
      ],
      "authors": [
        {
          "name": "Quionie Gaban",
          "url": "https://quionie.com"
        }
      ]
    },
    {
      "id": "https://quionie.com/blog/how-i-think-about-building",
      "url": "https://quionie.com/blog/how-i-think-about-building",
      "title": "Seven Things I Learned from Building",
      "summary": "I didn't come from a CS background. I just kept building long enough to start noticing patterns. Seven of them, mostly learned by doing the wrong version first.",
      "content_html": "<p>I didn't come from a CS background. I just kept building long enough to start noticing the same mistakes show up twice.</p>\n<p>Here are seven of them. Most I learned by doing the wrong version first, which is the slow way to learn anything and also the only way it sticks.</p>\n<h2>The hard part is deciding what to build</h2>\n<p>The building isn't the bottleneck anymore. Claude Code and Codex get you most of the way in an afternoon. So the constraint moved. Building got cheap. Deciding what's actually worth building didn't.</p>\n<p>My rule now: I have to write it as one sentence before I open anything. &quot;This does X for Y.&quot; If I can't finish that sentence, the idea isn't ready and I don't start.</p>\n<h2>Start smaller than feels reasonable</h2>\n<p>Now that building is fast, everything wants to be big, because a big idea feels basically free to start. It isn't. A bigger first version just means more things you can be wrong about, and you find that out later, after you've sunk the time.</p>\n<p>So I cut the idea down until it feels almost too small to bother with. One feature, one user, one thing it does. If that works I keep pulling. If it doesn't, I found out in a day instead of three weeks.</p>\n<h2>Ship the first time it works, not the first time it's good</h2>\n<p>Shipping teaches you something polishing can't: how someone actually uses the thing. No amount of tweaking gives you that. Every extra day of polish is a day you don't have the only information that matters.</p>\n<p>If I can get through it end to end once without it breaking, it goes out. Whatever's ugly after that, I fix against real feedback instead of guessing.</p>\n<h2>Notice when you're doing fake work</h2>\n<p>A lot of what feels like work is avoidance wearing a productive costume. Renaming variables. Redoing spacing that was fine. Reorganizing files that already worked. You feel busy and nothing moves.</p>\n<p>When that happens I stop and ask what I'm avoiding. The answer is almost always the actual task. Usually it's the hard part, or the part where I'd have to show it to someone.</p>\n<h2>Design is part of whether it works</h2>\n<p>I came from a creative and video background, so this one was obvious to me before the code was. People decide how they feel about something in about a second. If it looks broken, they assume it is, before they ever try it.</p>\n<p>At some point I stop adding features and go fix spacing, hierarchy, and what the eye lands on first. Fewer things, arranged clearly, usually beats more things in a pile every time.</p>\n<h2>If it's hard to explain, it's probably too complex</h2>\n<p>Every extra feature feels like progress and adds a new way for things to break. I've rebuilt simpler versions of almost everything I've made. The v2 is usually the v1 with half the parts taken out.</p>\n<h2>Build from friction, not from ideas</h2>\n<p>Everything I've actually finished started as something that annoyed me. Ideas are easy to have and easy to drop, because nothing pulls you back to them. A real annoyance keeps pulling.</p>\n<p>So I stopped collecting &quot;good ideas&quot; and started fixing small things that get in my way. Those are easy to scope, easy to test, and I already know at least one person wants it. Me.</p>\n<p>That's most of it. I'll probably rewrite half of these once I figure out why they're wrong.</p>",
      "image": "https://quionie.com/og/how-i-think-about-building.png",
      "date_published": "2026-02-26T12:00:00.000Z",
      "date_modified": "2026-02-26T12:00:00.000Z",
      "tags": [
        "building software",
        "shipping",
        "side projects",
        "product lessons"
      ],
      "authors": [
        {
          "name": "Quionie Gaban",
          "url": "https://quionie.com"
        }
      ]
    }
  ]
}
