March 5, 2026

《Everything Is Predictable》书摘

You may say that you assume perfect ignorance, but there are different kinds of “ignorance,” and you have to pick one.

Introduction: A Theory of Not Quite Everything

The general rule in psychiatry is: if you think you’ve found a theory that explains everything, diagnose yourself with mania and check yourself into the hospital.

What Bayes’ theorem does is tell you how much you should change your belief. But in order to do that, you have to have a belief in the first place.

For other situations, it’s much more difficult. If you want to know how likely it is that Russia will invade Ukraine, what’s your prior probability? How often Russia has invaded Ukraine per year? How often one country invades another? How often one country invades another when they have just sent a whole load of tanks to that country’s border?

And we’re Bayesian at a deeper level too. Our brains, our perception, seem to work by predicting the world—prior probabilities—and updating those predictions with information from our senses: new data. Our conscious experience of the world can be best described as our priors. I predict, therefore I am.

Chapter One: From ‘The Book of Common Prayer’ to the Full Monte Carlo

The family behaved as you’d expect a wealthy, educated family of the time to behave.

William Whiston, Isaac Newton’s successor as the Lucasian professor of mathematics at Cambridge, was another associate of Bayes, and at one breakfast the two men had together he asked Bayes whether the sermon at the local Anglican church that weekend would include the Creed of Athanasius, which lays out the doctrine of the Trinity. Whiston said he would leave the service if so, and Bayes reassured him it would likely not.

Bayes was a committed supporter of Newton. “Some [Nonconformists] were hesitant to teach mathematics,” says Bellhouse, “in case it led to Newtonian science, and from there to atheism. But a much larger group among the Nonconformists said that it’s important to study mathematics—you need to understand God’s universe.”

A century or so later, in 1654, Antoine Gombaud, a gambler and amateur philosopher who called himself the Chevalier de Méré, was interested in the same questions, for obvious professional reasons.

(We should take a moment, here, to recognize the absolutely heroic amount of gambling that Gombaud must have been doing in order to be able to tell that his 52 percent bet was coming off, but his 49 percent bet wasn’t. Apparently, he had deduced, correctly, that you need twenty-five rolls of the dice, not twenty-four, for it to be a good bet. Gombaud was a man who enjoyed his dice-rolling.)

This led Gombaud to raise another question with Pascal.

Pascal found this fascinating, and exchanged a series of letters discussing the problem with his contemporary, Pierre de Fermat, of Last Theorem fame.

This is where Pascal and Fermat come into the picture. They realized the key point: It’s not how close to the finish you are, or how far from the start you’ve come, that matters. It’s the number of possible outcomes that remain, and how many of those outcomes favor one player over the other.

But Bernoulli wanted to take it further than that. His insight was that there are three components to this question: how big a sample you take, how close to the true answer you need to be, and how confident in your answer you need to be. He realized that you can never be truly certain of the actual ratio. What you can have instead, he said, is “moral certainty”—that is, a given degree of confidence (Which is an inconveniently large sample size for someone working in early modern Europe without the benefit of a computer or ready access to psych undergraduates willing to take part in social science research for beer money.

The abrupt ending of Ars Conjectandi after that line, says Stigler, suggests that “Bernoulli literally quit when he saw the number 25,500, mustering strength only to add one further sentence.”)

Bernoulli actually did this stuff, which may be why his book took twenty years to write and he never actually finished it.

What neither Bernoulli nor de Moivre were able to answer was what came to be known as inverse probability, although it’s really the heart of what we want probability to do. What we want, or at least what science wants, out of statistics, is to answer: Given the results I’ve seen, what can I say about my hypothesis?

We know, then, that Bayes was thinking about probability, and specifically inferential or inverse probability—remember, that’s “How likely is it that a hypothesis is true, given this data?”—as opposed to sampling probability—which asks, “How likely am I to see this data, given this hypothesis?”—

At the time, Price was a much more well-known man than his older friend Bayes. He was well connected with radical thinkers: notably, he was friends with several of the Founding Fathers of the American Revolution. He exchanged letters with Thomas Jefferson and Benjamin Franklin, both of whom visited him in Newington Green, as did John Adams, the second president of the United States.

Price was also friends with the philosophers David Hume (of whom more later) and Adam Smith, and with William Pitt the Elder, the politician. In all he seems to have been a pretty cool and impressive guy with cool and impressive friends, famous across England and America, and it is strange how little known his name is now.

Part of the reason it took so long for him to publish it was that Price didn’t just go over it for typos and misplaced commas; he was more than, in the words of the historian of statistics Stephen Stigler, “a loyal secretary.” Price had visions of his own for the work: while Bayes wrote the first half of the paper, the second half, containing all the possible practical applications of the theorem, was all Price. Bayes had no interest in applied statistics: his work, in this paper and all others, was “all theory with not a hint of application.” But Price was—to quote Stigler again—“the first Bayesian.”

Hume saw Price’s work. In fact there was a rather lovely exchange between the two men, which I will tell you about just because it’s so nice to see two people who disagree so profoundly on something so important—God versus no God—behaving so civilly.

Then the two men met, and Hume, by all accounts a very affable and reasonable man, left Price both charmed and ashamed: in a second edition, he removed every disobliging comment and added a rather apologetic introduction saying that one shouldn’t accuse one’s opponent of bad faith or disingenuousness.

Condorcet, and then Laplace himself, went on to acknowledge that Bayes got there first, hence “Bayes’ theorem” rather than “Laplace’s theorem,” even though Laplace’s treatment of the problem was probably the more impressive.

You may say that you assume perfect ignorance, but there are different kinds of “ignorance,” and you have to pick one.

Again, I want to be careful here. Many at the time who were very much on the progressive, liberal end of society were also pro-eugenics. Marie Stopes, the campaigner for birth control, abortion, and women’s rights, was also a major supporter of eugenics. John Maynard Keynes, the great economist and liberal (and my great-great-uncle, so I ought to declare an interest here when I try to downplay the awfulness of it all), was another. Sidney and Beatrice Webb, George Bernard Shaw, Bertrand Russell, all heroes of the socialist and liberal movements, were in favor of selective breeding of humanity in order to create a better, more perfect society.

The term wasn’t, as it is now, associated so heavily with the right. (In fact, when I was writing stories in the late 2000s and early 2010s about things like embryo screening for disability, in vitro fertilization, and mitochondrial donation, misleadingly named “three-person babies,” it was mainly the religious right who criticized them as “eugenics.”)

This is as neat a description of the Bayesian-frequentist disagreement as you could ask for, I think. Bayesianism treats probability as subjective: a statement about our ignorance of the world. Frequentists treat it as objective: a statement about how often some outcome will happen, if you do it a huge number of times.

The trouble is that those two claims can’t both be true. If we’re uniformly ignorant of the length of the sides, then it’s more than 70 percent likely that the square will have an area less than 50 square centimeters. (A 7 cm × 7 cm square would have an area of 49 square centimeters, because 7 × 7 = 49.) Meanwhile, if we’re uniformly ignorant of the area, then it’s 75 per cent likely that the sides will be at least 5 cm long (5 × 5 = 25). Again, as with Boole’s criticism, there are different kinds of ignorance, and we are ignorant of which one to use.

Harold Jeffreys, a Cambridge geologist, was the key figure in early scientific Bayesianism—he wrote that Bayes’ theorem “is to the theory of probability what Pythagoras’ theorem is to geometry.” While Fisher worked with pea plants and mice, experiments that gave precise answers and could be repeated as many times as required, Jeffreys looked at the propagation of seismic waves through the Earth. It was Jeffreys who first showed, in 1926, that the core of the Earth is liquid—and, in fact, that the outer mantle is mainly silicon-based stone, while the inner core is mainly iron and nickel.

I mentioned this on Twitter and Sir David got in touch to say that, alas, “Our performance of ‘The Full Monty Carlo’ was before the smartphone era, so no recordings exist.” (“Who would want to see a video of six male professors of Bayesian statistics taking their clothes off in front of a screaming crowd in a Spanish nightclub?” he went on to ask, in my view entirely misjudging the nature of the modern internet.) There

We shouldn’t overstate the enmity between Bayesians and frequentists—especially nowadays, but even when the Valencia conferences were going on and Bayesianism was growing in confidence. Grieve remembers that after roaring debates at the Royal Statistical Society, “they’d slag each other off in public, in those arguments and disputes. But afterwards they’d be going to dinner in a taxi together. While it does sound tough, a lot of them were great friends.” He remembers reading a paper by Spiegelhalter, a fellow Bayesian, in the 1990s, and thinking, “They had joined a recent tradition of avoiding controversy and pursuing the practical benefits, even though the former was more entertaining. I do think there was a bit of entertainment in it.”

Chapter Two: Bayes in Science

Another found that exposing people to the smell of fish made them more suspicious, because—seriously—it “smells fishy.”

That is the idea, sure. But it’s not as straightforward as that. The easiest way to get a p < 0.05 result—that is, something that you’d only see by coincidence one time in twenty—is to do twenty experiments, and then publish the one that comes up. That’s exactly what the “False Positive Psychology” people did:

He called for researchers to “analyze the sexes separately, make up new composite indexes… reorganize the data to bring them into bolder relief…. The data may be strong enough to justify recentering your article around the new findings and subordinating or even ignoring your original hypotheses…. Think of your dataset as a jewel. Your task is to cut and polish it, to select the facets to highlight, and to craft the best setting for it.” It’s not intended as a call for p-hacking, but “recentering your data around the new findings” is exactly what both the “False Positive” guys and Wansink were doing, and as they demonstrated,

But you could go deeper and say that the underlying cause of the replication crisis is even more basic: it’s that science, like Jakob Bernoulli three hundred years ago, is doing sampling probabilities, not inferential probabilities.

If you were an unscrupulous researcher, or even if you were just a naive one, you could very easily find apparently statistically significant results in noisy data when there’s nothing really there, just by checking your data a few times before you originally intended to.

If you were a Bayesian, though, this wouldn’t be a problem. You’ve already got the data from your priors—whatever they are—so each new data point coming in moves your opinion much less. And, of course, each new result forms part of the new prior for your next bit of information.

But then there’s epistemic uncertainty, from epistēmē, the Greek word meaning “knowledge.” That’s what Cassie Kozyrkov was demonstrating above. If you flip a coin, then you catch it, but don’t look at it—then there’s no aleatory uncertainty. The result is there, it’s happened, that’s it. Still, though. You don’t have any new information. As far as you’re concerned, the question is no closer to being resolved than it was before.

To the extent that the brain is a Bayesian machine—another idea we’ll come back to later—this is pretty much what it’s doing, when it predicts the world around you and updates it with new information from your senses.

But that doesn’t mean people have to pluck priors out of the air. There are reasonable ways of finding them in different circumstances. And then, of course, if your data is any good, your priors will be rapidly washed away.

In fact, there is an alternative model already in place—Alexandra Freeman of the Winton Centre for Risk and Evidence Communication at Cambridge University has launched a program called Octopus. It is a free repository for hypotheses, data, code, and methods. Freeman, a former journalist, told me, “When I moved from the media to academia, it struck me that academics are being given the exact same incentives as journalists—they’re pushed toward telling good stories, instead of doing good science. Journals encourage people to have high-impact publications, which they define as having high readership, short and to the point, carrying a message. It acts directly against what you actually want in a primary research record—which is everything there, in detail, so people can follow it.”

But Eric-Jan Wagenmakers makes a point that I also agree with, which is that Bayesianism is aesthetically more pleasing. “There’s something in Bayes,” he says. “Everything is coherent; you don’t have internal inconsistencies. In frequentism you can find all these anomalous cases, and people say it’s an anomaly but only in this situation, but it always feels ugly.

Chapter Three: Bayesian Decision Theory

“Thus, in our reasoning we depend very much on prior information to help us in evaluating the degree of plausibility in a new problem,” says Jaynes. “This reasoning process goes on unconsciously, almost instantaneously, and we conceal how complicated it really is by calling it common sense.”

As a great Bayesian thinker put it: I beseech you, in the bowels of Christ, think it possible that you may be mistaken.

Again, this is unavoidable. If some evidence is strongly expected, then it can’t move your beliefs very much; it’s already part of the model of the world that you’ve built. But if something really unexpected happens—or, in this case, if something expected doesn’t happen—it should move your posterior belief significantly.

But the third string is truly random. If you wanted to describe it to the millionth digit, you’d have to write it out to a million digits. There is no shortcut that would let you do it any faster; you cannot compress it at all. The minimum message length of any given output is how short your description can be.

That’s how much you should trade off between complexity and good fit. If an extra bit of information in your program doesn’t allow you to halve the search space, then it’s not paying its way. It’s not compressing the data—you’re just shifting it into the program, rather than the data.

So in choosing between two or more hypotheses, you should (in theory) be able to look at which is the more complex, and—all else being equal—assign higher prior probability to the one that would be simpler to write as a computer program, and with each extra bit of information in the program reducing its probability by half. There are other ways of producing priors, but minimizing complexity like this is a key one.

Chapter Four: Bayes in the World

But if you reversed the framing—told people that the first program would definitely mean four hundred people would die, while the second program meant a one-third chance that nobody would die, and a two-thirds chance that all six hundred would—then the respondent numbers reversed too. More than 75 percent of people chose the gamble.

(In case you’re interested, Sanders got the highest average for trustworthiness, Clinton the highest for expertise, and Trump the lowest for both.)

No mathematics are involved at all. Experiments show that non-human animals use the same system—dogs catch Frisbees by keeping them in the same place in their eyeline, just as baseball outfielders do to catch fly balls.

The gaze heuristic is almost as accurate as actually calculating the trajectory, but computationally far simpler. The Royal Air Force realized during the Second World War that they could use the gaze heuristic to help guide fighters to intercept bombers, and that it would be much quicker than doing the math.

You can also think of it like this. Before Monty opens the door, there’s a one-in-three chance that you picked the correct one. In that case, it would be bad if you switched. But there’s a two-in-three chance that you picked a wrong one, and in that case, it would be good if you switched.

Or imagine that you play the game three hundred times. In one hundred of those, you pick the right door, and switching means that you lose. But in two hundred, you pick the wrong door, and switching means that you win.

But what’s crucial is that Monty, first, knows where the car is, and second, always opens a door with a goat behind it. Once that’s stipulated—or assumed—you can easily work out the odds with Bayes’ rule.

Not that he doubted their intelligence or their integrity, but he thought that perhaps everyone, when confronted with unexpected information, found ways of saying that it just showed that whatever they already believed must be true.

After several years, Tetlock assessed the results, and it turned out that the average forecaster did very little better than random guessing. In fact, in a memorable phrase that he would come to somewhat regret, Tetlock said they did no better than “a dart-throwing chimpanzee.”

He regretted it, as he wrote in Superforecasting three decades later, because people misunderstood it—they took it to mean that all experts were guessing randomly. But in fact there were distinct groups. Some thought the world was simple and could be explained (and predicted) simply—they had what Tetlock called “one big idea” that they rubber-stamped onto every situation. Others thought the world was complicated—that the specifics and details of each situation mattered, and that predictions were difficult and uncertain.

He called the big-idea people “hedgehogs” and the life-is-complicated people “foxes,” following a quote that Isaiah Berlin lifted from Archilochus, a Greek poet: “The fox knows many things, but the hedgehog knows one big thing.” And it was the hedgehogs who did no better than the chimpanzee, if I can mix my animal metaphors.

Most of us, though, don’t keep base rates in our mind like that, so our beliefs are swayed by every new bit of information. “If every time you get new data you start over,” says Manheim, “then obviously your estimates jump all over the place and will be way overfocused on whatever’s most recent.”

The average American works about forty hours, fifty weeks a year (says Tetlock; it sounds pretty bleak, but OK).

Fermi estimates are a way of employing the law of large numbers by yourself: you make several estimates of small things instead of one big estimate, and if there’s no reason why those errors should be systematically high or low, then they will tend to cancel each other out

I just did it and found, pleasingly, that my 80 percent guesses were correct 85 percent of the time, which isn’t bad. Years of reporting on this stuff have beaten the overconfidence out of me.

Chapter Five: The Bayesian Brain

That’s because, he said, the brain has to do a lot of work. The world as it appears on our retinas is messy: upside down and back to front, for a start (if you close your eyes and press the bottom-left of one eye, the resulting splodge of color appears in the top right of your visual field). It’s also distorted by the concave shape of the back of your eyeball, and it’s bumpy with the blood vessels that cover it. Worst of all, the eye is just badly designed, with the nerves from the retina pointing inward rather than out, so in order for them to get out to the brain the optic nerve has to come through the retina, leaving a big blind spot.

But both Seth and Frith cheerfully agree with me. “Consciousness is our model of the world, not the world,” says Frith. “The content of our perception is the content of these top-down predictions,” says Seth. So: consciousness is Bayesian.

And that idea has become central to what Friston, Seth, and others talk about when they talk about perception. The brain is not only passively perceiving, but actively seeking out information to reduce its uncertainty in the world. “You can frame it,” says Seth, “in terms of actions that are instrumental, to reach the desired goal now, or epistemic actions that maximize the information gain.”

So, tickling. The same applies. If you try to tickle yourself, your brain can predict the sensations it’s going to receive, with very high precision. If you were to stroke my palm and record my brain activity while you did so, you’d see a sudden spike in the number of neurons firing in the relevant bit of my cortex. But if I were to stroke it myself, there’d be very little increase. “When you touch yourself,” says Frith in his book, entirely deadpan, “your brain suppresses your response.”

That’s apparently what might be going on with depression. Your prior probability on some untrue belief, something like “I am a terrible person and everyone hates me,” is inappropriately high.

Now, psychedelics. They’re unusual drugs. They don’t particularly make you happy or energetic or anything; they just make you really interested in things. They make the world feel unfamiliar. “Have you ever, like, really looked at a tree, man?” That sort of thing.

What they do, in this model, is to flatten your priors. You never really look at a tree, or your hand, or whatever, because you have very strong prior beliefs about what trees are like, and those beliefs very successfully predict the information that will come from looking at a tree, so your brain basically discounts them. “Familiar thing, accurately predicted, move on.”

Conclusion: Bayesian Life

Once you’ve read the papers, if they both seem pretty good to you, then you’ll update toward accepting them. But unless you have total confidence in your ability to judge the paper entirely on its merits, then the new data won’t completely wash out your prior probability, and you’ll still judge the Einstein paper as more likely to be good.

The statistician George Box, he who sang “There’s No Theorem Like Bayes’ Theorem” at the first Valencia conference, had a saying: “All models are wrong, but some are useful.”

There’s a thing called a “flow state,” when you’re doing some activity, playing an instrument, playing sports or a video game, painting, whatever, and it just seems to work: that’s when your predictions are high-precision and they’re coming true every time.

It explains why, as we get older, we get more set in our ways. When we’re young, we have very little data about the world, so our priors are weak and new information can shift them easily.

First, you don’t need to think so much in terms of right and wrong, true or false. You can think in terms of how confident you are in a belief, and adjust it up and down, rather than rejecting or accepting it at some arbitrary threshold.