《连线》访谈
运气与技能解缠:成功这门科学 2012 年 11 月 16 日
Luck and Skill Untangled: The Science of Success 16 th November, 2012
我们周围的世界变化无常,且时常充满艰难。但伴随着我们数学工具的日益精密,我们反过来也提升了理解周围世界的能力。
The world around us is a capricious and often difficult place. But as we have developed our mathematical tools with increased sophistication, we have in turn improved our ability to understand the world around us.
这种现象出现在一个看似简单的地方:运气与技能之间的关系。我们很容易认识到,国际象棋大师战胜新手是技能使然,也会认为章鱼保罗预测世界杯比赛的能力纯属偶然。但除此之外的一切呢?
And one of the seemingly simple places where this occurs is in the relationship between luck and skill. We have little trouble recognizing that a chess grandmaster’s victory over a novice is skill, as well as assuming that Paul the octopus’s ability to predict World Cup games is due to chance. But what about everything else?
我的一位朋友迈克尔·莫布森(也是我一位合作者的父亲),很友善地通过电子邮件做了一次问答。
Michael Mauboussin, a friend of mine (and the father of one of my collaborators), was kind enough to do a Q&A via e-mail.
塞缪尔·阿贝斯曼:首先,技能和运气都是难以捉摸的东西。在书的开头,你试图为生活中这两个特征给出操作性定义。你会如何定义它们?
Samuel Arbesman: First of all, skill and luck are slippery things. In the beginning of the book, you work to provide operational definitions of these two features of life. How would you define them?
迈克尔·莫布辛:这是一个非常重要的起点,因为运气问题尤其会很快滑入哲学领域。所以我试着用一些实用的定义,足以让我们做出更好的预测。我直接从字典里拿来了技能的定义,它把技能定义为“在执行或表现中有效且自如地运用自己知识的能力”。这基本上就是说,你知道怎么做事,而且需要的时候能做得出来。明显的例子是音乐家或运动员——到了音乐会或比赛时间,他们随时准备表演。
Michael Mauboussin: This is a really important place to start, because the issue of luck in particular spills into the realm of philosophy very quickly. So I tried to use some practical definitions that would be sufficient to allow us to make better predictions. I took the definition of skill right out of the dictionary, which defines it as "the ability to use one’s knowledge effectively and readily in execution or performance." It basically says you know how to do something and can do it when called on. Obvious examples would be musicians or athletes — come concert or game time, they are ready to perform.
运气这事儿更微妙。我喜欢把运气想成有三个特征。第一,它发生在某个群体或个人身上。第二,它可以是好运,也可以是霉运。我并不是说它好坏对半开,而是说它确实有这两种味道。最后,当有理由认为本该发生别的结果时,运气就起了作用。
Luck is trickier. I like to think of luck as having three features. First, it happens to a group or an individual. Second, it can be good or bad. I don’t mean to imply that it’s symmetrically good and bad, but rather that it does have both flavors. Finally, luck plays a role when it is reasonable to believe that something else may have happened.
人们经常把运气与随机性混为一谈。我喜欢将随机性理解为系统层面的运作,而运气则发生在个体层面。如果我召集 100 个人,让他们猜抛硬币的结果,随机性告诉我们,有少数人可能连续猜对五次。如果你恰好是那五次全中的人之一,那就是你运气好。
People often use the term luck and randomness interchangeably. I like to think of randomness operating at a system level and luck at an individual level. If I gather 100 people and ask them to call coin tosses, randomness tells me that a handful may call five correctly in a row. If you happen to be one of those five, you’re lucky.
阿伯斯曼:技巧和运气在投资世界里极其重要。你书中大量体育案例让读者觉得,你其实相当懂体育。
Arbesman: Skill and luck are very important in the world of investing. And the many sports examples in your book make the reader feel that you’re quite the sports
粉丝。但这本书的想法是怎么产生的?有没有某个特定的时刻促使你动笔?
fan. But how did the idea for this book come about? Was there any specific moment that spurred you to write it?
莫布森:这个话题恰好横跨了我的诸多兴趣领域。首先,我向来热爱体育,既是参与者也是球迷。和许多人一样,我被迈克尔·刘易斯在《魔球》中讲述的故事深深吸引——奥克兰运动家队如何利用统计数据来更准确地理解场上表现。当你花时间研究运动员的数据统计时,很快就意识到,某些指标中运气的成分远比其他指标重要。例如,运动家队发现上垒率比打击率更能可靠地反映技能水平,同时他们也注意到,这种差异并未在球员的市场价格中得到体现。这便创造了一个机会——以低成本打造一支有竞争力的球队。
Mauboussin: This topic lies at the intersection of a lot of my interests. First, I have always loved sports both as a participant and fan. I, like a lot of other people, was taken with the story Michael Lewis told in Moneyball – how the Oakland A’s used statistics to better understand performance on the field. And when you spend some time with statistics for athletes, you realize quickly that luck plays a bigger role in some measures than others. For example, the A’s recognized that on-base percentage is a more reliable indicator of skill than batting average is, and they also noted that the discrepancy was not reflected in the market price of players. That created an opportunity to build a competitive team on the cheap.
第二,做投资这一行,实在很难不去想运气这件事。伯特·马尔基尔那本畅销书《漫步华尔街》,基本已经把这事说透了。
Second, it is really hard to be in the investment business and not think about luck. Burt Malkiel’s bestselling book, A Random Walk Down Wall Street, pretty much sums it up.
现在事实证明,市场并非真正的随机游走,但要区分市场的实际行为与随机性,还需要一定的洞察力。
Now it turns out that markets are not actually random walks, but it takes some sophistication to distinguish between actual market behavior and randomness.
第三,在我上一本书《思考,快与慢》中,我写过一章关于运气和技巧的内容,但觉得那部分处理得不够到位。所以我知道,这个话题还有很多可以深入探讨和展开的空间。
Third, I wrote a chapter on luck and skill in my prior book, Think Twice, and felt that I hadn’t given the topic a proper treatment. So I knew that there was a lot more to say and do.
最后,这个话题吸引我,是因为它横跨许多学科。虽然在不同领域里确实有一些相当出色的分析,但此前我并未真正看到过对技能与运气进行综合处理的著作。顺便提一下,
Finally, this topic attracted me because it spans across a lot of disciplines. While there are pockets of really good analysis in different fields, I hadn’t really seen a comprehensive treatment of skill and luck. I’ll also mention
我希望这本书非常实用:我不只是想告诉你世界上存在很多运气;我更想帮助你弄清楚如何以及为何能够应对它,从而做出更好的决策。
that I wanted this book to be very practical: I’m not interested in just telling you that there’s a lot of luck out there; I am interested in helping you figure out how and why you can deal with it to make better decisions.
阿贝斯曼:你们在一个纯运气与纯技能之间的连续光谱上,对几项运动进行了排名,篮球最靠技能端,冰球最接近运气端。
Arbesman: You show a ranking of several sports on a continuum between pure luck and pure skill, with basketball the most skillful and hockey the closest to the luck end:
而且这个排名并不完全显而易见,因为你提到询问了不少同事,很多人都猜得相当不准。(事实上,我记得你问过我这个问题,我也答错了。)你是如何得出这个排名的,而这些运动之间存在的结构性差异,又是什么导致了这种排名差异?
And the ranking is not entirely obvious, as you note that you queried a number of your colleagues and many were individually quite off. (I in fact remember you asking me about this and getting it wrong.) How did you arrive at this ranking and what are the structural differences in these sports that might account for these differences?
莫布森:我觉得这个分析很酷。我是从汤姆·坦戈那里学来的,他是一位备受尊敬的数据棒球分析师,在统计学里这叫“真分数理论”。它可以用一个简单的公式来表达:
Mauboussin: I think this is a cool analysis. I learned from Tom Tango, a respected sabermetrician, and in statistics it’s called "true score theory." It can be expressed with a simple equation:
观察到的结果 = 技能 + 运气
Observed outcome = skill + luck
背后的直觉是这样的:假设你参加一次数学考试。你的成绩会反映你的真实水平——你实际掌握的知识有多少——再加上一些误差,这误差来自老师出的考题。有些时候,你考得比实际水平好,因为老师恰好只考了你复习过的内容;有些时候,你考得比实际水平差,因为老师恰好出了你没复习的题目。所以,你的成绩反映的是你的真实水平加上一点运气。
Here’s the intuition behind it. Say you take a test in math. You’ll get a grade that reflects your true skill — how much of the material you actually know — plus some error that reflects the questions the teacher put on the test. Some days you do better than your skill because the teacher happens to test you only on the material you studied. And some days you do worse than your skill because the teacher happened to include problems you didn’t study. So you grade will reflect your true skill plus some luck.
当然,我们已知方程中的一项——观察到的结果——而且我们可以估算运气。对一支运动队来说,估算运气相当简单。你假设该队参加的每一场比赛都由一次抛硬币决定。联赛中各队的胜负战绩分布服从二项式分布。因此,在确定了这两项之后,我们就可以估算技能以及技能贡献的相对比例。
Of course, we know one of the terms of our equation — the observed outcome — and we can estimate luck. Estimating luck for a sports team is pretty simple. You assume that each game the team plays is settled by a coin toss. The distribution of win-loss records of the teams in the league follows a binomial distribution. So with these two terms pinned down, we can estimate skill and the relative contribution of skill.
说得更技术性一些,我们会考察这些结果的方差,但直觉上,你是从已发生的事情中剔除运气,剩下的便是技能。这样一来,你就能评估这两者的相对贡献。
To be more technical, we look at the variance of these terms, but the intuition is that you subtract luck from what happened and are left with skill. This, in turn, lets you assess the relative contribution of the two.
排名中的某些方面说得通,而另一些则不那么一目了然。例如,如果一项比赛是一对一的,比如网球,而且赛程足够长,你基本可以确信,更优秀的选手会赢。随着参赛选手增加,运气的成分通常会上升,因为互动次数急剧增加。
Some aspects of the ranking make sense, and others are not as obvious. For instance, if a game is played one on one, such as tennis, and the match is sufficiently long, you can be pretty sure that the better player will win. As you add players, the role of luck generally rises because the number of interactions rises sharply.
我要强调三个方面。第一个方面涉及球员人数。但关键不只是人数多少,而是谁在掌控比赛。以篮球和冰球为例。冰球每方同时上场六名球员,篮球每方同时上场五名球员,看起来差不多。但优秀的篮球球员可以打满大部分,甚至整场比赛。而且你每次进攻都可以把球传给勒布朗·詹姆斯。因此,有天赋的球员能带来巨大差别。
There are three aspects I will emphasize. The first is related to the number of players. But it’s not just the number of players, it’s who gets to control the game. Take basketball and hockey as examples. Hockey has six players on the ice at a time while basketball has five players on the court, seemingly similar. But great basketball players are in for most, if not all, of the game. And you can give the ball to LeBron James every time down the floor. So skillful players can make a huge difference.
相比之下,在冰球比赛中,最优秀的球员在冰上的时间也只占全场比赛的三分之一多一点,而且他们无法有效控制球权。
By contrast, in hockey the best players are on the ice only a little more than one-third of the time, and they can’t effectively control the puck.
在棒球运动中,最佳击球手也不过是每九次上场打击中能多出场一两次。足球和美式橄榄球在任意时刻上场的球员数量也相似,但美式橄榄球队中几乎所有的进攻轮次都要经过四分卫之手。因此,如果比赛节奏需要通过一名技术型球员来传导,就会对球队的动态产生影响。
In baseball, too, the best hitters only come to the plate a little more frequently than one in nine times. Soccer and American football also have a similar number of players active at any time, but the quarterback takes almost all of the snaps for a football team. So if the action filters through a skill player, it has an effect on the dynamics.
第二个方面是样本规模。正如你在统计课上学到的那样,同一系统内,小样本的方差大于大样本。举例来说,一家每天只接生几个婴儿的医院,其女婴与男婴出生比例的方差,会远远高于一家每天接生几百个婴儿的医院。由于大样本倾向于剔除运气的影响,它们能更准确地反映技能水平。在体育领域,我对比了一场大学篮球赛和一场大学长曲棍球赛的控球次数。尽管长曲棍球比赛时间更长,但控球次数在……
The second aspect is sample size. As you learn early on in statistics class, small samples have larger variances than larger samples of the same system. For instance, the variance in the ratio of girls to boys born at a hospital that delivers only a few babies a day will be much higher than the variance in a hospital that delivers hundreds a day. As larger sample sizes tend to weed out the influence of luck, they indicate skill more accurately. In sports, I looked at the number of possessions in a college basketball game versus a college lacrosse game. Although lacrosse games are longer, the number of possessions in a
篮球比赛的优势大约是长曲棍球比赛的两倍。这意味着,技术更娴熟的队伍获胜的概率会更高。
basketball game is approximately double that of a lacrosse game. So that means that the more skillful team will win more of the time.
最后还有一点:比赛是怎么计分的。再拿棒球打个比方。一支队伍可以靠安打和保送让很多球员上垒,但如果出局时机不对,可能一个人都回不了本垒。理论上,一支队伍可以打出 27 支安打,但一分未得,而另一支队伍仅凭一支安打就能以 1 比 0 赢下比赛。
Finally, there’s the aspect of how the game is scored. Go back to baseball. A team can get lots of players on base through hits and walks, but have no players cross the plate, based on when the outs occur. In theory, one team could have 27 hits and score zero runs and another team can have one hit and win the game 1-0.
当然,这种情况非常、非常罕见,但它会让你体会到评分方法的影响力。
It’s of course very, very unlikely but it gives you a feel for the influence of the scoring method.
篮球是技能含量最高的运动。橄榄球和棒球相差不大,但棒球队的比赛场次是橄榄球队的 10 倍以上。
Basketball is the game that has the most skill. Football and baseball are not far from one another, but baseball teams play more than 10 times the games that football teams do.
换句话说,棒球近乎是随机的——即便打了 162 场比赛,最强的球队也只能赢下大约 60% 的比赛。冰球运动同样充斥着大量的随机性。
Baseball, in other words, is close to random — even after 162 games the best teams only win about 60 percent of their games. Hockey, too, has an enormous amount of randomness.
一个值得玩味的思路是:美国职业篮球联赛(NBA)和国家冰球联盟(NHL)连续两个赛季都遭遇了停摆。两个联盟的常规赛赛程都是 82 场。NHL 的停摆至今尚未解决,而外界期望他们能像去年的 NBA 那样,打一个缩水赛季。
One interesting thought is that the National Basketball Association and National Hockey League have had lockouts in successive seasons. Both leagues play a regular schedule of 82 games. The NHL lockout hasn’t been resolved, and there is hope that they will play a shortened season as did the NBA last year.
但关键点在于:即便赛季缩短,我们仍能判断 NBA 中哪些球队最优秀、从而有资格进入季后赛。如果 NHL 赛季只进行正常场次的一小部分比赛,结果就会非常随机。也许最顶尖的球队会占据一定优势,但你几乎可以确定,一定会出现一些意外。
But there’s the key point: Even with a shortened season, we can tell which teams in the NBA are best and hence deserve to make the playoffs. If the NHL season proceeds with a fraction of the normal number of games, the outcomes will be very random. Perhaps the very best teams will have some edge, but you can almost be assured that there will be some surprises.
阿布斯曼:你花了不少篇幅讨论均值回归这个现象。
Arbesman: You devote some attention to the phenomenon of reversion to the mean.
我们大多数人以为自己懂这个概念,但实际上常常搞错。在这个概念上我们容易犯哪些错,又为什么会这么频繁地犯错?
Most of us think we understand it, but are often wrong. What are ways we go wrong with this concept and why does this happen so often?
莫布辛:你的观察一针见血——
Mauboussin: Your observation is spot on:
听到“均值回归”这个词时,大多数人都会心领神会地点头。但如果你观察人们的实际行为,就会发现一桩又一桩例子,表明他们在行动中根本没有把均值回归考虑进去。
When hearing about reversion to the mean, most people nod their heads knowingly. But if you observe people, you see case after case where they fail to account for reversion to the mean in their behavior.
举个例子。事实证明,投资者赚到的美元加权回报低于共同基金的平均回报。例如,在截至 2011 年的过去 20 年里,标普 500 指数的年化回报约为 8%,共同基金的平均回报约为 6% 到 7%(差额来自费用和其他成本),但普通投资者赚到的回报不到 5%。乍一看,很难理解投资者怎么会比自己投的基金表现还差。关键在于,投资者往往在市场上涨后买入——忽视了均值回归——又在市场下跌后卖出——同样忽视了均值回归。这种高买低卖的做法,正是导致美元加权回报低于平均回报的原因。这一模式已被充分证实,以至于学术界称其为“傻瓜资金效应”。
Here’s an example. It turns out that investors earn dollar-weighted returns that are less than the average return of mutual funds. Over the last 20 years through 2011, for instance, the S&P 500 has returned about 8 percent annually, the average mutual fund about 6 to 7 percent (fees and other costs represent the difference), but the average investor has earned less than 5 percent. At first blush it seems hard to see how investors can do worse than the funds they invest in. The insight is that investors tend to buy after the market has gone up — ignoring reversion to the mean — and sell after the market has gone down — again, ignoring reversion to the mean. The practice of buying high and selling low is what drives the dollar-weighted returns to be less than the average returns. This pattern is so well documented that academics call it the "dumb money effect."
还要补充一点:只要各时期的业绩不是完全相关,就会出现均值回归。换句话说,但凡运气对结果有贡献,就一定会出现均值回归。这是一个我们的思维难以把握的统计学原理。
I should add that any time results from period to period aren’t perfectly correlated, you will have reversion to the mean. Saying it differently, any time luck contributes to outcomes, you will have reversion to the mean. This is a statistical point that our minds grapple with.
均值回归会制造一些迷惑我们的错觉。其中之一是因果错觉。关键在于,你不需要用因果关系来解释均值回归,当结果并非完全相关时,均值回归就会自然发生。一个著名的例子是父亲与儿子的身高。高个子父亲会有高个子儿子,但儿子的身高会比父亲更接近所有儿子的平均水平。同样,矮个子父亲会有矮个子儿子,但儿子的身高也会比父亲更接近平均水平。很少有人听到这一点会感到意外。
Reversion to the mean creates some illusions that trip us up. One is the illusion of causality. The trick is you don’t need causality to explain reversion to the mean, it simply happens when results are not perfectly correlated. A famous example is the stature of fathers and sons. Tall fathers have tall sons, but the sons have heights that are closer to the average of all sons than their fathers do. Likewise, short fathers have short sons, but again the sons have stature closer to average than that of their fathers. Few people are surprised when they hear this.
但既然均值回归仅仅反映的是不完全相关的结果,时间的方向就不重要了。所以高个子儿子有高个子父亲,但父亲的身高更接近所有父亲的平均身高。显而易见,儿子不可能导致父亲的身高,但均值回归的表述依然成立。
But since reversion to the mean simply reflects results that are not perfectly correlated, the arrow of time doesn’t matter. So tall sons have tall fathers, but the height of the fathers is closer to the average height of all fathers. It is abundantly clear that sons can’t cause fathers, but the statement of reversion to the mean is still true.
我想关键是,均值回归这件事本身并没有多么特别,但我们的头脑总是急于编造一个故事,试图找出其中的因果关系。
I guess the main point is that there is nothing so special about reversion to the mean, but our minds are quick to create a story that reflects some causality.
阿布斯曼:如果我们正确理解了均值回归,它甚至能帮助我们教育孩子吗?比如,如何应对孩子在学业上的表现。
Arbesman: If we understand reversion to the mean properly, can this even help with parenting, such as responding to our children’s performance in school?
莫布辛:没错,你又点出了另一种谬误,我称之为“反馈幻觉”。我们暂且承认,你女儿数学考试的成绩反映的是能力加运气。现在假设她拿回来一个优异的分数,说明能力不错,运气也非常好。你自然的反应会是什么?
Mauboussin: Exactly, you’ve hit on another one of the fallacies, which I call the illusion of feedback. Let’s accept that your daughter’s results on her math test reflect skill plus luck. Now say she comes home with an excellent grade, reflecting good skill and very good luck. What would be your natural reaction?
你可能会表扬她——毕竟她的成绩值得称赞。但下次考试会发生什么?平均来看,她的运气会回归中性,分数会下降。
You’d probably give her praise — after all, her outcome was commendable. But what is likely to happen on the next test? Well, on average her luck will be neutral and she will have a lower score.
你的大脑会自然而然地把你获得的正面反馈和负面结果联系在一起。你可能会对自己说,也许你的评论鼓励了她偷懒。但最简洁的解释不过是均值回归发挥了作用,而你的反馈并没有起多大作用。
Now your mind is going to naturally associate your positive feedback with a negative result. Perhaps your comments encouraged her to slack off, you’ll say to yourself. But the most parsimonious explanation is simply that reversion to the mean did its job and your feedback didn’t do much.
同样的情况也发生在负反馈上。
The same happens with negative feedback.
如果你的女儿因为运气不好考砸了,你可能会责备她,并通过限制她用电脑时间来惩罚她。但下一次考试,她很可能会考得更好——不管你有没有训斥和惩罚她。
Should your daughter come home with a poor grade reflecting bad luck, you might chide her and punish her by limiting her time on the computer. Her next test will likely produce a better grade, irrespective of your sermon and punishment.
需要记住的主要一点是,均值回归的发生纯粹是随机性的结果,而给随机结果强加因果关系是没有意义的。不过,我不想暗示均值回归只反映随机性,因为其他因素肯定也会起作用。
The main thing to remember is that reversion to the mean happens solely as the result of randomness, and that attaching causes to random outcomes does not make sense. Now I don’t want to suggest that reversion to the mean reflects randomness only, because other factors most certainly do come into play.
以体育领域的年龄老化与商业领域的竞争为例。但关键在于,仅凭随机性本身就能驱动这个过程。
Examples include aging in athletics and competition in business. But the point is that randomness alone can drive the process.
阿布斯曼:您在书中主要聚焦于商业、体育和投资领域,但技能与运气显然更广泛地存在于世界之中。在其他哪些领域,正确理解这两个要素同样至关重要(且常常被忽视)?
Arbesman: In your book you focus primarily on business, sports, and investing, but clearly skill and luck appear more widely in the world. In what other areas is a proper understanding of these two features important (and often lacking)?
莫布森:这在医学领域有极大的相关性。约翰·约安尼迪斯在 2005 年发表了一篇题为《为什么大多数已发表的研究成果是错误的》的论文,这篇论文引起了一些关注。他指出,基于随机试验并有适当对照组的医学研究,往往能有很高的可重复性。但他也表明,观察性研究的结果中有 80% 要么是错误的,要么是被夸大了的。观察性研究能制造一些不错的头条新闻,这对科学家的职业生涯或许有好处。
Mauboussin: One area where this has a great deal of relevance is medicine. John Ioannidiswrote a paper in 2005 called "Why Most Published Research Findings Are False" that raised a few eyebrows. He pointed out that medical studies based on randomized trials, where there’s a proper control, tend to be replicated at a high rate. But he also showed that 80 percent of the results from observational studies are either wrong or exaggerated. Observational studies create some good headlines, which can be useful to a scientist’s career.
问题是,人们听到这些观察性研究的结果就盲目听从。伊奥尼迪斯本人就是医生,他对这类研究的价值极为怀疑,甚至自己根本不理会它们。我在书中举了一个例子:一项研究表明,吃早餐麦片的女性更可能生男孩而非女孩。这种故事正是媒体趋之若鹜的。
The problem is that people hear about, and follow the advice of, these observational studies. Indeed, Ioannidis is so skeptical of the merit of observational studies that he, himself a physician, ignores them. One example I discuss in the book is a study that showed that women who eat breakfast cereal are more likely to give birth to a boy than a girl. This is the kind of story that the media laps up.
统计学家后来仔细梳理了数据,得出结论认为,这一结果很可能是偶然所致。
Statisticians later combed the data and concluded that the result is likely a product of chance.
现在,Ioannidis 的研究并未像我所定义的那样直接讨论技能与运气,但它触及了因果关系这一核心问题【编辑无耻自荐:关于科学中这方面的更多内容,请参阅《事实的半衰期》!】。凡是难以归因因果的地方,你就有可能误解正在发生的事情。
Now Ioannidis’s work doesn’t address skill and luck exactly as I’ve defined it, but it gets to the core issue of causality [Editor's shameless plug: for more about this in science, check out The Half-Life of Facts!]. Wherever it’s hard to attribute causality, you have the possibility of misunderstanding what’s going
那么,尽管我主要围绕商业、体育和投资展开叙述,但我希望这些想法也能轻易应用到其他领域。
on. So while I dwelled on business, sports, and investing, I’m hopeful that the ideas can be readily applied to other fields.
阿布斯曼:在理解技能与运气时,抽样(包括欠采样、有偏采样等)可能通过哪些方式让我们误入歧途?
Arbesman: What are some of the ways that sampling (including undersampling, biased sampling, and more) can lead us quite astray when understanding skill and luck?
莫布森:我们再来看一下欠采样和偏差采样这两种情况。
Mauboussin: Let’s take a look at undersampling as well as biased sampling.
在商业领域对失败的抽样不足是一个经典案例。华威商学院教授杰尔克·登雷尔(Jerker Denrell)在其论文《替代性学习、失败抽样不足与管理的神话》中提供了一个绝佳例子。想象一家公司。
Undersampling failure in business is a classic example. Jerker Denrell, a professor at Warwick Business School, provides a great example in a paper called "Vicarious Learning, Undersampling of Failure, and the Myths of Management." Imagine a company
公司可以二选一:高风险策略或低风险策略。选择前者要么大获成功,要么彻底失败;选择后者虽不如成功的高风险公司表现亮眼,但也不会倒闭。换句话说,高风险策略的结果方差极大,低风险策略的方差则小得多。
can select one of two strategies: high risk or low risk. Companies select one or the other and the results show that companies that select the high-risk strategy either succeed wildly or fail. Those that select the low-risk strategy don’t do as well as the successful high-risk companies but also don’t fail. In other words, the high-risk strategy has a large variance in outcomes and the low-risk strategy has smaller variance.
假设有一家新公司想要判断哪种策略最好。经过分析,高风险策略看起来会很棒,因为选择它并存活下来的公司取得了巨大成功,而选择它却倒闭的公司已经销声匿迹,因此不再列入样本。相比之下,由于所有选择低风险策略的公司都还在,它们的平均表现看起来反而更差。这就是经典的采样不足偏差案例。问题在于:选择每种策略的所有公司,最终结果究竟如何?
Say a new company comes along and wants to determine which strategy is best. On examination, the high-risk strategy would look great because the companies that chose it and survived had great success while those that chose it and failed are dead, and hence are no longer in the sample. In contrast, since all of the companies that selected the low-risk strategy are still be around, their average performance looks worse. This is the classic case of undersampling failure. The question is: What were the results of all of the companies that selected each strategy?
你可能觉得这简直显而易见,任何有头脑的公司或研究机构都不该犯这种错。但这个毛病恰恰困扰着大量商业研究。传统上帮助企业的方法是这样的:先找出那些成功的公司,归纳它们共有的特征,然后建议其他公司也追求这些特征,以便获得成功。这就是很多畅销书的套路,包括吉姆·柯林斯的《从优秀到卓越》。举个例子,柯林斯发现成功公司的特征之一是做“刺猬”——专注自己的业务。关键问题不是:是不是所有成功公司都是刺猬?关键问题是:是不是所有刺猬都成功了?
Now you might think that this is super obvious, and that thoughtful companies or researchers wouldn’t do this. But this problem plagues a lot of business research. Here’s the classic approach to helping businesses: Find companies that have succeeded, determine which attributes they share, and recommend other companies seek those attributes in order to succeed. This is the formula for many bestselling books, including Jim Collins’s Good to Great. One of the attributes of successful companies that Collins found, for instance, is that they are “hedgehogs,” focused on their business. The question is not: Were all successful companies hedgehogs? The question is: Were all hedgehogs successful?
第二个问题毫无疑问会得出与第一个问题不同的答案。
The second question undoubtedly yields a different answer than the first.
另一个常见的错误是基于小样本来得出结论,这一点我已经提过。我从霍华德·韦纳那里学到一个例子,跟学校规模有关。研究人员研究中小学教育时,想弄明白如何提高学生的考试成绩。于是他们做了看似非常合乎逻辑的事——看看哪些学校成绩最高。结果他们发现,成绩最高的学校规模都小,这从直觉上也说得通,因为班级规模小等等。
Another common mistake is drawing conclusions based on samples that are small, which I’ve already mentioned. One example, which I learned from Howard Wainer, relates to school size. Researchers studying primary and secondary education were interested in figuring out how to raise test scores for students. So they did something seemingly very logical – they looked at which schools have the highest test scores. They found that the schools with the highest scores were small, which makes some intuitive sense because of smaller class sizes, etc.
但这落入了样本陷阱。接下来的问题是:哪些学校考试分数最低?答案同样是:小学校。从统计学角度看,这完全在意料之中,因为小样本的方差更大。所以,小学校既有最高分也有最低分,而大学校的分数则更接近平均值。
But this falls into a sampling trap. The next question to ask is: which schools have the lowest test scores? The answer: small schools. This is exactly what you would expect from a statistical viewpoint since small samples have large variances. So small schools have the highest and lowest test scores, and large schools have scores closer to the average.
由于研究人员只关注高分,他们忽略了问题的关键。
Since the researchers only looked a high scores, they missed the point.
这不仅仅是统计课的案例。教育改革者随后投入了数十亿美元来缩减学校规模。例如,西雅图的一所大学校被拆分成五所小学校。结果发现,缩小学校规模实际上可能成为问题,因为它导致专业化程度降低——比如,大学先修课程变少了。韦纳将样本量与方差之间的关系称为“最危险的公式”,因为多年来它绊倒了太多研究人员和决策者。
This is more than a case for a statistics class. Education reformers proceeded to spend billions of dollars reducing the sizes of schools. One large school in Seattle, for example, was broken into five smaller schools. It turns out that shrinking schools can actually be a problem because it leads to less specialization—for example, fewer advanced placement courses. Wainer calls the relationship between sample size and variance the "most dangerous equation" because it has tripped up some many researchers and decision makers over the years.
阿贝斯曼:你关于技能悖论的讨论——群体越有技能,运气就越发挥作用——让我有点想到了红皇后效应,在进化过程中,生物体不断与其他高度适应的生物体竞争。你觉得这两者之间有关系吗?
Arbesman: Your discussion of the paradox of skill—that more skillful the population, the more luck plays a role—reminded me a bit of the Red Queen effect, where in evolution, organisms are constantly competing against other highly adapted organisms. Do you think there is any relationship?
莫布辛:确实如此。我认为关键区别在于绝对表现与相对表现。在接二连三的领域里,我们都看到了绝对表现的提升。例如,在那些靠秒表衡量成绩的运动中——包括游泳、跑步和赛艇——今天的运动员比过去快得多,并且会继续进步,直到触及人类生理极限。类似的进程也发生在商业领域,产品的品质和可靠性随着时间推移稳步提高。
Mauboussin: Absolutely. I think the critical distinction is between absolute and relative performance. In field after field, we have seen absolute performance improve. For example, in sports that measure performance using a clock—including swimming, running, and crew—athletes today are much faster than they were in the past and will continue to improve up to the point of human physiological limits. A similar process is happening in business, where the quality and reliability of products has increased steadily over time.
但存在竞争的地方,我们在意的就不是绝对表现,而是相对表现。这一点可以
But where there’s competition, it’s not absolute performance we care about but relative performance. This point can be
令人困惑。例如,分析显示棒球运动中存在大量随机性,但这似乎与“击中每小时 95 英里的快球是任何运动中最困难的事情之一”这一事实不符。当然,击打快球需要高超的技巧,投出快球也同样需要。关键在于,随着投手和击球手的进步,他们的进步大致是同步的,彼此相互抵消。绝对的进步被相对的均衡所掩盖。
confusing. For example, the analysis shows that baseball has a lot of randomness, which doesn’t seem to square with the fact that hitting a 95-mile-an-hour fastball is one of the hardest things to do in any sport. Naturally, there is tremendous skill in hitting a fastball, just as there is tremendous skill in throwing a fastball. The key is that as pitchers and hitters improve, they improve in rough lockstep, offsetting one another. The absolute improvement is obscured by the relative parity.
这一点引出了我认为最违反直觉的一个结论:随着技能的提高,它在人群中的分布往往会变得更加均匀。假设运气的贡献保持稳定,就会出现技能提升反而导致运气对结果的贡献更大的情况。这就是技能悖论。因此,它与红皇后效应密切相关。
This leads to one of the points that I think is most counter to intuition. As skill increases, it tends to become more uniform across the population. Provided that the contribution of luck remains stable, you get a case where increases in skill lead to luck being a bigger contributor to outcomes. That’s the paradox of skill. So it’s closely related to the Red Queen effect.
阿伯斯曼:对于理解技能和运气之间的关系,你认为哪个单一概念或想法最重要?
Arbesman: What single concept or idea do you feel is most important for understanding the relationship between skill and luck?
莫布辛:最重要的单一概念是确定某项活动在“全凭运气、毫无技能”与“全凭技能、毫无运气”这一连续谱上的位置。给活动定位,是把握预测接下来会发生什么的最佳方法。
Mauboussin: The single most important concept is determining where the activity sits on the continuum of all-luck, no-skill at one end to no-luck, all-skill at the other. Placing an activity is the best way to get a handle on predicting what will happen next.
让我从另一个角度来谈这个问题。当被问及哪篇论文是他有史以来最喜欢的一篇时,丹尼尔·卡尼曼提到了 1973 年他与阿莫斯·特沃斯基合著的《论预测心理学》。特沃斯基和卡尼曼基本上认为,要进行有效预测,需要考虑三件事:基础概率、个案情况,以及 *如何对两者进行加权。* 用运气与技能的术语来说,如果运气占主导,你就应该主要依据基础概率;如果技能占主导,你就应该主要依据个案情况。介于两者之间的活动,其权重则是混合的。
Let me share another angle on this. When asked which was his favorite paper of all-time, Daniel Kahneman pointed to "On the Psychology of Prediction," which he coauthored with Amos Tversky in 1973. Tversky and Kahneman basically said that there are three things to consider in order to make an effective prediction: the base rate, the individual case, and *how to weight the two. *In luck-skill language, if luck is dominant you should place most weight on the base rate, and if skill is dominant then you should place most weight on the individual case. And the activities in between get weightings that are a blend.
事实上,有一个概念叫做“收缩因子”,它告诉你在进行良好预测时,应该将过去的结果向均值回归多少。收缩因子为 1 意味着下一个结果将与上一个结果相同,表示完全由技能决定;因子为 0 则意味着对下一个结果的最佳猜测就是平均值。生活中几乎所有有趣的事情都处于这两个极端之间。
In fact, there is a concept called the "shrinkage factor" that tells you how much you should revert past outcomes to the mean in order to make a good prediction. A shrinkage factor of 1 means that the next outcome will be the same as the last outcome and indicates all skill, and a factor of 0 means the best guess for the next outcome is the average. Almost everything interesting in life is in between these extremes.
为了让这一点更具体,我们来看看棒球中的两个统计数据:击球率和上垒率。运气在决定击球率方面所起的作用大于决定上垒率时。因此,如果你想预测一名球员的表现(暂时假设技能不变),那么击球率所需采用的收缩因子要比上垒率更接近 0。
To make this more concrete, consider batting average and on-base percentage, two statistics from baseball. Luck plays a larger role in determining batting average than it does in determining on-base percentage. So if you want to predict a player’s performance (holding skill constant for a moment), you need a shrinkage factor closer to 0 for batting average than for on-base percentage.
我想再补充一点,这不是分析层面的,而是心理层面的。你的大脑左半球有一部分专门负责理清因果关系。它接收信息并创造出连贯的叙事。它非常擅长这项功能,以至于神经科学家称之为“解释器”。
I’d like to add one more point that is not analytical but rather psychological. There is a part of the left hemisphere of your brain that is dedicated to sorting out causality. It takes in information and creates a cohesive narrative. It is so good at this function that neuroscientists call it the “interpreter.”
对于未来结果是技能和运气的结合这一观点,没有人会提出异议。
Now no one has a problem with the suggestion that future outcomes combine skill and luck.
但是,一旦某件事发生了,我们的大脑就会迅速而自然地创造一个叙事来解释这个结果。由于“解释器”的职责是寻找因果关系,它并不擅长识别运气。一旦事情发生,我们的大脑就开始相信这是不可避免的。这导致了心理学家所称的“渐进决定论”——一种我们早已知道将会发生什么的感觉。
But once something has occurred, our minds quickly and naturally create a narrative to explain the outcome. Since the interpreter is about finding causality, it doesn’t do a good job of recognizing luck. Once something has occurred, our minds start to believe it was inevitable. This leads to what psychologists call “creeping determinism” – the sense that we knew all along what was going to happen.
因此,尽管最重要的单一概念是了解你所处的活动在运气-技能连续谱上的位置,但相关的一点是,你的大脑并不会很好地识别运气的本来面目。
So while the single most important concept is knowing where you are on the luck-skill continuum, a related point is that your mind will not do a good job of recognizing luck for what it is.