智商对决策商:区分聪明与决策技能
GLOBAL FINANCIAL STRATEGIES www.credit-suisse.com
GLOBAL FINANCIAL STRATEGIES www.credit-suisse.com
智商与理性商:区分聪明与决策能力 2015 年 5 月 12 日
IQ versus RQ Differentiating Smarts from Decision-Making Skills May 12, 2015
Authors
Authors
迈克尔·J·莫布森 [email protected]
Michael J. Mauboussin [email protected]
智商 理性商 丹·卡拉汉,特许金融分析师 [email protected]
IQ RQ Dan Callahan, CFA [email protected]
“保持理性是一种道德责任。”
“Being rational is a moral imperative.”
Charlie Munger1
Charlie Munger1
智商(IQ)与理性商(RQ)是两个不同的概念。把智商想象成发动机的马力,理性商则是输出功率。
Intelligence quotient (IQ) and rationality quotient (RQ) are distinct. Think of IQ as the horsepower of an engine and RQ as the output.
我们分享一项经典校准测试的结果,这是理性的一个重要方面。校准良好的人知道自己知道什么,也知道自己不知道什么。
We share the results of a classic test of calibration, which is an important facet of rationality. Well calibrated people know what they know and know what they don’t know.
与以往的研究一致,我们发现参与者高估了自己的准确度,因为他们的主观概率估计往往高于实际正确率。
Consistent with past research, we find that participants overestimate their accuracy as their subjective probability estimates tend to be higher than the actual percent correct.
投资者和高管可以通过记录分数、询问他人意见、使用基础率以及更新概率来提升理性水平。
Investors and executives can improve their rationality by keeping score, asking about others, using base rates, and updating probabilities.
一项大规模预测项目表明,最好的预测者善于运用归纳推理和数值推理,具有认知控制能力和成长型思维,思想开放,且能有效地在团队中协作。
A large-scale forecasting project has shown that the best forecasters use inductive and numerical reasoning, have cognitive control and a growth mindset, and are open-minded and effective working as part of a team.
Introduction
Introduction
多伦多大学应用心理学教授基思·斯坦诺维奇将智商(IQ)与理性商(RQ)区分开来。心理学家通过特定测试(包括韦克斯勒成人智力量表)来衡量智商,它与 SAT 等标准化测试高度相关。
Keith Stanovich, a professor of applied psychology at the University of Toronto, distinguishes between intelligence quotient (IQ) and rationality quotient (RQ).2 Psychologists measure IQ through specific tests, including the Wechsler Adult Intelligence Scale, and it correlates highly with standardized tests such as the SAT.3
智商衡量的是真实存在的东西,并且与某些结果相关联。例如,在 SAT 数学部分得分位居顶尖 1% 的前 10%(即 99.9 百分位)的 13 岁儿童,获得数学或科学博士学位的可能性是那些得分在顶尖 1% 的后 10%(即 99.1 百分位)的儿童的 18 倍。
IQ measures something real, and it is associated with certain outcomes. For example, thirteen-year-old children who scored in the top decile of the top percent (99.9th percentile) on the math section of the SAT were eighteen times more likely to earn a doctorate degree in math or science than children who scored in the bottom decile of the top percent (99.1st percentile).4
理性商是指理性思考并因此做出正确决策的能力。尽管我们通常认为智力和理性是相辅相成的,但斯坦诺维奇的研究表明,智商与理性商之间的相关系数相对较低,仅在 0.20 到 0.35 之间。智商测试并非旨在捕捉导致明智决策的思维过程。
RQ is the ability to think rationally and, as a consequence, to make good decisions. Whereas we generally think of intelligence and rationality as going together, Stanovich’s work shows that the correlation coefficient between IQ and RQ is relatively low at .20 to .35.5 IQ tests are not designed to capture the thinking that leads to judicious decisions.
斯坦诺维奇感到遗憾,因为几乎所有社会都专注于智力,而非理性行为的代价却如此高昂。但如果你足够警觉,就能识别出理性思维的标志。据斯坦诺维奇说,这些标志包括适应性行为、高效的行为调节、合理的目标优先排序、反思性以及对证据的正确处理。
Stanovich laments that almost all societies are focused on intelligence when the costs of irrational behavior are so high. But you can pick out the signatures of rational thinking if you are alert to them. According to Stanovich, they include adaptive behavioral acts, efficient behavioral regulation, sensible goal prioritization, reflectivity, and the proper treatment of evidence.6
你的 SAT 分数对这些品质几乎没有任何揭示作用。因此,评估你自己或他人决策的第一个教训是,将智商和理性商分开考虑。伯克希尔·哈撒韦公司董事长兼首席执行官沃伦·巴菲特将智商等同于发动机的马力,将理性商等同于输出功率。我们都认识一些智商很高但理性商平平或很低的人。他们的效率很差。还有一些人智商并不出众,却能持续做出明智的决策。他们的效率非常高。
Your SAT scores shed little light on any of these qualities. So the first lesson in assessing your own decisions or those of others is to consider IQ and RQ separately. Warren Buffett, chairman and chief executive officer (CEO) of Berkshire Hathaway, equates IQ to the horsepower of an engine and RQ to the output. We all know people who are high on IQ but average or low on RQ. Their efficiency is poor. There are others without dazzling IQs but who consistently make sound decisions. They are highly efficient.
沃伦·巴菲特拥有充沛的马力和强大的输出功率。但当被问及成功的秘诀时,巴菲特强调,带来巨大差异的是理性商,而不是智商:
Warren Buffett has plenty of horsepower and output. But when asked about his success, Buffett emphasized that it was RQ that made the big difference, not IQ:7
对我来说,走到今天这一步其实很简单。不是智商,我相信你们听到会很高兴。关键因素是理性。我总是把智商和天赋看作发动机的马力,而输出功率——发动机运转的效率——取决于理性。很多人一开始拥有 400 马力的发动机,却只能输出 100 马力的功率。拥有一个 200 马力的发动机并让它全力输出要好得多。
How I got here is pretty simple in my case. It's not IQ, I’m sure you'll be glad to hear. The big thing is rationality. I always look at IQ and talent as representing the horsepower of the motor, but that the output—the efficiency with which that motor works—depends on rationality. A lot of people start out with 400-horsepower motors but only get a hundred horsepower of output. It’s way better to have a 200-horsepower motor and get it all into output.
斯坦诺维奇的心理学研究支持巴菲特的观察。虽然目前还没有一套全面的测试来衡量理性商——斯坦诺维奇正在研究——但我们将考察校准,这是理性的一个重要方面。作为这项研究的一部分,我们测量了数千人的校准水平。有趣的地方在于,你也可以参与这个练习,看看自己与他人相比表现如何。
Stanovich’s psychological research supports Buffett’s observation. While there is not yet a comprehensive test to measure RQ—Stanovich is working on it—we will look at calibration, one of the important facets of rationality.8 As part of this research, we measured the calibration of thousands of people. And part of the fun is that you, too, can participate in the exercise and see how you stack up versus others.
Measuring Rationality
Measuring Rationality
认知科学家和哲学家谈论“工具理性”和“认知理性”。工具理性是指在约束条件下,以能够最大化实现自身目标的方式行事。期望效用理论基于一系列公理,为如何做到这一点提供了一个规范性框架。如果你遵循这些公理,你的行为就是工具理性的。
Cognitive scientists and philosophers talk about “instrumental” and “epistemic” rationality. Instrumental rationality is behaving in such a way that you get what you want the most, subject to constraints. Expected utility theory, which is based on a series of axioms, provides a normative framework for how to do this. You’ll be instrumentally rational if you follow the axioms.9
认知理性描述了一个人的信念在多大程度上符合现实世界。例如,如果你相信牙仙的存在,就表明你缺乏认知理性。这里有个更容易记住两个术语的方法:工具理性是“该做什么”,认知理性是“什么是真的”。
Epistemic rationality describes how well a person’s beliefs map onto the world. If you believe in the tooth fairy, for instance, you are showing a lack of epistemic rationality. Here’s a catchier way to remember the two terms: instrumental rationality is “what to do” and epistemic rationality is “what is true.”10
我们将聚焦于对认知理性的观察。评估这种理性形式的一种方法是通过校准测试。想象一位天气预报员。如果在她预测降雨概率为 70% 的日子里,实际降雨概率正好是 70%,那她的校准就很好。反之,如果那些日子只有 30% 的时间下雨,那她的校准就很差。
We will focus on observations about epistemic rationality. One way to assess this form of rationality is through a test of calibration. Think of a weather forecaster. If it actually rains 70 percent of the time on the days she predicts a 70 percent chance of rain, she is well calibrated. She is poorly calibrated, on the other hand, if it only rains on 30 percent of those days.
图 1 展示了记录分数的一种方法。横轴测度个人的主观预测(“明天有 70% 的概率下雨”),纵轴记录实际结果(“下雨了”)。如果结果落在一条 45 度角的直线附近,那他的校准就很好。
Exhibit 1 shows one approach to keeping score. The horizontal axis measures an individual’s subjective forecast (“there’s a 70 percent chance of rain tomorrow”) and the vertical axis captures the actual outcome (“it rained”). You know that someone is well calibrated if their results fall close to the line at a 45-degree angle.
图 1:校准的度量
Exhibit 1: Measure of Calibration
100% 90% 80% 70%
100% 90% 80% 70%
Objective Probability 60%
Objective Probability 60%
50%
50%
40% 完美校准线
40% Line of perfect calibration
原件此处是表格,PDF 抽取时列结构已丢失,下面只剩按列读出的数字,行列对应关系无法还原。核对数据请打开来源正文。
30% 20% 10% 0% 0% 10% 20% 30% 40% 50% 60% 70% 80% 90%100% Subjective Probability
30% 20% 10% 0% 0% 10% 20% 30% 40% 50% 60% 70% 80% 90%100% Subjective Probability
资料来源:瑞信。
Source: Credit Suisse.
另一个重要考量是信念强度。信念强度衡量的是一个人在赋予极端概率时的表现。校准和信念强度相关但不同。例如,伦敦一年中大约有一半的日子会下雨。所以如果你每天早上醒来,抛一枚均匀的硬币,并根据结果标记当天的晴雨,那么一年下来,你的校准看起来会很好。
Conviction is another important consideration. Conviction measures how well people do when they assign extreme probabilities.11 Calibration and conviction are related but distinct. For example, it rains about half of the days in an average year in London. So if you wake up every morning, flip a fair coin, and mark your outcome, you will appear well calibrated over one year.
但这并不能帮你安排野餐。你需要的是与当天实际天气相符的一系列晴天或雨天的预测。这样的预测需要比抛硬币更高的信念强度。图 2 左侧展示了完美校准但信念强度差的情况,右侧展示了完美校准和完美信念强度的情况。
But that doesn’t help you plan picnics. What you want are a series of predictions for sun or rain that correspond with the actual weather that day. Those predictions require higher conviction than what the toss of a coin can offer. Exhibit 2 shows perfect calibration but poor conviction on the left, and perfect calibration and conviction on the right.
图 2:校准与信念强度 最佳 可能 校准, 最佳 可能 区分度, 信念强度差 最佳 可能 校准, 最佳 可能 区分度, 信念强度佳 1 1
Exhibit 2: Calibration and Conviction Best-Possible Best PossibleCalibration, Calibration, Best-Possible Calibration, Best Possible Calibration, Poor Discrimination Poor Conviction Best-Possible Discrimination Best Possible Conviction 1 1
0.8 0.8
0.8 0.8
客观概率 客观概率
Objective Probability Objective Probability
原件此处是表格,PDF 抽取时列结构已丢失,下面只剩按列读出的数字,行列对应关系无法还原。核对数据请打开来源正文。
0.6 0.6 0.4 0.4 0.2 0.2 0 0 0 0.2 0.4 0.6 0.8 1 0 0.2 0.4 0.6 0.8 1
0.6 0.6 0.4 0.4 0.2 0.2 0 0 0 0.2 0.4 0.6 0.8 1 0 0.2 0.4 0.6 0.8 1
主观概率 主观概率 资料来源:菲利普·E·泰特洛克,《专家政治判断:它有多好?我们如何知晓?》(普林斯顿,新泽西:普林斯顿大学出版社,2005 年),第 48 页。
Subjective Probability Subjective Probability Source: Philip E. Tetlock, Expert Political Judgment: How Good Is It? How Can We Know? (Princeton, NJ: Princeton University Press, 2005), 48.
这里我们展示一项经典校准测试的结果。参与者访问网站 http://confidence.success-equation.com,看到 50 道是非题。图 3 是该网站的截图。
Here we present the results of a classic calibration test. Subjects who participated went to the website, http://confidence.success-equation.com, and saw 50 true-false questions. Exhibit 3 is a screenshot of the site.
参与者接着回答“正确”或“错误”,并要求登记一个正确概率,从 50% 到 100%,以 10 个百分点为增量。如果你完全不知道答案是“正确”还是“错误”,你应该随机选择一个答案,并在正确概率中输入“50%”。如果你对提供的答案非常肯定,则点击“100%”。
The subjects then answered either true or false and were asked to register a probability of correctness, from 50 to 100 percent, in increments of 10 percentage points. If you have no idea whether the answer is true or false you should select an answer at random and enter “50%” as your probability of correctness. If you are certain of the answer you provide, you click “100%.”
最后,参与者提交答案,并获得结果,包括:
At the end, the subjects submit their answers and receive their results, which include:
所有问题的平均置信度与正确百分比
Mean, or average, confidence and percent correct for all questions
正确与错误答案的平均置信度
Mean confidence for correct and incorrect answers
低置信度(50-60%)、中等置信度(70-80%)和高置信度(90-100%)提交中的正确与错误数量
Number correct and answered for low confidence (50-60 percent, medium confidence (70-80 percent) and high confidence (90-100 percent) submissions
一个类似于图 1 的校准图
A calibration graph similar to exhibit 1
我们访问了 1985 名参与者的结果,所有参与者均为匿名。
We accessed the results of 1,985 participants, all of whom were anonymous.12
图 3:校准问题与正确概率
Exhibit 3: Calibration Questions and Probability of Correctness
资料来源:http://confidence.success-equation.com。
Source: http://confidence.success-equation.com.
图 4 显示了结果。这个模式与研究人员几十年来的发现一致:主观概率估计平均而言显著高于实际正确百分比。对于整个人群,平均主观概率为 70%,而实际正确率略低于 60%。
Exhibit 4 shows the results. The pattern is consistent with what researchers have found for decades: Subjective probability estimates are substantially higher, on average, than the actual percent correct. For the whole population, the average subjective probability was 70 percent and the actual percent correct was just under 60 percent.
图 4:1985 名参与者的校准情况
Exhibit 4: Calibration for 1,985 Participants
原件此处是表格,PDF 抽取时列结构已丢失,下面只剩按列读出的数字,行列对应关系无法还原。核对数据请打开来源正文。
100% 90% 80% 70% Correct 60% 50% 40% 30% 30% 40% 50% 60% 70% 80% 90% 100% Confidence
100% 90% 80% 70% Correct 60% 50% 40% 30% 30% 40% 50% 60% 70% 80% 90% 100% Confidence
资料来源:http://confidence.success-equation.com。
Source: http://confidence.success-equation.com.
请注意,良好的校准并不要求每次都正确(右上角的点表示要么有人作弊,要么是神祇参与了测试),而是要求结果接近 45 度线。
Note that proper calibration does not require being right all of the time (the dot in the upper right-hand corner indicates that either someone cheated or a deity took the test) but rather being close to the 45 degree line.
这意味着知道自己知道什么,也知道自己不知道什么。
It’s knowing what you know and knowing what you don’t know.
图 5 显示了回答的分布,横轴代表置信度减去正确百分比,纵轴代表频率。正态分布,或钟形曲线,很好地描述了这些数据,其均值和标准差约为 10%。
Exhibit 5 shows the distribution of responses, with the horizontal axis representing the confidence level minus the percent correct and the vertical axis the frequency. A normal distribution, or bell curve, describes the data well, with a mean and standard deviation of about 10 percent.
这个分布使我们能够将参与者分成三组:过度自信者、自信不足者和校准良好者。我们将校准良好定义为主观概率与实际正确百分比的偏差在 1 个百分点以内。
This distribution allows us to segregate the participants into three groups: those who are overconfident, underconfident, and well calibrated. We define well calibrated as a subjective probability within 1 percentage point of the percent correct.
根据这个标准,82.7% 的参与者过度自信(主观置信度超过实际正确百分比),11.7% 的参与者自信不足(主观置信度低于实际正确百分比),只有 5.6% 的参与者校准良好。
Based on that criterion, 82.7 percent of the participants were overconfident (subjective confidence exceeded actual percent correct), 11.7 percent were underconfident (subjective confidence less than actual percent correct), and only 5.6 percent were well calibrated.
图 5:大多数参与者过度自信 25
Exhibit 5: Most Participants Are Overconfident 25
20
20
Frequency (Percent)
Frequency (Percent)
15
15
校准良好 10 过度- 自信- 不足
Well-Calibrated 10 Over-Under- Confident
原件此处是表格,PDF 抽取时列结构已丢失,下面只剩按列读出的数字,行列对应关系无法还原。核对数据请打开来源正文。
5 Confident 0 <-20 (5)-0 5-10 10-15 15-20 20-25 25-30 30-35 35-40 >40 (10)-(5) 0-5 (20)-(15) (15)-(10)
5 Confident 0 <-20 (5)-0 5-10 10-15 15-20 20-25 25-30 30-35 35-40 >40 (10)-(5) 0-5 (20)-(15) (15)-(10)
置信度减去正确百分比(百分比)
Confidence Level Minus Percent Correct (Percent)
资料来源:http://confidence.success-equation.com。
Source: http://confidence.success-equation.com.
著名心理学家阿莫斯·特沃斯基据说曾说过,人类只能区分三种概率水平:“会发生”、“不会发生”和“也许”。图 6 显示了正确率主观概率的分布。
Amos Tversky, the renowned psychologist, is reported to have said that humans can only distinguish between three levels of probability: “it’s gonna happen,” “it’s not gonna happen,” and “maybe.”13 Exhibit 6 shows the distribution of subjective probabilities of correctness.
42% 的回答是 50%,这相当于说“我完全不知道答案是什么”。这对应于特沃斯基的“也许”。请注意,这里可能还有一个额外效应,因为 50% 是网站的默认设置。所以如果参与者没有更改分配的主观概率,它会自动显示 50%。
Forty-two percent of the responses were 50 percent, which is the equivalent of saying, “I have no idea what the answer is.” This corresponds to Tversky’s “maybe.” Note that there is likely an additional effect here because 50 percent is the default setting on the site. So if the participant doesn’t change the assigned subjective probability, it registers 50 percent automatically.
第二受欢迎的回答,占总数的近四分之一,是 100%。这就是特沃斯基的“会发生”或“不会发生”。因此,近三分之二的回答要么是“我不知道”,要么是“我知道”,剩下的部分则分布在介于这两个极端之间的四个选项中。
The next most popular response, nearly one-quarter of the total, was 100 percent. This is Tversky’s “it’s gonna happen” or “it’s not gonna happen.” So nearly two-thirds of the responses were either “I don’t know” or “I do know” with the balance split between the four choices in between those extremes.
图 6:正确率主观概率的分布
Exhibit 6: Distribution of Subjective Probabilities of Correctness
原件此处是表格,PDF 抽取时列结构已丢失,下面只剩按列读出的数字,行列对应关系无法还原。核对数据请打开来源正文。
45% 41.9% 40% 35% 30% 25% 23.2% 20% 15% 9.4% 9.5% 9.1% 10% 7.0% 5% 0% 50% 60% 70% 80% 90% 100%
45% 41.9% 40% 35% 30% 25% 23.2% 20% 15% 9.4% 9.5% 9.1% 10% 7.0% 5% 0% 50% 60% 70% 80% 90% 100%
资料来源:http://confidence.success-equation.com。
Source: http://confidence.success-equation.com.
这就引出了一个合乎逻辑的后续问题:对于每个分配的正确概率,结果有多好?图 7 显示了答案。当参与者选择 50% 时,他们回答正确的概率是随机的。这意味着他们的校准良好。他们不知道,知道自己不知道,并且像不知道那样回答问题。
This leads to a logical follow up question: How good were the results for each assigned probability of correctness? Exhibit 7 shows the answer. When the subjects selected 50 percent, their probability of being correct was random. This means they were well calibrated. They didn’t know, knew they didn’t know, and answered as if they didn’t know.
然而,随着分配的正确概率上升,参与者的校准程度下降。例如,当参与者选择 100% 时,他们只有 77% 的时间回答正确。在 90% 时,他们只有 65% 的时间回答正确。在分配的正确概率较高时,高估自身能力的情况最为严重。
However, as the assigned probability of correctness rose, the subjects became less calibrated. For instance, when the subjects selected 100 percent, they were only correct 77 percent of the time. At 90 percent, they were only correct 65 percent of the time. Overestimation of ability was greatest at the high levels of assigned probability of correctness.14
图 7:每个分配的正确概率的校准情况
Exhibit 7: Calibration for Each Assigned Probability of Correctness
原件此处是表格,PDF 抽取时列结构已丢失,下面只剩按列读出的数字,行列对应关系无法还原。核对数据请打开来源正文。
100% 90% 80% Correct 77% 70% 60% 65% 60% 56% 57% 50% 51% 40% 40% 50% 60% 70% 80% 90% 100% Confidence
100% 90% 80% Correct 77% 70% 60% 65% 60% 56% 57% 50% 51% 40% 40% 50% 60% 70% 80% 90% 100% Confidence
资料来源:http://confidence.success-equation.com。
Source: http://confidence.success-equation.com.
过度自信可能带来问题,原因有两个。第一个显而易见:如果你对某个结果高度自信,却在相当高比例的情况下判断错误,你就不会考虑其他可能性,最终会做出糟糕的决策。15 近一个例子是罗恩·约翰逊担任 J.C. 彭尼首席执行官那段时期。
Overconfidence can be a problem for a couple of reasons. The first obvious one is if you are highly confident of an outcome and are wrong a relatively high percentage of the time, you will fail to consider alternatives and ultimately make poor decisions.15 One recent example is Ron Johnson’s tenure as CEO of J.C. Penney.
约翰逊凭借直觉行事,迅速调整了这家零售商的定位,结果效果很差。他被赶下台后,这家零售商又恢复了许多过去的做法。16
Johnson, relying on his intuition, quickly repositioned the retailer to poor effect. Following his ouster, the retailer returned to many of its past practices.16
另一个问题是,那些高估自己认知水平的人,比那些清楚自身局限的人,更缺乏学习和改进的动力。17 事实上,一项研究表明,能力最差的人,他们自认为能做到的事与实际取得的成就之间,差距最大。18
Another problem is that people who think that they know more than they do are less motivated to learn and improve than those who understand their limitations.17 Indeed, one study showed that the least capable people have the largest gap between what they think they can do and what they actually achieve.18
投资者的教训与管理者的箴言
Lessons for Investors and Executives
所有这些的好消息是,我们可以训练自己变得更理性。以下是一些建议:
The good news in all of this is that we can train ourselves to be more rational. Here are some ideas:
保持记分。尽可能提出那些在已知时间内会有确定答案的问题,你就有了记分的依据。传统做法是通过布里尔评分(Brier score)来实现,我们会在附录中详细讨论。布里尔评分最初是为了帮助气象学家就天气预测获得反馈而设计的。通过改进天气建模技术并获得更精确的反馈,如今的气象学家比一二十年前要准确得多。19
Keep score. To the degree to which you can pose questions that will have a definite answer within a known period of time, you have a basis for keeping score. The classic way to do this is through a Brier score, which we discuss in detail in the appendix. Brier scores were originally developed to help give meteorologists feedback on their predictions for the weather. Through improvements in weather modeling techniques and sharper feedback, meteorologists today are vastly more accurate than they were a generation or two ago.19
问问别人。普林斯顿大学心理学教授艾米莉·普罗宁发现,人们能识别出别人思维中的偏见,却不知为何认为自己受到同样偏见的影响要小得多。 20 举个例子,医生们知道药企的礼物会对其他医生产生偏见影响,但却相信自己对此有免疫力。
Ask about others. Emily Pronin, a professor of psychology at Princeton University, has found that while people recognize biases in the thinking of others, they somehow think that they suffer less from the same biases.20 For example, physicians know that gifts from pharmaceutical companies have biasing effects for other doctors but believe they are immune from the effect.
这里有个应对方法。假如你是一位投资者,正在与一家公司的管理团队面谈,考虑是否买入这家公司的股票。你得明白,这个管理团队会有乐观偏见,所以对他们说的话要持保留态度。但他们对于其他公司的看法,则更可能比较准确。换句话说,别问别人他们自己怎么样,问他们关于别人的事。
Here’s a technique to deal with this. Say you are an investor interviewing a company’s management team, and you are considering buying the stock. You should know that the management team will have an optimism bias, and so you have to take what they say with a grain of salt. But their views about other companies are more likely to be accurate. In other words, don’t ask people about themselves, ask them about others.
运用基础概率。也许最有效的去偏差工具就是使用基础概率。²¹ 尽管我们都喜欢认为自己独一无二,但问问当别人处于相同情况时发生了什么,往往大有帮助。美国前哈佛大学校长、财政部长劳伦斯·萨默斯对他的研究助理有一条规则。他会问一个项目需要多长时间。然后他会把助理的答案乘以 2,并把时间单位往上升一级。²² 于是,“两小时”就被变成了“四天”。这或许只是为了避免失望的一种机制,但新估算的方向无疑是对的。
Use base rates. Perhaps the single most effective de-biasing tool is the use of base rates.21 While we all like to think of ourselves as unique, asking what happened when others were in the same situation can be very helpful. Larry Summers, the former president of Harvard University and Secretary of the Treasury of the U.S., had a rule he used with his research assistants. He would ask how long a project would take. And then he would take the assistant’s answer, double the estimate, and move up to the next unit of time.22 So “two hours” would be translated as “four days.” Perhaps this was simply a mechanism to avoid disappointment, but the direction of the new estimate was no doubt correct.
更新概率。我们向参与者提供的测试是静态的。而在真实世界中,概率无时无刻不在变化。理性思考的关键挑战之一,就是在获得新信息时准确地更新概率。事实证明,最顶尖的预测者在这方面做得非常出色,并且会使用非常精细的概率递增量。²³ 我们大多数人都陷入了确认偏误的陷阱,宁愿忽略或打折新信息,也不愿将其恰当地纳入我们的评估。
Update probabilities. The test we shared with the participants was static. In the real world, probabilities shift all of the time. One of the key challenges in rational thinking is to accurately update probabilities as new information arrives. It turns out that the very best forecasters do this very well and use very granular increments of probability.23 Most of us fall into the trap of confirmation bias, preferring to disregard or discount new information than to properly incorporate it into our assessment.
本项目的局限性
Limitations to This Project
虽然这个项目很有趣,而且得出的结果也与过往研究一致,但我们想赶紧指出,这项工作出于几个原因并不符合学术研究的标准。
While this project was fun and revealed results that are consistent with past research, we want to be quick to note that this work does not meet the standard of academic research for a few reasons.
首先,我们可以从多个角度定义过度自信。2 这种测试捕捉到的主要形式的过度自信是过度估计——你认为自己有 70% 的正确率,但实际上只有 60%——但过度精确,即提供过度狭窄的结果区间的倾向,也起了一定作用。可以说,过度自信没有一个简单的定义,因此也没有统一的测试方法。
To begin, we can define overconfidence in multiple ways.24 The primary form of overconfidence that this test captures is overestimation—you think you’re right 70 percent of the time but you’re only right 60 percent— but overprecision, the tendency to provide ranges of outcomes that are too narrow, also plays a role. Suffice it to say that overconfidence has no simple definition and hence there’s no uniform way to test it.
我们的参与者如何作答也反映了我们提出的问题本身。我们尽量让问题多样化,但很可能同一个人面对另一套问题时就会得出不同的结果。换句话说,一个人的回答可能因问题性质不同而在不同测试间有所变化。
How our participants answered also reflected the questions we posed. We attempted to have questions that were varied, but it is likely that an individual may have a different result for a separate set of questions. In other words, an individual’s results might vary from test to test based on the nature of the questions.
我们的默认百分比是按 10 个百分点递增的,这限制了参与者提供更细致回答的能力。可能也加剧了极端结果——大量 50% 和 100% 的答案。
Our default percentages were in increments of 10 percentage points, which limited the ability of the participants to provide greater subtlety in their responses. This may have also encouraged the result of the extremes—lots of 50 and 100 percent answers.
最后,我们的测试样本可能存在偏差。大多数受试者是在社交媒体上看到该网站后进入的。来自四个国家的参与者构成了样本的四分之三以上,其中包括 37% 来自美国、21% 来自澳大利亚、15% 来自荷兰和 5% 来自英国。几乎所有参与者都是在四种操作系统之一上完成测试的,其中包括 40% 使用 Windows、27% 使用 iOS、15% 使用 Mac OS 和 14% 使用 Android。
Finally, we may have a biased sample of test takers. Most entered the site after having seen it mentioned in social media communication. Participants from four countries constituted more than three-fourths of the sample, including 37 percent from the U.S., 21 percent from Australia, 15 percent from the Netherlands, and 5 percent from the U.K. Nearly all participants took the test on one of four operating systems, including 40 percent on Windows, 27 percent on iOS, 15 percent on Mac OS, and 14 percent on Android.
高 RQ 人群的共同特征
Characteristics of People with High RQ
作为一项大规模预测项目的组成部分,研究人员已经识别出最优秀预测者的特征。²⁵ 他们将相关技能拆解为三个变量:倾向性因素、情境性因素和行为性因素。我们认为这些因素与高 RQ(理性商数)相符。以下是一些要点:
As part of a large-scale forecasting project, researchers have identified the characteristics of the very best forecasters.25 They break down the skills into three variables: dispositional, situational, and behavioral. We believe these are consistent with high RQ. Here are some highlights:
倾向性 – 进行归纳推理 – 展现认知控制²⁶ – 善于数量推理 – 保持积极开放的心态 – 对封闭状态需求有限
Dispositional Engage in inductive reasoning Exhibit cognitive control26 Comfortable with numerical reasoning Actively open-minded Have a limited need for closure
情境因素——接受概率推理训练(理解基础概率)
Situational Trained in probabilistic reasoning (understand base rates)
有效融入团队协作
Effective working as part of a team
行为层面——成长型(与固定型相比)思维模式 27
Behavioral Growth (versus fixed) mindset27
附录:用布里尔评分法记分
Appendix: Keeping Score with Brier
心理学家通常使用 布赖尔评分 来衡量概率预测的准确性。
Psychologists commonly use the Brier score as a method for gauging the accuracy of probabilistic forecasts.
气象学家格伦·布赖尔在 20 世纪 50 年代提出了布赖尔评分。28 布赖尔评分最简单的形式是衡量预测误差的平方,即(预测值 − 结果值)²。对于二元事件,如果事件发生,结果取值为 1;如果未发生,则取值为 0。和高尔夫一样,评分越低越好。
Glenn Brier, a meteorologist, developed the score in the 1950s.28 In its simplest form, the Brier score measures the square of the forecast error, or (forecast − outcome)2. For binary events, the value of the outcome is 1 if the event occurs and 0 if it does not. As in golf, a lower score is better.
布赖尔分数可以按 0 到 1 或 0 到 2 的刻度来表示,具体取决于计算方法。我们遵循布赖尔的原始方法,将结果置于 0 到 2 的刻度上。按这种方式计算布赖尔分数时,你会同时考量事件和事件未发生时的预测误差平方。
You can express a Brier score either on a scale of 0 to 1, or 0 to 2, depending on the calculation. We follow Brier’s original approach and place our results on a scale of 0 to 2. When calculating the Brier score this way, you consider the squared forecast error for both the event and the non-event.
附件 8 展示了一位气象学家对未来四天是否下雨的概率预测。例如,在第 2 天,她预测下雨的概率为 80%。同样,我们可以说她预测不下雨的概率为 20%。由于当天确实下了雨,我们在结果列“下雨”下方标记 1,在“无雨”列标记 0。她当天的布赖尔得分为 0.08。对于多次预测,总体布赖尔得分是每次预测得分的平均值。这位气象学家的总体布赖尔得分为 0.25。
Exhibit 8 shows a meteorologist’s probabilistic forecasts for whether it will rain over the next four days. For example, on Day 2, she forecasts an 80 percent probability that it will rain. Likewise, we can say she forecasts a 20 percent probability that it will not rain. Because it did rain, we place a 1 in the outcome column below “Rain” and a 0 in the “No Rain” column. Her Brier score for that day was 0.08. For multiple forecasts, the overall Brier score is the mean of the scores for each forecast. The meteorologist’s overall Brier score comes to 0.25.
表 8:主观正确概率分布
Exhibit 8: Distribution of Subjective Probability of Correctness
| 天 | 下雨 | 无雨 | 布赖尔评分 | |||
|---|---|---|---|---|---|---|
| 预测 | 结果 | 预测 | 结果 | 计算 | 结果 | |
| 1 | 30% | 0 | 70% | 1 | = (0.3-0)² + (0.7-1)² | 0.18 |
| 2 | 80% | 1 | 20% | 0 | = (0.8-1)² + (0.2-0)² | 0.08 |
| 3 | 60% | 0 | 40% | 1 | = (0.6-0)² + (0.4-1)² | 0.72 |
| 4 | 100% | 1 | 0% | 0 | = (1.0-1)² + (0.0-0)² | 0.00 |
| 均值 | 0.25 |
Rain No Rain Brier Score Day Forecast Outcome Forecast Outcome Calculation Result 2 2 1 30% 0 70% 1 = (0.3-0) +(0.7-1) 0.18 2 2 2 80% 1 20% 0 = (0.8-1) +(0.2-0) 0.08 2 2 3 60% 0 40% 1 = (0.6-0) +(0.4-1) 0.72 4 100% 1 0% 0 = (1.0-1)2 +(0.0-0)2 0.00 Mean 0.25
来源:瑞士信贷。
Source: Credit Suisse.
从 0 到 2 的刻度有一个很好的特性。随机猜测的布里尔分数恰好是 0.50。图表 9 展示了对于某个发生的事件(“降雨”),从 0% 到 100% 主观概率所对应的布里尔分数。
The scale from 0 to 2 has a nice feature. Random guesses have a Brier score of exactly 0.50. Exhibit 9 shows the Brier scores for an event that occurs (“Rain”) for subjective probabilities from 0 to 100 percent.
表 9:不同主观概率下“事件发生”的布莱尔评分 2.0
Exhibit 9: Brier Scores of Event That Occurs for Various Subjective Probabilities 2.0
1.5
1.5
Brier Score 1.0
Brier Score 1.0
0.5
0.5
原件此处是表格,PDF 抽取时列结构已丢失,下面只剩按列读出的数字,行列对应关系无法还原。核对数据请打开来源正文。
资料来源:瑞士信贷。
0.0 1.0 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0 Forecast Source: Credit Suisse.
附录 10 展示了近 2000 名参加测试者的 Brier 分数分布情况。采用这一评分标准,Brier 分数低于 0.25 的表现非常出色。但正如我们所看到的,能达到这一水平的人口比例很小。
Exhibit 10 shows the distribution of the Brier scores for the nearly 2,000 people who took the test. Using this scale, Brier scores below 0.25 are very impressive. But as we can see, the percentage of the population that can operate at that level is small.
表 10:1985 名参与者的布里尔分数 25
Exhibit 10: Brier Scores of 1,985 Participants 25
20
20
Frequency (Percent)
Frequency (Percent)
原件此处是表格,PDF 抽取时列结构已丢失,下面只剩按列读出的数字,行列对应关系无法还原。核对数据请打开来源正文。
15 10 5 0 <0.20 .20-.25 .25-.30 .30-.35 .35-.40 .40-.45 .45-.50 .50-.55 .55-.60 .60-.65 .65-.70 .70-.75 .75-.80 .80-.85 >0.85 Brier Score
15 10 5 0 <0.20 .20-.25 .25-.30 .30-.35 .35-.40 .40-.45 .45-.50 .50-.55 .55-.60 .60-.65 .65-.70 .70-.75 .75-.80 .80-.85 >0.85 Brier Score
来源:瑞信。
Source: Credit Suisse.
尾注 1 Sam Ro,“沃伦·巴菲特和查理·芒格在伯克希尔股东大会上说的最重要的事”,《商业内幕》,2015 年 5 月 3 日。
Endnotes 1 Sam Ro, “The most important things Warren Buffett and Charlie Munger said at Berkshire’s annual meeting,” Business Insider, May 3, 2015.
2 基思·E·斯坦诺维奇,《智力测验漏掉了什么:理性思维的心理学》(纽黑文,康涅狄格州:耶鲁大学出版社,2009 年)。
2 Keith E. Stanovich, What Intelligence Tests Miss: The Psychology of Rational Thought (New Haven, CT: Yale University Press, 2009).
3 Meredith C. Frey 和 Douglas K. Detterman 合著的“学术评估测验还是 g?学术评估测验与一般认知能力之间的关系”,《心理科学》期刊,第 15 卷,第 6 期,2004 年 6 月,第 373-378 页。SAT 之前的全称是学术评估测验(Scholastic Assessment Test)。
3 Meredith C. Frey and Douglas K. Detterman, “Scholastic Assessment or g? The Relationship Between the Scholastic Assessment Test and General Cognitive Ability,” Psychological Science, Vol. 15, No. 6, June 2004, 373-378. “SAT” previously stood for Scholastic Assessment Test.
4 Kimberly Ferriman Robertson, Stijn Smeets, David Lubinski, and Camillia P. Benbow, “Beyond the Threshold Hypothesis: Even Among the Gifted and Top Math/Science Graduate Students, Cognitive Abilities, Vocational Interests, and Lifestyle Preferences Matter for Career Choice, Performance, and Persistence,” Current Directions in Psychological Science, Vol. 19, No. 6, December 2010, 346-351.
4 Kimberly Ferriman Robertson, Stijn Smeets, David Lubinski, and Camillia P. Benbow, “Beyond the Threshold Hypothesis: Even Among the Gifted and Top Math/Science Graduate Students, Cognitive Abilities, Vocational Interests, and Lifestyle Preferences Matter for Career Choice, Performance, and Persistence,” Current Directions in Psychological Science, Vol. 19, No. 6, December 2010, 346-351.
5 Keith E. Stanovich 和 Richard F. West,《智力测试漏掉了什么》,《心理学家》杂志,第 27 卷,第 2 期,2014 年 2 月,第 80-83 页。
5 Keith E. Stanovich and Richard F. West, “What Intelligence Tests Miss,” The Psychologist, Vol. 27, No. 2, February 2014, 80-83.
6 Stanovich, 15.
6 Stanovich, 15.
7 Brent Schlender,“比尔与沃伦秀”,《财富》杂志,1998 年 7 月 20 日。
7 Brent Schlender, “The Bill & Warren Show,” Fortune, July 20, 1998.
8 参见 http://www.templeton.org/what-we-fund/grants/the-development-of-a-test-of-rational-thinking。 9 理查德·H·泰勒,《错误的行为:行为经济学的形成》(纽约:W.W. 诺顿公司,2015 年),28-30。
8 See http://www.templeton.org/what-we-fund/grants/the-development-of-a-test-of-rational-thinking. 9 Richard H. Thaler, Misbehaving: The Making of Behavioral Economics (New York: W.W. Norton & Company, 2015), 28-30.
K.I. 曼克特洛,“推理与理性:纯粹与实践”,载于肯·曼克特洛、钟文祥编,《推理心理学:理论与历史视角》(纽约:心理学出版社,2004 年),第 157-177 页。
10 K.I. Manktelow, “Reasoning and rationality: The pure and the practical,” in Ken Manktelow and Man Cheung Chung, eds., Psychology of Reasoning: Theoretical and Historical Perspectives (New York: Psychology Press, 2004), 157-177.
11 Philip E. Tetlock,《专家政治判断:它的准确性有多高?我们又如何知道?》(新泽西州普林斯顿:普林斯顿大学出版社,2006 年),第 47-48 页。我们所说的“信念”,Tetlock 称之为“区分度”。12 感谢 Andrew Mauboussin 搭建了该网站并汇集了结果。
11 Philip E. Tetlock, Expert Political Judgment: How Good Is It? How Can We Know? (Princeton, NJ: Princeton University Press, 2006), 47-48. What we call “conviction,” Tetlock calls “discrimination.” 12 Thanks to Andrew Mauboussin for building the site and gathering the results.
13 菲利普·E·泰特洛克,“良好判断项目”,在瑞信思想领袖论坛上的演讲,2014 年 6 月 11 日。参见“2014 年思想领袖论坛纪要”,瑞信全球金融策略部,2014 年 11 月 12 日。
13 Philip E. Tetlock, “The Good Judgment Project,” talk at Credit Suisse Thought Leader Forum, June 11, 2014. See “2014 Thought Leader Forum Proceedings,” Credit Suisse Global Financial Strategies, November 12, 2014.
14 Don A. Moore、Samuel A. Swift、Angela Minster、Barbara Mellers、Lyle Ungar、Philip Tetlock、Heather H.J. Yang 和 Elizabeth R. Tenney,《多年期地缘政治预测竞赛中的信心校准》,工作论文,2015 年 4 月 16 日。这些研究者发现,其样本中的过度自信程度低于我们的样本。
14 Don A. Moore, Samuel A. Swift, Angela Minster, Barbara Mellers, Lyle Ungar, Philip Tetlock, Heather H.J. Yang, and Elizabeth R. Tenney, “Confidence Calibration in a Multi-Year Geopolitical Forecasting Competition,” Working Paper, April 16, 2015. These researchers found a lower level of overconfidence than in our sample.
这些参与者的主观概率是 65.4%,而实际结果则为 63.3%。
The subjective probabilities for these participants were 65.4 percent and the outcomes were 63.3 percent.
但他们确实发现了在高置信度水平时存在更大的差距。
But they did find a larger gap at high levels of confidence.
15 Max Bazerman,《管理决策中的判断》,第 4 版(纽约:John Wiley & Sons,1998 年),第 32-34 页。
15 Max Bazerman, Judgment in Managerial Decision Making, 4th Edition (New York: John Wiley & Sons, 1998), 32-34.
16 Susan Bernfield, “J.C. Penney 抹去罗恩·约翰逊几乎所有痕迹,”《彭博商业周刊》,2013 年 10 月 22 日。
16 Susan Bernfield, “J.C. Penney Erases Almost All Traces of Ron Johnson,” Bloomberg Business, October 22, 2013.
17 Stanovich, 108.
17 Stanovich, 108.
贾斯汀·克鲁格与戴维·邓宁,“无法胜任且不自知:难以识别自身无能如何导致自我评价膨胀”,《人格与社会心理学杂志》,第 77 卷,第 6 期,1999 年 12 月,第 1121-1134 页。
18 Justin Kruger and David Dunning, “Unskilled and Unaware of It: How Difficulties in Recognizing One’s Own Incompetence Lead to Inflated Self-Assessments,” Journal of Personality and Social Psychology, Vol. 77, No. 6, December 1999, 1121-1134.
19 Nate Silver,“天气预报员不是傻瓜”,《纽约时报杂志》,2012 年 9 月 7 日。
19 Nate Silver, “The Weatherman is Not a Moron,” New York Times Magazine, September 7, 2012.
20 Emily Pronin,《人类判断中偏见的感知与误感知》,《认知科学趋势》,第 11 卷,第 1 期,2007 年 1 月,第 37-43 页。
20 Emily Pronin, “Perception and Misperception of Bias in Human Judgment,” Trends in Cognitive Sciences, Vol. 11, No. 1, January 2007, 37-43.
21 Michael J. Mauboussin 与 Dan Callahan,《基准率手册——销售增长:结合过去以更好地预测未来》,瑞信全球金融策略,2015 年 5 月 5 日。
21 Michael J. Mauboussin and Dan Callahan, “The Base Rate Book – Sales Growth: Integrating the Past to Better Anticipate the Future,” Credit Suisse Global Financial Strategies, May 5, 2015.
22 见 http://gregmankiw.blogspot.com/2013/11/the-excessive-optimism-of-research.html。
22 See http://gregmankiw.blogspot.com/2013/11/the-excessive-optimism-of-research.html.
23 菲利普·E·泰特洛克与丹·加德纳,《超预测:预测的艺术与科学》(纽约:皇冠出版社,2015 年)。
23 Philip E. Tetlock and Dan Gardner, Superforecasting: The Art and Science of Prediction (New York: Crown Publishers, 2015).
唐·摩尔和保罗·J·希利,《过度自信的麻烦》,《心理评论》,第 115 卷,第 2 期,2008 年 4 月,502-517 页。
24 Don Moore and Paul J. Healy, “The Trouble with Overconfidence,” Psychological Review, Vol. 115, No. 2, April 2008, 502-517.
25 号发言人芭芭拉·梅勒斯、埃里克·斯通、帕维尔·阿塔纳索夫、尼克·罗尔博、S·埃姆伦·梅茨、莱尔·昂加尔、迈克尔·M
25 Barbara Mellers, Eric Stone, Pavel Atanasov, Nick Rohrbaugh, S. Emlen Metz, Lyle Ungar, Michael M.
毕晓普、迈克尔·霍罗威茨、埃德·默克尔和菲利普·泰特洛克,《情报分析心理学:世界政治预测准确性的驱动因素》,《实验心理学杂志:应用版》,第 21 卷,第 1 期,2015 年 3 月,第 1–14 页。
Bishop, Michael Horowitz, Ed Merkle, and Philip Tetlock, “The Psychology of Intelligence Analysis: Drivers of Prediction Accuracy in World Politics,” Journal of Experimental Psychology: Applied, Vol. 21, No. 1, March 2015, 1-14.
26 Shane Frederick,“认知反思与决策”,《经济展望杂志》,第 19 卷,第 4 期,2005 年秋季,25-42 页。
26 Shane Frederick, “Cognitive Reflection and Decision Making,” Journal of Economic Perspectives, Vol. 19, No. 4, Fall 2005, 25-42.
27 Carol S. Dweck,《心态:成功心理学新论》(纽约:兰登书屋,2006 年)。
27 Carol S. Dweck, Mindset: The New Psychology of Success (New York: Random House, 2006).
28 Glenn W. Brier,“用概率表达的预测的验证”,《每月天气评论》,第 78 卷,第 1 期,1950 年 1 月,1-3 页。
28 Glenn W. Brier, “Verification of Forecasts Expressed in Terms of Probability,” Monthly Weather Review, Vol. 78, No. 1, January 1950, 1-3.