智商对决策商:区分聪明与决策技能
全球金融策略 www.credit-suisse.com
GLOBAL FINANCIAL STRATEGIES www.credit-suisse.com
IQ 与 RQ:区分聪明才智与决策能力
2015 年 5 月 12 日
IQ versus RQ Differentiating Smarts from Decision-Making Skills May 12, 2015
Authors
Authors
迈克尔·J·莫布森,邮箱:[email protected]
Michael J. Mauboussin [email protected]
IQ RQ 丹·卡拉汉,CFA,[email protected]
IQ RQ Dan Callahan, CFA [email protected]
“保持理性是一种道德责任。”
“Being rational is a moral imperative.”
Charlie Munger1
Charlie Munger1
智商(IQ)与理性商(RQ)是两回事。把 IQ 想象成发动机的马力,RQ 则是实际输出。
Intelligence quotient (IQ) and rationality quotient (RQ) are distinct. Think of IQ as the horsepower of an engine and RQ as the output.
我们分享一个衡量校准能力的经典测试结果,而校准能力是理性的重要一面。校准良好的人知道自己知道什么,也清楚自己不知道什么。
We share the results of a classic test of calibration, which is an important facet of rationality. Well calibrated people know what they know and know what they don’t know.
与过往研究一致,我们发现参与者高估了自己的准确度——他们主观概率估算值往往高于实际正确率。
Consistent with past research, we find that participants overestimate their accuracy as their subjective probability estimates tend to be higher than the actual percent correct.
投资者和高管可以通过记录成败、向他人请教、参考基础概率以及更新概率判断,来提升自身的理性程度。
Investors and executives can improve their rationality by keeping score, asking about others, using base rates, and updating probabilities.
一个大规模预测项目显示,最优秀的预测者采用归纳和数字推理,具备认知控制能力和成长型思维,且思想开放,在团队合作中高效。
A large-scale forecasting project has shown that the best forecasters use inductive and numerical reasoning, have cognitive control and a growth mindset, and are open-minded and effective working as part of a team.
Introduction
Introduction
多伦多大学应用心理学教授基思·斯坦诺维奇(Keith Stanovich)对智商(IQ)与理性商数(RQ)做了区分。2 心理学家通过特定测试衡量智商,例如韦氏成人智力量表,而它也与 SAT 等标准化测试高度相关。3
Keith Stanovich, a professor of applied psychology at the University of Toronto, distinguishes between intelligence quotient (IQ) and rationality quotient (RQ).2 Psychologists measure IQ through specific tests, including the Wechsler Adult Intelligence Scale, and it correlates highly with standardized tests such as the SAT.3
IQ 衡量的是真实存在的能力,并且它与某些结果有关联。例如,在 SAT 数学部分得分处于前百分之一里最顶尖的十分之一(即第 99.9 百分位)的 13 岁孩子,获得数学或科学博士学位的可能性是那些得分在前百分之一里最底部的十分之一(即第 99.1 百分位)孩子的 18 倍。⁴
IQ measures something real, and it is associated with certain outcomes. For example, thirteen-year-old children who scored in the top decile of the top percent (99.9th percentile) on the math section of the SAT were eighteen times more likely to earn a doctorate degree in math or science than children who scored in the bottom decile of the top percent (99.1st percentile).4
RQ 是一种理性思维能力,进而能让人做出明智决策。虽然我们通常认为智力和理性是相伴而生的,但斯坦诺维奇的研究表明,智商(IQ)与理性商(RQ)之间的相关系数相对较低,仅为 0.20 到 0.35。智商测试的设计初衷并非衡量那些导向审慎决策的思考过程。
RQ is the ability to think rationally and, as a consequence, to make good decisions. Whereas we generally think of intelligence and rationality as going together, Stanovich’s work shows that the correlation coefficient between IQ and RQ is relatively low at .20 to .35.5 IQ tests are not designed to capture the thinking that leads to judicious decisions.
斯坦诺维奇感慨,几乎所有社会都把智力当作焦点,而理性缺失的行为代价却如此高昂。但如果你保持警觉,还是能识别出理性思维的标志。据斯坦诺维奇所言,这些标志包括适应性的行为举动、高效的行为调控、合理的目标优先排序、深思熟虑,以及对证据的正确处理。
Stanovich laments that almost all societies are focused on intelligence when the costs of irrational behavior are so high. But you can pick out the signatures of rational thinking if you are alert to them. According to Stanovich, they include adaptive behavioral acts, efficient behavioral regulation, sensible goal prioritization, reflectivity, and the proper treatment of evidence.6
你的 SAT 分数对这些品质几乎没有参考价值。因此,在评估自己或他人的决策时,第一个教训是要将 IQ(智商)和 RQ(决策商)分开考虑。伯克希尔·哈撒韦的董事长兼首席执行官沃伦·巴菲特将 IQ 比作引擎的马力,将 RQ 比作输出功率。我们都知道有些人 IQ 很高,但 RQ 一般甚至很低。他们的效率很差。还有一些人没有耀眼的 IQ,却能始终做出稳健的决策。他们的效率极高。
Your SAT scores shed little light on any of these qualities. So the first lesson in assessing your own decisions or those of others is to consider IQ and RQ separately. Warren Buffett, chairman and chief executive officer (CEO) of Berkshire Hathaway, equates IQ to the horsepower of an engine and RQ to the output. We all know people who are high on IQ but average or low on RQ. Their efficiency is poor. There are others without dazzling IQs but who consistently make sound decisions. They are highly efficient.
沃伦·巴菲特确实拥有充沛的智力和出色的业绩。但当被问及成功秘诀时,巴菲特强调,真正拉开差距的是 RQ,而非 IQ。
Warren Buffett has plenty of horsepower and output. But when asked about his success, Buffett emphasized that it was RQ that made the big difference, not IQ:7
我是怎么走到今天的,原因很简单。跟智商没关系——听了我这么说,你们肯定很高兴。关键在于理性。我一直把智商和天赋看作是发动机的马力,但输出——也就是这台发动机的工作效率——取决于理性。很多人一开始就装了一台 400 马力的发动机,但实际输出只有 100 马力。与其这样,还不如有一台 200 马力的发动机,却能把全部马力都转化成输出。
How I got here is pretty simple in my case. It's not IQ, I’m sure you'll be glad to hear. The big thing is rationality. I always look at IQ and talent as representing the horsepower of the motor, but that the output—the efficiency with which that motor works—depends on rationality. A lot of people start out with 400-horsepower motors but only get a hundred horsepower of output. It’s way better to have a 200-horsepower motor and get it all into output.
斯坦诺维奇的心理研究支持巴菲特的观察。目前还没有一套全面的测试来测量理性商数(Rationality Quotient, RQ)——斯坦维奇正在开发中——但我们将考察校准能力(calibration),这是理性的重要维度之一。作为该研究的一部分,我们测量了数千人的校准能力。而有趣的是,你也可以参与这个练习,看看自己与其他人相比表现如何。
Stanovich’s psychological research supports Buffett’s observation. While there is not yet a comprehensive test to measure RQ—Stanovich is working on it—we will look at calibration, one of the important facets of rationality.8 As part of this research, we measured the calibration of thousands of people. And part of the fun is that you, too, can participate in the exercise and see how you stack up versus others.
Measuring Rationality
Measuring Rationality
认知科学家和哲学家经常谈及“工具理性”与“认知理性”。工具理性指的是在约束条件下,采取最有利于达成你最渴望目标的行为方式。基于一系列公理的预期效用理论,为如何实现这一目标提供了规范框架。如果你遵循这些公理,你的行为就符合工具理性。9
Cognitive scientists and philosophers talk about “instrumental” and “epistemic” rationality. Instrumental rationality is behaving in such a way that you get what you want the most, subject to constraints. Expected utility theory, which is based on a series of axioms, provides a normative framework for how to do this. You’ll be instrumentally rational if you follow the axioms.9
认知理性描述一个人的信念与现实世界之间的契合程度。比如,如果你相信牙仙子的存在,那就说明你缺乏认知理性。这里有个更简单的办法来记住这两个概念:工具理性解决的是“该做什么”,而认知理性解决的是“什么是真的”。
Epistemic rationality describes how well a person’s beliefs map onto the world. If you believe in the tooth fairy, for instance, you are showing a lack of epistemic rationality. Here’s a catchier way to remember the two terms: instrumental rationality is “what to do” and epistemic rationality is “what is true.”10
我们将集中讨论认识论理性的观察。评估这种理性的一种方法是通过校准测试。想想天气预报员。如果在她预测降雨概率为 70% 的日子里,实际下雨的概率确实是 70%,那么她的预测就是校准良好的。相反,如果那些日子只有 30% 的时间下雨,那么她的预测就是校准不佳的。
We will focus on observations about epistemic rationality. One way to assess this form of rationality is through a test of calibration. Think of a weather forecaster. If it actually rains 70 percent of the time on the days she predicts a 70 percent chance of rain, she is well calibrated. She is poorly calibrated, on the other hand, if it only rains on 30 percent of those days.
表 1 展示了一种衡量打分的方式。横轴衡量的是个人的主观预测(“明天下雨的概率是 70%”),纵轴记录的是实际结果(“下雨了”)。如果一个人的预测结果落在 45 度线附近,那就可以说他校准得很好。
Exhibit 1 shows one approach to keeping score. The horizontal axis measures an individual’s subjective forecast (“there’s a 70 percent chance of rain tomorrow”) and the vertical axis captures the actual outcome (“it rained”). You know that someone is well calibrated if their results fall close to the line at a 45-degree angle.
附录 1:校准的衡量标准
Exhibit 1: Measure of Calibration
100% 90% 80% 70%
100% 90% 80% 70%
Objective Probability 60%
Objective Probability 60%
50%
50%
40% 完美校准线
40% Line of perfect calibration
原件此处是表格,PDF 抽取时列结构已丢失,下面只剩按列读出的数字,行列对应关系无法还原。核对数据请打开来源正文。
30% 20% 10% 0% 0% 10% 20% 30% 40% 50% 60% 70% 80% 90%100% Subjective Probability
30% 20% 10% 0% 0% 10% 20% 30% 40% 50% 60% 70% 80% 90%100% Subjective Probability
来源:瑞士信贷。
Source: Credit Suisse.
信念是另一个重要考量因素。信念衡量的是,当人们给出极端概率时,他们的实际表现如何。校准与信念相关,但两者截然不同。例如,伦敦平均一年中约有半数日子会下雨。所以,如果你每天早上醒来,抛一枚公平的硬币,并记录结果,那么一年下来,你的校准效果看起来会相当不错。
Conviction is another important consideration. Conviction measures how well people do when they assign extreme probabilities.11 Calibration and conviction are related but distinct. For example, it rains about half of the days in an average year in London. So if you wake up every morning, flip a fair coin, and mark your outcome, you will appear well calibrated over one year.
但这并不能帮你计划野餐。你需要的是能够与当天实际天气情况相符的一系列晴天或雨天的预测。这些预测所需的确定性要比投硬币所能提供的高得多。图表 2 左侧展示的是完美校准但低确定性,右侧则是完美校准且高确定性。
But that doesn’t help you plan picnics. What you want are a series of predictions for sun or rain that correspond with the actual weather that day. Those predictions require higher conviction than what the toss of a coin can offer. Exhibit 2 shows perfect calibration but poor conviction on the left, and perfect calibration and conviction on the right.
表格 2:校准与信心 最优可能 最优可能 校准、 校准、 最优可能 校准、 最优可能 校准、 区分能力差 信心差 最优可能 区分能力 最优可能 信心 1 1
Exhibit 2: Calibration and Conviction Best-Possible Best PossibleCalibration, Calibration, Best-Possible Calibration, Best Possible Calibration, Poor Discrimination Poor Conviction Best-Possible Discrimination Best Possible Conviction 1 1
0.8 0.8
0.8 0.8
客观概率
Objective Probability Objective Probability
原件此处是表格,PDF 抽取时列结构已丢失,下面只剩按列读出的数字,行列对应关系无法还原。核对数据请打开来源正文。
0.6 0.6 0.4 0.4 0.2 0.2 0 0 0 0.2 0.4 0.6 0.8 1 0 0.2 0.4 0.6 0.8 1
0.6 0.6 0.4 0.4 0.2 0.2 0 0 0 0.2 0.4 0.6 0.8 1 0 0.2 0.4 0.6 0.8 1
主观概率 主观概率 来源:菲利普·E·泰特洛克,《专家政治判断:有多准?我们如何知道?》(普林斯顿,新泽西州:普林斯顿大学出版社,2005 年),第 48 页。
Subjective Probability Subjective Probability Source: Philip E. Tetlock, Expert Political Judgment: How Good Is It? How Can We Know? (Princeton, NJ: Princeton University Press, 2005), 48.
以下是我们呈现的一项经典校准测试的结果。参与测试的受试者访问了网站 http://confidence.success-equation.com,看到了 50 道判断题。图 3 是该网站的截图。
Here we present the results of a classic calibration test. Subjects who participated went to the website, http://confidence.success-equation.com, and saw 50 true-false questions. Exhibit 3 is a screenshot of the site.
然后,被试需要回答“真”或“假”,并对自己答案的正确概率进行评级,范围从 50% 到 100%,以 10 个百分点为增量。如果你完全不知道答案是真是假,你应该随机选择一个答案,并将正确概率设为“50%”。如果你对自己的答案十分确定,就点击“100%”。
The subjects then answered either true or false and were asked to register a probability of correctness, from 50 to 100 percent, in increments of 10 percentage points. If you have no idea whether the answer is true or false you should select an answer at random and enter “50%” as your probability of correctness. If you are certain of the answer you provide, you click “100%.”
最后,受试者提交自己的答案并收到结果,结果包括:
At the end, the subjects submit their answers and receive their results, which include:
所有问题的平均置信度与正确率
Mean, or average, confidence and percent correct for all questions
正确答案与错误答案的平均置信度
Mean confidence for correct and incorrect answers
低信心(50%-60%)、中等信心(70%-80%)和高信心(90%-100%)提交的正确回答数。
Number correct and answered for low confidence (50-60 percent, medium confidence (70-80 percent) and high confidence (90-100 percent) submissions
一个类似于图 1 的校准图
A calibration graph similar to exhibit 1
我们获取了 1985 名参与者的结果,所有参与者均为匿名。12
We accessed the results of 1,985 participants, all of whom were anonymous.12
表 3:校准问题与正确概率
Exhibit 3: Calibration Questions and Probability of Correctness
Source: http://confidence.success-equation.com.
Source: http://confidence.success-equation.com.
附件 4 显示了结果。这一模式与研究人员几十年来反复验证的发现一致:主观概率的平均估计值显著高于实际正确率。就全体人群而言,主观概率的平均值为 70%,而实际正确率仅为接近 60%。
Exhibit 4 shows the results. The pattern is consistent with what researchers have found for decades: Subjective probability estimates are substantially higher, on average, than the actual percent correct. For the whole population, the average subjective probability was 70 percent and the actual percent correct was just under 60 percent.
附表 4:1985 位参与者的校准结果
Exhibit 4: Calibration for 1,985 Participants
原件此处是表格,PDF 抽取时列结构已丢失,下面只剩按列读出的数字,行列对应关系无法还原。核对数据请打开来源正文。
100% 90% 80% 70% Correct 60% 50% 40% 30% 30% 40% 50% 60% 70% 80% 90% 100% Confidence
100% 90% 80% 70% Correct 60% 50% 40% 30% 30% 40% 50% 60% 70% 80% 90% 100% Confidence
来源:http://confidence.success-equation.com
Source: http://confidence.success-equation.com.
请注意,正确的校准并不要求每次都准确无误(右上角的那个点说明要么有人作弊了,要么是神明参与了测试),而是要求结果接近那条 45 度线。
Note that proper calibration does not require being right all of the time (the dot in the upper right-hand corner indicates that either someone cheated or a deity took the test) but rather being close to the 45 degree line.
知道自己知道什么,也知道自己不知道什么。
It’s knowing what you know and knowing what you don’t know.
图 5 展示了回答的分布情况,横轴代表自信水平减去正确百分比,纵轴代表频率。正态分布(即钟形曲线)很好地描述了这些数据,均值和标准差均为约 10%。
Exhibit 5 shows the distribution of responses, with the horizontal axis representing the confidence level minus the percent correct and the vertical axis the frequency. A normal distribution, or bell curve, describes the data well, with a mean and standard deviation of about 10 percent.
这一分布让我们将参与者分为三类:过度自信者、自信不足者以及校准良好者。我们将校准良好定义为主观概率与正确百分比相差在 1 个百分点以内。
This distribution allows us to segregate the participants into three groups: those who are overconfident, underconfident, and well calibrated. We define well calibrated as a subjective probability within 1 percentage point of the percent correct.
根据这一标准,82.7% 的参与者表现出过度自信(主观信心超过了实际正确率),11.7% 表现出信心不足(主观信心低于实际正确率),只有 5.6% 的人校准良好。
Based on that criterion, 82.7 percent of the participants were overconfident (subjective confidence exceeded actual percent correct), 11.7 percent were underconfident (subjective confidence less than actual percent correct), and only 5.6 percent were well calibrated.
附件 5:多数参与者过度自信 25
Exhibit 5: Most Participants Are Overconfident 25
20
20
Frequency (Percent)
Frequency (Percent)
15
15
过度自信的精准十大
Well-Calibrated 10 Over-Under- Confident
原件此处是表格,PDF 抽取时列结构已丢失,下面只剩按列读出的数字,行列对应关系无法还原。核对数据请打开来源正文。
5 Confident 0 <-20 (5)-0 5-10 10-15 15-20 20-25 25-30 30-35 35-40 >40 (10)-(5) 0-5 (20)-(15) (15)-(10)
5 Confident 0 <-20 (5)-0 5-10 10-15 15-20 20-25 25-30 30-35 35-40 >40 (10)-(5) 0-5 (20)-(15) (15)-(10)
信心水平减去正确率(百分比)
Confidence Level Minus Percent Correct (Percent)
源地址:http://confidence.success-equation.com
Source: http://confidence.success-equation.com.
据报道,著名心理学家阿莫斯·特沃斯基(Amos Tversky)曾说过,人类只能区分三种概率等级:“会发生”“不会发生”和“也许”。¹³ 图表 6 展示了主观正确概率的分布情况。
Amos Tversky, the renowned psychologist, is reported to have said that humans can only distinguish between three levels of probability: “it’s gonna happen,” “it’s not gonna happen,” and “maybe.”13 Exhibit 6 shows the distribution of subjective probabilities of correctness.
42% 的回答是 50%,这等于在说,“我完全不知道答案是什么。”这和特沃斯基所说的“也许”相对应。请注意,这里可能还有一个额外效应,因为 50% 是该网站的默认设定。所以,如果参与者不去修改已分配的“主观概率”,系统就会自动记录为 50%。
Forty-two percent of the responses were 50 percent, which is the equivalent of saying, “I have no idea what the answer is.” This corresponds to Tversky’s “maybe.” Note that there is likely an additional effect here because 50 percent is the default setting on the site. So if the participant doesn’t change the assigned subjective probability, it registers 50 percent automatically.
接下来的回答中,近四分之一的选项是 100%。这就是特沃斯基说的“板上钉钉”或“绝无可能”。于是,近三分之二的回答要么是“我不知道”,要么是“我知道”,剩余的回答则分散在这两个极端之间的四个选项上。
The next most popular response, nearly one-quarter of the total, was 100 percent. This is Tversky’s “it’s gonna happen” or “it’s not gonna happen.” So nearly two-thirds of the responses were either “I don’t know” or “I do know” with the balance split between the four choices in between those extremes.
附件 6:正确性主观概率分布
Exhibit 6: Distribution of Subjective Probabilities of Correctness
原件此处是表格,PDF 抽取时列结构已丢失,下面只剩按列读出的数字,行列对应关系无法还原。核对数据请打开来源正文。
45% 41.9% 40% 35% 30% 25% 23.2% 20% 15% 9.4% 9.5% 9.1% 10% 7.0% 5% 0% 50% 60% 70% 80% 90% 100%
45% 41.9% 40% 35% 30% 25% 23.2% 20% 15% 9.4% 9.5% 9.1% 10% 7.0% 5% 0% 50% 60% 70% 80% 90% 100%
来源:http://confidence.success-equation.com
Source: http://confidence.success-equation.com.
这就引出了一个合乎逻辑的跟进问题:每个指定正确概率下的实际结果如何?图表 7 展示了答案。当受试者选择 50% 的概率时,他们答对的概率纯属随机。这意味着他们的校准效果很好。他们不知道,知道自己不知道,回答得也像不知道一样。
This leads to a logical follow up question: How good were the results for each assigned probability of correctness? Exhibit 7 shows the answer. When the subjects selected 50 percent, their probability of being correct was random. This means they were well calibrated. They didn’t know, knew they didn’t know, and answered as if they didn’t know.
然而,随着所设定正确概率的升高,被测试者的校准度反而下降。例如,当被测试者设定概率为 100% 时,他们的正确率仅为 77%。设定为 90% 时,正确率只有 65%。对自身能力的过高估计,在设定正确概率的高端区间最为显著。
However, as the assigned probability of correctness rose, the subjects became less calibrated. For instance, when the subjects selected 100 percent, they were only correct 77 percent of the time. At 90 percent, they were only correct 65 percent of the time. Overestimation of ability was greatest at the high levels of assigned probability of correctness.14
表 7:各正确性概率的校准结果
Exhibit 7: Calibration for Each Assigned Probability of Correctness
原件此处是表格,PDF 抽取时列结构已丢失,下面只剩按列读出的数字,行列对应关系无法还原。核对数据请打开来源正文。
100% 90% 80% Correct 77% 70% 60% 65% 60% 56% 57% 50% 51% 40% 40% 50% 60% 70% 80% 90% 100% Confidence
100% 90% 80% Correct 77% 70% 60% 65% 60% 56% 57% 50% 51% 40% 40% 50% 60% 70% 80% 90% 100% Confidence
Source: http://confidence.success-equation.com.
Source: http://confidence.success-equation.com.
过度自信会带来几个问题。第一个显而易见的问题是:如果你对某个结果高度自信,而实际出错概率又相当高,你就会忽略其他可能性,最终做出糟糕的决策。15 近期的一个例子就是罗恩·约翰逊(Ron Johnson)执掌 J.C. Penney 期间的表现。
Overconfidence can be a problem for a couple of reasons. The first obvious one is if you are highly confident of an outcome and are wrong a relatively high percentage of the time, you will fail to consider alternatives and ultimately make poor decisions.15 One recent example is Ron Johnson’s tenure as CEO of J.C. Penney.
约翰逊凭直觉迅速调整了这家零售商的定位,结果收效甚微。他被赶下台后,这家零售商又恢复了许多过去的做法。16
Johnson, relying on his intuition, quickly repositioned the retailer to poor effect. Following his ouster, the retailer returned to many of its past practices.16
另一个问题是,那些自认为比实际懂得多的人,比起清楚自身局限的人,更缺乏学习和改进的动力。¹⁷ 实际上,一项研究表明,能力最差的人,在他们自以为能做到的事与实际取得的成就之间,差距是最大的。¹⁸
Another problem is that people who think that they know more than they do are less motivated to learn and improve than those who understand their limitations.17 Indeed, one study showed that the least capable people have the largest gap between what they think they can do and what they actually achieve.18
投资者的教训与管理者的箴言
Lessons for Investors and Executives
所有这一切中的好消息是,我们可以通过训练让自己变得更理性。以下是一些建议:
The good news in all of this is that we can train ourselves to be more rational. Here are some ideas:
保持计分。如果你能提出那些在已知时间内会有明确答案的问题,你就有了计分的依据。经典的做法是通过 Brier 评分,我们在附录中会详细讨论。Brier 评分最初是为了帮助气象学家获得天气预报反馈而开发的。通过改进天气建模技术以及更精准的反馈,如今气象学家的预测准确度比一二十年前大幅提升。
Keep score. To the degree to which you can pose questions that will have a definite answer within a known period of time, you have a basis for keeping score. The classic way to do this is through a Brier score, which we discuss in detail in the appendix. Brier scores were originally developed to help give meteorologists feedback on their predictions for the weather. Through improvements in weather modeling techniques and sharper feedback, meteorologists today are vastly more accurate than they were a generation or two ago.19
问问别人。普林斯顿大学心理学教授埃米莉·普罗宁发现,人们能看到别人思维中的偏见,却不知自己同样深受其害。举例来说,医生们知道药企赠礼会对其他医生产生偏见影响,却认为自己对这类影响免疫。
Ask about others. Emily Pronin, a professor of psychology at Princeton University, has found that while people recognize biases in the thinking of others, they somehow think that they suffer less from the same biases.20 For example, physicians know that gifts from pharmaceutical companies have biasing effects for other doctors but believe they are immune from the effect.
这里有个应对方法。假设你是一位投资者,正在面试一家公司的管理团队,考虑是否买入这支股票。你应该知道,这个管理团队会有乐观偏见,所以对他们说的话要半信半疑。但他们关于其他公司的看法,却更可能准确。换句话说,别问人们自己怎么样,问他们别人怎么样。
Here’s a technique to deal with this. Say you are an investor interviewing a company’s management team, and you are considering buying the stock. You should know that the management team will have an optimism bias, and so you have to take what they say with a grain of salt. But their views about other companies are more likely to be accurate. In other words, don’t ask people about themselves, ask them about others.
运用基础概率。或许最有效的去偏误工具,就是运用基础概率。²¹ 尽管我们都喜欢认为自己独一无二,但问问当别人处于同样境地时发生过什么,会非常有帮助。美国前哈佛大学校长、前财政部长劳伦斯·萨默斯对他手下的研究助理有一条规矩:他会问一个项目需要多长时间,然后他会把助理的答案翻倍,再把时间单位升一级。²² 所以“两小时”会变成“四天”。也许这只是一个避免失望的机制,但新估算的方向无疑是对的。
Use base rates. Perhaps the single most effective de-biasing tool is the use of base rates.21 While we all like to think of ourselves as unique, asking what happened when others were in the same situation can be very helpful. Larry Summers, the former president of Harvard University and Secretary of the Treasury of the U.S., had a rule he used with his research assistants. He would ask how long a project would take. And then he would take the assistant’s answer, double the estimate, and move up to the next unit of time.22 So “two hours” would be translated as “four days.” Perhaps this was simply a mechanism to avoid disappointment, but the direction of the new estimate was no doubt correct.
更新概率。我们提供给参与者的测试是静态的。在真实世界中,概率每时每刻都在变化。理性思考的一大关键挑战,就是在新信息出现时准确地更新概率。事实证明,最顶尖的预测者在这方面做得非常好,并且会使用非常细化的概率增量。²³ 我们大多数人都会陷入确认偏误的陷阱,宁愿忽视或低估新信息,也不愿将其恰当地纳入我们的评估之中。
Update probabilities. The test we shared with the participants was static. In the real world, probabilities shift all of the time. One of the key challenges in rational thinking is to accurately update probabilities as new information arrives. It turns out that the very best forecasters do this very well and use very granular increments of probability.23 Most of us fall into the trap of confirmation bias, preferring to disregard or discount new information than to properly incorporate it into our assessment.
本项目的局限性
Limitations to This Project
尽管这个项目很有趣,且得出的结果与以往研究一致,但我们想迅速指出,这项工作由于几个原因并未达到学术研究的标准。
While this project was fun and revealed results that are consistent with past research, we want to be quick to note that this work does not meet the standard of academic research for a few reasons.
首先,我们可以从多个角度定义过度自信。24 这项测试捕捉到的主要过度自信形式是过度估计——你认为自己有 70% 的时间是对的,但实际上只有 60%——但过度精确,即倾向于给出过于狭窄的结果区间,也起到一定作用。一言以蔽之,过度自信没有简单的定义,因此也没有统一的测试方法。
To begin, we can define overconfidence in multiple ways.24 The primary form of overconfidence that this test captures is overestimation—you think you’re right 70 percent of the time but you’re only right 60 percent— but overprecision, the tendency to provide ranges of outcomes that are too narrow, also plays a role. Suffice it to say that overconfidence has no simple definition and hence there’s no uniform way to test it.
参与者如何作答,也反映了我们提出的问题。我们试图让问题多样化,但一个人在面对另一套问题时,很可能得出不同的结果。换句话说,同一个人在不同测试中的结果,可能因问题的性质而有所差异。
How our participants answered also reflected the questions we posed. We attempted to have questions that were varied, but it is likely that an individual may have a different result for a separate set of questions. In other words, an individual’s results might vary from test to test based on the nature of the questions.
我们的默认百分比以 10 个百分点为间隔,这限制了参与者在回答时提供更细致入微的可能性。这可能也助长了极端结果的出现——大量 50% 和 100% 的答案。
Our default percentages were in increments of 10 percentage points, which limited the ability of the participants to provide greater subtlety in their responses. This may have also encouraged the result of the extremes—lots of 50 and 100 percent answers.
最后,我们的测试样本可能存在偏差。大多数参与者在看到社交媒体上的提及后才进入这个网站。来自四个国家的参与者占了样本总数的四分之三以上,其中 37% 来自美国,21% 来自澳大利亚,15% 来自荷兰,5% 来自英国。几乎所有参与者都在四种操作系统之一上完成测试,包括 40% 使用 Windows,27% 使用 iOS,15% 使用 Mac OS,以及 14% 使用 Android。
Finally, we may have a biased sample of test takers. Most entered the site after having seen it mentioned in social media communication. Participants from four countries constituted more than three-fourths of the sample, including 37 percent from the U.S., 21 percent from Australia, 15 percent from the Netherlands, and 5 percent from the U.K. Nearly all participants took the test on one of four operating systems, including 40 percent on Windows, 27 percent on iOS, 15 percent on Mac OS, and 14 percent on Android.
高 RQ 人群的特征
Characteristics of People with High RQ
作为一项大规模预测项目的一部分,研究人员识别出了最优秀预测者的特征。25 他们将预测技能分解为三个变量:倾向性(dispositional)、情境性(situational)和行为性(behavioral)。我们认为这些变量与高 RQ 相符。以下是一些要点:
As part of a large-scale forecasting project, researchers have identified the characteristics of the very best forecasters.25 They break down the skills into three variables: dispositional, situational, and behavioral. We believe these are consistent with high RQ. Here are some highlights:
性情倾向——运用归纳推理——展现认知控制能力²⁶——对数值推理感到自在——保持主动开放的心态——对结论的需求有限
Dispositional Engage in inductive reasoning Exhibit cognitive control26 Comfortable with numerical reasoning Actively open-minded Have a limited need for closure
情境型——接受过概率推理训练(理解基础比率)
Situational Trained in probabilistic reasoning (understand base rates)
有效作为团队一员开展工作
Effective working as part of a team
行为——成长型(相对于固定型)心态²⁷
Behavioral Growth (versus fixed) mindset27
附录:用布里尔评分法(Brier Score)来衡量
Appendix: Keeping Score with Brier
心理学家通常用布里尔分数来衡量概率预测的精确度。
Psychologists commonly use the Brier score as a method for gauging the accuracy of probabilistic forecasts.
气象学家格伦·布里尔(Glenn Brier)在 20 世纪 50 年代提出了这一评分。最简形式下,布里尔评分衡量的是预测误差的平方,即(预测值 − 结果值)²。对于二元事件,事件发生时结果值取 1,未发生时取 0。与高尔夫类似,得分越低越好。
Glenn Brier, a meteorologist, developed the score in the 1950s.28 In its simplest form, the Brier score measures the square of the forecast error, or (forecast − outcome)2. For binary events, the value of the outcome is 1 if the event occurs and 0 if it does not. As in golf, a lower score is better.
布莱尔评分的取值区间既可以为 0 到 1,也可以为 0 到 2,具体取决于计算方法。我们沿用了布莱尔最初的做法,将结果置于 0 到 2 的区间内。以这种口径计算布莱尔评分时,需要同时考虑事件与非事件的预测误差平方。
You can express a Brier score either on a scale of 0 to 1, or 0 to 2, depending on the calculation. We follow Brier’s original approach and place our results on a scale of 0 to 2. When calculating the Brier score this way, you consider the squared forecast error for both the event and the non-event.
图 8 展示了一位气象学家对未来四天是否会下雨的概率预测。例如,在第二天,她预测降雨概率为 80%。同样地,我们可以说,她预测不降雨的概率为 20%。由于当天确实下了雨,我们在“降雨”一栏的结果列填入 1,在“无降雨”一栏填入 0。那一天的布赖尔评分是 0.08。对于多次预测,总体布赖尔评分是每次预测分数的平均值。这位气象学家的总体布赖尔评分为 0.25。
Exhibit 8 shows a meteorologist’s probabilistic forecasts for whether it will rain over the next four days. For example, on Day 2, she forecasts an 80 percent probability that it will rain. Likewise, we can say she forecasts a 20 percent probability that it will not rain. Because it did rain, we place a 1 in the outcome column below “Rain” and a 0 in the “No Rain” column. Her Brier score for that day was 0.08. For multiple forecasts, the overall Brier score is the mean of the scores for each forecast. The meteorologist’s overall Brier score comes to 0.25.
表 8:正确性主观概率的分布
Exhibit 8: Distribution of Subjective Probability of Correctness
| 日 | 有雨 | 无雨 | 布莱尔评分 | |||
|---|---|---|---|---|---|---|
| 预报 | 实际 | 预报 | 实际 | 计算 | 结果 | |
| 1 | 30% | 0 | 70% | 1 | = (0.3-0)² + (0.7-1)² | 0.18 |
| 2 | 80% | 1 | 20% | 0 | = (0.8-1)² + (0.2-0)² | 0.08 |
| 3 | 60% | 0 | 40% | 1 | = (0.6-0)² + (0.4-1)² | 0.72 |
| 4 | 100% | 1 | 0% | 0 | = (1.0-1)² + (0.0-0)² | 0.00 |
| 均值 | 0.25 |
Rain No Rain Brier Score Day Forecast Outcome Forecast Outcome Calculation Result 2 2 1 30% 0 70% 1 = (0.3-0) +(0.7-1) 0.18 2 2 2 80% 1 20% 0 = (0.8-1) +(0.2-0) 0.08 2 2 3 60% 0 40% 1 = (0.6-0) +(0.4-1) 0.72 4 100% 1 0% 0 = (1.0-1)2 +(0.0-0)2 0.00 Mean 0.25
来源:瑞士信贷。
Source: Credit Suisse.
0 到 2 的刻度有一个很好的特性。随机猜测的布里尔分数恰好在 0.50。表 9 显示了对一个发生的事件(“下雨”)从 0% 到 100% 的主观概率对应的布里尔分数。
The scale from 0 to 2 has a nice feature. Random guesses have a Brier score of exactly 0.50. Exhibit 9 shows the Brier scores for an event that occurs (“Rain”) for subjective probabilities from 0 to 100 percent.
附件 9:不同主观概率下事件发生的布赖尔评分 2.0
Exhibit 9: Brier Scores of Event That Occurs for Various Subjective Probabilities 2.0
1.5
1.5
Brier Score 1.0
Brier Score 1.0
0.5
0.5
原件此处是表格,PDF 抽取时列结构已丢失,下面只剩按列读出的数字,行列对应关系无法还原。核对数据请打开来源正文。
0.0 1.0 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0 预测来源:瑞士信贷
0.0 1.0 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 0.1 0.0 Forecast Source: Credit Suisse.
附件 10 展示了近 2000 名参加测试者的布莱尔评分分布情况。采用这一评分标准,低于 0.25 的布莱尔分数相当出色。但我们可以看到,能达到这一水平的人口比例很小。
Exhibit 10 shows the distribution of the Brier scores for the nearly 2,000 people who took the test. Using this scale, Brier scores below 0.25 are very impressive. But as we can see, the percentage of the population that can operate at that level is small.
表格 10:1985 名参与者的布莱尔评分 25
Exhibit 10: Brier Scores of 1,985 Participants 25
20
20
Frequency (Percent)
Frequency (Percent)
原件此处是表格,PDF 抽取时列结构已丢失,下面只剩按列读出的数字,行列对应关系无法还原。核对数据请打开来源正文。
15 10 5 0 <0.20 .20-.25 .25-.30 .30-.35 .35-.40 .40-.45 .45-.50 .50-.55 .55-.60 .60-.65 .65-.70 .70-.75 .75-.80 .80-.85 >0.85 Brier Score
15 10 5 0 <0.20 .20-.25 .25-.30 .30-.35 .35-.40 .40-.45 .45-.50 .50-.55 .55-.60 .60-.65 .65-.70 .70-.75 .75-.80 .80-.85 >0.85 Brier Score
来源:瑞士信贷。
Source: Credit Suisse.
尾注
1 Sam·罗,《沃伦·巴菲特和查理·芒格在伯克希尔股东大会上说的最重要的事》,《商业内幕》网站,2015 年 5 月 3 日。
Endnotes 1 Sam Ro, “The most important things Warren Buffett and Charlie Munger said at Berkshire’s annual meeting,” Business Insider, May 3, 2015.
2 基思·E·斯塔诺维奇,《智力测试遗漏了什么:理性思维的心理学》(纽黑文,康涅狄格州:耶鲁大学出版社,2009 年)。
2 Keith E. Stanovich, What Intelligence Tests Miss: The Psychology of Rational Thought (New Haven, CT: Yale University Press, 2009).
3 Meredith C. Frey 与 Douglas K. Detterman,“学业评估测试还是 g 因子?学业评估测试与一般认知能力之间的关系”,《心理科学》,第 15 卷,第 6 期,2004 年 6 月,第 373–378 页。“SAT”此前代表学业评估测试(Scholastic Assessment Test)。
3 Meredith C. Frey and Douglas K. Detterman, “Scholastic Assessment or g? The Relationship Between the Scholastic Assessment Test and General Cognitive Ability,” Psychological Science, Vol. 15, No. 6, June 2004, 373-378. “SAT” previously stood for Scholastic Assessment Test.
4 Kimberly Ferriman Robertson, Stijn Smeets, David Lubinski, and Camillia P. Benbow, “超越阈值假设:即便在天才与顶尖数学/科学研究生中,认知能力、职业兴趣与生活方式偏好仍对职业选择、表现与坚持具有影响”,《心理科学当前方向》,第 19 卷,第 6 期,2010 年 12 月,第 346–351 页。
4 Kimberly Ferriman Robertson, Stijn Smeets, David Lubinski, and Camillia P. Benbow, “Beyond the Threshold Hypothesis: Even Among the Gifted and Top Math/Science Graduate Students, Cognitive Abilities, Vocational Interests, and Lifestyle Preferences Matter for Career Choice, Performance, and Persistence,” Current Directions in Psychological Science, Vol. 19, No. 6, December 2010, 346-351.
Keith E. Stanovich 与 Richard F. West,“智力测试遗漏了什么”,《心理学家》,第 27 卷,第 2 期,2014 年 2 月,第 80-83 页。
5 Keith E. Stanovich and Richard F. West, “What Intelligence Tests Miss,” The Psychologist, Vol. 27, No. 2, February 2014, 80-83.
6 Stanovich, 15.
6 Stanovich, 15.
7 Brent Schlender,《比尔与沃伦秀》,《财富》杂志,1998 年 7 月 20 日。
7 Brent Schlender, “The Bill & Warren Show,” Fortune, July 20, 1998.
8 参见 http://www.templeton.org/what-we-fund/grants/the-development-of-a-test-of-rational-thinking。9 理查德·H·塞勒,《错误行为:行为经济学的形成》(纽约:W.W. 诺顿公司,2015 年),第 28-30 页。
8 See http://www.templeton.org/what-we-fund/grants/the-development-of-a-test-of-rational-thinking. 9 Richard H. Thaler, Misbehaving: The Making of Behavioral Economics (New York: W.W. Norton & Company, 2015), 28-30.
10 K.I. Manktelow, “推理与理性:纯粹与实践”,载 Ken Manktelow 与 Man Cheung Chung 编,《推理心理学:理论与历史视角》(纽约:Psychology Press,2004 年),第 157—177 页。
10 K.I. Manktelow, “Reasoning and rationality: The pure and the practical,” in Ken Manktelow and Man Cheung Chung, eds., Psychology of Reasoning: Theoretical and Historical Perspectives (New York: Psychology Press, 2004), 157-177.
11 Philip E. Tetlock,《专家政治判断:它的准确性如何?我们如何得知?》(普林斯顿,新泽西州:普林斯顿大学出版社,2006 年),第 47-48 页。我们称之为“信念感”的东西,Tetlock 称之为“辨别力”。12 感谢 Andrew Mauboussin 搭建网站并收集结果。
11 Philip E. Tetlock, Expert Political Judgment: How Good Is It? How Can We Know? (Princeton, NJ: Princeton University Press, 2006), 47-48. What we call “conviction,” Tetlock calls “discrimination.” 12 Thanks to Andrew Mauboussin for building the site and gathering the results.
13 菲利普·E·泰特洛克(Philip E. Tetlock),“优秀判断项目”演讲,瑞信思想领袖论坛,2014 年 6 月 11 日。参见《2014 思想领袖论坛会议纪要》,瑞信全球金融策略部,2014 年 11 月 12 日。
13 Philip E. Tetlock, “The Good Judgment Project,” talk at Credit Suisse Thought Leader Forum, June 11, 2014. See “2014 Thought Leader Forum Proceedings,” Credit Suisse Global Financial Strategies, November 12, 2014.
14 Don A. Moore、Samuel A. Swift、Angela Minster、Barbara Mellers、Lyle Ungar、Philip Tetlock、Heather H.J. Yang 和 Elizabeth R. Tenney,《多年地缘政治预测竞赛中的自信校准》,工作论文,2015 年 4 月 16 日。这些研究者发现的过度自信程度低于我们样本中的水平。
14 Don A. Moore, Samuel A. Swift, Angela Minster, Barbara Mellers, Lyle Ungar, Philip Tetlock, Heather H.J. Yang, and Elizabeth R. Tenney, “Confidence Calibration in a Multi-Year Geopolitical Forecasting Competition,” Working Paper, April 16, 2015. These researchers found a lower level of overconfidence than in our sample.
这些参与者的主观概率为 65.4%,而实际结果为 63.3%。
The subjective probabilities for these participants were 65.4 percent and the outcomes were 63.3 percent.
但他们确实发现,在高置信度水平下,两者之间的差距更大。
But they did find a larger gap at high levels of confidence.
15 Max Bazerman,《管理决策中的判断》,第 4 版(纽约:John Wiley & Sons,1998 年),第 32-34 页。
15 Max Bazerman, Judgment in Managerial Decision Making, 4th Edition (New York: John Wiley & Sons, 1998), 32-34.
16 Susan Bernfield, “J.C. Penney 几乎抹去了罗恩·约翰逊的所有痕迹,” 《彭博商业周刊》, 2013 年 10 月 22 日。
16 Susan Bernfield, “J.C. Penney Erases Almost All Traces of Ron Johnson,” Bloomberg Business, October 22, 2013.
17 Stanovich, 108.
17 Stanovich, 108.
18 Justin Kruger 和 David Dunning,《无技能且不自知:识别自身无能之困难如何导致自我评价膨胀》,《人格与社会心理学杂志》,第 77 卷,第 6 期,1999 年 12 月,第 1121-1134 页。
18 Justin Kruger and David Dunning, “Unskilled and Unaware of It: How Difficulties in Recognizing One’s Own Incompetence Lead to Inflated Self-Assessments,” Journal of Personality and Social Psychology, Vol. 77, No. 6, December 1999, 1121-1134.
19 Nate Silver,“天气预报员不是傻瓜”,《纽约时报杂志》,2012 年 9 月 7 日。
19 Nate Silver, “The Weatherman is Not a Moron,” New York Times Magazine, September 7, 2012.
20 Emily Pronin, “人类判断中偏见的感知与误感知,”《认知科学趋势》,第 11 卷,第 1 期,2007 年 1 月,第 37-43 页。
20 Emily Pronin, “Perception and Misperception of Bias in Human Judgment,” Trends in Cognitive Sciences, Vol. 11, No. 1, January 2007, 37-43.
21 Michael J. Mauboussin 和 Dan Callahan,《基础率手册——销售增长:整合过去以更好地预测未来》,瑞信全球金融策略,2015 年 5 月 5 日。
21 Michael J. Mauboussin and Dan Callahan, “The Base Rate Book – Sales Growth: Integrating the Past to Better Anticipate the Future,” Credit Suisse Global Financial Strategies, May 5, 2015.
22 参见 http://gregmankiw.blogspot.com/2013/11/the-excessive-optimism-of-research.html。
22 See http://gregmankiw.blogspot.com/2013/11/the-excessive-optimism-of-research.html.
菲利普·E·泰特洛克、丹·加德纳,《超预测:预测的艺术与科学》(纽约:皇冠出版社,2015 年)。
23 Philip E. Tetlock and Dan Gardner, Superforecasting: The Art and Science of Prediction (New York: Crown Publishers, 2015).
24 Don Moore 与 Paul J. Healy,《过度自信的麻烦》,《心理学评论》,第 115 卷,第 2 期,2008 年 4 月,第 502–517 页。
24 Don Moore and Paul J. Healy, “The Trouble with Overconfidence,” Psychological Review, Vol. 115, No. 2, April 2008, 502-517.
25 Barbara Mellers, Eric Stone, Pavel Atanasov, Nick Rohrbaugh, S. Emlen Metz, Lyle Ungar, Michael M.
25 Barbara Mellers, Eric Stone, Pavel Atanasov, Nick Rohrbaugh, S. Emlen Metz, Lyle Ungar, Michael M.
毕晓普、迈克尔·霍罗威茨、埃德·默克尔和菲利普·泰特洛克,“情报分析心理学:世界政治预测准确性的驱动因素”,《实验心理学杂志:应用》,第 21 卷,第 1 期,2015 年 3 月,第 1 - 14 页。
Bishop, Michael Horowitz, Ed Merkle, and Philip Tetlock, “The Psychology of Intelligence Analysis: Drivers of Prediction Accuracy in World Politics,” Journal of Experimental Psychology: Applied, Vol. 21, No. 1, March 2015, 1-14.
26 Shane Frederick, “认知反思与决策”,《经济展望杂志》,第 19 卷,第 4 期,2005 年秋季,25-42 页。
26 Shane Frederick, “Cognitive Reflection and Decision Making,” Journal of Economic Perspectives, Vol. 19, No. 4, Fall 2005, 25-42.
27 Carol S. Dweck,《心态:成功的新心理学》(纽约:兰登书屋,2006 年)。
27 Carol S. Dweck, Mindset: The New Psychology of Success (New York: Random House, 2006).
28 Glenn W. Brier,“以概率表述的预报验证,”《每月天气评论》,第 78 卷,第 1 期,1950 年 1 月,第 1-3 页。
28 Glenn W. Brier, “Verification of Forecasts Expressed in Terms of Probability,” Monthly Weather Review, Vol. 78, No. 1, January 1950, 1-3.