2014思想领袖论坛:预测的艺术与科学
迈克尔·莫布森 引言
Michael Mauboussin Introduction
我们所有人都身处预测行业。把未来猜对了收益巨大,猜错了代价高昂。事情重大时,我们自然会去求助专家。然而,专家预测未来的能力,在不同领域差异极大。例如,在处理稳定、线性系统的问题时,专家一直优于普通人。想想设计桥梁的工程师,或者与象棋大师对弈的新手。
We are all in the business of forecasting. Getting the future right has significant benefits, and getting it wrong can be very costly. Our natural tendency is to turn to experts when the stakes are high. Yet the ability of experts to predict the future varies widely across a range of realms. For example, experts consistently outperform the average person in dealing with problems in stable and linear systems. Think of an engineer designing a bridge or a novice at chess playing a grandmaster.
相反,在不稳定、非线性的领域,包括经济学、政治学和其他社会体系,专家预测未来的能力却很糟糕。而金融服务业面对的,恰恰就是这类问题。这正是我们瑞信持续寻找原创且富有创意的方式来提升预测能力的原因。这也是为何“预测的艺术与科学”成为 2014 年思想领袖论坛的主题。
In contrast, experts are poor at predicting the future in unstable and non-linear fields, including economics, politics, and other social systems. These are precisely the types of problems the financial services industry faces. That’s why we at Credit Suisse continue to look for original and creative ways to improve our ability to predict. It is also why “The Art and Science of Prediction” was the theme of the Thought Leader Forum in 2014.
2014 年的主题将三个相互关联的概念统一在一起,这三个概念也是当今一些最激动人心的研究的核心:如何融合人类与计算机的能力,何时依赖直觉、何时依赖统计,以及如何训练从而做出更好的决策。本次论坛汇聚了来自多个学科的顶尖思想家,每人探讨了其中一个或多个概念。
The theme in 2014 unites three inter-related concepts that are at the center of some of the most exciting research today: how to blend the ability of humans and computers, when to rely on instincts versus statistics, and how to train to make better decisions. The forum brought together leading thinkers from a range of disciplines, each of whom discussed one or more of these concepts.
第一个概念探讨人类与计算机之间的互动。挑战在于结合人类和计算机的最佳能力,同时避免各自的局限。“自由式国际象棋”就是一个很好的例子——选手之间对弈,但可以选择借助计算机输入。1997 年,“深蓝”击败了世界冠军加里·卡斯帕罗夫,这是机器对人类的决定性胜利。
The first concept explores the interaction between humans and computers. The challenge is to combine the best of human and computer capabilities while avoiding the limitations of each. Freestyle chess, where individuals play one another but also have the option to get input from computers, is a good illustration. In 1997, Deep Blue beat the world champion, Garry Kasparov, a decisive victory for machine over man.
但最顶尖的自由式棋手,实力超越了纯粹的人类或纯粹的计算机。这引出了一系列有趣的问题:我们何时该听从计算机,何时该否决它?如何在不被海量数据淹没的情况下理解它?又如何将看似来源各异的数据结合起来,帮助我们做出更优的决策?
But the top freestyle chess players are better than either man or machine. This raises a host of interesting questions: When should we defer to a computer and when should we override it? How do we make sense of massive amounts of data without getting bogged down by it? And how can we combine data from seemingly disparate sources to help us make better decisions?
第二个概念是何时相信直觉,何时依赖统计分析。当今许多领域,这两种方法之间的张力正在加剧。在棒球界,就是球探与数据极客的对立。
The second concept is when to trust our intuition and when to rely on statistical analysis. The tension between the two is rising in many fields today. In baseball, it’s the scouts versus the statistics nerds.
但在投资界,是基本面分析师与量化分析师;在政界,是传统竞选团队与新一代统计学家;在医学界,是医生与算法之间关于最佳治疗路径的争论。这种张力与专家何时有用的问题密切相关。一个关键结论是:在非线性、不稳定的领域,专家表现很差。结合这两种观点的一种方法,是权衡赋予基础比率(即过去发生的事情)多大权重,以及赋予个案的具体情况多大权重。运气或技能对最终结果的相对贡献,决定了这个权重。
But it’s also fundamental versus quantitative analysts in investing, the old-school political campaigns versus a new breed of statisticians in politics, and doctors versus algorithms for how best to approach treatment in medicine. This tension is closely related to the question of when experts are useful. A key takeaway is experts are poor in non-linear and unstable domains. One method to combine these viewpoints is to consider how much weight to place on the base rate, which is basically what’s happened before, and how much weight to place on the circumstances of an individual case. The relative contribution of luck or skill to the final outcome informs that weighting.
最后一个概念是关于如何准备,以做出更好的决策。飞行员常用飞行模拟器。在其他领域,有没有类似飞行模拟器的东西?这样的训练可以让我们更有效地为未来可能出现的不同环境做准备。当我们在组织内部进行选择时,比如设计架构和招聘人员,这种准备也很关键。我们还强调了认知多样性对于确保群体保持明智、避免陷入疯狂的重要性。
The final concept is about how to prepare to make better decisions. Airplane pilots commonly use flight simulators. What is the equivalent of a flight simulator in other realms? Such training would allow us to prepare more effectively for the different environments that will arise. Preparation is also relevant when considering the choices we make in our organizations, including the structures we set up and the people we hire. We also emphasize the importance of cognitive diversity in ensuring that the crowd is wise and in preventing it from acting mad.
迈克尔·莫布森 瑞信
Michael Mauboussin Credit Suisse
迈克尔·莫布森是瑞信投资银行部门的董事总经理,常驻纽约。他是全球金融策略主管,凭借其在估值与投资组合定位、资本市场理论、竞争战略分析以及决策制定等领域的专长、研究和著述,为外部客户和瑞信内部专业人士提供思想领导力和战略指导。
Michael Mauboussin is a Managing Director of Credit Suisse in the Investment Banking division, based in New York. He is the Head of Global Financial Strategies, providing thought leadership and strategy guidance to external clients and internally to Credit Suisse professionals based on his expertise, research and writing in the areas of valuation and portfolio positioning, capital markets theory, competitive strategy analysis, and decision making.
在 2013 年重新加入瑞信之前,他曾任美盛资本管理公司的首席投资策略师。迈克尔最初于 1992 年加入瑞信,担任包装食品行业分析师,并于 1999 年被任命为美国首席投资策略师。他曾任纽约消费者分析师协会主席,并多次入选《机构投资者》全美研究团队和《华尔街日报》食品行业全明星调查。
Prior to rejoining Credit Suisse in 2013, he was Chief Investment Strategist at Legg Mason Capital Management. Michael originally joined Credit Suisse in 1992 as a packaged food industry analyst and was named Chief US Investment Strategist in 1999. He is a former president of the Consumer Analyst Group of New York and was repeatedly named to Institutional Investor’s All-America Research Team and The Wall Street Journal All-Star survey in the food industry group.
迈克尔著有《成功方程式:解开商业、体育和投资中的技能与运气》、《三思而后行:驾驭反直觉的力量》以及《比你知道的更多:在非传统之处寻找金融智慧》。他还与阿尔弗雷德·拉帕波特合著了《预期投资:通过解读股价获取更高回报》。
Michael is the author of The Success Equation: Untangling Skill and Luck in Business, Sports, and Investing, Think Twice: Harnessing the Power of Counterintuition, and More Than You Know: Finding Financial Wisdom in Unconventional Places. He is also co-author, with Alfred Rappaport, of Expectations Investing: Reading Stock Prices for Better Returns.
自 1993 年起,迈克尔在哥伦比亚商学院担任金融学兼职教授,同时也是 Heilbrunn 格雷厄姆与多德投资中心的教员。他还是圣塔菲研究所的董事会主席,该所是复杂系统理论跨学科研究的领先中心。迈克尔拥有乔治城大学文学士学位。
Michael has been an adjunct professor of finance at Columbia Business School since 1993 and is on the faculty of the Heilbrunn Center for Graham and Dodd Investing. He is also chairman of the board of trustees of the Santa Fe Institute, a leading center for multi-disciplinary research in complex systems theory. Michael earned an AB from Georgetown University.
迈克尔·莫布森 欢迎致辞
Michael Mauboussin Welcome Remarks
大家早上好。对于还没认识我的朋友们,我叫迈克尔·莫布森,是瑞信全球金融策略主管。
Good morning, everybody. For those I haven’t met, my name is Michael Mauboussin, and I am head of Global Financial Strategies at Credit Suisse.
我谨代表瑞信的所有同事,热烈欢迎大家出席 2014 年思想领袖论坛。昨晚已经和我们相聚的朋友,希望你们度过了一个美好的夜晚。
On behalf of all my colleagues at Credit Suisse, I want to wish you a warm welcome to the 2014 Thought Leader Forum. For those of you who joined us last night, I hope you had a wonderful evening.
今天我们的议程安排非常精彩。
We certainly have a very exciting lineup for today.
在今天上午把时间交给各位演讲嘉宾之前,我想先做几件事。首先,我想强调一下与预测的艺术和科学相关的一些主题。然后我想谈谈这个论坛本身,以及你们每个人可以做些什么来帮助它取得成功。
I want to do a couple of things this morning before I hand it off to our speakers. First, I want to highlight some of the themes associated with the art and science of prediction. I then want to discuss the forum itself and what you all can do to contribute to its success.
让我先从关于预测的艺术和科学的几个宏观想法说起。我特别想谈三点想法。我想你们今天会反复听到这些主题。
Let me start with some high level thoughts on the art and science of prediction. There are three thoughts in particular that I'd like to make. I think you'll hear each of these themes throughout the day.
这三个主题是相互关联的,你们会看到。
Here are the three themes. They're somewhat interrelated, as you'll see.
第一个主题涉及人类与计算机之间的互动。问题是:“我们如何结合人类和计算机能做的最好的事,同时避免他们能做的最坏的事?”
The first relates to the interaction between humans and computers. The question is, “How do we combine the best of what humans and computers can do and avoid the worst of what humans and computers can do?”
第二个主题是经典区分:你应该相信直觉,还是依赖统计分析?知道你的直觉何时可能管用、何时会失效,这很重要。
The second is the classic distinction about whether you should trust your gut or go with statistical analysis. It's important to know when your intuition is likely to work well and when it's going to fail you.
一个相关的问题,我认为与各位所在的组织都息息相关,那就是:当一个组织在决策方法上的思维方式发生转变时,其内部是否必须发生文化变革?
A related issue, which I think is relevant to all of your organizations, is whether there are cultural changes that must occur within an organization as it shifts its thinking about approaches to decision making.
随着更好的方法或方法体系出现,一个组织做出改变的可能性有多大?
As better methods or sets of methods come along, how likely is it that an organization will change?
最后一个主题是关于我们如何准备,以做出更好的决策。这可以是我们如何在组织内部思考选择,也可以是招聘人员的方式。
The final theme is about how we prepare to make better decisions. This could be how we think about choices in our organization or how we hire people.
我们在学校里学习各种各样的主题,但关于如何做出更好决策的课程却少之又少。
We study all sorts of topics in school, but there are very few courses on how to make better decisions.
我们如何弥补这个差距?
How do we close that gap?
让我从第一个主题开始。我得告诉你们,过去六到十二个月里,我一直在思考这个概念。我觉得这很大程度上源于大约八个月前我与肖恩·古尔利的一次讨论。
So let me start at the top. I have to tell you this is a concept that I have been thinking about a lot the last six to 12 months. I think this was really kicked off by a discussion I heard from Sean Gourley about eight months ago.
解释这个问题最好的方式是通过国际象棋。我们可以说 1997 年是机器击败人类的时刻。那是加里·卡斯帕罗夫输给“深蓝”的时候。而如今,最好的计算机程序可以击败最好的人类棋手。
The best way to explain this is through chess. We can say 1997 is when machine beat man. That's when Garry Kasparov lost to Deep Blue. Now, today, the best computer programs can beat the best humans.
在过去大约 15 年里,出现了“自由式国际象棋”——选手之间对弈,但他们也可以借助计算机来指导自己的走法。
In the last 15 years or so, there's been the advent of "freestyle chess," where individuals play one another but they also can avail themselves of computers to help guide their moves.
事实证明,这些自由式棋手比纯粹的人类或纯粹的计算机都强。所以,人加机器,胜过人或机器。
It turns out that those freestyle chess players are better than either man or machine. So, man plus machine beats man or machine.
此外,这些优秀的自由式棋手中有不少并非顶尖棋手。他们不是经过增强的大师。他们只是水平尚可的棋手,但有了增强。
Further, a number of these great freestyle chess players are not great chess players. They're not grandmasters who've been augmented. They're decent players who are augmented.
于是问题就变成了:这些自由式棋手拥有什么样的技能?他们如何知道何时应该听从计算机,又如何知道何时应该否决它?
So the question becomes, what is the skill that these freestyle chess players have? How do they know when to defer to the computer, and how do they know when to override the computer?
迈克尔·莫布森 欢迎致辞(续)
Michael Mauboussin Welcome Remarks (Continued)
这几乎就像量化策略和基本面策略之间的对等关系。你何时知道应该跟着量化走?你何时知道应该跟着基本面走,从而否决量化?我们能否定义这种技能,并思考如何将其应用于公司或投资背景下?
It's almost like this equivalent between quantitative and fundamental strategies. When do you know when to go with a quant? When do you know when to go with a fundamental to override that? Can we define that skill and how we think about that in a corporation or investing context?
下一个想法与接口有关。计算机显然在 0 和 1 方面极为出色。人类则擅长识别模式。
The next idea has to do with the interface. Computers are obviously amazing at ones and zeros. Humans are great at seeing patterns.
我们如何将数据输入计算机,使其能够发挥自身优势?我们又如何处理计算机的输出,从而让我们作为人类能最高效地利用数据?
How do we feed data to the computer so that it can leverage its strength, and how do we compute output from those computers so that we can use the data most effectively as humans?
自然,像可视化这样的技术至关重要。但关键问题在于,在金融、商业、科学和政治领域,我们人类如何能够消化海量数据,进而理解它的意义?
Naturally, techniques such as visualization are essential, but the key question is how we humans will be able to ingest large amounts of data in a way that we can make sense of it–in the world of finance, business, science, and politics.
另一个我觉得也很有趣的相关想法是“混搭”。我们如何将看似来自不同来源的数据整合在一起,帮助我们做出更好的决策?我们今天也会听到一个例子,是关于追踪疾病的。
A related idea, which also I think is fascinating, is that of mash-ups. How can we bring together data from what are seemingly disparate sources to help us make better decisions? We'll hear about that today, as well, in the instance of tracking disease.
最后一点是,我们必须承认,随着技术的发展,我们将创造出比我们自身能理解得更复杂的系统。这引出了耶鲁大学退休社会学教授查尔斯·佩罗所说的“正常事故”。
The final point is that we have to acknowledge that as we develop technologies we're going to create systems that are more complex than we can understand. This leads to what Charles Perrow, who is a retired professor of sociology at Yale, called "normal accidents."
“正常事故”的一个例子是 2010 年 5 月的“闪电崩盘”。一笔看似无害的单笔交易,似乎引发了一系列意想不到的连锁反应。由于其中许多反应都是算法的结果,一旦开始级联,就很难阻止。
One example of a normal accident is the Flash Crash in May of 2010. A single, seemingly innocuous, trade appeared to trigger a series of unanticipated responses. Since many of those responses were a function of algorithms, they were very difficult to stop once they started to cascade.
同样的描述也适用于 2003 年的东北大停电。在俄亥俄州一次非常不起眼的故障,再次导致了这种级联式的失败。
Same description could be used for the Northeast power blackout in 2003. Again, a really innocuous failure in Ohio led to this cascade of failures.
我们所有人面临的问题变成了:在我们利用技术改善世界和提高效率的同时,我们是否也在为正在做的事情注入一定程度的脆弱性,从而导致这些大规模的影响?
The question for all of us becomes, as we use technology to improve our world and improve our efficiency, are we injecting a degree of fragility into what we're doing as well that will lead to these large scale effects?
下一个问题是,何时可以依赖直觉甚至分析,何时又应该依赖统计数据。换一种问法,专家在哪些领域能发挥作用?这显然是个关键问题,关系到我们何时该相信直觉,何时该相信统计数据。
The next issue is when we can rely on our intuition or even our analysis, and when we should rely on statistics. So, differently, where are experts helpful? This is obviously a crucial question as to where we should rely on intuition versus statistics.
有两位非常知名的心理学家,丹尼尔·卡尼曼和加里·克莱因,曾合写过一篇关于这个主题的论文。卡尼曼对专家的看法比克莱因要悲观得多,而克莱因则高度推崇专家的作用。论文的标题几乎说明了一切,叫《无法达成共识》。
There are a couple very well-known psychologists, Danny Kahneman and Gary Klein, who co-wrote a paper on this topic. Kahneman takes a much dimmer view of experts than Klein, who celebrates the role of experts. The title of the paper pretty much said it all. It's called “A Failure to Disagree.”
他们的观点是,思考这个问题的最佳方式之一,就是考虑哪些领域里专家最可能非常有效,哪些领域里他们不太可能有效。
Their point was that one of the best ways to think about this problem is to think about domains in which experts are most likely to be very effective and domains where they are unlikely to be effective.
在稳定且线性的领域里,你可以培养出非常强大的专业能力和直觉判断。
Where domains are stable and linear, you can develop expertise and intuition that is very powerful.
在那些非线性和不稳定的领域里,你会看到大量专家失手。再说一次,我确信这将是贯穿今天始终的另一个话题。
In domains that are non-linear and unstable, you see a lot of failure of experts. Again, this is another theme I'm sure we'll hear a great deal about throughout the day.
一个相关的观点——也是我深以为然的——就是思考什么时候该依赖基础概率,什么时候该依赖具体情境。
A related point, an idea that's very close to my heart, is the consideration of when you should rely on a base rate and when you should rely on the specific circumstances.
卡尼曼获得诺贝尔奖后,一位同事问他:“您最喜欢自己写过的那篇论文?”他提到了 1973 年那篇题为《论预测心理学》的论文。
After Kahneman won the Nobel Prize, a colleague asked him, "What is your favorite paper you ever wrote?" He mentioned the 1973 paper called, "On the Psychology of Prediction."
迈克尔·莫布辛致欢迎辞(续)
Michael Mauboussin Welcome Remarks (Continued)
在那篇与阿莫斯·特沃斯基合著的论文中,他们指出,做出一个预测需要三个要素。
In that paper, which he co-wrote with Amos Tversky, they said that there are three elements you need to make a prediction.
第一个是基础概率,本质上就是过去发生过什么的历史记录。第二个是某个具体案例的实际情况。第三个则是在你的实际预测中,如何权衡这两者的机制。
One is the base rate, which is basically a history of what's happened before. Second is the facts about an individual case. And the third is a mechanism to weight both of those in your actual forecast.
如果你想象一个从“全靠运气、毫无技巧”到“全靠技巧、毫无运气”的连续分布,那么一条简单规则是:如果结果主要由运气决定,你就应该把大部分注意力放在基础比率上。反之,当结果主要取决于技巧时,你就可以更多地关注具体情境。
If you think about a continuum of all luck / no skill to all skill / no luck, one of the simple rules is if an outcome is mostly determined by luck, you should place most of your emphasis on base rates. By contrast, when it's mostly skill, you can focus much more on the specific circumstances.
仅凭知道某项活动在这条连续谱上的位置,就能帮助你以更明智的方式思考如何权衡这两方面因素,并做出更为深思熟虑的统计预测。
Just knowing where an activity is on that continuum can give you an informed way to think about how to weight those two things and make a more thoughtful statistical prediction.
最后一点不具分析性,但在每个方面都同样重要。也就是说,我们如何让组织以不同的方式思考世界,并把统计分析融入它们所做的工作中?这在很大程度上是迈克尔·刘易斯(Michael Lewis)十年前在《点球成金》(Moneyball)中描绘的戏剧性场面——他刻画了球探与统计极客之间的对峙。
The last point is not analytical but it's in every way as important. That is, how can we have organizations think about the world in a different way and incorporate statistical analysis in what they do? This in large part was the drama that was created by Michael Lewis in Moneyball a decade ago, when he depicted this squaring off between the scouts and the statistical nerds.
当然,这场戏并不只在棒球中上演。它无处不在。投资中,是量化分析与基本面分析的对决;政治竞选里,是老派竞选方式与新一代统计专家的较量;广告界,是凭直觉行事的《广告狂人》与使用 A/B 测试的科学家的交锋;医院里,是医生与治疗方案算法之间的博弈。
The drama doesn't end with baseball, of course. It's everywhere you look. It's quantitative versus fundamental analysis in investing. It's old-school political campaigns versus a new breed of statistical people. It's the seat-of-the-pants Mad Men versus scientists using A/B testing in advertising. It's doctors versus algorithms for treatments in the hospital.
因此,从一套方法转向另一套方法的过程中,会催生出新旧派系的对立——组织将如何适应和应对这些变化?
So the degree to which moving from one set of methods to another set of methods creates an old and a new guard – how will organizations adapt and change to those things?
让我用三个要点中的最后一部分来总结,那就是:我们如何为将来做出更好的决策做准备?
Let me wrap up with a final section of my three points, and that is how do we prepare to make better decisions going forward?
每当我思考这个问题时,常常会想到飞行员如何通过训练来提升自己的表现。当然,他们最常用的方式就是使用飞行模拟器。一台飞行模拟器能让飞行员磨练技能,为现实世界中各种不同的环境做好应对准备。
Whenever I think about this, I often think about how pilots train to improve what they are doing. Of course, the way they typically do that is the use of flight simulators. A flight simulator allows a pilot to hone his or her skills in order to be prepared for different environments in the real world.
问题是我们如何组装属于自己的飞行模拟器?我们如何训练自己更有效地思考世界上正在发生的事情?
The question is how do we assemble our own flight simulators? How do we train ourselves to think more effectively about what's going on in the world?
有几点已经得到了充分证实。其中一点,我们昨晚从汤姆·西利的演讲中得到了精彩体现:集体能够非常高效地解决极其困难的问题。
There are a couple of things that have been pretty well established. One of them we saw beautifully demonstrated last night with Tom Seeley's talk: collectives can solve very difficult problems very effectively.
这就 是我们现在常说的“群体智慧”。群体智慧的关键在于——
That is something we now commonly call the wisdom of crowds. Here's the key to the wisdom of crowds
——这一点每次谈到都必须着重强调——它只在特定条件下才有效。你必须营造出合适的条件,群体智慧才能发挥作用。
– and this needs to be underscored every single time you talk about it – it only works under certain conditions. You have to have the proper conditions in place for the wisdom of crowds to be effective.
在所有重要条件中,多样性首当其冲。你可以像昨晚我们看到的蜜蜂那样,在群体内部实现多样性。你也可以让自己头脑中拥有多样性——这做起来要难得多,但同样可以做到,而且能帮你获得更好的结果。
Among the most important conditions is diversity. You can achieve diversity within a group as the honeybees did as we saw last night, or you can actually have diversity in your own head, which is much more difficult to do, but also can be done and help you with your better outcomes.
不久前,我和一位客户聊天,他说了一句话,让我深有同感。他说:“你知道吗,我在投资这行长大,听到的不是‘群体的智慧’,而是‘群体的疯狂’。”我也是如此。我读的第一本书讲的就是群体的疯狂,而不是群体的智慧。
Now, I was with a client not long ago, and he made a statement that completely resonated with me. He said, "You know, I grew up in the investment business hearing not about the wisdom of crowds, but rather hearing about the madness of crowds." That's true for me too. The first book I read was about the madness of crowds not the wisdom of crowds.
迈克尔·莫布森欢迎致辞(续)
Michael Mauboussin Welcome Remarks (Continued)
这就是我最后的想法:我们如何避免群体疯狂?如何避免成为某个群体的成员,尤其是当这个群体正在做一些有损利益的事情时?在我们的组织中,如何建立一种机制,当这种情况出现时能加以利用,或者反过来,在我们的组织内部避免它?
That's really my final thought, which is, how do we avoid the madness of crowds? How do we avoid being part of a group that especially is doing something that's suboptimal? In our organizations, how do we create structures that take advantage of that when it arises, or on the other side, avoid it in our organizations?
这只是我今天开始前的一些想法。我期待在接下来的时间里,与各位一起探讨这些以及其他话题。
Those are just a couple of my thoughts going into the day. I look forward to exploring these and other topics with all of you as we go through this.
正式开始之前,先说几件事。首先,今天到场的还有 Ink Factory 那几位才华横溢的朋友。达斯蒂或苏今天会为我们所有的演讲者做图形记录。昨晚在场的朋友已经看到达斯蒂在埋头工作了。他们要把各位发言人的话提炼成图像和文字,用视觉方式呈现今天你们将听到的每一场主题和每一段谈话的核心概念。
A couple of things before we start. First, I want to mention we have with us today the very talented folks from Ink Factory. Either Dusty or Sue today will be graphically recording all of our presenters today. For those who were here last night, you saw Dusty hard at work. They are synthesizing the words of our speakers into images and text in order to visually represent the key concepts of each of the themes and each of the talks that you'll hear today.
他们的口号——我很喜欢——是:“你来说,我们来画,效果绝佳。”我确实觉得效果极为出色。我想你们都会认同这一点。请随意给那些白板拍照。每次演讲结束后,你们都会看到白板被摆在外面。昨晚的那块白板现在就在外面。
Their slogan – which I love – is, "You talk, we draw, it's awesome." I really do think it is awesome. I think you'll agree with all of that. Please feel free to take pictures of those boards. You'll see after every presentation the boards will be outside. Last night's board is out there now.
最后说说我们今天想做的事情。第一,我们希望让你们见到一些在日常交往中不太可能碰到的演讲者,但他们确实有能力引发有益的思考和对话。我们知道你们的时间非常宝贵,也感谢你们愿意与我们分享这份时间,来开阔自己的视野。
Let me just finish with what we're trying to do today. First, we want to provide you access to speakers who you're unlikely to see in your day-to-day interactions but who are nonetheless capable of provoking useful thought and dialogue. We understand your time is very valuable. And we appreciate that you're sharing some of that time with us in order to expand your horizons.
其次,我们非常希望大家能够自由交流想法。注意,我们的演讲时段比平时稍长一些。每位演讲人之间都安排了休息时间,方便大家互动。我们真正想做的,是鼓励充分的你来我往。
Second, we really want to encourage a free exchange of ideas. Notice that our speaking slots are a little bit longer than normal. We have breaks between all the speakers allowing you to interact. What we're trying to do is really encourage a lot of back and forth.
我们特意称之为“论坛”而非“会议”,因为论坛是对话,是互动。我们鼓励这种氛围。探究、质疑、交流,这些才是真正的核心主题。
We purposefully call this a “forum” instead of a “conference” because a forum is a dialogue, an interaction. We would encourage that environment. Inquire, challenge, exchange are really the key themes.
最后,我们瑞信希望这次活动能给您带来绝佳体验。因此,请随时向我或我们团队的任何成员提出任何需求。我们定会竭尽全力为您提供便利。
Finally, we at Credit Suisse would like this to be a great experience for you. So please don't hesitate to ask me or anybody else on our team for anything. We will certainly do our very best to accommodate you.
宾夕法尼亚大学的菲利普·泰特洛克
Philip Tetlock University of Pennsylvania
菲利普·泰特洛克是宾夕法尼亚大学沃顿商学院和文理学院的安嫩伯格大学教授。他于 1979 年在耶鲁大学获得博士学位,随后虽偶尔在其他机构任职,但 1979 年至 2010 年间主要在加州大学伯克利分校工作。他获得了专业和科学组织颁发的众多奖项,包括美国艺术与科学院、美国国家科学院、美国心理学会、美国政治科学协会以及美国科学促进会的荣誉。他在同行评审期刊上发表了 200 多篇文章,并编辑或撰写了 9 本书。
Philip Tetlock is the Annenberg University Professor at the Wharton School and the School of Arts and Sciences at the University of Pennsylvania. He received his PhD in 1979 from Yale University and, with occasional appointments elsewhere, worked mainly at the University of California Berkeley between 1979 and 2010. He has received numerous awards from professional and scientific organizations, including the American Academy of Arts and Sciences, the National Academy of Sciences, the American Psychological Association, the American Political Science Association, and the American Association for the Advancement of Science. He has published more than 200 articles in peer-refereed journals and edited or written nine books.
菲尔长期关注人类在以下方面面临的巨大困难:“按时间维度思考”——即从历史中推导出逻辑上一致的因果结论,并做出基于经验的准确条件预测。他围绕历史学家如何思考“可能的过去”(如果……可能会、将会或本可以发生的事情)以及政治与经济专家如何思考“可能的未来”(可能或仍可能发生的事情),进行了大量研究。他 2005 年的著作《专家的政治判断:准确度如何?我们能知道吗?》(普林斯顿大学出版社)深入探讨了这些问题。
Phil has a long-standing interest in the enormous difficulty that human beings have in “thinking in time” in drawing logically consistent causal inferences from history and in making empirically accurate conditional forecasts. He has conducted numerous studies of how historians think about “possible pasts” (things that might or would or could have happened if…) and of how political and economic experts think about “possible futures” (things that might or could yet happen). His 2005 book, Expert Political Judgment: How Good Is It? How Can We Know? (Princeton University Press), explores these issues in depth.
他目前(与妻子芭芭·梅勒斯共同担任)是“良好判断项目”的共同首席研究员,该项目由美国情报高级研究计划局资助(2011-2015 年),是一项大规模的预测竞赛。有兴趣作为志愿者预测员参与该科学研究项目的读者,可访问 www.goodjudgmentproject.com。2014 年 2 月,菲尔担任良好判断有限责任公司的首席科学顾问,该公司致力于将“良好判断项目”的科学发现转化为适用于私营及公共部门组织的可落地建议。
He is currently a co-principal investigator (with his wife Barb Mellers) in the Good Judgment Project, a large-scale forecasting tournament sponsored by the Intelligence Advanced Research Projects Agency (2011-2015). Readers who are interested in participating as volunteer forecasters in the scientific research project are encouraged to visit www.goodjudgmentproject.com. In February 2014, Phil took on the role of Chief Scientific Adviser to Good Judgment LLC, a firm dedicated to translating the scientific discoveries of the Good Judgment Project into implementable advice for private- as well as public-sector organizations.
菲利普·泰特洛克的“良好判断力项目”
Philip Tetlock The Good Judgment Project
迈克尔·莫布森:我很高兴介绍我们今天的首位发言者,菲利普·泰特洛克教授。
Michael Mauboussin: I'm very pleased to introduce our first speaker today, Professor Philip Tetlock.
菲尔是宾夕法尼亚大学沃顿商学院和文理学院的双聘心理学教授,头衔为安嫩伯格大学讲席教授。
Phil is the Annenberg University Professor in Psychology at both the Wharton School and the School of Arts and Sciences at the University of Pennsylvania.
我第一次听说菲利普的作品,是在 2005 年与丹尼·卡尼曼会面之后。那次会面结束后,有人问丹尼他在读什么。他说,他很长一段时间里读过的最重要的书之一,是菲利普写的一本叫《专家的政治判断》的书。顺便说一句,对于还没拥有那本书的人来说,那是任何人必读的书。
I first heard of Phil's work following a meeting with Danny Kahneman in 2005. After that meeting, someone asked Danny what he was reading. He said that one of the most important books he'd read in a very long time was a book written by Phil called Expert Political Judgment. By the way, for those who don't own that book it’s an essential read for anybody.
你们中大多数可能熟悉这本书,或者听说过它。这本书表明,某些领域的专家很难做出准确的预测。但菲尔并没有就此止步。他下定决心要找出培养良好判断力的方法,于是“良好判断力项目”[GJP] 便诞生了。由于这个项目目前仍在推进中,我们所有人都将非常荣幸地听到菲尔的一些最新想法。
Now, most of you probably know the book, or know of the book, which shows that experts in certain fields struggle to make accurate predictions. But Phil didn't let it go at that. He became very determined to figure out how to cultivate good judgment, and so the Good Judgment Project [GJP] was born. Since this is a project that's very much a work in progress, we are all going to be very privileged to hear some of Phil's thoughts that are hot off the press.
我再补充一点个人感受:在研究我的著作时,我需要大量阅读关于决策制定的资料。几乎每翻开一处,都会看到菲尔出色的研究。这些年来,他贡献非凡,涵盖了许多我们今天无法一一讨论的课题。我从他身上学到了极多,对此我深怀感激。请大家和我一起欢迎菲尔·泰特洛克教授。
I'll add on a personal note that in doing research on my books I've had to do a lot of reading about decision making. Pretty much everywhere I'd look I’d run into Phil's outstanding research. He’s contributed a great amount over the years, including topics that we won't be able to cover today. I've learned an enormous amount from him for which I'm very grateful. Please join me in welcoming Professor Phil Tetlock.
[applause]
[applause]
菲尔·泰特洛克:谢谢迈克尔,这番介绍太客气了。
Phil Tetlock: Thank you, Michael, for that gracious introduction.
我研究专家判断已有大约 30 年。你可以把我这 30 多年研究专家判断的职业生涯分成两个阶段,第一阶段持续了大约 25 年,第二阶段到现在为止持续了大约 5 年。
I've been studying expert judgment for about 30 years now. You could divide my 30-year plus career on expert judgment into two phases, one of which lasted for about 25 years, the other which has lasted about five years now.
头 25 年,它多少是源于对专家判断的恼火。坦率讲,关注点在于专家判断在很多方面有多糟糕——专家有多不校准,多容易陷入信念固着。一系列错误和偏见,有点卡尼曼(Kahneman)的味道。
For the first 25 years, it was somewhat born of exasperation with expert judgment. Frankly, the focus was on how bad expert judgment is in many ways. How miscalibrated experts are. How prone to belief perseverance experts are. A variety of errors and biases, somewhat in the Kahneman spirit.
过去五年情况有所不同。我们不再那么忙着诅咒黑暗,而是更多地点亮蜡烛。重点转向了如何提高对未来各种可能性的概率判断。
The last five years have been different. We've been focusing less on cursing the darkness and more on lighting candles. The focus has been on how much we can improve probability judgments of possible futures.
这场演讲中,我要传达四个关键信息。第一,那种千篇一律、含糊其辞的预测,会拖慢学习周期。
There are four key messages I'm going to try to convey in this presentation. The first is that business-as-usual vague-verbiage forecasts slow learning cycles.
第二个启示是,你可以通过预测锦标赛来加速学习周期。这种锦标赛要求你对那些平时不会估算概率的话题,明确地给出概率估算。
And the second message is that you can accelerate learning cycles by using forecasting tournaments, which require making explicit odds estimates on topics you don't normally make odds estimates on.
我说的“学习周期”是什么意思呢?意思就是,你会更快地发现自己擅长什么、不擅长什么。你也会更快地了解,你的哪些顾问在特定领域、特定事情上表现更好或更糟。
What do I mean by learning cycles? Well, I mean you'll more quickly learn what you're good and bad at. You'll more quickly learn which of your advisors are better or worse at particular things in particular domains.
证据表明,你在细化运用概率方面也会变得更熟练,能够区分不同程度的不确定性。我稍后会详细解释这到底是什么意思。加速学习周期——我想我们都同意,这听起来是一件相当不错的事。
Evidence suggests that you'll also become more skilled at using probabilities in a granular fashion, in distinguishing many degrees of uncertainty. I will unpack exactly what that means later. Accelerating learning cycles I think we can all agree, sounds like a pretty good thing.
第三个信息:我将转而谈谈我们迄今为止一直在进行的锦标赛中得出的关键教训。
My third message: I will turn to the key lessons from the tournaments we've been running to-date.
菲利普·泰特洛克的“良好判断项目”(续)
Philip Tetlock The Good Judgment Project (Continued)
由于一个格外富有创新精神的政府机构——情报高级研究计划局(Intelligence Advanced Research Projects Activity),也就是国家情报总监办公室下属的研发部门——的推动,
Due to an unusually innovative government agency – the Intelligence Advanced Research Projects Activity, which is the R&D [research and development] branch of the Office of the Director of National Intelligence
– 我们搞了四年的预测锦标赛资金,有时和其他大学竞争,有时跟我们自己的预测市场竞争,有时还和情报机构内部分析师们搞的预测暗中较劲。关于如何培养良好判断力,我们学到了很多,这些经验基本上还热乎着刚出炉。
– we have had funding for four years now for running forecasting tournaments. Sometimes in competition with other universities, sometimes in competition with our own prediction market and sometimes in a shadow competition with predictions emerging from analysts working inside the intelligence community itself. We've learned a lot about how to cultivate good judgment, and it is more or less hot off the press.
我要谈的第四个也是最后一个信息是:我想说说在推动知识进步方面的下一个挑战是什么,以及下一代竞赛应该聚焦于什么,因为我们不仅了解到人们做出明确概率判断时能有多快地学习,也在学习如何设计更好的竞赛,从而以更根本的方式改善组织运作。
Fourth and final message I will be delivering: I want to talk about what the next challenge is in advancing knowledge, and what the next generation of tournaments should focus on, because we are not only learning things about how rapidly people can learn when they make explicit probability judgments, we're also learning something about how to design better tournaments that can improve organizational functioning in even more fundamental ways.
那么,这就是今天的议程安排。
So, that's the agenda.
第一部分,“模糊措辞预示缓慢的学习周期。”我想大家都知道,专家和顾问们往往偏爱模糊措辞。“俄罗斯入侵乌克兰东部是可能的。”“中国的新海上钻井平台可能、或者或许、也许标志着,在南海”,或者东海,或者在别的什么地方。“最近的希腊选举有可能……”你们全都听过这类说法。模糊措辞非常模糊,研究者已经量化了它到底有多模糊。
Now, part one, “vague verbiage forecasts slow learning cycles.” I think we all know that pundits and consultants tend to prefer vague verbiage. “A Russian invasion of the Eastern Ukraine is possible.” “The new Chinese oil rig could, or might, or may signal, in the South China Sea,” or the East China Sea, or wherever it might be. “The recent Greek election has the potential to...” You've all heard things of that sort. Vague verbiage is very vague, and researchers have quantified just exactly how vague it is.
他们估算了一个叫做“量化等价区间”的东西。当你让人们把这些短语翻译成概率时,得到的答案往往天差地别。
They've estimated things called quantitative equivalence ranges. When you ask people to translate some of these phrases into probabilities, you often get an enormous range of answers.
“有可能发生”这一表述的含义跨度从 0.09 到 0.64 左右——这是读者对其的理解方式。“它可能会发生”,范围在 0.02 到 0.56 之间。“存在这种可能……” “确实存在这种可能……” 这里面的跨度可真是够大的。
“It might happen” takes on meanings ranging from 0.09 to about 0.64. That's how readers decode it. “It could happen,” a range from 0.02 to 0.56. “It's a possibility…” “It's a real possibility…” There you've got a real range.
[laughter]
[laughter]
菲尔:“很有可能……”还有“也许”,甚至纯良无害的“或许”。“明显的可能性”——这个词在情报界惹过不少麻烦,无论是猪湾事件还是其他几桩事都如此。“有风险。”“有点机会。”而另外一些短语则几乎没什么回旋余地。
Phil: “It's probable…“Maybe,” even innocent old “maybe.” “Distinct possibility,” that's a phrase that's caused a great deal of mischief in the intelligence community – both the Bay of Pigs and in a couple of other situations. “Risky.” “Some chance.” And then there are a few phrases that there isn't that much wiggle room.
正如中央情报局前局长乔治·特内特所发现的那样,“铁定拿下”这个词,人们确实会将其理解为几乎等同于百分之百的把握。而一旦你判断失误,又表现得如此极度自信,那就会遭受巨大的声誉损失——而且可以说,你活该。
As George Tenet, a former director of the Central Intelligence Agency, discovered, “Slamdunk” people really do interpret it to mean something very close to 100 percent. And when you get it wrong, and you're that extremely confident, you take a huge reputational hit, and arguably you should.
这类含糊其辞的表述,让人根本无法衡量那些知名权威和收费昂贵的咨询机构(你可能还买过其中一些的服务)的真实业绩记录。也无从判断你在《金融时报》《华尔街日报》或任何你喜欢的刊物上读到的意见领袖,其观点到底有多准确。这些人确实都是智商极高之辈。他们能发表精彩绝伦的演讲,比我讲得还好。他们很棒,很有个人魅力,也很有感召力,但根本没有人清楚他们的实际业绩记录如何。他们自己更乐意把记录藏起来。
These properties of vague verbiage make it impossible to gauge the track records of famous pundits and lucrative consulting shops, some of whose services you may have purchased. Makes it impossible to know how accurate are the thought leaders you read in the Financial Times or the Wall Street Journal or whatever your favorite publications might be. These are all really high IQ guys. They can give great talks, they can give a better talk than I can. They're wonderful, they're personable, they're charismatic, but nobody has the faintest idea what their track records are. They prefer to keep their own track records.
含混的措辞还会让人无法学会在特定环境中做出最精微的概率区分。一位著名扑克玩家曾说过,区分真正的严肃扑克玩家和天赋型业余玩家的方法,是严肃玩家能看出 60/40 的机会和 40/60 的机会之间的差别。我们与我们当前研究中一些也属于“超级预测者”的严肃扑克玩家交流,他们说:“不,不,不。那还不够极致。真正的能力,是能分辨 55/45 和 45/55 之间的差异。”
Vague verbiage also makes it impossible to learn to make the subtlest possible probabilistic distinctions within a given environment. A famous poker player once said that the way to tell the difference between a really serious poker player and a talented amateur is the serious players can tell the difference between a 60/40 proposition and a 40/60 proposition. We talked to some other serious poker players who are among the "superforecasters" in our current research, and they say, "No, no, no. That's not aggressive enough. It's the ability to tell the difference between 55/45 and 45/55."
菲利普·泰特洛克《良好判断项目》(续)
Philip Tetlock The Good Judgment Project (Continued)
一个有意思的问题是,概率判断的区分度能有多大。伟大的心理学家阿摩司·特沃斯基在一次演讲中曾开玩笑说,他认为人类内心深处其实只能区分三种概率:“肯定会发生”、“肯定不会发生”和“也许。”
It's an interesting question how differentiated probability judgments can become. The great psychologist Amos Tversky once said facetiously in a talk that he thought that human beings deep down could really only distinguish three levels of probability: “It’s gonna happen,” “it's not gonna happen,” and “maybe.”
在座各位真正关注这类研究的人,当你们看到概率权重函数等概念时,它们表明人们会赋予极端概率非常大的权重。当概率从 0.3 变到 0.7 时,影响并不大。但当概率从 0.1 变到 0,或从 0.9 变到 1,乖乖,这会对决策产生重大影响。在这些预测竞赛中,人们确实会逐渐学习。这是一个缓慢的过程,但他们学会了在不确定性上做出越来越精细的区分评估。在那些平时你根本不会做这种评估的领域里。
Those of you who really follow that kind of work, when you see probability weighting functions and so forth, they imply that people give a lot of weight to probabilities at the extremes. When probabilities move from 0.3 to 0.7, there's not that much impact. But when they move from 0.1 to 0, or from 0.9 to 1, boy that has decision impact. In these forecasting tournaments, people do learn gradually. It's a slow process, but they learn to make increasingly differentiated assessments of uncertainty. In domains where you don't normally have such assessments.
那么,这引出了一个问题。显然,专家和顾问们不喜欢做出精确的预测,因为他们担心声誉受损。这就带来了一个问题:为什么在一个特定社会体系中地位极高的人,会自降身价去参加一个预测锦标赛?在这种锦标赛里,最好的结果大概也就是打个平手。
Now, there's a question that all this raises. Obviously, pundits and consultants don't like to make precise forecasts because they worry about taking a reputational hit. It raises the question why anyone who has very high status within a particular social system would ever lower him or herself to participate in a forecasting tournament in which the best possible outcome is probably breaking even.
想象一下,比如我是情报界内部的一名资深中国分析师,负责向总统每日简报和关于中国的国家情报评估提供中国方面的信息,我就是那个首席专家。
Imagine, for example, that I'm a senior China analyst inside the intelligence community, and I'm the go-to guy for the presidential daily briefing on China and the national intelligence estimates on China.
人们看重我,而现在我看到这些新锐开始在情报界内部打出概率判断的分数。那些新锐对我说:“顺便说一句,我们要你和那些没有接触任何机密情报的超级预测员比一比高下。”
People value me, and now I see these upstarts start scoring probability judgments inside the intelligence community. And the upstarts say to me: "By the way, we want you to compete against these superforecasters who don't have access to any classified information."
我会怎么应对这一切?不需要太多想象力就能猜到,我的反应不会太友好。事实上,在大多数组织内部,对衡量概率判断准确性的想法,都存在很大阻力。职位越高的人,越容易抗拒。这其实在微观经济学上很有道理。
How am I going to react to all of that? It doesn't take a lot of imagination to suppose I'm not going to react very favorably. In fact, there is a lot of pushback within most organizations to the idea of measuring the accuracy of probability judgments. The more senior people are, the more likely they are to be resistant. That actually makes good microeconomic sense.
那么,这一切都表明,你不该指望高高在上的“歌利亚”会帮助地位低下的“大卫”来消灭自己。这种事压根就不该指望。但它确实发生了——就发生在美国政府里。
So, all this implies that you should never expect high status Goliaths to help low-status Davids to kill them. You just shouldn't expect things like that to happen. But it did. It did happen in the U.S. government.
情报高级研究计划署(IARPA)每年的研发经费远超过 50 亿美元,该机构每年拿出约 500 万美元资助一批初创企业,并让它们在公平竞争的环境下比试,看谁能对政策制定者——也就是情报机构的客户——关心的事件给出更准确的概率判断。
The Intelligence Advanced Research Projects Activity [IARPA], which is the R&D branch of a way-over- $5-billion plus per-year bureaucracy, has funded a set of upstarts at about $5 million per year, and has engaged these upstarts, in a level playing field competition to see who could assign better probability estimates to events of interest to policymakers, who of course are the customers of intelligence agencies.
那么,究竟发生了什么?这很了不起。我说的这个项目,按照官僚体系基础法则(Bureaucracy 101)和理性政治参与者的常识来看,根本就不该发生。它之所以成为现实,背后的原因恐怕永远不会有人讲出来。只能这么说:若不是一小群极其敬业、甚至富有远见的公务员——他们的名字我甚至不能提及——这个项目根本不可能实现。
So, what happened? It's remarkable. I'm talking about a project that according to the most basic laws of Bureaucracy 101 with rational political actors should never have occurred. The story of why it occurred is one that will probably never be told. It must suffice to say this project would have been impossible but for the efforts of a small group of very dedicated, even visionary, civil servants, whose names I cannot even mention.
这是第一部分:空洞的言辞让人难以学习,也难以评估那些思想领袖、专家、各种顾问之类的人。
That's part one: vague verbiage makes it hard to learn, makes it hard to evaluate thought leaders, pundits, miscellaneous consultants, and the like.
第二部分是,预测竞赛能够加速学习进程。从这个意义上说,预测竞赛优于常规做法。我将预测竞赛视为一种颠覆性技术,它动摇了陈腐的现状层级——这些层级由那些缺乏远见却精于将“未曾料之事”转化为“事后看来必然之事”的巧言者主导。
Part two is that forecasting tournaments can accelerate learning. Forecasting tournaments are, in this sense, superior to business as usual. I see forecasting tournaments as a disruptive technology that destabilizes stale status hierarchies dominated by smooth talkers who lack foresight but who are adept at transforming the unforeseen into the retrospectively inevitable.
[laughter]
[laughter]
菲尔:这句话出自我们提交给 IARPA 的一份年报。但你能理解为什么那个假设性的中国高级分析师不会喜欢这样的表述——它有点太直白了。
Phil: That comes from one of our annual reports to IARPA. But you can see why the hypothetical senior China analyst might not like a line like that. It's a bit in your face.
菲利普·泰特洛克的“良好判断项目”(续)
Philip Tetlock The Good Judgment Project (Continued)
IARPA 锦标赛是什么样的?最初,有五个大学研究项目赢得了参赛资格,进入了这项竞争。加州大学伯克利分校……我曾在伯克利做了很多很多年的教授,最近刚搬到宾夕法尼亚大学的沃顿商学院……有一支伯克利/宾大、宾大/伯克利联合团队,有一支麻省理工学院团队,有一支密歇根大学团队,等等。我们彼此竞争,看谁能做出更好的概率估计。
The IARPA tournaments, what do they look like? In the beginning, there were five university-based research programs which won the competition to enter into the competition. The Berkeley…I was a professor for many, many years at Berkeley, and then I moved recently to Wharton at Penn…There was a Berkeley/Penn, a Penn/Berkeley team, there was an MIT team, a Michigan team, and so forth. We were competing with each other to generate better probability estimates.
比赛进行的每一天,我们的统计师和预测人员都会生成预测流,这些预测流会在美国东部时间上午 9 点提交给美国政府(及其咨询代理机构),所用的方法完全由学者们自行决定,他们认为哪种方法最有效就用哪种。如果你信占星术,你大可用占星术;你想用什么都可以。显然我们不会用那个。但这个机制的设计初衷就是为了激励方法上的创新。我认为它确实做到了这一点。
Each day the tournament was alive, our statisticians and our forecasters would generate forecast streams which would be submitted at 9 AM Eastern Time to the U.S. Government (and its consultancy proxies) using whatever methods the academics thought would work best. If you believed in astrology you could use astrology. You could use anything you wanted. Obviously we weren't using that. But it was designed to incentivize methodological innovation. I think it did exactly that.
过去三年里,IARPA 已经提出了近 400 个问题。其中一些确实直接落在金融领域,但大部分与国家安全更为核心相关。比如这些问题:朝鲜会在“某个时间点”之前试射核武器吗?马里奥·蒙蒂会在 X 日期之前辞去总理职务吗?在 Y 日期之前,克里米亚会发生致命冲突吗?在东海会发生致命的中日冲突吗?某个国家的主权债务会被降级吗?所以,问题范围非常广泛。
IARPA, over the last three years, has now posed almost 400 questions. Some of them actually do fall quite squarely in the financial domain, but most of them are more centrally relevant to national security. Questions like, Will North Korea test a nuclear weapon by “blup”? Will Mario Monti vacate the office of prime minister by date X? Will there be a lethal confrontation in the Crimea before date Y? Will there be a lethal Sino-Japanese confrontation in the East China Sea? Will there be a sovereign debt downgrade of country X? So, a wide range of questions.
运营一场预测锦标赛,需要先扫清一些科学上的障碍。要让比赛顺利进行,我们必须解决三个大问题。第一,我们需要证明我们能够可靠地衡量人类概率判断的准确性。第二,同样地,我们需要决定用什么样的基准来衡量人类才是合理的。我们应该拿他们和时间序列模型比吗?应该拿他们和彭博共识小组比吗?基准究竟应该是什么?我们应该只拿他们和扔飞镖的黑猩猩比吗?
Running a tournament requires clearing a certain amount of scientific underbrush. There were three big questions we had to clear away to make the tournament work. One is we needed to demonstrate we could reliably measure the accuracy of human probability judgments. Second, again, we need to decide against which benchmarks is it reasonable to measure human beings. Should we be measuring them against time series models? Should we be measuring them against Bloomberg consensus panels? Exactly what should the benchmarks be? Should we be measuring them just against the dart-throwing chimpanzee?
菲尔:第三个问题是,当我们未达到基准时,我们可以做些什么来增强个人、团队和组织的判断力。
Phil: The third question is, when we fall short of benchmarks, what can we do to augment the judgment skills of both individuals, teams, and organizations.
衡量概率判断的准确性。这是一项最初由气象学家和统计学家开发的技术,非常直接明了。我们还可以讨论许多其他衡量概率判断的技术。
Measuring the accuracy of probability judgments. This was a technique that was originally developed by meteorologists and statisticians. It is a very straightforward technique. There are many other techniques we could talk about for measuring probability judgments.
核心思路是这样的:假设有一位气象学家在预测降雨概率。他给出的数据是:这一处 90%,这一处 50%,这一处 50%,这一处 80%。实际结果则是:下雨、下雨、没下雨、下雨。
The basic idea is you've got a meteorologist who is making predictions, the probability of rain. He's at 90 percent here, 50 percent here, 50, 80. There's an outcome. It rains, rains, doesn't rain, rains.
然后你将现实编码为 0(如果没下雨)或 1(如果下雨),并使用一个简单的二次评分规则。你将概率相对于 1 或 0 进行偏离。
Then you code reality as either zero, if it doesn't rain, or as one, if it does rain. And you use a simple quadratic scoring rule. You deviate the probabilities against one or zero.
举例来说,如果你说有 90% 的概率会下雨,然后雨确实下了,你的布莱尔评分(Brier score)就会很漂亮。布莱尔评分越低越好,就像高尔夫一样,0.02 是极好的成绩。而在这里,你和那只扔飞镖的黑猩猩表现差不多,刚好落在不确定性最大的位置附近。
If you say, for example, there’s 90 percent probability of rain and it rains, you get a great Brier score. Low Brier score, it's like golf, low scores are great, 0.02. Here, you're with the dart-throwing chimpanzee, right at maybe maximum uncertainty.
在这里,同样的情况。你说有 80% 的把握会出事,结果 Brier 评分依然很不错。
Here, again, same thing. Here, you say 80 percent yes, and you get again a pretty good Brier score.
把这一切算个平均值,大约是 0.28。
You average it all out. It comes to about to 0.28.
当你对大量问题重复这个过程,你就能得到相当稳定的判断能力统计估值,这些估值与多种在心理学和组织学上具有意义的事物具有相关性。
When you do this over a large numbers of questions, you can get fairly stable statistical estimates of judgment skill that are correlated with a variety of things that are meaningful, psychologically and organizationally.
如果你有一套完全准确的确定性系统理论,你的布赖尔得分可以趋近于零。
If you had a perfectly accurate theory of a deterministic system, your Brier score could approach zero.
如果你只是胡乱猜测,你的布莱尔评分(Brier score)就会在 0.5 附近徘徊。而如果你是一个逆向
If you were just guessing, your Brier score would hover in the vicinity of 0.5. And if you were an inverse
菲利普·泰特洛克 良好判断项目(续)
Philip Tetlock The Good Judgment Project (Continued)
有预知能力,而你预测的所有事情全都没发生,那么你的 Brier 评分就会高得离谱,达到 2.0。
clairvoyant, and the opposite of everything you said happened, then you would have a ridiculously high Brier score of 2.0.
你可以将这些布里尔得分分解为两个关键指标:校准度和分辨度。
You can break these Brier scores down into two key metrics: calibration and resolution.
校准(calibration)是指你为事件赋予概率的能力,且这些概率长期来看要与这些事件发生的客观频率保持一致。
Calibration is your ability to assign probabilities that over the long term correspond to the objective frequencies with which those events occur.
对于那些你赋予 70% 概率的事件,它们往往大约有 70% 的时间会发生。
For those events you assign 70 percent probability to, they tend to occur about 70 percent of the time.
对于那些你赋以 30% 概率的事件,它们实际发生的概率也大约是 30%。这叫作校准。这是对你自身知识局限性的细致感知。
For those events you assign 30 percent probability to, they occur about 30 percent of the time. That's calibration. That's a nuanced sense for the limitations of your knowledge.
但是,你想要的不仅仅是校准,还需要分辨力。你希望赋予那些真正发生的事情比没发生的事情高得多的概率。我来解释一下这是什么意思。
But, you want more than just calibration. You also want resolution. You want the ability to assign much higher probabilities to things that occur than to things that don't occur. I'll show you what I mean by that.
这是一个预测校准极佳的预测者案例——因为当这位预测者说某事有 40% 的发生概率时,该事确实在 40% 的情况下发生了;当他说有 60% 的概率时,结果 60% 应验;当他说 50% 的概率时,结果正好一半对一半。
This is an example of a forecaster who has really good calibration, because when the forecaster says there's a 40 percent likelihood of things happening, things happen 40 percent of the time. When the forecaster says there's a 60 percent chance of things happening, things happen 60 percent of the time. 50 percent likelihood, things happen 50 percent of the time.
唯一的问题是,这位预测者从来说不出什么特别不同的东西,无非是各种微妙程度的“可能”。你的预测者从来不会跳出 0.4 到 0.6 这个区间。
The only problem is this forecaster never says anything very different from minor shades of “maybe." Your forecaster never moves outside the bracketed range between 0.4 and 0.6.
这是一个预测者的例子,他拥有出色的校准能力和很好的区分能力。可以看到,所有点都沿着那条对角线分布。
This is an example of a forecaster who has wonderful calibration and very good discrimination. You can see all of the points are along the diagonal there.
最后,完美的境界是这样:全知全能,或者说上帝。这样的人,问题刚一提出,就能立即告诉你零或一——“这事不会发生,这事会发生”——而且准确无误,从不犯错。不用说,从来没有人能接近这个水平,连边都沾不上。
And, finally, this is what perfection would look like. This is omniscience or God. This would be someone, who as soon as the question is asked, can either tell you zero or one. "It's not going to occur, it is going to occur," and do it with infallible accuracy. Needless to say, nobody ever even remotely approximates this.
布赖尔评分表的性能范围是多少?零分代表全知全能,0.5 分代表黑猩猩水平,两者之间有大量变化区间。
What's the performance range on the Brier scale? It's zero omniscience, 0.5 the chimpanzee, and lots of variation in between.
我们应该用什么基准来评判人类的表现?这是在比赛前我们必须解决的另一个问题。
What kinds of benchmarks should we use for judging the performance of humans? That was another question we had to wrestle with before the tournaments.
当时存在一些极为简陋的基准,比如投掷飞镖的黑猩猩,以及时间序列模型中的简单外推法——“维持现状”。你能打败一个“维持现状”的预测算法吗?
There were minimalist benchmarks, like the dart-throwing chimpanzee, simple extrapolation in time series models. More of the same. Can you out-predict a "more of the same" algorithm.
有一些中等激进的绩效评估基准,比如未经加权的人群智慧均值或中位数,这些基准极难被超越。
There are moderately aggressive benchmarks for assessing performance, like the unweighted mean or median of the wisdom of the crowd which is quite difficult to beat.
顺便提一句,这是我们最初在 IARPA 竞赛中决定使用的衡量标准。然后是专家共识小组,你们都很熟悉这些机构,比如费城联储、经济学人智库或彭博。
This was, by the way, the metric we decided to use initially in the IARPA tournament. And then expert consensus panels, which you guys are familiar with, like the Philadelphia Fed or the Economist Intelligence Unit, or Bloomberg.
最后来看最高标准的基准:最先进的统计技术和大数据模型,最终还要打败一个——我想在座各位看来——终极的试金石,那就是战胜深度、高流动性的市场。
Finally, maximalist benchmarks: the most advanced statistical techniques and Big Data models, and beating, finally, the ultimate litmus test, I guess, for people in this room, is beating deep, liquid markets.
我们最好的预测者和最好的算法,在各种基准面前表现得有多好?这个问题没有一个放之四海而皆准的答案。它取决于我们面对的是哪种类型的问题。
How well can our best forecasters and our best algorithms do against various benchmarks? There's no across-the-board answer to that question. It depends on the type of problem with which we're dealing.
用“可预测性连续谱”来思考这个问题很有帮助。在最具可预测性的一端,是傅科摆。如果你懂牛顿物理学,又对地球了解一点,你就能以几乎完美的准确率预测傅科摆的运动。你的布里尔分数将是零。
It's useful to think of a predictability continuum. At the most predictable end, Foucault's pendulum. If you know Newtonian physics and you know a little bit about the planet earth, you're going to be able to predict, with virtually perfect accuracy, the motion of Foucault's pendulum. Your Brier score will be zero.
菲利普·泰特洛克 良好判断项目(续)
Philip Tetlock The Good Judgment Project (Continued)
在另一端,轮盘赌桌上不存在专家。无论你训练多久都无济于事。
At the other end of the continuum, there are no experts at the roulette wheel. It doesn't matter how much you train, it's not gonna help.
最后还有一个饶有趣味的案例:预测飓风的轨迹。没有人能准确预知飓风在何时何地形成。蝴蝶翅膀之类的东西会介入其中,制造混沌和湍流。但一旦飓风发展到一定成熟阶段,就变得可以预测了。
Finally, there's the interesting case of predicting the trajectories of hurricanes. Nobody can predict when and where exactly a hurricane will form. There are butterfly wings and all that kind of stuff intruding and producing chaos and turbulence. But once a hurricane has reached a certain level of maturity, it is possible to predict.
四十年前,气象预报员在这方面并不太擅长。如今他们已经相当拿手了。这是一个成功的故事。气象预报员之所以变得更出色,部分原因是背后的科学进步了,但也因为气象预报员养成了定期给自己的校准度和分辨度——布赖尔评分(Brier score)的组成部分——打分的习惯。他们得到反馈,从而变得越来越强。
Forty years ago, meteorologists weren't very good at it. Now they are pretty good at it. It's a success story. Meteorologists have become better at it, partly because the underlying science has become better, but also because meteorologists are in the regular habit of scoring their calibration and resolution – the components of the Brier score. They’re getting feedback and getting better at it.
气象学家之所以在关于专业判断力的文献研究中,是被拿来作为校准度最高的一类职业群体,这背后是有原因的。他们和扑克高手、桥牌专家属于同一梯队。这是一个人数稀少、极为精尖的群体,这些人确实校准得相当出色。
There's a reason why meteorologists are among the best-calibrated professional groups studied in the literature on expertise. They're up there with expert poker players, expert bridge players. It's a small, rarified crowd of people who are really super well-calibrated.
我们能将预见能力提升多少?看看 IARPA 预测竞赛的结果,相比对照组中未加权的群体平均水平,我们的表现要高出 50% 到 70%。
How much can we improve foresight? When you look at the results of the IARPA tournament, we were able to do about 50 to 70 percent better than the unweighted average of the crowd in a control group.
关于 IARPA 竞赛,你们应该了解的一点是,我们会做实验。预测者被随机分配到不同的实验条件中,对照组也是随机分配的。我们计算对照组的平均预测值,那就是我们必须超越的基准线。
One of the things you should appreciate about what goes on in the IARPA tournament is that we run experiments. Forecasters are randomly assigned to experimental conditions. There's a random assignment to a control group. We compute the average forecast or prediction in that control group. That's the benchmark we have to beat.
如果我们宣称自己有一套能提升概率判断能力的训练体系,那它就必须比对照组做得好,好得超过那个基准。这有点像美国食品药品监督管理局(FDA)的药物试验——你必须打败对照组。
If we're going to say that we have a training system that improves probability judgment, it has to improve probability judgment above and beyond that control group. It's rather like an FDA [Food and Drug Administration] drug trial, you've got to beat the control group.
我们很高兴地看到,如今在三年时间里,通过随机对照试验,我们能够复制这一表现——在这些试验中,数千名预测者对数百个问题做出了预测。
We were pleased that we were able to replicate that performance now in three years, in randomized trials in which thousands of forecasters make predictions on hundreds of questions.
我们拥有大量数据——在第二年之后,我们的胜率足以将其他学术竞争者挤出这场竞争,现在让我们与自己的预测市场(由 Inkling 运营)直接竞争,并与情报界内部产生的预测形成间接竞争。
We have a lot of data—and our margin of victory was sufficient to knock the other academic competitors out of the competition after the second year, and it puts us in direct competition now with our own prediction market, run by inkling, and in indirect competition with predictions emerging from the intelligence community itself.
我们现在知道,我们甚至能比掌握机密信息的专家做得更好,因为戴维·伊格纳修斯在 2013 年《华盛顿邮报》的一篇文章中透露,据他的消息来源称,我们最好的预测者和最好的算法,其预测结果比情报机构内部生成的预测高出约 30%。
We now know that we were even were able to do better than experts with access to classified information because David Ignatius wrote a story in the Washington Post in 2013 in which he revealed that, according to his sources, our best forecasters and best algorithms were able to do about 30 percent better than the forecasts being generated within the intelligence community.
我觉得那相当了不起。NPR(美国国家公共电台)为此做了一篇报道,标题极尽俗套,叫“你比 CIA 特工聪明吗?”——这倒引来了一大波新志愿者。报道里重点介绍了一位我们最厉害的预测者,特拉华州的一名药剂师,他一直在追踪叙利亚难民的流动。这个故事迅速传开了。
I think that's pretty amazing. NPR [National Public Radio] ran a story on it, with a terribly tacky title “Are you smarter than a CIA agent?”—and that triggered a veritable avalanche of new volunteers. They featured one of our very best forecasters, a pharmacist in Delaware, who was tracking Syrian refugee flows. The story went viral.
从我的角度来看,我们在过去三年中学到的是,卓越表现有四大驱动因素。由于采用了实验性质的设计,我们对这一结论的信心非同寻常。这些都是严谨的、有对照的实验,而不仅仅是相关性研究设计。
From my point of view, what we've learned over the last three years is there are four categories of drivers of superior performance. We know this with unusual confidence because of the experimental nature of the design. These are rigorous, controlled experiments. These aren't just correlational designs.
菲利普·泰特洛克“良好判断项目”(续)
Philip Tetlock The Good Judgment Project (Continued)
我们知道,让合适的人上车至关重要,因为我们可以把我们的对照组与其他对照组进行比较。我们确实在招揽更聪明、思维更开放的预测者方面取得了一点成功。我们有了更好、更优质的人力素材。
We know it's important to get the right people on the bus, because we can compare our control group to other control groups. We know we had a little bit of success in getting smarter, more open-minded forecasters. We had better, raw human material.
我们从实验角度也了解互动的好处,因为我们可以将其与对照组进行对比评估。
We know about the benefits of interaction, experimentally again, because we can assess it against a control group.
我们的预测准确率会提高 10% 到 20%。预测者在以特定方式组建的团队中合作,或在预测市场中竞争时,表现会更好——这些方式我们稍后可以详谈。
We get about a 10 to 20 percent boost. Forecasters do better when they're working either collaboratively in teams that are structured in certain ways we could talk about, or competitively in prediction markets.
我们设计出了一些培训模块,能够将判断力提升大约 10%。这些是认知去偏误训练。
We've been able to design training modules that improve judgment in the range of 10 percent. These are cognitive de-biasing exercises.
最后,我们的统计学家在开发算法方面做得很出色,这些算法采用加权平均,给预测表现更好的人更高权重,然后再做“极端化处理”。我稍后解释极端化处理的逻辑。
Finally, our statisticians have done a wonderful job in developing algorithms, weighted averaging algorithms that give more weight to our better forecasters, and then "extremize." I'll explain the logic of extremizing in a minute.
转向算法层面,群体智慧的经典概括是:平均预测往往比构成该平均值的多数个体预测更为准确。
Turning to the algorithms, the classic wisdom of the crowd generalization is that the average forecast tends to be more accurate than most of the individuals from whom the average was derived.
当然,那就是由詹姆斯·苏罗维茨基推广的著名高尔顿公牛故事,其中蕴含着不少真理。我们反复以各种形式复制着这一发现。
That's, of course, the famous Galton ox story that was popularized by James Surowiecki, and there's a lot of truth to that. We replicate that finding over and over in various forms.
优良判断项目不使用等权平均法。我们的统计学家发现,采用加权平均法效果要好得多。我们对智商更高、思想更开放、预测历史记录更佳的预测者赋予更大权重。此外,我们加入了一项颇具争议的元素,称之为“极端化参数”。我来解释一下这是什么意思。
The Good Judgment Project does not use unweighted averaging. Our statisticians have found it's much more effective to use weighted averaging. We give more weight to higher IQ, more open-minded forecasters who have better track records. Plus we add something that is somewhat controversial, we called an "extremizing parameter." I'm going to explain what that is.
转换参数的作用是:在设定值为 0.5 时,它等同于实际表达出的概率,但它会进行极端化处理,因此将 0.6 转化为约 0.72,将 0.4 转化为约 0.27。它进行极端化处理。
What the transformation parameter does is, at 0.5, it's equal to the actual expressed probabilities, but it extremizes, so it turns 0.6 into about 0.72, and 0.4 into about 0.27. It extremizes.
从数学上看是这样的:这是一个带有收缩与噪声的 log-odds(对数胜率)。这里有一个关键参数——参数 a,它决定了转换的幅度。
It looks like this, mathematically. It's a log-odds with shrinkage plus noise. There's a key parameter here. It's a, parameter a. That determines the amount of the transformation.
根据我们的统计学家的说法——我认为他们为此提出了一个非常有说服力的论证——转化程度取决于预测者群体的精密度和多样性。预测者群体越是多样,你就能越是极端化。
The amount of the transformation, according to our statisticians – I think they've made a very compelling case for this – depends on the sophistication and diversity of the forecaster pool. The more diverse the forecasting pool, the more extremely you can extremize.
我来给你举个思想实验的例子。你们很多人可能看过电影《猎杀本·拉登》,已故演员詹姆斯·甘多菲尼在片中饰演了时任中情局局长莱昂·帕内塔,那正是追捕奥萨马·本·拉登的行动期间。
I'll give you a thought-experiment example. Many of you have probably seen the movie Zero Dark Thirty, in which James Gandolfini, the late actor, played Leon Panetta when he was Director of the CIA [Central Intelligence Agency] at the time of the operation to target Osama Bin Laden.
情报界内部曾有过一场争论:奥萨马·本·拉登是否住在巴基斯坦小镇阿伯塔巴德的一个特定院落里。在电影中……好莱坞总能同时把事情拍得极其准确和极其离谱,他们能拍出这种效果,真有意思,但这就是艺术,这是艺术与科学中属于艺术的那部分。
There was a debate inside the intelligence community about whether Osama Bin Laden was living in a particular compound in the Pakistani town of Abbottabad. In the movie…and Hollywood manages to get things deeply right and deeply wrong, simultaneously, it’s interesting how they manage to do that, but that's art, that's the art part of the art and science…
[laughter]
[laughter]
菲尔:他们围坐在桌边,局长问:“你们对这件事的最佳概率估计是多少?”每位顾问都给出了自己的估计。假设——电影并不是这样展开的——但假设帕内塔的顾问每个人都告诉他:“0.7,0.7,0.7。”听起来答案就是 0.7。
Phil: They're around the table and the director says, "What's your best probability estimate of this?" Each of the advisors gives an estimate. Just assume – the movie doesn't unfold like this – but assume that Panetta's advisors, each of them says to him, "0.7, 0.7, 0.7." Sounds like the answer is 0.7.
如果这些顾问彼此是克隆人,那么答案将是 0.7。
If the advisors were clones of each other, the answer would be 0.7.
菲利普·泰特洛克“精准预测项目”(续)
Philip Tetlock The Good Judgment Project (Continued)
但若其中一位顾问依赖人类智能,另一位依赖卫星情报,第三位依赖某种其他形式的情报——也就是说,他们各自依靠截然不同的智能来源,且每一位都得出了 0.7 的结论——那么正确答案应该是多少?
But if one of the advisors is drawing on human intelligence, another on satellite intelligence, another on some other form of intelligence, if they're drawing on very different forms of intelligence, and each of them is reaching a 0.7 conclusion, what's the right answer?
我们并不确切知道。这个问题没有数学上的答案,但你现在有相当强烈的直觉,它大于 0.7。
We don't know exactly. There's no mathematical answer to that, but you have a pretty strong intuition now that it's more than 0.7.
这就是极端化参数所捕捉的直觉。
That's the intuition the extremizing parameter captures.
这门技艺与学问的关键在于:这种极致化的参数应该是什么样的?
The art and science, what should this extremizing parameter look like?
你可以设定一个极端的参数,一旦上升到 51%,就直接变成 1;一旦下降到 49%,就直接归零。这种做法在某种程度上确实有效,但不如这个好。更温和的极端化参数才是更好的选择。
You could create an extremizing parameter that's so extreme that as soon as you go up to 51 percent, you go to one. As soon as you go to 49 percent, you go down to zero. That actually works to some degree, but not as well as this. The more moderate extremizing parameter is a better way to go.
让我简单说几句。我提到过超级预测者这件事。在 IARPA 竞赛中,我们从所有预测者中选出最顶尖的 2%,把他们编入精英团队。这些团队每组 12 人,彼此之间协作配合。
Let me say a few words. I mentioned superforecasters. In the IARPA tournament, we take the top two percent of our forecasters and we assign them to elite teams. They work with each other in teams of 12.
去年,我们有大约 10 支团队,每队 12 名超级预测者协同工作。他们的表现极为出色。
Last year, we had about 10 teams of 12 superforecasters working together. They have performed phenomenally.
超级预测者是唯一一种在生成概率方面接近击败我刚才描述的对数赔率转换的方法。
The superforecasters are the only method of generating probabilities that comes close to beating the log-odds transformation I just described to you.
人类无法战胜这种变革。但超级富豪们确实在竭力追赶,我认为这相当了不起。
Human beings cannot beat this transformation. But the supers give it a run for its money, which I think is quite remarkable.
超级预测者还优于 IARPA 竞赛中所有未经转化的数据源。他们优于未加权的群体智慧平均值,优于普通团队,也优于预测市场。他们甚至优于顶尖团队的聚合体。他们是一股自然之力。
The supers are also better than all untransformed sources of data in the IARPA tournament. They're better than the unweighted average of the wisdom of the crowd, better than average teams, and better than the prediction markets. They're better than the aggregated top teams. They are a force of nature.
当然,你们都在问的问题是:“它们什么时候会跌回地面?均值回归什么时候会起作用?” 我不知道这个问题的答案,但我认为确实会出现一定程度的回归。
Of course, the question you're all asking is "When are they going to fall down to the ground? When is regression toward the mean going to kick in?" I don't know the answer to that question, but I think there is going to be some regression.
个体的超级预测者确实在一定程度上会向均值回归。不知为何,超级预测者团队至今仍能抵御这种趋势。但我完全清楚,在 IARPA 竞赛中,我们试图预测的事件存在很大的随机成分,而且向均值回归在此处确实是一个真实存在的过程。
Individual supers do regress to some degree toward the mean. Somehow the superforecaster teams have been able to resist that thus far. But I'm fully aware that there is a large chance component to the events that we're trying to predict in the IARPA tournament, and that regression toward the mean is a real process here.
但其中也有技巧的成分。正如迈克尔[莫布森]近来的著作所强调的那样,你应该预期回归均值发生的幅度,取决于你对运气与技巧重要程度比率的直观判断。
But there is a skill component as well. And as Michael [Mauboussin’s] recent book underscores, the amount of regression toward the mean you should expect is going to be a function of your intuition about the ratio of importance of luck and skill.
你越是觉得这纯属运气,就越应该预期它会很快向均值回归。
The more you think it's just luck, the more you should expect very rapid regression toward the mean.
你越觉得这是技术活——能指望的回报就越少——而眼下,我们发现了大量“技术活”。
The more you think it's skill – the less you should expect—and right now, we're finding a substantial amount of skill.
再举一个例子,说明超级观察者(supers)有多厉害。这类曲线在工程学中被称为接收者操作特征曲线(receiver operating characteristic curves)。它衡量的是观察者以较低误报成本获得极高命中率的能力。
Here's another illustration of how super the supers are. These are called, in engineering, receiver operating characteristic curves. They refer to the ability of a perceiver or observer to achieve a really high hit rate at a low cost in false positives.
再说一遍,你越是像神灵一样,就越有能力预测一切经济衰退,准确无误,一次失策都没有,而且不会付出任何误报的代价。这将是完美的 ROC(接收者操作特征曲线),全知全能。
Again, the more godlike you are, the more capable you are of predicting all economic recessions, up to one, up to perfection, at no cost in false positives. It would be a perfect ROC, omniscient.
菲利普·泰特洛克 良好判断项目(续篇)
Philip Tetlock The Good Judgment Project (Continued)
这些超级系统远远优于顶级团队、顶级个人和其他组合。它们在 ROC 曲线上的表现之优秀,令人叹为观止。
The supers are way, way better than top teams, top individuals, and others. It's quite remarkable how well they manage to do on the ROC curves.
锦标赛的一个重要启示:预测是一项植根于认知风格与能力的技能。这对人员选拔具有启示意义。你可以通过某些测试来挑选人才。
One of the big key takeaway lessons from tournaments: forecasting is a skill grounded in cognitive styles and abilities. This has implications for personnel selection. There are tests you can use to select people.
预测是一项可以在你绝大多数员工中培养的技能。它可以边做边学,只需参与竞赛,逐渐细化,在反复练习中学会区分 40/60 概率与 60/40 概率的情形。它还可以通过认知去偏倚练习来实现。
Forecasting is a skill that can be cultivated across a broad swath of your workforce. It can be done learning by doing, simply participating in tournaments, becoming increasingly granular, learning how to distinguish 40/60 from 60/40 propositions with repeated practice. It can be done through cognitive de-biasing exercises.
可以通过将预测者分配到具有学习文化的团队来实现这一点。关于如何具体组织团队,使其像在 IARPA 竞赛中那样高效运作,还有另一场讨论可以展开。
It can be done by assigning forecasters to teams that have learning cultures. There's another talk that could be had about how exactly you organize teams to make them work as effectively as they have in the IARPA tournament.
最后,也是非常重要的一点,精心设计的算法能够提炼群体的智慧,从而大幅提高准确率。从统计上看,这是我们带来的最大效果。
Finally, and very importantly, well-designed algorithms for distilling the wisdom of the crowd can produce substantial boosts in accuracy. They are, statistically, our biggest effect.
但锦标赛无法创造奇迹。这你们心里都清楚。你们也许已经在想,“这好得让人不敢相信。”
But tournaments cannot deliver miracles. You all know that. You might be thinking already, "This is too good to be true."
我不是想自称当代诺查丹玛斯。我认为我们得避开那种寻找预测超级明星的谬误。我觉得预测产品消费者犯的最大错误,就是寻找所谓的远见卓识者。
I'm not claiming to be a modern day Nostradamus. I think we want to steer clear of the looking-for-forecasting-superstars fallacy. I think the biggest mistake that consumers of forecasting products make is looking for visionaries.
如果有人声称能在地缘政治、地缘经济问题上达到 90% 的命中率,你该问的正确问题是:“你的误报率是多少?谁能独立验证这些判断?”我认为这些都是非常合理的问题。
If someone promises you a 90 percent hit rate on geopolitical, geoeconomic questions, the right question to ask is, "What's your false positive rate and who can independently verify the claims?" I think those are very reasonable questions to ask.
在一个充满随机性的世界里,高命中率总是以某种误报为代价。这对我们的超级天才如此,对任何其他人也同样如此。这就像那个老笑话:经济学家预测了过去五次经济衰退中的十一次,或者诸如此类的话。
In a stochastic world, high hit rates always come at some price in false positives. That's true for our supers, that's going to be true for anybody else. That's like the old joke about economists having predicted 11 of the last 5 recessions or whatever.
我们能改善你的眼力吗?这与在座各位的关联性又有多大?细节里头,可是藏着魔鬼呢。
Can we improve your vision? How relevant is this to the people in this room? There, the devil is going to lurk in the details.
第一个问题,你们已经离最优预测前沿有多近?
First question, how close are you already to the optimal forecasting frontier?
如果你确信自己已经接近最优边界,那确实没有太多理由再去努力提高远见。从成本收益的角度来看,这么做几乎没什么回报。
If you're convinced that you're already close to the optimal frontier, there's really not much reason for you to be interested in trying to improve foresight. The cost benefit equation, there just isn't much payoff.
你可能会以为自己已经处在了预测能力的最高边界上,尽管你只是比丢飞镖的黑猩猩稍好一点,但那是因为你工作的环境极端特殊——在这样的环境里,没有人能比丢飞镖的黑猩猩做得好多少。
You might think that you're at the optimal forecasting frontier even though you're only doing slightly better than the dart-throwing chimpanzee, but that's because you're working in such an extremely different environment that nobody can do much better than a dart-throwing chimpanzee.
你可能会觉得已经接近最优预测边界了,因为你的表现远高于那条线。但问题是,除非你举办预测竞赛、做实验、评估是否有可能再提升一点准确度,否则你根本不知道最优预测边界到底在哪里。若不这么做,你就纯粹是在凭信心行事。
You might think that you're close to the optimal frontier because you're considerably above that. You don't know where the optimal forecasting frontier is however, until you run forecasting tournaments and run experiments and assess whether or not it is possible to get some incremental accuracy. Otherwise, you're simply taking it on faith.
这还取决于你的预测者是否具备正确的心态,以及是否获得了适当类型的团队支持。
It also hinges on whether your forecasters have the right mindset and the right types of team support.
两者都至关重要。
Both are crucial.
菲利普·泰特洛克 优质判断项目(续)
Philip Tetlock The Good Judgment Project (Continued)
正确的思维应该是怎样的?如果你作为一名预测者,认同运气起主导作用的理论,或者认同先天能力起主导作用的理论——就像“g 因子”,它深植于你的基因里,是一般智力——如果你认同其中任何一种理论,那么你就几乎没有动力去投入资源提升远见能力。
What is the right mindset? If you as a forecaster subscribe to the theory that it's predominantly luck or subscribe to the theory that it's predominantly innate ability – it’s like g, it's hard wired into your genes, general intelligence – if you subscribe to either of those theories, there's very little incentive for you to invest in improving foresight.
能够进步的人,是那些相信进步是可能的人。这里面有一种自我实现的预言成分。这确实存在,而且值得深思。
The people who improve are the people who believe that it's possible to improve. There is an element of self-confirming prophecy here. It's real and it's worth thinking about.
团队支持也是一个关键因素,因为这是一项艰苦的工作,得到队友的支持非常有价值。
The team support is also a crucial factor because this is hard work and having support from teammates is valuable.
另一个因素是,管理层在评估员工时如何平衡过程问责与结果问责。这是管理学文献中的一个重大议题。
Another factor is how does management balance process versus outcome accountability in evaluating employees. This is a big issue in management literature.
在情报界有一条铁律:如果分析师——尤其是基层分析师——遵循了正确的流程,即便他们犯下错误、判断失误,也不应追究其责任。
In the intelligence community, it is gospel that you don't hold the analyst, the lower level analyst accountable for errors, getting it wrong, if they went through the right processes.
如果他们在撰写报告时流程评分很高,但结论错了,他们不会受到惩罚。中情局局长可能会被解雇,但低层分析师在流程问责制下是受保护的。他们认为这很公平。
If they have a good process score in constructing their reports and they're wrong, they don't get nailed. The Director of the CIA might get fired, but the lower level analysts are insulated in a process accountability system. They think it's fair.
这就像员工与组织之间的一份契约。你遵循这些最佳流程实践,我们就为你提供保障,使你免受环境随机性的影响。这就是他们想出的解决方案。
It's like a contract between the employees and the organization. You follow these best process practices and we will indemnify you against the stochastic nature of the environment. That's the solution they have arrived at.
不幸的是,这种解决方案可能僵化为官僚式的形式主义和走过场——分析师能猜到经理们想听什么,经理们又希望尽快把事情办完,流程问责制会迅速退化。
Unfortunately, that solution can ossify into bureaucratic ritualism and box checking in which analysts can anticipate what managers want to hear and the managers want to get things done fast, and the process accountability can degenerate quite rapidly.
我们的研究表明,要让员工发挥最佳表现,最好的办法是尝试混合流程结果体系——既要让员工对最佳流程负责,也要让他们对创新、寻找超出组织现有流程建议的解决方案负责。究竟该如何平衡这种关系,我称之为一个棘手的取舍。这确实很难。
Our research suggests that the best way to get the best out of your employees is to experiment with hybrid process outcome systems where you want to hold people accountable for best processes, but you also want to hold them accountable for innovating and finding solutions beyond the existing process recommendations of the organization. Exactly how you balance that, I call it a nasty trade-off. It's a tough one.
此外,你们的统计专家是否懂得如何设计算法,并使其适应预测团队的专业水平和多样性。对于高能力、高多样性的团队,你们有充分理由采用极端化处理。
Then, also whether your statisticians know how to design algorithms and tailor them to the sophistication and diversity of your forecasting pool. High-ability, high-diversity pools, you have some substantial warrant for extremizing.
依我看,这事既不光鲜也不亮丽,但验光远胜预言。这有点像——就像看那张斯内伦视力表,你会慢慢看清一点,至于能看清多少,显然取决于你所在的那个预测环境本身是个什么脾气。
In my view, this is unglamorous, but optometry trumps prophecy. It's a bit like this. It's a bit like looking at the Snellen eye chart. You gradually become a little bit clearer, and how much clearer you can become is obviously going to hinge on the nature of the forecasting environment within which you're working.
GJP 的方法能够通过经过实验验证的工具、人员筛选、培训、团队协作、激励机制和算法来提升预见能力。这些方法被证明是有效的。
GJP's methods can improve foresight using experimentally tested tools, personnel selection, training, teaming, incentives, and algorithms. Those are the things that work.
局限性?GJP 的方法并不完美。我们远无法达到零和一,以及我之前展示的那种神一般的全知。平均而言,GJP 的最佳方法对未发生事件赋予的概率在 24% 到 28% 之间,对实际发生事件赋予的概率在 72% 到 76% 之间。这比五五开好得多,但离神或全知还差得远。
The limits? GJP's methods are not perfect. We can't get anywhere close to zero and one and the god-like omniscience I showed earlier. On average, GJP's best methods assign probabilities between about 24 and 28 percent to things that don't happen and probabilities of 72 to 76 percent to things that do. It's a lot better than 50/50. It's much short of god or omniscience.
好的,最后一部分,第四部分,推进知识的下一步。我认为我们学到了一些关于如何把预测竞赛做得更好的东西。我认为我们已经启动的第二代竞赛
OK, last section, Part Four, the next steps in advancing knowledge. I think we've learned some things about how to do forecasting tournaments better. I think the second generation of tournaments that we've
在“好判断项目”中,他们选取了 500 个具有不同背景的预测者——也就是那些能够从不同视角看待问题的人——然后从这 500 人中找出了前 100 名,接着又从中筛选出前 50 名、前 20 名,最后剩下前十几名。随后他们让这些最顶尖的预测者相互交流。他们没有向我透露这些人的交流内容,但你懂的,他们确实会在一起讨论。结果,这些顶尖预测者的预测表现远远超过了包括情报机构在内的其他所有机构。而我们则坐等这部分投资组合带来回报。
Philip Tetlock The Good Judgment Project (Continued)
我一直与 IARPA 和情报界沟通这些运作事宜,也可能与私营部门实体合作——我认为这些才是他们应该关注的方向。
been talking about running with IARPA and the intelligence community, and maybe with private-sector entities as well, these are the things I think they should focus on.
问题的质量与答案的质量——对预测竞赛而言,提出一个好问题意味着什么?我们之前对此做过一些讨论。
The quality of questions as well as the quality of answers – what does it mean to generate a good question for a forecasting tournament? We talked about that a bit.
对尾部风险的准确性和判断,以及你对主流风险的判断,正是纳西姆·塔勒布对预测竞赛所提出的批评。
The accuracy and judgments of tail risks as well as your judgments of mainstream risks, the Nassim Taleb critique of forecasting tournaments.
最后再谈一个主题:人类、模型与混合系统的相对表现。我希望 IARPA 会就此决定启动一场新的竞赛。但这事我并不确定。
And finally the performance of humans versus models versus hybrid systems, on which I hope that IARPA will decide to launch a new tournament. I don't know that for sure.
为锦标赛填充更好的问题。锦标赛需要有严谨且可裁决的问题,比如“东海将发生一起中日冲突,造成 10 人或以上死亡”。问题必须非常非常具体,一个具体的降级,从这个降级到那个。你不希望事后留有推诿空间,去争论谁对谁错。
Populating tournaments with better questions. Tournaments require rigorous and resolvable questions like, "There will be a Sino-Japanese clash in the East China Sea with 10 or more dead." There needs to be something very, very specific, a specific downgrade, a downgrade from this to this. You don't want room for after-the-fact wiggle room about who was right and who was wrong.
预测锦标赛理想上确实需要这种清晰度,但严谨的问题往往不是顶级政策制定者最关心的。顶级政策制定者对东海是否会发生意外冲突导致人员伤亡的兴趣,远不及对以下问题的关注:中国的地缘政治意图是什么?他们在各个战线上打算推进到什么程度?
Forecasting tournaments ideally have that kind of clarity, but rigorous questions aren't often the questions that top policymakers most care about. Top policymakers are less interested in whether there's an accidental clash in the East China Sea that kills some people than they are in what are Chinese geopolitical intentions? How far do they intend to push on various fronts?
很难提出一个诸如“中国的地缘政治意图是否变得更加对抗性?”这样的问题,然后过一年回来说,“啊,你是对的,你是错的。”这种做法并不太奏效。
It's very hard to ask the question, "Have Chinese geopolitical intentions become more confrontational?" and then come back in a year and say, "Ah, you were right, and you were wrong." It doesn't work very well.
所以,我们需要通过开发所谓“重大问题的多重微指标”并进行三角验证来弥合这个差距。我们把这些东西称为贝叶斯问题簇。我们有非常非常多的问题,每一个都针对一项重大政策议题。在试图变得更积极的过程中,今年我们有一组问题关于中国是否会拦截美国海军舰艇——这是一个大问题。这事儿确实发生了。不过,它没有上新闻。
So, we need to bridge the gap by developing what we call multiple mini-indicators of big issues and triangulating. We call these things Bayesian question clusters. We have many, many questions, each targeting a big policy issue. While trying to become more aggressive, we had a set of questions this year about whether the Chinese would interdict a U.S. naval vessel, a big issue. It actually happened. It didn't make the news though.
中日冲突、中国与越南和菲律宾在南海的渔业冲突、日本首相参拜二战神社、中国是否向其他国家出售无人机、以及中菲围绕仁爱礁的冲突——每一桩都是一种体现……如果所有这些问题都以对抗性的方向解决,那就会让政策辩论的天平多少偏向“中国正变得更具对抗性”这一看法。
A Sino-Japanese conflict, South China Sea fishing conflicts with Vietnam and the Philippines, the Japanese prime minister visiting a World War II shrine. Will China sell a UAV [unmanned aerial vehicle] to another country, and the conflict between China and the Philippines over the Second Thomas Shoal. Each of these is a manifestation…if all these questions resolved in the confrontational direction, that would tip the policy debate somewhat toward the view that China is becoming more confrontational.
这是一种改进方法,即提高问题的质量。另一种方法是增强对尾部风险的敏感度,并回应我所认为的“纳西姆·塔勒布之问”。
That's one way of going, improving the quality of the questions. The other way of going is creating more tail-risk sensitivity and addressing what I think of as the “Nassim Taleb question.”
我认为纳西姆是一位友善的批评者。事实上,多年前他曾在哥伦比亚大学和我合讲过一门课,当时迈克尔 [莫布森] 也一起参与了。纳西姆对这一切持怀疑态度,我认为他确实有不少充分的理由。
I consider Nassim to be a friendly critic. Actually he and Michael [Mauboussin] once jointly visited a class I was doing out at Columbia many years ago. Nassim has some very good grounds I think for being skeptical of all this.
想象一下,超级预测者告诉我们,在未来 16 个月或某个类似的时间段内,中日之间发生冲突并导致 10 人或更多人死亡的可能性为 4%。这个信息挺有趣,但算不上多么令人安心。它没有告诉我们,如果死亡人数不是 10 人,而是 1000 人甚至 100 万人、第三次世界大战——也就是那些极端的尾部风险——发生的可能性又有多大。
Imagine the superforecasters tell us that there's a four percent likelihood of a Sino-Japanese clash claiming 10 or more lives in the next 16 months or some such time period. That's interesting, but it's not all that reassuring. It doesn't tell us what's the likelihood of not 10 dead but how about 1,000 dead or a million dead, World War III, extreme tail risks.
我们需要开发对尾部风险更敏感的预测锦标赛。我们需要开发能衡量尾部风险领域技能的指标,这极为困难,因为评估
We need to develop forecasting tournaments that are more tail-risk sensitive. We need to develop indicators that are sensitive to skill in tail-risk domains, which is extremely difficult because assessing the
菲利普·泰特洛克的“优良判断项目”(续)
Philip Tetlock The Good Judgment Project (Continued)
预测概率精确到十万分之一的级别,可以说需要持续数万乃至数十万年的预测锦标赛,这恐怕连雷·库兹韦尔(Ray Kurzweil)的寿命都等不到。
accuracy of probabilities in the 0.00001 range arguably would require forecasting tournaments that would last tens of hundreds of thousands of years, which may be even beyond Ray Kurzweil's lifespan.
[laughter]
[laughter]
菲尔:更多尾部风险敏感度。以下是目前针对这一问题的两种补救措施:帮助预测者学会做出更细微的区分,以及帮助管理者提出更具探究性的问题。
Phil: More tail-risk sensitivity. Here are the two workarounds we have for that right now: by helping forecasters learn to make subtler distinctions, and by helping managers to pose more probative questions.
学习进行更精细的区分后,我现在要重温“训练”这个主题。
Learning to make subtler distinctions, I'm going to revisit now the theme on training.
我们一直在开发训练方法,让人在判断极小概率事件时逻辑上更一致,因为认知心理学文献告诉我们,人类处理极小概率事件的能力确实很差。说得客气一点,对低概率事件的估值非常不一致,往往极为不稳定。
We've been developing techniques for training people to become more logically consistent in their judgments of tiny probabilities because we know from the cognitive psychological literature that people are really bad at dealing with very small probabilities. Estimates of low-probability events are, to put it very kindly, inconsistent. They are often wildly inconsistent.
你问人们这样一个问题:“每天开车一小时的人,在一天内发生超过保险免赔额的事故的概率,与十年内发生这种事故的概率相比如何?”他们通常会低估十年内的概率,同时高估一天内的概率。
You ask people a question like, "Chance of someone who drives one hour per day having a crash above the insurance deductible in one day versus 10 years?" What they'll typically do is they'll underestimate the probability in 10 years. Then they'll overestimate the probability in one day.
他们常常在一天之内给出的概率,会暗示他们在十年期间会发生多起事故。这是一项很难学会的本领,人们需要经过训练才能掌握。
They'll often give probabilities in one day that would imply they would have many accidents in the 10-year period. This is a hard thing to learn how to do, and people need to be trained how to do it.
我们显然无法确定第三次世界大战或其他极端尾部风险事件(如流行病等)的概率。我们做不到这一点,但我们可以设计一些演练,让预测者在逻辑上更一致,并减少对分类依赖(partition dependence)等偏见的易感性。
We obviously can't pin down the probability of World War III or whatever other extreme tail-risk events, an epidemic or whatever it might be. We can't do that, but we can design exercises that make forecasters more logically consistent and less prone to biases such as partition dependence.
尾部风险的一个特点是:人们总是低估它,直到被迫直面它,那时又会严重高估它。所以,要在完全忽视和过度敏感之间把事情搞对,非常非常困难。在漠不关心和草木皆兵之间,你如何找到那个平衡点?
One of the things about tail risk is that people underestimate it until the tail risk is called to their attention, at which point they grossly overestimate it. So it's very, very difficult to get it right, between oblivion and hypersensitivity. How do you strike that balance between oblivious and hypersensitive?
学会提出更具诊断价值的问题:这是让推演竞赛对尾部风险更敏感的另一种方式。这里的“诊断价值”指的是揭示那些不可观测的尾部风险。一种做法是开发场景轨迹的早期预警指标。比如对中国而言,如果中国真的要变得极度对抗,那么专家告诉我们,在未来 12 到 18 个月里应该观察哪些迹象——设计这类问题。
Learning to pose more probative questions: that's the other way of making tournaments more tail risk sensitive. Probative means here shedding light on unobservable tail risks. One way to do it is developing early warning indicators of scenario trajectories. So the sort of thing with China, if China really were becoming extremely confrontational, what sorts of things do subject matter experts tell us we should be observing in the next 12 to 18 months. Developing questions of that sort.
还有一种非常有趣且微妙的方法,可以帮助判断谁在尾部风险方面表现更好或更差:不仅要让人们预测事件本身,还要让他们预测别人会做出什么样的预测。这是我从 GJP 的一位合作者德拉任·普雷莱茨(Drazen Prelec)那里学到的技巧——他在麻省理工学院开发了一种叫做“贝叶斯真相血清”的方法。
Another very interesting and subtle way of finding out who's better and who's worse with respect to tail risks, is by asking people not just to make predictions about events, but to make predictions about the predictions other people are making. This is a technique I owe to one of our collaborators in GJP, Drazen Prelec at MIT developed a technique called Bayesian truth serum.
结果发现,那些更擅长预测他人预测结果的人,同样也更擅长预测我们所关注的事件本身。因此,你事先就有很好的依据,知道谁会是一个更好的预测者——因为你立刻就能知道,谁在正确预测他人的预测。不仅预测他人当下的预测有用,预测他人未来的预测也很有帮助。所以,如果有人在一年前就说,主流观点认为一年后(也就是现在)俄罗斯会变得更加强硬,那显然就是一语中的。
It turns out the people who are better at predicting other people's predictions are also better at predicting the event of interest, and therefore you have a good basis ex ante, for knowing who's going to be a better predictor, because you know immediately who's correctly predicting other people's predictions. It's also helpful to be making predictions not only about other people's predictions now but other people's predictions in the future. So someone who said a year ago that consensus opinion was that Russia would be more confrontational a year later, like now, would obviously have hit a bullseye.
一旦我们掌握了哪些预测者在这些逻辑一致性测试中表现更优,以及哪些预测者更擅长预测早期预警指标、更擅长预测他人的预测——当你获得这些信息后,我认为你就有了一个统计依据,可以在尾部风险型锦标赛中对特定类型的人赋予更大权重。你不一定非要等上一万年。
Once we know which forecasters are better on these logical consistency tests, and the forecasters who are better at predicting early warning indicators and better at predicting others' predictions, once you've got that information then I think you have a statistical basis for giving greater weight to certain types of people in tail risk-style tournaments. You don't have to wait 10,000 years necessarily.
菲利普·泰特洛克 优秀判断项目(续)
Philip Tetlock The Good Judgment Project (Continued)
但你究竟减少了多少不确定性?对于尾部风险,你能获得的远见和清晰度始终极为有限。你或许只是从完全模糊不清,变得稍微不那么模糊。这才是看待这件事的正确方式——就像视力表上的连续刻度一样。
But how much uncertainty are you reducing? The amount of vision, the amount of clarity you can get on tail risk is always going to be extremely limited. You might be moving from totally blurry to just a slightly less than totally blurry. That's the right way to think about it on the eye chart continuum.
最后我来澄清几个错误的二分法。你怎么看待准确性?大多数人认为你要么对,要么错。
I'm going to close by dispelling some false dichotomies. How do you think of accuracy? Most people think you're either right or wrong.
按照“良好判断项目”的观点,正确的理解方式是:你应该把它看作一个连续体。如何解释预测的准确性?大多数人会归因于运气或技能,但正如迈克尔[莫布森]指出的,你应该把它视作一个连续体。存在一条“运气—技能连续谱”。技能的根本来源是什么?先天与后天之争?同样,这不是正确的思考方式。应该是先天与后天以不同方式共同作用。流体智力确实有遗传成分,这是事实,并且它对预测准确性很重要。
The right way to think of it, according to the Good Judgment Project, is that you should think about it as a continuum. How do you explain accuracy? Most people think luck or skill, and as Michael [Mauboussin] pointed out, you should think about it as a continuum. There's a luck-skill continuum. What are the underlying sources of skill? Nature versus nurture? Again, that's not the right way to think about it. It's nature and nurture in various combinations. Fluid intelligence does have a genetic component to it, that's true, and it's important for forecasting accuracy.
但有很多东西是你可以掌控的,同样也能提高准确性。所以,这不是一个非此即彼的问题。
But there are all sorts of things that are under your control that can also produce greater accuracy. So it's not all or none.
那么,我们和一些处于公众聚光灯下的其他人有何不同?首先,丹尼尔·卡尼曼——我在伯克利的老同事,现在也是大名鼎鼎的人物——认为人类存在认知缺陷。这本质上就是卡尼曼的论点。事实上,这些认知缺陷几乎是普遍存在的。
So how are we different from some of these other people who are in the public limelight? First, Daniel Kahneman, who's an old colleague of mine from Berkeley, as well as a mega-famous guy now, humans have cognitive shortcomings. That's the essential Kahneman argument. In fact, these are virtually universal cognitive shortcomings.
它们几乎是与生俱来地嵌在我们的感知认知系统里。我们同意丹尼的观点,即存在大量偏见,但我们对改善判断力的可行性更为乐观。如果说我们之间存在分歧,那就是这一点。我们对改进的潜力持更为谨慎的乐观态度。
They're very close to hardwired into our perceptual cognitive apparatus. We agree with Danny that there's a lot of bias, but we're more optimistic about the feasibility of improving judgment. If there's daylight between our positions, it's there. We're more cautiously optimistic about the potential for improvement.
纳西姆·塔勒布认为世界具有根本上的不可预测性,因此那些让人们误以为世界可能并非如此不可预测的竞赛,与其说是有益,不如说是有害。我们同意世界上存在大量根本性的不可预测因素,但同时,我们也看到概率性预判在帮助机构确定优先级方面发挥着关键作用。
Nassim Taleb says the world is radically unpredictable, so tournaments that lead people to think it might not be radically unpredictable are a disservice rather than a service. We agree that there's a lot of radical unpredictability out there. But we also see a key role for probabilistic foresight in helping institutions prioritize.
借用纳西姆的话来说,让机构变得反脆弱,让银行或其他机构变得反脆弱,代价极其高昂。所以你必须设定优先级,而设定优先级时,你需要考虑概率。
To use Nassim's terms about anti-fragilizing institutions, anti-fragilizing banks or other institutions is extremely expensive. So you have to set priorities, and when you set priorities, you need probabilities.
最后,还有一位你们可能听说过的人——纽约大学的布鲁斯·布埃诺·德·梅斯基塔。他写了《预测游戏》这本书,还经营着一家中等规模的咨询公司,声称凭分析工具和博弈论工具就能预测未来。
Then finally, someone else you may have heard of, Bruce Bueno de Mesquita, who's at NYU, wrote the book The Predictioneer's Game, and has a moderate size consulting firm: Analytical tools, game theory tools, can predict the future.
我们认为这话有一定道理。但我们也认为,你无需为每个问题构建极其详细的博弈论模型,用锦标赛方法就能以低得多的成本达到与他相当的精准度。
We think there's some truth to that. But we think that you don't need to construct really detailed game theoretic models for each problem, that you could deliver levels of accuracy comparable to what he can deliver much more cheaply using tournament methodologies.
这就是我在这个领域的定位。显然,我们还可以把自己和其他人比较,但我已经说了很久了,所以……我们应该开始提问环节了。请说?
That is where I position myself in this space. There are obviously other people we could have compared ourselves to, but I've talked a long time, so...we should open things up for questions. Yes?
问题:在投资环境中,你如何应用这个思路?也就是说,你有一支分析师和投资经理组成的团队,然后你设置了某种锦标赛机制,让他们做出不同的预测,那么在我们的世界里,你会如何实施呢?
Question: How would you apply this in an investment setting? So you have a team of analysts, portfolio managers, and you set up a tournament some way to have them make varying forecasts, and then how would you implement that in our world?
菲尔:我认为,具体如何实施,将取决于高层管理者认为哪些决策对概率足够敏感,值得纳入一场锦标赛(tournament)中。如果存在一些领域,管理者说,“天哪,如果我们对这些概率有更准确的把握,就会影响这项投资策略,或者如果我们知道珍妮特·耶伦真的是……”
Phil: I would say that exactly how you implement it is going to be a function of what top management considers to be the decisions that are sufficiently probability-sensitive that they're worth including in a tournament. If there are domains in which managers say, "Boy, if we had a better bead on these probabilities, it would influence this investment strategy, or if we knew that Janet Yellen was really
菲利普·泰特洛克《良好判断项目》(续)
Philip Tetlock The Good Judgment Project (Continued)
极其鸽派,而不只是适度鸽派,那会影响我们对 bup、bup、bup 采取的措施。
extremely dovish rather than just moderately dovish, that would influence what we would do with bup, bup, bup."
我不是财务专家。具体怎么做我会听你的意见,但我觉得,要让锦标赛发挥最大效果,高层管理者得相信存在某些对概率高度敏感的问题类别,值得为此投入资源去设立竞赛。再次强调,我们需要更好的问题——同时设立竞赛来催生更好的问题,以及更好的答案。
I'm not the finance guy. I would defer to you on how you would do that, but I think tournaments work best if senior management believes that there are categories of questions that are sufficiently probability-sensitive, that it's worth the investment to set up a tournament. Again, the need for better questions—and for setting up competitions to generate better questions as well as better answers.
Yes?
Yes?
问题:超级预测者事先是如何被评判的?换句话说,你们怎么识别他们?
Question: How are the superforecasters judged up front? In other words, how do you identify them?
第二,他们的特点是什么——你在考虑雇用这类人时,会看重哪些特质?
And second, what were their characteristics – in thinking of hiring these types of people, what characteristics do you look for?
菲尔:其中有几位是华尔街分析师。有两个领域——地域和职业——在超级预测者中占比异常高。我倾向于认为,他们是华尔街的金融专业人士和硅谷的程序员。当然,还有很多其他人也是超级预测者。比如特拉华州的一位药剂师,还有各式各样散居全国各地的奇人异士,个个有趣且才华横溢。
Phil: Some of them are Wall Street analysts. There are two domains that are overrepresented among superforecasters, geographically and professionally. I would say they are Wall Street financial professionals and Silicon Valley programmers. There are lots of other people who are supers too. There’s the Delaware pharmacist, and there are all sorts of eccentric people scattered across the country who are interesting and very talented.
但预测超人是怎样的?他们是特定预测年度排名前 2% 的人,依据最简规则——最佳的布赖尔得分(Brier score)——选出来。然后他们开始彼此合作,而且他们热爱合作。其中一位预测超人被分配到超级团队后说:“哎呀,我还以为自己是比赛中唯一聪明的人呢!”这些预测超人确实是精英范儿。
But what do supers look like? They were the top two percent in a given forecasting year. They got the best Brier scores, a very simple rule for choosing them. Then they get to work with each other, and they love to work with each other. One of the supers said upon being assigned to a super team he said, "Oh, I thought I was the only intelligent person in the tournament!" These supers are elitist.
[laughter]
[laughter]
菲尔:他们毫不掩饰自己是精英主义者,我们在这里使用的也是一种相当精英主义的策略。他们在流体智力上得分更高,在晶体智力上得分也更高,在认知风格上表现也更出色。他们不把自己的信念当作珍贵财物来对待,而是当作可检验的假设。他们在智力上非常灵活,拥有相当强的分析能力。
Phil: They are unapologetic elitists, and this is a pretty elitist strategy we're using here. They score better on fluid intelligence, they score better on crystallized intelligence, they score better on cognitive style. They don't treat their beliefs as precious possessions. They treat their beliefs as testable hypotheses. They're intellectually agile. They've got quite a bit of crunching power.
而且他们相信,这是一种可以培养的技能。即便你具备了超级预测者所需的所有智力与认知风格前提条件,但如果你认为这本质上靠的是运气,或是先天固化的能力,那你几乎找不到任何理由去投入培养这项技能所需的努力。
And they believe this is a skill that can be cultivated. You could have all the intelligence and cognitive style prerequisites a superforecaster has, but if you believe that it's essentially luck or it's essentially hardwired ability, there would be very little reason for you to invest the effort required to cultivate this skill.
遗憾的是,我没把迈克尔[莫布森]以前见过的那张图表带来。那是 1930 年代开发的一个测试项目。二战期间,军方用它来区分哪些农村青年有军官潜质、哪些只适合直接编入步兵。苏联红军显然也用过它——至少我是这么听说的。
Unfortunately, I didn't bring the graphic that Michael [Mauboussin] has seen before. It was a test that was developed in the 1930s. The militaries in World War II latched onto it as a basis for distinguishing farm boys who have officer potential from farm boys who just go directly into the infantry. The Soviet red Army also apparently used it—or so I have heard.
这是一种军官潜力测试方法,不依赖于读写能力。你不需要真的识字。你只需要观察图案,看清图案如何演变,然后根据驱动前序图案演变的生成规则,预测下一个图案是什么。这种测试叫做“瑞文标准推理测验”,是心理学文献中的经典测试方法。它确实有效。
It was an officer potential measure that doesn't depend on literacy. You don't have to be able to read, really. All you have to do is look at the patterns, see how the patterns are evolving, and make a prediction about what the next pattern will look like based on the generative rule that was causing the evolution of the previous patterns. This is called "Raven's Progressive Matrices." It's a classic measure in the psychological literature. It works.
我认为它在我们的比赛中能起作用,是因为我们提的问题极为多样化。看看我们在问的各种各样的问题,一般智力确实发挥了更大作用。
I think it works as well as it does in our tournament because our questions are so extremely heterogeneous. Look at the enormous variety of questions we're asking. General intelligence does play a bigger role.
如果你主持的竞赛领域更加集中,那么一般液态智力就变得不那么重要,而晶态智力则变得更为重要。
If you were running a tournament in a more concentrated domain, general fluid intelligence would become less important and crystallized would become more important.
菲利普·泰特洛克 优秀判断项目(续)
Philip Tetlock The Good Judgment Project (Continued)
提问:这组预测人士的平均年龄有多大?
Question: How old is the forecasters group, on average?
菲尔:这很有意思。相当大一部分人是早期退休者,正在寻找可做的事情。
Phil: That's interesting. There is a significant group who are early retirees who are looking for things to
确实。相对富裕的提前退休者占了较大比例。但我想说,平均年龄可能在 35 岁到 39 岁之间——中位数略低一些。
do. Relatively affluent early retirees are overrepresented. But I would say the average age is probably in the late 30s—and the median a bit lower.
提问:是什么让它们持续良好运转,而不变得功能失调、令人不快?
Question: What makes them continue to operate well and not become dysfunctional and disagreeable?
菲尔:我们会给他们一些团队协作的指导,就是安迪·格鲁夫那套建设性对抗的老调——“如何不失风度地表达分歧。”有些超级预测者的社交能力并不太好。今年确实有一个超级预测者团队彻底搞砸了。不过,总体而言,他们做得相当不错。
Phil: We give them some guidance on how to interact together in teams, the old Andy Grove constructive confrontation mantra – “How to disagree without being disagreeable." Some of the superforecasters, their social skills are not that great. One superforecaster team did crash and burn this year, actually. But, overall, they really did well.
我们会指导他们如何做到这一点。我们会指导他们一种名为“精准提问”的技巧。我们会指导他们如何相互协作、如何互相支持而非互相拆台、如何进行建设性批评而非破坏性批评。我认为这种培训确实有帮助。
We give them guidance on how to do this. We give them guidance on a technique called "precision questioning." We give them guidance on how to get along with each other, how to be supportive of each other and not to be destructive, how to be constructively critical rather than destructively critical. I think that training does help.
他们确实对群体思维的危害非常警觉。如果说超级团队有什么风险,那与其说是群体思维,不如说是破坏性竞争。不过,在大多数情况下,我们一直能够控制住这一点。
They're really attuned to the dangers of groupthink. If there's any danger in the super teams, it's not really groupthink as much as it is destructive competition. But, for the most part, we've been able to keep that under control.
让我看看。有什么事吗?
Let's see. Yes?
问题:情报机构掌握着更好的信息,它们的表现又如何解释?
Question: What explains the performance of the intelligence community, which has better information?
四位水平一般的预测者如果拥有更好的数据作为依据,其表现应当优于一位超级预测者。
Four average forecasters with better data to work with should do better than a superforecaster.
菲尔:问得好。如果你去问情报界的某些人,他们会说:“那是因为我们还没开始认真动手。我们根本没把你当回事。要是我们真把你当回事了,肯定能打得你满地找牙。”这是其中一种反应。
Phil: That's a great question. If you were to ask some people in the intelligence community, they would say, "That's because we haven't begun to fight. We weren't taking you seriously. And if we ever did take you really seriously, we would kick your butt." That's one reaction.
另一种可能是,我们打交道的这个情报圈子里的人还不习惯使用量化概率。他们只是需要一些时间来适应这种做法。
Another possibility is that the parts of the intelligence community we're dealing with is not accustomed to using quantitative probabilities. It's just taking time for them to get accustomed to doing it.
问:这是对数据的过度自信吗?您认为过度自信导致了这一点?
Question: Is this an overconfidence in the data? Do you think overconfidence is doing that?
菲尔:这是另一种解释——由戴维·伊格纳修斯提出。情报分析人员被大量可疑有用的机密信息淹没。这是一个信号/噪音问题。也许我们的预测者拥有的优势恰恰是没有机密信息,而非劣势。
Phil: That's another explanation that's been offered—by David Ignatius. It's that the intelligence analysts are overwhelmed by classified information of dubious usefulness. It's a signal/noise problem. It could be that our forecasters have the benefit of not having classified information, rather than the liability.
Yes?
Yes?
问题:还有哪些偏见会导致糟糕的预测?想想 1 月 1 日——人们觉得所有这些事情都会发生,其数量远远超过典型的 12 个月里实际可能发生的。
Question: What other biases lead to bad forecasts? Think about the first of January – people think all these things are going to happen, a lot more than probably do in a typical 12-month period.
菲尔:没错。或许存在一条极为简单的经验法则,能让那些更优秀的预测者免于这种错误。这是一条基础比率(base rate)启发式:在数以千计的预测问题中,平均而言,变化是一种概率更低的结果。
Phil: Yes. There could be a very simple heuristic that spares our better forecasters from that error. It is a base rate heuristic that change is a lower probability outcome, on average, across thousands of forecasting problems.
人们往往倾向于夸大变化的可能性。这也是一个更有趣的预测。预测者希望自己显得有趣。各种激励机制在起作用,鼓励着概率上的不准确。
There is a tendency to exaggerate the probability of change. It's also a more interesting forecast. And forecasters want to be interesting. There are incentives operating that encourage probabilistic inaccuracy.
问:当完全没有基础概率可用,当你面对的是一个纯粹独一无二、绝无仅有的事件时,预测者能做什么?
Question: When there are no base rates at all, when you're dealing with a pure sui generis, unique event, what can the forecasters do?
菲利普·泰特洛克“优秀判断项目”(续)
Philip Tetlock The Good Judgment Project (Continued)
菲尔:嗯,就像迈克尔 [莫布森] 谈到卡尼曼和特沃斯基那篇经典论文时说的,最后剩下的只有内在视角。你不得不依靠因果原理。你说:“历史上没有哪个人跟这个人相似。”
Phil: Well, as Michael [Mauboussin] said about the classic Kahneman and Tversky paper, the only thing that's left is inside view. You're going to have to draw on causal principles. You say, "There's nobody who resembles this person in history."
说实话,我很难想象有什么事是完全独一无二的。即便是极其罕见的事件,我们最优秀的预测者通常也能拼凑出合理的基准概率(base rate)。
It's very hard for me to imagine, by the way, that something is totally unique. Even for extremely idiosyncratic events, it's typically possible for our best forecasters to cobble together a reasonable base rate.
但在那些完全想不到任何基准概率的情况下,你将不得不依赖内部视角(inside view)和支撑该视角的因果概括。也许是来自博弈论或其他社会科学理论体系的因果概括。接下来的问题就是,“这种推理会导向什么样的预测?”但在那种情景下,你只能依赖纯粹的因果推理。
But in the cases where no base rate comes to mind at all, you will be required to rely on the inside view and causal generalizations that plug into the inside view. Maybe causal generalizations from game theory or from of some other body of theory in social science. Then the question is, "What kind of forecast does that lead to?" But you would have to rely on pure causal reasoning in that scenario.
我的经验是,这有点危险,但又常常必不可少。当然,我们的预测者始终在把统计推理和因果推理交织在一起运用。
My experience is that's a little dangerous. But it is often essential. Certainly, it is true that our forecasters weave together statistical and causal reasoning all the time.
提问者:你谈了很多关于多样性在聚集预测者和制定加权方案方面的价值。你能谈谈你在寻找预测者群体多样性时,会关注哪些精明的指标或特质吗?
Question: You talked a lot about the value of diversity in bringing together the forecasters and then the weighting schemes. Can you talk a little bit about the savvy metrics or attributes you look for in finding diversity across the forecaster groups?
菲利普:我使用“多样性”这个词时,不是用你们组织里人力资源部门使用的那种方式,那更多是一种法律概念,涉及受保护群体和歧视。我所说的多样性是指功能性多样性或实质性多样性。
Phil: When I use the term "diversity," I'm not using the term "diversity" in the way human resources in your organizations use diversity, which is more of a legal construct having to do with protected groups and discrimination. I'm using diversity in the sense of functional diversity or substantive diversity.
围绕关于中国的问题,桌上有多少种观点?它们是什么?桌上有没有懂中文、深度关注中国媒体、在中国生活过的人?有没有认识中国中央委员会成员的人?诸如此类。
How many points of view on China are represented around the table? What are they? Are there people around the table who know Chinese well, that are following the Chinese media in depth, who have lived in China? Are there people who have known people on the Chinese Central Committee? Things like that.
提问者:你招募这些想成为预测者的人。是通过试错来了解他们有多样的性,还是有特定问题可以问?
Question: You get these people who want to become forecasters. Is it like a trial and error to figure out how diverse they are, or are there questions you can ask?
菲利普:这是个很棒的问题。多样性让我们的统计学家抓狂,因为他们知道这在决定极端化算法(extremization algorithm)应该有多极端时至关重要:该公式中的 A 应该取什么值。他们知道这在逻辑和数学上非常重要。但他们一直无法确定预测者在不同领域中哪些具体特质对多样性至关重要。
Phil: This is a great question. Diversity makes our statisticians tear their hair out because they know it's important in determining how extreme the extremization algorithm should be: what value should A take in that formula. They know that's really important logically and mathematically. They haven't been able to pin down the specific attributes of forecasters that are essential for diversity in different domains.
我认为这是因为问题太异质了。随着我们转向一种有议题集群(question clusters)的模式,比如一个非常大、非常集中的中国议题集群,我们就能更有意义地定义多样性。但现在他们要从中国跳到缅甸,再到印度,再到希腊。这非常困难,是一个正在进行的工作。
I think that's because the questions are so heterogeneous. As we move toward a model where we have question clusters, a big, big question cluster on China, we'll be able to define diversity more meaningfully. But they flit from China to Myanmar, to India, to Greece. It's very difficult. It's a work in progress.
提问者:你能谈谈纳西姆·塔勒布关于路径依赖的观点如何在这里应用吗?
Question: Can you discuss how Nassim Taleb’s views on path dependency fit in here?
菲利普:纳西姆的部分论点具有深刻的哲学性。就拿现在人们争论的 1914 年与 2014 年的类比来说——一战前的英德对抗和今天的中美紧张局势。德国和中国都是崛起的力量。英国或美国则是既有的霸权。
Phil: Part of Nassim's argument is deeply philosophical. Take the argument you see nowadays between people who think there's an analogy between 1914 and 2014, the Anglo-German rivalry that preceded World War I and the Sino-American tensions today. You have rising power in Germany and China. You have an entrenched power in England or the United States.
其主张是,这里存在着某种因果规律。纵览历史,一直追溯到伯罗奔尼撒战争中的斯巴达和雅典。你会看到一个模式:当崛起的大国挑战霸主或主导力量时,战争的可能性会急剧上升。因此,亚洲发生非常糟糕事件的可能性就大大提高了。
The claim is that there's a causal law operating here. You look through history all the way back to Sparta and Athens in the Peloponnesian War. You see a pattern in which a rising power coming up against a hegemon, a dominant power, that the likelihood of war goes boom. Therefore, the likelihood of something really nasty happening in Asia is way up there.
菲利普·泰洛克,“良好判断项目”(续)
Philip Tetlock The Good Judgment Project (Continued)
那么,你提到的这个路径依赖论点非常有趣。假如我认为第一次世界大战并非不可避免。我认为一战是一个意外,因为某个疯子刺杀了萨拉热窝的大公。如果你把 20 世纪初的历史重演一千次,一战最多只会出现一次。
Now, the path dependency argument that you mentioned here is a very interesting one. Imagine I think that World War I was not inevitable. I think World War I was a fluke that occurred because some crazy guy shot the archduke in Sarajevo. And if you were to rerun early 20th century history a thousand times, World War I would emerge at most once.
如果我是一个持有这种极端路径依赖观点的人,我就不会对 1914 年和 2014 年的类比印象深刻。我认为 2014 年和 1914 年截然不同。
If I'm an extreme path-dependency person like that, I'm not going to be too impressed by the analogy between 1914 and 2014. I think 2014 is just very different from 1914.
你对反事实和路径依赖所做的假设,在你如何使用历史类比中扮演着关键角色,而历史类比的运用在你如何生成关于中国的预测中又扮演着关键角色。这说得通吗?
The assumptions you make about counterfactuals and path dependency play a pivotal role in how you use historical analogies, which play a pivotal role in how you generate predictions on China. Does that make sense?
这里存在一个断言与证据的比例问题。纳西姆确实提出了很多断言,但其中一些断言所依据的证据数量可能没有你期望的那么多。他提出的一些主张具有相当宏大的历史哲学特征。
There's an assertion-to-evidence ratio problem here. Nassim does make a lot of assertions, but the amount of evidence for some of the assertions is not as great as you might hope. Some of the claims he's making are of rather sweeping, philosophy of history character.
提问者:例如,路径依赖对金融市场可能非常重要。你可能正接近一个事件,但可能会朝一个方向或另一个方向发展。突然间,所有版本都开始起作用,从这个方向来一个,从那个方向又来一个。这就是我们在金融市场中经常处理的那类问题。
Question: Path dependency could be very important to the financial markets, for example. You might be approaching an event, but you could go in one direction or the other. All of a sudden, all the versions come into play, from this direction and another one from that direction. That's the sort of problem that we often deal with in the financial markets.
菲利普:我看到我的经济学家朋友们对路径依赖的重要性有很多争论。
Phil: I see a lot of argument among my economist friends about the importance of path dependency.
例如,经济学家会说,闪电崩盘(Flash Crash),“啪”地一下,就过去了。闪电崩盘的持久影响是什么?他们像看待一般均衡模型一样,认为市场具有很强的自我修正能力。
The economists, for example, say the Flash Crash, [snaps] it's gone. What's the lasting legacy of the Flash Crash? They see markets as very self-correcting, as in general-equilibrium models.
而另一些人,比如纳西姆和许多思考历史的人,则认为路径依赖性要强得多。我不知道在座各位对市场与路径依赖的看法分布是怎样的。
Whereas others, like Nassim and many other people who think about history, see it as much more path-dependent. I don't know what the distribution of opinion in this room would be about markets and path dependency.
提问者:你能谈谈这个观点吗——领域越狭窄……这是一种可迁移的技能吗?在预测体育方面做得好的人,也能成为更好的市场预测者,也能成为更好的天气预测者?它对领域不敏感,还是并非如此?
Question: Can you talk about this idea that more narrow domains...Is this a transferable skill? Somebody who does well at predicting sports can also be the better predictor in markets, can be the better predictor in weather? Is it agnostic to the domain, or not?
菲利普:我之前关于一般智力的讨论,即一般智力作为在 IARPA 那种极其异质性问题竞赛中预测绩效的重要性,暗示着存在一种“预见力 g”(foresight g),在一定程度上的确如此。
Phil: My earlier discussion about general intelligence, the importance of general intelligence as a predictor of performance in an extremely heterogeneous question tournament of the IARPA sort, implies that there is a foresight g, to some degree, yes.
但我们掌握的证据表明,通过实践、训练和团队协作,这种能力是可以培养的——这更多的是领域特定性的,而不是一种一般的 g 因素。你真正在意的事情,你会做得相当不错。我认为存在一些通用性,但也存在一些极端领域特定性的情况。
But the evidence that we have that is cultivatable--through practice, training, and teaming--suggests it's not so much a g as it is something that's going to be domain-specific. The things you really care about, you're going to get pretty good at. I think there is some generalizability, but I think there are also pockets of extreme domain-specificity.
提问者:你能谈谈一些针对尾部风险的去偏技术吗?有没有一些更普遍应用的去偏技术可以告诉我们?
Question: Can you talk about some de-biasing techniques, in the context of tail risks? Are there de-biasing techniques that are applied more generally you can tell us about?
菲利普:最基本的去偏技术是在做出预测之前找到恰当的参考和比较类别。过去有哪些情况与当前情况类似?当想到多个比较类别时,你可能会对这些比较类别进行某种平均或加权平均。这就成了你的初始概率估计,然后你再深入了解该案例的内部视角细节。作为粗略的第一近似值,这对预测美国 GDP 增长之类的效果还不错。
Phil: The most basic de-biasing technique is to find the right reference and comparison classes before you make a prediction. What kinds of situations are similar to this situation, from the past? When multiple comparison classes come to mind, then you probably do some kind of averaging or weighted averaging of those comparison classes. That then becomes your initial probability estimate, and then you go into the details of the inside view of the case. That works pretty well as a rough, first approximation for predicting, say US GDP [gross domestic product] growth.
菲利普·泰洛克,“良好判断项目”(续)
Philip Tetlock The Good Judgment Project (Continued)
它在我们所谓的 10/90 领域可能效果很好。一个问题要进入 IARPA 竞赛,领域专家必须认为其可能性至少为 10%,但不超过 90%。他们希望问题有一定难度。我们的竞赛刻意排除了尾部风险。这也是我为什么强调在第二代竞赛中关注尾部风险的重要性。
It may work well in what we call the 10/90 domain. For a question to make it into the IARPA tournament, subject matter experts have to consider that it's at least 10 percent likely, but not more than 90 percent likely. They want them to be somewhat difficult. Our tournament has deliberately excluded tail risks. That's one of the reasons I emphasize the importance of a tail-risk focus for second generation tournaments.
提问者:两个预测者之间是否存在某种相关性,与人常说的自闭症谱系有关?
Question: Is there any correlation between two predictors and what people refer to as the autistic spectrum?
菲利普:你是指更像是阿斯伯格综合征而非自闭症吗?
Phil: You mean more Asperger-ish than autistic?
Question: Yes.
Question: Yes.
菲利普:我猜测:相关性不太高。可能有一点,但不高。说“一点”的话,相关系数大概在 0.10,0.15。这只是猜测。我的直觉是你说到了一些东西,但影响很小。
Phil: My guess-timate: Not very high. It may be slight, but not very high. By slight, perhaps a correlation coefficient of 0.10, 0.15. That's just a guess. My hunch is that you're onto something, but it's small.
现场股东:我是最近一期“良好判断项目”的参与者。
Audience Member: I was a participant in the most recent Good Judgment Project.
菲利普:哦,谢谢。
Phil: Oh, thank you.
提问者:做出良好预测大致有两个要素——你的初始预测和更新你的观点。你能谈谈这两件事的相对重要性吗?
Question: There were basically two elements to making a good prediction – your initial prediction and updating your view. Could you talk about the relative importance of these two things?
菲利普:很多复杂的问题都与时机、更新以及随时间更新有关。我可以这样说。当你审视在问题刚提出时、在问题中期阶段以及最后阶段所作判断的准确性关联因素时,它们大致相同。
Phil: Lots of complicated issues having to do with timing, updating, and updating through time. Here's what I can say. I could say that when you look at the correlates of accuracy for judgments made when questions were first asked, in the middle phase of the questions, and in the final phase of the questions, they're more or less identical.
这没有太大区别。准确性的关联因素在整个预测窗口期内都惊人地一致。
It doesn't make a big difference. Accuracy correlates are remarkably consistent, across periods of the forecasting window.
提问者:如果我们审视评估概率的过程,你强调的因素之一是开放性思维,也就是愿意重新评估自己的观点。你在超级预测者和其他预测者中看到这个过程有不同模式吗?
Question: If we look at the process of assessing probabilities, one of the factors you emphasize is open-mindedness, a willingness to reassess one’s views. Do you see a different pattern in this process for superpredictors compared to other predictors?
菲利普:他们是纪律严明的信念更新者。超级预测者在其概率判断上的波动性,实际上比普通预测者要小。他们更新的频率更高,但调整的幅度更小,我认为这与超级预测者判断更具颗粒度的观点是一致的。
Phil: They are disciplined belief updaters. Supers actually show less volatility in their probability judgments, over time, than regular forecasters do. They update more often, but they update by smaller amounts, which I think is consistent with this idea of greater granularity of superforecaster judgments.
我告诉过你们阿莫斯·特沃斯基的那个老笑话,人们只区分三个概率水平。当你查看国家情报评估时,为国家提供最高级别评估的国家情报委员会(National Intelligence Council)会区分七个概率水平。
I told you about the old Amos Tversky joke, that people only distinguish three levels of probability. When you look at the national intelligence estimates, the National Intelligence Council, which does the top estimates for the government, they distinguish seven levels of probability.
我们的超级预测者能区分大约 20 个概率水平。国家情报委员会的人完全有能力做到这一点,但他们没有使用正确的测量工具。
Our supers are distinguishing about 20 levels of probability. People on the NIC, the National Intelligence Council, are perfectly capable of doing this as well, but they're not using the right measurement tools.
当我说我们的超级预测者能区分大约 20 个概率水平时,这相当于 0.05、0.10 和 0.15 这样的区分。
When I say our supers are distinguishing about 20 levels of probability, that's the equivalent to about 0.05, 0.10, and 0.15.
我怎么知道的?我们之所以知道,是因为如果你把他们预测中的小数位舍去,准确度就会下降。超级预测者在使用中间数字时所做的并非无意义的区分。他们实际上是在合理地提高他们的准确性。
How do I know that? We know it because when you round off their forecasts, there's degradation of accuracy. It's not that the superforecasters are making silly distinctions when they're using intervening numbers. They're actually doing a reasonable job of improving their accuracy.
这是一个非常有趣的测量问题:在不同的问题领域中,人们能够区分多少个不确定性的水平。在座的每个人显然都能做得比三个更好。
It's a very interesting measurement question of how many levels of uncertainty people are capable of distinguishing in different problem domains. Everybody in this room can obviously do better than three.
菲利普·泰洛克,“良好判断项目”(续)
Philip Tetlock The Good Judgment Project (Continued)
在像扑克这样的游戏中,你能得到非常快速的反馈,并且有明确定义的抽样范围,你可能能找到能区分到 0.5013 的扑克专家。当然,扑克电脑可以做到这一点。
In a game like poker, where you get very rapid feedback, and you a have a well-defined sampling universe, you can probably get poker experts who can go to 0.5013. Certainly the poker computers can do that.
恐怕你得有非常强的阿斯伯格特质才能走那么远。
Probably you'd have to be very Asperger-ish to be able to go that far.
在某些领域,我认为人们可以变得超级精细。在另一些领域,精细度则更难实现。
In some domains, I think people can become super-granular. In others, granularity is harder to achieve.
这是为你们这些同仁做的一个成本-收益问题:在一个特定知识领域投入精力去获取更精细的颗粒度,对你而言有多大价值?不过,颗粒度确实与准确性高度相关,其相关性远超阿斯伯格综合征。
It's a cost-benefit question for you guys. How valuable is it to you to invest in becoming more granular in a particular content domain? But granularity is definitely a strong correlate of accuracy, much stronger than Asperger's.
卡罗琳·巴基 哈佛大学公共卫生学院
Caroline Buckee Harvard School of Public Health
卡罗琳·巴基于 2010 年夏季加入哈佛大学公共卫生学院,担任流行病学助理教授。她的研究重点是揭示驱动疟疾寄生虫及其他遗传多样性病原体动态变化与进化背后的机制。
Caroline Buckee joined Harvard School of Public Health in the summer of 2010 as an Assistant Professor of Epidemiology. Her focus is on elucidating the mechanisms driving the dynamics and evolution of the malaria parasite and other genetically diverse pathogens.
在获得牛津大学博士学位后,卡罗琳作为亨利·韦尔科姆爵士博士后研究员,在肯尼亚医学研究所工作,分析疟疾的临床与流行病学特征。她的研究成果为她赢得了圣塔菲研究所的欧米迪亚研究员职位,在那里她发展了理解疟疾寄生虫进化与生态学的理论方法。在哈佛,她利用数学模型扩展了这些方法,以连接疟疾流行病学中不同生物学尺度;她与实验研究人员合作,理解宿主体内导致疾病与感染的分子机制,并使用基因组和手机数据将这些个体层面的过程联系起来,以理解群体层面的传播模式。
After receiving a PhD from the University of Oxford, Caroline worked at the Kenya Medical Research Institute to analyze clinical and epidemiological aspects of malaria as a Sir Henry Wellcome Postdoctoral Fellow. Her work led to an Omidyar Fellowship at the Santa Fe Institute, where she developed theoretical approaches to understanding malaria parasite evolution and ecology. Her work at Harvard extends these approaches using mathematical models to bridge the biological scales underlying malaria epidemiology; she works with experimental researchers to understand the molecular mechanisms within the host that underlie disease and infection, and uses genomic and mobile phone data to link these individual-level processes to understand population-level patterns of transmission.
她的研究成果曾发表在《科学》和《美国国家科学院院刊》等知名科学期刊,也出现在包括 CNN、 《新科学家》、美国之音、美国国家公共广播电台和美国广播公司在内的主流媒体上。
Her work has appeared in high profile scientific journals such as Science and PNAS, as well as being featured in the popular press, including CNN, New Scientist, Voice of America, NPR, and ABC.
卡罗琳·巴基:理解疾病传播的新方法
Caroline Buckee Novel Ways to Understand How Disease Spreads
迈克尔·莫布森:欢迎各位回来。我们这个时代的一个引人入胜的问题,无疑是理解诸如时尚潮流、投资理念,当然还有疾病等事物,是如何在一个相互关联的代理网络中传播的。
Michael Mauboussin: Welcome back, everybody. Certainly one of the fascinating questions of our time is understanding how things, including fads or fashions, investment ideas, and of course, diseases, propagate across a network of interconnected agents.
我们下一位演讲者,卡罗琳·巴基教授,正在进行一些令人着迷的工作,以深入了解这个问题。在她这样做的过程中,她有可能在这个过程中拯救数十万人的生命。
Our next speaker, Professor Caroline Buckee, is doing some fascinating work to gain insight into this question. As she does so, she has the potential to save hundreds of thousands of lives in the process.
卡罗琳·巴基是哈佛大学公共卫生学院的流行病学助理教授。她的工作主要集中在理解疟疾和其他多种病原体的动态变化与进化。
Caroline Buckee is an assistant professor of epidemiology at the Harvard School of Public Health. Her work focuses primarily on understanding the dynamics and evolution of malaria and other diverse pathogens.
我第一次见到卡罗琳是在圣塔菲研究所,当时她是一名博士后和欧米迪亚研究员。她有机会在一个非常跨学科的环境中探索解决这个问题的理论方法。卡罗琳的工作中有几个具体方面让我非常兴奋。
I first met Caroline at the Santa Fe Institute where she was a postdoc and an Omidyar fellow. She had the opportunity to explore some theoretical approaches to this problem in a very multidisciplinary setting. There are a couple specific things about Caroline's work that really excite me.
首先,通过跨越学科边界并利用新颖的数据集,卡罗琳和她的同事们正在对疾病的传播形成深刻的洞见。正是模型、计算能力和数据之间的相互作用,提升了我们对这一棘手问题组的理解。
First is that by transcending disciplinary boundaries and utilizing novel data sets, Caroline and her colleagues are developing deep insights into the spread of disease. It's the interaction between models, computing power, and data that are improving our understanding of this tricky set of problems.
其次,正如你们将看到的,这项研究有潜力拯救大量生命。这项工作本身当然就很有趣味,因为它将阐明这些事物传播的过程,但同时,它也能以非常重要的方式真正为公共政策提供信息。
Second is that this research, as you'll see, has the potential to save lots of lives. The work is, of course, inherently interesting as it'll illuminate the processes by which these things spread, but at the same time, it can truly inform in a very important way, public policy.
这是理论与实践的一种很酷的结合。请大家和我一起欢迎卡罗琳·巴基教授。
This is a cool melding of theory and practice. Please join me in welcoming Professor Caroline Buckee.
[applause]
[applause]
卡罗琳·巴基:我在哈佛大学公共卫生学院的流行病学系和传染病动态中心工作。实际上,我们最近参与了上个流感季的流感预测锦标赛,我团队中的成员赢得了那次比赛,你需要预测每年流感疫情的时机和规模。所以我们也在开始进行这类预测练习。
Caroline Buckee: I'm at the Harvard School of Public Health in the Department of Epidemiology and in the Center for Communicable Disease Dynamics. Actually, we recently participated in the flu forecast tournament for this last season, and members of my group won that, where you have to predict the timing and size of the flu epidemic each year. So we're starting to do these kinds of forecasting exercises as well.
我想谈谈我最感兴趣的人群,他们是地球上最贫困的群体之一。具体来说,是儿童。所以这本手册(指《在没有医生的地方》这本书)已经被翻译成几十种语言,在实地工作中非常宝贵,那里基本上没有医疗条件,这本手册包含了你可能想知道的一切答案。
What I want to talk about are the populations that I'm most interested in, which are among the poorest communities on the planet. Specifically, children. And so this handbook [refers to “Where There is No Doctor”] has been translated into tens of different languages, and it's invaluable in the field, where essentially if there's no medical care available, this has the answers to everything you might want to know.
作为流行病学家,我们在群体层面思考,试图理解疾病在人群中的传播。所以我们关心的是,如何填补那些没有数据的地方的知识空白。
As an epidemiologist, we think on a population level and we try and understand the spread of disease through populations. And so what we're concerned about is how we try and fill the knowledge gaps where there’s no data.
在低收入环境下的大多数地方,当我们谈论流行病学数据时,我们实际上想到的是人们坐上路虎车,在泥泞中行驶,去询问人们调查问卷、采集血液样本这类事情。我认为,这就是大数据和数字数据能够对公共卫生产生巨大影响的空白地带。危险在于,炒作太多,却没有研究问题。
In most places in low-income settings, when we talk about epidemiological data, we’re really thinking about people getting into Land Rovers and driving through the mud and going and asking people surveys and taking blood and these kinds of things. I think that this is the gap where big data and digital data can really have a massive impact for public health. The danger is that there’s too much hype and there’s no research question.
我认为有人做了一个类比,说大数据就像在干草堆里找一根针,而我的同事则认为,它可能更像是在一大堆针里找一根针,你需要知道你想要哪根针以及为什么想要它。没有这个,没有好的方法,你就会迷失。
I think that the analogy has been made that big data is like searching for a needle in a haystack, and my colleague suggests that maybe it's like searching for one needle among a giant pile of needles, and you need to understand which needle you want and why you want it. Without that, without good methods, you're lost.
卡罗琳·巴基:理解疾病传播的新方法(续)
Caroline Buckee Novel Ways to Understand How Disease Spreads (Continued)
我将谈谈我们的一些项目,试图理解大数据在传染病流行病学中的位置。为了把这一点放在背景中,我认为流行病学正在发生变化的最重要和最有意义的领域之一,是理解人类行为模式,特别是迁徙和流动性。
I'm going to talk about some of the projects that we have to try and understand where big data fits in infectious disease epidemiology. Just to place this in context, I think one of the most important and interesting places that epidemiology is changing is in understanding human behavior patterns and specifically migration and mobility.
在过去几代人的时间里,每个人一生中的旅行范围已经发生了巨大变化。这项研究追踪了一个家族四代、四位男性一生的移动范围。
The extent of travel that each individual undertakes in their lifetime has changed enormously over the last couple of generations. This study just tracks the lifetime extent of movements in four men from one family for four generations.
这是曾祖父,他的活动范围大约有几十公里。祖父去了英国的几个郡。父亲去了欧洲内的几个地区。当然,我们都熟悉这样的旅行模式——我们现在很快就能在全球范围内移动。这意味着我们处在一个全球互联的世界,从这里出现的传染病可以迅速传播。
So here's the great grandfather, and he went on the order of tens of kilometers. The grandfather went to a couple of counties in the U.K. The father went to regions within Europe. And of course, we're all familiar with these travel patterns – we quickly migrate globally now. What that means is that we are a globally connected world, and infectious diseases that emerge over here can quickly spread.
我们都熟悉这一点。2009 年,爆发了始于墨西哥的猪流感。在几个月的时间里,它传播到了全球。当时我们是幸运的,因为它传染性不强,并且没有引起那么严重的疾病,但全球大流行的可能性是真实存在的,我认为我们需要开始认真对待这些威胁。
We're all familiar with this. In 2009, there was the swine flu outbreak which started in Mexico. In the course of a few months, it spread globally. Now, we were fortunate that it wasn't that transmissible and it didn't cause that severe a disease, but the potential for global pandemics is real, and I think that we need to start taking these threats seriously.
当然,这是一种尾部风险类型的事件,即全球大流行,但考虑到过去十年中这类事件出现的频率,我认为公平地说,这是我们非常担心的事情。
Certainly, this is a tail risk type of event, a global pandemic, but, given the frequency of emergence of these types of events in the last decade, I think it's fair to say that this is something that concerns us a lot.
这种情况正在发生,就在加勒比海地区。如果你们中有人要去加勒比海度假,我们现在有一种新的病毒叫做基孔肯雅热,正在出现。我们对这种病毒了解不多。它由传播西尼罗河病毒的同一类蚊子传播,所以我们在美国也面临风险。上周我们在波多黎各报告了第一例病例。所以,如果你们要去加勒比海,请带上驱虫喷雾。
This is happening right now, in the Caribbean. If any of you are going on vacation in the Caribbean, we now have a new virus called Chikungunya that's emerging. We don't know much about this virus. It's spread by the same mosquito that spreads West Nile, so we are at risk of it here in the United States. We had our first reported case in Puerto Rico last week. So, if you're going to the Caribbean, wear bug spray.
这是一个真实且不断出现的问题,影响着每个人,但现实是,如果你查看全球传染病死亡率,几乎全部集中在撒哈拉以南非洲和亚洲的低收入环境,以及南美洲的一小部分地区。传染病仍然占所有死亡人数的四分之一,在儿童中,这一比例达到 65%。
This is a real and emerging kind of a problem that affects everyone, but the reality is that, if you look at global mortality due to infectious disease, it's almost all concentrated in low-income settings in Sub-Saharan Africa and Asia and a little bit in South America. Infectious diseases still account for a quarter of all deaths, and, in children, 65 percent.
在低收入环境中,它们在死亡率和经济负担方面确实产生了巨大影响。我向我的学生展示这张幻灯片,是为了论证除了统计模型之外,还需要动态模型。当我们考虑慢性病,比如糖尿病或癌症时,你可以想象很多人各自在玩自己的弹球机。预期结果是每个人独立游戏的平均值。
They really represent a huge impact in terms of the burden of mortality and economics in low-income settings. I show this slide to my students to argue for dynamical models in addition to statistics. When we think about chronic diseases, like diabetes or cancer, you can imagine many people playing their own pinball machines. The expected outcome is the average of everybody's individual game.
可能存在随机性和各种不同的情况,但事实是你的结果不一定与他人相关。但当你在谈论传染性感染,就像任何传播过程一样,每个人的结果都依赖于其他人,并且系统中存在反馈。所以你需要动态模型。理想情况下,你需要了解一些关于传播机制的知识。
There can be stochasticities and different things can happen, but the fact is that your outcome isn't necessarily related to somebody else's outcome. When you're talking about communicable infections, like any contagion process, the outcome of everybody depends on everyone else, and there's feedback in the system. So you need dynamic models. And ideally, you need to understand something about the mechanism of contagion.
我将简要介绍一下传染病的工作基础模型,即易感-感染-康复模型,也就是 SIR 模型。作为一个数学模型,我们把一切都想象成圆圈,这是一个人群。他们都是相同的。没有年龄结构,这是一个封闭的人群,所以我们假设没有出生、死亡,也没有人进出。
I'm going to just walk through the workhorse of infectious disease, the susceptible-infectious-recovered model, the SIR model. As a mathematical model, as we imagine everything as circles, this is a population of people. They're all the same. There's no age structure, and it's a closed population, so we'll imagine that there's no births, deaths, and nobody's coming in and out.
每个人都可以根据疾病被分类到这个单一的隔间里。每个人都是易感的。我们在群体中加入一个感染者,突然之间,我们就有了两类感染人群。大多数人处于易感隔间。现在有人因为某种原因变得具有传染性,然后这个人就可以将疾病传播给其他人。
Everybody can be classified in this one compartment with respect to the disease. Everyone is susceptible. We add one infectious person to the mix, and, suddenly we have two categories of infected people. Most people are in this compartment. Somebody has now become infectious for whatever reason, and that person can now transmit the disease to others.
卡罗琳·巴基:理解疾病传播的新方法(续)
Caroline Buckee Novel Ways to Understand How Disease Spreads (Continued)
这个箭头代表了这两个隔间之间的流动,我们就可以开始构建一个动态模型,来描述每个隔间中人数比例的变化。这里,第一个人把疾病传染给了一些人,然后这些人又继续传播下去,如此类推。
This arrow represents the flow between these two compartments, and we can start to build up a dynamic model that describes the change in the fraction of the number of people in each of these compartments. Here, the first person gave the disease to some people, and then these people passed it on, and so on.
最终,最初的感染者,在这个例子中他们变得免疫,动态过程仍在继续。我们通常用一些确定性的微分方程来建模,尽管这个主题有无数种变化。如果你写下这些方程并模拟模型,你会得到一条看起来有点像这样的流行病曲线。
Eventually, the original infections, those people become immune in this example, and the dynamic continues. We usually model this as some deterministic differential equations, although there are innumerable variations on that theme. If you write down those equations and simulate the model, you get an epidemic curve that looks a bit like this.
在疫情初期,每个人都是易感的。感染人数呈指数增长,并在某个点达到峰值。当易感人群耗尽时,疾病逐渐消失,康复人群增加,然后趋于平稳。尽管这是一个极其简单的框架,假设也极其简单,但它在很多情况下确实能很好地描述疫情爆发。
At the beginning of the epidemic, everybody is susceptible. There's an exponential growth in the number of infected people that peaks somewhere. As your susceptible population runs out, the disease dies out, and your recovered population increases and then plateaus. Although this is an incredibly simple framework, and the assumptions are incredibly simple, it actually describes epidemic outbreaks pretty well in a lot of cases.
这是来自巴基斯坦一个登革热项目的数据,我稍后会详细讨论,它发生在 8 月到 11 月之间,针对的是一个以前从未经历过登革热的人群,它显示了这种……不同的颜色代表不同的地区。
This is data from a dengue project in Pakistan that I'll talk more about in a minute, but this occurred between August and November in a population that had not experienced dengue before, and it shows this kind of...The colors are just different districts.
你们可以看到这条典型的流行病曲线。它呈指数级增长,达到一个峰值,然后,易感人群大概会被消耗殆尽,疫情也随之消退。我们现在观察到的是感染人群。我们实际上没有测量易感人群的手段,可以测量康复人群,但这很难。
You can see this classic epidemic curve. You get exponential growth. You get a peak and then presumably, your susceptible fraction is now being used up and the epidemic dies out. What we observed is the infected class. We don't really have a measure of the susceptible people and you can measure the recovered people, but it's difficult.
一般来说,这是医院报告的病例。这就是我们必须依赖的数据。我们可以拟合参数,然后就能理解传染性流行病背后的一些机制。
Generally, this is hospitals reporting cases. That's what we have to work with. We can fit parameters and then we can understand some of the mechanisms behind the infectious epidemic.
我只想指出一点,数学建模者始终面临一个权衡,实际上在任何领域都是如此。这两个都是汽车模型,都是玩具车,但这个模型非常复杂,细节很多。这个模型只有汽车最基本的特征,有几个能转的轮子和一个车架。
I just want to point out here that there's this perpetual trade-off that mathematical modelers face, in any field really. So both of these are models of cars. They're both toy cars, but this one is very complicated and there's a lot of detail. This one just has the very bare bones features of what a car is. It has some wheels that go around and it has a frame.
我之前展示的那个简单模型就跟这个一样。它是透明的,你的假设被清晰地编码在其中,你也知道驱动流行病的机制是什么。我们可以通过增加模拟细节来添加更多细节。我们可以把人放入不同的隔间,引入出生机制,考虑不同的年龄结构,让人群流动起来。
The simple model that I showed you is like this. It's transparent. Your assumptions are very clearly encoded, and you know what the mechanisms are driving the epidemic. We can add detail by increasing the simulation detail. We can put people in different compartments, have births in there, have different age structure, have people moving around.
美国有一些针对流感的模拟,其中甚至精确编码了人们早上遛狗、去上班、吃什么等等所有这类事情。问题是,你是否真的关心所有这些额外的细节?你是否愿意为了加入这些东西而牺牲透明度和普适性?
There are simulations for flu in the States where people are literally encoded walking their dogs in the morning, going to work, what do they eat, all of this kind of thing. The question is, do you care about all of that extra detail? Do you want to give up the transparency and generality by having all of that stuff in there?
一般来说,最简单的流行病模型往往能提供最深刻的洞见。简单的模型能引导我们找到一些有趣的阈值参数。基本再生数(basic reproduction number)是一个非常通用的概念:如果平均每个初始感染者感染的人数超过一人,那么疫情就会扩散;反之,疫情就会消亡。
Generally, the epidemic models that are simplest often provide the most insight. The simple models can lead us to some interesting threshold parameters. The basic reproduction number is this idea that's very general, that if, on average, the first infected person infects more than one other person, then your epidemic will spread. And if they don't, then your epidemic will die out.
这里有一个阈值参数 R0,这是我们试图为一场流行病测量的标准参数之一,用来判断其进一步传播的潜力。当然,这个概念在互联网梗传播、产品采用以及许多其他情境中也有类似的应用。
There's this threshold parameter R-naught, that's one of the stock parameters that we try and measure for an epidemic to figure out the potential for onward spread. Of course, this has analogies in Internet meme spread or product adoption and many other situations.
卡罗琳·布基:理解疾病传播的新方法(续)
Caroline Buckee Novel Ways to Understand How Disease Spreads (Continued)
我们要思考的是,R0 的组成部分是什么?我们如何测量它们?是什么因素决定了某种疾病会不会成为流行病?我们如何利用这些来预测某种疾病的传播方式?
We want to try and think about what are the components of R-naught? How can we measure them and what are they that make something epidemic or not? How can we use that to predict how something's going to spread?
第一个参数是,如果有人在我身上咳嗽,而他携带病毒,那么我被感染的概率是多少?这实际上非常难以测量,你必须了解病原体的一些内在生物学特性。
This first parameter is, OK, if somebody coughs on me and they have a virus, what's the probability that I become infected? That's actually very difficult to measure and you have to know something about the inherent biology of the pathogen.
接下来是传染期。假设发生了接触,传播概率会提高这个流动速率。传染期决定了这个隔间(感染人群)减少的速度。对于大多数病原体,我们都可以很好地测量这个参数,只要我们有足够大的样本量,并且知道要寻找什么。
We have the duration of infectiousness, so the probability of transmission given contact is going to increase this rate of flow. The duration of infectiousness determines how rapidly this compartment is diminished. We can measure this for most pathogens pretty well, once we have a reasonable sample size and we know what we're looking for.
然后,我们有一个关键参数:有效的接触概率。这个数据既依赖于感染者和易感者的数量,也依赖于他们之间接触的频率。
Then, we have this critical parameter, the probability of infectious contact. So that data relies on both infected people and susceptible people and how often they mix.
一般来说,政策制定者会问我们,而我们也会思考的问题是:如何将 R0 降到 1 以下,从而阻止一场流行病?我们如何进行干预?我们可以通过为人们接种疫苗来干预有效的接触率。这当然是经典的公共卫生措施之一,对我们的寿命产生了巨大的影响。
Generally, what we are asked by policymakers and what we think about, is how can we reduce R-naught below one, stop an epidemic? How can we intervene? We can intervene on the infectious contact rate by vaccinating people. This is of course, being one of the classic, public health measures that's made a huge difference to our life spans.
这是通过模型模拟的,最简单的模型就是直接把易感人群从易感隔间移动到康复隔间。有时,如果疫苗不是完美有效的,它不一定能把人直接转移到康复隔间,但可以降低接触后的传播概率,这同样有用。
This is modeled and the simplest model is just by moving people from the susceptible compartment straight to recovered. Now, sometimes, if it's an imperfect vaccine, it doesn't necessarily move them to the recovered compartment, but it can reduce the probability of transmission given contact. That's also useful.
我们也可以实施隔离,有时如果症状严重到让人待在家里,这也会自然发生,他们与其他人的接触率自然就会降低。我们可以通过教育计划和其他方式来减少易感人群和感染人群之间的接触。然后,我们可以治疗病人。
We can also quarantine people and this happens naturally sometimes if the symptoms are bad enough that you stay home. Your contact rate with other people is going to reduce anyway. We can use education programs and other ways to reduce that contact between our susceptible and infected populations. Then, we can treat people.
如果我们能识别出传染源人群并对其进行治疗,就可以减少这一类人群的数量,从而减缓疫情的扩散。
If we can identify the infectious reservoir and treat them, we can reduce the number of people in that category and we can therefore reduce the spread of the epidemic.
在使用这类基本模型来理解感染传播时,我们面临的最大问题之一是,缺乏关于人们如何移动的数据,因此也就无法了解这些接触率在空间上是如何变化的。
One of the biggest problems that we have in using these types of basic models to understand the spread of infection is that we are lacking data on how people move around and therefore, how these contact rates vary spatially.
所以,我们来看这项研究,它展示了感染的空间和时间决定因素。这里,空间尺度从社区级别开始,一直到城市内部、区域,最后是国际范围。这里则是时间尺度,包括每日移动、周期性移动、季节性移动——这在撒哈拉以南非洲非常重要——以及长期移动和迁移。
So we look here at this study showing the spatial and temporal determinants of infection. Here we have spatial scale from the neighborhood scale up to within a city, regional, and then all the way up to international. And here we have temporal scale. These are daily movements, periodic movements, seasonal movements, very important in sub-Saharan Africa, long term and migration.
它们对感染传播的影响,从仅在当地社区内部让人接触,一直延伸到全球传播和大流行病。但在这两者之间,还有区域传播以及所有支撑流行病在一个国家和一个地区传播的元种群动态。
The impact that they have on the spread of infection ranges from just exposing people within their local communities, all the way up to global spread and pandemics. But in between here, we have regional spread and all of the meta-population dynamics that underlie epidemic spread through a country and a region.
那么,数据缺口在哪里?我们有来自小规模调查研究的数据,比如你去一个村庄,问:“你昨天去了哪里?上周去了哪里?去那里的频率是多少?”诸如此类。这些数据可以告诉我们一些关于人们在社区中的日常暴露情况。
So, where's the data gap? We have data from small-scale survey studies where you go out to a village and you say, "Where did you go yesterday? Where did you go last week? How often do you go there?" and these kinds of things. That can tell us something about the daily exposure people have in their communities.
卡罗琳·布基:理解疾病传播的新方法(续)
Caroline Buckee Novel Ways to Understand How Disease Spreads (Continued)
我们也可以给人们戴上小型 GPS 设备,然后观察他们每天去哪里。有人这样做,但范围有限,遗憾的是,你没办法给几百万人戴上 GPS。我倒是很想这么做。
We can also strap little GPSs to people and then see where they go every day. People do that, but it's limited in scope, so you can't really strap GPSs to millions of people, unfortunately. I would love to.
[laughter]
[laughter]
我们有普查数据和微观普查数据,这些数据告诉我们关于长期移动和迁移的信息。通常的问题是:“去年这个时候你住在哪里?”这没问题,但时间分辨率确实达不到我们理解疾病所需的数量级。
We have census data and we have micro census data, and these tell us about long-term movements and migration. Generally, the question is, "Where did you live this time last year?" That's fine, but really the temporal resolution is not on the order of magnitude that we need to understand disease.
我们还有航空和航运网络。我们通过航空路线大致了解全球连通性,但当然,这些人群与驱动区域层面动态的人群会有所不同。
And then we have airline and shipping networks. We know kind of global connectivity through airline routes, but of course, those populations are going to be somewhat different than the populations driving dynamics at the regional level.
大数据登场了,我们都很兴奋。你可能见过这个图表的某些变种。这是过去十年移动电话用户数,以十亿计,当然,增长最大的部分是在发展中国家。
In comes big data and we're all really excited. You've probably seen some variation on this graph. This is the number of mobile phone subscriptions over the last decade, in the billions, and of course the biggest increases have been in the developing countries.
我们现在处于什么位置?突然间,我们在那些以前只能开着路虎、带着社区卫生工作者团队才能接触到人的地区,拥有了数十亿人的传感器。
Where are we? Suddenly, we have billions of sensors on a lot of people in regions where previously, we would only be able to go and find out about those people if we got into a Land Rover and we drove there and we had teams of community health workers and so forth.
我们自己的研究显示,这些用户在农村和低收入地区的渗透率惊人。我认为,很快,移动电话数据集就能提供真正令人惊叹的代表性样本。
What we're seeing from our own research is that the penetration of these subscribers in rural places and low-income settings is astonishing. I think that soon, we will have a really amazing representative sample among mobile phone data sets.
每当你打一个电话或发一条短信,移动运营商就会记录下你的电话号码、通话对象和基站 ID。如果你知道那个基站所在的经纬度,那么你就知道了那个人在那个时间点的位置。更具体地说,你知道了他们 SIM 卡的位置,这两者不总是一回事。
Every time you make a call or a text, the mobile phone operator logs your phone number, the person you called, and a cell tower ID. If you know the latitude and longitude for that cell tower, then you have a location for that person at that time. More specifically, you have a location for their SIM card. That's not always the same thing.
不过,现在我们想象一下,我们有了每个人随时间变化的这些记录。这是我们的一个人,在 A 基站打了些电话——我们知道 A 基站的位置。然后这个人移动到了 B 基站,在那里打电话,我们就能推断出地点之间的移动。最后,他们移动到了 C 基站。
But let's imagine now we have these records over time for each person. Here's our person, making some calls at Tower A – we know where Tower A is. Then that person moves to Tower B, makes calls there, we can infer movement between locations. Finally, they move to Tower C.
我们有了一个人位置随时间变化的纵向记录。这些日志由移动运营商定期保存,用于客户流失分析等目的。它们为我们提供了有史以来关于人口动态和位置的最惊人的数据集之一。
We have a longitudinal record of that person's location over time. These logs are kept routinely by mobile operators for churn analysis and other things. They provide one of the most amazing data sets for the dynamics and location of populations that we've had to work with ever.
如果你对这些地方的人口有所了解——我们主要使用卫星数据和其他类型的遥感信息——我们可以从中描绘出人口分布和密度,然后就可以建立模型,模拟每个子人群之间人口流动的流量和动态。
If you know something about the populations in those places – we use mainly satellite data, other types of remote sensing information – we can delineate population distributions and densities from them, and then we can build up models of the flow and dynamics of human movement between each sub-population.
这非常强大,尤其是考虑到……人们总对我说:“这很好,但你不知道它有多大偏差。”是的,但可能比零强,我们以前可就是零。
That's incredibly powerful, especially given the fact that...People always say to me, "That's good, but you don't really know how biased it is." Yes, it's probably better than zero, which is what we had before.
实际上,我们正在努力研究如何调整偏差和所有权问题,思考如何搞清楚我们在这类数据中捕获的是哪些人,尤其是儿童的移动显然非常重要。
We're working hard, actually, to figure out how to adjust for biases and ownership and think about ways that we can figure out who it is that we're capturing in these types of data and especially, the movement of children is obviously going to be important.
这真是一个了不起的数据集。突然间,我们拥有了整个领域。我们为整个领域找到了一个惊人的信息来源。我们可以获得人们每天甚至每小时的定位。至少在城市里,我们的分辨率可以达到一个街区的级别。我们知道每个人在哪里,我们知道我们的易感人群在哪里。如果我们有良好的临床数据,我们就可以开始构建真正优秀的模型了。
It's a pretty remarkable data set. Suddenly, we have this whole area. We have an amazing source of information for this whole area. We can get daily or even hourly locations for people. In cities, at least, we can get resolution on the order of a city block. We know where everyone is, we know where our susceptibles are. If we have good clinical data, we can start building really good models.
卡罗琳·布基:理解疾病传播的新方法(续)
Caroline Buckee Novel Ways to Understand How Disease Spreads (Continued)
这真的令人兴奋。我的大多数同行都对这一机遇感到非常兴奋。我们用它做了些什么呢?这里是肯尼亚。这张地图显示了肯尼亚的道路网络。你可以看到,道路网络与人口密度模式相当吻合。肯尼亚北部人口稀少,而中部和西部地区以及沿海地区人口相当密集。
This is really exciting. Most of my community is really excited about this opportunity. What have we been doing with this? Here's Kenya. This is a map that shows the road networks in Kenya. You can see they follow population density patterns pretty well. Kenya is sparsely populated in the North and pretty densely populated in the Central and Western regions and on the coast.
我在道路网络之上叠加了颜色,这里的颜色代表平均流动性测量值。
What I've mapped on top, the color on top of the road network here is an average measure of mobility.
高亮部分,红色代表出行频繁且距离远的人群,蓝色代表出行较少的人群。你可以看到,这些贫困地区的人们不得不走很远的路。在发展规划领域,关于农村地区出行负担以及偏远程度如何影响人们获取市场和医疗服务等问题一直存在长期争论。
High, the red, is people that travel a lot and they travel far. The blue is people that don't travel very much. What you can see is that people in these poor regions have to travel pretty far. There's been a long debate in the development world about the burden of travel in rural areas and the impact of remoteness on people's ability to access markets and healthcare and things like that.
现在,我们实际上可以观察这些人。我们可以验证这是否属实。农村地区的人们是否真的出行更远?看起来确实如此,尤其是当你将其与其他类型的信息结合时。
Now, we can actually observe these people. We can see whether it's true. Is it true that people in rural communities travel further? It seems like it is, especially if you can couple it with other types of information.
我们做的事情是想看看这些平均汇总指标是否与医疗保健有任何关联。我们采用了目前最标准的方法:在地图上标注道路,并计算到诊所的平均时间。这里是肯尼亚西部的一个区域。你可以看到,有些中心的就诊时间较短,而周边区域获取医疗服务的难度则稍大一些。
What we did is we wanted to see whether these kind of average aggregate measures had any bearing on healthcare. We took the standard state of the art measures where you map onto the roads, the average time to a health clinic. Here, this is this area in Western Kenya. You can see that there are centers where the travel time to health clinic is short, and then there's areas around them where it's a little bit harder to get access to healthcare.
我们想问的问题是:在预测健康结果方面,我们的估算是否比这些数据更好?这里是我们基于手机数据得出的估算结果,你可以看到,即使在那些到诊所的出行时间非常相似的偏远地区,人们的平均出行行为也存在相当大的差异。
We wanted to ask the question, are our estimates better than this for predicting health outcomes? Here are our estimates from the mobile phone data and you can see that even in places that have quite similar remoteness in terms of their travel times to health clinics, we have quite heterogeneous average behavior in terms of mobility.
在某种意义上,这捕捉到的不仅仅是出行时间这样一个静态指标。我们关注的是动态,而出行活动很可能更能反映人们的基本行为。
In some sense, this is capturing more than simply a travel time, a static measure. We're looking at dynamics, and we're looking at something mobility probably captures more about the fundamental behavior of people.
真正有趣的是,我们的指标在预测哪些家庭缺少产前护理和儿童疫苗接种方面显著更优。我们可以将这些地图叠加到当前的政策指南上,从而确定需要重点推进医疗保健可及性计划的区域。我认为这相当令人兴奋。
What's really interesting is that our measures are a significantly better predictor of which households are missing antenatal care and childhood immunizations. We can layer these kinds of maps onto current policy guidelines to figure out where you need to target your healthcare access initiatives. I think that's pretty exciting.
我们还可以考察季节性因素。显然,在低收入环境中,人们的出行方式与这里截然不同。我们拥有大量关于通勤和节假日出行的数据。长期以来有一种教条认为,学校放假和学期设置基本上决定了季节性流感或麻疹的传播规律。在低收入环境中,情况则不那么明确,因为存在大量农业迁移,人们出行的原因也各不相同。
We can also look at seasonality. People obviously travel very differently in low income settings compared to here. We have a ton of data here about commuting and travel around holidays. There's been this kind of dogma that school closures and school terms basically define our flu seasonality and/or measles seasonality. In low income settings, it's much less clear, because there's a lot of agricultural migration and there are different reasons for people moving around.
这里再次提到肯尼亚。我只是展示一个流动指标——人口移动。这是圣诞节期间的数据,由于我们的数据集,比例有些奇怪。圣诞节期间有大量人口移动,随后还有两个高峰。
This, again, is Kenya. I'm just showing a measure of flux, population movement. This is Christmas, and it's a weird scale because of our data set. Lots of movement around Christmas and then two more peaks.
我们与普林斯顿大学的研究人员合作,他们试图判断肯尼亚是否应该推行风疹疫苗接种。我们与他们讨论,他们表示:“我们无法解释肯尼亚风疹疫情出现的三个高峰。我们不知道原因。”
Working together with the researchers at Princeton who are trying to figure out whether Kenya should institute rubella vaccination, we're talking to them and they were like, "We can't figure out these three peaks of rubella that happen in Kenya. We don't know why."
这里是风疹疫情的数据,你可以看到持续出现的三个高峰。结果表明,我们基于出行活动的估算数据在预测这些风疹高峰方面,比学校学期数据或降雨与农业数据更准确。因为我们有了这种直接的实地测量数据,而不是依赖于学校学期或出行时间等替代指标,我们就可以真正开始理解其机制,以及驱动这类疾病的动态过程。
Here's rubella, and you can see these three peaks that keep happening. It turns out that our mobility estimates are better at predicting those rubella peaks than either school terms or rainfall and agriculture. Because we have this actual direct measure rather than proxies like school terms or travel times, we can start actually understanding the mechanism and understanding the dynamics that drive these types of diseases.
卡罗琳·巴奇:理解疾病传播的新方法(续)
Caroline Buckee Novel Ways to Understand How Disease Spreads (Continued)
这些是对整体出行指标的大规模流动观察。我们现在也可以将我们 SIR 模型框架应用于这类数据集。这是一张元种群模型(meta-population model)的示意图。同样,它只是圆圈,因为在我看来一切都可以这样表示。它代表的是这些相互关联、风险各异的人群。
Those are large scale flux in looking at aggregate mobility measures. We can now bring our SIR-type frame works to bear on this kind of data set as well. This is a picture of a meta-population model. Again, it's just circles, because that's how everything looks to me. It's meant to be these linked populations with different risk.
这是一个高风险人群,这些是低风险人群。他们之间存在我们可以测量的流动。那么我们要问的是:从这个高风险地区到我们的低风险居住区会有多少输入性感染?当低风险居住区的人们度假或访问这个高风险地区时,他们感染疾病并将其带回的频率有多高?
This is a high-risk population, and these are low-risk populations. We have movements between them that we can measure. Then what we want to ask is, how much imported infection is there going to be from this place to our low-risk settlement? And when our low-risk settlement guys go on holiday or they visit this high risk place, how often are they going to get something and bring it back?
在这些人群中的每一个群体内部,我们希望能够在最基本层面上参数化一个小型 SIR 模型,并在我们的感染项上添加来自另一人群的恒定输入风险。这样,我们就可以构建出由出行活动连接的网络,并使用手机数据来参数化这些流动。
Within each of these populations, we want to be able to parameterize a little SIR model at its most basic and add onto our infected term a constant risk of importation from the other one. So we can build up these networks connected by mobility and use the mobile phone data to parameterize the movements.
我们有一个高风险和一个低风险人群,我们想知道会发生什么。过去,我们只能依赖基于标准医院报告要求、人口普查数据估算等得出的粗略估计,而现在,我们有机会通过症状监测(syndromic surveillance)系统收集感染的实时数据——这基本上就是像“附近流感”(Flu Near You)、登革热这类新型手机应用,你可以与用户互动并询问他们的症状。我们如今也拥有出色的气候数据。
We have a high risk and a low risk population, and we want to know what's going to happen. Whereas before we had very crude estimates based on standard hospital reporting requirements and estimates from census data and things like that, now we have the opportunity to collect real-time data on infection rates using syndromic surveillance – that’s basically Flu Near You and dengue, these types of new mobile apps where you can engage with people and ask them about their symptoms. We have great climate data now.
对于任何受环境因素影响的疾病,我们都有海量数据可以进行整合以调整我们的风险评估。
If anything that's environmentally forced, we have huge amounts of data that we can integrate to adjust our risk.
此外,我们还有一些遥感技术,可以让我们非常精确地估算居住区的边界、面积和人口密度。我们想问传播速度的问题。疾病会传播到这里吗?传播速度有多快?我们能做些什么?如何有针对性地配置资源?成本是多少?
Then we have some remote sensing technologies that allow us to get really good estimates of the boundaries of settlements and their sizes and densities. We want to ask the rate of spread. Is it going to come to this place? How quickly is it going to come to this place? What can we do about it? How can we target our resources? How much is it going to cost?
现在,参数化这些模型的方法,就是利用我们掌握的这些关于人们去向以及他们如何交往的庞大数据集。另一个正在深刻改变该领域的大数据是寄生虫基因组数据和病原体基因组数据。我们还拥有所有这些病毒等的序列数据。我们可以判断这些序列与那些序列的相关性。我们能否也通过这种方式测量输入性感染?所有这些新的数据集都需要通过严谨的方式进行整合,才能得出一些可靠的估算并做出更好的预测。
The way to parameterize those models now is through these huge data sets that we have on where people are going and how they're mixing. Another kind of big data thing that's really transforming the field is also parasite genomic data, pathogen genomic data. We also have sequences from all of these viruses and things. We can say how related are the sequences here to these ones. Can we measure importation that way as well? All of these new data sets need to be integrated in a rigorous way to come up with some of these estimates and to make better predictions.
有趣的是,我们的手机数据在预测人口密度和分布方面似乎也略胜一筹。它与我们关于每个人所处位置的标准估算数据相关性很好。尤其是在像巴基斯坦这样的地方,它在统计那些已知在当前数据源中缺失的人群时,表现得相当出色。
What's interesting is that our mobile phone data seems to also be slightly better at predicting population densities and distributions. It correlates very well with our standard estimates of where everyone is. And especially in places like Pakistan, it seems to do a really good job of counting people that are known to be missing from the current data sources.
我现在要讲两个例子。我的初恋是疟疾。
I'm going to now talk about two examples. My first love is malaria.
[laughter]
[laughter]
它是一种非常有趣的真核寄生虫,通过蚊子传播。它已经伴随我们很久很久了。有人认为图坦卡蒙法老死于疟疾,拜伦勋爵也死于疟疾。它长期困扰着人类,至今仍每分钟夺走一个孩子的生命。
It's a very interesting eukaryotic parasite. It's spread by mosquitoes. It's been with us for many, many years. They think that King Tutankhamun died of malaria. Lord Byron died of malaria. It's been with us for a long time, and it still kills a child every minute.
可悲的事实是,它完全可以治疗,也完全可以预防。每一例死亡都是不必要的。我们知道如何解决这个问题。我们确实了解疟疾的控制方法。我们有了一些工具,虽然它们并不完美。但我们依然承受着这种负担,我认为这越来越不可接受。
The sad fact is that it's completely treatable and it's completely preventable. Every single one of those deaths is unnecessary. And we know how to fix this. We do understand malaria control. We have some tools, they're not perfect. We still have this burden, and I think it's increasingly unacceptable.
卡罗琳·巴奇:理解疾病传播的新方法(续)
Caroline Buckee Novel Ways to Understand How Disease Spreads (Continued)
如果你感染了疟疾,那会非常痛苦。你会出现发烧、寒战、剧烈头痛,肌肉和骨骼也会疼痛。我们在诊所见到的儿童死于严重的疟疾性贫血,这往往与营养不良相伴。他们死于呼吸窘迫和昏迷。这个小女孩曾陷入昏迷。发生的情况是,寄生虫被隔离起来,卡在毛细血管中,然后导致代谢性窘迫而死亡。
If you have malaria, it's pretty horrible. You'll get fever and chills, very severe headaches. Your muscles and bones ache. Children that we see in the clinic die of severe malarial anemia, often coupled with malnutrition. They die of respiratory distress, and they die of coma. This little girl had a coma. What happens is the parasites get sequestered away, stuck in capillaries, and then they go into metabolic distress and die.
疟疾影响的社区是地球上最贫穷、最偏远的社区。他们难以获得医疗服务。营养不良是一个大问题,合并感染也是一个重大问题。在许多案例中已经证明,如果一个社区能够根除疟疾,那么其整体健康、教育和经济机会都会立即得到改善。
The communities that malaria affects are the poorest, most remote communities on the planet. They have poor access to healthcare. Malnutrition is a big problem. Co-infections are a big problem. It's been shown in many cases that if you can eliminate malaria from a community, you instantly improve overall health, education, and the economic opportunities.
我认为这是一个迫切需要解决的问题。在过去一百年里,我们取得了一些进展。这是 1900 年疟疾的分布地图,撒哈拉以南非洲地区,但一直延伸到美国全境——我们这里曾有很多疟疾——然后进入欧洲、地中海地区,并横跨直至澳大利亚北部。
I just think it's a really urgent problem to be worked on. We have made some progress over the last 100 years. This is the map of the extent of malaria in 1900, sub-Saharan Africa, but all the way up throughout the U.S., we had quite a lot of malaria here, and into Europe, Mediterranean, and all across into Northern Australia.
到了 2007 年,这张地图的范围已经大大缩小。疟疾如今仍然是一个热带和真正意义上的低亚热带问题。撒哈拉以南非洲仍然是个大问题,南亚、东南亚和南美洲也是,但当然,我们已经在美国和澳大利亚根除了它。这在很大程度上归功于 20 世纪 50 年代启动的全球疟疾根除计划。
In 2007, we've now shrunk that map considerably. Malaria now remains a problem of the tropics and really lower sub-tropics. Sub-Saharan Africa is still a big problem and South Asia, Southeast Asia, and South America, but of course, we eliminated it from the U.S. and from Australia. That's in large part due to a global malarial eradication program in the 1950s.
这来自一本美国教科书。当时他们到处喷洒 DDT,效果非常好。但到了 60 年代,情况基本是资金耗尽了。他们认为在非洲实施根除计划并不可行。DDT 产生了抗药性,还存在各种后勤问题,计划完全停滞,最终他们放弃了。
This is from a U.S. textbook. They sprayed DDT everywhere. It worked really well. But in the 60's, basically what happened was money ran out. They didn't really consider it viable in Africa. There was resistance to DDT and there are all kinds of logistical problems and the program was totally stalled and they gave up.
此后的几十年里,基本上所有的重点都放在了控制上——控制发病率、控制死亡率,尽我们所能提供治疗和医疗服务,防止儿童死亡。但至于传播,我们无法根除。
For decades after that, basically, all of the emphasis was on control – control morbidity, control mortality, do what we can to treat and provide healthcare for people to stop children dying. But forget about transmission, we can't eliminate it.
然后,2007 年发生了这件事。比尔和梅琳达·盖茨敢于说出“根除”这个词,所有人都群起反对。但通过将这个议题重新提上议程,确实改变了疟疾学界思考这个问题的方式。如果我们回到我们的 SIR 模型问题上,你会看到基本模型几乎仍然完全相同,只做了一些修改。
Then, this happened in 2007, Bill and Melinda Gates had the audacity to say the word “eradication” and everyone was up in arms. But by putting it back on the agenda, it's really shifted how the malaria community thinks about this problem. If we go back to our SIR problem, you'll see that the basic model is still almost exactly the same with a few modifications.
事实上,这个模型在第一次根除计划中就被成功运用。现在,我们再次拥有我们的目标人群,他们可能易感,也可能已被感染。我们没有恢复期类别,因为对于疟疾,即使你对疾病症状产生了免疫力,你也不会真正对感染产生免疫。因此,在最简单的模型中,你只是在易感和感染状态之间来回切换。
In fact, this model was used to great success in the first eradication program. Now, again, we have our people, they can be susceptible and they can be infected. We have no recovered class because with malaria, you never really become immune to infection even though you become immune to disease symptoms. So you just shuttle back and forth between susceptible and infected in the most simple model.
然后,我们还必须对蚊子进行建模。蚊子也可能处于易感状态。它们有一个潜伏期,这被证明非常重要,之后它们便具有传染性。这个潜伏期接近蚊子的寿命,这正是病媒控制效果如此之好的原因,因为它能迅速杀死它们。
Then, we have to model mosquitoes also. Mosquitoes can also be susceptible. They can have a latent period which turns out to be very important, and then they're infectious. This latent period is close to the lifespan of a mosquito which is why vector control works so well because it kills them quickly.
那么,当然,这两组方程是通过人与受感染蚊子之间的接触联系起来的。因此,叮咬频率、我们自身的传染性程度,以及蚊子和人类的相对密度,都是关键因素。
Then, of course, these two sets of equations are linked by the contact between people and infected mosquitoes. So, feeding rates and how inherently infectious we are in the relative density of mosquitoes and humans.
我们如何干预?没有疫苗,而且你们最近可能读到的刚刚推出的一款疫苗,在阻断传播方面效果相当糟糕。它对重症似乎有一定效果,这倒是不错。而使用经杀虫剂处理的蚊帐,我们不仅能够阻断
How can we intervene? There's no vaccine, and the vaccine that you may have read about that's been rolled out recently is a pretty crappy vaccine from the perspective of transmission blocking. It seems to have an impact on severe disease, which is good. With insecticide-treated nets, we not only block
卡罗琳·巴基:理解疾病传播的新方法(续)
Caroline Buckee Novel Ways to Understand How Disease Spreads (Continued)
蚊子无法叮咬人,从而降低了接触率,同时因为杀虫剂的作用,我们还能提高蚊子的死亡率。这是一件好事。
mosquitoes from feeding on people, thus reducing the contact rate, we also increase the death rate because there's insecticide in there. That's a good thing to do.
如果我们治疗患者,就会把这些人重新送回这个群体。这可能不会对传播产生巨大影响,但至少能让患者好转。在感染者比例非常低的情况下,这很可能有助于整体的防疫努力。而控制病媒的另一个目标,就是直接降低病媒的密度。
If we treat people, we shove people back into this group. It may not have a huge impact on transmission, but at least people get better. And at very low prevalence of infected people, it's probably helping overall transmission efforts. Then, with vector control, the other goal is to just reduce the density of vectors.
That's straightforward.
That's straightforward.
我们拥有这样一套工具。它们并不完美,但组合使用时,能以非线性方式运作,实际上效果相当不错。问题在于,没有人能对感染完全免疫,而且绝大多数病例根本没有症状,所以你根本不知道他们是否被感染。此外,我们目前拥有的诊断工具灵敏度还不够,无法实际检测到那些密度极低、但仍然能传染给蚊子的感染。
We have this kind of suite of tools. They're not perfect, but in combination, they work in non-linear ways to actually be quite effective. The problem is that nobody is ever immune to infection and the vast majority of cases have no symptoms at all, so you don't know if they're infected. Plus, the diagnostic tools that we have right now aren't sensitive enough to actually measure those very low density infections that are nonetheless infectious to mosquitoes.
如果你想根除疟疾,真的想永远消灭疟疾,你就必须考虑到这些人,因为他们四处游走,传播感染,破坏防控计划,却全然不自知,而我们也很难精准锁定他们。
If you want to eradicate malaria, if you actually want to get rid of malaria forever, you have to think about these people because they are the ones that are traveling around spreading the infection and undermining control programs without even knowing it, and we can't target them very easily.
我们对此采取了什么措施?这又是一个肯尼亚的例子。背景中可以看到疟疾风险区域。维多利亚湖周边是红色高风险区,沿海地区也是高风险区。肯尼亚中部风险相对较低,但仍存在一些传播。
What are we doing about this? This is Kenya again. In the background, you can see the malaria risk. The red is high risk around Lake Victoria. We have high risk on the coast. And then we have fairly low risk in the middle of Kenya, although some transmission.
背景中这些灰色团块,代表我们通过卫星估算出的人口分布和位置。黑色的和蓝色的点,则是手机信号塔的位置。你可以再次看到,信号塔的分布是跟随人口密度的。我们的定位估算在内罗毕这样的地方效果最好,那里信号塔密度非常高;而在北部人口稀少的区域,效果就没那么好了。
These gray blobs in the background represent our satellite estimates of the distribution and location of populations. Then, the black and blue dots are the location of cell towers. You can see again that the cell towers follow population density. Our location estimates are best in places like Nairobi where we have a very high density of cell towers, and they're less good up here in the north where there aren't that many people.
我们有办法对此进行调整,但这确实是个问题。无论如何,我们掌握了所有手机数据,而我们想问的是,应该把干预措施集中在哪些地方?首先…… 肯尼亚没有太多资金,但他们想要彻底消除,那么,我们如何高效地做到这一点?需要多长时间?这对成本以及彻底消除是否可行至关重要。
We have ways for adjusting for this, but that is a concern. In any case, we have all of our mobile phone data and what we want to ask is, where do we focus interventions? For a start...Kenya doesn't have a lot of money, but they want to eliminate and so, how do we do that efficiently? How long will it take? That's critical to the cost and whether it's even feasible to consider elimination.
如果耐药性蔓延到沿海,它的传播速度会有多快?我待会儿会给各位看一张幻灯片,但所有疟疾药物的耐药性都是从东南亚出现,然后传播到东非,再通过人类旅行扩散到整个非洲大陆。目前,只剩下一种药物还没有出现耐药性,而且它是重症疟疾的一线疗法。
If drug resistance reaches the coast, how quickly will it spread? I'll show you a slide in a minute, but all drug resistance to malaria drugs has emerged in Southeast Asia, spread to East Africa, and then spread across the continent through human travel. At the moment, there's one drug left to which there is no resistance and it's the first-line therapy for severe malaria.
现在东南亚已出现耐药性。如果这种耐药性蔓延到非洲,我们就真的麻烦大了,儿童死亡率会急剧上升。这是一个切实的担忧,我们希望有能力预测这种情况将如何发生,以及如果我们没能彻底消除疟疾,哪些地区的疾病复发风险最高。这也是一个非常重要的问题,因为在上一次的根除计划中,计划失败后,许多地方随后爆发了大面积的疫情。
There's drug resistance now emerging in Southeast Asia. If it reaches Africa, we're in real trouble, and there' s going to be a huge amount of mortality among children. This is a real concern and we want to be able to predict how that's going to happen and which regions are at the highest risk for resurgence of disease if we do fail at elimination. That's an important problem, too, because in the last eradication program, when it failed, a bunch of places had massive epidemics afterwards.
这些正是我们想要回答的问题。我们所做的是建立数学模型,整合所有分层数据,然后模拟寄生虫在流动人群中的传播路径。我们进行了一种“源-汇”分析:人们从哪里来、要到哪里去,寄生虫又从哪里来、要往哪里去?这些都是不同维度的问题。但你看左侧这张图,人类出行的源头分散在这些高密度地区周围。内罗毕始终是一个旅行汇聚地,流入首都的人口数量惊人——它是全国各地的客流汇集终点。
These are the kinds of questions that we want to be able to answer. What we've done is we've built up our mathematical models. We've taken all of our layers of data and then we've modeled the spread of parasites in people traveling around. We've done a kind of a source-sink analysis: Where are people coming from and where are they going to and where are parasites coming from and where they are going to? Those are the different questions. But if you look here on the left, you can see that the source of human travel is kind of scattered around these high density places. Nairobi remains a sink for travel, so the sheer volume of people traveling into the capital is amazing. It's a sink from all across the country.
卡罗琳·巴奇:理解疾病传播的新方法(续)
Caroline Buckee Novel Ways to Understand How Disease Spreads (Continued)
如果你观察这些寄生虫——这相当于把数学模型映射到我们的移动矩阵上——可以看到,肯尼亚绝大多数的疟疾来自一个地区,然后从这里向外扩散。
If you look at the parasites – so this is kind of the mathematical model imposed on our mobility matrices – what you can see is the vast majority of malaria in Kenya is coming from one region and spreading this
方式。内罗毕是一个巨大的耗资池,同时也是一个传播率低的高原地区——这些地区仍然有蚊子。
way. Nairobi is a huge sink, but also this highland region and these areas that have low transmission, but they still have mosquitoes.
这提示我们,与其去摘那些低垂的果实——不断清理这些输入型的感染——不如从根本上解决这个问题,把所有干预措施集中在那里,这样一来问题可能就自行消失了。这实际上改变了我们对资源分配的思考方式,并且直接预测了你可以如何影响传播链。它还提供了一些具体的政策指导。
What this suggests is that instead of going for low-hanging fruit, where you're continually mopping up these imported infections, you've got to address this problem, put all your interventions there, and this may well just go away. It kind of changes how we think about resource allocation, and it provides direct predictions about how you can impact transmission. It provides some kind of specific policy guidelines.
在内罗毕,由于估算的粒度非常精细,我们可以更进一步追问:本地传播的可能热点在哪里。在内罗毕,历史上一直被认为完全没有本地传播,因为天气太冷,尽管我们知道蚊子确实存在,而且内罗毕在 20 世纪初曾有过疟疾。
In Nairobi, because the estimates are so finely grained, we can go a step further and we can ask, where are the likely hot spots of local transmission. So in Nairobi, it's been historically thought that there's no local transmission at all because it's too cold, although we know that the mosquitoes are there and Nairobi used to have malaria back in the early 1900s.
利用我们的模型,我们所做的是选取了这些代表城市各医院的数据点,并将其临床病例与模型预测的病例进行了对比。
Using our model, what we did is we took these blobs, which represent hospitals around the city, and we compared their clinical cases to our predicted cases from our modeling exercise.
色块的颜色代表实际输入性感染人数与预测输入性感染人数之比。这里的概念是,如果你预测基本没有感染,却看到大量临床病例,那么当你想找出本地传播正在发生的地方时,这些就是你该去查探的区域。
The color of the blob represents the ratio of observed-to-predicted imported infections. And the idea here is that if you predict basically no infections, but you're seeing tons of clinical cases, then when you want to find out where local transmission is happening, those are the places you should go.
如果你真的去看看那些地方的位置,它们就在城市周边的贫民窟里,这也许并不令人意外,但为那些实地调查传播媒介的团队提供了佐证。
If you actually look at where those places are, it's in the slums surrounding the city, which is maybe not surprising, but it gives some evidence to back up teams of people going out there and looking for vectors.
我们正在与肯尼亚疾控中心及该地区的其他研究人员合作,试图弄清楚如何将某些此类方法融入他们的政策指南中。
We're working with the Kenyan CDC and other researchers in the region to try to figure out how to integrate some of these types of approaches into their policy guidelines.
我认为疟疾研究面临的最大问题在于,正如我所说,柬埔寨现已确认对青蒿素类药物出现耐药性。这堪称一场灾难。因此紧急应对行动正在进行。最核心的疑问是:耐药性扩散速度有多快?缅甸改革开放对耐药性的出现有何影响?下一步会蔓延至何处?我们应在哪些区域重点监控?
I think the biggest problem for malaria research is this, as I said. Cambodia now, we have confirmed drug resistance to artemisinin-based drugs. That's kind of a disaster. So there's a lot of emergency action going on. One of the biggest questions is, how quickly is it spreading? What is the impact of Myanmar opening up on the emergence of drug resistance, and where will it go next? Where should we be looking for it?
我们知道它已经蔓延到了孟加拉国。由于孟加拉国和印度的人口密度极高,且流动劳工频繁往返于中东和非洲,这是一个巨大的问题。说到预测,他们或许应该举办一场预测比赛,专门预测耐药性的传播范围,但目前他们正手忙脚乱地应对一切。我认为这是一个非常紧迫的问题。
We know it's got as far as Bangladesh. And because of the high population densities in Bangladesh and India and the frequency of travel to the Middle East and Africa of migrant laborers, this is a huge problem. When we talk about prediction, they should probably have a forecasting tournament for predicting the spread of drug resistance, but they're scrambling to do anything right now. This I think is a very urgent problem.
这是一个模型,由我的合作者安迪·塔特姆建立,用来预测我们认为寄生虫从该地区扩散到哪里。结果显示,它们几乎无处不在。
Here, this is just a model that my collaborator Andy Tatem put together to look at where we think the parasites are going out of that region. It's basically everywhere.
最后,我要把话题从困扰人类数千年的固有难题转向新发疫情这个问题。我将以登革热为例来说明,因为其底层模型本质上与疟疾相似,只是参数不同。我们对登革热的了解不如疟疾那么深入,而且两者存在一些关键差异——我认为这些差异正是华盛顿方面对此格外担忧的根源。
Now, I'm going to end by shifting from a problem that's been endemic for thousands of years in human populations to this issue of emerging epidemics. I'm going to illustrate it with dengue because the underlying model is essentially the same as for malaria with different parameters. We don't know as much about dengue as we do about malaria and there are some key differences that I think underlie how worried people are about this, specifically in Washington.
过去 50 年里,我们看到登革热病例在空间范围和病例数量上都出现了增长。这是一种由不同媒介传播的病毒,与疟原虫的传播方式不同。实际上,它是由携带西尼罗病毒的蚊虫——埃及伊蚊和白纹伊蚊——传播的。这种蚊子在白天叮咬,这意味着蚊帐派不上用场,这就带来了麻烦。
Over the last 50 years, we've seen the emergence of dengue cases, both spatially and in the number of cases. It's a virus that's spread by a different vector than the malaria parasite. It's spread by the West Nile virus-carrying mosquito, aedes aegypti and aedes albopictus, actually. This mosquito bites in the day and that means bed nets are a no-go, which is problematic.
卡罗琳·巴基:理解疾病传播的新方法(续)
Caroline Buckee Novel Ways to Understand How Disease Spreads (Continued)
登革热没有特效疗法。目前有一款疫苗正在临床试验中,但它存在一些令人担忧的问题——比如可能反而加重病情,这显然不是好事。但关键问题在于,这种蚊媒属于城市型传播媒介。这是新加坡,他们正面临登革热困扰,而且问题日益严重。
There is no treatment for dengue. There's a vaccine that's in clinical trials. There are some troubling aspects to it, like it might cause more disease, which is bad. But essentially this vector is an urban vector. This is Singapore and they have problems with dengue, increasing problems with dengue.
因为这些超大城市位于亚洲,人口密度急剧攀升,对病毒的未来传播构成了严峻问题。疟疾是农村病,登革热却是城市病,而且根除病媒极其困难——它基本就在水洼、轮胎和垃圾堆里繁殖,所以几乎不可能消灭。
Because it's urban, the massive increases in population densities in these mega cities across Asia are really problematic in terms of the future of this virus. Whereas malaria is a rural disease, dengue is an urban disease and it's very, very hard to get rid of the vector because it basically breeds in puddles and tires and trash, so it's almost impossible to get rid of it.
他们担心的另一个原因当然是,美国存在这些传播媒介。全球一半人口面临风险……如果登革热蔓延开来,将影响全球一半人口。美国本土有很多地方非常适合登革热传播。因此,我认为登革热引发重大全球问题的可能性确实很大。
The other reason they're worried is because of course, we have these vectors in the States. Half the world is at risk of...If dengue spread, it would be half the world would be affected. There are many places locally that have high suitability for dengue transmission. And so, I think the possibility for a major global problem with regards to dengue is very real.
这是一个例子,说明传染病流行病学为什么可能充满挑战。所以我们正在与 Telenor 集团合作,这是一家大型移动运营商,Telenor 巴基斯坦公司。我们找他们聊,因为 2011 年,就像我在开始时说的那样,这个国家东北部爆发了一场大规模登革热疫情,那里之前从未出现过登革热,他们基本上认为,是轮胎行业把轮胎运到全国各地,同时把蚊子也带了过来。
This is an example of how infectious disease epidemiology can be challenging. So we are working with Telenor Group, it's a large mobile operator, Telenor Pakistan. We talked to them because in 2011, like I said at the beginning, there was this big dengue outbreak in the northeast of the country where they've never seen dengue before, and they think that basically the tire industry transported tires across the country and carried the mosquito with it.
所以现在巴基斯坦到处都有蚊子。那次大暴发导致超过 2 万例临床病例,可能还有更多轻症和无症状病例,主要集中在拉合尔。然后 2013 年又发生了一次暴发。于是我们和巴基斯坦 Telenor 公司沟通,试图证明通话详单(CDR)分析确实能派上大用场。
So now the mosquito is everywhere in Pakistan. There was this huge outbreak that caused over 20,000 clinical cases, so probably many more mild and asymptomatic cases, with a focus in Lahore. And then in 2013 there was another outbreak. And so we were talking to Telenor Pakistan and trying to make the case that CDR [call detail record] analytics could really help.
这里的思路并不太像疟疾防治那样——我们要与政策制定者合作、试图逐步推动政策改变——而是,我们能否相对实时地利用 CDR 来做出可供行动的预测?我们能否与一家移动运营商合作,这家运营商同样有能力接触这些人群,并通过教育信息触达他们——这其实是我们应对这类流行病为数不多的工具之一。
The idea here is not so much like in the case of malaria, where we'd be working with policymakers to try and like shift policy incrementally, but rather the idea here is, can we use CDRs in relatively real time to make forecasts that can be acted upon? And can we work in partnership with a mobile operator who also has the ability to engage with those populations and reach them through education messaging, which is really one of the only tools we have to address these types of epidemics.
我在白沙瓦大学的同事们帮助我们获取了数据,并提供了所需的本地专业知识。这花了一整年时间来组织、签署保密协议、召集相关人员,以及向卫生部提交申请。之后,我们还必须前往伊斯兰堡,在那里待了十天,获取数据,但因为隐私保护问题不能查看这些数据,而且数据不能带出该国——这倒也没什么问题。这些项目非常庞杂,但它们确实有潜力带来真正的变革。
My colleagues at the University of Peshawar helped us with accessing the data and giving us the local expertise that we need. This took a year of organizing, signing NDAs, getting people on board, and writing to Ministries of Health. Then, we had to go to Islamabad and spend ten days there, getting data, and not looking at it because of the privacy implications, and the data couldn't leave the country which is fine. These projects are very bulky, but they have the potential to be really transformative.
这就是 2013 年的疫情暴发。颜色再次代表巴基斯坦的不同地区。8 月初,该国西北部出现了一次小型暴发。然后,疫情又蔓延到了拉合尔。接着,我们又看到了另一次暴发,但时间比 2011 年那一次晚得多,出现在 10 月至 11 月。
Here's the 2013 outbreak. The colors again represent different places in Pakistan. And so there was this small, mini-outbreak that happened at the beginning of August in the northwest of the country. Then, it spread again to Lahore. Then, we have this other outbreak, but it's much later than the 2011 outbreak, it's October-November.
问题是,CDRs(通话记录数据)能对巴基斯坦未来登革热的预测能力增添什么?
The question is, what can CDRs add to our predictive capacity about dengue in Pakistan in the future?
这是初步数据,但为了让你们有个概念,这里有一张拉合尔的地图。红点是登革热病例,我们已将其地理位置编码得相当精准。黑点是手机信号塔,因此我们拥有非常高分辨率的数据,能掌握每个人在每三小时时段内的位置。这涉及大约 800 万人。
This is preliminary stuff, but just to give you an idea, here's a picture of Lahore. The red dots are cases of dengue, and we have them geo-coded pretty well. And the black dots are cell towers, so we have really great resolution on these and we have three hour chunks of where everybody is. It's like eight million people.
这是迄今为止以这种方式被分析的最大规模数据集,我认为也是首次有人尝试针对一场疫情做这样的事。绿色的点代表报告了病例的医院,而那里还有
This is the largest data set ever to be interrogated in this way, and I think the first time that anyone's tried to do this for an epidemic. The green dots are the hospitals where the cases were reported, and there's
卡罗琳·巴奇:理解疾病传播的新方法(续)
Caroline Buckee Novel Ways to Understand How Disease Spreads (Continued)
有一些传闻说,靠近医院的风险更高。我们想核实一下这一点。所以这就是我们正在处理的数据。你可以看到这些塔的集群,它们对应着村庄。
been anecdotes that you're at higher risk near the hospital. We want to check out that. So this is the data that we're working with. You can see these clusters of towers that correspond to villages.
我们能看出什么?好,我们还拿到了巴基斯坦所有气象站的完整天气数据。这张是拉合尔的温度图。你能看到气温很高,然后在这里开始下降。到某个时间点,气温会冷到让病媒无法存活。
What can we see? All right, so we also got all of the weather data for all of the weather stations across Pakistan. Here's the temperature one for Lahore. You can see it's hot and then tails off here. At some point, it's going to get cold enough that the vectors are no longer able to survive.
下雨了。八月间曾发生一场大洪水——强降雨及其引发的洪灾。我们认为这触发了第一个峰值,之后数据略有下降。这是登革热病例数。这里出现一个峰值,随后又出现一个非常明显的第二波峰值。问题在于,是什么原因导致了这一现象?我们能否通过人类行为指标来量化它?
Here's rain. There was a massive flood – rainfall and a flood associated with it – in August. We think that that kicked off that first peak, and then it goes down a little bit. Here's dengue cases. Here we have this peak, and then we have this very pronounced second peak. And the question is, what's driving this? Can we measure it with human behavior measures?
现在,这是第一轮 CDR 分析。我们观察的是拉合尔人口的流入与流出,可以看到,从七月到八月,所有人先离开,然后又返回。这是斋月期间的现象,非常惊人。在斋月期间,每个人的行为都完全改变了,我们在肯尼亚的圣诞节期间也曾部分观察到这一点,但这次要明显得多。
Now, here's the first pass CDR analysis. So we're looking at movements in and out of Lahore here and what you can see, July to August, everyone leaves and then comes back. This is Ramadan and it's amazing. Everybody does something completely different during Ramadan, which we saw in Kenya to some extent over Christmas, but this was much more pronounced.
一个可行的假设是:所有人都四处移动,而这里又发生了洪水。一例病例与下一例病例之间大约相隔一个月,所以我们认为这次洪水可能在此地引发了疫情,斋月期间人员流动导致病毒扩散,然后又出现了病例聚集增多的情况。拉合尔市的易感人群数量正在增长。我们已经根据订阅用户数量的增加等因素对此进行了调整。
A working hypothesis is that everyone moves around and there's flooding here. There's a time of like a month between one case and the next case, so we think that this flooding probably caused an outbreak here, Ramadan spread it around, and then we see this increase in the aggregation, basically. The number of susceptible people in Lahore is growing. We've adjusted this for increasing numbers of subscribers and stuff like that.
这两大因素——人们在斋月期间四处迁徙、随后聚集、再形成规模——为疫情暴发创造了完美条件。我们正与 Telenor 和巴基斯坦研究人员合作,试图弄清楚这类方法能带来什么效果。我们能否实时做出精准的预测?
These two factors, people moving around for Ramadan then coming together, then aggregating, have made the conditions perfect for this kind of outbreak. We're working with Telenor and researchers in Pakistan to try to figure out what these types of approaches can add. Can we do a really good job of prediction in real time?
如果爆发疫情——事实上,他们现在正面临一场脊髓灰质炎疫情——而我们有每两周更新的 CDR 数据,就能说,人群向这里、这里和这里流动了。实时人口分布是这样的。这是病例出现的位置。你需要在这五个地点部署监测系统,并需要向所有这些人群发送关于这些具体事项的教育信息。
If there was an epidemic – and they've got a polio epidemic right now, actually – and we had CDRs every two weeks, we could say people moved here, here, and here. Live times are like this. This is where the cases are. You need to put surveillance in these five places and you need to send educational messages to all of these people about these specific things.
这不只是加强病媒控制,而是说在你出行时,要确保做到这些。如果你要去轮胎市场,就穿长衣。清理这种特定类型的水容器。这些都是非常实际的事情,但它们正是我们现有的工具,我认为这确实是一个影响这些流行病、甚至可能阻止它们的机会。
It's not simply do more vector control, it's when you travel, make sure you do this. If you are going to the tire market, wear long clothes. Clean this particular type of water container. These are very practical things, but they are the tools that we have and I think that this is really a chance to impact these epidemics and stop them potentially.
最后我想说的是,疟疾每分钟都在夺走一个孩子的生命。我认为,考虑到这种疾病完全可以治疗和预防,这是非常不可接受的。这主要发生在最贫困的社区。而那些大流行病威胁,虽然发生概率很低,但一旦发生,会影响在座的每一个人。它们是真实存在的,所以我认为,现在我们拥有技术、方法和数据,能够真正在这些地方产生影响力。
I'll end by basically saying that malaria kills a child a minute. It's I think pretty unacceptable given that it's treatable and preventable. That's among the poorest communities. But these pandemic threats, they're very low probability events, but if they happen they're going to affect everyone in this room. They're real, and so I think that now we have the technology, the methods, and the data to start actually having impact on these places.
就像我刚提到的,这项工作没有资金支持——这让人无比沮丧——所以如果有人愿意在这方面帮助我们,我真的认为我们有机会为许多人带来巨大改变。我非常期待听到任何问题,也感谢你抽出时间。这是我的联系方式。
And just as I kind of pitched, this work isn't funded – it’s incredibly frustrating – and so if anyone would like to support us in this, I really think we have a chance to make a big difference to a lot of people. So I'd love to hear any questions, and thank you for your time. Here's my details.
[applause]
[applause]
问题:您能否详细谈谈斋月期间模式会发生怎样的变化?
Question: Can you talk more about how the patterns change around Ramadan?
卡罗琳·巴基:理解疾病传播的新方法(续)
Caroline Buckee Novel Ways to Understand How Disease Spreads (Continued)
卡罗琳:嗯,这挺有意思。斋月每年时间都不同,所以并不固定。我们现在做的就是回顾过去,问自己:“它有没有预测性?”——具体来说,不是简单地说通量变化会导致疾病,而是要问:它是否会导致某种特定方向的疾病,以及我们能不能建立一种机制,将来看到某种特定类型的运动时,就能据此做出预测。
Caroline: Well, it's interesting. Ramadan happens at a different time every year, and so it's not necessarily that it's consistent. What we're doing is we're trying to look back and say, "Is it predictive?" and the specifics, not just that flux causes disease, but rather, does it cause disease in a specific direction and can we make a mechanism that we can use to make predictions when we see particular types of movements in the future?
我们只有一年的数据。我不清楚年度之间的波动有多大,所以情况可能很规律,也可能不规律。要继续进行数据分析和方法开发才能弄清楚:我们是否真的需要每月的气候需求比率(CDRs)?现有的数据够不够?我们现在能否直接预测登革热?
We have one year's worth of data. I don't know how much inter-annual variability there is, and so it might be very regular, it might not. Continuing data analysis and continuing methods development is going to be required to figure out, do we really need monthly CDRs, is this enough? Can we just now predict dengue?
再说,这组数据的好处在于成本低。我们不用每年大搞调查,那种做法贵得离谱。要是能想出一个靠谱的预测机制,那就很棒了。我的工作假设是,这个机制应该会是人口流动、气候变量以及每年疫情爆发具体方式的综合结果。
Again, the nice thing about this data is it’s low cost. We're not going out and trying to do surveys every year, which is incredibly expensive. If we can come up with a plausible mechanism that turns out to be predictive, then that's great. My working hypothesis is that it's going to be the aggregation of movement, climate variables, and the specifics of how the outbreak happens each year.
提问者:看上去,您看那张幻灯片时说,“好,这就是源头。”就像您说的,这才是您真正下功夫的地方。那是什么阻止了肯尼亚政府直接去源头灭掉那些蚊子,或者做任何他们能做的事……?
Question: It seemed like when you looked at that slide you said, "OK, this is the source." As you said, this is where you're actually concentrating. What's stopping the Kenyan Government from going to the source and wiping out the mosquitoes, or whatever it is that they could do to...?
卡罗琳:因为这和人们通常想做的正好相反。这不是一个容易获得胜利的选项。
Caroline: Because it's kind of the reverse of what people tend to want to do. That's not an easy win.
你会注意到,实际上它与乌干达接壤,所以我没有谈到的是跨境移民问题,这是一个巨大的问题,尤其是在大湄公河区域。我认为我们有办法应对这个问题,因为运营商通常会记录手机号码。然后,无论 SIM 卡是什么,我们实际上都可以开始追踪到这一点。
You'll notice, actually it's on the border with Uganda, so what I didn't talk about is cross-border migration, which is a huge issue, especially in the Greater Mekong region. I think that we have ways to deal with that, because quite often, operators log handset numbers. Then we can actually, regardless of SIM card, we can start getting at that.
我认为不仅如此,更大的问题是庞大的惯性和推进事情的难度。我认为这正是与运营方合作的价值所在,因为我们可以对他们说:“我们来做一项随机对照试验,向识别出的高风险人群发送短信,提醒他们在出行时使用蚊帐。”
I think more than that, there's this issue that it's just huge inertia and a difficulty getting things done. I think that's why these partnerships with the operators is great, because we can say to them, "Let's do a randomized control trial where we try sending text messages to people that we identify as high risk, and asking them to use a bed net when they travel."
我们可以做到。我们能快速、可规模化地执行,取得结果,然后检验是否奏效。我觉得各国的卫生部,它们就像是行动极其缓慢、笨重的巨兽。
We can do that. We can do it quickly and scalably and get results and see if it works. I think ministries of health, they're very slow, cumbersome monsters.
提问:您对滴滴出行怎么看?
Question: What is your opinion on DDT?
卡罗琳:哦,滴滴涕(DDT),对,我觉得大家的共识是它有它的用处。在某些情况下它确实非常有用。滴滴涕的耐药性是个大问题。所以,在根除计划接近尾声时,有些地区的蚊子已经对它完全没有反应了,就是产生了耐药性,因此现在有很大一股力量在推动研发新型杀虫剂和不同的方法。
Caroline: Oh, DDT, yeah, so I think the consensus is that it has its place. It's incredibly useful in some situations. DDT resistance is a big problem. So, at the end of the eradication program, there were some populations where like it didn't work at all, the mosquitoes were just resistant, so there's a big push to actually come up with novel types of insecticide and different ways of doing things.
我和一位同事共事过,她发现了一种激素:如果你把它放进蚊帐里,它基本上能让雌蚊停止产卵——所以你可以想象出各种全新的、不同的办法来控制病媒。当然,还有转基因蚊子。
I worked with a colleague who, she's found this hormone, if you put it in a bed net, it basically stops the females from laying eggs, and so you could imagine all these new, different ways of manipulating the vectors. There's GM [genetically modified] mosquitoes of course.
滴滴涕自有其用武之地。抗药性是目前最大的问题,我认为,至少在撒哈拉以南非洲是这样。但我确实认为,这不是一种非此即彼的选择。
DDT has its place. Resistance is the biggest issue I think, at least in sub-Saharan Africa. But I definitely think it's not a kind of all-or-nothing kind of thing.
问:可否详细谈谈传染性,以及它与不同流行病的关系?一次接触事件中传播的概率是多少?
Question: Can you expand on infectiousness and how it relates to different epidemics? What is the probability of transmission given a contact event?
卡罗琳:好了,这事儿挺难的,对吧?我觉得……传染性或许不如传播方式和它对人的影响那么重要。比方说,拿非典(严重急性呼吸综合征)来讲,我们算比较走运,因为你先出现症状,然后才具有传染性。如果
Caroline: OK, it's a difficult thing, right? I think that's...The infectiousness is probably less important than the mode of transmission and the way it affects people. So, with SARS [severe acute respiratory syndrome] for example, we were kind of lucky because you get symptoms and then you're infectious. If
卡罗琳·巴基:理解疾病传播的新方法(续)
Caroline Buckee Novel Ways to Understand How Disease Spreads (Continued)
有一种情况是,你长期具有传染性,然后突然死亡——实际上有点像 HIV,你具有传染性,在传播病毒,自己可能还不知道——这才是真正麻烦的地方,也是让疾病快速扩散的方式。所以,最重要的未必是每次接触的传播率,而是传播方式以及该特定病原体的流行病学特征。
there was something that happened where you're infectious for ages and then suddenly you die – kind of like HIV, actually, you’re infectious and you're transmitting and you may not know it – that’s really problematic, and that's how you spread something really fast. So, it's not necessarily the per contact transmission rate that's the most important thing. It's mode of transmission and the kind of epidemiology of that particular bug.
对于这些流行病,登革热是其中之一,但动物宿主也是一个重大问题。我和很多合作者共同绘制过一张地图:如果以胡志明市为中心画一个四小时航程的圈,里面涵盖了全球四分之一的人口和 75% 的鸡。
The real worry for these epidemics, dengue is one, but the zoonotic reservoir is a big issue. There's a map that a lot of people I collaborate with put up, where if you draw a four hour flight around Ho Chi Minh City, you capture a quarter of the world's people and 75 percent of the world's chickens.
就像流感和其他存在于鸟类、鸭子、鸡以及家禽体内的病毒一样,人们和这些动物同住一屋,这类地方存在着一个巨大的病毒病原体和活体动物市场。因此,不确定性在于这类动物源性事件何时会发生。但我想,内在的传播性本身并非最大问题,流行病学特征才是关键。
Like flu and other viruses that live in birds, ducks, chickens, poultry that people live in their house with, there's a huge reservoir of viral pathogens and live animal markets in these types of places. And so the uncertainty lies in when those kind of zoonotic events are going to happen. But the inherent transmissibility itself is not the biggest issue so much as the epidemiological characteristics, I think.
提问:您对严重健康风险的担忧程度如何?
Question: How concerned are you about severe health risks?
卡罗琳:这很难回答。一场严重的流行病……我认为另一个问题是它会影响谁。
Caroline: That's difficult to answer. A severe epidemic...I think the other issue here is who it affects.
我们谈论的这些尾部事件,比如“你们所有人都死了”,人们往往更关心这类情况,而不是“非洲婴儿死了”。问题是,一直有一些可怕的疾病在发生,而且已经持续了很久,非常严重,但因为它们不是流行病,也因为不会影响到我们,就被认为不那么严重。
These tail-end events we're talking about like “all of you guys are dead” which people tend to care about more than like “African babies are dead.” The problem is that, there are horrible diseases that happen all the time and have been happening forever that are really bad, but because they're not epidemic and because they don't affect us, they're perceived as less bad.
就严重程度而言,情况千差万别。从进化角度来看,过去有一个假说认为,所有病原体都应该进化成无症状的。它们不该引发症状,因为最成功的病原体不会杀死任何人——如果你杀死宿主,基本上就等于切断了传播机会。我们现在知道这并不完全正确。并没有一个全球性的病原体协议认为这是个好主意。
In terms of severity, you can get the whole spectrum. In evolutionary terms, there was this old hypothesis that all pathogens should evolve to be asymptomatic. They shouldn't cause symptoms because the most successful one wouldn't kill anyone, because you're basically cutting off chances for transmission if you kill people. We know now that that's not really true. There's no kind of global pathogen agreement that that's a good idea.
单个病原体只是通过自然选择在进化,所以如果它们有效,那就有效;如果它们杀死了一些人,那就要在这一点与它们的传播能力之间取得平衡。有一种观点是,存在一种进化到稳定状态的中等毒力水平,如果这能回答你的问题的话。
Individual pathogens are simply evolving by natural selection, so if they work, they work, and if they kill some people, then they'll balance that with how good they are at transmitting. There's this idea that there's an intermediate level of virulence that has evolved to be stable, if that answers your question.
这也是新出现病原体之所以麻烦的部分原因——它没有与人类长期适应过。它可能直接横扫人群,杀死所有人。黑死病、西班牙流感等等就是这么回事。当时没有免疫力。
That's part of the reason that an emergent thing is problematic because it hasn't adapted with humans for a long time. It can just rip through the population and kill everyone. That's what happened with the Bubonic plague and Spanish flu and these kinds of things. There was no immunity.
提问:我们该对埃博拉病毒有多担心?
Question: How worried should we be about the Ebola virus?
卡罗琳:嗯,如果我住在西非,我会担心。所以,你说“我们”是指……?
Caroline: Well, I would if I lived in West Africa. Yes, so, when you say we, do you mean...?
提问:比如说,如果它上了飞机呢。
Question: Well, if it gets on a plane, for example.
卡罗琳:它的传播性目前也不清楚。所以我们担心的是高传播性与高毒力的组合。这就是那些经典的流行病案例。它们就是那样。麻疹是高毒力、高传播性的,但我们已经习惯了。
Caroline: It's not clear how transmissible it is either. So what we're worried about is this combination of highly transmissible and highly virulent. And that's these kind of classic examples of epidemics. That's what they were. Measles is highly virulent and highly transmissible, but we're kind of used to that.
提问:你提到盖茨用了“根除”这个词,这当时引起了争议。
Question: You mentioned that Gates used the word "eradicate" and how that was kind of controversial.
他们对疾病的动态有不同的看法吗,还是说这对研究类型有影响……他们是资助方吗,背后是什么情况?
Do they have a different view of the dynamics of the disease or does it have implications for the kind of research...Were they funded or what's the story around it?
卡罗琳·巴基:理解疾病传播的新方法(续)
Caroline Buckee Novel Ways to Understand How Disease Spreads (Continued)
卡罗琳:情况是这样的,整个领域基本上没有人在研究传播,我们打算永远承受这个负担,只管处理症状。这就是当时的主流框架。
Caroline: The story is that, the whole community, there was basically no research into transmission, and we were just going to live with this burden forever, but we’d try and manage symptoms. That was the kind of general framework of where we were.
而比尔和梅琳达·盖茨基本上说:“不,这不可接受。我们应该能够根除它。这种疾病的长期经济负担极其沉重,所以现在投入精力根除它,以后就再也没有这些成本了。”
They, Bill and Melinda Gates basically said, "No, that's not acceptable. We should be able to eradicate this. It's the kind of burden where the economic impact over a long time is just enormous, so let's put in effort now, eradicate it, and no longer have those costs.”
这之所以有争议,是因为整个领域都说:“嗯,不可能。”他们则挑战了这个想法,说:“有可能。”一个经典的思维实验是:疟疾是可以治疗的,对吧?如果我们搞一个“全球吃疟疾药日”,所有人同时吃药,也许每月一次,连续三个月,疟疾就没了。这是可能的。
It was controversial because the entire community was like, "Well, it's not possible." They sort of challenged that idea like, "It is possible." The classic thought experiment is: It's treatable, right, so if we had like World Take Your Malaria Pill Day, and everyone took it at the same time, maybe once a month for three months, there would be no malaria. It's possible.
他们把这个想法重新摆上了台面,而且坦率地说,因为他们的资金非常充裕,而疟疾领域没什么钱,所以他们能在某种程度上决定科学的方向。就像,“好,我们就研究传播。”
They kind of put that idea back on the table and because, frankly, they have so much money and there wasn't much money in malaria, they can dictate kind of how science goes. It's like, "Great, we'll work on transmission."
提问:不是研究药物?
Question: Less on pills.
卡罗琳:嗯,过去对严重疾病、发病机制等方面强调得太多,导致一些生物学问题研究严重不足,现在有一点回归趋势,转向基础工具开发,比如实地工具——科学家们往往喜欢花哨的分子技术。嗯?
Caroline: Well, there's been such an emphasis on severe disease and pathogenesis and those kinds of things, that there's been some really understudied biology, there's been a bit of a shift back, and towards basic tool development, like field tools that scientists tend to prefer fancy molecular stuff. Yes?
提问:疫苗是难开发,还是病原体对变化有免疫力?
Question: Were the vaccines difficult to develop or is that they immune to change?
卡罗琳:两者都有。我们对疾病的免疫机制了解不多,比如为什么有的孩子会这样……两个孩子的血液里有相同数量的寄生虫,一个在踢足球,另一个昏迷了。我们不知道为什么。
Caroline: Both, so we don't really understand immunity to disease, like why some kids have this...Two kids will have the same number of parasites in their blood and one will be playing soccer and one will be in a coma. We don't understand why.
疟原虫的基因多样性极其丰富,这也是为什么你永远无法获得免疫力——你每次遇到的都是不同的毒株。一个问题是,你的身体永远不会见到两次相同的疟原虫。
The parasite is enormously genetically diverse, so that's part of the reason why you never get immune because you get a different strain every time. One issue is that your body never sees the same malaria parasite twice.
另一个问题是,我们似乎不太擅长产生免疫反应。疟原虫会操纵宿主的免疫系统,我们对所有疟原虫共有的那些靶点无法产生有效的反应。而理想情况是,疫苗的靶点是所有疟原虫共有的东西,但我们对那些东西无法产生良好的抗体反应。
The other issue is that we just don't seem to be very good at making immune responses. So, the parasite manipulates the host immune system and we don't make very effective responses against the things that are shared by all malaria parasites. So that would be the ideal. You have a vaccine where the target is something that all malaria parasites have, but we don't make good antibody responses to those things.
Yes?
Yes?
提问:你们能方便地获取医院记录吗?
Question: Do you have good access to hospital records?
卡罗琳:临床数据实际上是迄今为止最大数据鸿沟的最大来源。我们得不到好的临床数据。实际上……对于肯尼亚,那里长期进行了大量研究,我们与疟疾地图项目合作。
Caroline: Clinical data is actually the biggest source of the big data gap, ever. We can't get good clinical data. We actually...For Kenya, there's been so much research there for a long time and so we work with the Malaria Atlas Project.
他们一直在使用相当复杂的地理空间技术构建高分辨率地图,利用一系列横断面研究和纵向研究,对估算进行平滑处理,最终得出一个一公里乘一公里的网格,估算感染者的比例。
They've been building high resolution maps using quite sophisticated kind of geo-spatial techniques to take a bunch of cross-sectional and longitudinal studies and smooth the estimates to come up with a one by one kilometer grid of the estimate of the fraction of infected people.
我们利用这个来为我们的旅行者分配感染概率,然后按这个方式运行。同样,你也可以思考:“那里的发病率是多少?我们对人口统计有什么了解?”我们可以把所有这些信息输入模型。另一个部分是什么?
We use that to basically assign our probabilities of infection for our travelers, and then we kind of run it that way. Similarly, you can think about, "Well, what's the kind of disease rate there? Do we know anything about the demographics?" We can put all of those things into the model. What was the other part?
卡罗琳·巴基:理解疾病传播的新方法(续)
Caroline Buckee Novel Ways to Understand How Disease Spreads (Continued)
提问:你们有感染者随时间推移的临床报告吗?
Question: Do you have clinical reports over time for infected people?
卡罗琳:哦,不。现阶段绝对没有。在尼日利亚,我们与“无疟疾”组织合作。在尼日利亚,他们刚刚通过一项法律,所有药品都必须通过个人移动设备进行验证。
Caroline: Oh yeah, no. Definitely not at this stage. In Nigeria, we're working with Malaria No More. And in Nigeria, they've just made a law that all drugs have to be authenticated through one’s mobile device.
有一种刮涂层的东西,你发短信到验证服务,他们会回复短信告诉你这是假药还是真药。
There's a thing you scratch off, you text it in to the authentication service, and they text back if it's counterfeit or real.
这意味着我们应当开始获取个体数据,比如你至少因疟疾发作接受过多少次治疗。
That means that we should be getting...We should start getting individual data on like how many episodes did you get treated for at least.
我们和他们讨论过建设后端系统,试图在保留匿名性的同时,向正确的人发送干预和教育信息。这非常困难,我认为现阶段还处于非常早期的阶段,没人找到解决办法。尝试过很多小办法,但没有一个真正奏效。没有一个能大规模推广。我认为未来会实现,但需要时间。嗯?
We've talked to them about sinking back-end, like trying to preserve anonymity while also sending intervention, educational messages to the correct people. It's really hard and I think at this stage, it's still very much in its infancy, and nobody's figured it out. There's been so many little things tried and nothing's really worked. Nothing's really scaled. I think it will, but it takes time. Yes?
提问:能再多谈谈使用杀虫剂的问题吗?还有,如果你能创造一种转基因蚊子,理论上它可以消灭其他蚊子,这会带来什么问题?
Question: Can you talk more about using insecticides? And, if you could create a genetically-modified mosquito which presumably could wipe out other mosquitoes, what are the issues associated with that?
这有多大可行性?
How plausible is that?
卡罗琳:我认为,是的,它们会这么做。这相当政治化,因为消灭一整个蚊子物种——我个人觉得这太好了——但有很多非常有趣的技术仍处于早期阶段。
Caroline: I think, yes they would. It is quite political, because getting rid of a whole species of mosquito, I would think it would be great, but there are lots of really interesting technologies here that are still in their infancy.
控制登革热的一个非常有趣的方法是让蚊子感染沃尔巴克体,这是一种细菌。沃尔巴克体实际上能使蚊子绝育,并能传递给后代。所以你在种群中传播不育性……有很多方法可以培育出大量雄性蚊子,然后它们显然就没什么用了,所以……
One really interesting one for dengue control is that you infect them with wolbachia, which is a bacteria, and the wolbachia effectively makes a mosquito sterile and it passes it on. So you're spreading sterility through the population...There are lots of ways to make a lot of males, and then it's obviously useless, so...
[laughter]
[laughter]
还有各种各样的技术。我会说它们令人兴奋,但仍处于早期阶段。我确实认为,如果试验和实地试验结果有效,它们就会被使用。事实上,目前澳大利亚东北部正在进行沃尔巴克体感染的登革热媒介的实地试验,他们释放的地点与当年释放甘蔗蟾蜍和兔子的地点相同。我是说,这真的很糟糕,但不管怎样,这些技术令人兴奋,我认为它们是真实的。我觉得它们很有趣。
Then, there are all these different techniques. I would say that they're exciting, but they're in their infancy still. I do think that they would be used if the trials and the field trials seem to be productive. In fact, there are currently field trials in Northeast Australia of the wolbachia infected dengue vectors, and they released them in the same place that they released the cane toad and the rabbits. I mean, it's really terrible, but anyway those are exciting and I think they're real. I think they're interesting.
提问:你在整理数据时有没有什么特别让你惊讶的发现?
Question: Is there anything that really surprised you when you sorted the data?
卡罗琳:比如在巴基斯坦的数据里,我真的很惊讶人们打电话的频率。非常多,这很好,因为我们有了更精细的分辨率估计,但对我来说有点奇怪。我认为,在这些低收入地区,进出大城市的旅行量之大令人惊叹。
Caroline: Like in the Pakistan data, I was really surprised how often people were calling. Like a lot and so that's great because we have finer resolution estimates, but it's kind of weird to me. I think the sheer volume of travel in and out of big cities in these low income settings is amazing.
从标准的经济学模型来看,人们常用引力模型来估算两地之间旅行的人数。但这些模型偏差太大了。所以我们一直用来思考感染空间传播的标准模型效果并不好,而这是我们第一次实际看到这一点,并看到偏差有多大。
From standard models in the kind of economics, people often use gravity models, which approximate the number of people that travel between places. They're just off by so much. So the standard models that we've been using to think about the spatial dissemination of infections are not doing well, and this was the first time that we could actually see it and see how.
这些大城市中心,这类巨型枢纽,对于感染的传播及其带来的所有社会和经济后果至关重要。这就是其中一件有意思的事。
These centers, these sort of megacity hubs, are incredibly important for the spread of infection and all of the social and economic repercussions. And so that was one of the interesting things.
提问:还有哪些主要机构在研究传染性疾病?你们之间有合作吗?
Question: What other major bodies are looking at research in contagious diseases? Are you working together?
卡罗琳·巴基:理解疾病传播的新方法(续)
Caroline Buckee Novel Ways to Understand How Disease Spreads (Continued)
卡罗琳:实际上拥有数据的团队非常少,因为获取数据极其困难。
Caroline: There are very few groups that actually have data, because it's incredibly difficult to get.
实际上,Flowminder 是一个松散的学者网络,我们尝试汇集资源,共享数据,以便能充分利用……我们有来自纳米比亚、卢旺达和肯尼亚的数据。我们在努力扩充数据集。
Actually, Flowminder is a kind of a loose network of academics, and we kind of try and pool our resources and share data so that we can leverage...We have data from Namibia, Rwanda, and Kenya. We're trying to build up our data sets.
问题在于,从运营商那里获得支持非常困难,所以实际上很少有人……他们不感兴趣给你数据,而且数据也很难获取。特别是,使用 CDR 分析的人真的很少。
The problem is that it's really hard to get buy-in from operators, and so very few people actually, they're not interested in giving you the data, and it's very hard to access it. There's really very few people using CDR analytics in particular.
有一些团队主要分布在美国和英国,他们在这方面做了大量理论工作,并使用数学方法研究空间传播。这是一个规模很小、联系紧密的社群,所以我们和他们有很多合作,比如风疹研究小组,就是普林斯顿的那个小组。我和牛津大学的人也有很多合作。
There are groups across the [United] States and in the U.K., mainly, who do a lot of kind of theoretical work on this and look at spatial transmission using mathematics. It's quite a small, close-knit community, so we work with them a lot, like the rubella group, that's the Princeton group. I work with people at Oxford a lot.
好处在于,他们在这些地区与临床医生和实地站点有广泛的联系网络,所以关键是把所有人聚在一起,弄清楚如何把数据送到合适的人手中,并进行本地能力建设。我不知道这是否回答了你的问题,但我想说,目前这个圈子还相当小。
The nice thing there is that they have far-reaching networks with clinicians and field sites across these places, and so it's really about bringing together everyone, figuring out how to get the data to the right people and do local capacity building. I don't know if that answered your question, but I would say it's a pretty small group at this point.
提问:你怎么看待围绕大数据和其他工具的热情?
Question: What do you think about the enthusiasm surrounding big data and other tools?
卡罗琳:显然,关于大数据,尤其是大数据促进发展,已经有了大量的炒作。我不满的地方在于,这一切都是自上而下的,就像是,“耶,大数据!让我们解决世界的问题。派数据科学家来吧。”
Caroline: There's obviously been a massive amount hype around big data, especially big data for development. My problem with it is that it's all been top-down, so it's like, "Yay, big data! Let's solve the world. Send in the data scientists."
而实际上,应该是:我们需要弄清楚耐药性会往哪个方向发展,这就是我们需要做的事,以及如何以严谨的方式找到正确的数据来实现它。
When in reality, it should be we need to figure out where drug resistance is going, this is what we need to do that, how can we find the right data in a rigorous way to make it happen.
我认为,对于参与式监测、把一切众包出去、大数据,甚至那种现在可以做无假设科学的想法,存在一种过度热情。
I think there's been a kind of over-enthusiasm for participatory surveillance, crowd-sourcing everything, big data, and even this idea that you can do hypothesis-free science now.
我不同意那个概念,我认为最终会发生的是,当所有泡沫都消散后,留下来的会是那些说,“是啊,我还是想治愈疟疾”的人。我认为谷歌流感趋势是未来趋势的一个预兆。我确实认为参与式监测可以被非常有效地利用。
I disagree with that concept, and I think what has to happen is this kind of, once the froth has all died down, you'll be left with people who are like, "Yeah, I still want to cure malaria." The Google Flu Trends thing, I think, is a harbinger of things to come. I do think that participatory surveillance can be used very effectively.
我们需要更好地理解它。我们需要更好地把它与真实数据严谨地联系起来。我们需要进行验证和测试,并且我们需要对它的能力以及它如何能影响实际控制工作更诚实一些。
We need to understand it better. We need to do a better job of rigorously linking it to real data. We need to do validation and testing, and we need to be a bit more honest about the capabilities and how it can impact the actual control.
肖恩·古尔利:Quid
Sean Gourley Quid
肖恩·古尔利是一位物理学家、十项全能运动员、政治顾问和 TED 研究员。他来自新西兰,在那里他曾竞选国家公职,并帮助创办了新西兰第一家纳米技术公司。
Sean Gourley is a physicist, decathlete, political advisor, and TED fellow. He is originally from New Zealand, where he ran for national elected office and helped start New Zealand’s first nanotech company.
肖恩作为罗德学者在牛津大学学习,并获得博士学位,研究方向是现代战争的数学模式。这项研究带他走遍世界各地,从五角大楼到联合国,再到伊拉克。
Sean studied at Oxford as a Rhodes Scholar where he received a PhD for his research on the mathematical patterns that underlie modern war. This research has taken him all over the world, from the Pentagon to the United Nations and Iraq.
此前,肖恩曾在 NASA 从事自修复纳米电路的研究,并两次获得新西兰田径锦标赛冠军。肖恩目前常驻旧金山,是增强智能公司 Quid 的联合创始人兼首席技术官。
Previously, Sean worked at NASA on self-repairing nano-circuits and is a two-time New Zealand track and field champion. Sean is now based in San Francisco where he is the co-founder and CTO of Quid, an augmented intelligence company.
肖恩·古尔利:增强智能
Sean Gourley Augmented Intelligence
迈克尔·莫布森:我很高兴介绍下一位演讲者,肖恩·古尔利。他一方面从事数学研究,另一方面也是 Quid 公司的联合创始人兼首席技术官,该公司帮助决策者消化和理解公开信息。
Michael Mauboussin: I’m pleased to introduce our next speaker, Sean Gourley, who splits his time between doing mathematical research and working as a co-founder and chief technology officer for Quid, which helps decision makers digest and understand public information.
肖恩在两个领域做过工作,我认为这些工作非常吸引人,并且与在座各位高度相关。第一个领域是使用数学和统计技术来理解社会科学中的模式和规律。他的工作聚焦于现代战争,但我们也看到了类似的方法应用于其他领域,例如城市研究。这项工作解决了一些基本问题,即如何在宏观层面上对社会系统进行建模。
Sean has done work in two areas I believe are fascinating and highly relevant for this group. The first is using mathematical and statistical techniques to understand patterns and regularities in the social sciences. His work is focused on modern war, but we've seen similar work on initiatives, for example, in the realm of cities. This work addresses fundamental questions about how we can model social systems at a macro level.
第二个领域,也是今天讨论的重点,是围绕增强智能的概念。我们知道,人类擅长某些事情,比如模式识别,但不擅长其他事情,比如快速、海量的计算。我们也知道,计算机擅长某些事情,不擅长其他事情。
The second area, which will be much more of the focus of today's discussion, is around the concept of augmented intelligence. We do know that humans are good at some things, such as pattern recognition, and bad at other things, such as rapid, massive calculations. We also know that computers are good at some things and bad at others.
增强智能的挑战,或者说增强智能的机会,在于弄清楚如何将人类最擅长的事情与计算机最擅长的事情结合起来,以便做出深思熟虑的预测,同时当然也要避免人类和计算机各自最不擅长的事情。
The challenge of augmented intelligence, or the opportunity of augmented intelligence, is to figure out how to combine the best of what humans do with the best of what computers do in order to make thoughtful predictions, while of course avoiding the worst of what both humans and computers do.
关于肖恩,最后一点是,他曾两次(2000 年和 2002 年)获得新西兰十项全能全国冠军。如你所知,奥运会十项全能的冠军通常被认为是世界上最伟大的运动员,所以我认为我们可以说,肖恩至少两次拥有新西兰最伟大运动员的头衔!
One final note about Sean is that he was New Zealand's national champion in the decathlon twice; 2000 and 2002. As you know, the winner of the Olympic decathlon is generally defined as the world's greatest athlete, so I think we can say that Sean held the title of New Zealand's greatest athlete at least twice!
请大家和我一起欢迎肖恩·古尔利。
Please join me in welcoming Sean Gourley.
[applause]
[applause]
肖恩·古尔利:很高兴来到这里,我想我有整整 75 分钟的时间,这大概是我一段时间以来需要做的最长的演讲了!从我看到的情况来看,根据你们的问题,我们最后会进行一些有趣的互动交流,你来我往。
Sean Gourley: Great to be out here and I think I've got a full 75 minutes, which is probably the longest talk that I've had to do for a while! I think, from what I've seen, with your questions, we're going to have some fun at the end of it as we kind of go backwards and forwards.
但在那之前,我想带你们回到我的童年;可能和你们很多人一样,我的童年是在电脑屏幕前度过的。我花了很多时间玩电脑游戏,尤其是《大蜜蜂》。我想在座一些特定年龄的人会记得《大蜜蜂》。
But before we get to all that, I want to take you back to my childhood; and probably, like a lot of yours, mine was spent in front of a computer screen. I spent so much time playing computer games, Galaga in particular, and I think some of you of a certain age will remember Galaga.
我确实变得很擅长玩这个游戏——好到能拿到各种高分。但无论我变得多好,我总会被外星人杀死,被子弹和导弹击中。我的反应不够快。我根本没法移动得足够快。电脑会赢。
I actually got pretty good at playing this game—so good that I could get in all sorts of high scores. But no matter how good I got, I'd always get killed by the aliens and dodging the bullets and missiles. My reactions were not quick enough. I simply couldn't move fast enough. The computers would win.
后来我才了解到,人类的反应时间大约是 120 毫秒,这差不多是我们的极限,而电脑当然可以比这思考得更快。这在高频交易领域是有影响的。
I'd later come to learn that the human reaction time is about 120 milliseconds and that's about as far as we go, and so computers, of course, can think faster than that. That has ramifications in the world of high frequency trading.
但我们比那强,对吧?我们可以进行战略性思考。我们可以思考更大、更概念性的框架。我们可以思考电脑无法思考的问题。
But we're better than that, right? We can think strategically. We can think about bigger, more conceptual frameworks. We can think about problems that computers can't think about.
1997 年,很多事情都变了,当时我们很多人目睹了卡斯帕罗夫对阵 IBM 的“深蓝”超级计算机,并以 1 比 2 落败。对许多人来说,这是一个人工智能战胜人类思维能力的传奇。凭借纯粹的计算机算力,计算机可以提前思考八步,而一位国际象棋特级大师实际上只能思考大约四步。
In 1997, a lot of that changed when many of us watched Kasparov take on, and get defeated 2-1 by, IBM's Deep Blue supercomputer. Now, for many of us, this was a story of artificial intelligence triumphing over the human ability to think. That through sheer computational horsepower, a computer could think up to eight steps ahead, whereas a grandmaster was really only about four.
这个故事引起了共鸣,并成了今天我们许多人关于人工智能的主流看法——它纯粹就是胜过人脑。
This story resonated and became sort of the dominant story that many of us think about today regarding artificial intelligence―that it simply trumps the human brain.
肖恩·古尔利:增强智能(续)
Sean Gourley Augmented Intelligence (Continued)
当然,对卡斯帕罗夫来说,事情没那么简单。卡斯帕罗夫,像他那么聪明的人,意识到计算机将在国际象棋上击败人类,但这未必就是最终结局。如果人类和计算机联手呢?如果人类和计算机开始合作呢?
Of course, for Kasparov, this wasn't so simple. Kasparov, being the smart guy that he is, realized that computers were going to beat humans at chess, but that didn't have to be the be all and end all. What if humans teamed up with computers? What if humans and computers started to come together?
他构思了一种新型的国际象棋,他称之为“自由式国际象棋”。自由式国际象棋是一个有趣的概念。你可以和其他特级大师组队,或者如果你觉得这是最佳方案,你也可以单独使用计算机。你可以人和计算机一起上。任何组合都可以。如果你觉得能获得更清晰的思路,你甚至可以给自己吃点迷幻药。真的,没有什么基本规则。你尽自己最大能力下棋。
He formulated a new type of chess, which he called freestyle chess. Freestyle chess was an interesting concept. You could team up with other grandmasters or you could use a computer by itself if you felt that was the best solution. You could use a human and a computer together. You could have any combination of that. You could put yourself on psychedelic drugs if you thought you'd get more clarity. Really, there were no basic rules. You played chess to the best of your ability.
我非常喜欢这种理念。我们下棋能下得多好?
I really like this kind of philosophy. How good can we get at playing chess?
2006 年的世界在线锦标赛,有来自世界各地的 48 支队伍参赛。它们大致分为两类。其中一组,用这里的 Hal 代表,你可以把它们看作是“无头计算机”。纯粹的计算机——没有人类。另一边,是所谓的“半人马”——人类和计算机协同工作,希望能下出更好的棋。
The world championships online in 2006, 48 teams entered from around the world. They were broadly split into two categories. One of the groups, represented here by Hal, you can think of as the headless computers. Pure computer―no humans. On the other side, you've got what they call centaurs―humans and computers working together to kind of create, hopefully, better chess.
随着比赛进行,竞争开始演变,有一件事变得非常清楚。半人马队伍在主导比赛。他们如此占优,以至于最后四强全部是半人马。
As the games went on and the competition started to evolve, one thing became very clear. The centaurs were dominating. They were dominating so much that the final four teams were all centaurs.
其中三支队伍我们认识:他们是俄罗斯特级大师,配备大型军用级别的超级计算机。但第四支队伍隐藏了身份。它叫 ZackS。ZackS 下了一些非常有趣的棋。有时挺大胆,有时挺非正统;但它实际上开始赢了,而且赢下了整场比赛。
Three of them we knew: they were Russian grandmasters and big military-grade type supercomputers. But the fourth team kept its identity hidden. It was call ZackS. ZackS played some really interesting chess. It was at times kind of bold, at times kind of unorthodox; but it actually started to win, and so much so that it won the whole competition.
到了这个时候,大家都以为卡斯帕罗夫会走上台,领取奖金,宣布胜利。那会是一个不错的故事——但事实要有趣得多。
Now, at this point, everyone sort of expected Kasparov to walk on stage, collect the prize and claim victory. That would be a nice story―except the truth is much more interesting.
事实是,有两名人类玩家,而且绝对不是特级大师。一个是数据库管理员,另一个是足球教练。他们用的是三台消费级的计算机,每台运行着不同版本的人工智能系统,其中一台还是他们从父母那里借来的计算机。
The truth is there were two human players, most definitely not grandmasters. One was a database administrator and the other a soccer coach. They were playing on three consumer grade computers, each running a different version of an AI system, with one being a computer they borrowed from their parents.
想想看。世界上最好的国际象棋团队竟然用着消费级笔记本电脑,而且成员也不是特级大师。他们击败了最好的棋手和最好的计算机。
Think about that. The best chess playing team in the world was running consumer grade laptops and weren't grandmasters. They beat the best chess players and the best computers.
他们成功的原因,不是因为任何一个组成部分是最好的;而是因为他们知道如何驾驭这些机器。他们知道什么时候该听 AI-1 的,而不是 AI-3 的。他们知道什么时候该听自己的。他们知道什么时候该回头说,“我们要让这一步更激进一些。”他们根据对手来下棋。
The reason they did this was not because any one of those pieces were the best; it was because they knew how to drive the machines. They knew when to listen to AI-1 and not AI-3. They knew when to listen to themselves. They knew when to go back and say, "We're going to make this a little more aggressive." They played the player they were playing.
我认为这一点确实极为关键。你必须了解算法是如何运作的。你必须知道什么时候该服从一种算法,什么时候不该;什么时候该听从自己的判断,什么时候该向你面前的那台机器低头。他们在这件事上做得最好,最终也登上了顶峰。
This I think is really, really key. You have to learn how the algorithms work. You have to know when to obey one algorithm and not another. When to listen to yourself, and when to defer to the machine that's in front of you. They did this the best and they ended up on top.
我认为我们面临的一些最棘手问题,答案不在于人工智能,而在于增强智能:人与机器协同配合,解决我们面对的最大难题。国际象棋当然是一个例子,但还有其他领域也是如此。
I think the story for some of the most difficult problems that we are facing is not one of artificial intelligence, but instead it's one of augmented intelligence: humans and machines interfacing to solve the most difficult problems we face. Chess is one example of that, of course, but there are others.
想象一下,一边是一个高维度、复杂庞大的世界,我们可以运用数学技巧来降低这个世界的复杂度。另一边,是我们与生俱来的人脑。我们可以借助可视化技术来增强人脑的能力。在这两者之间的某个地方,我们发现了一个增强智能(augmented intelligence)的空间——在这里,我们有望对世界获得足够清晰的把握,从而真正地驾驭它。
Think of a big complex world with high dimensionality on one side, and we can use mathematical techniques to reduce the complexity of that world. On the other side, we have the human brain that we were born with. We can use visualization techniques to enhance that. Somewhere in the middle, we find a space of augmented intelligence, where we hopefully can get enough grasp on the world to actually navigate it.
肖恩·古尔利谈增强智能(续)
Sean Gourley Augmented Intelligence (Continued)
它实际用途之一就是天气预报。天气预报是一个引人入胜的难题——我不知道在座有没有气象学家,但这是目前最精妙的预测游戏之一。
One of the things that it's actually used for is weather prediction. Weather prediction is a fascinating problem. I don't know if there are any meteorologists in the room, but that's one of the best prediction games out there.
与其他大多数处理数据的人不同,气象学家实际上是在构建世界的模型。他们建立了空气粒子相互作用的模型,一个覆盖美国和加拿大的完整模型,用来预测天气。
Meteorologists, unlike most people that deal with data, actually build models of the world. They build models of the interaction of air particles, an entire model of the United States and Canada to predict the weather.
我们做的远不止是简单的线性回归。这本质上是一个混沌系统,存在反馈效应,高度非线性。非线性程度之高,以至于在 20 世纪 50 年代,三天天气预报都被认为是不可靠的。
It's not simply linear regression that's being done. It's a chaotic system, it has feedback effects. It's highly non-linear. So much so that back in the 1950s, a three day weather forecast was kind of considered untenable.
另外,气象学家还有两个优点:一是他们很执拗——但这并没阻止他们继续预测;二是他们确实会记录自己的预测。当他们说下雨概率是 83% 的时候,他们会测量实际是否下雨——我觉得这种做法很美,我们应该多做这类事。这样做的好处是,我们可以长期追踪这些预测的准确率。
Also, a couple nice things about meteorologists are: one, they're stubborn – that didn't stop them; and two, they actually keep records of their predictions. When they say it's going to rain with 83 percent probability, they measure if it rained or not, which I think is a beautiful thing that we should do more of. The nice thing about that is we can track the success of those predictions over time.
你可以看到,无论是 36 小时还是 72 小时的预报,准确率都在稳步上升。20 世纪 50 年代,36 小时预报的准确率大约只有 25%。到了 2005 年,这一数字已经升至约 78%。
You can see that the percentage accuracy, for both 36- and 72-hour forecasts, rose steadily. In the 1950s, the 36-hour accuracy is about 25 percent. By 2005, it’s about 78 percent.
有趣的是,你在 72 小时预报的改进上也能看到同样的情况。再说一次,这些都是非常、非常棘手的非线性问题。
Interestingly, you see the same kind of thing with the improvement on the 72-hour forecast. Again, these are very, very difficult non-linear problems to solve.
如今,这一切当然是由计算能力的巨大飞跃所驱动的。你当年用的可能是 IBM 701 那样的机器,而现在,IBM 的超级计算机已经达到当年性能的许多倍——计算能力提升了大约 150 亿倍。
Now, of course, it's being powered by massive changes in computational power. You're dealing with something like an IBM 701 back in the day. Today, they've now got IBM supercomputers running at many times greater than that―about a 15 billion-fold increase in computational power.
我们获得的计算能力增长了 150 亿倍,用来驱动一个非线性问题大致呈线性增长。想想还挺有意思。但人类在这个过程里处于什么位置?如果我们有 150 亿倍的计算能力用在解决这个问题上,感觉人类可能没什么立足之地。天气预报员的好处在于,他们对这点也有跟踪记录。
We've got a 15 billion-fold increase in computational power to drive a roughly linear increase in a non-linear problem. That's kind of a nice thing to think about. But where are humans in all of this? If we've got 15 billion times the computational power being thrown at this problem, it feels like humans probably don't have a place. The nice thing about the weather forecasters is they keep track of that as well.
蓝色代表仅使用机器的方案,红色代表人类与机器协作的方案。我们获得了约 16% 的跃升。在测量的整个时间段内,这一提升幅度保持相对恒定。期间计算能力发生了巨大变化。
In the blue, we've got the machines only, and in the red we've got humans and machines. We're getting about a 16 percent jump. That stays relatively constant through the time period that was measured. There were big changes in computational power there.
就像扎克斯在象棋例子中的表现一样,他们知道什么时候该相信机器,什么时候不该相信。他们知道什么时候说“我们用极端天气算法,别用干旱算法”,“用降雨算法”,或者“我知道这个算法在西风条件下总是出错,所以我手动修正一下”。气象学家坐在那里干的就是这个。这能让准确率提高大约 16%。
Much like ZackS on the chess example, they knew when to listen to the machine and when not to. They knew when to say, "We're going to use the extreme algorithm and not the drought algorithm,” “We're going to use the rainy algorithm," or, "I know that this algorithm always gets it wrong with the westerly winds, so I'm going to correct for that." That's what the meteorologist is sitting and doing. They get about a 16 percent improvement.
然而,这一景象很像那种典型的由一个人和一条狗管理的工厂——人在那里负责喂狗,
However, much like the archetypal factory being run by a man and a dog, the man is there to feed the
狗。狗的作用是让人远离机器。
dog. The dog is there to keep the man away from the machines.
[laughter]
[laughter]
如果你不是气象学家,很可能不该去碰这个模型,你多半改进不了它。如果你不是国际象棋特级大师——或者至少不懂相关算法——也不该去改动它。对你们大多数人来说,机器告诉你怎么做,多半就是对的。
If you are not a meteorologist, you probably shouldn't mess with this model. You're probably not going to improve it. If you're not a chess grandmaster―or at least if you don't know the algorithms―you shouldn't mess with it. For most of you, what the machine tells you to do will probably be right.
在我们这些人中,有一小部分人确实能与 AI 互动并在此基础上增加价值。这非常引人注目,因为这些问题是最难的,而在一些最困难的问题上获得 16% 的提升,意义极其重大。
There’s a small percentage of us around that can actually interact and add value on top. That's fascinating because they are the hardest problems, and 16 percent on some of the hardest problems is a massive
肖恩·古尔利谈增强智能(续)
Sean Gourley Augmented Intelligence (Continued)
请记住这一点,我们在一系列不同的系统中都看到了这一点。人加机器,大约能提升 10% 到 15%。
jump. Keep that in mind that we see that across a range of different systems. Human plus machine, around 10 to 15 percent improvement.
要做到这一点,我们需要理解界面,因为尽管机器非常擅长读取特征向量和线性代数等内容,但我们大多数人并不擅长在这种界面下工作。
To do that, we need to understand the interface, because whilst machines are very good at reading eigenvectors and linear algebra and so on, most of us don't work very well in that kind of interface.
我们视觉感受强得多,所以得弄懂那个交互界面。
We're much more visual, so we need to understand that interface.
要理解这个交互界面,我们需要了解我们每个人脑袋里都有的这台计算机。
To comprehend that interface, we need to understand this computer that we all have up in our heads.
我们想谈谈直觉。
We want to think about intuition.
我认为直觉是一种我们视为专家才具备的能力。他们在看似毫无可能知晓的情况下,却知道该怎么做。
I think intuition is something that we deem an expert to have. They know what to do when it seems like there's no way they could have known it.
现在,专家拥有直觉。这听起来很神秘,但我们如今知道它其实并非如此。我们从 Wan 等人在 2011 年《科学》杂志上发表的一篇真正引人入胜的论文中了解到这一点,该论文研究了直觉的神经回路,以及人们产生直觉时大脑中发生的情况 [Wan 等,“棋盘游戏专家直觉性最佳下一步生成的神经基础”,《科学》,2011 年 1 月 21 日,第 341-346 页]。
Now, an expert has intuition. It's mystical except we know now that it's actually not. We know that from Wan et al. in Science 2011, who put together a really fascinating paper on this looking at the neural circuitry for intuition and what goes on in the brain when people have it [Wan et al., “The Neural Basis of Intuitive Best Next-Move Generation in Board Game Experts,” Science, January 21, 2011, 341-346.].
回到国际象棋——它就像是所有这类系统的玩具模型一样。研究者重新回到棋盘前,将棋盘投影在 fMRI 脑功能成像机上,把人放进扫描仪,测量大脑哪些区域会亮起来。他们用两个群体做了这个实验:一类是国际象棋专家,一类是业余棋手。
Going back to chess, it is like the toy model for every kind of system like this. They went back to chess, and they projected a chessboard up onto an fMRI machine. They put the people into their fMRI machine and measured the parts of the brain that lit up. They did this with two groups: expert chess players and amateur chess players.
他们将那张图像只展示了一秒钟,因为不想让显意识大脑介入。
They showed the image for just one second because they didn't want the conscious brain to kick in.
大脑的楔前叶区域——位于后上方——开始活跃起来。在专家身上,这个区域的活跃程度是业余爱好者的两倍。
What happened was the precuneus part of the brain, which is up and to the back, lit up. It lit up twice as strong in the experts as it did the amateurs.
给非神经科学背景的朋友解释一下:前楔叶(precuneus)是大脑负责模式识别的区域。它的功能就是在别人只看到噪音的地方,发现信号。棋盘只展示一秒时,专家就看到了信号,而业余选手什么都没看到。接下来,他们把展示时间延长到仅仅两秒……
For those that aren't neuroscientists, the precuneus is the part of the brain associated with pattern recognition. It's the thing that sees a signal where everyone else sees noise. In one second with the chess board shown the experts saw a signal and the amateurs didn't see anything. The next thing they did was flash up something for just two seconds . . .
[laughter]
[laughter]
……然后说:“选出你最好的那步棋。”这真的很有趣,因为现在大脑的不同区域被点亮了。这是尾状核。尾状核是大脑中更原始的部分,位于下方,几乎靠近基底神经节。它是大脑中与习得反应功能相关的区域。
. . . and said, "Choose your best move." This was really fascinating because now a different part of the brain lit up. This was the caudate nucleus. The caudate nucleus is a much more primitive part of the brain down and to the bottom just about at the basal ganglia. It's the part of the brain associated with learned response functions.
这个信号对专家们亮了,对业余人士倒没怎么亮,而且亮的方式非常、非常有意思。
It lit up for the experts. It didn't really light up for the amateurs, and it lit up in a really, really interesting way.
在这里,纵轴代表活跃强度,横轴代表正确回答的百分比。这张图实际在说明的是:当那张图像闪过的两秒钟里,尾状核发出的信号越强,他们答对的可能性就越高。
Here you can see the strength of activity on the vertical and the correct response percentage on the horizontal. What this is actually saying is that the stronger the signal from the caudate nucleus when that image was flashed up for two seconds, the more likely they were to be correct.
仿佛那些专家在潜意识层面装了一个内部的热-冷开关,只要感觉某件事正确或错误,开关就会自动切换。
It's as if the experts at a subconscious level had an internal hot-and-cold switch, that they would get a feeling that a thing was right or wrong.
具体来说,直觉图像的神经回路会通过视觉皮层传递。它们被楔前叶抽象简化为更简单的维度,以此来识别信号。
Breaking that down, the neural circuitry for intuition images come through the visual cortex. They're abstracted to a simpler dimension with the precuneus to identify a signal.
肖恩·古尔利增强智能(续)
Sean Gourley Augmented Intelligence (Continued)
该信号会被拿来与我们拥有的所有其他信号进行比对,此时尾状核会启动冷热反应机制,从中选出最佳的那个。构建这一系统的成本确实极其高昂。如果你今天还不是国际象棋特级大师,那可能就没什么机会了。
That signal is compared back against all the other signals that we have with the caudate nucleus kicking in with the hot-and-cold response function to pick the best one. It's really expensive to build this. If you're not a chess grandmaster today, you're probably out of luck.
训练这些神经回路需要耗费大量时间,才能在不到一秒内识别出模式是什么、最优走法是什么、以及该如何应对。在计算机编程的所有任务中,构建直觉的神经回路大概是最耗费计算资源的任务之一。
It takes a huge amount of time to train these neural circuits to identify in less than a second what the pattern is, what the best move is, what they should do about it. Probably one of the most computationally expensive tasks in all of computer programming is to build neural circuitry for intuition.
其价值在于,作为专家,他们能将思考中的计算部分卸载给潜意识,从而腾出所有时间和意识来思考这些计算的含义,而业余人士还在那儿琢磨这里面到底有没有什么信号。
It's valuable because as an expert they're able to offload the computational part of their thinking to the subconscious, and that frees them up with all the time and the consciousness to think about the implications of that, whilst the amateurs are still trying to figure out if there's some sort of signal here or not.
如果我们能够开发软件来完成第一部分的工作呢?如果我们能做出这样的软件——识别模式、建立关联、抽象出结构——让你把任何一组数据放进去,都能获得专家那种直觉般的感知呢?这正是大约四年前我们开始着手做的事情。
What if we could build software to do this first part? What if we could build software to identify patterns, to make connections, to abstract structure so that you could put any data set in front of it and get that same feeling of intuition that an expert has? This is exactly what we set out to do about four years ago.
四年前,我们着手创建一家专攻此事的公司,名叫 Quid。公司总部在旧金山,目前有 65 名员工。待会儿我会带你们回溯到旧金山看一看,但在此之前,我想先聊聊我们是怎么走到那一步的。
Four years ago, we set out to build a company to do just that, and it's called Quid. We're based in San Francisco. There are 65 of us working there. I'm going to take you back to San Francisco in a little bit, but before we get there I just want to talk about how we got there.
我有幸获得机会前往牛津攻读博士学位。最初去那里,本意是想在物理实验室研究生物分子马达。在那间实验室待了几周后,我很快意识到,牛津本身远比实验室里有趣得多。
I was lucky enough to get a chance to go up to Oxford and study to do my Ph.D. work. I originally went over there really to study in a physics lab to do biomolecular motors. I spent the first couple of weeks in that lab, and soon realized Oxford was a much more interesting place outside of it.
[laughter]
[laughter]
我暗自琢磨:“攻读物理学博士这条路,一定还有更好的法子。我在物理实验室之外,还有别的事情要做。”
I thought to myself, "There's got to be a better way to do this whole physics Ph.D. I've got something to do outside of the physics lab."
我早年有一次参加牛津大学那种高桌晚宴,记得自己坐在前中情局局长旁边。几杯酒下肚后,我们开始辩论即将到来的伊拉克入侵——或者你想怎么描述都行。不用说,我们对那项决策可能产生的结果看法相左,但我能理解他的出发点。我绝对能看到他的道理。
I was sitting down at one of these high-table dinners they do at Oxford early on, and I remember sitting next to the former head of the CIA. We got into a debate after a few glasses of wine about the upcoming invasion or however you want to frame that of Iraq. Suffice to say we disagreed about the likely efficacy of that decision, but I could see where he was coming from. I could absolutely see his points.
我在很多方面不同意他,但我没有什么论据来支撑自己的立场。那对我来说极其迷人。我们正在做出一个重大决定——进入一个国家,却并不真正知道事情会如何发展,而且根本没有一个分析框架来指导。对我这个物理学家来说,我觉得这触动了我脑中很多模式识别的开关。
In many ways I disagreed with him, but I didn't have anything to back that up. It was incredibly fascinating for me. We're making a massive decision here to go into a country and not really know exactly how this thing's going to unfold, and there's no real framework to do so. For me being a physicist on this, I think it set off a lot of the pattern recognitions for me.
我说:“天哪,如果我们能拿到数据呢?如果我们能开始分析这个呢?如果我们能提出一个理论框架来理解叛乱如何开始展开呢?”
I said, "Well, jeez. What if we could get data? What if we could start to analyze this? What if we could come up with a theoretical framework for understanding how insurgencies start to unfold?"
这正是我着手去做的事,后来成了我的博士论文:研究伊拉克等地的叛乱,以及世界各地发生的混乱与不确定性。这些是非常难以理解把握的系统,按理说也不该适合这类分析。
That's exactly what I set out to do, and that became my Ph.D. work, studying the insurgencies in places like Iraq and to study the chaos and uncertainty that occurs all around the world. These are very difficult systems to wrap our heads around and things that shouldn't really lend itself to this kind of analysis.
当然了,作为一名物理学家,做这件事的第一步就是获取数据。你需要从实证角度获得数据流,但当你像我这样带着口音时,对方会说:“抱歉,你不会拿到我们的数据的。”
Of course, the first thing to do that, as a physicist, is to acquire data. You need data streams as an empirical perspective, but when you have an accent like mine, they say, "Well, I'm sorry. You're not going to get our data."
美国——无论对错——对流入的信息进行了管控。这几乎从一开始就把我拦在了门外。我只能看福克斯新闻来安慰自己。
The U.S., rightly or wrongly on that, had controls on the information that was coming in. That almost stopped me dead in my tracks at the start. I consoled myself by watching Fox News.
肖恩·古尔利 增强智能(续)
Sean Gourley Augmented Intelligence (Continued)
[laughter]
[laughter]
有一晚我在牛津看深夜电视,我记得听着、看着那些专家评论。但我同时还看到别的东西——屏幕底部有一个小滚动条。
When I was watching TV late one night at Oxford, and I remember listening to and seeing the talking heads. But I remember seeing something else, this little ticker down at the bottom.
那是一条条滚动播出的新闻,列着爆炸和袭击事件。我记得自己当时想:“如果我们能把那个提取出来呢?如果我们能开始训练计算机去读取那些信息呢?如果我们能建立一个系统来监控它、提取那些数据呢?”这就能绕开所有机密限制,对吧?
This little ticker of news that came across, scrolling little events that happened with bombs and attacks that were going off. I remember thinking, "What if we were able to extract that? What if we were able to start training computers to read that? What if we could have a system that monitored that, and pulled out that data.”? That would get us past the whole classified thing, right?
我们就是这么做的。我们建立了一个系统,监控新闻报道和当时开始涌现的博客,来源有几百个,然后训练它们去识别袭击。
That's exactly what we did. We built out a system that would monitor news reports and the blogs that were starting to emerge from hundreds of different sources, and train them to find attacks.
我们训练计算机去识别哪里有人员死亡:什么时间发生、在哪里发生、死了多少人等等。
We trained them to find people that had been killed, when it happened, where it happened, how many people were killed and so on.
我们做得相当成功。有了这个系统,我们就有了一个所有袭击事件的数据库。然后我们可以分析这些数据,看看里面是否真的包含某种信号。
We got quite successful at it, and when we did that we had a database of all of the attacks that were there. We could then look at that to see if there was actually a signal contained inside it.
我们一直不知道自己的数据到底有多好,直到[朱利安]·阿桑奇发布了维基解密的信息。
We didn't know how good our data was until [Julian] Assange released the WikiLeaks information.
那件事非常有意思。我们的公开数据覆盖了美军信息的 81%,这个结果相当不错。但美军只覆盖了我们信息的 70%。你想想看。
It was quite interesting. The open data had 81 percent coverage of what the U.S. military had, and that's a pretty good result. But they only had 70 percent of what we had. Now, let that sink in for a second.
2006 年的信息格局下,牛津大学三台笔记本电脑上训练计算机读取环境新闻源,收集到的重要事件信息,竟然比整个驻当地美军还要多。
Three people on laptops in Oxford training computers to read ambient news sources pulled in more significant event information than the entire U.S. military on the ground, with an information landscape of 2006.
信息格局已经发生了翻天覆地的变化。无论你组织内部有多少信息,组织外部的信息都要多得多。
The information landscape has massively changed. However much information you have inside your organization, there is much more outside of it.
当然,这对安全领域也有影响,但这只是第一步——可以从环境信息源中获取海量信息。第二步是,这些信息内部存在统计特征。
Of course, this has implications for security as well, but that's the kind of first step – the massive amount of information that can be claimed from ambient sources. The second is that there are statistical signatures that exist inside this.
如果你把整个伊拉克冲突期间袭击的频率对袭击规模画成一个对数-对数图,你会看到一阶近似下是一条直线,符合幂律分布。这里的指数是这条线的斜率,2.3。在所有这些混乱中,在我们攻击和演化的不同团体的行动中,我们观察到了这样一个强烈的统计特征。
If you look at the frequency of attacks versus the attack size for Iraq on a log-log plot across the entire conflict, you see that the first order of a straight line approximates a power law distribution. The exponent here is the slope of the line of 2.3. Out of all the chaos unfolding across all of the different groups that we're attacking and evolving, we see a strong statistical signature like this.
有趣的是,这个特征在全世界许多不同的冲突中一再出现。你在阿富汗、哥伦比亚、塞拉利昂都能看到同样的特征。甚至在北爱尔兰也是如此。
The interesting thing is this is repeated across many different conflicts around the world. You see the same signature in Afghanistan, Columbia, and Sierra Leone. You even see it in Northern Ireland.
就好像当人们聚在一起互相残杀时,他们是以数学上精确的方式进行的。这听起来不可思议,直到你仔细思考。作为叛乱力量,要与强大的对手对抗是极其困难的。
It's as though when people get together to kill each other they do so in mathematically precise ways, which doesn't really make sense until you think about this. It's incredibly hard as an insurgent force to take on a strong opposition.
实际上如此困难,以至于可行的方法寥寥无几。如果你在演化过程中没有找到解决方案,你就不复存在了,也就没有足够的数据可测量。对于任何有足够数据可测量的系统来说,这意味着叛乱分子已经找到了一种组织形态,而这种形态的特征正是这些统计特征。
It's so hard, in fact, there are very few ways to do this. And if you evolve and you don't find the solution, you don't exist. There's not enough data to be measured. For any system where there's enough data to be measured, it means that the insurgents have found an organizational structure that is typified by these statistical signatures.
五角大楼的反应很有意思。当然,他们第一句话就问:“你能预测吗?”
The Pentagon was quite interesting, and of course, the first thing they start to say is, "Can you predict?"
肖恩·古尔利 增强智能(续)
Sean Gourley Augmented Intelligence (Continued)
没有预测能力是得不到这类统计特征的——这就是它的本质——但这个问题本身也很奇怪。我为什么要预测巴格达北郊会有三个人被炸死,而不坐下来想办法改变它呢?
You don't get these kinds of statistical signatures without predictive ability, that's the nature of it, but it's also a very strange question. Why would I want to predict that three people are going to killed in the northern suburbs of Baghdad without wanting to sit down and actually change it?
当我们坐在五角大楼的会议桌旁,思考伊拉克的增兵行动——是派 3 万、6 万还是不派兵时,问题的关键根本不是“我们能预测袭击吗?”,而是“我们要如何改变这个系统,让这个特征完全不存在?”要做到这一点,你不能只靠一个简单的统计模型。你需要开始思考叛乱分子是如何组织的。你需要建立模型,模拟他们如何配置资源、如何做出决策。
As we sat around the Pentagon table and thought about this surge going into Iraq and whether to send 30,000, 60,000, or no troops, it wasn't really a question of "Can we predict the attacks?" It was, "How are we going to change the system so that this signature doesn't exist at all?" To do that, you need more than a simple statistical model. You need to start to think about how the insurgents are organizing. You need to create models of how the insurgents are allocating resources and create models of how they're making decisions.
这就是我们做的事。我们决定在计算机内部建立模型来模拟叛乱分子,模型的输出必须与我们在现实中观察到的统计特征相吻合。
This is what we did. We decided to create models inside of our computers to simulate the insurgents, and the output of the models had to come back and agree with the statistical signatures that we're observing.
然后你就可以开始调整这些变量。你可以问:“如果我们改变驻军人数会怎样?如果我们打击大型组织而不碰小型组织会怎样?
You could start to change these. You could say, "Well, what happens if we change the number of troops on the ground? What happens if attack the large groups and not the small groups?
如果我们试图切断通信会怎样?如果我们什么都不做会怎样?”我们于 2009 年将这项研究发表在《自然》杂志封面[胡安·卡米洛·博霍克斯、肖恩·古尔利、亚历山大·R·迪克逊、迈克尔·斯帕加和尼尔·F·约翰逊,《共同生态学量化人类叛乱》,《自然》杂志,2009 年 12 月 17 日]。
What happens if we tried to break communication? What happens if we did nothing at all?" We published this in 2009 on the cover of Nature [Juan Camilo Bohorquez, Sean Gourley, Alexander R. Dixon, Michael Spaga and Neil F. Johnson, “Common Ecology Quantifies Human Insurgency,” Nature, December 17, 2009.].
对我来说,能把这项研究发表出来是一件很棒的事。在跟中情局局长那次餐桌对话之后的七年,我终于可以坐下来,说:“好了,这里有一个框架。这里有一些我们可以开始用来在这个世界做出重大决策的工具。”
For me, it was a great thing to get this research out. To finally, after seven years at that point since that conversation I'd had at that dinner table with the CIA, sit down and say, "Well, actually here's a framework. Here's something that we can start to use to make some very significant decisions in our world."
当我在伊拉克四处游荡、思考这一切意味着什么的时候,我在脑子里做了一些算术。
As I was in Iraq, kind of wandering the hills thinking about what this all meant, I was doing some math in my head.
我们当时有大约 30 万欧元的欧盟资助,六个人在攻关这个问题。我在想:“我们想建立一个系统,能自动追踪全世界所有的信息,从而发现这些统计特征、这些模式,让我们能做出一些非常重要的决策。要做到这一点,可能还需要更多的资金。”
We had about 300,000 dollars in European Union funding, six people working on the problem, and I was thinking, "We want to build a system that ambiently tracks all the information in the world to start to find these statistical signatures, these patterns, that allows us to make some very important decisions, and it's probably going to take a little more money to do that."
我估算了一下,1 亿美元和 1000 个人应该是个不错的起点。当然,到那时我继续往前走,觉得这件事可能成不了,直到我意识到有一个神奇的地方叫硅谷,那里的人们怀揣疯狂的想法做着这类事情,而且他们真的能得到资金去实现。这一切真是太神奇了。
I figured $100 million and 1,000 people should be a good start, and of course at that point I kept walking and decided that that might not happen, until I realized there was a magical place called Silicon Valley where people with crazy ideas do these kinds of things. They're actually given money to do them. It's just all very magical.
我后来见到了彼得·蒂尔,他是 PayPal 的创始人之一,也是 Facebook 的早期投资人,当时身价堪比神明。
I actually met up with Peter Thiel, who is [a] founder of PayPal, but also an early investor in Facebook, and worth more than God at the moment.
他给了我最初的 250 万美元来启动这个项目,所以我们实际上在 2009 年底就开始了,并开始组建团队来开发软件。
He gave me the first $2.5 million to get this started, so we actually began at the end of 2009, and started to build a team that would build the software.
我们 65 个人在硅谷旧金山市中心,办公室里培训着非常年轻的数据科学家。团队大约三分之二是技术人员,大约 10% 是纯粹的数据科学家,其余的是销售和运营人员。
With 65 of us out in Silicon Valley, downtown San Francisco, we train our data scientists very young in our office and we’re about two-thirds technical. There are about 10 percent of us that are pure data scientists, and the rest are kind of sales and operations.
现在我们正在向世界各地许多不同的团体交付软件。这种软件可以集成在你的网页浏览器里,接入不同的数据流,利用人工智能来寻找信息。
We're shipping software today to many different groups around the world, a software that you can plug into inside of your web browser to tap into these different data streams to make use of the artificial intelligence to find the information.
肖恩·古尔利 增强智能(续)
Sean Gourley Augmented Intelligence (Continued)
我们的客户范围很广:从微软和 IBM 这样的大型科技公司,到红牛和乐高这样更有趣的消费类公司,以及介于两者之间的各种企业。客户还包括对冲基金、广告公司以及政府决策部门。
There's a range of clients: everyone from the big tech companies of Microsoft and IBM, all the way though to some of the more interesting consumer companies like Red Bull and Lego, and sort of everyone in between. A huge range of customers from hedge funds to advertising agencies through to government decision-making.
原因在于,一旦你尝试用计算机像人类一样去阅读并在事物之间建立联系,你就会发现这在各行各业都有很多应用。
The reason for that is because once you try to compute it to read and make connections between objects like a human does, it actually turns out there are quite a lot of applications for that across a whole range of industries.
他们提出的问题大致是这样的:“印度关于气候变化的主流叙事是什么?不同年龄段的人有何差异?”或者“我的竞争对手在先进的柔性显示技术上做了什么?我是应该与他们合作还是与之竞争?”
They ask questions kind of like this: "What are the dominant narratives about climate change in India, and how does it vary within age groups," or "What are my competitors doing with advanced flexible display technology, and should I partner or compete with them?"
这些都是专家通常会问的问题,但如果要一个四人分析团队花六周时间去阅读资料才能给你答案,那就……如果你恰好有幸与麦肯锡合作,做这件事要花掉你 120 万美元。
These are the kinds of questions that experts would generally ask, but would take a team of four analysts six weeks of reading to get this information back to you. If you happen to be so lucky to work with McKinsey, that will cost you $1.2 million to do that.
你看,人类阅读这些信息的成本很高。我们为美国宇航局做的一个项目是分析商业太空探索系统(commercial space seeker system)的结构。
Again, humans are expensive to read this information. One that we did here with NASA looked at the structure of the commercial space seeker system.
商业航天行业已不再是政府的专属领地。Space X 很可能即将发射火箭,并具备将人类送入空间站的能力,但这将彻底改变整个行业的格局。
The commercial space industry is no longer a government monopoly. They're likely to be launching Space X and taking humans capabilities now up to the space station, but this changes the landscape quite dramatically.
如果你想要了解商业航天产业,一个可以看看的地方就是谷歌。我们可以在新闻流中搜索一下,过去几个月里看到了 4200 篇关于航天产业的文章,这还挺不错的。
If you want to understand the commercial space industry, one place you might look is Google. We might do a search on the news stream and in the last few months we see 4,200 articles on the space industry, and that's kind of great.
只需 0.17 秒就能返回结果,多亏了谷歌,你接下来六周的阅读材料都有了。
It returns it back in 0.17 seconds, and thanks to Google, you have your reading for the next six weeks.
再说一遍,有哪一篇文章是我应该读的?
Again, what's the one article that I should read?
这个问题听起来有点奇怪,“这周有没有一篇文章特别值得我读?” 当你开始看到人工智能说“这是最适合你的内容”时,现实是我想读所有东西。我想看到全部。
That seems like kind of a strange question, "Is there one article I should read about this week?" You start to see the issues with AI saying, "Here's the best thing for you," when the reality is I want to read everything. I want to see all of it.
我们能否将那 4200 篇文章变成一个动态交互式用户界面,让人类去探索其中究竟发生了什么?答案是,我们能。我们的做法是,提取文章中的文本。
Can we turn those 4,200 articles into a dynamic interactive user interface to let humans explore what's going on? We can do that. What we do is we take the text from the article.
我们对每一篇发布的新闻文章、博客和推特帖子都做这样的处理,提取出内容,然后从中寻找关键概念。
We do this for every single news article, blog, Twitter post that's published, extract that out, and start to look for key concepts within that.
我们拥有一些算法,可以通读这些内容来提取关键概念,从而生成我们所谓的数字签名——一种每篇文章的"指纹"。
We have algorithms that will read through these things to extract out the key concepts to create what we thought of as a digital signature―a sort of fingerprint of every article.
现在,有了这个基础,我们就可以开始与其他每一篇文章进行比对。如果找到内容相似的,就能建立关联,并将其投射到一个二维空间中——一个我们终于能够与之互动的空间。
Now, with that, we can then start to make comparisons to every other article. If we find one that's similar, we can connect it and project it onto a two-dimensional space, something that we can finally interact with.
理论上是这样运作的,但我们现在可以看看实际中是如何操作的。这里我们在网页浏览器中查看 Quid 软件。
That's how it works in theory, but we can have a look now to see how it works in practice. Here we're looking at the Quid software now inside the web browser.
你可以看到,“No”在这里代表了一个关于全球航天产业的故事,但故事并非孤立存在,它们之间有联系。这些联系指向相似的故事,而我们实际上可以开始在那里发现这一点。
You can see "No" represents a story here about the global space industry, but stories aren't just by themselves, they have connections. They have connections to similar stories, and we can actually start to see that there.
肖恩·古尔利增强智能(续)
Sean Gourley Augmented Intelligence (Continued)
目前这方面还没有空间,也不重要,所以我们可以按程度来调整规模。我们还可以运用一种力导向算法,开始为我们所观察的内容带来一定的空间分辨率。
Now, there's no space in this at the moment and there's no importance, so we can size by degree. We can also play a force directed algorithm to start to bring some spatial resolution into what we're looking at.
当我们设定一些参数并加以应用时,就能看到相似的故事聚在一起,而不同的故事则四散开来,一幅关于全球航天产业当前讨论话题的地貌图便逐渐浮现。
As we set some parameters and apply that, we can now see the similar stories cluster together, and the different stories kind of fly apart, and a landscape starts to emerge of the different topics that are now being discussed in the global space industry.
我们可以开始应用聚类算法。颜色现在大致代表了相似的社群,我们可以像在谷歌地图上那样放大和环顾四周,只不过这不再是地理地点的地图。
We can start to apply a clustering algorithm. The colors now represent similar communities at a high level, and we can zoom in and take a little look around, much like we do on a Google map, except this now is not a map of geographic places.
这是一张全球航天产业正在讨论的各种想法和概念的蓝图。沿着它移动,在黄色区域你能看到欧洲航天局对遥远行星进行的一些探索任务。
This is a map of all the different ideas and concepts that are being discussed around the global space industry. You move across there and in the yellow you see some of the European Space Agency explorations to far out planets.
我们可以往中间看下去,从网络安全到底部这里,波音和雷神公司的 GMD 杀伤载具和拦截弹。关于那里发生的事情,已经有过不少报道了。
We can come down through the middle, everything from cybersecurity to down at the bottom here, the GMD kill vehicles and the interceptors from Boeing and Raytheon. A bunch of stories written about what's going on there.
短短几分钟内,你就能把成千上万篇文章理出头绪,看清正在讨论的不同话题和概念。这正是你希望自己的分析师团队在六周后能交付给你的成果,但也许他们还未必能完全拿捏准。应用聚类算法后,我们还能看到它开始为不同的元素命名。
Within the space of a few minutes you can orientate yourself to all those thousands of articles to see the different topics and concepts that are being discussed. The same thing that you'd hope that your set of analysts would return to you after six weeks, but maybe they wouldn't quite get it right. Applying the cluster, we can see it will start to name the different elements as well.
有趣的是——您可以看到左上角——预算中包含了 NASA。对我们很多人来说,航天业就是 NASA,但您能看到,这个话题实际上只占了正在发生的事情的一小部分。
What's interesting here is―you can see up into the left―you’ve got NASA in the budget. For many of us, the space industry is NASA, but you can see that the conversation is actually a very small part of actually what's going on.
更大的部分是其他相关议题。例如,连接在地球观测旁边的美国军事、军事防御、商业卫星图像,以及中间相连的网络战和太空碎片。
Much larger is the other topics that are going there. For example, the U.S. military, connected next to Earth observation, military defense, commercial satellite imagery and also connecting up into the middle there, cyber war and space debris.
迄今为止,最大的一类(或者说一组业务)来自卫星宽带、移动宽带和运载火箭。
By far the largest cluster, or group of things, is what comes around from a satellite broadband, mobile broadband, launch vehicles.
你开始意识到,商业航天行业其实并非我们通常想象的那样——火箭和宇航员。它关乎通信系统,关乎在卫星之间来回发送数据,为你的移动生活方式提供动力——这差不多是最先跳出来的东西。
What you start to realize is the commercial space industry is not really about space as we generally imagine it with rockets and astronauts. It's about communications systems. It's about sending data backwards and forwards between satellites to power your mobile lifestyle – sort of the first thing that pops up.
如果让一位专家在酒过三巡时,随手在餐巾纸背面勾画出当时的情形,他可能会画出这样一张图,把不同元素都标出来,一边解释一边带你走一遍。
Now, if you were to ask an expert to sketch out what was going on on the back of a napkin over drinks, they might put something like this together to draw up the different elements and explain and kind of walk you through that.
这就是它们的认知模型。实际上,我们可以让机器也拥有相当类似的认知模型,把同类概念聚合、聚类在一起,用连接表示相似性,然后我们去探索这个模型。我们可以跟它互动。
This is their cognitive model. Well, we can actually have a pretty similar kind of cognitive model from the machine, the same kinds of concepts aggregated and clustered together, their connections representing similarities, and we can go and explore that. We can interact with it.
这个领域的技术含量很高。本质上,我们是在把那些信息上传到我们的大脑中。
It's got a tech top field to it. In essence, we're sort of uploading that information into our brain.
问题:一个简短的问题,软件是自动选择类别名称,还是由分析师来操作?
Question: Quick question, does the software select category names or is it an analyst?
肖恩:只需要三次猜测,然后分析师就可以编辑它了。它可能像是观测/地球/成像之类的。
Sean: It takes three guesses and then an analyst can edit it. It might be like observation/earth/imaging.
然后,编辑处理完成,数据就会发送出去。这就是地球观测。
Then, the editing will come through and will go off. That's earth observation.
肖恩·古利 增强智能(续)
Sean Gourley Augmented Intelligence (Continued)
是的,它擅长猜测,但还比不上人类。再说一次,这里面有一种编辑的成分。
Yeah, it's good at guessing, but it's not quite as good as a human would be. Again, there's a sort of editing component.
有了这个,我们就可以再次应用该算法。这是一种非常高层的聚类算法。我们称它大致分为两组。你可以看到,这个网络的结构确实让信息的各个部分分布在两侧。
With that we can again apply the algorithm. This is a very high level clustering algorithm. We say it splits into broadly two groups. You can see the structure of that network does lend itself with different bits of the information across either side.
你可以把这事看作美国与世界其他地区的对决。有意思的是,当我们谈论美国内部的事情时,和谈论世界其他地区时,完全是两套不同的对话。这很有意思。
You can kind of think of this as United States versus the rest of the world. What's interesting there is the different conversation when we're talking about stuff in the United States compared to when we're talking about the rest of the world. That's interesting.
真正有意思的事情恰恰发生在边界区域——这两个小地方。我们放大观察美国和数字世界之间正在发生什么。你会看到这里有一堆文章,讲的都是出口许可制度的彻底改革。
What's really interesting is the stuff that happens right at the boundaries―these two little places. We go and zoom and take a look at what's happening between the United States in the digital world. You see a bunch of articles here about the overhaul of the export licensing system.
当然,还有 ITAR(《国际武器贸易条例》)的管制规定,禁止任何太空技术出口到世界其他地区,以防苏联可能开始使用这些技术。算法已经准确识别出,这正发生在这两个集群的交界处。这是一个相当强大且微妙的洞察。
Of course, there's the ITAR [International Traffic in Arms Regulations] regulations that forbid any space technologies to be exported to the rest of the world for fear that the Soviets might start to use them. The algorithms have correctly identified that that's something that's happening at the interface of these two clusters. Quite a powerful nuanced insight.
我们转到另一边,看看这里在这个界面里发生着什么——更多是围绕 NASA 的,我们看到 NASA 局长(查尔斯·博尔登)正在准备一趟中国之行。这同样是因为他们开始向其他有太空能力的国家伸出橄榄枝,寻求建立联系。
We go up to the other side and we see here what's happening in this interface, which is more around NASA, as we see [Charles] Bolden, the chief administrator of NASA, preparing for a trip to China. Again, as they start to make overtures to connect with other countries that have space capabilities.
当你开始探索这些领域时,你能观察到的不仅是主要集群,还有它们之间的互动方式——这些洞察可能相当微妙。
These, as you start to navigate it, can be quite nuanced insights that you can observe, not just the major clusters, but how they're interacting with each other.
这是高耸的塞拉山脉。如果有人去过那里,就知道它有多美。有一年夏天,我跟一位生态学家在那里待了几周,四处走访那些高山草甸,研究其中蕴藏的食物系统。
This is the High Sierra. If anyone has been there you know how beautiful it is. I spent a few weeks there one summer running around with an ecologist, looking at these high alpine meadows to see the food systems that were contained within them.
我们正在引入褐鳟,并观察将一个新物种放入该体系后会形成的食物网。
We're introducing the brown trout and looking at the food web that would emerge from putting a new species into that system.
那些繁荣兴旺的物种是彩色的,而遭受重创的则是灰色的。你可以把同样的逻辑套用到将 SpaceX 置入某个生态系统中,思考其他所有参与者将如何开始进化。我们在阅读文本时,就能开始看到其中的关联。
The species that thrived were in color. The ones that were decimated were in gray. You can think of the same thing of putting SpaceX into an ecosystem and how the ecosystem of every other player would start to evolve. We can start to see the connections between that when we read the text.
我们能阅读这样一段文字,然后说出这里发生了一个事件:一份 1.6 亿美元的合同,涉及轨道科学公司(Orbital Sciences)和 SpaceX。我们知道事件发生的时间,知道泰科姆公司(Thaicom)参与其中,还知道物体之间存在一个事件。我们可以把它在视觉上呈现出来,但用三维方式呈现会好得多。
We can read a text like this and say that there was an event. It was a $160 million contract, Orbital Sciences and SpaceX. We know when it happened. We know the company Thaicom was involved. We know that there's an event that occurred between objects. We can represent that visually, but much better to do that in three dimensions.
在这张红色图表中,SpaceX 位于顶端,其首次订单连接通过不同的金融事件与其他实体相连。绿色部分则是二次订单连接。这就是全球商业航天产业生态系统的食物网。
On the red here, we've got SpaceX at the top and you've got its first order connections through different financial events with other entities. In the green are the second order ones. This is the food web of the ecosystem of the global commercial space industry.
然后我们就可以操纵它。但这项能力适用于任何系统——汽车、想法、模式、科学论文。突然间,专家们所掌握的这些关联空间,变得人人可及。
Then we can go and navigate it. But we can do this for any system. We can do this for cars. We can do it for ideas. We can do it for patterns and scientific papers. All of a sudden, this space, these kinds of connections that experts hold becomes accessible to any of us.
肖恩·古尔利 增强智能(续)
Sean Gourley Augmented Intelligence (Continued)
此时,这位专家其实无法在脑海里同时装下所有这些物件。我认为这一工具相当具有革命性,因为它能让我们在一个极为复杂的世界里探索、上传、发现模式并做出决策。
At this point, the expert can't actually hold all these objects in the head. I think this tool is quite revolutionary insofar as what it enables us to do to explore, upload, find patterns and make decisions in a very, very complex world.
这些的分析一个有趣应用是叙事分析。这在几个领域相当有意思。一是广告行业和全球公关集团。对上市公司而言,叙事分析同样可用于衡量头条风险这类事项。
One of the interesting uses of this is narrative analysis. This is really interesting for a couple of places. One is the industry of advertising and the publicists groups of the world. It's also interesting for narrative analysis for things like headline risk in publicly traded companies.
我们围绕这些公司在讲述怎样的故事,这些故事是否正确?这些故事又是如何传播开来的?当然,从政治视角来看,这也颇有意思。
What are the stories that we're telling about the companies and are they the right stories? How do these stories propagate through? Also interesting of course from a political perspective.
这里我们可以拿“占领华尔街”来打个比方。当然,关于这群松散聚集的人群,已经写出了成千上万篇报道。但实际情况是怎样的呢?
We can take something here like Occupy Wall Street. Of course, thousands of stories written about an amorphous agglomeration of people. What does it look like?
我们可以再次进行预测,从这里看到所有故事的走向。我们可以看到这些聚类。我们可以看到,黄色部分显示的是所有提及不平等问题的故事,它们正从政治讨论中向外辐射。
We can take a projection of that again and we can see here the projection of all the stories. We can see here the clusters. We can see here, in yellow, all the stories that mentioned inequality as it starts to radiate out from the political discussion.
我们还能看到“占领纽约”与“占领奥克兰”之间的分化,这两个群体的对话内容不同,反映出两者间的裂痕。在此我们看到,共和党与民主党之间并无明确界限,双方都在艰难地寻找政治对策来应对实际发生的情况。
We can also see the separation between Occupy New York and Occupy Oakland, as different conversations were had representing a schism between those two parts. We see here no clear distinction between the Republicans and the Democrats as they struggle to come up with a political response to what was actually going on.
这是事后的分析,当然,你也可以实时操作——随着报道不断涌现、局势持续演变、格局不断变化,你完全可以同步进行。我们可以看到,观点在发布后会像涟漪一样扩散传播。这当然非常强大:无论你是广告公司、政治组织,还是试图监控金融市场动态的人,你都能看到这些故事,甚至能做更多——你可以自己“投放”故事来改变局势……
This is looking at a post hoc, but of course, you can do all the stuff in real time as the stories come in, the landscape evolves, the topology changes. We can see the ideas radiating as they're published. This, of course, is very powerful; whether you're an ad agency, a political group, or you're trying to monitor things within the financial market. You can see the stories, and even do more than that―seed the stories yourself to change that...
提问:肖恩,你有没有回溯历史,然后观察演变过程?
Question: Sean, did you go back in history and then watch it?
肖恩:有的,你可以浏览《纽约时报》的整个档案库。你可以倒带回去,看故事如何演变。你可以对多种不同的数据源做同样的事情。所有内容都有时间戳,所以非常有趣。
Sean: Yes, you can go through a bunch of the whole New York Times archive. You can rewind and see the stories evolve. You can do that with a range of different sources. Everything is time stamped, so really fascinating.
许多对冲基金正在用这个功能,研究 2007 年苹果推出 iPhone 时的市场反响,再与今天微软 Surface 的发布进行对比。
A lot of the hedge funds that are using that are looking at what the launch looked like in 2007 when Apple put out its iPhone and how does that compare to the launch today from the Microsoft Surface?
我们的世界很复杂。我们必须在这个世界中做决策。我们人类的大脑并没有足够的能力应对这一切,但也不存在一个神奇的机器按钮可以按下去就解决。我认为,我们必须走向“增强智能”,借助它来驾驭这一切,把机器最好的部分和人类最好的部分结合在一起。
Our world is complex. We do need to make decisions within it. We don't have all the capabilities within our human brain, but nor is there a magical button that we can push from a machine. We have to reach towards this augmented intelligence, I believe, to allow us to navigate that and to take the best of the machines and the best of the humans and put them together.
我想把这个话题放到背景里来说。当然,我认为 2013 年是数据量有史以来最大的一年。情况已经发展到荒谬的地步,甚至出现了关于数据科学家的咖啡桌书(Coffee table books)和各种古怪的东西。
I want to contextualize this. We have, of course. I believe 2013 was the biggest year for data that we've ever had. It got to crazy points of coffee table books being produced about data scientists and all these kind of weird things.
这种热度如此之高,以至于我最喜欢的一篇论文是在物理学档案网站上发表的,标题是:“我写了一篇关于 Twitter 预测 X 的论文,得到的只有这篇烂论文”。似乎 Twitter 可以预测一切,但最终却只产出了一堆论文。
It was so much so that I think my favorite paper that came out was an article on the physics archive called... ‘I wrote a paper about Twitter predicting x and all I got was this lousy paper’. It seemed that Twitter could predict everything and yet it was only producing papers.
肖恩·古尔利——增强智能(续)
Sean Gourley Augmented Intelligence (Continued)
这股热潮如此之大,以至于大卫·布鲁克斯(David Brooks)不得不撰写一篇专栏,题为“数据做不到的事”。因为数据似乎无所不能。他提出了一些相当有道理的观点。他列出了数据难以应对的四件事。
It got so big, of course, that it got to the point where David Brooks was compelled to write a column titled "What Data Can't Do," because it seems it can do everything. He had some pretty good points. He outlined four things that data struggles with.
第一,数据难以捕捉人类情感的细微差别。它能看出一个人从 A 处向 B 处发送了一封邮件,但它极难捕捉到父亲对儿子之间的爱。
The first was that it struggles to capture the nuance of human emotion. It can see that a person sent an email from A to B. But it really struggles to capture the sense of love between a father and his son.
第二,数据将自己标榜为真实性和客观性的徽章。它只是数字,却常常抹杀掉数据的来源、所依赖的模型,以及其中包含的任何偏见。我们只是盯着数字看。
The second one was that it holds itself up as a badge of authenticity, as a badge of objectivity. It's numbers, whilst oftentimes brushing away where the data came from, any of the models that underlay it, or any of the kind of the biases that were contained. We looked at the number.
第三,非常关键的一点是,随着我们获得越来越多的数据,我们往往依赖相关性这根“拐杖”,而牺牲了对因果关系的探寻。变量越多,相关关系出现的机会就越多。如果我们不开始寻找因果关系,就可能会走上错误的道路。
The third thing was, very much so, as we get more and more data, we tend to lean on the crutch of correlation at the expense of looking for causation. The more variables we have, the more opportunities we have for correlation to occur. If we don't start looking for causation, then we maybe lead ourselves down the path that's wrong.
但也许他最大的批评是:大数据解决不了大问题。我们手握一项了不起的技术,却用它来解决鸡毛蒜皮的小事。
But perhaps his biggest critique was that big data can't solve big problems. That we're sitting here with an amazing technology and we're using it to solve trivialities.
假设有一位数据科学家,坐在硅谷的实验室里,优化儿童早餐麦片及其广告包装——颜色、大小、形状、款式、货架上的摆放位置,再用眼动追踪软件看看人们看哪里。
The hypothetical data scientist sitting in their lab in Silicon Valley optimizing children's breakfast cereal and advertising packaging on top of it to take the color, the size, the shape, the form, the position on the store shelf, to get the eye tracking software to see where people were looking.
我们能让该产品的收益率提高 3%,却从不静下心想一想:“天哪,那肥胖问题呢?糖尿病呢?”那些用简单的信息流优化结构难以真正理解的大问题,却被忽视了。
Also, that we can get three percent better yield on the product without sitting back and saying, "Well, jeez, what about obesity? What about diabetes?" Things that are a lot harder to actually understand with a simple kind of structure of information streams that we seek to optimize.
我认为原因在于——我们需要倒带回到原点——为什么我们手握这么了不起的技术,却只用来优化儿童早餐麦片那 3% 的收益率?我们必须意识到,在过去 6 年里,作为社会,我们进行了一场也许是有史以来最大规模的量化实验。
I think the reason for that, and we need to rewind the clock – why are we here with this amazing technology optimizing archetypal children's breakfast cereals at three percent? I think we have to appreciate that over the last six years as a society we've undertaken one of the biggest experiments in quantification that perhaps we'll ever do.
我们把我们的人际关系、个人信息,通通上传到了 Facebook 这样的结构化数据库中,好让它们挖掘这些信息并向我们投放广告。这是一场巨大的实验。
We've all taken our relationships, our personal information, and we've uploaded it into a structural database from the likes of Facebook, so that they can mine that information and serve us ads. This is a massive experiment.
我们所有人都参与了进来。我们以数十亿人的规模参与其中。我不知道这到底会给我们带来什么,但我们正在做这件事。
We've all run it. We've run it on the scale of billions. I don't know what exactly it's going to do to us, but we're running it.
正是从这场实验中,诞生了我们今天所知的数据科学。它出自两个创造这个术语的人——杰夫·哈默巴赫(Jeff Hammerbacher)和 D.J. 帕蒂尔(D.J. Patil)——在数据科学定义的早期阶段,正是他们塑造了我们今天所看到的一切。
Out of this experiment really came data science as we know it. It came out of that with the two people that coined the phrase, Jeff Hammerbacher and D.J. Patil, and in the very instrumentally early stages of data science in defining what we have today with it.
杰夫曾是 Facebook 的数据分析团队负责人。D.J. 曾是 LinkedIn 的数据分析团队负责人。这些人所定义的数据科学,反映的是他们自身的经验。这很合理。
Jeff was the head of the data analytics team at Facebook. D.J. was the head of the data analytics team at LinkedIn. Data science that's been captured by these people reflects their experience. That's fair enough.
它必然会反映你的经验。
It's going to reflect your experience.
你的经验是:数以百万计的数据点、结构化的信息流、A/B 测试——你可以判断某个东西是好是坏——最终目标就是投放一则广告。
Your experience is millions of data points, structured information streams, A/B testing, so you can know whether something's better or worse with the ultimate goal of serving an ad.
公平地说,他们确实很擅长做这些事。擅长到仅仅通过公开的“点赞”信息,我们就能预测你是否单身、是否处于恋爱关系、是否吸毒、你的种族、宗教信仰,甚至性别。
To their credit, they're really good at doing this stuff. So good that just through the public likes we can predict if you're single or in a relationship, whether you use drugs, your race, your religion, even your gender.
肖恩·古尔利——增强智能(续)
Sean Gourley Augmented Intelligence (Continued)
即使你选择不透露这些信息,我们也能以惊人的准确度预测出来。即使你的朋友也不透露,也一样能预测。我们在所有这些数据中发现,“喜欢薯条”是预测智力的首要指标。
We can predict this information with startling accuracy even if you don't choose to reveal it. Even if your friends don't reveal it either. What we found out of all of this was that liking curly fries is a top predictor of intelligence.
[laughter]
[laughter]
这是个真实的研究结果。这很有趣。但我们对此发笑,是因为它跟智力毫无关系。我们用这些数据来找相关性,却不理解其背后的真实机制。
That's an actual result. That was interesting. But we laugh at that because it says nothing about intelligence. We use this data to find correlations without understanding what's actually going on underneath it.
问题的关键是:如果你是 LinkedIn 的数据科学家,你根本不在乎。你在乎的只是,如果有人想向聪明人投放广告,而你的主页上出现了“喜欢薯条”这个标签,那你很可能就会收到他们卖的任何东西的广告。我们根本没有努力去理解背后发生了什么。
The point of that is if you're a data scientist at LinkedIn, you don't care. All you care is that if someone says I want to advertise to smart people and if you've got curly fries on your home page, you're probably going to get an ad for whatever they're selling. We haven't strived to understand what's happening.
我认为,数据科学在它被塑造的方式、以及它开始谈论数据能做什么的方式上,是相当有局限性的。我认为“数据智能”是更有趣的思考方式。数据科学与数据智能的区别,大概可以这样总结。
I think data science is quite limiting in the way it's being cast and in the way it starts to talk about what data can do. I think data intelligence is a much more interesting way to think about this. Data science versus data intelligence can sort of be summed up like this.
在数据科学中,你在寻找 10% 的改进;而在数据智能中,你在寻求改变游戏规则。数据科学的目标是预测和优化。数据智能的目标是创造和改变。
In data science you're looking for 10 percent improvements, whereas in data intelligence you're looking to change that game. The goal of data science is to predict and optimize. The goal of data intelligence is to create and change.
在数据科学中,决策主要由算法做出,无论是 Facebook 上投放的广告,还是高频交易。算法在做这些决策;而在数据智能中,做决策的主要是人。
The decisions are primarily made by algorithms in data science, whether it's the ads being served on Facebook or high-frequency trades. Algorithms are making those decisions, whereas in data intelligence it's very much the humans.
数据通常是海量且干净的;而在数据智能中,数据是少量的、杂乱的,并且最初并非为当前问题而设计。数据科学中的沟通方式是通过方程式,而数据智能中的沟通方式是通过故事。大体上你可以认为,数据科学是战术性的,而数据智能本质上是战略性的。
The data is often big and clean, whereas in data intelligence it is small, messy, and it wasn't necessarily designed for the problem in front here. The communication in data science is done through equations, whilst in data intelligence it’s stories. You could think, broadly speaking, of data science as being tactical, whilst data intelligence is being strategic in nature.
我们中的一些人,在过去十年处理数据的工作经历中,已经走过了这个阶段。我想和你分享五个我用来思考数据和数据智能,特别是如何用数据解决大问题的启发式原则(heuristics)。
Some of us have sort of gone through that over the last decade of experience working with data. I want to share with you five heuristics that I use to think about data and data intelligence and particularly how to use data to solve big problems.
第一个原则是:数据需要为人类交互而设计。人类能轻松阅读左边的数据,并与它交互、探索它。机器则只能读取右边的数据。我们需要记住我们到底在跟什么打交道。我们需要为我们的数据设计以人为本的用户界面。
The first of these is data needs to be designed for human interaction. A human will read this very well on the left and interact with and explore it. A machine will read the one on the right. We need to remember what we're actually dealing with. We need to design human-centered user interfaces for our data.
第二个原则是:我们需要理解人类处理能力的极限。我们需要知道,我们无法比 650 毫秒更快地做出战略决策;我们大脑中大约只能同时处理 5 到 7 个对象;我们只对大约 150 段人际关系感到舒适。这就是生物进化让我们处于的位置。
The second is we need to understand the limits of human processing. We need to know that we can't make a strategic decision faster than 650 milliseconds, that we hold about five to seven objects inside of our head and that we're comfortable with 150 relationships. That's the kind of biological evolution that brings us to that point.
我们还应该知道,超出人类能力范围之外,就会是算法的天下。
We should also know that beyond the human capability, there will be open game for the algorithms.
这就是高频交易环境的现状——美国股票市场 90% 的交易由算法完成。
We get the high-frequency trading environment, where 90 percent of trades in US equity markets are made by algorithms.
它们会利用这一点,因为它们的思考速度更快。所以要理解人类的偏见,理解人类的局限,但也要理解,当我们无法超越这些局限时,机器就会占据主导。
They'll exploit that because they think faster. So appreciate the human biases, appreciate the human limitations, but also appreciate that when we can't go past it, the machines will dominate.
肖恩·古尔利——增强智能(续)
Sean Gourley Augmented Intelligence (Continued)
数据是杂乱的、不完整的、有偏见的。我们需要知道这些数据到底是什么。我经常说一句话:如果你拿到了数据,却没有用文本编辑器打开过它,如果你用 Excel 打开它并且至少读过一行,那你就不应该使用任何从它推导出的模型。
Data is messy. It's incomplete and it's biased. We need to know what that data is. One of the things I always say is if you've got data and you haven't opened it up inside of a text editor, if you opened it up inside of Excel and read at least one roll of it, you shouldn't be using any models that are derived from it.
有一种数据,比如我们从阿富汗获得的数据——基地组织特工的手写笔记——它们不适合大规模分析,但非常重要。你需要使用它们,但你未必能把它放到数据库里,然后对它做回归分析。我们必须处理好这种情况。
There's a kind of data that we get out of places like Afghanistan―handwritten notes from al Qaeda operatives that don't lend themselves to a lot, but they're very important. You need to use them, but you're not going to necessarily be able to look that up inside of a database and run a regression analysis on top of it. We need to deal with that.
数据需要理论支撑,否则你就会遇到这样的情况:掉进游泳池淹死的人数,与尼古拉斯·凯奇(Nicolas Cage)出演的电影数量高度相关。
Data needs a theory, because if it doesn't, you have things like this: The number of people who drown by falling into a swimming pool is highly correlated with the number of films that Nicolas Cage appeared in.
[laughter]
[laughter]
我知道尼古拉斯·凯奇是个坏人。这是最糟糕的相关性。随着我们获取越来越多的变量、越来越多记录在案的变量,你几乎能为任何事情找到这种相关性。这就是问题的核心所在。
I knew Nicolas Cage was a bad man. That's correlation at its worst. As we get more variables, more variables recorded, you can find this for pretty much anything. This is at the heart.
《连线》杂志的克里斯·安德森说过:“理论已死。”克里斯,这个世界就是你创造的。我们需要理论,是因为我们要理解正在发生的事情。我们需要理论,还因为我们也要做预测。如果未来与过去相似,预测固然很好。
Chris Anderson of Wired magazine said, "Theory is dead." This is the world you've created, Chris. We need theories because we need to understand what's going on. We need theories because we also need to make predictions. Predictions are great if the future's going to look like the past.
但如果你想看到一种不像过去的图景——如果你真的想改变它——你就需要一套理论。这正是物理学登场的地方。我们开始构建理论。物理学的核心,正是用理论来解释经验证据。
But you need a theory if you want to see what it might look like if it doesn't look like the past, if you're actually going to change it. That's where physics comes in. We start to build theories. Very much to the heart of physics are theories to explain empirical evidence.
最终,数据需要故事,但故事也需要数据。我们讲故事,而且讲得很好,因为故事令人难忘。它们能够沟通信息,并在组织内部渗透传播。它们是编码复杂环境信息的极为强大的工具。当然,从我们有记忆以来,我们就一直在讲故事。
Finally, data needs stories, but stories also need data. We tell stories and we tell them very well because they're memorable. They communicate and they percolate through organizations. They're very powerful tools to encode information about complex environments. We're telling them, of course, for as long as we can remember.
尽管我可以向你展示一个描述叛乱动态的方程式,但我也完全可以用九头蛇的神话故事来讲述——这只多头怪兽,你砍掉它一个头,它就会长出两个新头。这种非线性函数蕴含在这个故事中的方式,与它蕴含在那个方程式中的方式是一样的。不过,这个故事或许会比那些方程式流传得更久远。
Whilst I could show you an equation for the dynamics of insurgency, I could equally well tell you a story about the myth of the Hydra, the multi-headed beast that comes in and cuts off one head only for two more heads to appear. This non-linear function is encoded in that story in the same way as it's encoded within that equation. But this story will be probably told a lot longer than those equations will be.
大数据能够解决大问题。确实,它必须解决大问题,但它将在一个更类似于数据智能的框架内来完成这件事。要做到这一点,我们需要理解人机之间的交互界面。
Big data can solve big problems. Indeed, it must solve big problems, but it's going to be doing it with a framework more akin to data intelligence. To do that, we need to understand that interface between humans and machines.
简而言之,我认为我们需要打造更好的半人马。半人马是神话中生活在森林里的生物;半人半马,头部汇聚了智慧与洞察,你可以去向它们咨询面临的重大疑问。
In short, I think we need to build better centaurs. The centaurs were the mythical creatures that lived in the forest; half man, half horse, the head piles of wisdom and insight that you would go and consult for the big questions that were around.
[laughter]
[laughter]
你甚至可以直接输入那个搜索关键词,针对尼古拉斯·凯奇的电影做优化,然后得出一个专为社交媒体点赞设计的地缘政治策略。再把数据接入无人机,由它来做决策。
You could actually type that search query, optimize it against Nicolas Cage films and come up with your geopolitical strategy optimized for social media likes. Plug into a drone that would make decisions.
我和迈克尔 [莫布森] 在 [圣塔菲研究所] 聊过这个。他说,“这有些影响。”我说,“没错,确实有。”我们聊了聊那些影响具体是什么。我想稍微梳理一下,尤其是算法设计方面的。
I was chatting with Michael [Mauboussin] up at [the Santa Fe Institute] about this. He said, "That's got some implications," I said,"Yeah, it does." We talked about what some of those implications were. I want to run through a little bit of that and particularly algorithm design.
肖恩·古尔利 增强智能(续)
Sean Gourley Augmented Intelligence (Continued)
如果我们为在线互动设计算法,可以取两个非常简单的向量。你的历史记录——我是记住你全部的历史,还是完全不记住?以及你的身份——你能成为任何你想成为的人吗?能是分身吗?还是你必须做真实的自己?
If we design algorithms for online interaction, we could take two really simple vectors. Your history – do I remember all of your history or do I remember none of it? And your identity – can you be anything you want to be? Can you be an avatar? Or do you have to be you?
用这两个维度可以很直观地理解。我可以仅凭这两个维度来设计一套在线社交互动模式。结果会怎样?右上角是 Facebook——身份和历史。事实证明,这是一个利润非常丰厚的广告销售空间。
You can use these two vectors quite simply. I can design a set of social interactions online with two vectors. What do I get? Up onto the right is Facebook―identity and history. Turns out that's a very profitable space to sell ads to.
往左上角看,这是 Reddit。了解 Reddit 的人都知道,你在上面可以随心所欲扮演任何角色,但它确实会记住你的真实身份,并根据你的身份分配你的声望值。
Up into the left, we've got Reddit, and for those that know Reddit, you can be anyone you want on Reddit, but it does remember who you are and it assigns your karma based on that.
往下看左手边,那是《财富》杂志。如果你没去过《财富》杂志,千万别去,因为你可以想象得到:当你不让人拥有身份标识、也不记得他们做过什么的时候,会发生什么。
Down to the left, you've got Fortune. If you haven't been to Fortune, don't go to Fortune because you can imagine what happens when you don't let people be an identity and you don't remember anything they do.
你得到的是一个古怪的 12 岁孩童的游乐场,它偶尔闪现天才的光芒,而且大部分手段我们在《财富》杂志上都见过,但它很快就变得不屑一顾了。
You get a kind of a weird 12-year-old playground, which is at times genius and most of the means that we ever see come out of Fortune, but it denigrates pretty quickly.
在糟糕的一端,我们有 Snapchat——一个真实的身份,却没有历史。我们都知道这样做的后果。或者,也许你的孩子知道。
On the bottom side, we've got Snapchat, a real identity and no history. We all know what happens when you do that. Or maybe your kids do.
[laughter]
[laughter]
你最好别知道。这四个系统分别占据四个不同的象限。仅通过操控历史与身份这两个维度,我们就得到了四种截然不同的行为模式。从“我的身份永远会被铭记”到“它永远不会被铭记”这两种极端之间,我们的行为方式会大相径庭。
You don't want to know. These four systems occupy these four different quadrants. By just manipulating history and identity, we get four very different types of behaviors. From the kind of behaviors we have where my identity is always going to be remembered for me, versus it's never going to be remembered and we behave very differently.
我们如何设计这些系统,以及我们想要什么样的系统?因为它们正在被推广给数十亿人,并且确实影响了我们的思维方式、互动方式和社交方式。
How do we design these systems and what kind of systems do we want? Because they are being rolled out to billions of people and they do affect how we think and how we interact and socialize.
我觉得这确实很有意思。我们一直在思考该如何设计算法。但我也想问,当这些算法真正能帮助我们更好地思考时,会发生什么?当你有了增强智能,又会怎样?这会在平等或不平等上造成什么样的局面?
I think that's really interesting. We think about the stance of how do we design. But I will also say what happens when you have these algorithms that actually allow us to think better? What happens when you have augmented intelligence? What sort of equality or inequality that starts to create?
我们可以把这个模式推广开来,探讨算法和人类互动的不同方式。我刚才用 Facebook 的例子展示的,正是大多数人目前在世界上与 Facebook 和 Google 互动的主要方式——通过过滤算法获取信息。
We can roll that out and think about the different ways the algorithms and humans can interact. What I showed you there with the Facebook example is the dominant way most of us are interacting in the world through Facebook and Google to get information through these filtering algorithms.
这些都是免费系统。我们不用给 Google 付费,它们基本上是免费的。交换条件是,我们被当作产品来对待。我们被算法实时买卖,就像华尔街那些高频算法一样,来回交易、互相竞价,就为了向你展示一则广告。
These are free systems. We don't pay for Google and they're sort of free. The tradeoff is that we're being treated as a product. We're being bought and sold literally in real time by algorithms, much like the high frequency algorithms that we have on Wall Street, trading backwards and forwards to bid against each other to show you an ad.
你们所看到的信息就是被那个机制筛选过的。这带来了两个结果。一个是它给运行这套算法的成本设了个上限。Facebook 每位用户每年的系统成本是 1.50 美元。谷歌大约是 12 美元。它们必须控制在这个上限以下,因为能从广告里赚到的钱就这么多。
The information that you've got there is being mediated by that. That does a couple of things. One is it puts a cap on how much it can cost to run that algorithm. The entire system per user per year for Facebook costs $1.50. Google costs about $12. They have to sit under that cap because that's all the money you can get from advertising.
要从广告赚到几百美元是极其困难的。算法的复杂度有其上限,但算法的目标未必是为你提供最优的信息,帮你养成正确的思维方式,从而拥有最优质的人际关系、朋友、动态更新,以及关于太空新闻的准确信息。算法的目的是向你推销广告,这一点我们应该牢记。
It's very hard to get into the hundreds of dollars from advertising. The complexity of the algorithms is capped, but the goal of the algorithm is not necessarily to give you the best information, to make you think the right way so that you have the best personal relationships, friends, and updates, the right information about space news. It's to sell you ads and we should remember that.
肖恩·古尔利 增强智能(续)
Sean Gourley Augmented Intelligence (Continued)
我们将会看到这些算法的所有权开始发生转移。不再是谷歌拥有它,并且为了谷歌的广告需求而优化,而是会变成这样:“你知道吗,12 美元,我大概会花这个钱。事实上,我可能甚至愿意付几百美元,但我要拥有这个算法。”
What we're going to see is the ownership of these algorithms start to switch. Instead of it being owned by Google, and being optimized for Google's advertising needs, we're going to go "You know what, $12, I'm probably going to spend that. In fact, I might even pay a couple of hundred dollars, but I want to own the algorithm."
我希望能把它为我优化。我希望它能把我推向我未曾思考过的地方,提醒我与未曾联系的人保持联系,如果我在给自己塞太多垃圾信息,它能约束我。我希望它能为我量身优化,我准备好为此付费,因为信息以及我们思考的方式,价值无比巨大。
I want it optimized to me. I want it to push me to places to think about things I haven't thought of, to make sure to remind me to connect with people I haven't connected with, to kind of constrain me if I'm feeding myself too much junk food type of information. I want it to optimize for me, and I'm prepared to pay for that because information and how we think is incredibly valuable.
当然,总有些人穷到根本不值得广告算法为他们费心,而这些人实际上将为算法打工,成为所谓的“技术农奴”。这听起来有点科幻,但你看亚马逊土耳其机器人(Mechanical Turk),再看看许多这类微支付网站——你根本赚不到足够的钱。于是算法会说:“你知道吗?
Of course, there are going to be people who don't have enough money to be interesting at all for advertising-based algorithms, and they're actually going to be working for the algorithms, the sort of techno serfs of the word. This seems kind of science fiction, but you look at Mechanical Turk, you look at a lot of these micropayments sites, you can't earn enough money. So an algorithm says, "You know what?
我还没聪明到能判断这张图是否适合工作场合。人类,这张图安全吗?这是某个人的生殖器照片吗?
I'm not smart enough to determine whether or not this picture is work appropriate. Human, is this a safe picture? Is this a picture of someone's genitalia?"
它们确实做得到。这类工作你只需花一两美分就能完成——那些算法无法做到的工作。
They do that. That's the sort of thing you pay one or two cents to do -jobs that the algorithms can't do.
当然,在它的另一端,增强智能正在兴起——人类不仅拥有算法,还主动驱动算法,与它交互。人们不只是接受预测,而是真正开始操控和理解它。
Of course, on the far side of that, you've got augmented intelligence where humans are not just owning the algorithms, but they're actively driving the algorithms. They're interfacing with it. They're not just taking the prediction, but they're actually starting to manipulate and understand.
我认为我们将从我们习以为常的模式——每个人都从这些广告驱动的平台上获取信息——转向一些不同的东西。你可以看到,这也随之带来了巨大的不平等。
I think we're going to move from what we've seen; where everyone gets information from these advertising-driven models, to these different things. You can see here that it brings with it a huge amount of inequality as well.
其中一个问题是,随着我们构建并推出这些系统,我们需要思考这样一个社会将如何运转——一部分人拥有的算法能让他们真正看得更远、思考得更清晰,而另一部分人却在被投喂着垃圾食品般的算法,从谷歌这类平台上接收着信息。
One of the things is, as we build these systems and roll them out, we're going to need to think about what do with a society where some people can own algorithms that make them literally see further and think better, whilst other people are being fed junk food algorithms, serving them information on the likes of Google.
就这一点而言,争议性大概已经足够大到可以开放提问了。谢谢大家,我很乐意回答任何问题。
On that, that's probably controversial enough to open up to questions. Thank you, and I'm happy to field any and all.
[applause]
[applause]
问题:在你向我们展示的界面中,你稍微提到了让分析师有能力提供一些监督。那么,所有这些监督数据是否有能力去训练整个系统,从而让它更好地……你是在上传所有这些信息吗?
Question: In the interface that you've shown us, you talked a little bit about the ability for analysts, to make the analysts then provide a little bit of supervision. Is there any capacity for all that supervision and all of that to then go and train the total system, so that it could better...you're uploading all that information?
肖恩:完全正确。随着用户增多,他们能添加或删除关联,也能修改我们给出的猜测。系统在运行中不断学习。我们现在拥有一个非常完善的架构。
Sean: That's exactly right. As we get more users, they're able to add and delete connections. They're able to change the guesses that we make. The system is learning as it goes through. We have a very nice setup now.
问题:但这样一来,你们软件的所有用户不都能从那些改进中受益吗?
Question: But all users of your software then benefit from those changes?
肖恩:没错。
Sean: That's right.
问题:这有点像维基百科。
Question: It's sort of like Wikipedia.
肖恩:没错。你的搜索记录别人看不到,但算法会自我调整:“我给词性标注这部分赋的权重太高了,应该调低一些。在猜测名字时,我应该……”
Sean: That's right. They won't see your searches, but the algorithm will say, "I weighted too much the part of speech tagging, I should bring that weight down a little bit. When I'm guessing names, I should be
肖恩·古尔利增强智能(续)
Sean Gourley Augmented Intelligence (Continued)
“对于使用形容词就不那么开心了”,诸如此类。这会改变用于命名和链接的猜测权重。
a little less happy to use adjectives," and so on and so on. It will change the weights of the guesses that are made for both the naming and the links.
提问者:善于与电脑协作的人类,是天生的还是后天培养的?
Question: Are humans who are good at collaborating with computers born or made?
西恩:这是个好问题。要回答它,我觉得你得思考一下,什么能让一个人擅长与计算机协作。我想到了那两位在国际象棋上赢了计算机的人,他们一个是数据库管理员,一个是足球教练。你可以从两个侧面来看这件事。你确实需要理解数据,需要知道数据从哪里来、在哪里、正在发生什么。
Sean: That's a good question. To answer that, I think you need to think about what makes someone good at collaborating with a computer. I think of the two people that won the computer chess were the database administrator and a soccer coach. You can kind of think of two sides of that. You do need to understand the data, you need to have a sense of where the data has come from, where it is, what's going on.
你需要理解算法是怎么执行的,是怎么编程的。理解它怎么编程,才能知道它会模拟爬上局部最高点,但通常会错过全局最高点。你需要掌握那套语言才能明白这一点。但我认为,你还需要懂策略,并在自己内心形成一种感觉。所以你对自己得有充分的认知。
You need to understand how the algorithm is being performed, how it's being programmed. To understand how it's been programmed, to know that it's going to do simulate to climb a local maximum, but it generally misses the global maximums. You need to have that language to know that. But I think you also need to know, in a sense, a strategy and gain internal sense yourself. So you have to have a good awareness yourself.
同时具备数据知识,以及对自身和周围世界的认知,这种情况相当罕见。
That's quite rare to have both a knowledge of data, and a knowledge of yourself and the world around you.
我倾向于画一个韦恩图,一边是理解数据和算法的人,另一边是理解世界的人。这个韦恩图中间的交集,也许就是擅长人工智能的人,但这个交集非常小。我认为我们可以训练人们去理解算法,也应该这样去做。
I think I'd sort of draw a Venn diagram of people that understand data and algorithms, and people that understand the world. The Venn diagram is maybe the people that are good at artificial intelligence, but that's pretty small. I think we can train it and we should train people to understand algorithms.
我觉得我们可以把很多被推向数学和科学的孩子拉出来,确保他们也理解政治、哲学和经济学。这样我们也许能在这个过程中帮到自己。
I think we can take a lot of the kids that we're pushing into math and science to get out and make sure they understand politics, philosophy, and economics. We probably can help ourselves along that way.
我认为我们还应该认识到,不是每个人都能为算法做贡献。这会产生相当深远的经济影响。如果我们有一群人能为算法做贡献,另一群人的工作正被算法取代,那我们会落到什么境地?
I think we should also appreciate that not everyone is going to be able to add to an algorithm. That's got pretty profound economic implications. If we've got a set of people that can add to algorithms, and a set of people whose jobs are being replaced by algorithms, where does that leave us?
迈克尔·莫布森:菲尔,你对这个有什么想法?
Michael Mauboussin: Phil? Do you have a thought on that?
菲尔·泰特洛克:这个问题问得真好。要回溯到航天工业的计算语言学有多难?
Phil Tetlock: A really good question. How hard would it be to go back to the computational linguistics of the space industry?
我在情报界做过一些关于开源指标的工作,我很好奇我们离某些基准还有多远。
I have done a little bit of work with the intelligence community on Open Source Indicators, and I am curious how close we are to certain benchmarks to others.
其中一个问题是,你们的系统能否回答这个问题:我知道 NASA 的高层高管访问过中国,而不实际提及,也不用那些新上映电影里一模一样的词。
One question would be whether your system could answer the question, I know high level NASA executives visited China, without actually mentioning, not using the exactly same terms as those in the newly-released films.
西恩:我们的系统能回答“NASA 的高层高管访问过中国吗”这个问题吗?这里面有几个要点,一个是我们知道查尔斯·博尔登,知道他是高管,知道他是 NASA 局长,知道中国是一个实体,知道去那里需要地理上的移动。我们并没有把系统设计成那样来回答这类问题。
Sean: Could our system answer the question, has a high level NASA executive visited China? There are a couple of points in there, one is we know [Charles] Bolden, we know that he is an executive, we know he’s the chairman of NASA, we know that China is an entity, we know the geographic thing to have traveled. We haven't designed the system to answer those questions like that.
对我们来说,我们把它设计得非常面向人类交互,这样你可以搜索关于 NASA 的一切,很快就能看到像你刚才发现的那样一个聚类。那里面就有 NASA 去中国的聚类。现在,我们可以让这个过程更快,让计算机也知道这一点,但这也有可能把我们引向人工智能的方向。
For us, we have designed it very much for humans to interface to it, so that you'll search for everything that's going around NASA, and you'll see a cluster there very quickly, like you did. There was cluster of NASA going to China. Now, we can make that quicker and have the computer know that but also, it might lead us down the path of artificial intelligence.
西恩·古尔利谈增强智能(续)
Sean Gourley Augmented Intelligence (Continued)
提问者:有趣的是,这可能正好和费鲁奇的演讲联系起来——沃森就能做到这一点。另一个问题,我不确定这是增强智能还是明天就能用到的问题是:我们离能够就未来可能性提问,比如向“明天的职业区”提问,还有多远?
Question: What is interesting is that it probably ties into the Ferrucci talk – Watson could do that. The other question, and I am not sure whether this is augmented intelligence or this will be used tomorrow, would be: how far are we from being able to pose questions about possible futures, to the occupational district tomorrow?
中国会在 2020 年前实现载人登月吗?还是说它会再测试一次反卫星武器?
Is China going to have a man on the moon by 2020, or is it going to test another anti-satellite weapon?
西恩:我们离在这些模型之上做预测还有多远?中国会在 2020 年实现载人登月吗?我们会看到反卫星武器吗?
Sean: How far are we away from predictions on top of these models? Is China going to have a man on the moon in 2020, are we going to see anti-satellite weapons?
你可以向计算机提出一个问题,计算机可以吐出一个答案。但我要说,你不会只是觉得“哦,酷,它说了这个”。我们很可能围绕那个预测建立地缘政治政策。你得再仔细看看。即使今天它能做到,我们会信任它吗,我们会真的花几个小时去研究一个计算机输出吗?
You could ask a question that to the computer and the computer could spit back an answer. I would contend that you wouldn't be like, "Oh cool, it said that." We'll probably build around geopolitical policy on top of that prediction. You got to want to look a little closer. Even if it could today, would we trust it, and would we really kind of spend hours for just a computer output?
就算这一点已经实现了,我要说的第二点是,这是一个情报问题。我们有一堆情报分析员在做这个。他们现在是通过阅读海量信息,把微弱的信号拼接起来,建立模型来预测他们认为会发生什么。
That's if it was there, the second I would say is, that's an intelligence question. We have a bunch of intelligence analysts doing that. They do it at the moment by reading through scores of information, to plug weak signals to create a model about what they think will happen there.
考虑这个工具的作用,它把所有那些阅读工作压缩起来,提取出分析员会建立的那些关联,并开始替他们建立这些关联。
The point of thinking what this thing does is, it takes all of that reading and compresses it right down, and takes the connections that they would make, and starts to make them for them.
他们可以拿到一周的信息,我们知道这两件事情都有发生的迹象,然后他们就能开始整合这些信息。分析员肯定还有用武之地。分析员应该接受这些信息。我不认为计算机已经到了能做这个的地步,也不认为未来十年内能达到。
They can get the week's information, and we know that there's evidence of both those things happening and then they can start to assimilate that. There's definitely still a role for the analyst. The analyst should take that information. I don't think the computer's at a point where they can do that nor do I think they will be there in the next 10 years.
提问者:我也是这种感觉。我们可以在一个大型组织内部做增强智能实验,告诉你结果。一些分析员能接触到你的模型,一些不能,这种能力能带来多大的价值?
Question: That was my sense too. We could do augmented intelligence experiments within a large organization that tells you. Some analysts get access to your model, some don't, how much value is there in the ability to perform?
西恩:当然,我们宁愿先从三大咨询公司开始,选了其中两家。你可以猜猜是哪两家。我们对此做了 A/B 测试。他们真的组建了一个四人分析团队,来做一项并购活动的市场细分。他们问,要花多长时间?能找到多少目标公司?做了什么样的聚类?
Sean: Absolutely, we would rather start with the three major consultancies, and two of the three. You can guess which two of the three. I did A/B test on that. They actually set up a team of four analysts to do market segmentation for an M&A activity, they said, how long will it take them? How many targets do they find? And what sort of clustering did they do?
他们花了大约四周时间,一位负责人加三位助理。他们找到了 200 家目标公司,分成了五个不同的小组。而运行 Quid 的十个人做了同样的实验。他们在三天之内完成,找到了 1500 家公司,分成了 25 个小组。
It took them about four weeks; a principal and three associates. They found two hundred target companies, and they split it into five different groups. The 10 that are running Quid did the same experiment. They did it in three days and found 1,500 companies, and they split them into 25 groups.
这些人都不便宜,你不仅做得更快,而且做得分辨率更高。
These are not cheap people, and you are not only doing it faster, you are doing it with a higher resolution.
这是事情的一方面。我要提到的第二个轶事是一家大型搜索引擎公司,他们把 Quid 用于竞争情报流程。他们做了 A/B 测试,因为他们原本一周能做完的事,现在一天就能完成。
That's on one side of things. The second anecdote I’ll note is a major search company that runs this for their competitive intelligence process. They ran the A/B tests, since they could do what they did in a week, they could do it in a day.
提问者:我很好奇你对那三家的评估。为什么恰好是这三家?因为你描述的一部分主要是基于偶然情境的,知道什么时候该激进,什么时候不该。
Question: I was curious about your assessment of the three. Why the exact three? Because part of your description was mainly circumstantially based, knowing when to be aggressive and when not to.
大概来说,那些国际象棋特级大师就是这么做的。这些人应该也能成为特级大师……
Presumably, the grandmasters, that's the way they do it. These guys should be able to be the grandmasters...
西恩:这可不是一次性的事件。它一直在被重复验证。
Sean: This was not a one-off. It's consistently being done again.
西恩·古尔利谈增强智能(续)
Sean Gourley Augmented Intelligence (Continued)
我认为,人们之所以赢的原因,当然开始有各种推测。有趣的是,如果他们单独与其他人对弈,他们根本赢不了任何一个对手。
I think the thing is why people won, of course, there starts to be conjecture on that. The interesting thing was that they wouldn't have beaten any of the other players had they played them by themselves.
我们知道这一点。他们做了别人没做的事是什么?我的猜测是,当你被训练成特级大师时,你接受的是与跟机器交互截然不同的训练过程。知道如何像人类一样下棋,与知道一组算法如何下棋,完全是两码事。
We know that. What did they do that the other people didn't do? My guess is that when you become trained as a grandmaster, you are trained in a very different process compared to interacting with a machine. There is no sense that knowing how to play chess as a human is knowing how a set of algorithms play chess.
要知道这个,你需要真正理解算法是怎么下棋的,这意味着你需要理解背后的代码。你还需要知道它们是怎么被训练的,在残局中是怎么训练的,在开局中是怎么训练的。你想知道它们什么时候在爬局部最高点,什么时候你很有把握那是全局最高点。
To know this, you need to really understand how algorithms play chess, which means you need to understand the code behind it. You also need to know how they've been trained, how they've been trained in their endgames, how they've been trained in their opening games. You want to know when they're climbing local maximums and when you're pretty confident it’s a global maximum.
我认为,了解这些,了解算法的运作方式,我并不认为特级大师们一定都懂这些。也可能有懂这些的特级大师。那会很有意思。
I think knowing that, and knowing workings of algorithms, I don't think grandmasters necessarily know all of that. There can be grandmasters that know that. It would be interesting.
其中一件有趣的事情是,通过电脑游戏,人们现在下的棋不一样了。他们都是在网上与其他人、与机器对弈中长大的。这实际上改变了开局风格,我们可以追踪计算机引入前后国际象棋开局风格的变化。
One of the interesting things is through computer games, people are playing different chess now. They've all been brought up playing online computer games against each other, against machines. It actually changes the opening styles and we can track the opening styles of chess pre- and post-computer induction.
我们下棋的方式发生了显著变化。我认为这也很有趣。就理解这到底是什么而言,这还处于非常早期的阶段,也回扣了前面那个问题:是什么让人与机器协作得好。毫无疑问,随着这在经济上变得越来越重要,我们会越来越理解它。
There is a significant change in how we play. I think that's also interesting. This is pretty early stuff in terms of understanding what it is and to gather the earlier question, what it is that makes a human work well with a machine. No doubt, as that becomes more economically important, we're going to get better at understanding it.
提问者:关键难道不是特级大师处理信息组块的方式吗?
Question: Isn’t the key the way grandmasters deal with the chunking of information?
西恩:我认为信息组块是一件非常有趣的事情。早期,我们大脑的前楔叶,也就是模式识别区域,会抽象出一系列复杂事物,然后说——是的,我在这里看到了一个信号。这是对那个高维信息的一个低维表示。我们天生就具备这种数据组块的能力。
Sean: I think the chunking of information is something that's really interesting, early on the precuneus, the pattern recognition of our brain that abstracts a series of complex things to say- yes, I've seen a signal in this. That's a low dimensional representation of that high dimensional information. Something really innate to us is that chunking of data.
这是人类几乎必然会做的事情之一。几乎总是五到七个组块。从来没人回来说,你知道吗,我在伊拉克发现了 123 个叛乱组织。
One of the things that humans almost invariably will do. There will almost invariably be five chunks or seven chunks. No one ever comes back and goes you know what, I've seen 123 insurgent groups in Iraq.
总是七个组块。
There's always seven groups.
那是因为这差不多是我们能处理的极限。大多数人只能处理三个。我们会把信息组块抽象到我们能处理的层面。
That's because it's about as many as we can handle. Most people are dealing with three. We chunk information at an abstraction that we can deal with.
计算机也做同样的事。它们把棋局组块化。我认为有趣的一点是,虽然计算机在解决国际象棋方面做得相当不错,但它们在解决围棋方面却做得一塌糊涂。
Computers do the same as well. They chunk the game. One of the interesting things I think is computers, although they've done a pretty good job of solving chess, they've done a terrible job of solving Go.
国际象棋的状态空间大约是 10 的 40 次方步。围棋则是 10 的 150 次方步。
Chess has 10 to the power of 40 moves as a sort of state space. Go has 10 to the power of 150 moves.
这看起来没差多少,但实际是一个巨大的跨越。它们甚至基本上赢不了围棋新手。围棋就是那个用棋子下的棋,如果你玩过的话。它的状态空间大得惊人。
That doesn't seem like a lot, but that's a big jump. They can't even really beat a novice Go player. Go is the game with the stones, if you've played that. It's got a huge state space.
他们解决了这个问题。他们目前最优秀的系统运行在本地笔记本电脑上,而不是巨型超级计算机上,因为它们本质上是在调用启发式算法。它们像人类一样下棋,因为状态空间太大了——即便用我们现有的最大计算机,去探索那种状态空间也就像用茶杯舀干大海一样。
They solve that. Their current best things run on local laptops. They don't run on massive supercomputers because they're basically imploring heuristics. They're playing the game like a human plays it because the state space is so big it would be like going into the ocean with a teacup, even with the biggest computers we have we can't explore that kind of state space.
肖恩·古尔利 增强智能(续)
Sean Gourley Augmented Intelligence (Continued)
观察围棋的发展,以及人工智能将如何开始解决这个系统,我们目前处于什么阶段,这很有意思。我认为有趣的一点是寻找人工智能发展方向的各种信号。它能做什么?不能做什么?在这次演讲之后,我们会看到一些围绕自然语言处理和问答系统的真正有趣的东西,我觉得那将会非常吸引人。
It's interesting to see Go develop and how AI will start to solve that system and where we are with that today. I think one of the interesting things is to look for signals of where AI is going. What it can do? What it can't do? Straight after this talk, we'll see some real interesting stuff around natural language processing and question and answers which I think are going to be fascinating.
我们也应该清楚它现在所处的位置。我们还没有一台机器能够准确判断伊拉克的地缘政治正确行动。考虑到我们连围棋都还没攻克,就不应该指望会有那种能力。我们也许会在解决那个问题之前先攻克围棋,也许吧。
We should also be cognizant of where it is. We do not have a machine that can accurately determine the right geopolitical move for Iraq. We shouldn't expect to have that given that we haven't solved Go. We'll probably solve Go before we solve that, maybe.
问:这对阅读实践来说意味着什么?
Question: What’s this say about the practice of reading?
肖恩:弗兰科·莫雷蒂(译注:Franco Moretti,文学学者)提出了“远距离阅读”这个概念,我觉得这是一个非常有意思的想法。
Sean: Franco Moretti kind of coined the term "distant reading," which I think is a really interesting thing.
我们通常认为阅读就是读文字,但你其实可以在一种抽象层面上去理解正在发生的事情。这之所以重要,是因为每天有 120 万篇英文新闻文章发表。
We tend to think about reading words, but you can actually read at sort of an abstracted level of what's going on. That's important because there are 1.2 million news articles in English published every day.
如果你把阅读当作全职工作,你或许能从中抽取 200 篇。那 200 篇是合适的样本吗?这个样本里隐含了哪些偏见?这其中有多少是金融行业出版物,又有多少是反映消费者真实想法的文章?当你同时读到消费者喜欢那只股票的文章,却看到另一篇文章说那只股票要跌,你有多难摆脱那篇文章带来的偏见?
If it's your full-time job, you may sample 200 of them. Is that the right sample? What are the biases implicit in that sample? How much of that is the financial trade publications versus reading what the consumers are actually thinking? How hard is it for you to unwind that bias about that one article that said that stock was going down even though you’re reading at the same time that the consumers love it.
当我们从 120 万篇里只抽取 200 篇时,我们必然会带有偏见。传统阅读方式下,我们的阅读速度不会变得更快。如果我们想获得那种涵盖全局的事件信息,机器就必须来读这些东西。它们能回答很多事情。比如有一位 NASA 高级官员要去中国了。
When we're sampling 200 from 1.2 million, we're going to be biased. We're not going to get faster at reading traditionally. If we want to get that encompassing event information, machines will have to read that. They'll be able to answer a lot of things. There's a high level NASA official going to China.
更复杂的事情仍然需要人类去与之互动。这就是那个人机交互界面变得极其重要的领域。
The more complex stuff is still going to need a human to interact with it. That's where that interface stuff become really important.
问:这对企业来说会如何运作?
Question: How might this work for a corporation?
肖恩:想象一下,有一家大型广告公司正与一家除臭剂公司合作。
Sean: Imagine if you will there's a large advertising company that is working with a deodorant company.
他们想强调更浓的男性气概。你会在网上看到所有关于男子气概的讨论。这些讨论有不同的话题簇。有一簇故事是关于克里斯·克里斯蒂(译注:Chris Christie,时任新泽西州长)是不是过于阳刚而不适合当州长。还有一些更严肃的,比如 NFL 中的霸凌现象是否可以被接受,还有胡须这个话题,它总是会出现。
They want to be more manly. You get a conversation of all the stories about manliness. There are different clusters on manliness. There's a cluster of stories around whether Chris Christie is too manly to be the governor. Or some more serious, bullying in the NFL and is that acceptable or not, and facial hair, that's always going to be there.
这是在一亿美元广告投入之前的状态。无论这家公司如何进入那个领域,你几乎可以确定的是,那个网络、那个空间、那种投射会变得不同。他们可以预测讨论将如何演变。但他们选择怎么做将极大地塑造未来演变的方向。
That's pre-hundred million dollars of advertising spent. However this company moves into that space, you can be sure as hell that network, that space, that projectionis going to be different. They can predict how the conversation is going to evolve. But what they choose to do will massively shape what evolves.
我可以在不做决策的情况下进行预测,但这会有点奇怪。更好的做法是问自己,我应该传递消息中的哪一部分。我应该合并两个不同话题簇吗?我应该占据一个空白地带吗?我应该拿下被竞争对手占据的一个话题簇,然后自己开始主导它吗?你可以对政治做同样的事情,无论你站在哪一边。
I can predict without me making a decision, but it would be kind of strange. A much better thing to say is what part of that message should I take. Should I combine to get two different clusters? Should I take a white space? Should I take one that's owned by a competitor and start to own that myself? You can do the same thing for politics, which side of the fence you're on.
我们进入那个领域时所做的决定会改变事情。这在任何你构建的系统中都必须被考虑到。
The decisions that we make as we go into that space will change things. That has to be accounted for in any kind of system that you're building.
问:你会如何改进这个系统?
Question: How would you improve this system?
肖恩:我会说,获取更多观察数据。你读了一篇关于胡须彰显男子气概的故事,并不了解全貌,所以你需要更多。我要说的是,这个系统反映的是世界上已经发生的一切——所有关于 X 的故事,所有关于 Y 的事件。
Sean: What I say to this is get more observations. You read one story about the manliness of facial hair and you don't understand the whole thing, so you get more. What I would say on this here is it is reflecting back everything that's happened in the world. All the stories about X, all of the events about Y.
肖恩·古尔利 增强智能(续)
Sean Gourley Augmented Intelligence (Continued)
它不是向前做预测。哪些数据是重要的?我认为这非常关键。它不是向前走并说我们认为这将要发生。针对如何更好地决策未来、知道未来会有什么这个问题,我的回答是:让做决策的人能够接触到更多信息。给他们一个更好的世界图景,更高的分辨率,更快的速度,这样他们就可以在此基础上更深入地思考影响和后果。
It's not making projections forward. What data is significant? I think that's really important. It's not going forward and saying we think this is going to happen. To the point of how do we make better decisions about the future and what's going to be there, I'd say is make more information accessible to the people that are making the decisions. Given them a better view of the world and higher resolution quicker so they can think more about the implications on top of that.
今天,做决策的仍然是人类。当年坐在五角大楼会议桌前决定向伊拉克派兵的就是一群人。他们带着偏见,有自己的意识形态。他们没有掌握所有信息。那些信息并不在他们指尖。我们就是这么做的。
Humans are making these decisions today. There were a group of people that sat around the Pentagon table that decided to send troops into Iraq. That was biased and they had their own ideologies. They didn't have all the information. They didn't have all that stuff at their fingertips. We are doing this.
我把这个问题想成:“在未来十年里,我们如何才能获得一套工具,帮助人类仍然能更好地做出这些决策?”
I think of the question as, "Over the next decade, how can we get a set of tools that can help humans still make those decisions better?”
问:你如何比较不同的模型?
Question: How do you compare models?
肖恩:我认为这归结于预测是如何记录你的预判,并将其与实际发生的情况进行对比的。如果你能做到这一点,那么给你足够的数据点,你就可以算出你的分析师在使用这个系统与另一个系统时预测的差值。
Sean: I think that comes down to the way the forecast is recording your predictions and recording them against what actually happened. If you can make that, then you should be able to do a delta given enough data points of your analyst predictions with this system versus another system.
我想争议点在于,如果你能给他们一个更好的对当今天世界的呈现,他们就能在明天做出更好的决策。
I think the contention would be if you can give them a better representation of the world as it is today, to make better decisions on top of that about tomorrow.
问:目前正在做的不同的事情有哪些?
Question: What are some of the different things being done today?
肖恩:一个是全球宏观方面。我需要了解克里米亚正在发生什么。那里发生了很多事情,而我如何理解这些事情将改变我对石油期货或天然气的定价方式。
Sean: One is global macro stuff. I need to understand what's going on in Crimea. There is a lot of stuff going on and how I understand that is going to change the way I price my oil futures or my natural gas.
目前做这件事的方法是阅读大量数据流,而这个系统的作用是让这个过程更快、更高效、更准确。另一个是看多/看空公开股票。你会获取所有信息,特别是围绕标题风险的信息。如果一个事件发生了,并在社交领域爆发,那是真实的事件还是捏造出来的?
The way that’s done at the moment is reading a lot of data streams and what this does is it starts to make that quicker and more efficient and more accurate. Another one is looking at long-short public equity. You take all the information, particularly around headline risks. If an event happens and it blows up in the social sphere, is that something that is real or is it something that's actually just fabricated?
围绕不同主题的标题风险会呈现出哪些不同的统计模式?这不仅仅是说“这个情绪是正面的还是负面的”,而是要真正去问:“这些被讲述的故事核心是什么?”
What are the different statistical patterns that emerge from headline risks around different topics? It's not like just saying, ‘Is this sentiment positive or is sentiment negative,” but actually getting it to say, “What are the heart of the narratives that are being told”?
当你稍微深入到私募股权领域,你就开始审视世界上所有不同的硬件公司。
Once you go down a little bit into private equity, you are starting to look at the selection of all different hardware companies in the world.
全世界可能有 10000 家不同的硬件公司。我应该投资哪一家?或者应该收购哪一家?
There might be 10,000 different hardware companies. Which one should I invest in or which one should I acquire?
对我来说,在金融领域能用这个做的最有趣的事情是,审视我们给自己讲述的关于世界的故事,然后拿一个关于实际正在发生的故事的正交数据源,开始利用两者之间的差值进行交易。
I think for me, the most interesting stuff that you can do with this in finance is to look for the stories that we are telling ourselves about the world and then to take an orthogonal data source about stories that are actually happening and start to trade off the delta between them.
这可以非常简单:我接入印度新闻,只看印度人讲述的关于 X 的故事,再接入美国新闻,寻找两者之间的差异。或者,我也可以拿关于我们无人机行动的故事,再拿波音公司提交的所有专利的故事,寻找波音正在做但我们没在谈论的事情,或者我们谈论得太多但根本没有发生的事情之间的差异。
That could be very simply saying, I'm going to plug in the Indian news and just look at the stories the Indians are telling about X, and put in the American news, and look for the differences between the two of them. Or, I might take the stories about what we're doing with drones, and take the stories of all the patents that are being filed by Boeing, and any differences between what Boeing is doing that we're not talking about or what we're talking about too much that isn't being done.
肖恩·古尔利 增强智能(续)
Sean Gourley Augmented Intelligence (Continued)
从这个意义上说,这就像是在此基础上做一波动音的仓位。两个不同的数据源,也许没有人完整阅读过它们,但现在确实可以同时读取两者,把机器用上去,寻找差值并据此进行交易。
That in a sense like a trade on Boeing on top of that. Two different data sources that probably no one has read in their entirety is certainly now read across both of them, put the machine on that, and look for deltas and trade against it.
问:你有什么个人经历可以分享吗?
Question: What are some of your personal experiences?
肖恩:有几件事。第一是,如果你觉得自己什么都知道。你会想,“我全搞定了。”
Sean: There are a couple of things. One is if you think you know everything. You think, “I've got it all.”
这是一件非常难的事情。
That's a really hard thing.
有一件事一直让我感到谦卑:我会调出过去两年所有关于自然语言处理的专利申请。我总会发现一些我完全不知道的东西正在被做出来。
One of the things I'm consistently humbled by is, I'll throw up all of the patents that have been filed in the last two years on natural language processing. I would always find stuff that's being done that I wasn't aware of.
我就拿上周的科技新闻来说。我住在硅谷。我生活、呼吸、喝咖啡、全身心投入所有事情。我把这些新闻拿来,在这个数据流上运行一个事件提取算法,我会看到一些东西然后想,“我靠,我怎么会错过那个?”
I'll just take the last week of technology news. I live in the Valley. You live, breathe, drink, do the whole thing. I'll take that, I'll do an event extraction algorithm on top of that data feed, I will see things and go, "How the hell did I miss that?"
那会是一个大事件。比如英特尔出了新芯片。我怎么就错过了呢?那些自认为了解一切的行家们很难接受这一点。我认为始终记住这一点很重要:世界上有太多你不知道的事情在发生。你永远不可能跟上所有事情,所以别试着去跟上。当你需要的时候,让机器来干这个。
It will be a big event. It will be a new chip from Intel. How did I miss that? Experts that come to think they know it all struggle with that. I think it's always important to remind, there is stuff going on in the world you don't know about. You can't ever keep up with it all, so don't try. Plug into machines when you're needed to do it.
另一件经常出现的事情是,那些在寻找答案的人。机器告诉我该做什么?它什么也不会告诉你去做。它只是向你展示这个世界。它告诉我该做什么?
I think the other one that comes back are people that are looking for that answer. What does a machine tell me to do? It doesn't tell you to do anything. It just shows you the world. What does it tell me to do?
你得自己做那个决定。
You're going to have to make that decision yourself.
这对人们来说真的很难跨过去。他们根深蒂固地认为机器应该给我们答案。
That's really hard for people to get past. They are so ingrained that the machine should give us an answer.
我们如此根深蒂固地认为存在一个正确答案,以至于我们没能意识到,事情可能更微妙一些,也许我们仍然需要参与其中。
We're so ingrained that there's a right thing that we don't appreciate that perhaps it's a little more subtle and that maybe we need to still be involved.
问:有哪些不同的选择可用?
Question: What are the different choices available?
肖恩:你可以有一系列不同的关联方式,也有很多种选择。从明确的关联,比如 A 与 B 有过互动,到隐含的关联。隐含关联可能是 A 的概念与 B 的概念相似,也可能是 A 的情感语言与 B 的情感语言相似。
Sean: You can have a range of different connections and there’s a range of options. Anything from explicit connections like A interacted with B, through implicit connections. An implicit connection could be the concepts of A are similar to the concepts of B. That could also be the emotional language in A is similar to the emotional language in B.
你可以有一系列不同的东西。你选择投射出来的,就是你认定具有价值的那类关联。我们试图给用户提供选项,让他们选择想要这种关联还是那种关联。当你逐步深入时,你会想到一个高维度的物体,它在很多很多方面都有联系。
You can have a range of different things. What you choose to project is what you choose to deem to be the valuable connections. We try and give the user options to say, do you want this kind of connection or that type of connection. As you go down, you think of a high dimensional object that's connected in many, many ways.
你必须把它投射到二维或三维空间,这样我们才能导航。你是在压缩一个维度空间,并试图猜测该做哪种压缩。你可以做得合理,但更好的做法是把它作为选项提供给用户,这样他们就能看到对自己有意义的关联。
You've got to project that to two dimensions or three dimensions so that we can navigate it. You're collapsing a dimensional space and trying to guess which collapsing to do. You can do reasonable jobs, but it's much better to give that as an option to the user so that they can have the connections that are meaningful to them.
戴维·费鲁奇,桥水联合基金
David Ferrucci Bridgewater Associates
戴维·费鲁奇是一位获奖的人工智能研究员。他于 2012 年 12 月加入桥水联合基金的研究部门,致力于开发在市场和管理中捕获和应用知识的新方法。
David Ferrucci is an award-winning Artificial Intelligence (“AI”) researcher. He joined Bridgewater in December 2012 as part of the research department to work on new approaches to capture and apply knowledge in markets and management.
戴维在人工智能和软件系统架构方面拥有超过 25 年的经验。在 IBM 研究院,他是构建 Watson 的首席研究员。Watson 在人工智能和自然语言处理方面取得了里程碑式的成就,并获得了许多公众赞誉,包括 AAAI 费根鲍姆奖。由于戴维对科学和商业的广泛影响,他于 2011 年成为 IBM 会士。
Dave has over 25 years of experience in AI and software systems architecture. At IBM Research, he was the principal investigator who built Watson. Watson resulted in a landmark achievement in Artificial Intelligence and natural language processing and was awarded many public accolades including the AAAI Feigenbaum Prize. As a result of Dave’s broad impact on science and business, he became an IBM Fellow in 2011.
戴维在人工智能、自然语言处理、文本生成以及软件架构与工程领域发表过多篇论文。他于 1994 年从伦斯勒理工学院获得计算机科学博士学位。
Dave is published in AI, Natural Language Processing, Text Generation, and Software Architecture and Engineering. He earned a PhD in Computer Science from Rensselaer Polytechnic Institute in 1994.
戴维·费鲁奇,人工智能
David Ferrucci Artificial Intelligence
迈克尔·莫布森:我很荣幸为大家介绍今天的收尾演讲者,戴夫·费鲁奇。
Michael Mauboussin: It's my real pleasure to introduce our cleanup speaker today, Dave Ferrucci.
戴夫目前在对冲基金桥水联合基金的研究部门工作。
Dave currently works in the research department at Bridgewater Associates, the well-known hedge fund.
在我认识戴夫之前很久,我就非常欣赏他的工作以及他所做的那类工作。不一定直接与戴夫相关,但我确实密切关注了 IBM 打造一台能击败加里·卡斯帕罗夫的计算机的努力,正如大家所知,他们在 1997 年成功了。
Long before I met Dave, I was a big fan of his work and the kind of work he does. Not necessarily directly related to Dave, but I certainly followed closely IBM's efforts to build a computer that could beat Garry Kasparov, which, as you all know, they succeeded in doing in 1997.
但当我听说 IBM 正在打造一台能够在《危险边缘!》游戏中挑战冠军的机器时,那真是令人兴奋,我对这项工作更加着迷了。Watson 这台击败了冠军、能玩《危险边缘!》的计算机,其背后的男人就是戴夫·费鲁奇。
But when I heard that IBM was building a machine that would take on the champions in the game of Jeopardy!, that was exciting and I was even more hooked on the work. The man behind Watson, which is the champion-beating, Jeopardy!-playing computer, was Dave Ferrucci.
国际象棋基本上是一种简单的游戏,但它确实有巨大的可能组合结果。不过,这也是一种计算机天生就擅长玩的游戏。基本上,它只是数字运算。
Chess is basically a simple game but it does have enormous possible combinations of outcomes. But it's also a game that computers are well designed to be good at. Basically, it’s number crunching.
《危险边缘!》则是完全不同的一回事,我们待会儿会听到。这些问题需要理解明喻、双关语、谜语,语言处理能力必须非常出色。因此,开发一个程序,在人类真正擅长的领域——同时还要具备所有其他信息——击败人类,这确实是非凡的成就。
Jeopardy! is a very different ball of wax, as we'll hear. The questions require an understanding of similes, puns, riddles, language processing has to be excellent, and so developing a program to beat humans in that realm where humans are really strong, as well as having all the other information, that is truly extraordinary.
戴夫·费鲁奇在人工智能领域有超过 25 年的经验。他获得了该领域众多著名奖项。他将向我们介绍这个领域,分享他多年来学到的经验,以及未来会怎样。请大家和我一起欢迎戴夫·费鲁奇。戴夫。
Dave Ferrucci has more than 25 years in artificial intelligence. He's received numerous prestigious prizes in the field. He's going to tell us about this field and what he's learned over the years, and what's to come. So please join me in welcoming Dave Ferrucci. Dave.
[applause]
[applause]
戴维·费鲁奇:谢谢。听到肖恩的演讲真是太好了,因为他演讲的主题与我接下来要讲的非常契合。他也做了很多基础工作,所以我就能跟大家聊聊哲学,而不是深入探讨那些深奥的东西了。[笑] 两者非常吻合。
David Ferrucci: Thank you. It was great to hear Sean's talk because the theme behind that talk is going to be very compatible with the theme that you're going to see in my talk as well. He also did a lot of the heavy lifting, so I'll be able to chat with you about philosophy instead of going into all of that deep stuff. [laughs] It was very compatible.
我想一开始就说两点,大致是关于肖恩的演讲让我如何思考人工智能与增强智能的区别。此刻,你和我正在交谈,我增强你的智能,你也在增强我的智能,因为我们正在对话。
I will say a couple of things right at the outset, sort of how Sean's talk left me thinking about artificial intelligence versus augmented intelligence. Right now, you and I talking, I'm augmenting your intelligence, you're augmenting mine as we we're dialoguing.
这之所以能发生,是因为我们都有智能,我们说同一种语言,拥有兼容的流程和表征。你可以想象,当两个人合作时,他们是在互相增强对方的智能。但另一方面,要做到这一点,也需要一定水平的智能。
The reason that can happen is because we're both intelligent, we both speak the same language, we have compatible processes and representations. You can think of when two people collaborate, they're augmenting each other's intelligence. But on the other side of that fence, it requires a level of intelligence to do that.
我觉得非常有趣的是,他谈到有些人能很好地与计算机互动,而有些人则不能。我认为这是人工智能领域从业者的责任,对吧?人工智能的责任是构建让这种互动更容易、让大多数人都能坐下来,并从与机器的互动中获得价值的计算机。
I thought what was really interesting was when he talked about how some people can interact with computers really well and some people can't. I think it's the responsibility of people in AI, right, it's the responsibility of artificial intelligence to build computers that make that interaction easier, that make that interaction possible, for the majority of people to sit down and actually get value out of interacting with that machine.
不管怎样,这是一个有趣的视角,有很多有趣的东西。我会简要谈谈人工智能从理论驱动系统到数据驱动系统的演变,以及最佳结合点在哪里。你会听到肖恩思想的回响,对 IBM Watson 的反思,这个人工智能领域的里程碑及其所处位置,以及我对人工智能未来的看法和它将如何演变。
Anyway, it's an interesting perspective, lots of interesting stuff here. I'll talk a little bit about the evolution of AI from theory driven to data driven systems, and this idea of where the sweet spot is. You'll hear the echoes from Sean's thoughts there, reflections on IBM Watson, a landmark in artificial intelligence and where that fit in, and my view of the future of AI, and how that will evolve.
就从最开始说起吧。人工智能这个词是谁创造的?它从哪来?是什么含义?对此有两种有趣的观点。一种是要求非常高的,即图灵的观点:计算机系统的交互行为最终与人类无法区分。
Just to go from the very beginning, artificial intelligence, who even coined the word? Where does it come from? What does it mean? Two interesting perspectives on this. One, very demanding, which was Turing's perspective, which was computer systems whose interactive behavior is ultimately indistinguishable from
戴维·费鲁奇,人工智能(续)
David Ferrucci Artificial Intelligence (Continued)
你们当中有多少人读过关于图灵测试被解决了的报道?你们当中有多少人相信?哦,很好。
the humans. How many of you have read that the Turing test, the Turing thing was solved? How many of you believe it? Oh, good.
[laughter]
[laughter]
有人看过那个据称通过了图灵测试的机器的对话吗?真是可怜。有多少人听说过 20 世纪 60 年代的 ELIZA 程序?不错?我感觉它比那好一点,但也好不到哪去。
Has anybody seen the dialogue from the machine that supposedly passed the Turing test? It's really pathetic. How many people have heard of the ELIZA program from the 1960s? Good? It's a little bit better than that I think, not that much.
麦卡锡在 20 世纪 50 年代创造了这个词,他说,人造系统是执行那些如果由人类执行就会被认为与智能相关的任务的计算机系统。这是一个非常不同的定义。在某种程度上,它很巧妙,因为它完全避免了真正定义“智能”,而是把标准放在了任务上。
McCarthy in the 1950s coined the term actually, and he said artificial systems are computer systems that perform tasks that if performed by a human would be associated with intelligence. It's a very different kind of definition. It's sort of ingenious in a way because what it does is it avoids really defining intelligence entirely, and it really puts it on the task.
所以,我们有深蓝,我们说:“哇,一台计算机,深蓝,击败了国际象棋特级大师。如果有人能打败他,我会认为他……如果他/她能做到,那一定是很聪明的。”一个计算机程序在《危险边缘!》中击败了最强者。好,我认为这需要某种智能,所以这些可以被认为是人工智能。
So, we have Deep Blue and we say "Wow, a computer, Deep Blue, beat a grand chessmaster, and people who can do that, I would associate them...They'd be intelligent if they were able to do that." A computer program beat the best at Jeopardy! OK, I associate with that a certain kind of intelligence, so those things are considered artificially intelligent.
但当然,深蓝不会走路,Watson 也不能走下台,开始和你聊天,谈论经济学,讲笑话。这些都不会发生。所以只是它们完成的任务,我们说:“哎呀,如果人类做到了,我们会认为那个人很聪明。”这是一种非常不同的定义。
But of course, Deep Blue couldn't walk nor can Watson walk off the stage, start conversing with you, talking about economics, making jokes. Neither of those are going to happen. So it's just that task that they did, we say "Gee, if a human did that, we would think that human was intelligent." Sort of a very different kind of definition.
人工智能实际上已经取得了长足的进步。我不知道有多少人熟悉正在进行的许多渐进式进展。划分世界的一种方式是将智能分为“知道”——智能意味着知道某物,比如解释、理解、推理、学习、预测——以及“行动”——能够采取行动,并以有效的方式执行该行动,比如走路、看、感知、抓住、驾驶、飞行,你们知道的,诸如此类。
AI has come a long way actually. I don't know how many people are familiar with a lot of the incremental advances that have been going on. One way to divide the world is into the idea of knowing―intelligence as it means to know something, like to interpret, to understand, to reason, to learn, to predict―and then doing―being able to take some action, and do that action in an effective way, walk, see, sense, catch, drive, fly, you know, that kind of stuff.
我想从这里开始。有多少人知道谷歌无人驾驶汽车?每个人肯定都知道谷歌无人驾驶汽车。我不知道你们是否见过市面上的一些东西。如果没见过,我打算点开几个看看。我准备了一些 YouTube 视频给不熟悉这类东西的朋友。噢,这里有一个。
I'd like to start here. How many people know about the Google car? Everybody's got to know about the Google car. I don't know if you've seen some of the things that are out there. If you haven't, I was going to click on a couple of these things. I have a couple of YouTube videos for those of you not familiar with this kind of stuff. Oh, there's one.
视频:抓取历史包括抛和接,这与传统保持接触状态的方法形成对比。
Video: History Grasping consists of throwing and catching, which contrasts with traditional methods that maintain contact state.
戴维:这相当不错。你能做到吗?
David: That's pretty good. Can you do that?
[laughter]
[laughter]
视频:……在空中,然后移动去拦截它。这是一个极其动态的任务,充满了很多不确定性,动态速度非常快,然而这种设计控制系统的方法非常稳健,并且……
Video: ...In the air, and then move to intercept it. This is an extremely dynamic task with a lot of uncertainty, with very fast dynamics, yet this technique for designing control systems is very robust, and...
戴维:有多少人看过这个?这是“大狗”。基本上这条狗可以在冰上行走并保持平衡。它学会了如何自我平衡。实际上它相当……我做不到这个。你能做到吗?我做不到。我会直接摔倒。
David: How many people saw this? This is the big dog. Basically the dog can walk on ice and balance itself. It’s learned how to balance itself. It's actually quite...I can't do that. Can you do that? I can't do that. I would just fall.
这是某种智能。你在感知。你接收这些感知数据,并对它们做出反应。
That's a certain kind of intelligence. You're sensing. You're taking in that sense data, you're responding to
这东西你实际上可以买到。
it. This you can actually buy.
戴维:声音定位,所以它有听觉,有压力传感器,它学会在不同地面上行走。
David: Sound localization, so it's got hearing, it's got pressure sensors, it learns how to walk over different surfaces.
戴维·费鲁奇,人工智能(续)
David Ferrucci Artificial Intelligence (Continued)
戴夫:不管怎样,当然还有谷歌无人驾驶汽车,它将改变每个人的生活。太神奇了。在解释、理解、推理、学习和预测方面,这在很多方面都更具挑战性,因为它要求你能够触及“意义”,而“意义”最终与人类如何思考和解释事物相连,因为你接收的是人类数据,你要问:“这意味着什么?我如何理解它?我如何以有意义的方式推理它?”
Dave: Anyway, and then of course the Google car, which is going to change everybody's life. It's amazing. In interpreting understanding, reasoning, and learning, and predicting it's a little bit more challenging in many ways, because what it requires you is to kind of get at meaning, and meaning is ultimately connected with how other humans think and interpret things, because you're taking human data, and you're saying "What does this mean, and how do I understand it, and how do I reason about it in a meaningful way?"
有多少人看过 Watson 参加的《危险边缘!》比赛?有多少人?不多。
How many people saw the game against Watson, the Jeopardy! game? How many people? Not a lot.
你们当中哪些人看了比赛,或者没看?有多少人支持计算机赢?
Those of you who watched it, or didn't watch it, how many people were rooting for the computer?
[laughter]
[laughter]
戴维:有多少人希望人类赢?他们可以留下来吗?
David: How many were rooting for the humans to win? Are they allowed to stay?
[laughter]
[laughter]
戴维:这挺有意思的。那些没看过的人。这里有一段小片段。
David: It's kind of interesting. Those of you who didn't see it. Here's a little clip.
[视频片段开始]
[video clip begins]
亚历克斯·特里贝克:Watson?
Alex Trebec: Watson?
Watson:什么是角蛋白?
Watson: What is keratin?
亚历克斯:回答正确。
Alex: You are right.
Watson:选择“偷窃的艺术”类别,1200 美元。
Watson: The Art of the Steal for $1,200.
亚历克斯:答案,另一个“每日双倍”机会。
Alex: Answer, the other Daily Double.
[applause]
[applause]
亚历克斯:Watson,你打算下注多少?
Alex: Watson, what are you going to wager?
Watson: $1,246, please.
Watson: $1,246, please.
[laughter]
[laughter]
亚历克斯:1246 美元。这里是“偷窃的艺术”类别中的线索。“古代尼姆罗德之狮于 2003 年从这座城市国家博物馆失踪,同时还有大量其他物品。”
Alex: $1,246. Here is the clue in The Art of the Steal. The ancient Lion of Nimrod went missing from this city's national museum in 2003, along with a lot of other stuff.
Watson:我猜一下。什么是巴格达?
Watson: I'll take a guess. What is Baghdad?
亚历克斯:尽管你对自己的回答只有 32% 的把握,但你是对的。
Alex: Even though you are only 32 percent sure of your response, you are correct.
[applause]
[applause]
[clip ends]
[clip ends]
戴维:这背后很有意思。最有趣的一点是那个答案面板。在我看来……我费了很大力气,这里有个幕后故事,我就不讲了——我据理力争,才让那个答案面板最终出现在电视上。因为那个面板显示了排名前三的答案及其置信度,我认为它彻底改变了每个人对当时情况的理解。
David: A lot going on there. One of the most interesting things about that was that answer panel. In my view…I fought very hard, there's a back story, I won't tell it, I fought very hard to make sure that answer panel ended up on television. Because that answer panel that showed the top three answers with the confidence, I think changed everyone's impression about what was going on.
如果你只是把那台计算机放在那里,它回答问题时才亮起,不回答时你什么都看不到,我想人们就会完全凭想象去揣测计算机是如何工作的。他们会把自己的想法投射到计算机上。
If you would have just put that computer up there and it answered when it did and you didn't see anything when it didn't answer, I think it would have been completely left up to people's imagination about how they thought computers worked. They would have projected that onto the computer.
戴维·费鲁奇 人工智能(续)
David Ferrucci Artificial Intelligence (Continued)
你知道大多数人认为计算机内部是怎么运作的吗?是一张巨大的电子表格,一张大表。所有问题都已经在那里了。问题来了,你查一下,就找到了答案。人们原本会那样想象。但是,一旦你放上那个答案面板,你就会说,等等,它正在考虑多个答案。
And you know what the majority of people think is going on inside of a computer? It's a rich spreadsheet, a big table. All the questions are already there. The question comes up, you look it up and you look up the answer. That's how people would have imagined it. But as soon as you put that answer panel up there, you say wait a second, it's considering multiple answers.
它正在计算某一个答案正确的概率。它一定在思考。你对它内部运作的印象——不管你是否称之为思考——一旦放上那个答案面板,就彻底改变了。我认为这确实改变了人们对当时情况的理解,使之更贴近现实。
It's coming up with a probability that one might be right. It must be thinking. The impression of what's going on, whether you call it thinking or not, but it completely changes once you put that answer panel up there. I think that really changed people's perception of what was going on to actually become more aligned with reality.
我们之前讨论了很多关于国际象棋的事——有限、数学上精确定义的搜索空间,所有回应都基于精确无误的规则,庞大但有限的走法集合——有人提到了 10 的 40 次方种可能,后果清晰明确。计算机是完美的。真正令人惊讶的是人类竟然也能做到。
We talked about chess a lot―finite, mathematically well-defined search space, all responses grounded in precise unambiguous rules, a large but finite set of moves―somebody mentioned 10 to the 40th, clear consequences. It's perfect for a computer. The really surprising thing is that humans can do it.
我们之前了解到,人类实际上是以某种不同的方式做到这一点的,不是仅仅浏览所有的走法可能性,而是记住各种不同的棋型和类似的东西。人类语言则完全是另一回事。词汇可以是图像或语音,我说的语言指的是我们相互交流的方式。
We learned earlier about how they actually do this somehow differently than just looking over the broad possibilities of moves, but they actually remember these various kinds of patterns and things like that. Human language is a completely different story. Words, they can be images or speech, what I mean by language is just how we communicate with one another.
我们用词汇交流。我们用图像交流。它们缺乏精确的解释。几乎有无穷无尽的表达映射到各种各样的含义上。含义仅仅植根于人类共同的经验。我们倾向于认为……当我们开始研究计算语言学时,我们会说,“哦,有语法规则,有词汇,词汇有精确的含义。”
We can communicate with words. We communicate with images. They lack precise interpretation. There are nearly an infinite number of expressions mapped to a huge variety of meaning. The meaning is only grounded in shared human experience. We tend to think that…when we started with computational linguistics we'd say, "Oh, there are grammars and there's words and words have precise meaning."
但现实是,你可以认为词汇本身没有内在含义。你可以把它们更多地看作是一种指向我们共同经验的复杂索引系统。我用一堆词对你们说话,它们点亮了你们大脑的不同部分。它们在唤起你们不同的经历。而你们正在梳理这些经历。
The reality is you can think of words as having really no intrinsic meaning. You can think of them more as this intricate indexing system into our shared experience. I use a bunch of words and I'm speaking to you and they're lighting up different parts of your brain. They're reminding you of different experiences. And you're sorting out those experiences.
当我正在对你们说话时,你们在点头。我想,“哦,好的。这些词似乎点亮了相同的经历。那个人似乎听懂了。” 但据我所知,你根本不知道我在说什么。事实上,你可能有一些概念,但你没有和我完全相同的解读,这甚至可能更糟,因为现在你在点头并说,“我懂了。” 然后到了紧要关头,我期望你去做某件事。我预测你会以某种方式行事,因为我看到你理解了我的话。但你却做了件疯狂的事。结果发现你根本没理解我。尽管我用了那些词,但它们在你头脑中点亮了不同的经历。
And as I'm speaking to you, you're nodding. And I'm thinking, "Oh, OK. Those words seem to be lighting up the same experiences. That person seems to be understanding.” For all I know, you have no idea what I'm talking about. In fact, you might have some idea, but you don't have exactly the same interpretation I have, which might even be worse, because now you're nodding and you're saying, "I understand." And then push comes to shove and I expect you to do something. I predict that you're going to behave a certain way because I saw you understanding my words. And you do something crazy. It turns out you don't understand me at all. Even though I used the words, but they were lighting up different experiences in your head.
这是我生活中的一个很好的例子。我在家做科学实验。我有两个小女儿。她们会听到我说,“嘿,姑娘们,快下来。我在做这个,真的非常非常有趣。” 她们会下来,我正在水槽里做某种疯狂的事,或者别的什么。
This is a great example this from my own life. I do these scientific experiments at home. I have two young daughters. They'll hear me say, "Hey guys, come down here. I'm doing this thing. It's really, really interesting." They would come down, and I'm doing something crazy in the sink or whatever.
我一而再再而三地这样做,“姑娘们,你们得下来。这个真的很酷。这个真的非常非常有趣。” 有一次,她当时七岁,我说,“快下来。” 她停在楼梯顶端说,“爸爸,有趣的事很无聊。”
I'm doing this over and over again, "Guys, you've got to come down here. This one's really cool. This one's really, really interesting." One time she was seven at the time and I go, "Come on down here." She stopped at the top of the stairs and she said, "Daddy, interesting things are boring."
[laughter]
[laughter]
戴维:因为她对这个词的经验,她把“无聊”这个词赋给了它——她根据自己对那个词的经验,赋予了它完全相反的意思。语言就是如此演变的。它通过语境以及将共同经验赋予那些词的过程来演变。
David: Because of her experience with that word, she assigned “boring,” she assigned a completely opposite meaning to that word based on her experience with that word. That's really how language evolves. It evolves through context and through the assignment of that shared experience to those words.
戴维·费鲁奇 人工智能(续)
David Ferrucci Artificial Intelligence (Continued)
这正是在追求那种含义,而那种含义真正的源头是人类,以及人类如何体验世界。
It's getting at that meaning and that meaning is really sourced in humans and how humans experience the world.
这里,我们有这张图片。无论你看到了什么,先别说。我可以用很多种不同的方式来解读它。计算机存储了那张图像,我们可以移动它;用相机拍下来,然后通过短信或邮件发给你。
Here, we have this picture. Whatever you see there, don't even tell me. I can interpret this in lots of different ways. A computer stores that image and we can move it around; take it with a camera and then send it to you in a message or an email.
但它只是一堆 0 和 1。它到底意味着什么?当我开始观察一个人时,因为他的视觉系统能识别颜色并说,“有白色,有黑色,还有红色。” 现在,我可以进行物体检测。我可以说,“我知道那是什么。那是一个计算器,上面放着几个骰子。”
But it's just a bunch of zeros and ones. What does it really mean? When I start looking at a human, because it's optical system can identify colors and say, "There's white, there's black and then red." Now, I can do object detection. I can say, "I know what that thing is. It's a calculator with some dice on top."
我可以思考,“你知道他们想告诉我什么吗?他们在计算概率。这张图像实际上由人类标注了这个词:胜利。” 我甚至都不会建立那种联系,但一群人就这么做了。
I can think, "You know what they meant to tell me? They're calculating the odds. This image was actually labeled by humans with this word here, winning." I wouldn't even have made that connection, but a bunch of humans did.
另一种方式是人类的工具使用。也许那就是它的含义。不同的解读层次,只是一堆 0。这些解读从何而来?它来自我们与那个事物的经验。
Another way is human use of tools. Maybe that's what it means. Different levels of interpretation, just a bunch of zeros. Where do those interpretations come from? It comes from our experience with that thing.
那个事物,那个人造物成了我们语言的一部分。我们用它在你的大脑中点亮那些不同的经历。
That thing, that artifact becomes part of our language. We use it to light up those different experiences in your head.
最终,含义是主观的,而我们人类就是主体。我们如何解决含义问题?正是语境在根据那些词来三角定位含义时变得如此重要。你可以把含义看作是从符号到共同经验的概率性映射。更丰富的语境缩小了这些可能性,并提高了这种映射的置信度。你在说什么?
Ultimately, meaning is subjective and we the humans are the subjects. How do we resolve meaning? It's the context that becomes so important in triangulating that meaning given those words. You can think of meaning as this probabilistic of mapping from symbols to a common experience. Richer context narrows those possibilities and improves confidence in that mapping. What are you talking about?
“蝙蝠朝他飞来。” 有几种可能。那些词开始点亮一些东西。有很多可能性。我们只展示了两种相互竞争的可能性——一群吸血蝙蝠或果蝠之类的,朝一个人飞去。我们可以想象在另一种情况下,一根棒球棒向我飞来;一个是我的前女友,另一个是有人打了一个本垒打,球棒飞了出去。我当时在观众席上。不同的可能性。
“The bat was flying toward him.” A couple of ideas. Those words start lighting up things. There's lots of possibilities. We're just showing two competing possibilities – a bunch of vampire bats or fruit bats or whatever, flying toward a person. We could imagine a baseball bat flying toward me under different conditions; one is my ex-girlfriend, another one is maybe someone hit a home run through the bat. I was in the audience. Different possibilities.
“比利拼命地跑。” 现在我有两种相互竞争的可能性。一种是比利跑着躲避蝙蝠的追赶。另一种是比利刚挥棒,球棒飞了出去。
“Billy ran as fast as he could.” Now I have two competing possibilities here. One is running away from the bats chasing Billy. The other one is Billy just swung and the bat went flying.
“他安全到家了。” 还是不确定。随着语境的增多,我仍然有两种相互竞争的可能性。
“He made it home safe.” Still not really sure. I still got two competing possibilities with more and more context.
最后,“他得分了。” 我现在很确定是棒球比赛了。还有一种解读我就不提了,但那时我非常确定是棒球比赛。
Finally, “He scored.” I'm pretty sure now it's the baseball game. There's another interpretation I won't mention, but I'm pretty sure it's the baseball game at that point.
随着越来越多的信息进来,我接收那些词,并开始将它们与那种情境联系起来。
As more and more information comes in, I take those words and I start connecting them to that meeting.
不过非常有趣的是,表达方式上的细微差异可以戏剧性地改变它所点亮的含义。
Really interesting though because subtle differences in the expression can dramatically change the meaning that it happens to light up.
这些微小的变化,取决于这些词是如何被使用的——不仅是词本身,还有它们出现的顺序。这里,我仅仅拿了这些词;“safe home”和“at”,稍作打乱。我把它们输入到谷歌图片搜索里。结果如下。别忘了,谷歌搜索很大程度上是由人们如何标注事物驱动的。这正是人们说“这对我来说是这个意思。不,这对我来说是这个意思。”的来源。将含义与词汇联系起来。
These small changes, it depends on how those words are used, not just the words themselves, but what order and sequence they come in. Here, I just took these words; "safe home" and "at," and I just mixed them up a little bit. I put them into Google image search. Here's what came out. Don't forget, Google search is very much driven by how people annotate things. It's the source of people saying, "This is what this means to me. No, this is what it means to me." Connecting the meaning with the words.
“Safe at home”,我得到这些图片,大部分是棒球,除了这张是一个乐队之类的,我不认识。“Home safe。”“Safe home。” 完全不同的解读。看到这一点真是
"Safe at home," I get these images, mostly baseball, except for this one right here of some band or something. I don't know what that is. “Home safe.” “Safe home.” Really different interpretations. It's
戴维·费鲁奇 人工智能(续)
David Ferrucci Artificial Intelligence (Continued)
令人惊叹。这完全取决于词汇如何使用,不仅是词本身,甚至包括词的顺序。
amazing how it makes that point. It's all about how the words are used, not just the words, but even the ordering of the words.
当你回顾人工智能的早期起源时,就是“我想让计算机理解事物、回答问题并做出预测。我该如何着手?” 想法是构建理论,来模拟世界如何运作。要说,“如果 A(x) 为真,那么 B(x) 为真。如果条件改变,我有 A(x) 和 C(x)…… 并且模拟所有这些规则。” 识别概念及它们之间的关系,通常是在少量数据上工作。
When you look at AI from its very early beginnings, it's, "I want the computer to understand things, answer questions, and make predictions. How am I going to go about that?" The idea was to build theories, to model how the world works. To say, "If A(x) is true, then B(x) is true. If the conditions change and I have A(x) and C(x)…and to model all those rules. To identify concepts and relationships between them, usually working on small data.
同样有趣的是,肖恩谈到了这些挑战之间的区别,即面对小数据时需要更多的智能。需要有更多解读和理解的方式,而不仅仅是拥有大数据。
It was also interesting, Sean talked about the difference between the challenges, that you need more intelligence with small data. More ways to interpret and understand than if you just have big data.
不管怎样,科学家们当时在小数据上工作,是因为那时他们没有大数据。他们进行解读和泛化,提出一套本体论,区分不同的概念——通用概念、更具体的概念——将它们与规则联系起来,最终形成理论。然后提出一个问题,他们就能得到预测和答案。
Anyway, scientists worked with small data because they didn't have big data in those days. They interpreted and generalized, came up with an anthology, different concepts, general concepts, more specific concepts, connected them to rules, came up with the theory. Then would ask a question, and they would get these predictions and answers.
这些是可以解释的。为什么是可解释的?换句话说,它是一个“开箱”。意思是你可以说,“你为什么得出那个结论?” 计算机可以回溯它所有的推理过程,识别概念和规则,并给你一个演绎证明。
They were explicable. Why were they explicable? In other words, it was an open box. In the sense that you could say, "Why'd you come up with that?" The computer can go back and trace all its reasoning, identify the concepts and the rules, and give you a deductive proof.
我刚从纳特·西尔弗的书里翻出这段话,他在描述为什么有人认为会出现衰退:“消费者为购买住房泡沫下已变得不可负担的房产而过度借贷。其中许多人停止了还款,而杠杆体系的程度会加剧问题,如此种种。你在展示,‘这是我对世界如何运转的理解,以及我为何可能预测到会有一场衰退。’”
I just pulled this out of Nate Silver's book when he was talking about the description of why somebody thought there was a recession: “Consumers have extended too much credit to pay for homes that the housing bubble had made unaffordable. Many of them had stopped making their payments to the degree of leveraging the system would compound the problem and so forth and so on. You're demonstrating, "Here's my understanding of how the world works and why I might predict that there's going to be a recession."
以数据为驱动,过去 15 年间情况开始发生变化,海量数据变得越来越多、越来越容易获取,再加上极为庞大的计算能力来处理这些数据。各种不同的统计机器学习技术层出不穷。大数据时代来临了。A 与 B 存在相关性,这可能暗示某种因果关系,所以我觉得没问题。给我看 A,我就能预测 B。
Data driven, since things started changing over the last 15 years, huge amounts of data becoming more and more available, tremendous amount of compute power to munge over that data. Lots of different statistical machine learning techniques. We had the advent of big data. A is correlated with B and that may suggest some causation, so I'm good. Show me A and I'll predict B.
在医疗健康、经济、电子商务、奈飞(Netflix)、亚马逊、经济、选举等多个领域,确实出现了非常有趣的结果。我们今天大概把所有这些都聊到了。你得到的反馈是预测结果,但这些预测有点难以解释。背后并没有任何与人类思维兼容的理论作为支撑。
Really interesting results in a broad number of areas: health care, economics, e-commerce, Netflix, Amazon, economics, elections. We talked about I think all of them today. What you get back is you get back predictions, but there sort of inexplicable. There's no human compatible theory underlying there.
我可以告诉你所有我可能考虑过的不同特征。可能有一大堆,但它们之间有什么关联,又意味着什么?对于预测,我得到的是截然不同的一种预测或解释。目前最可靠的前瞻性指标集体表现出的状态,就像全面衰退前夕那样。对类似的预测,给出的解释却非常不同。
I can tell you all different features I may have considered. There may be a ton of them, but how are they related and what do they mean? I get a very different kind of prediction or explanation for prediction. The most reliable forward looking indicators are now collectively behaving as if they did on the cusp of a full blown recession. Very different kind of explanation for a similar prediction.
我认为这一点之前也被提过,我觉得这是个很有意思的讨论。实际上,这场争论已经持续了很多年,并不是什么新鲜事。人们一直在谈论归纳法和演绎法,只是现在它才进入了主流视野。
I think this was mentioned, too, I think it's an interesting debate. It's actually been going on for years. It's not that new. People have been talking about induction and deduction forever. It’s just kind of made it into the mainstream.
2008 年,克里斯·安德森在担任《连线》杂志主编时表达了一种极端观点:“相关性胜过因果关系,即便没有连贯的模型、统一的理论,甚至没有任何机制性的解释,科学也能进步。”这是一种颇为极端的观点,相当引人深思。
One extreme represented in 2008, Chris Anderson when he was editor-in-chief at Wired, "Correlation supersedes causation, and science can advance even without coherent models, unified theories, or really any mechanistic explanation whatsoever." It's a fascinating point of view, very extreme.
2012 年,纳特·西尔弗(Nate Silver)说过,“把理论弃之不顾,绝对是错误的做法。
We had, 2012, Nate Silver saying, "Throwing the theory out is just categorically the wrong attitude.
当有理论支撑时,统计推断会更有说服力。而我们在几个案例中也看到了这一点。
Statistical inferences are much stronger when backed up by theory." And we saw that theme with a couple
戴维·费鲁奇 人工智能(续)
David Ferrucci Artificial Intelligence (Continued)
今天发言者当中,信息确实蕴含在数据里。但当数据越来越大,从中获取信息的难度也越来越大。
of the speakers today. Yes there's information in data. It gets harder and harder to grab it when the data gets bigger and bigger.
这是一个值得思考的有趣关系:当数据呈指数级增长时,我的计算方法与算法是否也增长得足够快,足以从中提取出有意义的数据?我需要多少理论来判断这些相关性是否有意义?这既是一个巨大的机遇,也是一个巨大的挑战。
It's an interesting relationship to think about, as the data is growing exponentially, are my computational methods and my algorithms growing fast enough to extract the meaningful data out of there? How much theory do I need to identify whether or not the correlations are meaningful? This becomes a huge opportunity, but also a huge challenge.
接下来几分钟,我想带你们过一遍一个简单到让人痛苦的例子——我已经向所有对此深有研究的人道过歉了。我就用这个简单到让人痛苦的例子,来精确定位一个要点:数据驱动方法与理论驱动方法之间的区别。我要做一个简单到让人痛苦的决定:走去吃午饭时,我该不该穿雨鞋?
What I want to do for the next few minutes is take you through a painfully simple, and I already apologized for all of you who understand this in great depth. I'm just going to take you through this painfully simple example, to fine tune the point, the difference between a data driven approach and a theory driven approach. I have a painfully simple decision to make: should I wear galoshes on my walk to lunch?
因为一直在下雨,所以情况有点复杂。
Because it's been a bit raining, so it's complicated.
首先我来运用我的理论驱动方法,我要为世界建模。这是准通用逻辑,并不完美,因为我记不全所有的通用逻辑。存在一些表面。存在一些被称为路径的东西。
First I'm going to start my theory driven approach. I'm going to model the world. These are quasi-versatile logic, it's not perfect because I don't remember all my versatile logic. There exist surfaces. There exist things called paths.
所有路径都是表面。表面可以被遮盖,也可以处于干爽或潮湿的状态。世间存在各种事件。下雨是事件的一种。如果一个表面正在下雨且未被遮盖,那它就是湿的。世界上有一种叫做人的事物。人会穿橡胶套鞋。当需要行走且路面潮湿时,人就会想穿橡胶套鞋。
All paths are surfaces. Surfaces can be covered. Surfaces can be wet or dry. There are events. Raining is a type of an event. Something is wet if it's raining and not covered. There are things called people. People wear galoshes. When it's walking and when it's wet, people want to wear galoshes.
我已经对世界建模了。这就是模型。然后,我构建一个系统,对模型进行推理。我问系统:“我要走去吃午饭,该穿雨靴吗?”
I've modeled the world. Here it is. Now, I build the system that does deduction over that. I ask the system, "I will be walking to lunch, should I wear galoshes?"
系统问:“在下雨吗?”它回答:“我正在使用规则九和七。”它说:“是的。”
The system asks, "Is it raining?" It says, "I'm using rules nine and seven. It says, "Yes."
沃伦·巴菲特:“这段路有顶棚吗?”再套用那些规则,用户回答:“没有。”系统说:“你应该穿上雨鞋。”用户问:“为什么?”因为如果正在下雨,而这段路没有顶棚,并且路面是湿的,如果路面是湿的,人们就会穿雨鞋。我假设你是一个人。
"Is the path covered?" Using again those rules, user, "No." System, "You should wear galoshes." User, "Why?" Because if it's raining, and the path is not covered, and the path is wet, if the path is wet, people wear galoshes. I'm assuming you're a person.
[laughter]
[laughter]
太棒了。太精彩了。我爱这本书。它讲清楚了你应该做什么,而且是基于理论之上的描述。
It's fabulous. It's fantastic. I love it. It gives you this theory based description of what you should do.
我想再推进一点。如果没下雨,但路上还是湿的呢?现实生活总是比你构建理论时想象的更复杂。但有一套理论,能让我们调动人类的思考、直觉、感知——所有这些人类天生具备、却很难让机器学会的东西——不过,这也会变得很有挑战。
I want to advance it. What if it's not raining, but the path is still wet? Life gets always more complicated than you'd like to think it is as you build out these theories. But having a theory allows us to engage human thought, intuition, perception, all the things humans are aware of that is very difficult to make the machine aware of, but it can get challenging.
如果我想做这件事,就得开始思考:“降雨事件持续多久?什么时候开始?什么时候结束?需要多久才能干?”我得操心地面的保水能力、温度和湿度。地表可能有的凹陷处会积水,那些叫水坑。
If I want to do this, I have to start thinking, "What's the duration of the raining event? When did it start? When did it end? How long does it take to dry?" I have to worry about ground retention, temperature, and humidity. There might be depressions in the surface. They fill up. They're called puddles.
我必须根据地形的拓扑结构来思考水分的保持情况。这些影响需要时间才能消散,诸如此类。最终,我面对的是一个极其复杂的问题,试图用理论驱动的方式来建模。我原以为能给你一个非常扎实的解释,但要做到这一点,并持续维护、发展它,变得异常困难。
I have to think about the topology and the retention of water depending on the topology. It takes time for these effects to dissipate and so forth. I end up with an extremely complex problem to try to model in a theory-driven way. I imagine giving you a really robust explanation, but this becomes exceedingly difficult to do and to maintain, and to grow.
这就像数据驱动的方法。我有我的雨靴数据,我有所有这些观察记录。我知道是否在下雨,还有人标注并说明在那些情境下他们是否会穿雨靴。
It's like the data driven approach. I have my galoshes data. I have all these observations. I have whether it's raining or not. I have people annotating and say whether or not they would be wearing galoshes under those circumstances.
“我该穿雨靴吗?”
"Should I wear galoshes?"
戴维·费鲁奇人工智能(续)
David Ferrucci Artificial Intelligence (Continued)
"Yes."
"Yes."
"Why?"
"Why?"
“85% 的时间里,只要下雨,人们就会穿雨鞋。”
"85 percent time it rains, people wear galoshes."
"Great." "Why?"
"Great." "Why?"
我还可以给这个系统加更多特征。比如这样说:“好吧,它的预测能力没那么出色,所以我打算加更多特征。我可以加入湿滑路面、有遮盖的路面,这样说不定能得到更好的预测结果。”如果我想升级系统呢?效果还是不够好,那就出个 1.1 版本,再加些特征。我可以加入鞋型、地点、树种、季节、气温、湿度、开始时间、红袜队赢了没有、路面是否有遮盖——各种各样可能相关的因素。
I can add features to this. I can say, "Well, its prediction qualities are not that great, so I want to throw in more features. I could throw in wet path, I could throw in covered, I could get maybe better predictive results." What if I want to advance my system? It's still not as good, so I want my version 1.1, so I throw in more features. I can just, shoe type, location, tree type, season, temperature, humidity, start time, whether the Red Sox won or not, whether the paths are covered, all kinds of possible things.
我可以把这个问题扔进同一个机器学习算法里。它大概会说:“85% 的情况下,那口井你可能最好还是穿上雨靴。”“为什么?”“只要看看这张巨大的表格,你自己算算就行了。”
I could throw this into the same machine learning algorithm. It should say, "85 percent of the time, that well you should probably just wear galoshes." "Why?" "Just look at this giant table and you do the math."
[laughter]
[laughter]
这些关系、意义以及我是如何思考这些不同特征的,它们如何构建出一个世界运行模型——这些恰恰是这种特定方法所缺失的,即便它可能具有预测性。如果我必须现在为自己的决策作出解释或承担责任,我能拿出什么来作为依据?对于那个特定情况下的预测,我又该如何判断自己是否相信它?我给你讲一个我自己生活中的绝佳例子。
The relationships and the meaning and how I think about these various features, and how they build out a model of how the world works is just missing from this particular approach even though it may be predictive. If I have to now explain myself or be liable for my decisions, what do I have to point to? How do I now engage in whether or not I believe that prediction in that particular case? I'll give you a great example from my own life.
稍微正经点说,但也没事。我已经看开了。很多年前,我父亲在一家餐厅突发心脏骤停。救护车花了 10 分钟才赶到,火速把他送进医院。他终于到了医院。
A little bit serious but it's OK. I'm over it. My dad many years ago went into cardiac arrest in a restaurant. An ambulance took 10 minutes to get there, rushed him to the hospital. Finally he gets there.
医生出来说:“听着,你爸爸已经脑死亡了。我们需要你签一份‘放弃急救同意书’。”我说:“你怎么知道他脑死亡了?”医生说:“鉴于他的情况、救护车到达的时间,从统计上讲,有 98% 的几率他已经脑死亡了。”我说:“好吧。”
The resident comes out and says, "Look, your dad's brain-dead. We want you to sign a "Do-Not-Resuscitate." I said, "How do you know he's brain-dead? He said, "Given where he was, how long the ambulance took, statistically speaking, there's a 98 percent chance he's brain-dead." I said, "OK.
餐馆里有一百个人。其中只有两个人能撑过这一关。你凭什么说他不是那两人之一?” “好吧,你得去找心内科主任谈谈。”
There were a hundred of people in the restaurant. Two of them would have survived this. How do you know he's not one of those two?" "OK, you need to speak to the chief cardiologist."
[laughter]
[laughter]
于是,我打电话给首席心脏病专家。他说:“我知道这非常艰难,但我们了解情况。我们有统计数据。你父亲脑死亡。负责任的做法就是签署一份‘不实施心肺复苏术’协议。”我说:“你们有任何推定性证据吗?”
So, I got on the phone with the chief cardiologist. He said, "I know this is very difficult, but we understand these things. We have the statistics. Your dad's brain-dead. The responsible thing to do is to sign a "Do-Not-Resuscitate." I said, "Do you have any deductive evidence?"
他说:“你说什么?”
He said, "What?"
我问:“你怎么知道他脑死亡了。你做脑电图了吗?”
I said, "How do you know he's brain-dead. Do you have an EEG?"
不,我们没有脑电图。
"No we don't have and EEG."
“那么,你怎么知道呢?”
"Well, how do you know?"
“他的瞳孔放大了。”
"His pupils are dilated."
我说:“做一下推理,他瞳孔放大的原因还有其他可能吗?”
I said, "Doing the deduction, are there other reasons why his pupils might be dilated?"
“嗯,有的,救护车上会给一种药,可能会让他的瞳孔放大。”
"Well, yes, there's a drug they give at the ambulance that might dilate his pupils."
[laughter]
[laughter]
“你还有什么其他的?”
"What else you got?"
戴维·费鲁奇人工智能(续)
David Ferrucci Artificial Intelligence (Continued)
“我必须过来跟你谈谈这件事。”他来了之后,我们又争辩了一会儿。最终他放弃了,对我非常失望。长话短说,大约 18 小时后,我父亲坐在床上,没有任何脑损伤。我没有签 DNR(放弃心肺复苏同意书)。
"I have to come down there and talk to you about this." He came down and we debated this a little bit further. Eventually, he gave up. He was very frustrated with me. Long story short, about 18 hours later, my father was sitting up in bed with zero brain-damage. I would not sign the DNR.
当时的情况是,虽然统计数据总体适用,但在那个特定案例中,有非常具体的证据会引导你朝某个方向走。你必须仔细梳理,而我也确实想通过推理一步步来。这个过程非常漫长,因为那位医生反对我的思路,所以逼着我做出每一个决定。他问:“你现在打算怎么办?”
What went on there was while the statistics apply in general, in that particular case, there's very specific evidence that would take you one direction or another. You have to go through and I wanted to go through the deduction. It was a very long process because what the doctor made me do was, because he was against the way I was thinking, he made me make every decision. He said, "What are going to do now?"
我一个劲儿地跟他来回提问,然后不得不做出决定。有意思的是,我发现自己只想跟着统计证据走。
I would ask him questions back and forth, then I would have to make a decision. It was interesting how I just want to go with the statistical evidence.
最终极的圣杯,当我们谈论人工智能时,是建立能够获取、理解、预测和解释的系统。值得注意的一点是,人类在整个过程中参与度有多高。
The Holy Grail ultimately is to build systems, when we talk about artificial intelligence, is to acquire, understand, predict, and explain. One thing to observe is how engaged the human is in all of this process.
要回答某类问题,首先需要获取哪些重要数据?我需要构建什么样的模型来体现理解?需要用哪些算法来做预测?什么才是值得预测的?
What data is important to even acquire to answer certain kinds of questions? What model do I have to build to represent an understanding? What algorithms do I need to predict? What is worth predicting?
最终,如果我要做一个预测,我能否根据理解来解释它?我希望把整个过程自动化,而你可以把人类视为意义之源——由人类提出目标函数,定义什么是兼容的理解、什么是有意义的解释——人类参与了这个完整的过程。
Ultimately, if I make a prediction, can I explain it with regard to the understanding? I’d like to automate this whole thing, and you could think about the human in the source of meaning, coming up with the objective function, coming up with what it means to have a compatible understanding and a meaningful explanation, is engaged in this whole process.
一方面,我认为要构建人工智能系统,我们绝对必须在流程的每一个环节都让人类参与其中。关键在于,我们以何种方式让他们参与?系统能否像孩子一样,通过反复与环境及其他人类互动,逐渐理解事物的含义,并开始模仿、延伸这些含义,从而从这种参与中逐步学习?
On one hand, I think that to build artificially intelligent systems, we absolutely have to engage the human in every step of the process. The idea is, in what way do we engage them? Can the system incrementally learn from that engagement just like a child interacts over and over again with its environment and with other humans and starts to learn what the meaning of things are and starts to mimic and extend that meaning through those interactions?
世界在变。我们好像走了一个轮回,从理论驱动转向了强力数据驱动——这当然有效,令人兴奋——但你也看到这次会议上的观察,还有另一篇 2014 年《连线》杂志的文章也在讲,你需要的远不止数据:
The world is changing. We're coming like full circle, we went from theory-driven to power data-driven, which works, of course, exciting, but you're seeing the observations made at this conference as well as you see this other now, 2014 Wired, talking about how you need more than just data:
“若有人声称计算机终有一天能整理我们的数据,或让我们完全理解流感、健身、社交关系乃至其他任何事物,那么他们从根本上贬低了数据与理解的意义。”
"By claiming that computers will ever organize our data or provide us with a full understanding of the flu or fitness or social connections or anything else, for that matter, they radically reduce what the data and understanding means.”
而这里讲到的是必须让员工参与进来,并建立那种理解模式。
And it talks here about having to engage people and build the model of that understanding."
我认为人们越来越清楚地认识到:海量数据确实存在,机会也确实巨大,但仅仅找到这些相关性,却无法将它们对应到理论和因果链条上,这到底意味着什么?
I think there's a growing and growing awareness that yeah, huge amounts of data out there, huge opportunities, but what does it mean to just find these correlations and not be able to map them to theories and causation?
回顾这个领域里什么算作困难,是件有意思的事。当你想到那些非常底层的计算机应用,比如传统信息技术时,这一层级的困难在于一切都结构清晰、确定无疑。所有东西都在数据库里,我确切知道每一列的含义。
It's interesting to reflect on what represents difficulty in this space. When you think of very, very low level computer applications like traditional information technology, this one level of difficulty was everything was structured and deterministic. Everything’s in databases, I know exactly what all the columns mean.
基本上,我在做的事情就是从数据库里选取信息,可能再算点数学,然后把它发送到某个地方。
Basically, what I'm doing is I'm selecting information out of the database, maybe I'm doing some math on it, and I'm shipping it somewhere.
当你考虑预测性分析时,挑战会稍微大一些,因为你现在关注的是寻找某种模式……你并不知道那些模式到底是什么,所以你是在试图发现模式。而在你刚才说的情况下,数据是完全被理解的,几乎只是从中做选择。
It gets a little bit more challenging when you think of predictive analytics, because now what you're looking at, is you're looking for patterns that are...You don't know what the patterns are, so now you're trying to discover the patterns. Here, the data is completely understood and it's almost selecting from it. You're
戴维·费鲁奇 人工智能(续)
David Ferrucci Artificial Intelligence (Continued)
实际上你是在观察或发现模式,但并不知道这些模式意味着什么。你在建立相关性,却又在思考因果。
actually looking or discovering patterns, but you don't know what the patterns mean. You're forming correlations, but you're wondering about causation.
当问题变成非结构化和概率性的,意味着“我甚至不确定该如何解读这些数据”,难度就更大了。这就好比语言——正如我们之前所见,语言具有极高的概率性。
Then it gets even harder when it's unstructured and probabilistic, meaning, "I'm not even sure how to interpret the data." That's like with language. Language is highly probabilistic as we saw earlier.
这些方法可以用来分级衡量让计算机完成这类事情的难度。沃森在这个领域中处于什么位置?沃森是光谱中一个有趣的点。我们当时为推动人工智能发展设定了一个宏大挑战,具体来说,是在自然语言理解领域,挑战大致是这样的。
These are ways to stratify difficulty in getting the computer to do these kinds of things. Where does Watson fit into this space? Watson's an interesting point in the spectrum. Here we had this grand challenge for advancing AI, but specifically, in the areas of natural language understanding and the challenge went something like this.
我的提问范围非常开放,也就是说,你不会提前告诉我问题是什么。
I have this broad open domain, meaning you're not going to tell me ahead of time what the questions are.
没有一个电子表格,里面列着我要问你的所有问题、你必须得去查答案。我不知道问题会是什么。我不知道它们是关于什么的。
There isn't a spreadsheet where all the questions I'm going to ask you are in there, and you have to look up the answer. I don't know what they're going to be. I don't know what they're going to be about.
复杂的措辞,有各种各样的写法。“如果你站着,这就是你应该查看护墙板的方向。”有人知道吗?这题很简单。200 美元。向下,答案就是向下。
Complex language, it's written all sorts of different ways. “If you're standing, it's the direction you should look to check out the wainscoting.” Anyone? It's an easy one. $200. Down, down is the answer, down.
这很有意思,因为如果你是个理论驱动型的人,坐在那儿就会想:“好,我懂了。”别忘了,你再也碰不上这个同样的问题了。你在构建沃森系统时,投入了多少精力去理解那个问题?你永远不会再遇到一模一样的问题。他就会想:“好吧。”那我怎么把它一般化呢?
That's interesting, because if you sat there as a theory-driven guy, you're like, "OK, I got it." Don't forget. You'll never see this question again. How much you do invest in understanding that question in building Watson? You're never going to see this question again. He's like, "OK." How do I generalize?
这个问题我不会再问了,但可能会有人问方向方面的问题。让我示范一下方向是什么。
I won't get this question again, but I might get a question about direction. Let me model what a direction
是的。让我想想。我理解了上下,也理解了左右。也许我应该像罗盘一样做 360 度全方位思考,我能做到。我身后有相对的方向,身前也有。我搭建了这个庞大的复杂模型,但很可能再也碰不到关于它的问题了。
is. Let's see. I got up-down. I got left-right. Maybe I should do like 360 degrees around a compass, I could do that. There's relative direction behind me, in front of me. I build this big complex model, probably never see a question about it again.
哦,来了个问题。它问的是织物(fabric)的方向。答案是绿色。你们会想到去为这个建模吗——看那个问题的话?很可能你们猜不到。很难想象,如果领域太宽泛的话。我们看看。“在细胞分裂中,有丝分裂分裂细胞核,胞质分裂分裂这层缓冲细胞核的液体。”你们是什么人,搞金融的吗?
Oh, a question shows up. It refers to the direction of fabric. The answer is green. Would you have thought to model that, having looked at that question? Probably, you would've missed it. Very challenging to imagine, if the domain is too broad. Let's see. “In cell division, mitosis splits the nucleus and cytokinesis splits this liquid cushioning the nucleus.” What are you, finance guys?
[laughter]
[laughter]
细胞质是答案。600 美元那条,你漏掉了。一旦有人答对,你就要付钱,对吧?“看来这个凶手是《圣经》里第一个杀人犯,更绝的是,他干掉了自己的亲兄弟。”我没问题了。好了,800 美元那条。
Cytoplasm is the answer. 600, you've missed that. You're paying if anybody gets it, right? “Seems this perp was the first murderer in the Bible and to top it all off, he iced his own brother.” I'm good. There we go, 800.
为什么这是个难题?有意思的是,如果让计算机来做,它得……稍微给你透露一下沃森的运作方式——它得开始寻找单词之间的关联,从而推测这些词共同指向的常见含义。看来这个“凶手”——什么叫“凶手”——是《圣经》里的第一个杀人犯。更要命的是,他干掉了自己的亲兄弟。
Why is this a hard question? It's interesting, because if a computer was doing this, it has to...Hinting to you how Watson works, but it has to start finding connections between words to come up with, what's the common expected meaning of these words. Seems this perp―what’s a perp―was the first murderer in the Bible. To top it off, he iced his own brother.
我翻了翻词典,韦氏词典那类的,想看看“iced”这个词能跟什么联系起来,但里面没提到谋杀。那我该去哪里找呢?结果发现,要是去查“都市词典”的话——“iced”在那个词典里的第一个解释就是“谋杀”。问题是,“都市词典”的内容差不多有一半是色情内容,在面向全国观众的电视节目里用这玩意儿可不太合适。
I go to a dictionary, Webster's or something, to find what "iced" might connect to, and I don't see murder there. Where do I have to go? It turns out if I went to the Urban Dictionary―it’s the first definition in the Urban Dictionary―“iced” is murder. The problem is, the Urban Dictionary’s like 50 percent pornographic, bad source to use on national television.
[laughter]
[laughter]
“我跟你说,今天可真是冷。有多冷?冷得我都希望咱们回到 64 年,那会儿他当皇帝。那时候热乎着呢,你懂的。”说得不错,2000 美元。
“I tell you, it was so cold today. How cold was it? It was so cold, I wish we were back in '64 when he was emperor. Hot times, if you know what I mean.” You are very good, $2,000.
大卫·费鲁奇 \_ 人工智能(续)
David Ferrucci Artificial Intelligence (Continued)
那东西里塞了这么多词。什么是有意义的?它在问什么?你作为人类能够理解这些,因为“64 年的罗马”或“皇帝”在你脑海里亮了,可如果我没有共同经历,我收到的只是一堆词,我怎么知道哪些该亮、哪些不该亮?
There's so many words in that thing. What's meaningful? What is it asking? You as a human connect to that, because the '64 in Rome, or the emperor lights up for you, but how do I know which should light up and what shouldn't light up if I don't have that shared experience, that I just got a ton of words?
“美国未与之建立外交关系的四个国家中,最靠北的那个。” 朝鲜,没错。这些问题你必须答得上来。你必须做到高度精准。要想与最优秀的人竞争,我跟你说说你在精准度上得达到什么水平。
“Of the four countries in the world that the U.S. does not have diplomatic relations with, the one that's farthest north.” North Korea, that's right. You have to able to get these questions. You have to be able to have high precision. To compete with the best, I'll tell you how good you have to be in terms of precision.
你还需要有准确的信心,也就是说,你必须清楚自己知道什么。因为如果你抢答却答错了,就会损失那美元金额,而你的竞争对手则获得它。
You also have accurate confidence, meaning you have to know what you know. Because if you buzz in and then get it wrong, you lose that dollar value and your competitor gets it.
所以你必须要准确预判自己是否判断正确,并且要有一个合理的概率、准确的概率与之对应。你还必须做得非常非常快。大概几秒钟之内——大约两到三秒——才能跟最顶尖的人竞争。
So you have to have a good prediction of whether or not you're right, and you have to have a decent probability, an accurate probability associated with that. You have to do it really, really quickly. Within just a couple of seconds; about two to three seconds to compete with the best.
他们问五花八门的问题。下面这张柱状图显示的就是所有问题的类别分布——以《危险边缘》节目的问题为例,我们随机抽取了 2 万道题,涵盖了电影、乐队、女性、歌手、人物、创始人、疾病、制造者、物品(不管物品是什么)、女英雄、蔬菜、月份、宠物等等。这是一种典型的“长尾”现象。
They ask about all kinds of things. Here’s a histogram of all the kinds of things, Jeopardy! questions, we look at 20,000 randomly-sampled Jeopardy! questions, and they ask about a ton of stuff like films and groups and women and singers and persons and founders and diseases and makers and objects, whatever an object is, a heroine, vegetables, months, pets. It's a very long- tail phenomenon.
即便你试着把对排在最前面的几样东西知道的所有情况都模型化,也只能覆盖 10% 的问题。你根本没法竞争。你得用比那笼统得多的办法。你不可能坐下来,构建一套关于所有常识知识的复杂理论。
Even if you tried to model everything you knew about the top few things here, you'd only cover 10 percent of the questions. No way you're going to compete. You have to do something a lot more general than that. You're not going to sit down and build this complex theory of all common sense knowledge.
13% 的问题甚至没告诉你类型是什么,只写了“这个”或“它”。大家都以为,“哦,类别能告诉你情况如何。”并非如此。在这里,类别对答案类型的指示作用很弱。
13 percent of the questions didn't even tell you what the type was. It just said "this" or "it." Everybody thinks, “Oh, the categories tell you what's going on.” Not true. Here, the categories were weak indicators of what the answer type was.
不过,在这一题里,我们给出的线索是美国城市,答案却是沙狐球。另一题里,答案是伊利运河。这一题,我们给出的线索是乡村俱乐部,答案是权杖、警棍;还有一个答案是国际联盟。关于作者,答案是《约伯记》。答案是罗马尼亚。
Here, if we have U.S. cities but the answer was shuffleboard in this one. The answer was the Erie Canal in this one. Here, we have country clubs. The answer was a mace, a baton; League of Nations was another one there. Author, the answer was the book of Job. The answer was Romania.
这些分类并不能准确告诉你答案的类型是什么。它们只能大致暗示答案可能涉及的方向,并不像大多数人以为的那样是强有力的指示。光是理解问题本身就已经是个挑战。来看这个问题:"这位演员,奥黛丽 1954 年至 1968 年间的丈夫,曾执导她在《绿厦》中饰演鸟女里玛。"这个问题问的是什么?问的是电影《绿厦》的导演。你知道吗?
The categories didn't tell you exactly what the answer type was. It was generally what the answer type might have something to do with. It wasn't as strong of an indicator as most believed. Just understanding the question's a challenge. Here's a question. This actor, Audrey's husband from 1954 to 1968, directed her as Rima the Bird-Girl in Green Mansions. What's it asking for? It's asking for the director of the film Green Mansions. Did you know that?
你只需要拆解这些问题,把句法成分划出来就行,就像你小学上语法课写句子成分那样——但接下来你还得开始解读这些成分。一个“丈夫”是个词,它修饰的是“演员”,可它到底是什么意思?我又不能去一个“丈夫数据库”里查。我甚至不知道在这儿它指的是什么。“直接”这个词呢?“直接的男人”,直接,就像“导演一部电影”里的动词“导演”那样,引导?那又是什么意思?Rima 是谁:一个人、一只鸟、一个女孩、还是一个角色?“绿府”是什么:一个地方、一部电影、一出戏、一本书、还是一栋房子?
You just have to parse the questions, just the syntactic parts. Like back in your grammar school when you wrote out the grammar, but then you got to start interpreting the parts. A husband's a word. It modifies actor but what does it mean? It's not like I can go look it up in a husband database. I don't even know what this means in this case. What about direct? Direct guy, direct like direct a film, lead? What does that mean? Director of...Who's Rima? A person, a bird, a girl, a character. Green Mansion's a place, a movie, a play, a book, a house?
我必须对所有事情做出预测。这些预测将是概率性的。我有许多这样的预测。我同时推进它们,开始收集证据、构建背景,不断缩窄范围来决策:“哪一个可能性更高?这里正确的解释是什么?哪一个解释相比其他更可能是正确的?”
I have to make predictions on everything. Those predictions are going to be probabilistic. I have many of them. I have to pursue all of them and start gathering evidence, building context, narrowing it in to decide, "Which one is more likely? What's the right interpretation here? Which one's more likely the right interpretation over another one?"
如果我是谷歌,我可以直接丢给你一大堆包含这些词语的文档,你就能从中找到答案。但沃森不可能递给亚历克斯·特雷贝克 1500 份文档,然后说,“你自己找答案。”
If I was Google, I could just throw you a whole bunch of documents that contain those words and you could find the answer. But Watson couldn't hand Alex Trebek 1,500 documents and say, “You find the answer.”
大卫·费鲁奇 人工智能(续)
David Ferrucci Artificial Intelligence (Continued)
[laughter]
[laughter]
有个问题。“财政部长蔡斯刚刚第三次把这份东西提交给了我——猜猜看,
Here's a question. “Treasury Secretary Chase just submitted this to me for the third time―Guess what,
好吧。这次我接受了。”认命,挺好的。
pal. This time, I'm accepting it.” Resignation, very good.
你这样做是因为了解历史,还是基于合理推断?B 对吧?是合理推断,没错。因为你坐在那里想:“我不了解这段历史。”但我坐在那里会想:“什么样的事情会被提交上来?”一大堆事情。我会利用更多背景信息——什么样的事情会被提交给一位总统?很可能还是很多事情。然后我会进行元推断,一种元层面的合理推断。什么样的事情值得在电视节目上谈论,并且会被提交给总统?你说:“辞职信。”非常好,这就是正确答案。
Did you do that because you know the history or because you use plausible inference? B? Yeah, plausible inference, that's right. Because you're sitting there, you're saying, "I don't know the history," but I sit there, "What kinds of things get submitted?" A bunch of things. I use more context. What types of things get submitted to a president? Probably still a lot of things. Then I use a meta inference, a meta plausible inference. What types of things get submitted to a president that's worth talking about on a television show? You say, "Resignation." Very good, that's the right answer.
现在,你要是拿这个问题去问一群六年级学生,他们同样会运用合理推断,但他们的共同经历不同,得出的答案也就不一样。
Now, you ask this question to a bunch of sixth graders, and they use plausible inference as well, but they have a different shared experience and they come up with a different answer.
[屏幕上显示“好友请求”,随后笑声响起]
[“Friend request” shown on screen followed by laughter]
这很合理!沃森并不总能解读出问题,从而找到完全匹配的原文段落,这时它就基本得靠预测。它在提升自身智能时采取的做法之一,就是分析大量文本,构建这些句法框架,比如在主语、谓语、宾语以及所有修饰成分前面做解析,然后结合一定的语义知识加以分析,再做出泛化。
It’s reasonable! Watson couldn't always interpret the question such that it could find an exact passage that match, then it would have to basically make predictions. One of the things it did to enhance its intelligence was it would analyze a whole bunch of text, and it would do these syntactic frames, like do these parsing in front of the subject, the verb, the object, and all the modifiers, and then it would analyze them with some semantic knowledge and then generalize.
它会完成像发明家申请专利、官员递交辞呈、人们在学校获得学位这类事情。沃森(Watson)甚至连这些都不懂,它本身一无所知。我没有编写过任何一条规则,团队里也没有人编写规则……
It would get things like inventors patent inventions, officials submit resignations, people earn degrees at schools. Watson didn't even know that, Watson knew nothing on its own. I didn't write a single rule, or no one in the team wrote rules…
它的基本方法是通过分析文本和查找模式来进行推断。官员提交辞呈,人们在学业上获得学位。你会问:“它为什么非得学习这个?”它学到了各种各样的东西。我们没有规定它要学习什么,基本上是让数据驱动它所做的归纳。
It basically inferred things from analyzing text and looking up patterns. Officials submit resignations, people earn degrees at school. You say, "Why did it have to learn that?" It learned all kinds of things. We didn't dictate what it would learn, we basically let the data drive the inductions that it would make.
这挺有意思的,因为我问别人“你在哪儿拿的学位?”时,我可能会回答城市名——那很合理。但更多时候,人们回答的是学校名。
That's kind of the interesting thing, because I asked you, "Oh, where did you get your degree?" I might answer with the city. That would be reasonable. But more often than not, people answer with schools.
这告诉了我一些东西。它告诉我,如果有人问你从哪里获得了学位,我可能更倾向于回答一所学校,而不是其他东西,因为这是基于模型——从所有这些人类文本中推断出的语言模型——所预期的答案。
That tells me something. It tells me, if somebody asks a question about where you got a degree, I might more likely answer with a school than something else, because that's what's expected based on the model, the language model that was inferred from all of this human text.
“流体就是液体,液体就是流体。” 这话对吗?在场有物理学家吗?液体是流体的一种,因为流体还包括气体和等离子体。我们在日常使用语言时,常常混用这两个词。如果你去看正式的学术分类,会发现液体是流体的一种类型,但人们还是混着用。
“Fluid is a liquid, liquid is a fluid.” Which is true? Any physicists out there? Liquid is a fluid, because there are other fluids like gases and plasmas. When we use language, people use it interchangeably. If you went and you looked at a formal taxonomy, you would see a liquid is a type of fluid, but people use it interchangeably.
船只沉没,但人们在台球游戏中打进了 8 号球。看到这些词了吗?台球游戏代表了一个语境。那组词就构成一个语境,从而限制了这些归纳推理的进行方式。
Vessels sink, but people sink 8 balls in the game of pool. See those words there? Pool game, that represented a context. That bag of words would represent a context, so that would limit how these inductions were made.
解析句子、尝试寻找答案、归纳式的技巧——我还会再讲几种技巧。但让我们退一步想想,这项任务有多难。这是一条衡量我们进展的标尺。这上面是一张被称为“胜者云图”的图表。横轴是比赛胜出者答出的问题占比。每个圆点代表一场《危险边缘》比赛。横轴上标绘的,是获胜选手按下蜂鸣器抢到答题机会并得以作答的问题数量。
Parsing sentences, trying to find answers, inductive techniques, I'll talk about a few more techniques. But let's step back and think about how hard was this task. This is a way to measure our progress. This was called the winner's cloud graph up here. We have on the x-axis, is the percent of questions answered by the winner of the game. Every dot here is a Jeopardy! game. What I'm plotting on the x-axis is the number of questions that the winning player buzzed in for and got a chance to answer.
大卫·费鲁奇人工智能(续)
David Ferrucci Artificial Intelligence (Continued)
如果你们看云图中心位置——具体数字我不确定——大概在 48% 左右。平均来看,获胜选手反应够快、也够自信,抢答了大约 48% 到 50% 的题目。
If you look at the center of the cloud, I don't know, something around 48 percent. The winning player, on average, was fast enough and confident enough to buzz in for about 48 to 50 percent of the answers.
在这些问题中,获胜参赛者答对了多少道?大约在 80% 到 90% 之间,大概是 85%。黑点代表的是肯·詹宁斯(Ken Jennings)。肯·詹宁斯本身就是一个现象级人物。他连续赢了 72 场比赛,大概是 71 或 72 场。在他的平均每场比赛中,他能拿下 62% 的题目板。平均下来!这太疯狂了。他还打过一场比赛,拿下了 82% 的题目板。你甚至都搞不清还有谁在跟他同场竞技。他就那么坐在那儿,像这样,答题。
Of those questions, how many did the winning player get right? Somewhere between 80 and 90 percent, about 85 percent. The black dots are Ken Jennings. Ken Jennings is a phenomenon unto himself. He won 72 games in a row, something like that, 71, 72. His average game, he acquired 62 percent of the board. On average! That's crazy. He played a game where he acquired 82 percent of the board. You don't even know who else is playing. He's just sitting there, going like this, answering questions.
82% 游戏时间,他获得了……平均来看,他简直是碾压……看那朵云,简直难以置信。2007 年,这是当时最先进的开放域问答系统。基于所有最先进的技术,我在 IBM 的团队在参与《危险边缘》之前就为此研究了多年。在此之前,这被称为置信度曲线。读图方法是:系统根据置信度对所有问题排序,这是它最有把握的那 5% 的问题。
82 percent of the game, he acquired...On average, he just clobbered...Look at that cloud. That's unbelievable. In 2007, this was the state of the art, open domain, question-and-answering system. Based on all state of the art techniques, my team at IBM was working on this for a number of years before we did Jeopardy! Before, this was called the confidence curve. The way to read this is if it sorted all the questions based on its confidence; this would be the five percent of the questions it was most confident in.
所以,那个系统在《危险边缘》上运行,这是它的表现。最自信的 5% 准确率约 48%。随着它需要回答的问题越来越多,自信心越来越低,最后稳定在约 13%。差距挺大,对吧?问题是啥?
So, that system turned on Jeopardy!, this was its performance. 5 percent that was its most confident did about 48 percent accuracy. As it had to answer more and more, it's less and less confident, and it plateaued at about 13 percent. That’s a big gap, huh? What was the question?
提问:问题出在哪里?
Question: What was the issue?
大卫:问题在于它的置信度曲线很差,预测效果很不好。一条真正优秀的置信度曲线应该是什么样?那就是当它认为自己做得好的时候,确实做得好——曲线应该沿着顶部走;当它预期自己会做得很差的时候,确实就开始表现得差。这才是一条更准确的置信度曲线,形状应该是那样。而这条曲线在任何层面都没做好预测,因此问题很多。
David: What happened was its confidence curve stunk. It wasn't predicting very well. What would a really good confidence curve look like? Where it thought it was doing a good job it was. It'd be along the top. As it predicted it was doing badly, it would start to do badly. That would be a more accurate confidence curve, it would be shaped like that. This was not doing a good job in predicting in any way, so lots of problems.
我们怎么解决那个问题?我之前在给你们方向提示时就谈到过这一点,但当时我们碰到的问题是:“在细胞分裂中,有丝分裂分裂细胞核,胞质分裂则分裂这层缓冲细胞核的液体。”我当时坐在这里,对管理层说:“我会在三到五年内解决《危险边缘》。”你能想象我拿着这类问题,逐个构建出像这样的模型,把整个生物学都表示出来,就为了解一道我再也碰不到的题目吗?那根本行不通。
How are we going to solve that problem? I talked about it before when I gave you a hint with the direction question, but we saw this question: “In cell division, mitosis splits the nucleus and cytokinesis splits this liquid cushioning the nucleus.” I'm sitting here, I told the executives, "I'll solve Jeopardy! in three to five years." Can you imagine me taking each one of these questions and building out a model like this, representing all of biology to get this question I'll never see again? That just wasn't going to happen.
这是一个有趣的解决问题的方式,因为我可以这样说:“你知道吗,我们来看看,它分裂了细胞核的缓冲层。什么液体缓冲了细胞核?缓冲意味着什么?让我搞清楚细胞的结构是什么。”这本来会很棒,因为我会解释我的答案,但这做起来很难。
It's an interesting way to solve it, because then I could say, "You know what, let's see, it splits the nucleus cushioning. What liquids cushion the nucleus? What does it mean to cushion? Let me find out what the structure of the cell is.” It would be great, because I explain my answers, but this is tough to do.
我给你一点提示,这个提示更接近沃森解决那道题的方式:“1898 年 5 月,葡萄牙庆祝了这位探险家抵达印度 400 周年。”有人知道吗?
Let me give you a hint that's closer to how Watson solved the problem in that. “In May 1898, Portugal celebrated the 400th anniversary of this explorer’s arrival in India.” Anybody?
戴维:很好。现在,我可以把这些拆成一组关键词,然后做一次搜索。因为谷歌的 PageRank 算法以及这个主题在高中人人都学,所以非常热门,你很可能搜出麦哲伦和瓦斯科·达伽马,但克里斯托弗·哥伦布也可能出现,因为他匹配了很多关键词,比如葡萄牙、时间范围和印度——反正他当时以为自己要去的是印度。
David: Very good. Now, I can break this down to a bunch of keywords and I can do a search. Because of PageRank on Google and how popular this is because everybody learns it in high school, you're probably going to get Magellan up there and Vasco da Gama up there, but you might get Christopher Columbus, too, because Christopher Columbus hits a lot of these keywords like Portugal and the time frame and India. He thought he was going to India, anyway.
但如果你只看关键词,而不理解上下文,你可能会抓住这样一句话:“五月份,加里在葡萄牙庆祝完结婚纪念日之后抵达印度。”这句话里所有的关键词都对上了,简直太棒了。关键词推理确实很厉害。加里是个探险家。谁能说加里不是探险家?我们每个人在某种意义上都是探险家。下一句你也猜得到。
But if you just strictly looked at keywords and no understanding of what's going on, you could pick up this passage. “In May, Gary arrived in India after he celebrated his anniversary in Portugal.” It’s got all the keywords in common, that's fantastic. Keyword inference is amazing. And Gary's the explorer. Who's to say Gary's not an explorer? We're all explorers in our own way. You can imagine the next sentence.
David Ferrucci 人工智能(续)
David Ferrucci Artificial Intelligence (Continued)
下一句可能有据可查地证明加里是一个探索者,因为那句话可能会说:“加里回到家,探索他的阁楼,找一本相册。”在这里,我们看到“加里”是动词“探索”的主语,这本身就是证据——加里是个探索者。这是好的证据吗?大概不是。
The next sentence could legitimately give evidence that Gary's an explorer, because it might say, "Gary returned home to explore his attic, looking for a photo album." Here we are, we have Gary as the subject of the verb "to explore." That's evidence right there. Gary's an explorer. Is it good evidence? Probably not.
现在,我们遇到了同样的问题。这里写着,“1498 年 5 月 27 日,瓦斯科·达·伽马在卡帕德海滩登陆。”这挺有意思,因为现在,我唯一能匹配的关键词就是 5 月。这条信息甚至根本不会出现,但它里面确实包含了不少有趣的内容。
Now, we have the same question. Now, we have, “On the 27th of May 1498, Vasco da Gama landed in Kappad Beach.” That's interesting, because now, the only keyword I have in common is May. This is not even going to come up, but it's got a lot of interesting stuff in it.
只要我有一整套能够做匹配的算法,我就能把 1898 年的 400 周年纪念和 1498 年匹配上吗?我能把“arriving”和“landed”匹配起来,而不必为每两个词都写一条规则吗?
As long as I had a whole bunch of algorithms that can do the matching. Can I match a 400th anniversary in 1898 with 1498? Can I match arriving with landed without having to write a rule for every two words?
通过统计式改写,我就能发现,“到达(arriving)”和“登陆(landing in)”出现在非常相似的上下文里。从这儿我得到一些信号,从那儿我也得到一些信号,它们可能表示同一件事。然后对于“卡帕德海滩(Kappad Beach)”,借助地理空间推理,我想我可以查一下数据库,然后说:“卡帕德海滩在印度,所以如果我到达卡帕德海滩,我就到了印度。”
With statistical paraphrasing, I could find that gee, “arriving” and “landing in” appear in very similar contexts. Maybe I get some signal that they might mean the same thing. I get some signal here, I get some signal here. And then “Kappad Beach,” with geospatial reasoning, I guess I could look this up in a database, here I can say, "Kappad Beach is in India, so if I arrived in Kappad Beach, I've arrived in India.”
好了,这么看来,瓦斯科·达伽马还算靠谱。我现在能找到的关于瓦斯科·达伽马是探险家的证据,恐怕比我能找到的关于加里的证据还要多。
Now, I got Vasco da Gama looking pretty good. Now, I kind of find out evidence that Vasco da Gama is an explorer, probably more evidence than I can for Gary.
如果我能权衡这些证据,我开始觉得这个答案比加里的更好。这更像是沃森的工作方式。又是那个关于细胞质的问题。它实际上并没有构建一个生物学模型,而是开始找到大量段落,并判断是否存在路径,能从问题中的词语和短语连接到它读到的不同内容。
If I could weigh that evidence, I'm starting to feel like this is a better answer than Gary. This is more of the way that Watson works. Here is that cytoplasm question again. It doesn't actually build a model of biology, but it starts finding lots of passages and figuring out if there are paths for making the connections from the words and phrases in the question to the different things that it reads.
“分裂”能意味着“分开”吗?“减少”能意味着“分裂”吗?“减数分裂”能暗示“有丝分裂”吗?“液体”能暗示“流体”吗?如果我找到资源、其他段落、词典、同义词库,任何能帮我建立这些联系的东西,我就能真正开始构建证据,证明这段文字可能支持那个答案。这正是它做到的事。
Can split mean divide? Can reduce mean split? Can myotic imply mitosis? Can liquid imply fluid? If I could find resources, other passages, dictionaries, thesauri, anything that can help me make these connections, I can actually start to build evidence that this passage may support that answer. This is exactly what it did.
我们拿到这个问题,对它进行分析、解析,提炼出关键词、答案类型(我在找“流动的”这一类的),以及一些关联关系。然后我们会进行大量搜索,找出大量文档。
Here we have that question. We analyze it, we parse it, we come up with the key terms, the answer type, I'm looking for liquid, some of the relations. We then do lots of searches to come up with lots of documents.
从那些资料里,我们找出可能的答案,其实啥也拿不准。嗯,可能是细胞器、液泡、细胞质、细胞膜、线粒体……全都有可能。对每一种可能,我都假设它就是正确答案。你可以这么想:这些就成了互相竞争的假设。
From those documents, we pull out possible answers, not really knowing anything. Well, maybe it's organelle, vacuole, cytoplasm, plasma, mitochondria,…all possibilities. For each one of those possibilities, I assume that's a right answer. Think of it this way. These become competing hypotheses.
我借助搜索,再用一些算法技术,从那些文档中找出可能的答案。现在,每一个答案都成了一个相互竞争的假设,我会单独去验证,说:“我能为这个找到证据吗?我能为这个找到证据吗?我能为这个找到证据吗?”针对每一个假设,我都去收集证据。
I've used search to then use some algorithmic techniques to pull out possible answers from those documents. Each one now becomes a competing hypothesis, and I can go off independently and say, "Can I get evidence for this? Can I get evidence for this? Can I get evidence for this?” For each one of those, I get evidence.
假设我有一个问题,我提出了 100 个可能的答案。针对每个答案,我找到了 100 条可能的支撑证据,这样我就有了 1 万对答案—证据组合。接下来,我想用算法来分析每一条证据。我有 100 种算法来审视每一段文字,判断“这段内容以这种方式或那种方式,是否与另一项内容匹配”。
Let's say I have a question, I come up with 100 possible answers. For each answer, I get 100 possible supporting pieces of evidence, so I have 10,000 now answer-evidence pairs. For each piece of evidence, I now want to analyze it with algorithms. I have 100 algorithms that look at each passage and go, "Does this match this other thing in this way or that way or this way or that way?"
现在面对 1 万张餐桌上的 100 种算法,这就相当于 100 万个不同的评分。我需要对它们进行排序,这正是机器学习发挥作用的地方。我凭什么相信这个评分?我凭什么用这种方式而不是那种方式给出这个评分?我必须做一场庞大的多元回归分析,而我是用机器学习来完成这件事的。
Now with 100 algorithms on 10,000 pieces, it's like a million different scores. Now I've got to rank them, and this is where the machine learning comes in. Why should I believe this score? Why should I give this score this way versus that way? I have to do this giant multiple regression, and I use machine learning to do that.
我在过往的《危险边缘》节目题目和答案上训练模型,这些题目的正确答案和错误答案我都清楚。然后我得到一个大型模型——一组庞大的权重参数——可以应用到所有这些不同的算法技术中。
I train on previous Jeopardy! answers and questions for which I know the right and wrong answer, and I get a big model, a big set of weights that I can apply to all these different algorithmic techniques that I
戴维·费鲁奇 人工智能(续)
David Ferrucci Artificial Intelligence (Continued)
我评估了所有这些答案,从而能够对它们进行排序,并构建一个置信度值,然后将其映射为概率。
evaluated all these answers with. That allows me to rank the answers and to build a confidence value that I can then map to a probability.
难点在于拿出所有算法。这是基础架构:提出问题并进行分析,引入所有这些来源,生成不同的假设,为每个假设找到证据,依据训练结果应用各种模型,然后再对答案进行排序。
The hard part was coming up with all the algorithms. This is the basic architecture. Do the question and analysis, bring in all these sources, generate the different hypotheses, find evidence for each hypothesis, apply the various models based on training, and then rank the answers.
于是,25 名 AI 科学家和软件工程师,来自 IBM 研究院及大学合作伙伴,在数十年铺垫的基础上继续推进。我们用到了搜索引擎,用到了机器学习技术,用到了自然语言处理。我们主要在自然语言处理技术及其与机器学习融合的方式上取得了进展。
So, 25 AI scientists, software engineers, at IBM research and university partners, built on decades of groundwork. We used search engines. We used machine learning techniques. We used natural language processing. We advanced, principally, natural language processing techniques and ways to combine them with machine learning.
我们在四年里完成了超过 8000 次记录在案的实验,积累了海量的实验数据,其规模远超我们用于驱动机器本身的数据——后者大约相当于 200 万本书的内容,不算太多,大约 100 GB 的量级。在分析了 100 GB 数据之后,我们又得到了大约 1 TB 的原始分析数据。
We did, during a four-year period, over 8,000 documented experiments, tons and tons of experimental data, dwarfing the actual data we used to drive the machine itself, which was the equivalent of about two million books – not a lot, in the order of 100 gigabytes. After we analyzed 100 gigabytes, we came up with about a terabyte of raw, analyzed data.
这就是我们的进展。有人问到过信心曲线的问题。这是我们起步时的位置。四年之后——看看这里的顶线——你看那条曲线的形状,多漂亮。随着我们的信心越来越低,我们实际上错得越来越离谱。
This is our progress. Somebody asked about the confidence curve. This is where we started. After four years – this is top line here – look at the shape of that curve. It's beautiful. As we were getting less and less confident, we were actually getting more and more wrong.
如果你看看肯·詹宁斯参赛时的表现,我们当时的水平跟他差不多,接近 90% 的胜率,所以我们能上场。我们能参加《危险边缘》!我们能跟肯·詹宁斯好好较量一场。我们一定会赢吗?不一定。进入比赛时,我们的胜率大概是 73%,也就是说,我有 25% 到 30% 的可能会丢掉饭碗。
If you look at where Ken Jennings was competing, we were doing just about what he was doing, close to 90 percent, so we could play. We could play Jeopardy! We can give Ken Jennings a good game. Would we necessarily win? Not necessarily. Went into the game with about a 73 percent chance of winning, so that's a 25 to 30 percent chance of me losing my job.
[laughter]
[laughter]
大卫:别人告诉我,“你必须赢。”
David: I was told, "You must win."
[laughter]
[laughter]
竞争很激烈。实际上,因为每日双倍奖金的存在,竞争比你想象中还要激烈,但我们最终赢了。这真是非常非常出色的表现。我们几乎花满了那整整四年时间,才把那些算法进化到足够出色的水平。
It was competitive. It was actually more competitive than you might have thought because of the daily doubles, but we ultimately won. This was very, very good performance. It took every bit of those four years to evolve those algorithms to get them to the point that they were doing that well enough.
现在从人工智能的角度回顾这件事很有趣:刚才发生了什么?我们用机器击败了人脑。我们是用什么击败它的?是一台拥有约 2880 个核心、并行运行的机器。
It's interesting to step back now from an AI perspective and say, "What just happened here?" We just beat a human with a brain. What did we beat it with? We beat it with a machine, about 2,880 cores, running in parallel.
还记得我们生成所有那些答案的时候,每个答案都得生成一大堆段落。
Remember when we generated all those answers, and each answer had to generate a bunch of passages.
每一个算法都必须运行这些文本段落。我基本上把它们全部分割开来。所有算法都并行运行,因为如果你用一台机器——一台 3 GHz、64 GB 内存的机器——来运行这个计算,回答一个问题就要花两个小时。那样会让《危险边缘》节目变得非常无聊。
Every one of those algorithms had to run those passages. I basically divided all that up. They all ran in parallel, because if you ran that computation on a single machine, a single three-gigahertz machine with 64 gigs of RAM, two hours to answer a question. It would have made for a very boring Jeopardy! game.
[laughter]
[laughter]
大卫:有了这种并行处理能力,我们把它降到了两到三秒。2880 个核心——一个大脑。体积大约相当于 10 台冰箱——却能装进一个鞋盒。IBM 现在已经把它做得小多了,不过还没到鞋盒那么大。驱动它需要 80 千瓦电力——大约相当于一份金枪鱼三明治和一杯牛奶的能量。20 吨的制冷量——一只手摇扇就能搞定。4 年时间,大约 200 万本书的内容——相当于 30 年的学习量。
David: With that parallelization, we were able to get that down to two to three seconds. 2,880 cores―one brain. Size of about 10 refrigerators―fits in a shoebox. IBM's made it much smaller now. Not quite the size of a shoebox, though. 80 kilowatts of electricity to power it―tuna fish sandwich and a glass of milk, maybe. 20 tons of cooling―hand fan. 4 years and about 2 million books of content―about 30 years of learning.
大卫·费鲁奇 人工智能(续)
David Ferrucci Artificial Intelligence (Continued)
沃森必须有一只“手”——我不确定你是否知道这点。设计思路是,电脑与游戏系统要有相同的交互接口,人类需要用手按那个小塑料按钮,所以电脑也必须物理按下那个小按钮。实际上他们真给它造了一只“手”,一个装在蜂鸣器上、能把按钮按下去的小型机械手。
Watson had to have a hand. I don't know if you know that. The idea was that the computer would have equivalent interfaces to the game system, and humans had to push that little plastic button down physically, so the computer had to push that little button down physically. It actually had to be made a hand, a bit of a hand that's on top of that buzzer that would push it down.
在比赛过程中,电脑好几次在抢答环节被人抢先。所以,虽说整体上它的速度更快——这取决于问题类型、回答速度以及它的自信程度——但在抢答上它还是输了。事实上,它被抢答的次数还不少。
The computer was beat to the buzz several times during the game. So, while in general it was faster depending on the questions and how quickly it could answer and how confident it was, it was beat to the buzz. In fact, it was many times beat to the buzz.
我想谈一个非常有意思类型的问题。《终极危险!》的题目就是这样。这类题目对电脑来说更棘手。观察这些题目很有意思。题目是这样的:“听闻发现乔治·马洛里的遗体时,他告诉记者,他仍然认为自己才是第一个登顶的人。”
I want to talk about a question that was a very interesting type of question. Final Jeopardy! questions were like this. Final Jeopardy! questions were harder for the computer. It was interesting to notice about these questions. Here's the question: “On hearing of the discovery of George Mallory's body, he told reporters, he still thinks he was first.”
顺便问一句,有人知道这个问题的答案吗?是埃德蒙·希拉里,没错。我们一直觉得这是个缺失的环节。你得从乔治·马洛里以及所有跟他与珠穆朗玛峰相关的事情出发,然后想:“谁是第一个登上珠穆朗玛峰的人,然后才想到埃德蒙·希拉里?”
Does anybody know the answer to the question, by the way? It's Edmund Hillary, right. We thought of it as a missing link. You had to go from George Mallory and all of the things he's connected to Mount Everest and then think, "Who was first to Mount Everest and get Edmund Hillary?"
这个问题里缺了一个环节。无论你的自然语言评分有多好,都不太可能出现某一段落直接告诉你:“这段话正在回答这个问题,我有 65% 的把握认为这就是答案,依据是这段话、这段话和这段话。”
There was a missing link in the question. It’s unlikely that there was any one passage, no matter how good your natural language scores were, that can sit there and say, "This passage is answering this question, I’m 65 percent sure that this is the answer based on this passage and this one and this one and this one."
它真正需要结合两段信息,完成两次跳跃。从这一段跳转到乔治·马洛里和珠穆朗玛峰,再从珠穆朗玛峰链接到埃德蒙·希拉里。这类问题的难度要大得多。实际上,这直接引出了我离开 IBM 前所做的下一件事——沃森路径系统(WatsonPaths),而他们此后一直在不断迭代这个系统。
It really had to combine information from two passages. It had to make two hops. At the hop from here to George Mallory, Mount Everest, and then make the link from Mount Everest to Edmund Hillary. It was a much harder kind of question to solve. In fact, that led to the next thing, the thing I did right before I left IBM was this WatsonPaths stuff which they've been evolving ever since.
基于这一观察,医学领域的思路是:当你审视医学问题时,情况要复杂得多。你并不知道路径是什么,而且可能需要考虑多条路径。前提是问题确实集中在某个点上。你必须将问题拆解开来,并建立这些关联。这正是 WatsonPaths 所做的——它通过文献定义出路径。
Based on that observation, the idea was, in medicine, when you look at medicine, it's much more complicated. You don't know what the passage is and you may have to consider many of them. The presumption is in that it's really in one place. You have to break the question down and make these connections. That's what WatsonPaths does. It defines paths through the literature.
这里有一个 63 岁的病人被送去神经科医生那里。临床表现为两年前开始出现的静息性肿瘤,诸如此类。这其实是美国医学执照考试的一道题。题中给你一堆信息,然后问:“他的神经系统哪个部位最有可能受到影响?”这就是那道题。
Here, we have a 63-year-old patient is sent to neurologist. The clinical picture of a resting tumor that began two years ago and so forth. This is actually a United States Medical Licensing Exam question. It gives you a bunch of information and says, "What part of his nervous system is most likely affected?" It was the question.
沃森路径(WatsonPaths)的做法是将它拆分成不同的片段。然后基于不同的问题,它会试图寻找关联,问题非常宽泛,比如:“这与什么有关?它指向什么?它包含什么?它影响了什么?它由什么引起?”
What WatsonPaths would do is break it up into different segments. Then based on different questions, it would try to find links, very general questions like, "What is this associated with? What does it indicate? What does it contain? Was has it affected? What is it caused by?"
它会基于自己的置信度对这些连接进行加权,然后落到不同的地方。它可能会落到帕金森病上,但这不是问题的答案。然后它可能会落到基底神经节上,落到这里,又落到那里。接着,它会从那里脱颖而出。
It would weight those connections based on its confidence, and then it would land in different places. It would go to Parkinson's disease, but that's not the answer to the question. Then it would go to basal ganglia, it would go here, and it would go there. Then it would stand out from there.
当它从输入朝这个方向进行时,它还会从可能的答案出发——要么是一道多选题,要么生成一组可能的答案,就像最初的沃森系统那样——然后从这组答案反向推导,看它们在中间什么地方相遇。
As it was going from the input in this direction, it would go from the possible answers, which is either a multiple choice question or it would generate a set of possible answers, just like the original Watson did. And it would go from the set of questions backwards and see where it met in the middle.
这样一来,你就能回答问题了——不是假设答案来自某个固定位置,而是可以沿着文献的脉络,去阅读这一段和那一段的内容,从中找到答案的解释。
This now allowed you to answer questions by not assuming that the answer was described in one place, but that you could find paths through the literature that if you read the passages here and the passages here, you'd see an explanation for what the answer was.
戴维·费鲁奇人工智能(续)
David Ferrucci Artificial Intelligence (Continued)
没有一条 if-then 规则,没有任何人编写过领域专属的 if-then 规则,WatsonPaths 却能找到从输入到答案之间的关联,尽管它并未构建任何模型。这说的正是增强智能,也是我喜欢肖恩那场演讲的原因。它的意思是:“你,人类,我其实不太明白这里发生了什么,但你,作为人类,如果读了这段文字,又读了那段文字,你很可能会相信‘黑质’就是答案。”
Without a single if-then rule, without a single human writing a domain-specific if-then rule, WatsonPaths would be able to find connections from the input to the answers even though it didn't build a model. Talk about augmented intelligence, which is why I liked Sean's talk. It said, “You human, I don't really understand what went on here, but you, the human, it's likely that if you read the passage here and you read the passage here, you'd be convinced that Substantia Nigra is the answer.”
你可以想象这个链条会延伸三四步甚至五步之深。对于一个来自那种基于理论的 AI 方法、构建了那些规则和学到的知识、然后发现“哇,我从这里就是走不通”的人来说,这截然不同。拥有这样一个系统来增强人类如何通过文献建立这些联系的能力,是一个相当有力的想法,也是我认为沃森(Watson)最令人期待的东西之一。
You could imagine this going three or four or five steps deep. This is very different for a person who came from sort of that theory-based approach in AI, building all those rules and learning, "Wow, I just can't get there from here." Having a system like this to augment how humans can make these connections through the literature, sort of a very powerful idea, one of the very promising things that I think came out of Watson.
那么,谈谈一些反思。在构建系统架构理念上取得的巨大成功,对 Watson 的成功至关重要。换句话说,不是坐在那里说“我要为《危险边缘》做些具体的事”,而是提出一个通用架构:你进行问题分析,生成假设,让假设相互竞争,收集证据,为证据打分,利用机器学习来权衡证据。这种通用架构极其强大。它将让我们能够构建许多、许多算法,插入到这个架构中。
So, some reflections, huge triumph on building this notion of systems architecture, was hugely important to succeeding at Watson. In other words, not sitting there and saying, "I'm doing something specific for Jeopardy!," but coming up with a general architecture, you do question analysis, you generate hypotheses, you let the hypotheses compete, you gather evidence, you score the evidence, use machine learning to weigh the evidence. That general architecture was extremely powerful. It will let us build many, many algorithms to plug into that architecture.
多元方法融合也是一大胜利。我的团队有很多研究人员,他们都在以下意义上独立工作:“我要研究一个生成候选答案的算法。”“我要研究一个评分否定句的算法。”“我要研究一个做这个的算法。”“我要研究一个做那个的算法。”很多、很多不同的想法。
A triumph for combining a diversity of methods. My team, I had a bunch of researchers, and they all worked independently in the following sense: "I'm going to work on an algorithm to generate candidates.” “I'm going to work on an algorithm for scoring negation.” “I'm going to work on an algorithm for this.” “I'm going to work on an algorithm for that." Many, many different ideas.
其中一些算法在意图上会有所重叠,但那没问题。我不需要微调这些算法之间的关系。我只是把它们放进这个架构,然后用机器学习对以往的问题和答案进行训练,来确定这些权重应该是什么。
Some of those algorithms would overlap in their intent, but that was OK. I didn't have to fine tune the relationships between those algorithms. I just put them into this architecture, and then I used machine learning to train on former questions and answers to figure out what those weights should be.
这一点非常让人联想到马文·明斯基(麻省理工学院人工智能的创始人之一)的观点。他写了一本书叫《心智社会》,其中主张智能的本质就是这样。大脑中并没有一个单一的、知道答案的系统。而是有很多不同的计算,这些计算输入信息,然后被组合起来。
And this is very reminiscent of Marvin Minsky's view, one of the founders of artificial intelligence, at MIT. He wrote a book called The Society of Mind, where he argued that intelligence was of this nature. There wasn't one system that knows the answer in your head. There are lots of different calculations that input and get combined.
例如,Watson 中没有任何一个组件了解生物学,或者知道如何回答某个问题,而是许多、许多组件产生信号,这些信号被组合起来。
For example, there wasn't any one component of Watson that knew anything about biology or that knew how to answer the question, or many, many things, that produce signals that those signals got combined.
另一个观察是,我们尚未让机器学会理解并解释。Watson 并没有为它读到的内容建立一个逻辑模型——那种你我能够坐下来,开始对模型进行推理并说“你为什么这么认为?逻辑联系是什么?你能对你刚读到的东西进行演绎推理吗?”——我们还没做到这一点。我们尚未构建出这样一个系统:它能够读取任何内容,并为其建立一个与人类兼容的逻辑模型,并且能够通过演绎推理对其进行思考。
The other observation was we have yet to make machines learn to understand and explain. Watson did not build a logical model for what it read, that you and I can sit down and start reasoning over that and saying, "Why do you think that? What are the logical connections, and can you do deduction over what you just read?" We have yet to do that. We have yet to build a system that can read anything and build a human compatible, logical model for that and be able to reason over it with deduction.
我认为事态将如何演变,在以下方面与肖恩的观点非常相似。我不认为你可以把人从这个方程中移除,因为人类认知是意义的来源。
My view of how things are going to evolve is very similar to Sean's in the following way. I don't think you can take humans out of this equation, because human cognition is the source of meaning.
正如我在演讲开头所说:“词语是关于什么的?”词语只是指向共享人类经验的指针。
As I said in the beginning of the talk, "What are words about?” Words are just pointers into a shared human experience.
每当我们获取一条数据,我们都必须根据我们的价值体系、我们的目标函数、以及我们对数据背后意图或意义的理解来对其进行解释。
Every time we take a piece of data, we have to interpret it with regard to what our value system is, what our objective function is, what our intent behind or meaning behind that data is.
大卫·费鲁奇 人工智能(续)
David Ferrucci Artificial Intelligence (Continued)
然后,如果我们想与其他人类交流,我们就必须将其映射到一个理论。我们必须把盒子打开,而不仅仅是拥有所有特征。你可以在这里放入很多特征,我们可以让机器产生归纳假设,基本上是相关性:“这个相关性看起来有趣。那个相关性看起来有趣。”但我现在如何解释这些数据,将其映射到理论,以便我能对其进行推理,并在这个过程中融入人类对世界的理解?
Then if we want to be able to communicate with other humans, we're going to have to map that to a theory. We're going to have to open the box up, not just have all the features. You can put a lot of features in here, and we could have the machine produce inductive hypotheses, basically correlations: "This correlation looks interesting. This correlation looks interesting." But how do I now interpret that data, map it to a theory, so that I could reason about it, and engage the human understanding of the world in that process?
你可能会说:“如果我能预测出正确答案,人类理解就不重要了。”但你真的相信这一点吗?因为如果你想让人类参与其中,增强他们扩展理论、解释理论、与其他人类交流的能力,那么人类将如何扩展他们对世界的理解?他们将如何扩展他们的价值体系、对周围世界的解读,如果他们只是让机器处理数据的话?更不用说,你看到了所有那些伪相关的现象,它们到底意味着什么?
You could say, "The human understanding doesn't matter if I can predict the right answer.” But do you really believe that? Because if you want humans to engage in that, to augment their abilities to extend the theory, to explain the theory, to communicate with other humans, how are humans going to extend their understanding of the world? How are they going to extend their value system, their interpretation of the world around them if they're just letting machines process the data? Not to mention, you saw all the spurious correlations that happen, what do they really mean?
我对未来的看法是,作为系统,机器智能最终将演变为这样一种理念:“我们如何利用计算机来加速这个良性循环?我们如何利用计算机来快速处理数据,以产生并证明归纳假设?我们如何让人类机器界面参与进来,以理解这一点并有效地处理它?我们如何支持理论的快速形成和测试?”
My view of the future is that as systems, ultimately, machine intelligence, is going to evolve into this notion of, "How do we use computers to accelerate this virtuous cycle? How do we use computers to rapidly process data to produce and prove inductive hypotheses? How do we engage the human machine interface to understand that and be effective with processing that? How do we support the rapid formulation and testing of theories?"
我非常全面地看待计算机的角色,试图加速这个良性循环。我想我就讲到这里。
I look at the role of computers very holistically in trying to accelerate that virtuous cycle. I think I'm done.
[applause]
[applause]
提问:在某个时刻,你展示了游戏中的这部分,你说这对你非常重要。答案是概率,然后是两个备选项(如果你愿意这么说),也带有概率。
Question: At one point, you showed that this part of the game, you said it's very important to you. The answer was a probability, and then two covers if you will, with probabilities.
我试图跟上你的对话。在我看来,很明显你经历了这些步骤。我本以为在每个步骤中,你都在通过这些步骤推导出最可能的答案,但让人惊讶,或者也许需要补充说明的是,它最终是如何得出那个概率估计的,以及为什么它很重要?它是否有一个完整的分布,比如第二佳和第三佳的答案?
I'm trying to follow through your conversation. It seems to me, it's clear that you go through these steps. I would think at each step, you're deducing what's the most probable answer through these steps, but it's surprising or maybe to add some color on, how does it arrive at that probability estimate at the end, and why is it important? Does it have a whole distribution that there's a second best and third best answer?
感觉你编译了很多东西……
It just seems like a lot of things you're compiling....
大卫:在管道(pipeline)的末端之前,它并不会计算这些概率。它做的是,接受问题,然后利用整个过程来生成可能的答案。因此它会变成相互竞争的假设,比如:“我认为答案是瓦斯科·达·伽马,是麦哲伦,是哥伦布,或者加里。”我得到了这些相互竞争的假设。
David: It doesn't compute those probabilities until the end of that pipeline. What it's doing is it's taking the question and it's using the whole process for generating possible answers. So it would become competing hypotheses, like, "I think it's Vasco da Gama, it's Magellan, Columbus, or Gary." I got these competing hypotheses.
现在,对于每一个假设,我会去说:“让我们假设这是正确答案。我能找到支持它的证据吗?”现在你可以把这个过程想象成分成四个独立的流,每个流各奔东西,甚至还没考虑其他流,每个流都去尝试找到支持该假设的证据。
For each one of those now, I go and I say, "Let's assume this were the right answer. Can I find evidence for it?" Now you could think of the process literally breaking up into four independent streams, each one going off and, not even thinking about the other one yet, and each one going off and trying to find evidence in support of that.
但现在,这些证据是独立评分的,意思是,如果我发现一段支持哥伦布的段落,将会有 100 个算法来评估这段段落,并给出一个分数。这个分数表达的是:“这段段落支持或反驳哥伦布作为答案的可能性有多大。”我会得到所有这些分数。
But now the evidence is independently scored, meaning that if I find the passage supporting Columbus, I'm going to have 100 algorithms evaluating that passage, and it's going to give a score. That score is saying, "Here's the likelihood this passage either supports or refutes Columbus as the answer." I'm going to get all these scores.
每个答案都有所有这些段落,每个段落都有一个分数。现在,首先,我必须把它们组合起来,然后说:“我从许多不同的段落和许多不同的算法那里得到了关于这个答案的所有分数,我该如何组合它们?我该如何权衡它们?”
Each answer has all these passages and each one has a score. Now, first of all, I have to combine it and say, "I have all these scores for this answer from lots of different passages and lots of different algorithms, how do I combine those? How do I weigh them?"
大卫·费鲁奇 人工智能(续)
David Ferrucci Artificial Intelligence (Continued)
这就是你会在之前的问题和答案上进行训练的地方,以确定:“我会给我的类型化算法(typing algorithms)赋予比句法语法算法(syntactic grammar algorithms)更高的权重,比这个更高,比那个更高,无论我打算如何权衡它们。
That's where you would train over previous questions and answers to figure out, "I'm going to weigh my typing algorithms higher than my syntactic grammar algorithms, higher than this one, higher than that, however I'm going to weigh them.
我得出了所有权重,然后得到了康伦布的最终分数。我得到了麦哲伦的最终分数;我得到了这个的最终分数。这些分数在当时只是置信分数。
I come up with all the weights, and I get this final score for Columbus. I get a final score for Magellan; I get a final score for this. These scores are just confidence scores at that point.
基于大量的训练数据,我可以看到我的置信分数——它们的效果如何,以及它们如何映射到回答一个问题的正确或错误的可能性。既然我可以在很多答案上运行它,我现在就可以根据那个置信分数,将其映射到概率上。现在每个答案都有一个概率,即正确的可能性。
Based on a lot of training data, I could see my confidence scores – how well they work, how they map to a likelihood of getting a question right or wrong. Since I can run it on lots of answers, I can now take that confidence score, and I can map it to a probability. Now each answer has a probability, the likelihood of being right.
提问:这些概率通常都那么接近吗?比如 32%,26%。
Question: Are they typically that tight? It was like 32 percent, 26 percent.
大卫:是的,系统会这么说。它到概率的映射相当不错。换句话说,如果系统说“我有 60% 的机会答对”,那么当它说 60% 的可能性时,60% 的情况下它确实会答对。
David: Yeah, that's what it would say. It was a pretty good mapping to probability. In other words, if the system said, "60 percent chance I got this right," then when it says 60 percent chance, 60 percent of the time it would get it right.
这就是我们做的方式。我们有足够的数据来做这个。你并不总是能这么做,但我们有足够的数据。是吗?
That's how we did it. We had enough data to do that. You can't always do that, but we had enough data to do that. Yeah?
提问:我认为这是一个精彩的演讲。非凡的工作。如果 Watson 能够处理一个关于俄罗斯高级领导人之间职位互换的问题——梅德韦杰夫和普京——那么要让 Watson 处理一个关于梅德韦杰夫和普京是否会再次进行职位互换的问题,你需要做多大的改动?
Question: I thought that was a beautiful presentation. It's extraordinary work. If Watson could handle a question about a job swap between senior Russian leaders―Medvedev and Putin―how much would you have to change for it to handle a question about whether Medvedev and Putin will have another job swap?
大卫:让我看看我是否理解对了。我的意思是,你可以读到关于一次职位互换是如何发生的……具体是谁并不重要,两位领导人。问题是,Watson 如何预测是否还会有另一次职位互换?
David: Let me see if I got this right. What I'm saying is you can read about how there was a job swap between...It doesn't really matter who, two leaders. The question was how would Watson predict whether there was going to be another job swap?
简短的回答是,Watson 只会告诉你已经写下的东西。这就是我之前做出的区分。它没有在构建一个关于正在发生之事的内部逻辑表征。它没有构建一种理解,以便之后能够进行推理并说:“如果是这样,那么就会那样。这些是可能导致这次职位互换的原因。因此,如果我再次看到那些原因,我就会那样做。”
The short answer is that Watson is only going to tell you what has been written. That was the distinction I was making. It's not building an internal logical representation for what's going on. It's not building an understanding that it can then reason over and say, "If this, then that. These are the reasons that likely caused this job swap. Therefore, if I see those reasons again I'm going to do that."
你可以构建一个类似的系统,说:“我要寻找在发现职位互换之前出现的语言上的先行指标。”那将是一个非常独立的、特定的应用。
You could build a similar system to say, "I'm going to look for linguistic prior indicators that show up before I then see a job swap." That would be a very separate, narrow application.
提问:这类东西似乎在消费市场有很多应用,然而 IBM 却没有对此进行任何开发。这背后的原因是什么?
Question: There seems to be a lot of application with something like this in the consumer market, yet IBM's done nothing to develop that. What's behind that?
大卫:IBM 正在很多领域大力投资 Watson 的应用。具体细节,我现在不在那里工作。我无法告诉你商业策略是什么。抱歉,但你说得对。我认为它在那个领域确实有应用。
David: IBM's investing a lot in the application of Watson into a bunch of areas. The specifics, I don't work there right now. I couldn't tell you what the business strategy is. Sorry, but you're right. I think it does have applications in that space.
提问:你能谈谈决定在 Daily Double 中下注多少的算法所涉及的因素吗?
Question: Can you talk a little bit about the factors involved in the algorithm for determining how much to bet in the Daily Double?
[laughter]
[laughter]
大卫:这是一个很好的问题。人类在自然语言处理方面做得太出色了,以至于他们对投注策略更着迷。我们和之前的《危险边缘》参赛者玩过 100 多场练习赛,之后我们会采访他们,他们总是对投注更感兴趣。因为那是他们觉得困难的部分。[笑]
David: That's a great question. Humans do the natural language processing so well that they're much more fascinated with the betting strategy. We played over 100 practice games against former Jeopardy! contestants, and then we would interview them after, and they were always much more interested in the betting. Because that's the part they found hard. [laughs]
大卫·费鲁奇 人工智能(续)
David Ferrucci Artificial Intelligence (Continued)
我来回答你的问题……不过先让你有个概念,当时大概只有 1 个人在做投注,而有 24 个人在搞那些算法。这事挺有意思的,那一个人甚至还不是全职的。
I'm going to answer your question…but just to give you a sense, there was like one person working on betting, and there were 24 working on the algorithms. It's interesting. It wasn't even a full person.
尽管如此,他们还是做出了一些有趣的事情。
Nonetheless, they did some interesting things.
有位学者就此发表了一篇论文,用神经网络方法做了些研究。他们分析了人类的下注模式。
A paper was published on it. They did some neural network stuff. They looked at human betting patterns.
当时涉及一点儿博弈论。
There was a little bit of a game theory going on.
人类就是这样下注的,那我该怎么做?他们用神经网络技术精确计算出,怎样才能以最优方式下注,最大化获胜概率。他们发现了一些有趣的事,比如人类在“每日双倍”环节的下注金额远低于应有水平。人们下注不够激进,诸如此类。
This is how humans are likely going to bet so what will I do? They used neural network techniques to figure out exactly what's the best way to bet to optimize your chance of winning. They learned interesting things, like humans bet way too low on Daily Doubles than they should. They don't bet aggressively enough, things like that.
不管怎样,我们用了那些算法。在“每日双倍”上的下注和“最终危机”环节的下注非常不同。
Anyway, we used those algorithms. We'd come up with a betting…Daily Double betting was very different than the Final Jeopardy! betting.
我们会运用先验概率——判断我们答对或答错某道题的可能性,这个概率在“每日双倍奖”环节与“最终危险!”环节并不相同。如果我们了解到自己在某个分类中的表现情况,这些先验概率就会随之改变。
We would use our priors, our likelihood that we would get a question right or wrong, which was different from Daily Doubles than it was for Final Jeopardy! Those priors would change if we learned something about how we were doing in that category.
最后一点我要说的是,我记得负责这个项目的 IBM 员工格里·特索罗(Gerry Tesauro)走进来跟我说:“戴夫,我们在算出像 126 美元和 967 美元这类相当怪异的数字。你希望我们四舍五入到 10 的倍数,对吧?”
The final bit that I'll share is that, I remember when the guy working on this, Gerry Tesauro at IBM, came in and said, "Dave, we're coming up with really funky numbers like $126 and $967. You want us to round it to the nearest 10, right?"
我当时心想:“我为什么要让你这么做?那只会让人类算起来更方便罢了。(笑)保持原样别动。”
I was like, "Why would I want you to do that? That's just going to make the math easier for the humans. [laughs] Leave it the way it is."
问:如果要校准完美数字记忆与人类大脑模拟系统各自的作用,在最终结果中它有多重要?
Question: If you were to calibrate the role of perfect digital recall versus the analog system for the human brain, how important was that in the results?
大卫:那次召回事件有多重要?
David: How important was the recall?
提问:作为人类,拥有完美记忆和模糊记忆之间的差别。“我知道答案,就在嘴边,但就是想不起来,”这种事不会发生在 Watson 身上。
Question: The ability to have perfect recall compared to fuzzy recall for a human being. "I know the answer. It's on the tip of my tongue, but I can't get it," which doesn't happen to Watson.
戴维:这个嘛,那得看你指的是什么。有意思的是你提到了召回这件事。我想从两个角度来回答这个问题。
David: Well, it depends what you mean by that. It's interesting that you bring up the point of recall. I want to answer the question two ways.
首先我想谈谈检索。精确度是指我能否答对问题。而检索能力是“我能不能记住……我的记忆库里有答案吗”,可以这么说。沃森系统会考虑顶端的结果——一整套可能的答案。假设是从不同来源提取出的前 500 个答案。在这些前 500 个、甚至前 1000 个答案中,我们试验了不同数量的候选答案。
First, I just want to talk about recall. Precision is whether I get a question right or wrong. Recall is, "Can I even remember...Do I have the answer anywhere in my memory bank," if you will. Watson would consider the top end result, a whole bunch of possible results. Let's say the top 500 answers that it was able to extract from various sources. In those top 500, even 1,000 answers, we experimented with different numbers of competing answers to consider.
只有 85% 的问题被召回。换句话说,有 15% 的问题我们哪里都找不到。
Only 85 percent were in the recall. In other words, 15 percent of the questions we couldn't find anywhere.
我们根本没什么好考虑的。那种比赛里,如果你看到自己的正确率,不管具体多少,达到 85%,那已经做得非常、非常好了,因为我们抢答的总数本身没有多高——那是你按下抢答器的那些题里的正确率。
We had nothing to consider. When you see getting, whatever the accuracy was, 85 percent right, that's doing really, really well in that game because our recall total was not that high, that's of the questions you buzzed in for.
方法,这是第一点。第二点——所谓“知道”、“不知道”或“努力回忆某事”,究竟意味着什么?显然,沃森系统如果没有答案,是不会继续推进的。真正的难点在于:“我对这个答案有多大的把握?”
Anyways, that's one. The other one – what does it really mean to know or not know or to struggle remembering something? Clearly, Watson's not going to go forward if it didn't have an answer. The real struggle was, “how confident am I in that answer?”
所以我想问,所谓的“召回能力”到底意味着什么?沃森坐在那里说,“我不知道。也许
That's why I question what does it mean to have recall. Watson's sitting there going, "I don't know. Might
可能不会。该按铃吗?不该按铃吗?
be. Might not be. Should I buzz? Should I not buzz?"
大卫·费鲁奇人工智能(续)
David Ferrucci Artificial Intelligence (Continued)
我明白这个。确实如此。但这种事有多确定?我同意有区别,但区别到底在哪?因为人类会坐在那里嘀咕,“哦,就在嘴边了”,却怎么也想不起来。而沃森在纠结的是:“这到底对不对?”我不清楚。
I got a handle on that. That's true. But how sure is it? I agree there's a difference, but what is exactly the difference? Because humans will sit there and go, "Oh, it's on the tip of my tongue." They're not being able to remember it. Watson struggled with, “Is it even right?” I don't know.
最后,我想澄清一下我当时认为这话是什么意思,但连我自己也不知道答案。
In the end, I wanted to clarify what I thought that meant, but I don't know the answer.
提问:如果肯·詹宁斯看到沃森给出的概率列表,他能否比沃森取得更好的成绩?
Question: If Ken Jennings were to see the list of probabilities that Watson offered would he have improved on Watson's results?
戴维:这个问题问得好。大概是会的。这是个好问题。我觉得应该会。你问的那个召回问题,从这个角度想挺有意思的。
David: That's a great question. Probably yes. That's a great question. I think probably yes. That's an interesting way to approach your question about the recall thing.
以下是一些可能的情况。考虑到这些你可能无法回忆起来的情况,你现在能否提取出足够区分这两者的额外信息?这倒是挺有意思的。
Here's a bunch of possibilities. Given those possibilities that you might not have been able to recall, can you now retrieve the additional information you need to distinguish between those two? That's kind of interesting.
如果你观察过 Watson 的运作方式,就会发现它一开始会生成海量的可能性。然后它会问:“我能不能找到更多的信息,来帮助我对它们进行区分?”如果你错过了最开头的那一步……
If you look at the way Watson worked where initially it would generate tons of possibilities. Then it would say, "Can I find additional information that will allow me to differentiate them?" If you missed that first part…
Your question?
Your question?
问题:当沃森用不同方式回答问题时,你是如何设定其阈值规则的?
Question: How did you write the rules for the threshold for when Watson answered a different way?
大卫:这个问题问得好。实际上,这个标准是会变的——可能在比赛过程中变化。我们可以调整那个门槛的上下限。
David: That's a great question. That actually changed. That could change during the game. We could move that threshold around.
事实是,如果沃森大幅落后,它会变得更加激进,并降低门槛。它会一直盘算着:“戴夫完蛋了。我索性冒个大风险。”
What happened is if Watson was way behind, it would actually be more aggressive, and it would lower the threshold. It would have been sitting there thinking, "Dave's screwed. I might as well take a bigger chance."
然后,如果它遥遥领先,它就会对自己说:“何必表现得像个傻瓜呢?”实际上,它会提高门槛,在高门槛下作答。它会在游戏过程中根据自身所处的局势进行校准。
Then if it was way ahead it would say, "Why look stupid?" And it would actually raise the threshold and answer at a high threshold. It would calibrate during the game depending on where it was.
问题:那它调整了算法吗?
Question: Did it adjust the algorithms?
大卫:不,那样做并不会调整算法。它只是在做权重分配。当处于某个类别中时,它进行的是另一种学习。它会动态地完成这一过程。
David: No, it would not adjust the algorithms. It was just doing weighting. There was a different kind of learning going on when it was in a category. It would do this dynamically.
在《危险边缘》节目里,真正聪明的玩法是一开始就去找“每日双倍”。不找“每日双倍”而从类别顶端开始的好处是,你能以低成本摸清这个类别到底是怎么回事。
The really smart way to play Jeopardy! is to hunt for Daily Doubles right from the beginning. The advantage of not hunting for Daily Doubles and starting at the top of the category is because at a low cost, you can learn what the category is about.
第一个问题对于预测答案类型的准确性比类别本身更高。这就是为什么沃森一开始会以较低的置信度作答,然后根据情况调整可能性。如果它答对了这些问题,就说明它正在学习答案类型。当它碰上“每日赌注”题时,它会据此调整它的先验概率。
The first question is a better predictor for what the answer type is than the category itself. That's why Watson would go in low and then it would adjust its likelihood. If it was getting those questions right, it was learning the answer type. When it hit a Daily Double, it would adjust its prior for that.
问题:如果你落后了,而且这是孤注一掷的局面,沃森是否会接受一个负期望值的赌注?
Question: Does Watson ever take a negative expectation bet if you're behind and it's all or nothing?
大卫:我不知道。至少,让我这么说吧,不是我设计的。不过,如果它偷偷溜进去了的话【笑】。
David: I don't know. At least, let me put it this way, not that I designed. Now, if that snuck in there [laughs] .
问题:《危险边缘》里有不同类型的问题,你展示的有些问题有多条路径能通向答案。“绿苑春浓”这道题,你可以知道奥黛丽·赫本的丈夫是谁,也可以知道梅尔·费勒执导过这部片子。
Question: There are different types of questions in Jeopardy! and some of the ones you showed have different paths to get to the answer. The Green Mansion is one you could know who Audrey Hepburn's husband was or you could know that Mel Ferrer directed.
戴维·费鲁奇 人工智能(续)
David Ferrucci Artificial Intelligence (Continued)
大卫:没错。
David: That's right.
问 题:你觉得“沃森”是在只有一条路径通向答案的问题中,还是在有多种路径通向答案的问题中,对人类表现更出色?
Question: Do you think Watson would do better against humans in the question where there's just one route to the answer or where there are multiple routes to the answer?
大卫:在有多条路径且每条路径的证据都足够充分的情况下,沃森很可能表现更好。这其实就是它寻找的目标。它在这个意义上并不区分不同路径。如果某条路径——也就是问题的某个部分——带它找到了得分极高的文段,它就会直接判定为“良好”。
David: Watson would likely do better where there were multiple routes provided that the evidence for any given wrote was sufficient enough. That's really what it was looking for. It didn't distinguish between routes in that sense. If some route, which is one portion of the question, got it to a passage that had scored very high, it would just say good.
问题:你能说说沃森什么时候会抢答吗?
Question: Can you talk about when Watson would buzz in?
大卫:所以它唯一做的事情,就是实施了我称之为“信心加权按铃方案”的计划。
David: So the only thing it did there was it had what I would call the confidence-weighted buzzer scheme.
换句话说,它会得出一个答案,并计算该答案的置信度。
In other words, it would come up with an answer, and it would compute the confidences with that answer.
随后,如果它的信心处于中间区域,它就会假设——而且我们在比赛中并未动态调整这一点,但本可以这么做,因为我们是在与顶级选手对弈——它就会假设这些家伙确实非常厉害。
Then, if its confidence was somewhere in the middle area, it was assuming – and we didn't dynamically adjust this during the game, but we could have, because we were playing the top champions – it was just assuming that these guys were really good.
情况是,它的信心处于较低水平。根据它在比赛中的表现,它在蜂鸣器响时会变得更加怯懦。
What happens is its confidence was on the low side. Depending on how it was doing in the game, it would be more timid on the buzzer.
原因是我旁边这家伙实在太厉害了。如果他气势汹汹、准备扑上去抢,那我就干脆让给他,不然我肯定抢不到,最后还是他得手。
The reason was because the guy next to me is really good. If he's really aggressive and is going to jump in and get it, I'm going to let him have it, because otherwise I'm going to lose it and he's going to get it.
如果他抓住了,那挺好。要是他没抓住,我就冒险一试。实际上,这玩意儿当时在运行一套置信度加权的抢答机制,模拟实验中的表现往往更胜一筹。
If he grabs it, fine. If he doesn't grab it, I'll take a chance. It actually was doing a confidence-weighted buzzer scheme, which in simulations actually tended to work better.
这种做法还带来了另一个效果,这也正是我们在权衡的另一个因素。那是一个次要目标:不要显得愚蠢。我们本可以采取一种截然不同的策略——对每一个问题都迅速回答——但那样我们会搞错很多事情。
It also had another effect, which is another thing that we were balancing. It was a secondary objective, which was not to look stupid. We could have taken a very different strategy, which was buzz on every single question, but we would have gotten a lot of things wrong.
这个门槛设在哪里,会改变你玩这场游戏的方式。我们一直在“看起来像个傻瓜”和“赢”之间走钢丝。我们其实有一个指标,叫“傻瓜指数”,是我们努力想要最小化的东西。
Where you set that threshold changed the nature of how you played the game. We were walking that line between looking stupid and winning. We actually had a metric, the "looking stupid metric," that we were trying to minimize.
迈克尔·莫布森闭幕致辞
Michael Mauboussin Closing Remarks
到此结束。
We're going to call it a day.
我来做一个非常简短的总结。以下是我对过去 24 小时的要点提炼。顺便说一句,感谢各位的坚韧坚持。我希望我的总结能对得起所有演讲者的精彩分享。以下是浓缩成 30 秒的精彩片段。
I'm going to do a very quick wrap-up. These are my CliffsNotes versions of the last 24 hours. Thank you all, by the way, for your hardiness. I hope I do all the speakers justice in this. Here are my 30-second versions of some of the highlights.
首先,回顾一下昨晚汤姆·西利(Tom Seeley)关于蜜蜂的研究。对我而言,一个关键启示在于,这些蜜蜂是多样化的个体。它们彼此独立、各不相同。它们不断揭示或浮现出各种选项。它们飞出去寻找大量可能性。然后,它们基本上通过投票来筛选出最佳选项,从而找到新家。
First, a recap from last night from Tom Seeley and the work on the bees. To me, one of the key takeaways for us is the fact that these bees are diverse agents. They're independent and different. They are revealing, or surfacing, options. They're going out and trying to find a lot of alternatives. And they're basically voting on those best options in order to find a home.
一个关键点在于他们有一个明确的目标函数。他们清楚自己理想中的家是什么样子。他们也拥有统一的目标和价值观。所有人都想找到最好的家。显然,他们彼此之间有血缘关系。
One of the essential points is that there is an objective function. They know what their ideal home looks like. They also have unified goals and values. They all want to find the best home. Obviously, they're genetically related to one another.
这些基本原则听起来都很有道理。但在真实的组织中,一旦引入激励机制和人事政治之类的因素,事情就会变得复杂得多。
The basic principles all make sense. In real organizations it gets a lot more complicated when we introduce things such as incentives and politics.
菲尔·泰特洛克的所有成果都很出色,但他有一张幻灯片让我觉得格外精彩。它涉及一个可以立即应用于组织内部的道理。他指出,他现在能得出比群体不加权平均结果好上 50% 到 70% 的结论。对此有四个促成因素。
All of Phil Tetlock’s stuff was great but he had one slide that I thought was extraordinary. And it addressed something that can be applied to organizations right away. He pointed out that he now can find results that are 50 to 70 percent better than the unweighted average of a group. There were four contributing factors to that.
一个是“人”的因素。这部分带来了约 10% 到 15% 的提升。这得归功于那种流体智力——那种应对新事物和保持开放心态的能力,当时人们确实心态开放。
One was the people. That was about a 10 to 15 percent boost. That was a function of this fluid intelligence, this ability to deal with novelty and open-mindedness, people were open-minded.
顺便说一句,作为认知智能的一种形式,流体智力在 20 岁出头达到顶峰,然后随着年龄增长逐渐下降。我们大多数人的流体智力都在减弱,尽管我们同时在积累晶体智力。
By the way, fluid intelligence as a form of cognitive intelligence peaks in your early 20s and then drifts lower through life. Most of us are losing our fluid intelligence even though we're picking up crystallized intelligence.
第二个是交互效应:这方面能带来 10% 到 20% 的提升。这就是“我们如何有效进行团队协作?”这显然是个非常有趣的话题。我们有一份材料,在其中稍微提到了这一点。
The second is the interaction effects: 10 to 20 percent boost from that. That's "How do we effectively work in teams?" That's obviously a very interesting topic. We have a piece out there where we wrote a little bit about this.
团队可以非常有帮助,但在某些方面也可能分散注意力。关键不在于“成为团队”,而在于“以建设性的方式组成团队”。
Teams can be very helpful, but they can also be distracting in some ways. It's not just being a team. It's being a team in a constructive way.
第三项是培训:来自培训的增益约为 10%。这涉及认知去偏见的理念。
Third was the training: about a 10 percent boost from training. That is this idea of cognitive de-biasing.
菲尔提到了一两点相关的想法。很多组织可能都可以尝试去做这件事。
Phil mentioned one or two things on that. That's something that many organizations can probably try to do.
最后,我认为非常吸引人的是这些极端化算法。这部分占比很大,达到 15% 到 30%。
Finally, which I thought was fascinating, were these extremizing algorithms. That was a big chunk of this, 15 to 30 percent.
让我来确认一下我理解得对不对。如果一个群体内部的多样性很高,那么你想把两头的极端差异拉得更大,才能做得更好。是这样吗?
Let me see if I get this right. If you have high diversity in the group, you want to increase the extremes to get better. Is that right?
我想引述一段理查兹·霍耶(Richards Heuer)的话,他写过一本《情报分析心理学》。霍耶说:“面对重大的范式转换时,对某一领域了解最深的分析师,需要清除的旧知识也最多。”
There's a quote I want to read from Richards Heuer, who wrote a book called The Psychology of Intelligence Analysis. Heuer says, “When faced with a major paradigm shift, analysts who know the most about a subject have the most to unlearn.”
我们都有一种天然倾向,习惯去依赖专家——某个特定行业、某个地区或其它领域的专家。但当世界发生改变时,这类人其实最为不利,因为他们必须先放下已有的认知,才能去学习新东西。
We all have a natural tendency to default to experts, an expert in this particular sector or this geography or what have you. But when the world changes, that person is actually the most disadvantaged because he or she has the most to unlearn before he can learn the new stuff.
迈克尔·莫布森闭幕致辞(续)
Michael Mauboussin Closing Remarks (Continued)
我认为这也是一个非常值得记住的有趣思考——我们本能地会去找最信任的人,但这些人可能恰恰处于最糟糕的处境之中。
I think that's a really interesting thought to also bear in mind – our natural inclination is for our go-to people, but they may be the person or the individuals in the worst situation.
至于卡罗琳·巴奇的演讲,值得深思的是这样一个事实:每分钟就有一名儿童死于疟疾。这令人震惊。当然,这在座各位看来似乎十分遥远,但意义重大。
As for Caroline Buckee’s talk, it is worth dwelling a moment on the fact that a child a minute dies from malaria. This is remarkable. It seems quite remote to us here, of course, but very important.
卡罗琳的演讲中我欣赏的一点是,我们拥有这些成熟的模型,比如 SIR 模型,它已经存在了一段时间,但现在可以借助新的数据源来提供信息。
What I loved about Caroline’s talk is that we have these established models, the SIR model for example, which has been around for some time but can now be informed with new sources of data.
卡罗琳和她的同事们正在做的事情极其令人振奋。他们正在改进现有方法,并催生出新的洞见。她还提出了一个观点——关于理论的作用——这个观点也得到肖恩·古尔利和戴夫·费鲁奇二人的呼应。
What Caroline and her colleagues are doing is extremely exciting. They are improving on current approaches and are generating new insights. She also made a point, which was echoed by Sean Gourley and Dave Ferrucci, on the role of theory.
几年前,我看到克里斯·安德森写的一篇文章,标题是“理论的终结”。我简直不敢相信。当我们谈论数据时,难道现在就可以放弃因果关系,只依靠相关性了吗?
I saw an article by Chris Anderson a few years ago called, “The End of Theory.” I couldn't believe it. When we think about data, can we now release causality and use only correlations?
对于某些领域,你也许能这么做,但在我们涉及的领域,因果关系是根本。所以你需要好的理论来理解大型数据集。话虽如此,我们明年午餐时间会供应卷薯条。
For some fields, you may be able to do that, but for the things that we're involved with, causality is essential. So you need good theory to understand large data sets. That said, we will be serving curly fries next year at the lunch time.
[laughter]
[laughter]
肖恩·古利谈到了自由式国际象棋。我反复追问的问题是:“我们能否明确界定,这些自由式国际象棋冠军所具备的,究竟是怎样一种技能?”
Sean Gourley talked about freestyle chess. The question I keep coming back to is, "Can we pin down what is this skill of these freestyle chess champions?"
什么时候该依赖电脑?什么时候该否决电脑?要具备哪些技能才能知道什么时候该否决电脑,这又是否需要我理解算法和编程?
When do I rely on my computer? When do I override my computer? Which skills are required to know when to override my computer, and does that require me to understand the algorithms and programming?
我可以说自己确实不懂编程和这些算法。我认为大多数组织都会把基本面分析和量化分析分开处理。有没有什么办法能让我们更有效地把这两者结合起来?
I can say I certainly am not versed in programming and these algorithms. I think most organizations separate fundamentals and quantitative stuff. Is there a way for us to be more effective at bringing those two things together?
然而,有一个非常开放的问题——对此抱持怀疑态度也有充分理由——但这仍是一个值得探讨的好问题。问题是:“我们能否利用技术,本质上扮演我们潜意识的角色?”这能帮助我们加速直觉和理解力的提升。
And there's this very open question – there's a reasonable reason to be skeptical about it – but it's a good open question. It's, "Can we use technology to essentially play the role of our subconscious?" That helps us accelerate our sense of intuition and understanding.
你刚听了 戴夫 的演讲,所以我没什么要补充的。我觉得最后那张幻灯片,确实又是对许多想法的一个总结:人类认知如何成为意义的来源。它是整个方程中不可或缺的一部分。数据的作用,又是如此根本,是理论的基础。然后是逻辑的应用,聚焦在因果关系上。
You just heard Dave's talk, so there's not much for me to add. I thought that last slide was really, again, just a recap of many of these ideas of how human cognition is really a source of meaning. It's essential to be part of this whole equation. The role of data, again, so essential as a foundation for theory. And then the application of logic, focusing on causality.
我还想补充一点——这或许有些牵强——但即便在他描述沃森(Watson)如何做出选择的过程中,实际上又把我们带回到了关于蜜蜂的那段讨论,也就是说,独立地挖掘出多种不同可能性,形成一套独立的投票机制,然后选择最优方案。从某些方面来看,我觉得刚才最后几句话让我们的讨论又回到了原点。
I will say this too – this may be a bit of a stretch – but even as he described how Watson makes its choices, it actually returned me back to the discussion about the bees, which is, again, independently unearthing these different possibilities, coming up with an independent voting system, and then going with the best alternative. In some ways I felt like the last couple of comments brought us full circle.
再次感谢各位从百忙之中抽出一天时间来参加股东会。希望大家玩得开心,也希望 2015 年还能见到大家再来一次。非常感谢。
I just want to thank you again for taking a day out of your valuable schedules to join us. Hope you had fun. And I’m hopeful we'll see you again in 2015 for another round. Thanks a lot.
[applause]
[applause]
托马斯·西利 康奈尔大学
Thomas Seeley Cornell University
生物学家兼作家托马斯·西利(Thomas Seeley)是康奈尔大学生物学霍勒斯·怀特教授,并担任神经生物学与行为学系主任。他的研究聚焦于动物(包括人类)群体中的集体智能,尤其是蜜蜂蜂群。
Thomas Seeley, biologist and writer, is the Horace White Professor in Biology at Cornell. He is the chair of the Department of Neurobiology and Behavior. His research focuses on collective intelligence in animal (including human) groups, especially honey bee colonies.
他在纽约州伊萨卡长大,高中时就开始养蜂——那时他用一个木匣子把一群蜜蜂带回了家。汤姆在达特茅斯学院获得化学学士学位,在哈佛大学获得生物学博士学位。在 1986 年一路打拼回到家乡伊萨卡/康奈尔之前,他于 1980 年加入耶鲁大学任教。
He grew up in Ithaca, New York and began keeping bees while a high school student, when he brought home a swarm of bees in a wooden box. Tom earned his AB in chemistry from Dartmouth College and his PhD in biology from Harvard University. Before working his way home to Ithaca/Cornell in 1986, he joined the faculty at Yale in 1980.
因在科学研究领域的成就,他获得了亚历山大·冯·洪堡杰出美国科学家奖、古根海姆研究基金,并当选为美国艺术与科学院院士。汤姆面向大众的首部科普著作《蜜蜂民主》于 2010 年 10 月由普林斯顿大学出版社在美国出版,并销往全球六个市场。
In recognition of his scientific work, he has received the Alexander von Humboldt Distinguished US Scientist Award, been awarded a Guggenheim Fellowship, and elected a Fellow of the American Academy of Arts and Sciences. Tom’s debut science book for the general public, Honeybee Democracy, was published by Princeton University Press (October 2010) in the US and in six markets worldwide.
注意:此处无可用记录。
Note: No transcript available.
丹·瓦格纳 Civis Analytics
Dan Wagner Civis Analytics
丹·瓦格纳(Dan Wagner)是 Civis Analytics 的首席执行官兼创始人,也是公司外部合作事务的负责人,代表公司在接洽客户业务机会时维护其至关重要的核心利益。
Dan Wagner is Chief Executive Officer and Founder of Civis Analytics and ambassador for the firm’s external partnerships, representing the company’s vital, core interests as it engages client business opportunities.
丹在 2012 年奥巴马竞选团队中担任首席分析官,负责管理一支由 54 名分析师、工程师和组织者组成的团队,为选民联络、数字营销、付费媒体、筹款和传播提供了分析支持与技术支撑。他的部门工作被赞为“彻底重塑了全国性竞选的运作方式”,并得到了《时代》杂志、《麻省理工科技评论》、《华尔街日报》、彭博社、《洛杉矶时报》和《哈珀》杂志的专题报道。
Dan served as the Chief Analytics Officer on the 2012 Obama campaign, overseeing a 54-person team of analysts, engineers and organizers that provided analytics and technologies for voter contact, digital, paid media, fundraising, and communication. His department’s work was credited with “reinventing how national campaigns are done” and has been highlighted in Time, MIT Technology Review, The Wall Street Journal, Bloomberg, the Los Angeles Times, and Harper’s Magazine.
丹的从政生涯始于 2008 年艾奥瓦州党团会议期间,担任巴拉克·奥巴马(Barack Obama)的副选民档案经理,这段经历定义了数据的相关性,以及客户对数据的理解。在 2012 年选举周期之前,他曾于 2010 年选举周期担任民主党全国委员会全国目标总监。丹此前曾在 FTI 咨询公司经济业务部门工作。他毕业于芝加哥大学,获得经济学与公共政策学位。
Dan began his work in politics as the Deputy Voterfile Manager for Barack Obama during the 2008 Iowa caucuses, which defined the relevance of data, and client comprehension of data. Prior to the 2012 cycle, he worked as the National Targeting Director for the Democratic National Committee during the 2010 election cycle. Dan has previously worked with the FTI Consulting Economics Practice. He graduated from the University of Chicago with a degree in Economics and Public Policy.
注意:此处暂无转录文本。
Note: No transcript available.
瑞士信贷致力于营造一个专业且包容的工作环境,让所有个体都能获得尊重与尊严。瑞士信贷是机会均等的雇主。
Credit Suisse is committed to a professional and inclusive work environment where all individuals are treated with respect and dignity. Credit Suisse is an equal opportunity employer.