2016思想领袖论坛:犯错如何教会你对错之道
2016 年 5 月 19 日,纽约州塔里敦,会议纪要
Proceedings May 19, 2016 Tarrytown, NY
2016
2016
Thought Leader
Thought Leader
Forum
Forum
目录
Table of Contents
迈克尔·莫布森,瑞士信贷主办方致介绍词与欢迎词
Michael Mauboussin, Credit Suisse Introduction and Welcome .................................................................................................................................3
比尔·格利,基准资本管理合伙人 《差之千里》…9
Bill Gurley, General Partner, Benchmark Capital How to Miss by a Mile.......................................................................................................................................9
华盛顿大学计算机科学教授佩德罗·多明戈斯,《终极算法》…………………………………………………………26
Pedro Domingos, Professor of Computer Science, University of Washington The Master Algorithm .....................................................................................................................................26
宾夕法尼亚大学运营、信息与决策学教授凯德·梅西谈算法厌恶 ………… 48
Cade Massey, Professor of Operations, Information and Decisions, University of Pennsylvania Algorithm Aversion ..........................................................................................................................................48
迈克尔·莫布森 导言
Michael Mauboussin Introduction
在投资与商业的领域中,信息和环境无时无刻不在变化。因此,我们必须不断审视自己的信念,审视这些信念是否真实反映世界,同时思考有哪些工具能帮我们优化决策。因为我们所处的是一个只能以一定概率走向成功的世界,所以我们必须从错误中学习。正因如此,2016 年思想领袖论坛的主题定为——“从错误中学会正确”。
Information and circumstances change constantly in the worlds of investing and business. As a consequence, we have to constantly think about what we believe, how well those beliefs reflect the world, and what tools we can use to sharpen our decisions. Because we operate in a world where we can succeed only with a certain probability, we have to learn from our mistakes. Hence, the theme for the Thought Leader Forum in 2016 was “What Being Wrong Can Teach You About Being Right.”
今年的论坛邀请了风险投资人、计算机科学家、专注于决策领域的经济学家以及一位顶级体育界高管。每位嘉宾都深入探讨了我们在思维与决策中可能偏离理想状态的某个方面。我们听到了关于假设如何深刻影响你对一家公司潜力的评估,以及原本善意的激励机制如何可能适得其反。论坛探讨了计算机如何通过机器学习成为一种新的知识来源,与演化、经验和文化形成互补。尽管通过计算机增强智能有潜在好处,我们也讨论了为何人类对算法存在抵触心理,以及如何克服这种心理。随后还有新旧阵营的问题:如何才能说服那些在旧体制下取得成功的部分人,接受更新、更好的做事方式?
This year’s forum featured a venture capitalist, a computer scientist, an economist who focuses on decisions, and a leading sports executive. Each explored an area of how our thinking and decisions can come up short of the ideal. We heard about how assumptions deeply shape how you assess a company’s potential and how well-intentioned incentive systems can go awry. There was an exploration of how computers, through machine learning, can serve as a new source of knowledge, complementing evolution, experience, and culture. Notwithstanding the potential benefits of augmenting our intelligence through computers, we discussed why we humans have an aversion to algorithms and how to overcome it. And then there is the issue of the old and new guard: how can we convince some who have been successful in an old regime to accept new and better ways of doing things?
“错误能教你什么才是正确”这一主题,在朴素实在论、人机对比以及变化的作用方面,都能给我们带来启发。朴素实在论指的是那种认为我们对世界的看法就是正确看法的感觉。但当现实摆在面前时,我们有必要重新审视自己的信念。
The theme of “what being wrong can teach you about being right” has lessons to teach us about naïve realism, man versus machine, and the role of change. Naïve realism is the sense that our view of the world is the correct one. But when confronted with reality, we need to revisit our beliefs.
例如,当我们面对与自己观点不同的人时,往往采取三种态度之一,以便固守己见。第一种,我们认为对方只是不了解事实,所以只要简单分享信息,就能让对方站到我们这边。第二种,我们相信即便掌握了事实,对方也缺乏足够的心智能力,无法像我们一样看清后果——这种人可以直接忽略。第三种,有些人明明和我们一样理解事实,却故意背弃我们所认定的真相——我们就把这些人归为恶人。
For example, when we face someone who has beliefs different than ours, we tend to adopt one of three attitudes so that we can perpetuate our position. First, we might assume the other person is merely unequipped with the facts, so simple sharing will swing them to our side. Next, we believe that even with the facts, the other person lacks the mental capacity to see the consequences as we do. We can write off those people. Finally, there may be people who understand the facts as we do but turn their backs on what we perceive to be the truth. We categorize those people as evil.
机器学习和人工智能再次成为热门词汇。谷歌 DeepMind 的 AlphaGo 程序便是一个标志,它在围棋这项棋盘游戏中击败人类冠军的时间比大多数专家预测的要早得多。问题在于我们如何将认知工作分配给机器和人类判断。如果你身处信息行业——而你正在读这篇文章,这种情况很可能属实——那么你必须仔细考虑如何整合计算机和人类。
Machine learning and artificial intelligence are again hot terms. Google DeepMind’s AlphaGo program, which beat a human champion in the board game of Go much sooner than most experts had predicted, is emblematic. The question is how we divide the cognitive work between machines and human judgment. If you are in the information business—and the chances are good this is true if you are reading this—then you must consider carefully how you might integrate computers and humans.
所有这些都意味着变化,而我们是极不情愿变化的。改变想法需要时间、精力和谦逊。当你在自己的领域取得过成功时,这一点尤为切中要害。体育中的策略是一个很好的类比。一些传统的做法,而且通常这些做法是有效的。但更细致的分析揭示出了一些明显优于传统智慧的策略。棒球中的防守移位就是一个例子。说服守旧派去改变——最终,我们所有人都会成为守旧派——是一个难以逾越的障碍。
All of this implies change, something we are loathe to do. Changing your mind takes time, effort, and humility. This is especially pertinent when you have been successful in your domain. Strategy in sports is a good analogy. There are traditional ways to do things, and often those ways are effective. But more careful analysis has revealed strategies that fly in the face of conventional wisdom that are clearly better. Defensive shifts in baseball are but one example. Convincing the old guard to change—and eventually, we are all part of the old guard—is a difficult hurdle.
以下演讲记录不仅记录了会议过程,还提供了深刻见解,告诉你如何提高从错误中学习的能力,并增加未来做出正确判断的胜算。比尔·格利指出,一些科技初创公司的高估值(所谓的“独角兽”)与低流动性之间的平衡难以维系。佩德罗·多明戈斯解释了计算机如何可能完成人类力所不能及的任务。凯德·马西表明,我们并不轻易接受算法,但有一种方法可以克服这种抵触情绪,从而改善决策。保罗·德波戴斯塔则提出,对变革的偏见更多与人类自身的思维方式有关,而非你所从事的特定活动。
The following transcripts not only document the proceedings, they also provide insights into how you can improve your own ability to learn from mistakes and improve your odds of being right in the future. Bill Gurley suggested that the high valuations for some technology startups (so-called “unicorns’) and the low level of liquidity is a balance that is not tenable. Pedro Domingos explained how computers might be able to complete tasks that are out of the grasp of humans. Cade Massey showed that we don’t readily embrace algorithms but that there is a way to overcome this aversion and improve decisions. And Paul DePodesta suggested that the bias against change has less to do with the game you are playing and more to do with how we humans think.
下面即是对所给段落的译文,严格按照逐段全文翻译、不增减、不解释、保留原文语气与格式的要求完成,单独输出该段落的译文。
迈克尔·莫布森 瑞士信贷
Michael Mauboussin Credit Suisse
迈克尔·莫布森是瑞信集团全球市场部门的董事总经理,常驻纽约。他担任全球金融策略主管,凭借在估值与投资组合配置、资本市场理论、竞争战略分析及决策制定等领域积累的专业知识、研究成果与著述,为外部客户及瑞信内部专业人士提供思想领导力与策略指导。
Michael Mauboussin is a Managing Director of Credit Suisse in the Global Markets division, based in New York. He is the Head of Global Financial Strategies, providing thought leadership and strategy guidance to external clients and internally to Credit Suisse professionals based on his expertise, research, and writing in the areas of valuation and portfolio positioning, capital markets theory, competitive strategy analysis, and decision making.
在 2013 年重返瑞士信贷之前,他曾担任美盛资本管理公司的首席投资策略师。
Prior to rejoining Credit Suisse in 2013, he was Chief Investment Strategist at Legg Mason Capital Management.
迈克尔最初于 1992 年以包装食品行业分析师的身份加入瑞士信贷,后被任命为美国首席分析师。
Michael originally joined Credit Suisse in 1992 as a packaged food industry analyst and was named Chief U.S.
1999 年的投资策略师。他曾任纽约消费品分析师协会主席,并多次入选《机构投资者》全美研究团队以及《华尔街日报》食品行业全明星调查。
Investment Strategist in 1999. He is a former president of the Consumer Analyst Group of New York and was repeatedly named to Institutional Investor’s All-America Research Team and The Wall Street Journal All-Star survey in the food industry group.
迈克尔是《成功方程式:解构商业、体育与投资中的技能与运气》《三思而后行:驾驭反直觉的力量》以及《出乎意料:在非常规之处发现金融智慧》三本书的作者。他还与阿尔弗雷德·拉帕波特合著了《预期投资:通过解读股价获取更高回报》。
Michael is the author of The Success Equation: Untangling Skill and Luck in Business, Sports, and Investing, Think Twice: Harnessing the Power of Counterintuition, and More Than You Know: Finding Financial Wisdom in Unconventional Places. He is also co-author, with Alfred Rappaport, of Expectations Investing: Reading Stock Prices for Better Returns.
迈克尔自 1993 年起担任哥伦比亚商学院金融学兼职教授,同时也是海尔布伦格雷厄姆与多德投资中心(Heilbrunn Center for Graham and Dodd Investing)的教员。他还担任圣塔菲研究所(Santa Fe Institute)董事会主席,该所是复杂系统理论跨学科研究的顶尖机构。迈克尔在乔治城大学获得文学士学位。
Michael has been an adjunct professor of finance at Columbia Business School since 1993 and is on the faculty of the Heilbrunn Center for Graham and Dodd Investing. He is also chairman of the board of trustees of the Santa Fe Institute, a leading center for multi-disciplinary research in complex systems theory. Michael earned an AB from Georgetown University.
迈克尔·莫布森,瑞士信贷
Michael Mauboussin Credit Suisse
早上好。在座各位中还没见过面的,我叫迈克尔·莫布森,是瑞士信贷全球金融策略主管。我代表瑞士信贷的所有同事,热烈欢迎大家参加 2016 年思想领袖论坛。昨晚已经加入我们的朋友,希望你们度过了愉快的夜晚。我们对今天的议程非常期待。
Good morning. For those of you whom I haven’t met, my name is Michael Mauboussin, and I am head of Global Financial Strategies at Credit Suisse. On behalf of all of my colleagues at Credit Suisse, I want to wish you a warm welcome to the 2016 Thought Leader Forum. For those who joined us last night, I hope you had a wonderful evening. We are very excited about our lineup for today.
今天上午,在把麦克风交给演讲者之前,我想先做两件事。第一,我想提一下,你们可以这样理解今天的讨论主题——犯错如何帮你走向正确。第二,我想谈谈这个论坛本身,包括你们可以做些什么来让它办得更成功。
I’d like to do a couple of things this morning before I hand it off to our speakers. First I want to highlight the levels at which you might consider today’s discussion about the idea of how being wrong can inform you about being right. I then want to discuss the forum itself, including what you can do to contribute to its success.
今天的讨论,你可以从三个不同层次来听。有些观点会跨越多个层次,但这些都是我们今天全天将会听到的一些思路。
You might listen to today’s discussion at three different levels. Some of the points will span multiple levels, but these are some of the ideas that we’ll hear about throughout the day.
第一点关乎朴素的实在论(naïve realism)这个概念。在心理学中,这是人类的一种倾向,即相信自己客观地看到了周遭的世界,而那些与我们意见相左的人,必定是信息不足、不理性或心存偏见。
The first relates to the ideas of naïve realism. In psychology, this is the human tendency to believe that we see the world around us objectively and that people who disagree with us must be uninformed, irrational, or biased.
第二个是人与机器的较量。这是一个无处不在的主题。算法擅长什么,人类又擅长什么?我们如何利用算法来提升我们的表现?为什么在许多场景下,我们很难信赖算法?
The second is man versus machine. This is a theme that is popping up everywhere. What are algorithms good at and what are humans good at? How do we use algorithms to augment our performance? Why do we struggle to defer to algorithms in many settings?
最后一点是变革问题。机构惯性是许多公司面临的巨大难题。企业如何才能跟上时代?我们如何整合新信息?变革的心理机制是什么?
The final is the issue of change. Organizational inertia is a huge issue in many firms. How can firms keep up? How do we integrate new information? What is the psychology of change?
先说说“天真实在论”。我特别喜欢这么一幅漫画:画面里两支军队正剑拔弩张,准备开战。图上的引文是:“除非他们放弃他们的兔子神,皈依我们的鸭子神,否则和平绝无可能。” 而更妙的是,这两支敌对军队的旗帜居然一模一样。这幅漫画借用了“兔子-鸭子错觉”——那幅既可以看成兔子、也可以看成鸭子的双关图。
Let’s start with naïve realism. Here’s a cartoon I love: as you can see, there are two armies preparing to square off, and the quote is: “There can be no peace until they renounce their Rabbit God and accept our Duck God.” The picture shows that the flags of the competing armies are the exact same. This is based on the rabbit-duck illusion, an ambiguous picture that can be interpreted either as a rabbit or a duck.
心理学中的“朴素实在论”认为,我们都以为自己对世界持有一种客观真相。于是,我们很难接受别人有不同的观点。所以我们每个人都带着自己认为正确的信念四处行走——否则我们也不会紧握这些信念。而当这些信念与现实世界正面碰撞时,事情就开始变得有意思了。
The idea of naïve realism in psychology is that we all think that we have an objective reality of the world. As a consequence, we have a hard time accepting that others have different points of view. So we all walk around with beliefs that we think are true. Otherwise we wouldn’t hold onto those beliefs. Things become interesting when those beliefs confront the world.
有个众所周知的实验能说明这一点。一位名叫伊丽莎白·纽顿的心理学家设计了一个实验,分为“敲击者”和“听者”两组。敲击者拿到一份包含 25 首知名歌曲的清单,比如《祝你生日快乐》,然后被要求用手指在桌子上敲出这些歌曲的节奏。听者的任务则是根据敲击的节奏来辨认是哪一首歌。
Here’s a well-known experiment that demonstrates this point. A psychologist named Elizabeth Newton set up an experiment whereby there were “tappers” and “listeners.” The tappers were given a list of 25 well-known songs, such as “Happy Birthday to You,” and were asked to tap the rhythm of the song on the table. The task of the listener was to identify the song based on the taps.
她一共进行了 125 次试验。听众只能识别出其中 3 首歌,成功率大约 2.5%。但当研究人员问敲击者,他们认为听众能正确识别出多少比例时,答案竟是 50%!这跟知识的诅咒有关,它也是沟通中的巨大障碍。同样,我们很难理解别人看世界的方式与我们不同。
She ran 125 trials of this. The listeners were able to identify only 3 of the songs, a success rate of about 2.5 percent. But when the researchers asked the tappers what percent they thought the listeners would be able to identify correctly, the answer was 50 percent! This is related to the curse of knowledge, which is also a huge impediment to communication. Again, we struggle to understand that others don’t see the world as we do.
所以,如果你用一种方式看待世界,而别人用另一种方式,你就必须调和这些观点。而我们这么做的时候,通常倾向于预设三种情况之一。第一种是对方只是不知道你所掌握的事实,因此是无知的。解决办法很简单,告诉他们事实,这样他们就能明白你的观点。第二种是对方知道事实,但实在太蠢,没法正确理解。最后一种假设是,人们知道事实也能理解,但就是故意对真相视而不见。
So if you see the world one way and others see it a different way, you have to reconcile the views. And as we do so, we tend to assume one of three things. The first is that the other person simply doesn’t know the facts that you do, and hence is ignorant. The answer is simply to inform them so that they will then see your point of view. The second is that the person has the facts, but they are just too stupid to understand them properly. The last assumption is that people know the facts and can comprehend them, but they just turn their backs on the truth.
不信宗教的人就是一个例子。
Unbelievers in religion are an example.
迈克尔·莫布森 瑞士信贷
Michael Mauboussin Credit Suisse
现在想想你是怎么评价那些和你不一致的人的——你是不是也会用这些假设来调和他们的看法,好让自己的观点站得住脚?
Now consider how you assess people who don’t agree with you. Do you evoke one of these assumptions to reconcile their beliefs with yours?
现在我们来谈谈一个将贯穿全天的主题。我把它称为人机对决,但也许更准确的说法是人类 vs 算法。我想提出的第一点,与我所说的“专家挤压”有关。
We now turn to a theme that will spread through the day. I am calling it man versus machine but it may be just as accurate to say humans versus algorithms. The first point I want to make refers to what I call “the expert squeeze.”
可以把这个问题看作一个连续谱。在谱系的一端,是那些基于规则且一致的问题。在这方面,专家通常很熟练,但计算机更快、更便宜、也更可靠。当然,在今天,你不得不提到 AlphaGo 的成功——谷歌 DeepMind 的程序击败了围棋世界冠军。
The way to think of it is as a continuum. On one side there are problems that are rules-based and consistent. Here, experts are often proficient but computers are quicker, cheaper, and more reliable. Today, of course, you have to point to the success of AlphaGo—Google DeepMind’s program that beat a champion in Go.
在连续谱的另一端,是那些概率性的问题,并且处于不断变化的领域中。这里的证据表明,在某些条件下,集体表现得比专家更好。确保这些条件到位,对决策者来说至关重要。
At the other side of the continuum are problems that are probabilistic and in domains that change constantly. Here, the evidence shows that collectives do better than experts under certain conditions. Making sure those conditions are in place is crucial for a decision maker.
我现在要稍微抢一下第二位演讲者的风头,介绍几种机器学习的方法。但我强调的重点有所不同。如果你的组织依赖基础研究,这些方法中的任何一种看起来眼熟吗?
I’m now going to steal a bit of thunder from our second speaker and introduce various approaches to machine learning. But my point of emphasis is somewhat different. If your organization relies on fundamental research, do any of these approaches seem familiar?
例如,很多投资者喜欢诉诸类比:这个投资就像过去的那个投资。那么有趣的问题就变成了:作为基本面分析师,我们能从机器学习的发展中学到什么?下一步是考虑如何将机器学习技术整合到决策过程中。如果你依赖量化方法,你如何看待内嵌在算法中的偏差?
For example, lots of investors like to appeal to analogies: this investment is like that investment from the past. The interesting question then becomes: what can we, as fundamental analysts, learn from what’s going on in machine learning? The next step is considering how we can integrate machine learning techniques into a decision-making process. If you are relying on quantitative methods, how do you think about the biases built into the algorithms?
我要谈的最后一个关于人与机器的议题是,我们人类往往不太愿意让自己的命运由算法决定,即使有大量证据表明算法比人类更优秀。
The final issue I’ll mention for man versus machine is that we as humans tend to be uneasy letting our fate be decided by an algorithm, even if there’s abundant evidence that the algorithm is better than a human.
《点球成金》里的这个场景抓住了那种情绪:老派的人很难从统计分析中理解信号。这有几个原因。他们从自己的经验中归纳总结。他们过分强调近期表现。他们依赖他们看到的东西,而非因果关系。我们今天会讨论如何克服算法厌恶,但这确实是一个巨大的问题。
This scene from Moneyball captures the tone: the old timers have a difficult time grasping the signal from the statistical analysis. This is true for a few reasons. They generalize from their own experience. They overemphasize recent performance. And they rely on what they see versus cause and effect. We’ll talk today about how to overcome algorithm aversion, but it’s a huge issue.
最后一个话题是关于变革,这很困难。第一个障碍是组织惯性。回到过去,我是一名食品行业分析师,我记得一个故事,很好地说明了这一点。
The final topic is that of change, which is hard. The first impediment is organizational inertia. Back in the day, I was a food industry analyst, and I recall a story that captured this well.
大约 25 年前,大卫·约翰逊接手担任金宝汤公司的首席执行官时,公司的业绩落后于同行。于是他进行了全面审查,以了解如何改善运营。
When David Johnson took over as CEO of Campbell Soup about 25 years ago, the performance of the company lagged its peers. So he did a full review to understand how to improve operations.
他注意到,公司每年秋季都会对番茄汤进行一次大规模促销。番茄汤是它们规模最大、利润最高的产品之一。当他问一位高管为什么要这样做时,那位高管回答说:“我不知道,我们一直这么做。”
He noticed that the firm did a huge annual promotion of tomato soup in the fall every year. Tomato soup was one of their largest and most profitable products. When he asked the executive why they did it, the executive responded, “I don’t know, we’ve always done it.”
在第一次世界大战期间,金宝汤的策略是自己种植西红柿,收获后加工成罐装汤。由于库存增加,而汤的销售旺季还要等好几个月,金宝汤便通过促销来清空库存。
In World War I, Campbell’s strategy was to grow its own tomatoes, harvest them, and them convert them to canned soup. With inventory up and the soup season still months ahead, Campbell used a promotion to clear its inventory.
但当然,公司很久以前就转向了全年供应的供应商,消除了收获后供应激增的情况。这引出了彼得·德鲁克的一句名言:“如果我们还没有做这件事,以我们现在的了解,我们还会进入这个领域吗?”
But of course the company long ago went to year-round suppliers, eliminating the post-harvest spike in supply. This evokes a quote from Peter Drucker: “If we did not do this already, would we go into it” now, knowing what we now know?
也许最具挑战性的事情是在收到新信息时更新你的信念。
Perhaps the most challenging thing to do is to update your beliefs when you receive new information.
这是丹尼尔·卡尼曼《思考,快与慢》中一个著名的例子 [第 166 页]。
Here’s a famous example from Thinking, Fast and Slow by Daniel Kahneman [page 166].
迈克尔·莫布森 瑞信
Michael Mauboussin Credit Suisse
“某城市夜里发生了一起肇事逃逸事故。该城市有两家出租车公司,绿色和蓝色。给你如下数据:
“A cab was involved in a hit-and-run accident at night. Two cab companies, the Green and the Blue, operate in the city. You are given the following data:
该城市 85% 的出租车是绿色的,15% 是蓝色的。
85% of the cabs in the City are Green and 15% are Blue.
一位目击者指认出租车是蓝色的。法庭测试了目击者在事故当晚情况下的可靠性,并得出结论,目击者正确识别两种颜色的概率是 80%,识别错误的概率是 20%。
A witness identified the cab as Blue. The court tested the reliability of the witness under the circumstances that existed on the night of the accident and concluded that the witness correctly identified each one of the two colors 80% of the time and failed 20% of the time.
请问,事故中涉及的出租车是蓝色而非绿色的概率是多少?”
What is the probability that the cab involved in the accident was Blue rather than Green?”
最常见的答案是 80%,基于目击者的可靠性。但正确答案仅略高于 41%。在菲尔·泰特洛克那本出色的书《超预测》中,他有一句很棒的话:“信念是待检验的假设,而非待守护的宝藏。”这话说起来很容易,但在实践中却非常困难。改变我们的想法需要时间、精力,有时还需要专业技能,并且可能令人难堪。我们大多数人宁愿继续相信自己相信的东西。
The most common response is 80 percent, based on the reliability of the witness. But the correct answer is just a little over 41percent. In Phil Tetlock’s terrific book, Superforecasting, he has a great line: “Beliefs are hypotheses to be tested, not treasures to be guarded.” This is really easy to say and very difficult to do in practice. Changing our minds takes time, effort, in some cases technical skills, and can be embarrassing. Most of us would prefer to keep believing what we believe.
我最后的想法是关于损失厌恶。在座的每位都要处理那些只有一定概率才会成功的决策。
My final thought is on loss aversion. Everyone in this room deals with decisions that work only with some probability.
我们承受损失的痛苦胜过获得同等收益的快乐。所以我们倾向于坚持传统的做事方式,因为如果我们失败了,会有很多同伴。
We suffer losses more than we enjoy comparable gains. So we tend to stick to conventional ways of doing things because if we fail, we have lots of company.
体育领域有很多这样的例子。一个例子是在橄榄球比赛中决定在第四次进攻中强攻。大多数教练更喜欢更保守的路线,即使这会降低他们获胜的概率,因为在第四次进攻中被拦截的潜在痛苦比获得新一次进攻机会的收益要糟糕得多。
There are lots of instances of this in sports. One example is the decision to go for it on fourth down in football. Most coaches prefer the more conservative route even if it gives them a lower probability of winning, because the potential pain of getting stopped on fourth down is a lot worse than the upside of a fresh set of downs.
在谈论本次论坛的目标之前,我想提一下 Ink Factory 才华横溢的团队成员。
Before I speak about the goal of the forum, I want to mention the talented folks from Ink Factory.
达斯迪和瑞安今天会为所有演讲者做图形记录。这意味着他们会将演讲者的话语综合成图像和文字,以捕捉关键概念。他们的口号是:“你讲,我们画,效果超棒。”我们觉得你们会同意的。请随意拍摄这些作品并在推特上分享图片。
Dusty and Ryan will be graphically recording all of our presenters today. This means they will be synthesizing the words of our speakers into images and text to capture the key concepts. Their slogan is “you talk. we draw. it's awesome.” And we think you will agree. Please feel free to take pictures of the artwork and to tweet the images.
我们鼓励你们向他们提问——当然是在他们画完之后!
And we encourage you to ask them questions—after they are done drawing of course!
最后,我想强调一下我们今天的几个目标。首先,我们希望为你们提供接触一些演讲者的机会,这些人你们在日常工作中可能遇不到,但他们能够激发思考和对话。其次,我们希望鼓励自由的思想交流。请注意,我们的演讲时段比通常要长。这很大程度上是因为我们想留出时间进行互动。
Let me end by highlighting what our goals are for the day. First, we want to provide you access to speakers whom you may not encounter in your day-to-day interactions but who are nonetheless capable of provoking thought and dialogue. Second, we want to encourage a free exchange of ideas. Note that our speaking slots are longer than normal. This is in large part because we want to leave time for back-and-forth.
第二,我们特意称之为“论坛”而非“会议”,正是出于这个原因。我们希望营造一个鼓励探究、挑战和交流的环境。
Second, we purposefully call this a “forum” instead of a “conference” precisely for this reason. We want to encourage an environment of inquiry, challenge, and exchange.
最后,我们希望这对你们来说是一次精彩的体验,所以请随时向瑞信团队的任何人提出任何需求。我们会尽力满足你们。
Finally, we want this to be a wonderful experience for you, so please don’t hesitate to ask anyone on the Credit Suisse team for anything. We will do our best to accommodate you.
比尔·格利 Benchmark Capital
Bill Gurley Benchmark Capital
比尔·格利在 Benchmark Capital 担任普通合伙人已超过 10 年。在加入 Benchmark 之前,比尔是 Hummer Winblad Venture Partners 的合伙人。
Bill Gurley has spent over 10 years as a General Partner at Benchmark Capital. Prior to Benchmark, Bill was a partner with Hummer Winblad Venture Partners.
进入风险投资行业之前,比尔在华尔街做了四年顶级分析师,包括在瑞信第一波士顿的三年,专注于个人电脑硬件和软件。他的研究覆盖了戴尔、康柏和微软等公司,并且是亚马逊公司 IPO 的首席分析师。在 1995 年和 1996 年,比尔都是《机构投资者》全美研究团队的成员。
Before entering the venture capital business, Bill spent four years on Wall Street as a top-ranked research analyst, including three years at CS First Boston focusing on personal computer hardware and software. His research coverage included such companies as Dell, Compaq, and Microsoft, and he was the lead analyst on the Amazon.com IPO. In both 1995 and 1996, Bill was a member of Institutional Investor’s All-America Research Team.
在他的投资生涯之前,比尔是康柏计算机公司的设计工程师,曾参与 486/50 和康柏第一台多处理器服务器等产品的开发。过去十五年,比尔撰写了“Above the Crowd”博客,专注于高科技企业的演变和经济学。
Prior to his investment career, Bill was a design engineer at Compaq Computer, where he worked on products such as the 486/50 and Compaq’s first multi-processor server. For the past fifteen years, Bill has authored the “Above the Crowd” blog, which focuses on the evolution and economics of high technology businesses.
比尔是德克萨斯大学麦库姆斯商学院的顾问委员会成员,也是 KIPP 湾区学校的董事会成员。他于 1993 年获得德克萨斯大学工商管理硕士学位,并于 1989 年获得佛罗里达大学计算机科学学士学位。
Bill is on the advisory board of the McCombs School of Business at the University of Texas and a board member at KIPP Bay Area Schools. He received his MBA from the University of Texas in 1993 and a BS in computer science from the University of Florida in 1989.
比尔·格利 Benchmark Capital
Bill Gurley Benchmark Capital
迈克尔·莫布森:我非常高兴地介绍我们今早的第一位演讲者,我的好朋友兼前同事,比尔·格利。比尔是 Benchmark Capital 的普通合伙人, Benchmark Capital 是全球领先的风险投资公司之一。比尔结合了许多我推崇的品质:他极其聪明;求知欲旺盛;是出色的战略思考者;精通财务;并且他不把自己太当回事。你可以在 abovethecrowd.com 找到他的文章,我建议你阅读他写的每一篇文章。
Michael Mauboussin: I’m very pleased to introduce our first speaker this morning, my good friend and former colleague, Bill Gurley. Bill is a general partner at Benchmark Capital, one of the world’s leading venture capital firms. Bill combines a lot of attributes that I admire: he’s wicked smart; intellectually curious; a terrific strategic thinker; financially savvy; and he doesn’t take himself too seriously. You can find his writings at abovethecrowd.com, and I recommend reading everything he writes.
事实上,比尔的一篇文章深刻地启发了今天论坛的主题。故事是这样的:比尔投资的公司之一是 Uber。纽约大学的教授、估值专家阿斯沃斯·达摩达兰认为,这家公司的价值不应超过一个相当适度的数字,因为它是出租车和豪华轿车的替代品。
In fact, one of Bill’s essays deeply inspired the theme for today’s forum. Here’s the story: one of Bill’s portfolio companies is Uber. Aswath Damodaran, a professor at NYU (New York University) and valuation expert, suggested that the company’s value should not exceed a fairly modest amount because it was a substitute for cabs and limos.
比尔的回应——整个过程非常文明——标题是“How to Miss by a Mile”。比尔重新定义了 Uber 的可寻址市场总量(TAM),并暗示达摩达兰可能低估了 25 倍。
Bill’s response—and this was all very civil—was called “How to Miss by a Mile.” Bill reframed the TAM (Total Addressable Market) for Uber and suggested that Damodaran could be off by a factor of 25 times.
在进入风险投资界之前,比尔是一名华尔街分析师。我有幸在他刚从商学院毕业时就与他共事,他在很短的时间内就确立了自己作为顶尖分析师的地位。
Prior to joining the venture capital world, Bill was a Wall Street analyst. I had the fortune of working with him in his early days out of business school, where he established himself, in very short order, as a go-to analyst.
请大家和我一起欢迎比尔·格利。
Please join me in welcoming Bill Gurley.
比尔·格利:在开始之前,我想感谢迈克尔,不仅是因为他邀请我来这里,更是因为 23 年前当我开始我的投资生涯时,他给予了我指导,并且在过去的 23 年里,他一直是思想领袖。如果当时没有遇到迈克尔,我可能不会有今天的成就。他是食品行业分析师,而我是个人电脑行业分析师,不知怎么地,他找到了一种方式来塑造我所做的一切,并对我的职业生涯产生了巨大的影响。所以,谢谢你,迈克尔。
Bill Gurley: Before I get started, I want to thank Michael not just for having me here but for being there for me 23 years ago when I started my investment career, and for being a thought leader for the past 23 years. I would not be where I am today probably if I hadn’t met Michael back then. He was the food analyst and I was the PC (personal computer) analyst, and somehow he found a way to shape everything I did and have a huge impact on my career. So, thank you, Michael.
当我第一次见到迈克尔时,我们都爱上了这本书,《复杂性》,作者是米切尔·瓦尔德罗普,讲述了圣塔菲研究所的崛起,我知道在座的有些人曾在那里待过很久。我告诉人们,这本书对我的思维方式影响超过了任何我读过的书,因为我们作为投资者处理的大多数事情都是复杂系统,我最后会回到这一点,但这就是我为什么想在这里指出来。
When I first met Michael, we both fell in love with this book, Complexity, by Mitchell Waldrop, about the rise of the Santa Fe Institute, and I know some of the people in the room have spent a lot of time out there. I tell people this book has affected how I think more than any other book that I’ve ever read because most of the things we deal with as investors are complex systems, and I’m going to come back to this at the end, but that’s why I wanted to point it out here.
现在,这是迈克尔提到的那位教授。两年前我并不知道他是谁。原来他是一位相当有思想的估值教授,这是他在 Uber 上发表的博客文章,而我是 Uber 的董事会成员。他不仅发表了这篇长文,还为纳特·西尔弗的 FiveThirtyEight 网站做了一个摘要。
Now here’s the professor Michael mentioned. I didn’t know who he was two and a half years ago. It turns out that he’s a rather thoughtful valuation professor, and here’s the blog post that he published on Uber where I sit on the board. Not only did he publish this long post, but he did a summary for Nate Silver’s FiveThirtyEight.
这篇文章的标题是“优步不值 170 亿美元”,他花了不少功夫去计算,认为其价值应该在 50 亿美元左右。依我看,他在思考中犯了好几处关键性错误,迈克尔今天请我来谈谈这些。他说:“你不如讲讲‘如何错得离谱’?”他希望我结合三四个我亲身见过、这些年也聊过的事例。
The title of this one was “Uber Isn’t Worth $17 Billion,” and he did a lot of work to calculate that he thought the value should be around $5 billion. He made a number of critical errors in his thinking from my point of view, and Michael asked me to talk about them today. He said, “Why don’t you talk about ‘How to Miss by a Mile’?” And he wanted me to incorporate three or four different things that I’ve seen and that we’ve talked about over the years.
这只是第一个例子,但他犯的错误我认为很有意思,值得提出来。第一个错误是他没有考虑到网络效应的存在。这张餐巾纸上的示意图其实是大卫·萨克斯画的,他是 Uber 的天使投资人之一。这张图简单明了地展示了:更多的需求会吸引更多的司机。
So this is just the first one, but he made a number of errors that I think are interesting and I want to highlight. The first one was he didn’t theorize that there might be a network effect. This napkin drawing was actually done by David Sacks, one of the angel investors in Uber. It basically shows that more demand drives more drivers.
更多司机带来更广的地域覆盖。这又缩短等待时间,从而可以降低价格,因为你实际上提高了利用率,就像飞机一样。这还让接单速度更快。更低的价格和更快的接单反过来刺激更多需求,于是就形成了这样一个循环。他连这一点都没考虑到,但我觉得这甚至还不是他犯下的最大错误。
More drivers creates more geographic coverage. That leads to less downtime which can lead to lower prices because you in effect have more utilization, like an airplane. It also leads to faster pickups. Lower prices and faster pickups drives more demand, and so you have a circle here. He didn’t even consider this, but I don’t even think that was his biggest mistake.
他最大的错误与 TAM(总可用市场)有关。他假设优步当时攻击的市场是出租车和豪华轿车。基本上,他看了当时出租车和豪华轿车的收入,并假设优步能从中分得一杯羹。这是他两年半前写下的,而为了让你看看这个判断有多离谱——
His biggest mistake related to TAM, total available market. He assumed the market that Uber was attacking was taxis and black cars. Basically he looked at the revenue for taxis and black cars at the time, and he assumed that Uber could get some fraction of that. He wrote this two and a half years ago, and just to show you how far off
比尔·格利基准资本(Benchmark Capital)
Bill Gurley Benchmark Capital
他当时已经基于这一假设:2011 年旧金山出租车和黑色专车市场的规模是 1.2 亿美元,而如今旧金山网约车市场的规模是 12 亿美元,相差 10 倍,这完全超出了他的分析范畴。这个市场仍在增长,而我们只覆盖了 13% 的人口,所以光是他对总体可及市场(TAM)的假设,实际误差可能就超过 10 倍。
he was on that assumption already, the market for taxis and black cars in San Francisco in 2011 was $120 million, and the market for ridesharing in San Francisco right now is $1.2 billion, a 10 times differential and totally outside of his scope of analysis. It’s still growing, and we’ve only touched 13 percent of the population, so it’s likely going to be off by more than 10 times, just his TAM assumption.
那这是怎么回事?他错过了什么?他根本没想过这些事。他没想过,约车等待时间比叫出租车快得多——如果你在休斯顿这样的城市叫过出租车,得等 30 分钟才到。也没想过覆盖密度更大。优步在那些出租车历来不服务的区域,密度已经更高了。
So how does this happen? What did he miss? He didn’t think about any of these things. He didn’t think about the fact that the pickup times were much quicker than a taxi. If you’ve ever ordered a taxi in a city like Houston, it takes 30 minutes for it to show up. Greater coverage density. Uber already has more density in areas where taxis have never historically served.
更便捷的支付,更高的文明程度。由于双向评分机制以及更高的信任与安全水平,这实际上是一种更好的体验。如今,从信任角度来看,人们普遍认为这种体验远优于出租车。而且事实上,网络上有很多博客文章,讲述人们通过这个系统找回了手机、钥匙和钱。
Easier payment, higher civility. It’s actually a better experience because of the dual rating and higher trust and safety. People routinely now rate this experience as way better than taxis from a trust standpoint and, in fact, there’s numerous blog posts that you can find on the web of people getting their phones back, their keys back, their money back all through the system.
作为一名金融学教授,他犯了个更大的错误——压根没考虑价格弹性。这玩意儿我记得在初级微观经济学课上就教过。他写那篇文章的时候,优步在很多城市的定价已经只有出租车的一半左右了,这些价格都是公开的。他本来可以查到的,但没查。他根本没去研究。
He made an even bigger mistake for a financial professor which is he didn’t think about price elasticity. Now this is taught I think in early microeconomics courses. At the time he wrote it, Uber was already about half of taxi prices in many cities, and these prices were published. So he could have found them, but he didn’t. He didn’t look at it.
他也没有考虑新的使用场景、扩大地理覆盖范围、租车替代方案,或是情侣之夜。
He also didn’t consider new use cases, expanded geographic coverage, a rental car alternative, couples night
优步的高峰期在周五和周六晚上,已经成了防止酒后驾车的巨大屏障。在我住的帕罗奥图附近,门洛帕克的一对夫妻会去帕罗奥图吃饭。他们以前开车去,但现在不开了。他们直接打优步。不用找车位,也不用担心喝酒的问题。
out. Uber peaks on Friday and Saturday night and has become a huge DUI deterrent. Where I live near Palo Alto, a couple in Menlo Park will go out to eat in Palo Alto. They used to drive, but they don’t anymore. They just take Uber. They don’t have to park. They don’t have to worry about drinking.
运输儿童、老年人以及作为公共交通补充——这些都是他没有考虑到的因素。UberPOOL(拼车服务)如今大约占乘车量的 20% 到 25%,正在和公交车、地铁竞争。
Transporting kids, seniors, supplement of mass transit are all other things that he didn’t consider. UberPOOL, where you share rides in a car, now comprises about 20, 25 percent of rides and is competing with bus and subway.
然后,一个非常重大的替代方案是作为汽车拥有权的替代品,这彻底改变了视角。那位教授没有考虑租车替代方案。赫兹的 CEO 也没有考虑它,因为他一直说优步是“出租车替代品”——这是他用的说法。这是赫兹过去 12 个月的股价。
Then, a really big one is as an alternative for car ownership, which changes the perspective entirely. The professor didn’t think about car rental alternatives. The CEO of Hertz didn’t think about it either because he constantly says that Uber’s a “taxi alternative,” a phrase he uses. This is Hertz’s 12-month stock price.
[播放 CNBC 上吉姆·克莱默的视频片段]:为什么赫兹就是不肯向优步认输呢?我是说,他们又一次让人失望了。他们又一次说,这个季度的形势比上个季度更糟。他们又一次在否认优步是问题所在,但实际上是矢口否认,不过我觉得我们每个人手机里都有优步。
[Shows video clip of Jim Cramer on CNBC]: Why doesn’t Hertz just own up to Uber? I mean again they disappointed. Again they said that things had gotten worse this quarter than last quarter. Again they denied without really saying it’s denial that Uber’s the problem, but I think we all have Uber in our cell phone.
优步是我所说的声望品牌之一。我提到化妆品,是因为如今你出门必须打扮得体,穿戴整齐。我说 iPhone 是声望品牌,优步也是声望品牌。
Uber is one of the prestige brands that I talk about. I talk about cosmetics because when you walk outside you’ve got to be dressed up these days. Put your stuff on. I talk about the iPhone as being a prestige brand and Uber being a prestige brand.
我认为这直接影响到了赫兹。你可能觉得,租车行业经过这么多整合之后,价格应该会涨上去才对,但你猜怎么着?根本涨不了,因为它们实际上是在跟 Uber 竞争。
I think that is directly impacting Hertz. You would think that after all the consolidation in the rental business that rates would go up, but you know what? They can’t because they’re really competing against Uber.
比尔:我不知道你们最近有没有人租过车。我以前经常去西雅图和洛杉矶出差,差别大得惊人,因为租车的话得提前 30 到 45 分钟到。
Bill: I don’t know if any of you have recently used rental cars. I used to go to Seattle and L.A. frequently for business, and the difference is astounding because with a rental car you’ve got to arrive 30 to 45 minutes earlier.
你必须登上航天飞机。乘着它飞向那里。无论你去往何处,都得带上地图。
You’ve got to get on the shuttle. Ride the shuttle out there. You have to have maps for everywhere you’re going.
你得知道车该停哪儿。你得提前 15 分钟到停车场,才能走到车跟前。顺便说一句,等你到了酒店,还得交 40 美元停车费。
You have to know where you’re going to park. You have to get there 15 minutes earlier to get to the parking structure to get to your car. And by the way, when you get to the hotel you get charged $40 to put it there.
比尔·格利 基准资本
Bill Gurley Benchmark Capital
如今优步的体验已经完全胜出了,我愿意为优步的出行体验支付租车三倍的价格。如果你考虑汽车拥有权的替代方案……哦,等等。
Now Uber’s just completely better, and I would pay three times as much for the experience I get on Uber as I would for renting a car. If you think about a car ownership alternative . . . oh, wait.
这是来自追踪企业费用支出的公司的费用数据。因此,你不仅能够推测这可能会影响租车行业,还能实际看到这一变化。上面那部分是 优步(Uber),而这个蓝色部分——之前占 50% 但现在已经降到 30%——则是租车。
This is expense data that’s now coming out from companies that track corporate expenses. So not only can you presume that it might affect rental cars, you can actually see it. That top section is Uber, and this blue section which was at 50 percent but is now at 30 percent is rental cars.
现在有了硬数据来支撑这一理论。迈克尔说,我可以展示另一家投资银行制作的资料。我想让你们对比一下,那些认为 TAM(总可及市场)只有出租车和黑色轿车的人,与这段视频所呈现的视角有何不同。这是我最后一段视频,之后我们继续往下讲。
So now there are hard data to support the theory if you will. Now, Michael said it’s okay if I show something produced by a different investment bank. I want you to contrast someone who thought the TAM was just taxi and black car with the perspective in this video. This is my last video and then we’ll move on.
[Video Clip]
[Video Clip]
比尔:我觉得那很巧妙,不过只是突出了一个差异。从顶层出发,把产品定位为汽车的替代品,这样得到的市场空间会跟教授计算 TAM 的方式大相径庭。
Bill: I thought that was very clever, but just highlighting a difference. Coming at an approach from the top down of being a car alternative gets you a vastly different result than if you think about the TAM the way that the professor did.
现在我认为他犯了好几个错误,有意思的是,在他的文章里花了很多篇幅谈论投资者会犯的判断错误。(笑)
Now I think he made a number of errors, and what’s interesting is in his piece he spent a lot of time talking about judgment errors that investors make. [Laughs]
这些引述——我不会逐字念给你们听——都是他用来论证那个为 Uber 估值的投资者群体犯了错误的所有理由。所以,他思考的是投资判断中的错误,只是没有在自己身上思考过这些问题罢了。
These quotes which I won’t read to you were all his reasoning as to why the investor group that had valued Uber had made mistakes. So he thinks about errors in investment judgment. He just didn’t think about them in himself.
除了他可能带有的诸多偏见,这篇博文实际上还附了一份声明,我觉得相当搞笑。他承认自己从未使用过这个产品。(笑声)事实上,他住在纽约,只坐地铁,没有车。所以他的立场并不足以做出什么靠谱的判断。
In addition to many biases he may have had, the blog post actually had this disclosure which I think is quite hilarious. He admitted that he had never used the product. [Laughs] In fact, he lives in New York and only takes subways and doesn’t own a car. So he didn’t have a great position to make a judgment from.
我来举另一个例子,这个例子有点更偏门,但跟投资和薪酬有关。我把漫画里的引用写在这里,方便你们看得更清楚,它说的是股票期权的重新定价,这件事在 2000 年、2001 年很常见。而股票期权因为种种原因遭到了污名化,安然和世通是其中很大的原因。我一直对此很恼火,因为这跟硅谷没关系。
Let me move on to another example that’s a little more esoteric but relates to investing and compensation. I wrote the quote from the cartoon here so you can see it larger, and it’s talking about the re-pricing of stock options, which was a common thing that happened in 2000, 2001. And stock options became vilified for a number of reasons. Enron and WorldCom were a big part of it. This always pissed me off because this wasn’t Silicon Valley.
承认,硅谷在 99 年确实干过一些相当蠢的事,但没干过这个。安然和世通都不在硅谷,可在我们经历这两家公司的案例和重新定价的过程中,期权被妖魔化了。《华尔街日报》来了,《哈佛商业评论》来了,连自称股东权益捍卫者的 ISS(机构股东服务公司)也跳出来说期权是坏东西。
Now admittedly Silicon Valley did some really silly things in ’99, but they didn’t do this. Enron and WorldCom were outside of Silicon Valley but, as we move through the re-pricing and these two examples, options became vilified. Here’s The Wall Street Journal. Here’s Harvard Business Review. Even ISS (Institutional Shareholders Services), who claims to be the champion of the shareholder, came out and said options are bad.
以下是原因。第一条是它们会鼓励过度冒险。人们认为这种薪酬方案效力太强,实际上会导致你做违法的事。这是其中之一。
Here are the reasons. The first one was they encourage too much risk-seeking. And so it was felt that it was just way too potent of a compensation scheme and that it actually causes you to do illegal things. That was one of them.
这些期权被认为具有高度稀释性。人们对每年百分之三到四的稀释程度感到不满。这些期权通常会被重新定价——这正是那幅漫画所讽刺的——而且也没有得到妥善的会计处理。
They were considered highly dilutive. People were upset about three to four percent a year dilution. They were routinely re-priced, which was what the cartoon talked about, and then they weren’t properly accounted for.
所以,这就是我们决定处理掉它们的四个原因。作为替代,我们采用了一种叫做 RSU(限制性股票单位)的新东西,至少硅谷绝大多数这类股票单位都是零基数的。
So these were the four reasons that we decided to get rid of them. In their place we’ve used a new thing called an RSU, a restricted stock unit, and at least in Silicon Valley the vast majority of these are zero basis stock units.
通常的做法是,先看你会给某人多少期权,做一个布莱克-舒尔斯对比,然后换成发放限制性股票单位(RSU)。从 2001 年开始,这种做法即使是在微软、英特尔、思科这类大公司也变得很常见。
Typically you look at what you would have given someone in options, run a Black-Scholes comparison, and then give out an RSU. Starting in 2001 this became common even at the big Microsoft, Intel, Cisco, those types of companies.
比尔·格利 Benchmark Capital
Bill Gurley Benchmark Capital
我在这里做的是,为布莱克-斯科尔斯模型设定了一些假设,但我要向你们展示的是,期权持有者的相对收益,或者通过这种限制性股票单元的计算方式,转换成等值价值的限制性股票单元的实际变动情况。
What I’ve done here is I’ve made some assumptions for the Black-Scholes model, but I’m showing you the comparative payoff for an option holder, or an equivalent value amount in the way they actually move to RSUs through this math of an RSU.
所以期权持有人,如果公司业绩不佳,就赚不到钱,显然价值为零。而如果公司股价上涨,价值则是这个走势。RSU 持有人的赔付方案与股价表现的关系是这样的。
So the option holder, if the company doesn’t perform, makes no money, it’s worth zero obviously. And this is what it’s worth if it goes up. The RSU holder has this payout scheme relative to the stock performance.
现在我要说的是,假设公司股价完全没有变化,公司业绩原地踏步,那会是什么情况。在这种情况下,拿 RSU 的高管依然能拿到和股价上涨 50% 时拿期权的高管一样多的报酬。在所有股价表现不佳的情况下,他们照样拿钱。
Now here I’ve said let’s look at what happens when the company just stays even, when the stock price doesn’t move. The RSU executive makes the same compensation when the stock price doesn’t change as the option executive makes if the stock went up 50 percent. In all these cases where the stock underperforms, they’re still getting paid.
所以,基于这一点,你觉得高管们是更喜欢 RSU 还是期权?是还是不是?绝对喜欢。他们爱死 RSU 了。
So based on this, do you think executives prefer RSUs over options? Yes? No? Absolutely. They love RSUs.
他们恨不得全吞下去。那么结果呢?嗯,我认为这是一次判断失误,因为我觉得并没有人真正想清楚之后会系统性地出什么问题。
They eat them up. So what has happened? Well, I think this was an error in judgment because I don’t think everyone thought through the systematic things that would happen afterwards.
首先,员工不会持有 RSU。如果你和任何研究薪酬的高管、董事会成员或 CFO 聊起你投资的那些用 RSU 的公司,他们会告诉你,差不多 97% 的 RSU 会在归属当天就被卖掉。所以它们根本不是股权形式,而是现金工资的一种形式。没人持有它们,这一点至关重要,因为期权曾是股权形式,能起到激励对位的作用。
First of all, employees don’t hold RSUs, and if you talk to any compensation executive, board member, or CFO about the companies you’re investing in that use RSUs, they will tell you some number close to 97 percent of RSUs are sold on the vest date. So they’re not a form of stock ownership. They’re a form of cash compensation. No one’s holding them, and that’s huge because options were a form of stock ownership and aligned incentives.
第二,它们经常在股东看不到回报的时候照发不误。这一点我不知道 ISS 当初有没有想到,但我真心希望他们想到过。在那些表现不佳的公司里,高管拿到的钱比以前拿期权的时候多得多。多得多。
Second, they routinely pay out when shareholders do not see a return, and this is something that I don’t know if ISS thought about, but I really wish they had. Executives are making way more money in companies that don’t perform than they did when they were holding options. Way more money.
它们随时都在稀释股本。期权也有稀释作用,但只有在股价上涨时才会稀释。我觉得期权比 RSU 更能让利益对位。
They dilute all the time. Options were dilutive, but they were only dilutive in upside scenarios. I think we were more aligned with options than we are with RSUs.
第三,我们以前说期权可能刺激高管过度冒险。我觉得 RSU 反而刺激高管过度保守。我真正意识到这一点,是在我为我们的一家初创公司招揽一位高管时,当时我正和一家大型上市公司抢人。
Thirdly, we talked about options maybe creating incentive for too much risk-seeking. I think RSUs create too much incentive for super-conservative executive behavior, and my mind really understood this when I was recruiting an executive into one of our startups, and I was competing with a large public company.
我和那位高管坐下来谈,我说:“行,把那家公司给你的方案给我看看。” 他做了一个电子表格,把各种东西和对方给他的 RSU 方案都放了进去。我说:“这是啥?” 他说:“哦,我就是假设未来四年股价不涨不跌。”
I sat down with the executive and said, “Well, show me what package they’re offering you at the other company,” and he had built a spreadsheet, and he had put in all these different things and the RSU package they had given him. I said, “Well, what’s this?” And he goes, “Oh, I’m just assuming the stock’s flat for the next four years.”
所以他接受那份工作时,脑子里想的就是这笔报酬。他已经打定主意不在乎股价动不动了。对我来说,这可不是你想要的东西。我是说,如果你要那种回报,你直接买债券不就行了。
So he was accepting a job, thinking about the compensation. He had already decided he didn’t care if the stock moved or not. For me, that’s just not what you want. I mean you can hold debt if you want that kind of return.
然后会计处理变得更糟了,这在类似情况下似乎很常见。1999 年以前,价内期权毫无疑问是要记作费用的。你想都不会多想。
Then the accounting became even worse, which is typical I think in these situations. An in-the-money option prior to 1999 would have unquestionably been expensed. You just wouldn’t think about it.
现在,大部分股权激励是通过零行权价、完全价内的期权来发放的,可华尔街和其他人还是希望你去看经调整后的非 GAAP(公认会计准则)经营利润,这种口径把 SBC(股权激励支出)剔除了,比期权时代更加像现金了。
Now the majority of stock compensation is through a zero basis, totally in-the-money option, and yet Wall Street and others still want you to look at non-GAAP (Generally Accepted Accounting Principles) operating income, which excludes SBC (stock-based compensation) and is even more like cash now than it ever was with options.
我就举几个例子——我知道这从科学统计上不算严谨,但我还是要说。
Just to highlight a few examples—this isn’t scientifically statistically significant I understand, but I’m going to do it anyway.
这是思科从 2000 年开始全面改用 RSU 之后,15 年来的情况。如果有人去算一下
This is Cisco since they went to RSUs in 2000, 15 years ago. If someone were to go and calculate the amount
比尔·格利 Benchmark Capital
Bill Gurley Benchmark Capital
思科管理层在这段时间通过 RSU 拿到的股权激励收益,那肯定是,我不知道,比他们通过期权能拿到的多出 100 倍。一个非常大的数字。
of stock-based earnings that Cisco management has gotten via RSUs over this timeframe, it’s going to be, I don't know, 100 times what they would have gotten from options over this time frame. Some really large number.
这是英特尔。类似的情况。我以前在一些公开投资者论坛上跟人说过:“我觉得你们根本不了解这种薪酬方案是什么样子的。它跟期权完全不是一回事。”
Here’s Intel. Same kind of thing. I’ve been at public investor forums before, and I’ve said to people, “I don’t think you understand what this compensation scheme looks like. It’s not anything like options were.”
情况甚至变得更糟了,所以我要带你们看看现在很多薪酬委员会是怎么做决策的。
It got even worse, so I’m going to walk you through how a lot of compensation committees make decisions these days.
首先,他们找一个高管。他们做一份薪酬调研,然后说这个高管的年度目标总薪酬应该是 X,我这里用 CFO 来举例。他们可能会说,这个上市公司的 CFO 每年应该赚 200 万美元。然后他们拿出工资,也许是 50 万美元,加上奖金 25 万美元,然后说剩下的应该是股权激励。于是他们给这个人价值 125 万美元的 RSU,每年都这样做,也许还会提前给他。
First they take an executive. They go run a compensation survey and they say the total target compensation for this executive should be X, and I use the example here of a CFO. They might say this public CFO should expect to earn $2 million a year. And so they take the salary, maybe $500,000, and the bonus, $250,000, and they say the rest should be stock comp. And then they assign that person $1.25 million worth of RSUs, and they do this every year, and maybe they put it out in front of them.
这基本上导致我们现在是按金额而不是按比例来设计薪酬方案。那如果股价大幅下跌会怎样?
This has basically led us to where we’re building compensation programs dollar-based not percentage-based. So what would happen if share prices fell dramatically?
2015 年,许多中盘互联网股的市值大幅缩水——LinkedIn、Twitter、Yelp——这些公司中有些每年的股权激励支出占市值的比例飙升到了 6% 到 9%。以前我们担心期权每年会造成 3% 的稀释。
In 2015 many of the mid-cap Internet stocks’ market caps fell dramatically—LinkedIn, Twitter, Yelp—and stock-based compensation as a percentage of market cap shot up through the roof for some of these companies six to nine percent a year. So we were worried about options being diluted at three percent a year.
现在我们每年是 6% 到 9% 的稀释,而且因为是零行权价的东西,期权有现金回流的部分。所以情况比这更糟。我们本来想减少稀释,结果却落到了稀释程度大幅增加、激励错位的地步。
Now we’re at six to nine percent a year, and because it’s a zero-based thing, the option had the cash part that was coming back. So it’s even worse than this. We wanted to get to a place where there was less dilution, but we’ve gotten to a place where there’s dramatically more dilution and a misaligned incentive.
我本职是风险投资人,但有人请我看看几只他们想买的股票。我真的是去建了几个模型,然后算了出来。然后我去下载了华尔街研究分析师的模型,发现他们根本没算出来。
I’m a venture capitalist by day, but someone asked me to look at a couple stocks they were thinking about buying. I literally went and built a few models, and I figured this out. Then I went and downloaded the models from Wall Street research analysts, and they hadn’t figured it out.
他们没算出来自己的股本数里没有包含这部分新增的稀释。我当时心想:“我的天啊。” 差不多两个月后,报告就满天飞了,大家都发现了这个问题,开始看这玩意有多离谱。
They hadn’t figured out that their share counts didn’t have the incremental dilution in them. I was like. “oh my god.” Literally two months later reports start flying out where people have figured this out and are starting to look at how excessive it is.
这是我最喜欢的一段话,就是查理·芒格说的。他列了一份人们在投资时容易犯的 25 种心理偏误清单,这是我最喜欢的一个。他说排第一的是,人们对动机和薪酬的考虑不够。我认为,如果你想和股东利益对位,RSU 真是一种糟糕透顶、糟得不能再糟的报酬方式。
This is one of my favorite quotes, this [Charlie] Munger quote. He had a list of [25] biases that people bring to the table around investing, and this is my favorite one. He says his number one is that people don’t think enough about motivation and compensation. I think the RSU is just a horrible, horrible way to compensate somebody if you want to align interest with shareholders.
现在我来讲讲独角兽的诞生。这不会像《权力的游戏》里龙妈那么戏剧化,但算是金融界的翻版。
Now I’m going to move on to the birthing of unicorns. This won’t be quite as dramatic as Game of Thrones, Mother of Dragons, but it’ll be the financial equivalent.
很久以前,硅谷的 IPO(首次公开募股)是这样的。苹果 1980 年上市时市值 18 亿美元。微软 1986 年上市时市值 7.8 亿美元。思科上市时市值 2.24 亿美元。
This is how IPOs (Initial Public Offerings) used to work a long time ago in Silicon Valley. Apple went public in 1980 with a $1.8 billion market cap. Microsoft in1986 at $780 million. Cisco went at $224 million.
最近情况不一样了。谷歌上市时市值 270 亿美元,Facebook 则一直等到估值 1000 亿美元才上市。这让人感到焦虑。
More recently something different has happened. Google went public at $27 billion and then Facebook waited until it was worth $100 billion to go public. This is causing anxiety.
这两个人你们可能认识,也可能不认识。这位是 DST 的尤里·米尔纳。
These are two people you may or may not know. On this side is Yuri Milner of DST (Digital Sky Technologies).
尤里在 Facebook 估值约 100 亿美元时作为一家私有公司投了将近 10 亿美元。我不知道他具体什么时候退出的,但 Facebook 现在值 3000 亿美元,你们自己算。
Yuri invested almost a billion dollars in Facebook at about a $10 billion valuation as a private company. I don’t know exactly when he got out, but it’s worth $300 billion today, so you run the math.
比尔·格利 Benchmark Capital
Bill Gurley Benchmark Capital
这位是老虎环球基金的斯科特·施莱弗。他在京东作为私有公司时投了 1.3 亿美元,给基金带回了大约六七十亿美元的回报。
This is Scott Shleifer of Tiger [Global Management]. He invested $130 million in JD.com as a private company and returned about 6 or 7 billion dollars to shareholders.
现在,如果你在这样的一家基金工作,看到公司越来越晚才上市,一上市就值 1000 亿美元,看到尤里和斯科特这种人通过投资私有公司赚了大钱,你就会得 FOMO,也就是害怕错过。
Now if you’re working for one of these firms and you see companies going public later at $100 billion and you see people like Yuri and Scott making tons of money by investing in private companies, you have FOMO, fear of missing out.
这是一个被研究得很透的行为科学问题,会导致判断失误,而这些基金以及许多其他基金(很可能包括你们中的一些)会决定去做斯科特和尤里做过的事,一头扎进私有市场投资,这就催生了独角兽。独角兽就是这么来的。
It’s a well-studied behavioral science problem that causes misjudgment error, and these firms and many others and probably many of your firms decided they were going to do what Scott and Yuri did and dive into private market investing, which created the unicorn. That’s exactly what created the unicorn.
现在有超过 200 家私有公司估值超过 10 亿美元,其中大多数每家的融资额都超过 1 亿美元。
Today there are over 200 private companies that are valued at over a billion dollars, and most of them have raised over $100 million each.
我把这称为从未有过的伟大实验,因为我们以前从未把这么多私有资本塞进不成熟的私有公司里。我们向这些公司投入的资金远远超过 1999 年时能想到的任何水平,后果一定会出现。
I call this the great experiment that’s never been done before because we’ve never crammed this much private capital into immature private companies ever. We’ve put way more money into these companies than was ever considered in 1999, and there will be consequences.
我来说说哪些因素在推波助澜,让情况可能变得更糟。风投把钱投给创业者。然后创业者向后期投资者融资,推高公司估值。这个价格被展示给他们的有限合伙人(LP),LP 看到他们干得这么好,就给他们更多钱。
Let me tell you a little bit about what’s reinforcing this and making it worse than it could possibly be. Venture Capital has put money in entrepreneurs. They then raise money from late-stage investors and mark up what the company is worth. That price is then shown to their limited partners (LP) who then give more money to these guys because they’re doing such a great job.
他们甚至能赚到管理费。他们的报酬基于这个人让那个人以更高的价格开出支票,而且就是用这种方式来衡量的。现在让我给你们看我整个演讲中最重要的一张幻灯片。
They even get paid. They get paid on the fact that this guy got this guy to write a check at a higher price, and it’s measured that way. Now let me show you the most important slide in my entire deck.
我们这个行业现在的 IRR(内部收益率)创下历史新高,但流动性几乎为零,没有并购,没有 IPO。这难道没问题吗?这是个经典的按市值计价会计问题,终归要出问题的。而且已经开始出问题了。
My industry has record high IRRs (internal rate of return) right now and almost zero liquidity, no M&A (mergers and acquisitions), no IPOs. Is anything wrong with this? This is a classic mark-to-market accounting problem, and it’s going to come unwound. It has already started to come unwound.
你们可能读过这个。我注意到今天《华尔街日报》封面又提到它了。以 90 亿美元估值融资。LP 按 90 亿美元估值记账。这家公司叫 Zenefits。他们当时的营收大约是 2500 万美元,以 45 亿美元估值融资。
You’ve probably read about this. I noticed it’s on the cover of the Wall Street Journal again today. Raise money at $9 billion. LPs marked it at $9 billion. This is a company called Zenefits. They were doing about $25 million in revenue and raised money at $4.5 billion.
这 45 亿美元被用于向 LP 按市值计价。他们不仅能拿到奖励。他们还能根据这些回报进行分配。他们告诉别人 2015 年能实现 1 亿美元的销售额。实际做了大约 6000 万美元。后来他们不得不解雇销售团队。
That $4.5 billion was then marked to the LPs. Not only do they get the bonus. They make distributions based on these returns. They told people they’d do $100 million in sales in 2015. They did about $60 [million]. They’ve since had to fire their sales force.
一位董事会成员告诉这位创始人帕克·康拉德,说他不够有野心。于是康拉德写了一个程序来自动化处理销售人员完成合规考试的流程,这很可能招来刑事调查。它不值 45 亿美元。
A board member told this guy, Parker Conrad, he wasn’t being ambitious enough. So he wrote a program to automate filling out compliance exams for their sales force, and there are probably going to be criminal investigations. It’s not worth $4.5 billion.
Palantir 的财务数据一个半星期前泄露了。这家公司以 200 亿美元估值融资,并且也按这个价格向 LP 计价和分配了。
Palantir‘s financials leaked a week and a half ago. This company’s raised money at $20 billion, and it’s been marked that way to the LPs and distributed.
数据显示他们 2015 年销售额 4 亿美元,增长率 50%,但没赚钱。他们大部分工作更像是咨询业务。那么,一家 4 亿美元销售额、增长 50%、还没赚钱的咨询公司,到底值多少钱?
It says they did $400 million in sales in 2015 with a 50 percent growth rate and were unprofitable. Most of their work is really consultant-like work. So what is a consultancy that’s unprofitable at $400 million growing 50 percent worth?
我问了一群投资者。在 15 亿美元之后,没有人继续举手。所以,在 200 亿美元的估值标记下,1500 万美元是个巨大的差异。顺便说一句,过去三个月里,已有次级交易价格在 170 亿美元左右。
I asked a group of investors. No one kept their hand up after $1.5 billion. So $1.5 billion marked at $20 billion is a huge differential. By the way, within the past three months there were secondary trades around $17 billion.
这就是正在发生的事情。这就是泡沫如何破裂的。结果就是,你看到估值全面下调
This is what’s happening. This is how it’s coming undone. As a result you’re seeing markdowns across the board
比尔·格利(Benchmark Capital)
Bill Gurley Benchmark Capital
而且在许多公司中都是如此,独角兽也不再那么受宠了。
and across many firms, and then unicorns aren’t so beloved anymore.
有趣的是,又一次,人们在试图寻找回报的过程中,得到了完全相反的结果——亏损。
What’s interesting once again is that in an attempt to find returns, you get the exact opposite. You get losses.
这是怎么发生的?他们忽略了什么?
How does this happen? What did they miss?
首先,如果只有 DST 和老虎基金自己在做,这个策略效果很好。但当所有人都跳进这场游戏,难度就大大增加了。这只是典型的共识思维与非共识思维、逆向投资的对比。
First of all, this strategy works great if DST and Tiger are doing it by themselves. But when everybody jumps in the game, it gets a lot harder, and this is just classic consensus versus non-consensus thinking, contrarian investing.
更大的因素是:资金影响了生态系统。他们假设这样做不会有任何后果,但实际上存在很多很多问题。这些公司融到的钱比以往任何时候都多。
Here’s the bigger factor: the money affects the ecosystem. They assume that if you do this, there won’t be any ramifications, but there are many, many problems. These companies have raised more money than ever.
这些公司没有资本支出。他们不开店。他们不建工厂。如果你给他们更多钱,他们会雇更多人,烧钱率更高,距离核心的单位经济效益越来越远。很难知道他们是否还能恢复过来。
These companies don’t have capex. They don’t build stores. They don’t build factories. If you give them more money, they hire more people and they create bigger burn rates, and they get further away from their core unit economics. It’s hard to know if they’ll ever bring it back together.
有一件事我没列在这里:很多人认为史蒂夫·乔布斯对设计的最大贡献是他创造了约束。他不会告诉团队要做什么。他会告诉他们必须做到这么薄、这么高、或者这么重,然后你们自己去想办法。
One thing I didn’t put on here: a lot of people think that Steve Jobs’s greatest contribution to design was that he created constraints. He wouldn’t tell the team what to do. He would tell them it has to be this thin or this tall or weigh this amount, and you go figure it out.
我认为这对创业公司也一样。如果你有约束,你会做出更好的决策。如果你有无限的约束,你会什么都做。你会做出更差的决策,我认为这些公司的执行质量,比没有撒出所有这些钱时要差得多。
I think the same thing is true with startups. If you have a constraint, you make better decisions. If you have unlimited constraints, you do everything. You make worse decisions, and I think the quality of execution in these companies is way worse than it would have been had you not handed out all the money.
然后这又制造了过度的竞争,我认为这现在正在影响中盘公共互联网股票,因为市场上的钱实在太多了。
Then it creates excess competition, and I think this is now impacting the mid-market public Internet stocks because there’s just such an excessive amount of money.
所以,如果你们当中有人关注新兴借贷公司,OnDeck 的股价是 5 美元,Lending Club 是 3 美元。而 SoFi 和 Avalon 已经筹集了 10 亿美元的私募股权,SoFi 的 CEO 说 Lending Club 之所以不成功,是因为他们不够有野心。
So if any of you track the new lending companies, OnDeck’s stock is at $5 and Lending Club’s is at $3. Well, SoFi and Avalon have raised a billion dollars of private equity, and the SoFi CEO said the reason Lending Club’s not working is because they’re not ambitious enough.
我个人不认为借贷和野心是两件你想结合起来的东西。但问题是过度竞争。几乎在每个垂直领域,总有人愿意一年亏损 2 亿美元、4 亿美元,在这个案例中是 6 亿美元,如果你的竞争对手能做到这一点,你想盈利就非常困难。所以这只是在过度资助这个类别。
Now I personally don’t think lending and ambition are two things you want to combine, But the problem is excessive competition. You can go through almost every vertical and there’s someone out there who’s willing to lose $200 million, $400 million, in this case $600 million in a year, and it’s very hard to be profitable if your competitor’s able to do that. So this is just overfunding the category.
最后,IPO 过程提供了人们没有意识到的巨大价值。公司会为此做准备。他们非常认真对待这件事。他们实现 GAAP 合规。审计师也更认真对待。
Then lastly the IPO process provided immense value that people didn’t think about. Companies prepare for this event. They take it remarkably serious. They get GAAP compliant. The auditors take it more seriously.
这可能是一件难以理解的事。审计师在公司即将上市时比在公司私有时会更加努力。我向你保证。这就是为什么在上市前你会得到这样的声明,因为他们会把它提交到全国层面。
This is probably just an amazingly hard thing to fathom. The auditors try harder when you’re about to go public than they do when you’re private. I guarantee you. And that’s why you get statements right before you go out and this kind of thing because they send it to national.
大多数这些公司筹集资金时都没有经过审计的财务报表。其中一些公司,他们会拿到审计报告。在私有公司,你可能在 2016 年 11 月才拿到 2015 年的审计财务报表,所以没有人知道 PowerPoint 上写的数字如果按照正规会计处理是否准确。
Most of these companies are raising money without audited financials. Some of them, they get the audit. In a private company you might get audited financials for 2015 done in November of 2016, so no one knows if the math that’s put on the PowerPoint is even going to be accurate, if it were properly accounted.
银行家们,每个人都确保你准备好了,公司、管理层和董事会会考虑成为上市公司的份量,但他们没有考虑到这与私有状态相比的差异。
The bankers, everyone makes sure you’re ready, and companies and management and boards consider the weight of being public, and they don’t think about this relative to being private.
这是来自红杉资本我的同行迈克·莫里茨的一段话。他写了一篇关于“次贷独角兽”的精彩文章。文章不长,如果你想读,它发表在《金融时报》上。他说:“作为一家私有公司,可以更容易地掩盖弱点,呈现不可战胜的光环,并通过减少披露来迷惑投资者,
Here’s a quote from Mike Moritz who’s one of my peers at Sequoia. He wrote a fabulous article about ”subprime unicorns.” It’s not very long if you want to read it, it was in the Financial Times. And he says, “It is easier to conceal weaknesses, present an aura of invincibility and confound investors as a private company that can
比尔·格利(Benchmark Capital)
Bill Gurley Benchmark Capital
而不是作为一家上市公司。”
escape by making fewer disclosures than as a publicly traded one.”
我想说,我们创造了一个有利于过度宣传的竞技场,因为每个人都被鼓励成为独角兽,人们制作的 PowerPoint 内容天马行空,而且不遵守任何你通常会在 IPO 过程中看到的原则。我认为在接下来的 18 个月里,你会看到更多像我强调的那三家公司一样的情况。
I would say that we created a playing field ripe for over-promotion because everyone was being encouraged to be a unicorn, and people were putting together PowerPoints that were fantastical and didn’t adhere to any of the same principles that you normally would see in an IPO process. I think you’re going to see a lot more companies like the three that I highlighted over the next 18 months.
这段很短,然后我就要进入问答环节了。这其实是我很熟悉的一个人,叫尼克·哈诺尔。他是亚马逊的早期投资者,后来又投资了这家公司——Avenue A/aQuantive,微软以 60 亿美元收购了它用于广告领域。
This one’s really short, and then I’ll turn it over to Q&A. This is actually someone I know well, a guy named Nick Hanauer. He was an early investor in Amazon.com and then invested in this company, Avenue A/aQuantive that Microsoft bought for $6 billion in ad space.
所以他做得非常好。但他现在成了一个在最低工资问题上非常直言不讳的人,他推动了西雅图将最低工资提高到 15 美元。现在他希望西雅图提高到 28 美元。其他人也纷纷效仿。你可能读到过纽约和加州也将提高到 15 美元。
So he’s done really well for himself. But he’s become a very outspoken person on minimum wage, and he pushed to have Seattle go to $15. And now he wants Seattle to go to $28. Others have followed suit. You probably read about New York and California now going to go to $15.
当尼克谈到这一点时,我只想强调他说的其中一点:大部分就业增长来自服务业岗位,他说其中 39% 来自餐饮业。他谈到了服务员(男女都有)。
When Nick talks about this, I’ll just highlight one of the things he says is most of the job growth has been in service jobs, and he says 39 percent is in food. He talks about waiters and waitresses.
我住在硅谷,我是 OpenTable 的投资者,有一些看起来有点像 OpenTable 的新公司经常来找我。其中一家叫 E la Carte,他们的产品叫 Presto。这家公司一直表现平平。还有一家在达拉斯,叫 Ziosk。这些东西放在桌子上,取代了男女服务员。
I live in Silicon Valley, and I was an investor in OpenTable, and there are some new companies that look a little bit like OpenTable that often call on me. One of them is called E la Carte, and this is their product called Presto. and the company’s just been kind of middling along. And this one is in Dallas called Ziosk. These things sit on tables and replace waitresses and waiters.
两周前我去拜访他,他的客户线索在西雅图、纽约和加州从未如此之高。我不认为尼克本意是要加速服务员的消亡,但他可能确实通过他的努力做到了这一点。这些都是关于这项技术如今正在兴起的文章。我对低端和中端服务餐厅非常有信心,在未来 5 到 10 年内,这些设备将变得无处不在。
I went and visited him two weeks ago, and his inbound leads have never been higher there in Seattle, New York, and California. I don’t think Nick intended to accelerate the demise of waiters and waitresses, but he may have in fact done so with his efforts. And these are articles about this technology now taking off. I am highly confident in low- and mid-service restaurants that these will become pervasive in the next five or ten years.
所以这是我对此的总结。在每一个案例中,甚至在 Uber 的案例中,我认为都有一个有偏见且充满激情的倡导者,他们有一个想要解决的问题,对此他们感到极其热情。
So here’s my summary on this. In each of these cases, even in the Uber case, I think you had a biased and passionate advocate who had a problem they wanted to solve that they just felt extremely passionate about.
在教授的那个案例中,我认为他只是认为投资者很愚蠢,他想证明他们很愚蠢。在其他三个案例中,我认为人们要么觉得某些事情不公平,要么觉得这是一个需要解决的问题,所以他们寻找一个非常生硬的手段来解决它。
In the professor’s case, I think he just thought that the investors were being silly, and he wanted to prove that they were being silly. In the other three cases, I think people felt like either something wasn’t fair or it was a problem they needed to fix, and so they were looking for a very blunt instrument to fix it.
我认为所有这些事情都存在于非常复杂的系统中,这就是为什么我引用《复杂性》这本书,因为你有关联的状态机、多个变量。你有人类行为。我认为你需要系统性思维来看待这些方法。你必须思考所有输入。A 如何导致 B,B 又如何导致 C?当你转向一个新解决方案时,所有随之而来的后果是什么?
I think all of these things exist in very complex systems, and that’s why I reference the book Complexity, because you have interconnected state machines, multiple variables. You have human behavior. And I think you need systematic thinking to look at these approaches. You have to think through all the inputs. How is A going to cause B going to cause C? And what are all the ramifications that are going to play out when you move to a new solution?
我认为在最后三个例子中,他们都没有做到这一点。最讽刺的是,在三个案例中,最终的局面比你原本想解决的问题还要糟糕,我认为这种情况经常发生。以上是我准备好的发言,时间 42 分钟,所以我比预定时间少了 3 分钟。
I don’t think they did that in any of those last three examples. Then the ultimate irony is in three of the cases you end up in a worse situation than the problem you intended to fix, which I think happens frequently. Those are my prepared remarks at 42 minutes, so I'm three minutes under where I was supposed to be.
现场股东:非常感谢。你提到你是在两年半前才听说达摩达兰教授,但你可能不知道的是,在量化分析师中间,他有着巨大的声誉,不是因为他像吉姆·克莱默那样是个激动人心的选股者,而是因为他在线发布了估值模型和电子表格,并证明了你不需要非常了解公司。一些相当简单的指标就能给出不错的答案。
Question: Thanks very much. You mentioned you’d only heard of Professor Damodaran two and a half years ago, but what you may not know is among quants he has a tremendous reputation, not because he’s as exciting a stock picker as Jim Cramer, but because he’s put his valuations online, in spreadsheets, and has demonstrated that you don’t have to understand the companies a lot. Some fairly simple metrics can give you good answers.
所以我有一个假设性问题。假设你想知道所有私有科技公司的总价值,
So I have a hypothetical for you. Let’s say you wanted to know the total value of all private technology
比尔·格利(Benchmark Capital)
Bill Gurley Benchmark Capital
你会相信教授的计算结果,还是这些公司董事会成员调查的总和?
companies. Would you trust the professor’s calculation or the sum of a survey of the board members of these companies?
比尔:教授。是的,我会相信教授。我这么说的原因是,在我的行业里存在惊人的乐观主义,也许这是这份工作的一个必要特质。
Bill: The professor. Yeah, I would trust the professor. The reason I say that is there’s an amazing amount of optimism in my industry, and maybe it’s a requirement for the job.
我认为如果你天生就是一个怀疑论者,你在风险投资领域不会做得很好,但这里就是有一种固有的乐观主义。这有周期性的,但在过去 5 到 10 年里,我的行业里大量资金被交给了那些从未做过投资者的人。
I don’t think if you were just inherently a skeptic you’d do very well in venture capital, but there’s just an inherent amount of optimism. And this goes in cycles, but in my industry in the past five or ten years, a lot of money’s been given to people who have never been investors.
他们吹嘘自己的运营履历,但这并不反映他们是否具备任何投资资质或可信度,我认为他们甚至都不了解资本市场是如何运作的。这也是导致问题的一部分原因。
They celebrate their operating history, but they don’t reflect whether they have any investing credentials or credibility, and I don’t think they understand even how capital markets work. That’s part of what’s led to the problem.
现场股东:这非常有趣。谢谢。你举了一些例子,说明人类判断错误并且相差甚远。你结论中的一个观点是,这些情况需要系统性思维。
Question: This is really interesting. Thanks. You gave examples where humans got it wrong and missed by a mile. One of your points in your conclusion was these instances require systematic thinking.
你认为机器会做得更好吗?特别是,在你给出的一些例子中,我很难想象能构建一个机器,可以弄清楚那些人类没有弄明白的事情。
Do you believe that machines would have been better? And in particular in some of the examples you gave, I find it hard to construct a narrative where we could build a machine that would have figured out what the people didn’t.
比尔:这是一个我认为其他演讲者会比我更有发言权的话题。我对此有所接触,因为我的行业往往会对某些主题过度兴奋,有点发疯,目前其中之一就是机器学习和人工智能(AI),我们内部研究过哪些类型的问题适合用机器解决,哪些不适合,以及哪些可能具有投资价值、哪些没有。
Bill: This is a subject that I think other speakers are going to have a lot more knowledge on than myself. I’m exposed to it a little bit because my industry tends to get hyper-excited about certain themes and go a little nuts, and one of them right now is machine learning and AI (Artificial Intelligence), and we’ve studied internally the types of problems that work well and those that don’t and ones that might be investable or not.
我认为我指出的这类问题,很可能属于最后一批可能被人工智能解决的事情,原因就在于我之前提到的各种复杂性。部分决策是基于其他人如何反应来作出的,而这种反应又总在变化。这片领域的复杂程度实在太高了。不过,我还是把这个问题的回答权留给更懂行的人吧。
I think that the types of problems that I highlighted would probably be some of the last things that would be possible just because of all the complexity that I mentioned. Some of these decisions are based on how other humans react, and that’s constantly changing. It’s just a very complex surface area but, once again, I’d maybe re-pose the question to someone who knows more.
迈克尔:佩德罗,你对这个有什么想法?
Michael: Pedro, do you have a thought on that?
佩德罗:哎呀,我对这问题可真有想法。[笑声] 我的意思是,我认同你的看法。这些东西今天对机器来说,可能确实是比较难做到的事情之一。
Pedro: Boy, do I have thoughts on that. [Laughter] I mean my thought is that, yes, I agree with you. These are probably some of the harder things for machines to do today.
我认为机器学习系统可能会注意到很多人类不会注意到的事情,原因很简单——它在查看大量人类不留意、甚至早已遗忘的信号。突然之间,某个信号开始变得非常清晰,我觉得有些情况下已经发生了这种事,这在某些领域或许也是可能的。
I think there are a lot of things that a machine learning system might notice that people wouldn’t just because it’s looking at a lot of signals that people aren’t, that people have forgotten about. Suddenly one of those signals starts to be really clear, and I think there are cases where that’s happened, and that might be possible with some of these things.
这些机器的一个优势在于,它们能比人类看到更多整体图景。如果这意味着它们变得更困惑,那就不妙了。但这也意味着它们实际上能更好地理解正在发生的事情。
One advantage that the machines have is that they can look at more of the picture than the humans can. If that means they get more confused, then it’s bad. It also means that they actually understand what’s going on better.
你提到米奇·沃尔德罗普的那本书,我觉得这是个很有意思的切入点,因为这一行归根结底就是这个道理。复杂系统的运行方式,和它内部各个组成部分的运行方式是不一样的,可人们偏偏只盯着去模拟那些局部行为。
Your reference to Mitch Waldrop’s book I think is a very interesting one because this is really what it’s all about. The behavior of the complex system is different from the behavior of the individual parts, and people just look to model the behavior of the individual parts.
如今,大多数机器学习算法也只是在模拟个体部分的行为。但更优秀的算法实际上开始把系统作为一个整体来建模。我认为,一旦做到这一点,效果就能好得多。
Today most machine learning algorithms are also just modeling the behavior of the individual parts. But the better algorithms actually start to model systems as a whole. And I think once you do that you can do a lot better.
比尔:在所有这些之中,我觉得薪酬有潜力搞明白并敲定,只要我的判断没错。
Bill: Of all of them, I think that compensation would be the one that it could potentially figure out and nail if I’m right about my thesis.
比尔·格利 基准资本
Bill Gurley Benchmark Capital
问题:您如何看待您规模天平上的不平衡如何演变,这又如何影响您正在做的事情?
Question: How do you see the imbalance on your scale slide playing out, and how does that impact what you’re doing?
比尔:我每天都在想这件事。在独角兽的故事中,我漏掉了一个我骨子里认为完全真实的部分:全球范围内的低利率环境,正在对愿意涌入我所在的领域及其他领域的资金量,产生极为显著的影响。
Bill: I think about that every single day. The part I left out in the unicorn story that I think is inherently true is that the low interest rate environment on a global scale is having a remarkable impact on the amount of money that’s willing to come into my field and others.
即使到了今天,当初那些融资故事已经被吹得天花乱坠,我大概还能在接下来一周里跟 20 个想把更多钱投进我们这个行业的人见面,而且这些都是主动找上门来的,我们还在往后推。很多资金来自全球各地,俄罗斯、中东、中国,都在寻求全球多元化配置。钱多得简直吓人,所以我真不知道得什么情况才能让这股势头停下来。
Even today after those stories have blown up, I could probably take 20 meetings with people who want to put more money in my industry in the next week, and these are inbounds that we’re deferring. A lot of it’s coming globally, out of Russia, the Middle East, China, looking for global diversification. There’s just vast amounts of money, and so I don’t know what it will take.
这些公司中有很多也筹集了巨额资金,短期内不会耗尽。所以这一轮洗牌是从 2015 年 10 月/11 月开始的。可能要等到整整 18 个月之后,当这些公司真把钱烧完了,你才会看到灾难性的后果或行为转变。
A lot of these companies also have raised so much money that they don’t run out right away. And so this shake-up started October/November 2015. It’ll probably take a full 18 months after that when people literally run out of money before you start to see catastrophic effects or behavior change.
不过我之前说过,有些独角兽变成了僵尸。比喻方式有很多种。像 Zenefits 这种公司就对销售团队进行了大规模裁员。我所谓的僵尸,是指它们放弃了当初用来以超高估值融资的那条增长轨道,但它们手里还有现金,于是降档运转,不再去追逐曾经的金矿了。那家公司可能会维持很长时间。
Now some of the unicorns I’ve said, though, become zombies. There’s a lot of different metaphors. I think companies like Zenefits have done massive layoffs of their sales force. What I mean by zombies is they’ve given up on the growth trajectory that was used to raise the money at the super high valuation, but they have the cash, and so they downshift, and they’re no longer chasing the gold they once were. That company might stay around for a very long time.
另一个起作用的因素,尤其是在你们中有些人或许正考虑加入这场精彩舞蹈的情况下,是私募投资中的一个术语,叫做“清算优先权”。
Another factor that played a role, especially because some of you might be considering joining this wonderful dance, is a term in private investing called “liquidation preference.”
这个条款的基本意思是,如果你愿意,你可以拿回你的钱,而不转换成普通股。所以,如果公司卖出的价钱低于你当初支付的价格,你仍然可以拿回你的钱。条款就是这么运作的。
It basically says that if you want, you can get your money back and not convert to common. And so if a company sold for less than you paid, you can still get your money back. And that’s how the term works.
大多数进入市场的人,就像那张图表上的许多名字一样,认为股市会像债券一样运作,因此觉得自己的钱很安全。价格其实并不重要,因为他们只是想押注下一个谷歌或脸书。
Most of the people that came into the market, like many of the names on that chart, thought that it would act like debt, and so they thought their money was safe. Price doesn’t really matter because I just want an option on this being the next Google or Facebook.
他们不明白的是,因清算优先权而获得的回报其实只是极小部分。当这些公司出现问题时的实际情况是,它们会被资本重组——像 Foursquare 和 Jawbone 这样的公司已经被重组过了——你的清算优先权也就随之消失。董事会只需投票表决,或者股东就能把你洗掉,而大多数独角兽交易中的条款都没有针对这种情况的保护措施。所以这需要一些时间。
What they don’t understand is that returns paid out due to liquidation preference are a vast minority. What happens when these companies stumble is they get recapitalized, and a couple of them like Foursquare and Jawbone have already been recapped, so your liquidation preference goes away. The board just votes and elects or the shareholders wipe you out, and most of the terms in these unicorn deals don’t have protection against that. So it’ll take a while.
我原本希望这事能像 1999 年或 2001 年那样发生,因为我有责任和义务,去兑现那些我们的有限合伙人已经拿到业绩奖金的内部收益率。 (笑) 在这种环境下,要进行收益分配非常困难。要实现流动性也非常困难,因为预期存在怪异的错配,而且以远高于公开市场的私募价格来筹集资金的能力还在。在预期转变之前,并购交易不会发生。
I had hoped that it would happen like in 1999 or 2001 because I feel a duty and obligation to perform on those IRRs that our LPs have already been bonused on. [Laughs] It’s very hard to do payouts in this environment. It’s very hard to get to liquidity with the weird expectation mismatch, with the ability to raise money at private prices that are vastly different than public prices. The M&A’s not going to happen until the expectations shift.
提问者:关于 Lyft / 滴滴 / Uber 这些现象,目前公开市场上是否存在一些被低估的方面——你对此高度确信 Uber 会成功?第一个问题,如果还能偷偷塞进第二个:WeWork 这个现象。
Question: Are there things that are underappreciated publicly now about the Lyft/Didi/Uber phenomenon where you feel highly confident that Uber succeeds? One, and then if could sneak in two, the phenomenon of WeWork.
比尔:(笑)这算是个问题吗?
Bill: [Laughs] Is that a question?
Question: Yeah.
Question: Yeah.
比尔:(笑)好,那咱们先停一下,聊聊 WeWork。优步那边在竞争上遇到的情况——你
Bill: [Laughs] All right, we’ll pause for WeWork. The thing that’s happening with Uber’s competition—you
比尔·格利,基准资本
Bill Gurley Benchmark Capital
我刚才提到的中国滴滴出行和北美地区的 Lyft,就与我所说的资本环境密切相关,Lyft 当时正在大举烧钱。
mentioned Didi in China and Lyft here in North America—is tied to the capital environment that I’m talking about, and so Lyft was burning.
他们在 2015 年决定猛踩油门,我不知道当时的烧钱速度是多少,但很快就涨到了每月 2500 万美元。后来从通用汽车拿到一笔钱,又涨到了每月 5000 万美元。
They made a decision in 2015 to aggressively push the gas, and I don't know what their burn rate was but they went up to $25 million a month. Then when they got some money from General Motors, they went to $50 million a month.
这是公开信息,已经见诸报端。所以他们的年化营收规模达到了 6 亿美元,而且他们开始了被我称之为“租赁市场份额”的行为,也就是花钱购买市场份额。他们或许会把这叫做“挣得”,那我就不知道了。
This is published. It’s out in the press. So there’s a $600 million run rate, and they started what I call renting market share, buying market share. They might call it earning. I don't know.
实际情况是,如果你拥有系统性或规模优势——在一家普通公司里,可能一家公司没有利润,而另一家却能赚到 30% 的利润——那你就说它们有规模优势或网络效应。
What’s going on is if you had a systematic or scale advantage, in a normal company you might have one company with no profits and one with 30 percent profits, and you say they have a scale advantage or a network effect.
在这种情况下,可用的资本实在太多了,以至于这些事情都在亏本进行。一家公司每单业务亏的钱远多于另一家公司。这些数据都不公开,所以每个人都在猜测这些数字究竟是多少,而你根本无法知道。
In this case there’s so much capital available that those things are happening with losses. One company loses way more money per ride than another company loses per ride. None of this information’s public, so everybody’s guessing what these numbers are, and you just won’t know.
巴菲特有句名言,大意是“只有当潮水退去,你才能看出谁在裸泳”。我想,现在的情况正是如此。
There’s a famous Buffett quote I think where he says something like, “You don’t know who’s naked until the water goes out.” And I think that’s the situation we’re in, in this case.
过去发生的一件事情,我觉得谁也预料不到,那就是有多少人最终转而认同了视频里的观点,而不是教授的观点,结果有不少利益相关方因此卷入其中。
One of the things that’s happened that I think was unpredictable was how many people have come around to the view in the video as opposed to the professor’s view, and there are lots of interested parties as a result.
第二个,WeWork。(笑)在座大多数人可能根本不知道 WeWork 是谁。我们是一家叫 WeWork 的公司的投资人,这笔投资作为风险投资来说,对我们而言非常不典型。
On the second one, WeWork. [Laughs] Most of you probably don’t even know who WeWork is. We’re an investor in a company called WeWork that’s very atypical for us as venture capitalists.
这位名叫 亚当·诺伊曼 (Adam Neumann) 的企业家,确实是一位了不起的企业家,他断定我们在人们喜欢如何工作和生活方面正经历一场文化革命,而你可以通过重构办公空间来满足这一需求。他的做法基本上是整体租下高层建筑的整层楼,然后进行改造,设置更小的独立工作单元,但配以大量玻璃幕墙、公共区域和公共会议室。
This entrepreneur, who’s really an amazing entrepreneur, named Adam Neumann, decided that we were undergoing a cultural revolution in how people like to work and live, and that you could restructure office space to meet that need. He basically rented whole floors of tall buildings, and he retrofitted them with much smaller individual work units but a lot of glass and common areas and common meeting rooms.
如果你走进 WeWork 的某个办公空间,你真的需要理解它究竟是什么——那里跟传统的共享办公空间完全是两码事。如果你在一家只有两个人的公关公司上班,你会感觉自己像是在跟一大群人一起工作,而不是走进去关上门谁也不认识。
If you ever go into one of these places, and you really have to understand what WeWork is, there’s a totally different vibe than in a historic shared office space environment. And if you work for a two-person PR firm, you feel like you’re going to work with a lot of people as opposed to just going in and closing the door and not knowing anyone else.
千禧一代正大规模向城市迁移,越来越多的人从事自由职业。他认为自己的业务正好契合这一趋势。他在填补这些空间方面相当成功,每平方英尺收取的租金也远远超过他当初的买入成本。
There’s a big urban shift with millennials and a lot more people doing independent work And he believes this ties in to all that. And he’s been quite successful at filling these and extracting a rent per square foot that’s dramatically above what he paid.
这家公司收到过很多问题——关于它应该如何估值,它是类似租赁租户那样的业务,应该那样估值,还是应该像科技公司那样估值,我想亚当会试图争辩后者。
The company gets a lot of questions—about how it should be valued, s it just a leased tenant kind of thing, should it be valued like that, or should it be valued as a tech company which I think Adam would try and argue.
他最近推出了一款新产品,名为 WeLive,基本上就是为大学毕业后的年轻人准备的集体宿舍。你看看旧金山的租金,我认为大多数起步月租都要 3000 美元。所以他准备提供一款月租约 1100 美元的产品来竞争,而且社交性也比其他产品强。所以我也不清楚它到底该怎么估值。在座的各位,我想甚至有人已经给它估出了相当高的价格。
He recently launched a new product called WeLive that is basically a dorm for post-college people. And if you look at rents in San Francisco, I think most starting rents are three grand a month. And so he’ll be providing a product that’ll compete at like $1,100 a month and also will be more social than the others. So I don't know exactly how to value it. Different people in this room I think even have valued it fairly highly.
问:问个关于 IPO 这事的问题,因为我从自己角度看,现在公司上市其实比以往任何时候都容易。这些公司尽可能长时间保持私有,这种两极分化根本说不通。一边是有大量公司选择不上市,另一边是……
Question: A question on this IPO situation because from where I sit it’s actually easier to IPO a company than it’s ever been. This dichotomy of these companies staying private for as long as possible just doesn’t make sense. On one hand you have an enormous amount of companies that are staying private, and on the other
比尔·格利 Benchmark Capital
Bill Gurley Benchmark Capital
另一方面,上市从未像现在这样容易。
hand, it’s never been easier to go public.
比尔:所以你看。《就业法案》的修改,特别是允许首次申报保持非公开的那一条,效果极好。你可以和美国证券交易委员会(SEC)来回两三轮修改,前提是公司尚未上市,比如 Groupon 和一些其他高调公司就这么做过。然后从打法转向定价的时间窗口大大缩短,这非常正面,全都是好处,因为那段时期的风险变小了。
Bill: So look. The change to the Jobs Act, especially the one that allowed the first filing to be private, has been awesome. And you get two or three iterations with the SEC (Securities Exchange Commission) for those that aren’t public, like Groupon had and some of these other high-profile ones, and then your window from flip to price is a lot shorter, which is way positive, all good because there’s less risk in that period.
大约四五年前,硅谷开始有人鼓吹,说保持私有状态更好。我觉得有些创业者容易听信这种论调。这些创业者不喜欢接受审视,认为任何讲究规则的东西都是官僚主义。
There are people that started a rhetoric four or five years ago in Silicon Valley that staying private was better. I think there are certain entrepreneurs who were open to that message. Entrepreneurs who don’t like scrutiny, , and think that anything that’s rule-oriented is bureaucratic.
他们围绕这个编了一套说辞,说什么公开上市太难了……把责任推到股市的短期思维上,说自己想做长期的事,没人允许他们做。我想强调一点,这种说法从未阻止过杰夫·贝佐斯、马克·贝尼奥夫或里德·哈斯廷斯去做任何事情。
They put a whole narrative around it like it’s harder . . . they blame it on the stock market’s short-term thinking, saying they want to do long-term things and no one’s letting them. I’d like to highlight that it never stopped Jeff Bezos or Marc Benioff or Reed Hastings from doing anything.
但我认为他们骨子里是害怕,而且其实不必如此。硅谷有些董事会,哪怕在你们帮忙融资的独角兽公司里,也纯粹是在搞二次发行套现。
But I think they’re basically afraid, and they don’t have to be. And some of these Silicon Valley boards, even in the unicorns that you guys have helped fund, are just printing secondary.
Palantir 做了二次发行,估值远超 100 亿美元,创业者拿到了钱,还不用经历严苛的审查。我跟他们说过:“听着,从你接受股东、向员工授予期权的那一刻起,你就在一条既定的轨道上。你有义务,如果你不这么想,就该让出这个位置。”这就是我的看法。
Palantir did a secondary way above $10 billion, and so the entrepreneur’s getting paid, and they don’t have to go through the scrutiny. I’ve told them, “Look, the minute you took on shareholders and the minute you granted options to your employees, you were on a course. You have an obligation, and if you don’t feel that way, you should get out of the chair.” That’s how I feel about it.
想象一下,一个大学职业生涯极其出色的四分卫,就在选秀前一周突然说:“我决定不打职业了。”大家会问:“为什么?”他说:“唉,周日的比赛太难熬了。他们得盯着每一档进攻、每一次传球,记录每一项数据。那太可怕了。我没办法打球,在这样的环境下我无法规划我的长期职业生涯……”
Imagine a quarterback who had a great college career saying, just a week before the draft, “You know I’ve decided I’m not going to play.” And they go, “Why?” And he goes, “Well, the scrutiny on Sunday is going to be horrible. They’re going to track every play, every pass. They’re going to record every metric. It’s going to be horrible. I can’t operate. I can’t plan my long-term career with that kind of . . .”
如果真有一个四分卫这么干,肯定会被整个联盟笑话。但有些人现在就在干类似的事。过去三四年简直就是个游乐场。顺便说一句,你们可以去看看那些文章。
If a quarterback did that, he’d be laughed out of the game. But there are people that are acting that way. And so it’s been a playground for the past three or four years. And by the way you can go read these articles.
Evernote 新上任的 CEO 发现,公司居然给每个员工提供两天免费的家政清洁服务。这些福利简直离谱。没人在乎利润,没人在乎。
A new CEO went into Evernote and found out they were giving away two days of house cleaning services to every employee. The perks are off the hook. No one’s looking at profits. No one cares.
问:回到薪酬这个话题,查理·芒格说过一句大致这样的话:“从来没有这么多人,拿这么多钱,却创造这么少价值。”我认为这是我这一生中金融领域最大的变化之一。我刚入行的时候,赚大钱的人都是那些职业生涯末期、展现了诚信和勤奋、并取得成就的合伙人。当然,如今这种以短期业绩为导向的薪酬模式也让我自己的银行账户受益不少。而且我觉得,照您说的,科技行业尤其如此。
Question: To return to the compensation discussion, Charlie Munger said something to the effect of, “Never before have so many that earn so much earned so little.” And I think it’s one of the things that has changed in my lifetime in finance. When I first came in, the people who made a lot of money were the partners at the end of their career who had shown integrity and hard work and developed things. It certainly has benefited my bank account today where pay for performance in the short-term has become a lot more, and I think through your discussion it’s really in the tech sector too.
二十世纪末,很多人都发了大财。我认为我们这个领域需要找回的一点是:社会真正尊重的是那些打造出卓越组织的人,比如杰夫·贝佐斯、沃伦·巴菲特、史蒂夫·乔布斯。但大家反感的是那些恰好出现在对的地方、对的时间,或者玩“我赢了你输”游戏的人。
A lot of people got really rich at the end of the century. I think one of the things in our space that we need to come back to is I think society really respects people who create fantastic organizations: Jeff Bezos, Warren Buffett, Steve Jobs. But I think there’s this resentment of people who came in at the right place, right time, or heads I win, tails you lose.
我们是不是应该把薪酬体系改成十年期支付?让它重新把你和公司绑在一起,让你与公司同呼吸共命运。同时,要设立基准,人们应该对标所在行业来考核。
Don’t we need to move the compensation system to ten-year payouts? Turn it so that you’re partners again in the business, and that you live and die with this business. To that end, you create benchmarks, and people should be benchmarked to that industry space.
在我们这个领域,我认为这些基准已经变得越来越普遍了。所以,我只是对薪酬这个话题感到好奇。
In our space, the benchmarks I think have become much more prevalent. So I was just curious about the question around compensation.
比尔·格利 Benchmark Capital
Bill Gurley Benchmark Capital
比尔·格利:关于这个问题,我想说几点。第一,在非科技行业,我偶尔会看其他人的委托书,我注意到一种趋势:它们开始转向基于业绩的限制性股票单位(RSU)。
Bill: A couple different things I’d say to that. One, in non-tech sectors I’ve noticed from looking at other proxy statements, which I do occasionally, is that they’ve moved to performance-based RSUs.
这事的问题在于,基于业绩的 RSU 有两个缺陷,尽管我认为它们比简单的 RSU 要好。首先,它要求董事会知道什么能创造股票价值,至于这事是否可能,迈克尔可以主持一个讨论会。其次,经常出现的情况是,当业绩指标未达成时,董事会还是会决定照付不误。这是这种方法的两个问题。
The thing about that is there are two problems with performance-based RSUs, although I think they’re better than just simple RSUs. They requires the board to know what creates stock value, and Michael could lead a discussion on whether that’s possible or not. Then second, often when they miss the metrics, you’ll see the board decide to pay it anyway. Those are two problems with that approach.
你在硅谷面临的问题,或者说薪酬体系的普遍问题,是薪酬只能涨不能跌。它基本上是单向上升的。这是个竞争性问题,主要是由谷歌推动的。谷歌现在付给高管的薪酬,我称之为“亚历山大·罗德里格斯级别”的薪水。他们用的叫 GSU(Google Stock Units),但本质跟 RSU 一样。
The problem in Silicon Valley that you have, and maybe it’s a problem with compensation in general, is it’s hard to go backwards. It kind of only goes forwards. And that’s a competitive thing driven mostly by Google, where Google is now paying their top executives what I like to call Alex Rodriguez money. And it’s all what they call GSUs (Google Stock Units), but it’s the same thing.
这基本上相当于每年 1500 万到 2000 万美元的现金支票。我们这些初创公司要和能开出这种薪酬的公司争夺同样的人才。所以我不知道该怎么纠正这个问题。这是硅谷另一个问题的一部分——这些竞争对手手里有太多钱了。
It’s essentially a cash paycheck of 15 to 20 million dollars a year, and our startups compete for the same talent with companies that are paying that out. So I don’t know how you correct it. It’s part of another issue in Silicon Valley with all these competitors having so much money.
我们投了一家叫 Hortonworks 的公司,已经上市了。它的私有竞争对手 Cloudera,有一天融了 9 亿美元。你问,那你能怎么办?整个竞争环境变得一团糟。
We’re in a company called Hortonworks, which is public. Their private competitor, Cloudera, one day raised $900 million. You say, well, what do you do? The playing field gets messy for everybody.
你不能选择不玩。如果不玩,你就等于放弃阵地,失去所有客户。薪酬问题也是如此。我大可以理想主义一点,搞十年期的期权计划,但那结果可能是在硅谷一个员工都招不到。
You can’t choose not to play. If you do, you yield the field and you lose all the customers. The same thing is happening with this compensation issue. I could be idealistic and try and have a ten-year option period, and I might not hire anybody in Silicon Valley.
问:谢谢你,比尔。你最近写了篇文章叫《在资本重组的道路上》,列举了一系列挑战。我有两个相关联的问题。
Question: Thank you, Bill. You recently wrote “On the Road to Recap”, where you listed a bunch of challenges. Two related questions.
第一,你写这篇文章之后,从其他风投或其他方面收到了什么样的反弹或批评?
One, what pushback or criticism have you received on that from VCs or otherwise?
第二,你是否接受你今天说的某些话的前提,比如你想对冲你的独角兽敞口或者做空?除了公开市场操作或者像有些人说的租用场地之外,你有没有考虑过……
Two, do you accept the thrust of the premise in some of the stuff you said today in saying you want to hedge your unicorn exposure or get the other side, other than public market stuff or people who say lease space? Have you guys thought of . . .
比尔·格利:考虑过对冲?
Bill: Thought of hedging?
问:……对冲,或者你是怎么做空的?
Question: . . . Hedging, or how do you take the other side?
比尔·格利:我考虑过。(笑)那篇文章我收到的反馈大部分是正面的。也有一些来自风险投资人的负面反馈,他们正急着按自己纸面上的估值加快募资,这恰好是我文章里指出的一个问题,而且这种状况确实一直在发生。
Bill: I’ve thought about it. [Laughs] I’ve gotten mostly positive feedback on that piece that I wrote. I’ve gotten some negative feedback from VCs that were out trying to accelerate their fundraising on their paper marks which is one of the points that I made, and it’s definitely been going on.
2016 年第一季度,新进风投基金的资金创了历史新高。我甚至听说,有些基金全年的承诺出资额已经用完了,因为他们能分配的额度太多了。所以现在是一窝蜂地募资,这本身就是个信号——他们知道情况不妙。对冲。是啊,我考虑了很久。你可以看看某些有相关敞口的股票。
Q1 of 2016 had record high new LP money into new venture firms. I’ve even heard some of them are done with their commitments for the year because they have so many allocations they can do. So there’s a rush to raise which is just a signal that they know what’s going on. Hedging. Yeah, I’ve thought about it a lot. There are certain stocks that you could look at that have exposure.
比如我关注 Rocket Internet 这个标的。如果我们投的那些公司表现不好,这支股票几乎不可能有好的表现。但我从未真正建仓,我们内部也从未做过对冲。
I look at something like Rocket Internet. There’s very little chance that that stock will do well if our companies don’t do well. But I’ve never actually put anything on, and we’ve never hedged internally.
有些机构已经开始尝试卖出。Founders Fund 卖了一些头寸。就在通用汽车收购 Lyft 的时候,他们卖掉了 Lyft 的头寸,Andreessen Horowitz 也卖了。在我们这个行业,董事会成员通常不会干这种事,但现在已经有人开始这么做了。一旦所有人都开始干,就会加速下跌。其他人有问题吗?
Some firms have started trying to sell. Founders Fund has sold a few positions. They sold a Lyft position and Andreessen sold a Lyft position right when GM was buying. It’s not something that’s typically done in our industry by people that sit on boards, but there are people starting to do it now. The minute everyone starts to do it, it’ll accelerate the fall. Anyone else?
比尔·格利 Benchmark Capital
Bill Gurley Benchmark Capital
迈克尔:我有问题。什么是让你感到兴奋的?我其实看过一个你的访谈,你聊了一点自己在医疗健康领域做的工作。更积极一点来看,未来几年哪些事情让你感到兴奋?
Michael: I do. What’s exciting? I actually saw an interview you did and you talked a little bit about some work you did looking at health care. To be more positive, what kinds of things are exciting to you in the next few years?
比尔·格利:(笑)让我积极一点。我是个乐观主义者。我先说说医疗健康领域吧。有其他投资者在这个领域花的时间比我多得多,我大概花了两年半时间研究它。
Bill: [Laughs] To be more positive. I’m an optimist. Let me just mention the health care sector. So there are other investors who have spent way more time in it than me, and I probably spent two and a half years looking at
我现在才刚刚觉得有信心下注,而且实际上我已经投了几个非常早期的项目。
it. I just now feel comfortable making bets, and I’ve actually made a few very early-stage ones.
这是一个技术尚待发掘巨大价值的领域。我参与过其他一些垂直领域的公司,比如 Zillow、GrubHub、OpenTable 或 Uber。你会想:“天哪,你明明可以用智能手机和所有技术来解决这个问题啊。”但这个行业简直是一团糟。
It’s a segment that’s ripe for technology to add value to. I’ve been involved with other vertical players like Zillow or GrubHub or OpenTable or Uber, and you say, “Boy, you should be able to use this smart phone and all this technology and fix this problem.” It’s a ridiculously messed up industry.
我有一个理论:民主和资本主义,如果放任不管,会互相摧毁;而这种衰败最严重的恰恰是在监管最严密的电信、医疗和金融行业。
I have this theory that democracy and capitalism destroy one another if you give them time, and that decay is happening most in telecom, health care, and finance where you have the most regulation.
现有企业非常善于利用监管来阻挠颠覆——比如 HIPAA(健康保险便携性与责任法案)。人们以为 HIPAA 是在维护他们的最大利益。实际上,HIPAA 是在维护现有企业的最大利益。它让数据共享几乎不可能实现,而你们要构建的所有解决方案都必须依赖数据共享,结果就陷入了这样一片泥沼。
The incumbents are very adept at using the regulation to prevention disruption—things like HIPAA (Health Insurance Portability and Accountability Act). People think HIPAA is looking after their best interests. HIPAA’s looking after the incumbent’s best interests. It makes it almost impossible to share data, and all the solutions you would build to fix the problem require sharing of data, and there’s just all this kind of morass that happens.
我认为每一位创业者采取的创业路径都基于市场化方式,但我们的医疗体系并非市场化运作。为医疗买单的人并不是真正的消费者。我觉得大多数人对《经济复苏法案》一无所知。我们付钱给医生,让他们去实施电子健康档案(EHR)系统,每人付了 4.4 万美元。
I think every entrepreneurial endeavor someone takes on assumes a market-based approach, and our health care system is not a market-based approach. The people paying for it aren’t the buyers. I think most people are ignorant of the Recovery Act. We paid doctors to implement electronic health record (EHR) systems, and we paid them 44 grand each.
这对医生来说就像企业资源计划(ERP)系统。一想到这项决策的无知程度,我就觉得匪夷所思。花 4.4 万美元,不是给最顶尖的 1%,而是给最顶尖的千分之一的人,让他们去装自己根本不想要的软件。他们不想要的原因,是因为他们身处的不是一个需要进化的竞争环境。
It’s like an ERP (Enterprise Resource Planning) for doctors. To think about the ignorance of this decision is just mind-numbing to me. Forty-four thousand dollars to not the top one percent but the top thousandth of a percent to put in software that they didn’t want to put in anyway. And the reason they don’t want to put it in is because they’re not in a competitive environment that requires them to evolve.
然后更蠢的事情发生了,如果“更蠢”是个词的话。你本来最担心的是,付钱让医生装软件,结果他根本不用,对吧?所以两年后,他们又承诺,如果这些医生能证明自己在使用你花 4.4 万美元让他们装的软件,就再给他们 1.7 万美元。而这一切都来自我们联邦政府。他们向医生——这群“可怜的”医生——开出了几十亿、上百亿美元的支票,就为了让他们去安装软件。
Then it gets stupider if that’s a word. The thing you’d worry about is if a doctor was paid to put in something is that he wouldn’t use it. Right? So two years later they give them $17 thousand if they can prove they’re using the software you paid them $44K to use. And this came from our federal government. They wrote billions and billions of dollars of checks to doctors, the downtrodden doctor, to put in software.
顺便提一句,这就是初创公司做好的难度有多大。为了让你的软件有资格获得电子健康档案支付,它必须具备一组特定的功能,而你在网上能找到的 Excel 表格里,政府列出了软件解决方案所需的产品功能。那么,有没有可能以这种方式打造出一个颠覆性的软件解决方案呢?零。而这正是它有多混乱的体现。
By the way, this is how hard it is for a startup to do well. In order for your software to qualify for the EHR payments, it had to have a certain set of features, and there are Excel spreadsheets you can find on the Internet where the government lists the product features required for a software solution. Now is there any chance a disruptive software solution would be built that way? Zero. And that’s how messed up it is.
所以,当系统以那种方式构建时,想建立能自我变革的机制很难。说到激励机制,我对新加坡的医疗体系极为推崇。我们的医疗支出占国内生产总值的 17% 到 18%,而新加坡只有 4%。而且,在任何宽泛的健康指标上,你都找不到差别。他们的做法是,每个人都是付费方。
So it’s just hard to build systems that change when they’re structured that way. I'm enamored with the—talk about incentives—the Singapore health care system. We’re at say 17, 18 percent of GDP (Gross Domestic Product). Singapore’s at four. And on any broad-based health metric you can’t find a difference. And what they do is everyone’s a payer.
富人支付账单的 80% 到 90%,穷人只支付 10% 到 20%。但没有人会在自己不需要的情况下接活。我坦白说,我们得把雇主从这个游戏里踢出去。雇主根本没有理由掺和进来。需要做的事情还有很多,这件事很棘手。
The rich pay 80 or 90 percent of their bill. The poor pay 10 to 20 percent of their bill. But no one’s taking on work that they’re not shopping for. I quite frankly think we need to get the employer out of the game. There’s no reason the employer’s in the business. There’s all kinds of stuff that needs to happen. It’s difficult.
迈克尔:除了医疗健康之外,还有哪些领域更乐观?
Michael: Any other areas that are more positive but not health care?
比尔:有一家我们正在合作的公司,昨天我提过,叫 Stitch Fix,它正在采取一项
Bill: There’s a company that we’re working with called Stitch Fix that I mentioned last night that is taking a
比尔·格利基准资本
Bill Gurley Benchmark Capital
把“点球成金”的数据分析方法用到了女装行业,这原本是你不会想到能行得通的事。每位进店的顾客都要填写一份长达 15 页的个人档案,内容包括她们的尺码、风格偏好、所在地区,以及买衣服主要是为了上班还是外出社交等。
Moneyball approach to women’s fashion, which is something you wouldn’t think would be possible. Each customer that comes in fills in a 15-page profile where they talk about their size, their style, their geographic area, whether they buy clothes more for work or for going out and that kind of thing.
每件进入他们库存的商品,我们都会采集 67 个指标。肩宽多少?柔韧度如何?
Every item that comes into their inventory we collect 67 metrics on. How wide is the shoulder? How pliable is it?
The colors.
The colors.
然后我们开始把产品送到用户手中,看他们留下什么、退回什么。公司里有超过 60 位数据科学家。他们研究的不只是个人行为模式,还包括群体模式,以及这些模式如何应用到不同场景。现在的情况是,负责开发新品的采购人员,在产品还没生产出来之前,就会先用我们的算法做测试。
Then we start sending products to people and looking at what they keep and don’t keep. And there are over 60 data scientists in the company. They study the patterns, not only of an individual but also of groups and how they apply to different things. It’s at the point now where the merchandisers that are creating new products will test against our algorithms before they even build it.
所以,在我看来,这是一种令人振奋的机器学习应用方式,切实可行。其他人都在说,它会给你一个能聊天的人工智能机器人,那确实会很有趣。
So that’s an exciting use to me of machine learning in a way that’s applicable. Everybody else is saying that it will give you an AI bot that you can chat with. It’ll be a lot of fun.
网上流传一个很棒的梗。如果你关注 AI 和聊天机器人,推特上有个很火的 meme:两个完全不相交的圆圈,一个写着“机器人能做的事”,另一个写着“人类需要的事”。[笑声] 两者毫无交集。
There was this great meme. If you’re into AI and chatbots, there’s this great meme going around Twitter of two circles that were completely separate, and one of them said, “Things bots do,” and the other one said, “Things humans need.” [Laughter] There’s no overlap.
提问者:你在最后提到了要关注劳动收入占比、GDP 以及技术专家之类的问题,现在我们看到一些人开始用投票来表达意见,这显然会在一定程度上左右政策走向。你觉得未来五到十年里,这种转变会如何演进?
Question: At the end you mentioned a little bit about paying attention to labor share and GDP and technologists, and we’re kind of seeing some folks vote, and obviously that dictates a little bit of policy. How do you feel that transition evolves over the next five to ten years?
比尔:这个问题远超出了我的权限范围,因为它涉及太多不同方面。我想回到一个事实:短短几百年前,我们人口中大约 98% 是农民,而今天这一比例还不到 1%。我们完成了这种转变,而且我不认为有谁对此感到遗憾。
Bill: It’s a question way above my pay grade because of all the different things it involves. I go back to the fact that a few short couple hundred years ago like 98 percent of our population were farmers, and today it’s less than one. We made that transition, and I don’t think anyone regrets that we made the transition.
自动化将让很多新领域出现这种情况。我不知道解决办法。我认为你无法阻止它,而且如果你试图阻止它,实际上会拖慢全球的进步——不是对受影响的个人而言,而是对所有人而言。
That’s going to happen in a lot of new fields because of automation. I don’t know the solution. I don’t think you can stop it, and I think if you tried to stop it, you would actually slow progress on a global basis—not for the individual that’s impacted but for everyone else.
有一本书我极其推崇,叫《理性乐观派》,马特·里德利在书中谈到生活水平的最大提升,通常都围绕着创新、思想的分享以及开放的资本主义。
There’s a book that I’m a huge fan of called The Rational Optimist where Matt Ridley talks about the biggest increases in the standard of living. They’re typically around innovation and the sharing of ideas and open capitalism.
中国过去 20 年的开放对生活水平的影响,可能是最大的。所以,技术变革放缓,我不认为对全球生活水平有好处,但可能对个体的处境有利。不过反过来说,这可能又拖累了其他所有人的生活水平。所以这是个难题。
China’s unlocking for the past 20 years has probably been the biggest impact to standard of living. So slowing technological change I don’t think helps the global standard of living, but it might help an individual’s situation. But again that might slow it for everybody else. So it’s a difficult problem.
有一件事我们大家都能做的——实际上西雅图有一群企业家成立了一个非营利组织,我记得叫 Code.org——就是在各地推广编程。[笑]
One thing that we could all do—and there’s actually a group of entrepreneurs out of Seattle that have created a nonprofit, I think it’s called Code.org—is just promote programming everywhere. [Laughs]
现在硅谷公司招大学毕业生工程师,起薪就是 17.5 万美元,而职位供给严重不足。中国大概有 35% 的学生学工程,我们只有 5% 左右。
There are engineers being hired into Silicon Valley companies out of university at $175K right now, and there is a complete undersupply of jobs. China puts 35 percent I think of students into engineering, and we put five or something.
我们社会和文化层面形成了一种偏见,那种“书呆子”的刻板印象。这对我们的国家来说是个实实在在的问题,对这件事更是如此,所以要尽一切努力把编程推广到初中、高中这些阶段。
We’ve created a social and cultural bias, the whole nerd thing. That’s a real problem for our country and for this issue, and so do everything you can to get coding pushed into middle schools, high schools, that kind of thing.
迈克尔:我想我们就到此为止吧。非常感谢你,比尔。很棒。
Michael: I think we’ll call it there. Thank you very much, Bill. Great.
华盛顿大学 佩德罗·多明戈斯
Pedro Domingos University of Washington
佩德罗·多明戈斯是全球机器学习、人工智能和大数据领域的顶尖专家之一。他是西雅图华盛顿大学的计算机科学教授,也是《终极算法:机器学习和人工智能如何重塑世界》一书的作者。他是 SIGKDD 创新奖(数据科学领域的最高荣誉)的获得者,同时也是美国人工智能协会(AAAI)的会士。他还曾获得富布赖特奖学金、斯隆研究奖、美国国家科学基金会杰出青年学者奖(CAREER Award),以及多项最佳论文奖。
Pedro Domingos is one of the world’s leading experts in machine learning, artificial intelligence, and big data. He is a professor of computer science at the University of Washington in Seattle and the author of The Master Algorithm: How the Quest for the Ultimate Learning Machine Will Remake Our World. He is a winner of the SIGKDD Innovation Award—the highest honor in data science—and a Fellow of the Association for the Advancement of Artificial Intelligence. He has received a Fulbright Scholarship, a Sloan Fellowship, the National Science Foundation’s CAREER Award, and numerous best paper awards.
佩德罗是 200 多篇研究论文的作者或合著者,并在会议、大学和研究机构发表过 150 多场特邀演讲。他于 1997 年在加州大学尔湾分校获得博士学位,并于 2001 年共同创立了国际机器学习学会。他曾在斯坦福大学、卡内基梅隆大学和麻省理工学院担任访问学者。他的研究涵盖广泛的主题,包括将学习算法扩展到大数据、最大化社交网络中的口碑效应、统一逻辑与概率,以及深度学习。
Pedro is the author or co-author of over 200 research publications, and has given over 150 invited talks at conferences, universities, and research labs. He received his PhD from the University of California at Irvine in 1997 and co-founded the International Machine Learning Society in 2001. He has held visiting positions at Stanford, Carnegie Mellon, and MIT. His research spans a wide variety of topics, including scaling learning algorithms to big data, maximizing word of mouth in social networks, unifying logic and probability, and deep learning.
佩德罗·多明戈斯 华盛顿大学
Pedro Domingos University of Washington
迈克尔·莫布森:我非常高兴地介绍下一位演讲者,佩德罗·多明戈斯。他是华盛顿大学的计算机科学教授,也是机器学习、人工智能和大数据领域的顶尖专家。
Michael Mauboussin: I’m thrilled to introduce our next speaker, Pedro Domingos. Pedro is a professor of computer science at the University of Washington and a leading expert in machine learning, artificial intelligence, and big data.
去年秋天,我在伦敦与一系列资深投资经理进行了会面。令我印象深刻的是,在每一场谈话中,机器学习这个主题都会毫无征兆地出现。我决心进一步了解它,于是读了佩德罗的著作《终极算法》。我发现这本书在多个层面上都很有用且富有启发性,并强烈推荐。顺便提一下,书中最有价值的页码是 240 页。记下来。过一会儿你们还会听到更多关于这本书的内容。
Last fall, I had a series of meetings with senior investment managers in London. What struck me was that in every conversation, the topic of machine learning came up unsolicited. Determined to learn more about it, I read Pedro’s book, The Master Algorithm. I found the book useful and illuminating on multiple levels, and recommend it highly. By the way, the money page is 240. Make a note of that. You’ll hear more about the book in a few moments.
正如我刚刚所提到的,你可以听佩德罗讲讲机器学习的基础知识,并了解这些能力可能将我们带向何方。但你也可以从聆听中理解,机器学习中的各种方法与挑战,如何应用于日常思考。机器学习和人工智能代表着一种获取知识的全新且非常激动人心的方式。
As I mentioned a moment ago, you can listen to Pedro for a primer on machine learning and to get a sense of where these capabilities may take us. But you can also listen to understand how the various approaches and challenges in machine learning apply to everyday thinking. Machine learning and AI represent a new and very exciting means of attaining knowledge.
请大家和我一起欢迎佩德罗·多明戈斯教授。
Please join me in welcoming Professor Pedro Domingos.
佩德罗·多明戈斯:让我从一个问题开始:知识从哪里来?直到不久以前,知识只来自三个来源。第一个是进化。那是包含在基因里的知识。第二个知识来源是经验。那是包含在神经元里的知识。第三个知识来源是文化。那是我们通过与人交谈、阅读书籍等方式获得的知识。
Pedro Domingos: Let me start with a question: where does knowledge come from? Until recently, knowledge came from just three sources. The first one is evolution. That’s the knowledge that’s included in your genes. The second source of knowledge is experience. That’s the knowledge that’s included in your neurons. And the third source of knowledge is culture. It’s the knowledge that we acquire by talking with other people, reading books, and so on.
过去短短几十年出现的新情况是,地球上多了一个全新的知识来源——机器学习。这是计算机的力量。计算机正从数据中发掘新知识。
Now, what’s new in just the last few decades is that there is a new source of knowledge on the planet, and that’s machine learning. It’s computers. Computers are discovering new knowledge from data.
我发现,这些新知识发现方式的每一次出现,都是地球生命史上的重要里程碑。我的意思是,进化就是地球生命本身。从经验中学习,是哺乳动物区别于昆虫的关键;而文化,则是人类取得如此成功的原因——它造就了我们是谁。
I noticed that the emergence of each of these new ways of discovering knowledge was a major landmark in the history of life on Earth. I mean, evolution is life on Earth itself. Learning from experience is what distinguishes mammals from insects, and culture is what makes humans as successful as they are. It’s what makes us who we are.
我认为计算机作为知识的来源,其重要性将与上述这三者中的每一个不相上下,而我们才刚刚起步。我们已经看到了许多影响。还要注意,这些新的知识来源在最初出现在地球上时,其运行速度比之前的来源快了好几个数量级。因此,从经验中学习比从进化中学习快了好几个数量级,而通过听别人重复告诉你某些东西来从文化中学习,又比从经验中学习快得多。
I think that computers as a source of knowledge are going to be every bit as momentous as every one of these three, and we are just getting started. Already, we see a lot of the impact. Notice also that each of these new sources of knowledge, operated orders of magnitude faster than the previous sources when they first appeared on the planet. So learning from experience is orders of magnitude faster than learning from evolution, and learning from culture by just hearing something that somebody tells you again is a lot faster than learning from experience.
计算机带来的学习速度会更快。计算机发现知识的速度是人类无法想象的。伴随这种更高的速度,每一种新方法所发现的知识量,也比以往高出好几个数量级。
And learning from computers is going to be even faster. Computers can discover knowledge at a rate that is unimaginable for human beings. Corresponding to that greater speed, you also discover orders of magnitude more knowledge with each of these new ways than you did previously.
事实上,杨立昆(Yann LeCun),这位知名机器学习研究者、现任 Facebook 人工智能研究总监,曾表示在未来,世界上大部分知识将由计算机发现,并驻留在计算机之中。
In fact, Yann LeCun, who is a well-known machine learning researcher and now the Director of AI Research at Facebook, says that in the future, most of the knowledge in the world will be discovered by computers and will reside in computers.
所以我认为,我们所有人都到了这样一个时刻:不一定需要了解机器学习的细节,但在概念上,需要明白机器学习是什么、它能做什么。这正是我在这场演讲中试图做的事情。
So I think we’re at a point where all of us need to understand, not at a necessarily very detailed level, but conceptually, what machine learning is and what it does. That’s what I am going to try to do in this talk.
以下是一张幻灯片上的机器学习。在传统的编程中,信息时代其实经历了两个阶段。第一阶段,我们编程让计算机做事。当我们想让计算机做某件事时,必须事无巨细地写下一个算法,解释清楚计算机该如何执行。
Here is machine learning in one slide. In traditional programming, there have really been two stages in the Information Age. The first stage was where we programmed computers to do things. When we want a computer to do something, we have to write down an algorithm in painstaking detail explaining how that computer is
华盛顿大学的佩德罗·多明戈斯
Pedro Domingos University of Washington
应该这样做。
supposed to do it.
然后,情况是这样的:计算机接收数据,输入算法,算法对数据进行处理来产生输出。例如,输入一张 X 光片,输入一个诊断癌症的算法,然后它说“哦,这里有个肿瘤”,输出就是“是,这里有肿瘤”或“否,没有肿瘤”。直到最近,信息时代的一切都是这样构建的。
And then, this is what happens: there is the computer, in goes the data, in goes the algorithm, and the algorithm does something to the data to produce the output. For example, in goes an X-ray, in goes an algorithm to diagnose cancer. Then it says oh, there’s a tumor here, and the output is either, yes, there is a tumor here, or no, there isn’t one. This is how everything in the Information Age has been built until recently.
机器学习把这个逻辑颠倒了过来。机器学习中发生的事情乍看非常奇怪,但其实很有道理——希望听完这场演讲后,你也会觉得有道理。输出现在变成了输入,而最终产出的则是算法。我们提供给计算机的,是我们希望它处理的数据,以及我们希望它产生的输出结果。计算机自己摸索出如何把这些数据转化为这个输出,然后以算法的形式把它呈现出来。
Machine learning turns this around. What happens in machine learning is something that is very strange at first sight but makes a lot of sense. Hopefully it will to you after this talk. The output is actually now going in, and what comes out is the algorithm. So what we give to the computer is the data that we want it to operate on and the output that we would like it to produce. The computer figures out how to turn this data into this output, and then it produces that in the form of an algorithm.
然后那个算法就开始做它平常的工作。举个例子,这里可能有一堆 X 光片,而这是每张 X 光片对应的诊断结果。诊断结论说,是的,这里有个肿瘤,或者不,这里没有肿瘤,而机器学习通过观察图像中的像素,来学会如何判断到底有没有肿瘤。这个过程不断重复,然后一大批新病人来了,算法就开始自动诊断,成本只有人类病理学家的一个零头,效果还更好。一个只需要半小时就能学会做这件事的算法,做得比上了好多年医学院的人还要好。
Then that algorithm goes and does its usual job. For example, this might be a bunch of X-rays and this might be the diagnosis for each of those X-rays. The diagnosis says yes, there was a tumor here, or no, there wasn’t a tumor here, and the machine learning figures out how to decide whether there is a tumor or not by looking at the pixels in the image. That goes on and on, and then a bunch of new patients come in and it starts doing this automatically, at a fraction of the cost of a human pathologist and better. An algorithm that learns to do this in half-an-hour does better than someone who was in med school for many years.
令人惊叹的是,在旧的方式中,每当你想要做一件不同的事情,你都得编写新的算法。如果你想让计算机做医疗诊断,你得把它编程成会做医疗诊断。如果你想让计算机开车,你得把它编程成会开车。
The amazing thing is that in the old way of doing things, for every different thing that you wanted to do, you needed to write the new algorithm. If you wanted the computer to do the medical diagnosis, you had to program it to do the medical diagnosis. If you wanted the computer to drive a car, you had to program it to drive a car.
但在机器学习领域,同一个学习算法根据你提供给它的数据,可以完成无穷无尽的不同任务。学习算法堪称一种元算法,因为它是一种能够生成其他算法的算法。
But in machine learning, the same learning algorithm can do an infinite array of different things depending on what data you give to it. A learning algorithm is a master algorithm in the sense that it’s an algorithm that makes other algorithms.
原则上,同样的学习算法,只要足够强大,就能学会你希望它学会的任何东西——前提是你给它足够多、足够正确的数据。那么,这究竟是如何发生的呢?其实,机器学习中有多种不同的范式,多种学习新知识、从数据中提取程序和模型的方法。
In principle, the same learning algorithm, if it’s powerful enough, can learn absolutely anything that you want it to learn provided you give it a sufficient amount of the right data. So how does this actually happen? Well, there are a number of different paradigms in machine learning, a number of different ways of learning new knowledge, of extracting programs and models from data.
另外一点——机器学习,除了当今极具实用价值和经济重要性之外,也充满乐趣、非常迷人,因为机器学习的核心思想全都来自不同的学科领域。
One other thing—machine learning, in addition to being very useful and economically important these days, is also a lot of fun and very fascinating because the main ideas in machine learning all come from different fields.
事实上,机器学习主要有五个思想流派。每个流派都源于不同的学科领域,各自有一套独特的做事方式。每个流派都有自己专属的“主算法”:这种算法,原则上只要你提供某个问题的数据,它就能学会针对那个特定问题去完成所需的任务。
In fact, there are five main schools of thought in machine learning. Each one of them has its roots in a different field, and each one has its own version of how to do things. Each one has its own master algorithm: an algorithm that if you give it data from any problem, in principle, it can then learn to do what needs to be done for that particular problem.
我们要看的第一族是符号主义者,他们的根源在逻辑与哲学。在这五族之中,他们与计算机科学关联最紧密,他们的主算法是一种叫做逆演绎的东西,该算法将归纳视为演绎的逆过程。
The first tribe that we’re going to look at are the symbolists and their origins are in logic and philosophy. They are the most linked to computer science of the five tribes, and their master algorithm is something called inverse deduction, which sees induction as being the inverse of deduction.
然后是联结主义者,他们的想法是我们要通过逆向工程破解人脑来学习。你的大脑是地球上最伟大的学习机器,所以我们要弄清楚它如何运作,然后在计算机上实现它。他们被称为联结主义者,因为这全都基于一个理念:你的知识蕴含在神经元之间的连接中,而他们的主算法是一种叫做反向传播的东西。
Then there are the connectionists whose idea is that we’re going to learn by reverse-engineering the human brain. Your brain is the greatest learning machine on Earth, so let’s figure out how it works and do that on the computer. They’re called connectionists because it’s all based on this idea that your knowledge is included in the connections between your neurons, and their master algorithm is something called backpropagation.
然后是那些进化论者,他们说主算法并非人类的大脑,而是进化本身。他们
Then there are the evolutionaries who say that the master algorithm is not your brain, but rather evolution. They
华盛顿大学的佩德罗·多明戈斯
Pedro Domingos University of Washington
我想弄清楚进化是如何运作的,并在计算机上模拟这个过程,而他们的核心算法就是遗传编程。
want to figure out how evolution works and simulate that on the computer, and their master algorithm is genetic programming.
还有一类是贝叶斯学派,其根源在统计学领域。他们最关心的是所学知识的不确定性。所有从数据中习得的知识——这是归纳过程的一部分——都必然带有不确定性。因此,贝叶斯学派的核心思想是将这种不确定性量化,他们使用的核心算法是概率推理。这种方法根据你掌握的证据,计算不同假设成立的概率。
Then there are the Bayesians who have their origins in statistics, and their biggest concern is with the uncertainty of learned knowledge. All knowledge that is learned from data that is the part of induction is necessarily uncertain, so the idea of Bayesians is to quantify that uncertainty, and so, their master algorithm is probabilistic inference. It’s computing the probabilities of different hypotheses based on the evidence that you have.
最后还有一类人,他们的根基深植于多个不同领域——类比派。这些领域中最重要的,大概要数心理学。他们的核心理念是:人类大部分的学习与推理,都是通过类比完成的。也就是说,我们会寻找与自己当前处境相似的场景,然后尝试从一个场景推演到另一个场景。这类人使用最广泛的算法,是一种叫作“核机器”的方法,也叫支持向量机。
And finally, there are the analogizers who actually have their roots in many different fields. The most important of these fields is probably psychology. The idea here is that most of the learning and reasoning that we do is by analogy. It’s by finding similar situations to the ones that we are in now, and then trying to extrapolate from one to the other. Their most widely used algorithm is something called a kernel machine, also known as a support vector machine.
那么,我们先走进符号主义学派,看看他们提出了什么。以下是全球最著名的几位符号主义者:卡内基梅隆大学的汤姆·米切尔(Tom Mitchell)、英国的史蒂夫·马格尔顿(Steve Muggleton),以及澳大利亚的罗斯·昆兰(Ross Quinlan)。
So let’s start by visiting the symbolists and seeing what they have to propose. Here are some of the most prominent symbolists in the world: Tom Mitchell at Carnegie Mellon, Steve Muggleton in the U.K., and Ross Quinlan in Australia.
这种学习方式背后的核心理念——学习即归纳——实际上可以追溯到 19 世纪一位名叫威廉·杰文斯的哲学家兼经济学家。归纳是从具体事实推导出一般规则的过程,而与之相对的是演绎,即从一般规则推导出具体事实。
The basic idea behind this type of learning, that learning is induction, actually goes back to a 19th century philosopher and economist named William Jevons. Induction is going from specific facts to general rules whereas the inverse is deduction, which is going from general rules to specific facts.
我们有些人已经懂得如何进行归纳推理——就像数学家当年发现减法是加法的逆运算、积分是微分的逆运算那样,等等。
Some of us can figure out how to do induction in the same way that mathematicians, for example, figured out how to do subtraction because it was the inverse of addition, or how to do integration because it’s the inverse of differentiation, and so on.
在数学领域,这一点的历史源远流长且成就斐然。例如,加法能回答“二加二等于几”这个问题。答案当然是四。这可不是我在这场演讲中要讲的最深奥的内容。
In mathematics, this has a very long and distinguished history. So for example, addition gives us the answer to the question what is two plus two? It’s four, of course. That’s not the deepest thing I’m going to say in this talk.
减法给出了那个逆向问题的答案,那就是:在 2 的基础上需要加多少才能得到 4?所以,我对逆向推演的理解,其实就是做同样的事,只不过用的是归纳法。
Subtraction gives us the answer to the inverse question, which is what do I need to add to two in order to get to four? And so, my idea of inverse deduction is actually to do the same thing but with induction.
例如,演绎推理能帮我们回答这样一个问题:“如果我已知苏格拉底是人,而人终有一死,那么我能推断出关于苏格拉底的什么结论?”答案自然是:苏格拉底会死。人类非常擅长这种推理。在逻辑学、计算机科学和哲学领域,演绎推理几十甚至几百年来都已被研究得十分透彻。
For example, deduction gives us the answer to a question such as, “If I know that Socrates is human and that humans are mortal, than what can I infer about Socrates?” And of course, it’s that Socrates is mortal. We know how to do this very well. Deduction has been very well understood in logic and computer science and philosophy for decades or even centuries.
现在,棘手的问题在于归纳法。归纳法要回答的问题是:“如果我知道苏格拉底是人,还需要知道什么才能推断出他是必死的?”答案当然就是:人终有一死。
Now, the tricky problem is induction. Induction is the answer to the question, “If I know that Socrates is human, what else do I need to know in order to be able to infer that he’s mortal?” The answer, of course, is that humans are mortal.
但你得自己去琢磨这个。要是你能填上这个缺口,那你就有了一个新的一般性规则,可以和其他规则组合起来,应用到很多其他事情上,去回答那些你也许从未想过的问题。
But you try and figure that out. If you can fill in this gap, now you have a new general rule that you can go and apply to many other things in combination with other rules to answer questions that you may have never thought of.
通过这种用逆向演绎填补知识空白的过程,你就能构建起一套非常强大的规则知识库。尤其值得一提的是,你可以用不同方式组合不同规则——这种能力只有符号主义者才能实现。机器学习领域其他学派都不具备这一能力,而它恰恰至关重要。
By this process of filling in the gaps in your knowledge using inverse deduction, you build up a knowledge base of rules that is very powerful. In particular, this idea that you can combine different rules in different ways is something that only the symbolists can achieve. None of the other schools of machine learning actually have that capability and it’s a very important one.
当然,我以上全部是用英文写的。当然,计算机不理解自然语言,所以在计算机里,这实际上通常是用形式化语言来做的,比如一阶逻辑。但思路是一样的。
Now, I wrote all of this in English. Of course, computers don’t understand natural language, so in a computer, this is actually usually done using a formal language, such as first order logic. But the idea is the same.
华盛顿大学的佩德罗·多明戈斯
Pedro Domingos University of Washington
象征主义学习有点像模仿科学方法。它的意思是:这里有一些数据,让我先形成一些假设来解释这些数据,然后在新数据上检验这些假设,接着可能会修正或抛弃它们,如此循环往复。
Symbolist learning is a bit like emulating the scientific method. It’s saying: here’s some data, let me formulate some hypotheses to explain that data, and then test those hypotheses against new data, and then maybe refine them or discard them, and so on.
我们在这里真正做的事情,其实就是把科学方法自动化——只是我们做得更快、规模更大而已。一个具体的例子就是你在图上看到的这一幕。照片里的生物学家其实不是那个穿白大褂的人,穿白大褂的是一位名叫罗斯·金的机器学习研究员。
What we’re really doing here is just automating the scientific method, except we’re doing it much faster and on a much larger scale. And one concrete instantiation of this is what you see in this picture. The biologist in this picture is actually not the guy in the lab coat. The guy in the lab coat is a machine learning researcher by the name of Ross King.
这张照片里的生物学家其实就是这台机器。这台机器人是一台基于逆推原理、装在一个箱子里的完整生物学家。他们最初用一台名叫“亚当”的机器人起步,现在这台机器人的名字叫“伊芙”。目前全世界只有这一台,但有意思的是,一旦你拥有了这样的机器人,就没有什么能阻止你制造出数百万台,而它能让科学进步的速度加快数百万倍。
The biologist in this picture is actually this machine. This machine is a complete robot biologist in a box based on the principle of inverse deduction. They started out with a robot named Adam, and now the name of this robot is Eve. There is only this one in the world so far, but of course, what’s interesting is that once you have a robot like this, nothing stops you from making millions of them, and it can make progress in science millions of times faster.
例如,伊芙(Eve)会观察某种特定细胞类型的生物学特性。它从数据出发,通过逆向推理提出假设来解释这些数据,然后设计实验来检验这些假设,接着利用基因测序仪和 DNA 微阵列来完成实验——这正是图中展示的过程。之后,它会重复这一流程。所以,它确实是一个完整的机器人科学家。2014 年,伊芙发现了一种新的疟疾药物。这就是用这类学习方式能够做到的事情。
Eve looks at, for example, the biology of a particular cell type. It starts out with data, formulates a hypothesis to explain that data by inverse deduction, then it designs experiments to test those hypotheses, and then it carries out the experiments using gene sequencers and DNA microarrays, which is what’s going on here. And then, it repeats the process. So it really is a complete robot scientist. In 2014, Eve discovered a new malaria drug. So this is the kind of thing that you can do with this type of learning.
现在,联结主义者对这些观点持怀疑态度。他们说这种学习方式过于抽象、过于干净。大多数学习发生的方式,并不像科学家、逻辑学家、甚至哲学家那样运作——它要凌乱得多。
Now, the connectionists are skeptical about all of this. They say this type of learning is too abstract, too clean. The way most learning happens is not the way a scientist or a logician or even a philosopher works. It’s messier.
这涉及犯错、有身体感知,以及诸如此类的一切。联结主义者的观点是,我们在机器学习领域的竞争对手是人类大脑,而我们远远落后于对手。
It involves making mistakes and being embodied and all sorts of things like that. The idea of the connectionist is that our competition in machine learning is the human brain, and we are far behind the competition.
所以科技行业里,当你落后于竞争对手时,你做的事情就是逆向工程。你从复制开始。
So what you do in tech when you’re behind the competition is reverse engineering. You start out by copying it.
你拆开芯片,就能看到电路结构,然后想办法做出一样的东西。联结主义者对大脑也试图这么干。大脑藏在头骨里,里面有各种线路之类的东西,我们可以试着弄清楚它的工作原理。确实,这一方法对机器学习来说一直极为有效。
You open up their chip and you see what the circuit is, and you figure out how to do the same thing. The connectionists try to do that with the brain. The brain is inside the skull and it’s got circuits and whatnot, and we can try to figure out how it works. And this has indeed been a very productive approach to machine learning.
世界上最著名的联结主义者是杰夫·辛顿。他 70 年代其实是从心理学起家的,如今更像是计算机科学家。他现在一半时间在多伦多大学,一半在谷歌。
The most famous connectionist in the world is Geoff Hinton. He actually started out as a psychologist in the ’70s, and these days, he’s more of a computer scientist. He actually splits his time between the University of Toronto and Google.
杰夫相信大脑的学习方式可以归结为单一的算法,过去 40 年里,他一直在试图发现这个算法。实际上,他讲过这样一个故事:有一天他兴冲冲地回到家,喊道:“我成功了!我搞清楚大脑是怎么工作的了!”他女儿回答说:“哦,爸爸,又来了一次。”(笑声)
Jeff believes that the way the brain learns can be captured in a single algorithm, and he has spent the last 40 years trying to discover that algorithm. In fact, he tells the story of coming home from work one day very excited saying, “I did it! I figured out how the brain works!” And his daughter replied, “Oh, Dad, not again.” [Laughter]
他率先坦言自己经历过起起落落,但他的探索正开始见到成效。尤为关键的是,他是反向传播算法的发明人之一,该算法如今无处不在。比如,如果你用安卓手机,它就在做语音识别。反向传播的用途五花八门,确实令人难以置信。在 80 年代,它的一个杀手级应用是在金融领域,预测股票波动、外汇波动之类。另外两位突出的联结主义者是,我已经提到过的杨立昆,还有约舒亚·本吉奥。
Now, he is the first to say that he’s had his ups and downs, but his quest is starting to pay off. In particular, he is one of the inventors of this backpropagation algorithm which is used everywhere. Like, if you have an Android phone, for example, it’s what’s doing the speech recognition, and the variety of things that backprop is used for is truly mindboggling. One of its killer applications in the ’80s was in finance, predicting stock fluctuations and foreign exchange fluctuations and whatnot. Two other prominent connectionists are Yann LeCun, who I already mentioned, and Yoshua Bengio.
那么,我们来简单看看这一切是如何运作的。生物学家大致了解神经元的工作方式,而要想实现我们的目标,我们并不需要比这更精确的知识。神经元是一种非常有趣、极其独特的细胞;它看起来像一棵微观的树。主干被称为轴突,枝条被称为树突,根部也被称为树突。
So let’s see in a nutshell how this all works. Biologists know roughly how neurons work, and we don’t need to know more than roughly how they’ll work in order to do what we want. A neuron is a very interesting, it’s a very unique type of cell; it’s a cell that looks like a microscopic tree. The trunk is called the axon, the branches are called dendrites, and the roots are also called dendrites.
佩德罗·多明戈斯 华盛顿大学
Pedro Domingos University of Washington
神经元有趣且与树木不同的地方在于:一个神经元的枝突会与其他神经元的树突产生接触。这些接触点被称为突触,而突触的效能可强可弱。
The thing that’s interesting about neurons that makes them different from trees is that the branches of one neuron make contact with the dendrites of other neurons. The points of contact are called synapses, and the synapses can be more or less efficient.
神经元会在其胞体(soma)中积蓄电荷。如果电荷超过阈值,神经元就会沿轴突发放所谓的动作电位——毫不夸张地说,就是一道小闪电。此刻你正在工作的大脑,正是这场遍布四处的小闪电交响曲。
Neurons build up a charge in their body, or soma. If the charge exceeds the threshold, then the neuron fires what is called an action potential down the axon, which is quite literally a little lightning bolt. Your brain at work right now is a symphony of these little lightning bolts going all over the place.
接着,电荷沿着轴突传递,到达突触,效率更高的突触能通过一个化学过程更好地传递电荷,这个过程现在不必深究,但关键是神经科学家认为,你学过的所有东西,都编码在突触的强弱之中。
Then the charge goes down the axon and it goes to the synapses, and the more efficient synapses transmit the charge better using a chemical process that is not important to discuss now, but the point is neuroscientists believe that everything you’ve ever learned is encoded in how strong the synapses are.
粗略来说,当两个神经元同时放电时,它们之间的突触会变得更强——这意味着下一次,第一个神经元会更轻松地激发第二个神经元。于是,我们要建立一个神经元的数学模型,在计算机上运行它,构建一个由这些神经元组成的大型网络,这就是我们学习事物的方式。
Roughly speaking, when two neurons fire together, the synapse between them grows stronger, meaning that next time around, the first neuron will have an easier time firing the second neuron. And so, we’re going to build a mathematical model of a neuron, run it on the computer, build a big network of these neurons, and this is how we’re going to learn things.
以下是我们构建的一个神经元数学模型。注意,图中的每个模块都与之前展示的神经元图像一一对应:这是细胞体,这些是输入的树突,而这是轴突。那么,假设我们有一个神经元——这里就是你的视网膜——输入就是像素点。一般来说,中间可能还有其他神经元层,但我们现在就假定输入是像素点。现在,每个输入都会被乘以一个权重。有些输入的权重会比另一些更大,而这些权重正是学习发生的地方。神经网络中的学习,本质上就是在不断调整这些权重。
Here’s our mathematical model of a neuron. Notice that there’s a one-to-one correspondence between the blocks in this diagram and the image of the neuron that I had before. This is the cell body, these are the dendrites coming in, and this is the axon. So let’s suppose that we have a neuron, and this is your retina right here. The inputs are just pixels. In general, there could be other layers of neurons, but let’s say they’re pixels. Now, each one gets multiplied by a weight. Some will get multiplied by a larger weight than others, and these weights are where the learning is going to happen. Learning in neural networks is basically just twiddling those weights.
如果这些加权输入的总和超过阈值,输出就是 1。假设我正在看一张猫的图片,而神经元正在正常工作。某些属于猫的特征超过了阈值,因此神经元激活。否则,神经元会说,嗯,不,这实际上看起来不像猫,于是神经元不激活。
If the sum of these weighted inputs exceeds the threshold, then the output is one. Let’s say I’m looking at an image of a cat and the neuron is doing its job. Some of the features, which are features of cats, exceed the threshold so the neuron fires. Otherwise, it says, well, no, this actually doesn’t look like a cat, and the neuron doesn’t fire.
这部分很容易。真正有意思的地方在于,当我们拥有这样一个庞大的神经网络时,该怎么做。我们如何训练它?仔细想想,这是个非常棘手的问题,根本没有显而易见的答案。我有一个巨大的神经网络,某个小神经元在这里带有一个权重,而误差发生在输出端。网络现在说这是一只猫,但它并不是猫。谁错了?哪些权重需要调整?
This part is easy. When things get interesting is when we have a big network of neurons like this. How do we train it? If you think about it, this is a very hard problem with no obvious answer. I have a huge network of neurons and there is some little neuron here with some weight and the error is happening at the output. The network is saying that this is a cat but it’s not a cat. Who is wrong? What weights need to change?
人们最早在 50 年代就提出了神经网络的概念,但当时不知道如何解决这个问题,所以这事就渐渐无人问津了。
People first thought of neural networks in the ’50s, but they didn’t know how to solve this problem, and so, things kind of died out.
但在 80 年代,他们想出了这种反向传播算法,本质上是一种解决该问题的方法。
But in the ’80s, they figured out this backpropagation algorithm, which is in essence a way to solve this problem.
反向传播的工作方式在概念上非常简单。它实际所做的就是依次微调每一个权重,问自己:如果我把这个权重调大,输出端的误差是会减少还是不会?
The way backpropagation works is conceptually very simple. All it’s really doing is tweaking each weight in turn saying if I increase this weight, will the error at the output go down or not?
假设这实际上是 1,本应输出 1,但它只输出了 0.3。那么误差就是 0.7,我需要减小这个误差,让输出更高。如果我改变这里的这个权重,让它稍微增大一点,它会让我的误差减小吗?又或者,如果这个权重稍微减小一点,反而会让误差减小。
Let’s say like this was a cat, this should have been firing, it should have been one, but it was just 0.3. So the error is 0.7, and I need to reduce that error. I need to make the output higher. If I change this weight over here, if it goes up a little bit, does that make my error go down? Or maybe if this weight goes down a little bit, that will make the error go down.
当然,一次只处理一个权重的方式显然效率极低,所以反向传播算法采取的是分层推进的策略。这是我的输入层,这些紫色圆圈就是神经元。每个神经元先算出自己的数值,然后下一层的神经元就可以基于这些数值一路计算下去,直到输出层。当
Now, of course, doing it like this one weight at a time would be ridiculously inefficient, so what backpropagation does is handle things in layers. So here is my input, and these purple circles are the neurons. Each neuron computes its value and then the next layer of neurons can compute their values all the way to the output. When
佩德罗·多明戈斯 华盛顿大学
Pedro Domingos University of Washington
我们拿到输出后,把它和应有的结果做比较:你说它是一只猫,它到底是不是一只猫?
we get the output we compare it with what it should have been: did you say that it was a cat and was it a cat?
然后我们观察偏差的幅度,判断每个权重需要调整多少。具体做法是,我将误差(即那些黄色圆圈)从后往前逐层传播,并计算每个权重需要变动的数值。
And then we see by what amount we were wrong and determine how much the weight should change. So what I do is I propagate the errors, that’s the yellow circles, backwards through the layers, and I compute how much each weight needs to change.
首先,我计算出这些权重需要调整多少,然后根据这里的误差,一路往回计算出所有这些权重需要调整多少,直到网络的起点。也就是说,我通过网络反向传播误差,从而决定每个权重需要调整多少,这就是为什么这个算法叫做误差反向传播,或者简称为反向传播。
First I compute how much these weights need to change, and then based on the errors here, I compute how much these weights need to change all the way back to the beginning. So I’m propagating the errors back through the network in order to decide how much the weights need to change, and that’s why this algorithm is called error backpropagation, or backprop for short.
反向传播被证明是一种极其强大的方法,就在过去短短几年里,它已经在计算机视觉、物体识别、视频理解以及语音理解等领域引发了一场革命。
Backprop turns out to be an incredibly powerful way of doing things, and just in the last few years, it’s caused a revolution in things like, for example, computer vision, object recognition, video understanding, and speech understanding.
微软有一套系统——更准确地说,是在 Skype 里——能为你做实时翻译。你可以在电话里用英语和中国的某人通话,而对方听到的是中文,反过来也一样。这套系统是靠一堆神经网络用这类学习方式实现的。
Microsoft has a system, in Skype to be more precise, where it does simultaneous translation for you. You can be speaking English on the phone with someone in China and they’re hearing Chinese, and vice versa. This is done by a bunch of neural networks using this type of learning.
谷歌、微软、亚马逊和 Facebook 这类公司,利用这些类型的网络不仅用于物体和视频识别,还用来挑选搜索结果、选择向你展示的广告,等等。近年来,在媒体上这常被称为深度学习。
Companies like Google, Microsoft, Amazon, and Facebook use these types of networks not just for object and video recognition but also to choose search results, to choose ads to show you and whatnot. In the press, this is often called deep learning these days.
为什么它被称为深度学习?因为它是在训练拥有多个层级的网络。在 80 年代,人们搞明白了反向传播算法,但当时他们确实没法训练带有一个所谓隐藏层(既非输入层也非输出层)的网络。而现在,人们实际上已经知道如何训练带有更多层级的网络了。
Why is it called deep learning? Because it’s training networks with many layers. In the ’80s, people figured out backprop but they couldn’t really train networks with one so-called hidden layer, which is a layer that’s neither the input nor the output. But now people actually know how to train networks with more layers.
或许深度学习最知名的例子就是所谓的谷歌猫网络,几年前它登上了《纽约时报》头版。谷歌猫网络做的就是从观看 YouTube 视频中学会识别各类物体。它确实看了无数个小时的 YouTube 视频,所以或许应该叫它沙发土豆网络。
Perhaps the best-known example of deep learning is what has come to be known as the Google cat network, which was on the front page of the New York Times a couple years ago. What the Google cat network does is learn to recognize all sorts of objects from watching YouTube videos. It literally watches hours and hours and hours of YouTube videos, so maybe it should be called the couch potato network.
有些人真的以为这个网络只认识猫,其实不是的,它认识猫、狗、老鼠、人等等各种东西。那位记者之所以拿猫来举例,是因为这个类别是这个网络识别效果最好的。原因在于——我不知道你是否知道这件事——人们真的很喜欢上传自己家猫的视频。所以,猫相关的数据比其他任何实体的数据都要多。
Some people actually think that all it recognizes is cat, but no, it recognizes cats and dogs and mice and people and whatnot. The reason the reporter picked cats as the example is that this is the category on which the network does best. This is because, I don’t know if you know this, but people really like to upload videos of their cats. And so, there is more data on cats than on any other entity.
现在,进化论者会说,好吧,反向传播或许适合调整大脑的权重,但真正塑造大脑的是进化。进化不仅创造了大脑,还创造了地球上所有的生命。所以,进化而非反向传播,才是主算法。
Now, the evolutionaries say, well, sure, backprop might be good for tweaking the weights of the brain, but what made the brain was evolution. Evolution didn’t just make the brain, it made literally all life on Earth. So evolution, not backprop, is the master algorithm.
进化论者在计算机上模拟进化,只不过他们演化的是程序,而非动物或植物。最先着手这个领域的是约翰·霍尔,而他实际上就在去年夏天去世了。很久以前,他在 50 年代末、60 年代初刚起步时,人们常开玩笑说,进化计算学派就只有约翰和他的学生,以及他学生的学生。但到了 80 年代,情况开始起飞,许多不同的人开始在各类领域进行进化计算研究。
The evolutionaries simulate evolution on the computer except that instead of evolving animals and plants, they evolve programs. The person who first ran with this was John Hall, and he died actually just last summer. For a long time, when he started out in the late ’50s, early ’60s, people used to joke that the school of evolutionary computing consisted of just John and his students and their students. But then in the ’80s, things took off and a lot of different people started doing evolutionary computing in all sorts of different areas.
接着,约翰·科扎(John Koza)真正开发出了这一技术的变体,称为遗传编程,我们很快就会接触到。此外,霍德·利普森(Hod Lipson)是另一位目前正用这类学习方式做非常有趣事情的人,我们也会有所了解。
And then John Koza actually developed this version of it called genetic programming that we’re going to meet shortly. And then, Hod Lipson is another person doing very interesting things with this type of learning today that we will also look at.
这种学习方式的基本理念,约翰·霍尔称之为“遗传碳基”,因为这些算法本身就像是碳基生命一样。
This is the basic idea of this type of learning, which John Hall called genetic carbons, because they’re algorithms
华盛顿大学的佩德罗·多明戈斯
Pedro Domingos University of Washington
模仿基因学功能的数字化工厂。
that imitate what genetics does.
在任何给定的时间点上,你都拥有一个个体群体,而每个个体都由一个基因组来定义。
At any given time you have a population of individuals, and each of these individuals is defined by a genome.
但我们这里的情形是,基因组不是 DNA 碱基对,而只是比特。在某种意义上,DNA 的比特就是计算机的 DNA,它们只是比特的强度。
Except that in our case, the genome isn’t going to be the DNA base pairs, it’s just going to be bits. In some sense, DNA bits are the DNA of computers, so they’re just going to be bit strengths.
因此,这些比特强度中的每一个都定义了一个程序,然后这个程序进入世界去做它的事情。它试图执行我们想要它执行的任何任务,并得到一个适应度评分。再次强调,这与生物学非常接近。
So each of these bit strengths defines a program, and then that program goes out into the world and does its thing. It tries to perform whatever task we want it to perform, and it gets a fitness score. Again, this is very close to biology.
表现更好的程序会比表现不佳的程序获得更高的适应度评分。然后,适应度最高的个体实际上会繁衍下一代。你真的是在父本程序和母本程序的基因组之间进行交叉,得到它们的子代基因组。
The programs that do better get a higher fitness score than the ones that don’t do so well. And then, the fittest individuals actually get to produce the next generation. You literally do crossover between the genome of a father program and the mother program and you get genomes of their children.
然后,在此基础上,你又像真实进化那样进行随机突变,从而得到一个新的种群。令人惊奇的是,你可以从一个随机个体组成的种群开始这个过程,经过几千代之后,它们实际上就能做出非常有用且非常不显而易见的操作。
And then, on top of that, you do random mutations again just like in real evolution, and you have a new population. The amazing thing is that you can start this process with a population of random individuals and after, say, a few thousand generations, they’re actually doing very useful and very non-obvious things.
例如,该领域的人已经能够从一堆随机元件开始,进化出像收音机和放大器这样的东西。在这个过程中,他们实际上积累了大量专利。他们为那些由遗传算法发明的设备获得了专利。例如,他们有放大器、低通滤波器,其性能优于人类工程师设计的同类产品。
For example, people in this area have been able to evolve things like radios and amplifiers starting literally from piles of components. And along the way, they’ve actually amassed a lot of patents. They’ve gotten patents for devices that were invented by genetic algorithms. They have, for example, amplifiers, low pass filters that work better than the ones that were designed by human engineers.
但约翰·科扎的想法是,将程序表示为比特串太底层了。当我做交叉时,我会随机选择一个点,然后使用一个基因组到那个点为止的部分,再使用另一个基因组在点之后的部分。自然就是这样运作的。但这非常混乱。很可能我已经有了一个相当不错的程序,然后在随机位置把它切开,它就再也没什么用了。
But John Koza’s idea was that representing programs as bit strings is too low level. When I do crossover, I pick a random point and then I use one genome up to that point and then the other genome after that point. This is how nature does things. But it’s very messy. It’s very likely that I have something that is already a pretty good program and then I cut it at a random place and it doesn’t do anything useful anymore.
所以约翰·科扎的想法是一个程序。说到底,我们是在试图进化程序,而程序实际上是一个子程序调用的树形结构,一直下沉到执行加法、乘法、与运算、或运算这样的简单操作。因此,他发明了遗传编程,实际上使用程序树本身作为基因组。
So John Koza’s idea was a program. At the end of the day, we’re trying to evolve programs and a program is really a tree of subroutine calls all the way down to simple things like doing additions and multiplications and and’s and or’s. So he invented genetic programming to actually use the program tree itself as the genome.
这是一个非常简单的操作树。在树根处,是 C 乘以某个东西的平方根,假设我从我的交叉中选取了高亮显示的节点。其中一个子代树将是由白色音符组成的树。
Here is a very simple tree of operations. At the root, there is the multiplication of C by the square root of something else, and let’s say I have picked the highlighted note from my crossover. One of the child trees is going to be the tree with the white notes.
那个树实际上就是开普勒定律中的一条。它是描述行星年的平均持续时间是其到太阳平均距离的函数的那条定律。它实际上与距离的立方的平方根成正比。
That tree is actually one of Kepler’s laws. It’s the law that gives the average duration of a planet’s year as a function of its average distance from the Sun. It’s actually going to be proportional to the square root of the cube of the distance.
遗传算法可以以不同方式从类似开普勒使用的第谷数据中归纳出这一点,但它也能归纳出复杂得多的东西。它可以归纳出完整的机器人程序,也能归纳出执行非常复杂任务的程序。
A genetic algorithm can variously induce this from something like Tycho Brahe’s data that Kepler used, but it can also induce much, much more complex things. It can induce whole robot routines, and it can induce programs that do very nontrivial tasks.
事实上,如今这些进化论者不仅在进化程序,还在进化真正的实体机器人。这里的这只小蜘蛛实际上来自霍德·利普森实验室,是一个被进化出来的机械蜘蛛。
And in fact, these days, the evolutionaries are doing things like evolving not just programs but real hardware robots. This little spider here is actually a mechanical spider from Hod Lipson’s lab that was evolved.
发生的事情是,机器人在仿真中从随机的一堆元件开始。一旦它们表现得足够好,就会被 3D 打印出来,开始在现实世界中行走和爬行。其中有蜘蛛、有会飞的蜻蜓、有看起来像你从未见过的东西,但它们实际上是能爬行、行走、从伤害中恢复等等的。在每一代中,最适应环境的机器人会编程 3D 打印机来生产下一代机器人。
What happens is that the robots start out as random piles of components in simulation. Once they’re doing well enough, they get 3D-printed and they start to walk and crawl in the real world. There are spiders, there are dragonflies that fly, there are things that look like nothing that you ever saw before, but they actually crawl and walk and recover from injury and so forth. And in each generation, the fittest robots get to program the 3D printer to produce the next generation of robots.
佩德罗·多明戈斯,华盛顿大学
Pedro Domingos University of Washington
所以这很令人兴奋,也许还有点吓人,对吧?如果终结者真的来了,(笑声)也许就会是这样发生的。当然,这些小蜘蛛还没有准备好接管世界,但与它们最初那堆随机零件相比,已经走过了漫长的道路。
So this is exciting and maybe also a little scary, right? If the Terminator comes to pass, [Laughter] maybe this is how it’s going to happen. Of course, these little spiders are not ready to take over the world, but they’ve come a long way from the random pile of parts that they started out as.
现在,大多数机器学习研究人员实际上并不相信模仿生物学是通向主算法之路。所以进化论者模仿进化,连接主义者模仿大脑。但大多数机器学习研究人员的态度是,嗯,自然界就是这么做的,谁知道为什么,谁知道它到底有多好。让我们试着从第一性原理出发,弄清楚如何最优地学习。贝叶斯学派非常倾向于这个范式,即找出最优的学习方式,正如我所说,贝叶斯学派在统计学中历史悠久,在机器学习中也是如此。
Now, most machine learning researchers actually don’t believe that imitating biology is the way to get to the master algorithm. So the evolutionaries emulate evolution, the connectionists emulate the brain. But most machine learning researchers have the attitude that, well, nature did things that way, who knows why and who knows how good it really is. Let’s just try to figure out from first principles how we can learn optimally. Bayesians are very much in this paradigm of figuring out what the optimal way to learn is, and as I said Bayesians have a long history in statistics but also in machine learning.
最近,也许最著名的贝叶斯学派学者是朱迪亚·珀尔,他因发明了名为贝叶斯网络的东西,在 2011 年获得了图灵奖,即计算机科学的诺贝尔奖。这是一种非常强大的贝叶斯模型,现在被用于许多不同的事情。另外两位杰出的贝叶斯学派学者是大卫·赫克曼和迈克·乔丹。
More recently perhaps, the most famous Bayesian is Judea Pearl, who won the Turing Award, the Nobel Prize of Computer Science in 2011, for inventing something called Bayesian networks. This is a very powerful type of Bayesian model that is used for many different things now. Two other prominent Bayesians are David Heckerman and Mike Jordan.
贝叶斯学派的名字来源于贝叶斯定理。贝叶斯学习完全基于贝叶斯定理,事实上,贝叶斯学派如此热爱贝叶斯定理,以至于有一家贝叶斯机器学习的初创公司,真的制作了一个贝叶斯定理的霓虹灯标志,挂在办公室外让全城人都能看到。所以他们确实、确实非常相信贝叶斯定理。
Bayesians take their names from Bayes’ theorem. Bayesian learning is all based on Bayes’ theorem, and in fact, Bayesians love Bayes’ theorem so much that there is a Bayesian machine learning startup that actually had a neon sign of Bayes’ theorem made and hung outside its offices for the whole city to see. So they really, really believe in Bayes’ theorem.
在机器学习界,贝叶斯学派以五个部落中最狂热的一支而闻名。他们对他们的范式真的有着巨大的宗教般的依恋,而且他们是第一个承认这一点的人。
Bayesians are actually known in the machine learning community as being the most fanatical of the five tribes. They really have enormous religious attachment to their paradigm and they’re the first ones to say so.
他们不得不这样,因为在统计学中长达 200 年里,他们是被迫害的少数派。统计学一直被频率主义所主导,贝叶斯学派必须变得非常强硬才能生存下去,而他们这么做是件好事,因为他们有很多东西可以贡献。如今,随着计算机和更好的算法出现,他们实际上在统计学内部也占据了上风。
They have to be because for 200 years in statistics, they were a persecuted minority. Statistics was dominated by Frequentism, and the Bayesians had to get very hardcore in order to survive, and it’s a good thing they did because they have a lot to contribute. These days with computers and better algorithms, they’re actually in the ascendant even within statistics.
那么,贝叶斯定理和贝叶斯学习到底是怎么回事呢?其核心思想是,贝叶斯学派最关心的是不确定性问题。他们痴迷于这样一个事实:我所知道的任何事情,我都永远无法确切知道。任何不是从数据中得出的东西,我永远不能完全确定它是正确的。
So what is Bayes’ theorem and Bayesian learning all about? The idea is that Bayesians, above all, are concerned with the problem of uncertainty. They are obsessed with the fact that nothing I know do I ever know for sure. Anything that isn’t used from data, I can never be completely sure is right.
因此我们需要用我们量化概率的方式来量化不确定性。然后,我们有一系列正在考虑的假设,当我们看到证据时,我们会更新每个假设的概率。
So we need to quantify the uncertainty in the way we quantify probability. Then, we have a spate of hypotheses that we’re considering and as we see evidence, we’re going to update the probabilities of each hypothesis.
粗略地说,与数据一致的假设会变得更可能,与数据不一致的假设会变得更不可能,最终会有一个胜出者。但也可能没有唯一的胜出者,那时你就需要对假设按其置信度加权取平均。
Roughly speaking, the hypotheses that are consistent with the data will become more likely, the hypotheses that are inconsistent with it will become less likely and eventually, there will be a winner. But there may not be a single winner, and then you just have to average the hypotheses weighted by the confidence that you have in them.
贝叶斯定理实际上就是一小段数学,它告诉你如何做到这一点。它其实简单到几乎不配被称为一个定理,除了它的重要性之外。它所做的事情是帮助计算每个假设的后验概率,即我在看到证据之后对这个假设的相信程度。
Bayes’ theorem is really just the little piece of math that tells you how to do this. It’s actually so simple that it’s barely worth being called a theorem except for the fact that it’s so important. What it does is to help compute the posterior probability of each hypothesis, which is how much I believe in that hypothesis after seeing the evidence.
但我从我的先验概率开始,即我在看到任何证据之前对每个假设的相信程度。这正是贝叶斯主义极具争议之处。大多数统计学家,实际上是大多数科学家会说,嗯,你没有依据来编造这些先验定义。你只是在假装量化了一些你一无所知的东西。
But I start with my prior probability, which is how much I believe in each hypothesis before I even see any evidence. And this is what makes Bayesianism very controversial. Most statisticians, in fact, most scientists will say, well, you have no basis to make up these prior definitions. You’re just pretending that you’ve quantified something that you know nothing about.
然而,贝叶斯学派对此的回答是,你无论如何都必须以这样或那样的方式做出这些假设。你可以隐含地做出它们,也可以明确地做出它们,而我们至少会让它们变得明确,这是健康的。
The Bayesian answer to that, however, is that you have to make those assumptions one way or another. You
所以从先验开始,然后随着证据开始出现,你问的问题是,如果我的假设为真,那么我看到这个证据的可能性有多大?如果我的假设使得证据变得可能,那么反过来,证据使得假设变得可能——如果我的模型使得我看到的世界变得可能,则模型本身是可能的。这个量,即在假设为真时看到数据
Pedro Domingos University of Washington
佩德罗·多明戈斯,华盛顿大学
can make them implicitly or you can make them explicitly, and we at least are going to make them explicit, and that’s healthy.
的概率,被称为似然,也是频率主义统计学家所使用的,以及我们在统计学入门课程中学到的。
So you start with the prior, and then as the evidence starts coming in, the question that you ask is if my hypothesis is true, then how likely am I to see this? And if my hypothesis makes the evidence likely, then conversely, the evidence makes the evidence likely. My model is likely if it makes the world that I am seeing likely. This quantity, the probability to the hypothesis of the data given the hypothesis, is called the likelihood and, it’s also what frequentist statisticians use and what we all learn in Stats 101.
当你将两者相乘,即先验和似然,你就得到了后验概率。还有一个归一化常数来确保所有概率加起来等于一,但这对我们的目的来说并不太重要。
And when you do the product of the two, the prior and the likelihood, you get the posterior probability. There is a normalization constant to make sure that everything adds up to one, but it’s not too important for our purposes.
你可以用贝叶斯学习做各种令人惊奇的事情,其中之一是驾驶汽车。你的第一辆自动驾驶汽车很可能会内置一个视觉网络。谷歌使用一个拥有数亿个连接的庞大视觉网络来决定向你展示哪些应用程序。
You can do all sorts of amazing things with Bayesian learning, one of which is driving cars. Your first self-driving car is probably going to have a vision network inside it. Google uses a massive vision network with hundreds of millions of connections to decide which apps to show you.
我们都很熟悉的一个贝叶斯学习应用是垃圾邮件过滤器。在垃圾邮件过滤器中,两个假设是:这封邮件是垃圾邮件,或者这封邮件不是垃圾邮件。一开始,你有一个先验概率,比如说,90% 的邮件是垃圾邮件。
One application of Bayesian learning that we are all familiar with is spam filters. In a spam filter, the two hypotheses are: this email is spam or this email is not spam. You start out with a prior probability that, let’s say, 90 percent of emails are spam.
然后,证据就是邮件的内容。例如,如果邮件中包含全部大写的单词“免费”,那就会使它更可能是垃圾邮件。如果它包含单词“伟哥”,那会使它更可能是垃圾邮件。(笑声)如果它包含“免费伟哥”以及四个感叹号,那么几乎可以肯定它是垃圾邮件了。
Then, the evidence is the contents of the email. So for example, if the email contains the word “free” in all capitals, that makes it more likely to be spam. If it contains the word “Viagra,” that makes it even more likely to be spam. [Laughter] And if it contains “free Viagra” with four exclamation marks, then it’s almost certain to be spam.
[Laughter]
[Laughter]
另一方面,如果它在签名行包含你最好朋友的名字,那就会使它不太可能是垃圾邮件。所以最终,在看了证据之后,你得到一封邮件是或不是垃圾邮件的概率,然后你使用某个概率阈值来决定是将邮件扔掉还是放入用户的收件箱。
On the other hand, if it contains the name of your best friend on the signature line, that makes it a lot less likely to be spam. And so finally, after looking at the evidence, you get a probability that the email is spam or isn’t, and then you use some threshold of probability to decide whether to throw out the email or put it in the user’s inbox.
大卫·赫克曼多年前就有了做这个的想法,这实际上是某个研究生在微软研究院实习时的一个夏季项目。如今,人们使用各种不同的机器学习算法进行垃圾邮件过滤,视觉学习仍然是最广泛使用的、也是最优秀的方法之一。
David Heckerman had the idea of doing this many years ago, and this was literally a grad student’s summer project when he interned at Microsoft Research. These days, people use all sorts of different machine learning algorithms for spam filtering, but vision learning is still one of the most widely used methods and one of the best.
最后还有类比派,他们的核心理念是——一切学习都是类比。我们之所以能在新情境中做出正确行为,是因为注意到当前情况与过往经验的相似之处,然后根据过去做过什么、哪些做法有效,来判断眼前该怎么做。类比派是六个部落里最松散的一支,其实只是一群各自以类比理念为基础做研究的学者。但这套思想在机器学习中占据着非常核心的地位。
Finally, we have the analogizers, whose idea is that all learning is analogy. The way we learn, the reason we’re able to do the right thing in new situations is that we notice there are similarities to our previous experiences, and then, based on what we did or what worked in those experiences, we figure out what to do in this new case. The analogizers are a less cohesive tribe than the other five. They’re really just a bunch of different people that all do learning based on this idea. But this is a very central idea in machine learning.
最重要的类比大师当属弗拉基米尔·瓦普尼克。他发明了支持向量机,也称核方法,在深度学习达到顶峰之前,这组算法一直是机器学习的主流。即使在今天,对于大量问题而言,支持向量机依然是更优解的方案,而非深度学习。
The most important analogizer is probably Vladimir Vapnik. He invented support vector machines, also known as kernel machines, which until the height of deep learning, were the dominant machine learning algorithms. Even today support vector machines are still the best method for tackling a lot of problems, not deep learning.
彼得·哈特(Peter Hart)是最早开创类比学习雏形的人之一,他提出的方法被称为最近邻算法。我们稍后会看到这个算法的具体内容。此外,还有一些著名的类比研究者,比如《哥德尔、埃舍尔、巴赫》的作者道格拉斯·霍夫施塔特(Douglas Hofstadter),他实际上创造了“类比者”(analogizer)这个术语。
Peter Hart was one of the people who started the very earliest form of analogy-based learning, called the nearest neighbor algorithm. We’re going to see what that algorithm is shortly. And then, there are famous analogizers like Douglas Hofstadter, the author of Gödel, Escher, Bach. He actually coined the term analogizer.
他说自己是个“类比思考者”,爱因斯坦也是。所有那些伟大发现,全是靠类比推理实现的。所以他深信类比才是终极算法。事实上,他最近那本 500 页的书就在论证一件事:所有学习、所有智能,不过就是类比,别无其他。
He says he’s an analogizer, and Einstein was an analogizer, and that all these great discoveries and things all happened through reasoning by analogy. So he very much believes that analogy is the master algorithm. In fact, his most recent book is 500 pages arguing that all of learning, all of intelligence, is just analogy and nothing else.
佩德罗·多明戈斯 华盛顿大学
Pedro Domingos University of Washington
有意思的是,《哥德尔、埃舍尔、巴赫》这本书其实更多是关于符号主义的机器学习、逻辑以及诸如此类的东西。但整本书就是哥德尔定理、巴赫的音乐和埃舍尔的艺术之间一个延展开来的类比。所以,类比式学习这个思路早在那个时候就已经潜伏于他的思想之中了。
Interestingly, Gödel, Escher, Bach is really a book that has more to do with the symbolist type of machine learning, with logic and whatnot. But the entire book is an extended analogy between Gödel’s theorem, the music of Bach, and the art of Escher. So the whole learning by analogy was already latent in his thinking, even back then.
那么,类比学习是如何运作的呢?让我通过一个简单的谜题来说明。我会给你两个国家的地图,我夸张一点把它们分别叫做 Posistan 和 Negaland,因为其中一个将展示正面案例,另一个则展示反面案例。
So how does learning by analogy work? Let me explain it by way of proposing a simple puzzle to you. I’m going to give you a map of two countries, and I’m going to fancifully call them Posistan and Negaland because one is going to have the positive examples and the other is going to have the negative examples.
所以,举个例子,当我们学习识别猫的时候,我们把猫的图片叫做正样本,把狗以及其他所有东西的图片叫做负样本。那么我要告诉你的,是波西斯国主要城市在地图上的位置。这儿有一个波西斯国的城市,那儿还有一个,那里是首都波西城。而这边,是内加兰的主要城市。
So for example, when we’re learning to recognize cats, we call pictures of the cat the positive examples and pictures of dogs and everything else negative examples. And what I’m going to tell you is where the main cities in Posistan are on the map. So here’s a Posistan city, here’s another one, there is the capital, Positiville. And the same thing for the main cities in Negaland.
我再问你一个问题:这两个国家的边界在哪儿?我刚才告诉了你主要城市的位置,当然,你没法确定边界到底划在哪儿,因为城市本身并不决定国界。
Another question that I’m going to ask you is where is the border between these two countries? I just told you where the main cities are, and of course, you can’t know for sure where the boarder is going to be because the cities don’t determine the border.
但是,如果我给你一张纸,上面标着这些点,你大概能大致划出边界应该在哪儿。而最近邻算法其实就是用下面这个思路来干这件事的:我假定,地图上的一个点,如果它离波西斯坦的某个城市比离内加兰的任何城市都近,那这个点就属于波西斯坦。
But if I give you a piece of paper with these things marked on it, you can probably roughly put down where the frontier should be. And the nearest neighbor algorithm is really just using the following idea to do this: I am going to assume that a point on the map is part of Posistan if it’s closer to a city in Posistan than to any city in Negaland.
因此,我将把地图划分为每个城市的邻域。一个城市的邻域是指比其他任何城市都更靠近该城市的点所构成的区域。例如,在这个示例中,这个区域就是该城市的邻域。而正类别的区域,则正好是正类别城市的邻域之并集。
So I’m going to break up the map into the neighborhood of each city. The neighborhood of a city is the points that are closer to it than to any other. For example, the neighborhood of this example, here is this area. And then, the region of the positive class is just going to be the union of the neighborhood with the positive cities.
尽管近邻算法确实非常简单,但请注意,它在学习阶段完全什么都不做——你一点功课都不用做。这种算法有时也被称为惰性学习。就好比说:“哎呀,我太懒了,不想复习备考,等看到考题时再临时现编答案。”你妈妈告诉你拖延不好,可实际上,在机器学习里懒惰可能非常强大。之所以强大,是因为那条只被隐式形成的决策边界,其实可以变得极其、极其复杂。
Even though it’s a really simple algorithm, notice that at learning time the nearest neighbor algorithm consists of doing exactly nothing. You do no work. Sometimes this is also known by the name lazy learning. It’s like, oh, I’m lazy, I’m not going to study for the exam, and then, when I see the questions, I’ll make something up. Like, your Mom told you that procrastination is bad, but actually, in machine learning it can be very powerful. It can be very powerful because this frontier that is only implicitly being formed can actually get very, very intricate.
事实上,彼得·哈特早在 20 世纪 60 年代就证明了,你只需要用最近邻法就能学会世界上任何函数。说得更精确一点,只要你给它足够多的数据,它就能学会任何东西。
In fact, what Peter Hart did was prove back in the ’60s that you can learn any function in the world just by using the nearest neighbor. To be more precise, if you give this enough data, it can learn absolutely anything.
这个“最近邻”方法有几个缺陷,其中一个问题是,如果你观察它,这条前沿线有些参差不齐。真正的前沿线可能比这更平滑。
Now, the nearest neighbor has a couple of shortcomings, one of which is that if you look at it, this frontier is kind of jagged. The real frontier is probably smoother than that.
另一个问题是,如果你仔细想想,我在这里记住一些根本不需要的城市,其实是在浪费大量时间和空间。比如,如果我把这座城市拿掉,实际上,如果我把 “正数” 这个概念本身也拿掉,直接从地图上抹去它们,一切都不会改变。
The other is that if you think about it, I’m actually wasting a lot of time and space here by remembering cities that I don’t need to. For example, if I took out this city and, in fact, if I took out positive itself, if I just erased them from the map, nothing would change.
什么都不会改变的原因在于,这个地带只会被邻近城市的区域吸收,而边疆本身并不会发生变化。我真正需要记住的,只是那些让边疆保持原样、停留在原地的例子。举个例子,如果我把这个拿掉,那么边疆就会移动。
The reason nothing would change is that this neighborhood would just get absorbed by the neighborhoods of the nearby cities and the frontier itself would not change. The only thing that I really need to remember are the examples that keep the frontier as it is, where it is. For example, if I took this out, then the frontier would move.
这些例子被称为支持向量(support vectors)——称为向量,是因为机器学习中的样本通常被表示为向量;称为支持向量,则是因为它们支撑着决策边界。
Those examples are called the support vectors—vectors because examples in machine learning are usually represented as vectors, and support vectors because they’re supporting the frontier.
弗拉基米尔·瓦普尼克发明了支持向量机,这一方法从本质上解决了上述两个问题。它能够精确判断哪些样本需要保留,同时还能学习到一条更加平滑的分界线。这条分界线可以
Vladimir Vapnik invented support vector machines, which in essence solve both of these problems. They figure out exactly which examples you need to keep, and they also learn a smoother frontier. The frontier can come
佩德罗·多明戈斯 华盛顿大学
Pedro Domingos University of Washington
从远比分段直线更为广泛的曲线类别中得出。
from much more general classes of curves than just a piecewise straight line.
支持向量机实现这一点的方式也非常直观。假设我让你从地图的南端出发,你必须一路走到北端,始终保持正类城市在你的左侧,负类城市在你的右侧。
The way support vector machines do this is also quite intuitive. Suppose I told you to start on the south end of the map, and you have to walk all the way to the north, always keeping the positive cities on your left and always keeping the negative cities on your right.
我们懂该怎么走。我们先往前走,一直走到那儿。但有个转折——你必须尽可能远离所有城市。想象一下,这些城市是地雷,整片区域都是雷区。你不会随便乱走。在完成任务的前提下,你会尽量离地雷越远越好。
We know how to do this. We start walking, we go all the way up there. But there is a twist. You have to give all the cities the widest possible berth. Imagine that the cities were mines and this whole thing was a landmine. You wouldn’t just walk anywhere. You would stay as far away from the mines as you could while still doing the job.
你会尽力最大化你的安全边际,实际上,这正是支持向量机的工作原理——它们试图最大化分类边界与每类样本之间的距离。通过避免过于接近这个边界,你实际上规避了可能落入负面区域的风险,即便你并不确定那是否真的就是负面区域。
You would try to maximize your margin of safety, and this is, in fact, how support vector machines work. They try to maximize the margin between the frontier and the examples of each class. By avoiding going close to here, I actually avoid going into a region that perhaps actually is negative even though I am not sure.
基于类比的学习被用于各种场景。它是历史最悠久、最成熟的几种学习方式之一。但我们都熟悉的一个例子是推荐系统。例如,Netflix 需要决定向你推荐哪些电影。
Analogy-based learning has been used for all sorts of things. It’s one of the oldest and best established types of learning. But one that we are all familiar with is recommender systems. Netflix, for example, needs to decide which movies to recommend to you.
早期的时候,人们尝试根据受众的特征来推荐内容。比如,你喜欢动作片,但不喜欢阿诺德·施瓦辛格的片子,而喜欢某位导演的作品——
In the early days, people tried to recommend things based on the properties of the audience. It says, well, you like action movies but you don’t like movies with Arnold Schwarzenegger but you like movies with this director,
等等。事实证明,这个办法效果不太好,因为品味是很微妙的。你喜不喜欢一部电影,并不能简单地用它的特性来预测。利用别人作为参考资源,效果就好得多了。
etc. This turned out to not work very well because taste is subtle. Whether or not you will like a movie is not a simple function of its properties. Using other people as a resource works much better.
我需要做的,就是想推荐一部电影给你时,去找那些跟你有相似品味的人。我之所以知道他们的品味跟你相似,是因为你们给电影打了相近的分数。如果某个人在我给五星的时候也给五星,在我给一星的时候也给一星,然后他对一部我没看过的新电影打了五星,系统就会推测我也很可能喜欢它。这其实就是最近邻概念在这个特定领域的一个应用,效果惊人地好。
What I need to do to recommend a movie to you is look for people who have similar tastes. And I know that their tastes are similar to yours because you’ve given similar ratings to movies. If there is someone that gave five stars when I gave five stars, gave one star when I gave one star, and gave 5 stars to a new movie that I haven’t seen, the system hypothesizes that I’m going to like it as well. This is really just an application of the nearest neighbor idea in this particular domain, and it works shockingly well.
人们在 Netflix 上观看的电影中,四分之三来自推荐系统。这就是它对 Netflix 业务的重要性。当然,亚马逊也有一个推荐系统,你们都用过。亚马逊三分之一的销售额来自推荐系统。
Three-quarters of the movies that people watch on Netflix come out of the recommender system. This is how important it is to their business. And Amazon, of course, also has a recommender system that you’ve all met. A third of what Amazon sells comes out of the recommender system.
这对它们的净利润产生巨大差异,尤其是,这个系统的准确度至关重要。所有像样的电商网站都有一套这样的系统。如今人们使用各种不同的算法来完成这项任务,但最早的一种就是这种基于相似性的学习,而它至今仍是最好的之一。
This makes a huge difference to their bottom line and, in particular, how accurate this is makes a huge difference. Every e-commerce site worth its salt has one of these systems. These days people use all kinds of different algorithms to do this, but the earliest one was this type of similarity-based learning, and it’s still one of the best.
现在让我们退一步来看。我们已经认识了机器学习的五个主要流派。我们看到,每一个流派都有一个它能比其他流派解决得更好的问题,而每一个流派都有自己用来解决那个问题的算法。
Let’s take a step back now. We’ve met the five main tribes of machine learning. We’ve seen that each one of them has a problem that it can solve better than the others, and each one has an algorithm it uses to solve that problem.
对于符号主义者而言,他们真正关心的核心问题是学习知识,并能够以不同方式对知识进行重组。通过这种重组方式,他们可以具备高度的灵活性,而他们发现知识的途径则是逆向演绎——在演绎推理的过程中填补缺失的环节。
For symbolists, the problem that they really care about is learning knowledge that they can then compose in different ways. They can be very flexible that way, and they discover that knowledge through inverse deduction by filling in the gaps in deductive reasoning.
联结主义者模拟大脑,他们所要解决的问题被称为信用分配问题。它可能更应该叫作追责分配问题,因为当出现差错时,需要判断谁应当做出改变。他们用来完成这一任务的算法是反向传播。
Connectionists emulate the brain, and the problem that they solve is called the credit assignment problem. It probably should be called the blame assignment problem because it’s deciding who needs to change when something goes wrong. Their algorithm for doing that is backprop.
进化论者发现了结构。联结主义者必须先预设一个架构,然后再调整权重,但进化论者实际上首先知道如何演化这种结构。最……
The evolutionaries discover structure. The connectionists have to start with a predefined architecture and then change the weights, but the evolutionaries actually know how to evolve that structure in the first place. The most
华盛顿大学的佩德罗·多明戈斯
Pedro Domingos University of Washington
他们实现这一目标的高阶手段是基因编程。
sophisticated way they do this is via genetic programming.
然后是贝叶斯学派的信徒们,他们关心的是不确定性这个问题。他们对此的答案就是概率推断。
Then there are the Bayesians, who care about the problem of uncertainty. Their answer to that problem is probabilistic inference.
最后,类比推理者可以通过相似性进行推理。因此,他们在各种其他方法完全失效的情况下都能发挥作用。例如,如果我只给出一正一反两个例子,其他方法就不知道该怎么办了,但用最近邻法,你只需在它们之间画一条直线,这相当合理。
And then finally, the analogizers can reason by similarity. As a result, they can function in all sorts of situations where the others would completely fail. For example, if I just have one positive and one negative example, the other methods don’t know what to do, but with nearest neighbor, you just put a straight line between them, which is quite sensible.
他们还能进行更远的泛化。例如,尼尔斯·玻尔最初的量子力学理论就是基于原子与太阳系之间的类比:原子核相当于太阳,电子相当于行星,诸如此类。凭借类比,你实际上可以比其它方法走得更远。而如今,这些部落的每一个都深信自己掌握着主算法。
They can also generalize farther. For example, Niels Bohr’s original theory of quantum mechanics was based on an analogy between the atom and the solar system, where the nucleus was the sun and the planets were the electrons, and so on. So with analogy, you can actually generalize much farther than the other methods can. And now, each of these tribes very much believes in its own master algorithm.
例如,近来深度学习确实势如破竹,一些联结主义者认为反向传播就是他们永远需要的一切。但我认为事实是,正因为这些问题每一个都是真正的问题,所以没有哪个学派能给出全部答案。真正的答案是一个能实际解决所有五大问题的算法。到那时,我们才算真正拥有了一个主算法。
For example, these days, deep learning is really going like gangbusters and some of the connectionists think backprop is all they’re going to ever need. But I think the truth is that precisely because each of these problems is a real problem, none of the tribes has the whole answer. The whole answer is one algorithm that actually solves all five. That’s when we will truly have a master algorithm.
我们需要一个机器学习的大统一理论,正如标准模型是物理学的大统一理论——因为它统一了不同的力——或者中心法则生物学的大统一理论,以此类推。
We need a grand unified theory of machine learning in the same sense that the Standard Model is a grand unified theory of physics because it unifies the different forces, or the central dogma is a grand unified theory of biology, and so on.
所以那会是什么样子?我知道我们很多人已经在这方面研究了一段时间,最终取得了很大进展,并且已经相当接近了。我所描述的这些学习算法看起来都截然不同,因此似乎完全不清楚你如何才能将它们统一起来。事实上,有些人甚至认为这是不可能的。
So what might that look like? I know a lot of us have been doing research on this for a while and we eventually made a lot of progress and are getting fairly close. The learning algorithms that I describe all look very different, so it seems very unclear how you could possibly unify them. In fact, some people have argued that it’s impossible.
但一旦你注意到所有学习算法其实都由三个相同部分组成,事情就变得简单多了。因此,我们要做的只是逐一统一这些部分。
But it becomes a lot easier once you notice that all learning algorithms are really composed of the same three parts. And so, all we have to do is unify each of these parts in turn.
第一部分是表示。这是你为学习写程序所选择的语言。现在,人类程序员会使用 Java 和 Perl 之类的语言。通常,机器学习的人使用更抽象的语言,比如一阶逻辑,但原理是一样的。如果你在建模一个需要微分方程的物理系统,那也可以选微分方程。这就是表示的选择。
The first part is representation. It’s the choice of language in which you are going to write the program that you learn. Now, human programmers will use languages like Java and Perl and whatnot. Typically machine learning people use more abstract languages like, for example, first order logic, but the principle is the same. It could be differential equations if you’re modeling a physical system where you need to choose that. So that is the choice of representation.
这里自然要做的,就是把符号学派已经使用的一阶逻辑统一起来。但我们还需要把它与贝叶斯学派的方法统一,因为我们得处理概率问题,而逻辑做不到这一点。
A natural thing to do here is unify first order logic, which the symbolists already use. But we need to unify that with what the Bayesians do because we need to handle probability, which logic doesn’t.
我们已经实现了这一点,所以现在我们拥有了这种概率逻辑。例如,最著名的一种被称为马尔可夫逻辑网络,它将一阶逻辑与图模型结合在一起。这方面的例子包括视觉网络和马尔可夫网络。
We have done that, so now we have this probabilistic logic. For example, the best known one is called a Markov logic network, which combines first order logic with graphical models. Examples of this are vision networks and Markov networks.
这种语言本质上与一阶逻辑中的公式相同,一个公式就像一句英语句子。实际上,人们有时也把它们称为“句子”,只是表述得更形式化,以便计算机能够处理。但关键之处在于,我们现在要给每个公式附加一个权重。权重高的公式是你确实非常笃信的,因此如果世界违背了那个公式,其概率就会受到严重打击。
This language is essentially the same formulas as in first order logic, and a formula is just like a sentence in English. In fact, sometimes people call them sentences except it’s stated more formally in a way that the computer can handle. But the twist is that we are now going to attach a weight to each formula. A formula that has a high weight is a formula that you really believe in, so if the world violates that formula, then its probability really takes a hit.
所以你有所有这些带权重的公式,世界满足的公式越多、程度越高,
So you have all these formulas with all their weights, and the more formulas the world satisfies and the higher
华盛顿大学 佩德罗·多明戈斯
Pedro Domingos University of Washington
它们的权重越大,世界就越有可能如此。有了这一点,我们几乎可以代表任何领域中我们想要代表的一切。
weight they have, the more likely the world is. And with this, we can represent pretty much anything that we might want to represent in any field.
现在,第二部分是评估。我们需要一个评分函数来告诉我们一个候选程序有多好。其中一个可用的评分函数是后验概率,这个你们已经听过了,但更一般地说,评估函数实际上不应该是算法的一部分。它应该来自用户。是你,用户,应该告诉学习算法它需要优化什么。所以如果你是一家公司,评估函数可能是投资回报率。
Now, the second part is evaluation. We need a score function to tell us how good a candidate program is. One of the score functions to use is posterior probability, which you already heard about, but more generally, the evaluation function actually should not be a part of the algorithm. It should come from the user. It’s you, the user, who should tell the learning algorithm what it’s supposed to be optimizing. So if you’re a company, the evaluation function might be return on investment.
如果你是一名消费者,那或许是衡量你幸福感的一种指标。这由你自己来定义。最后,算法要做的就是优化——在那个由语言定义的巨大空间里,找到能获得最高分数的程序或模型。
If you are a consumer, it might be some measure of your happiness. It’s for you to say. And then, finally, what the algorithm has to do is optimize. Finding the program or the model in that big space defined by the language achieves the maximum score.
如今,进化论者与联结主义者的思想在这里发生了极为自然的融合。
And now here, there’s a very natural combination of ideas from the evolutionaries and from the connectionists.
我们需要找到这些公式,但一个公式本质上只是一棵由合取和析取构成的子公式树。
We need to discover the formulas, but a formula is just a tree of sub-formulas with conjunctions and disjunctions.
我们可以用遗传编程来演化我们的公式。
We can use genetic programming to evolve our formulas.
然后,我们可以利用反向传播来学习这些公式中的权重。我有一套完整的推理链条,用来解释这些数据,而公式中不同位置都有权重,我只需通过反向传播来优化这些权重即可。
Then, we can use backpropagation to learn the weights within the formulas. I have my big chain of reasoning that I use to explain the data, and with weights within the formulas in various places, I can just backprop through that to optimize my weights.
所以,我们实际上已经相当接近将这五种范式完全统一了。有些人认为或相信,这就是我们所需的一切。但我个人感觉、我的直觉是,事实并非如此——即便我们成功统一了这五种范式,仍然会有一些关键的新想法等着有人去提出。
So we are actually pretty close to having a complete unification of the five paradigms. And some people say that or believe that that’s all we’re going to need. My sense, my intuition is that actually that’s not the case. It’s that even after we have successfully unified these five paradigms, there will still be key new ideas that somebody has to come up with.
有些洞见我们至今尚未获得,而在某些方面,非机器学习领域的研究人员反而比我们这些业内人士更有可能捕捉到这些洞见。机器学习研究者已经沿着特定范式的路径在思考,这使得他们很难跳出这个范式去观察。
There are some insights that we haven’t had, and in some ways, someone who is not a machine learning researcher is better placed to have those insights than we in the field. Machine learning researchers are already thinking along the tracks of a particular paradigm, which makes it hard for them to see outside that paradigm.
我写这本书的一个隐秘动机,是希望更多人关注这个问题——说不定他们会想到我们没想到的主意。所以如果你搞明白了怎么做,请告诉我,我也好把它发表出来。
One of my secret motivations in writing my book was to get other people interested in the problem because maybe they’ll have those ideas that we’re not having. So if you figure out how to do this, let me know so I can publish it.
让我最后提一下,我认为借助大师算法将能实现而今天还做不到的一些事情。首先是家用机器人。我们都希望有机器人能洗碗、做饭、铺床,甚至照看孩子。为什么今天还没有呢?
Let me conclude by just mentioning some of the things that I think will be possible with the master algorithm that are not possible today. The first one is home robots. We would all like to have robots that do the dishes and do the cooking and make the beds and maybe even look after the children. Why don’t we have them today?
首先,大家一致认为,如果没有机器学习,你根本造不出家用机器人。我们连让汽车自动驾驶的程序都写不出来,更别提家用机器人了。第二个问题是,家用机器人在一天的日常活动中,会反复遇上这五类问题中的每一种,这意味着五主算法中的任何一个都不足以应付。如果我们能把它们统一起来,那么就有希望获得所需的能力。
Well, first of all, everyone agrees that you can’t build a home robot without machine learning. We don’t know how to program even a car to drive itself let alone a home robot. The second problem is that a home robot, in the course of an ordinary day, runs into every single one of those five problems, multiple times, which means that no single one of the five master algorithms is enough. If we unify them, then hopefully we will have what we need.
再说一个例子:所有大型科技公司都有一项计划,要把万维网变成计算机能理解、能推理的知识库。谷歌有知识图谱(Knowledge Graph),微软有萨托里(Satori),等等。核心思路是——我不想只输入关键词然后得到一堆网页,我想做的是提出问题,然后得到答案。
Here’s another one: all of the major tech companies have a project to turn the worldwide web into a knowledge base that computers can understand and reason with. Google has the Knowledge Graph, Microsoft has Satori, et cetera. The idea is that I do not want to just type in keywords and get back pages. What I want to do is ask questions and get answers.
但要让计算机做到这一点,它必须先理解网络上的各类文本,需要把这些文本转换成某种一阶逻辑式的表达。然而,网络上的知识本身就杂乱无章。
But for the computer to be able to do that, it has to understand the text that’s out there on the web. It needs to transform that text into something like first order logic. On the other hand, the knowledge on the web is messy.
它模糊不清、自相矛盾、支离破碎,而且不完整。未来将充满不确定性,因此你需要——
It’s ambiguous, contradictory, broken, and incomplete. It’s going to be full of uncertainty, so you need the
华盛顿大学的佩德罗·多明戈斯
Pedro Domingos University of Washington
概率也是如此。所以在一天结束时,我们还需要将这五种范式统一起来,才能做到这一点。
probability as well. And so again at the end of the day, we’re going to need to unify those five paradigms in order to be able to do this.
一个非常重要的问题——我们希望计算机有朝一日能够解决——是治愈癌症。从某种程度上说,诊断疾病是机器学习的完美应用场景。事实上,对于大多数疾病,学习算法已经比医生做得更好。从某种程度上说,你需要做的是为患者推荐一种药物,就像你可能推荐一部电影或一本书那样,只是这个问题要困难得多,非常困难。
A very important problem that we hope computers will be able to solve one day is curing cancer. In a way, diagnosing disease is a perfect application for machine learning. And indeed, for most diseases, learning algorithms already do this better than doctors. In a way, what you need to do is recommend a drug for the patient in the same way that you might recommend a movie or a book except that the problem is much, much harder.
我们之所以还没攻克癌症,是因为癌症不是一种疾病。每个人的癌症都不一样,同一种癌症也会随着发展出现变异,所以几乎不可能找到一种单一药物就能治愈癌症。
The reason we haven’t cured cancer is that cancer isn’t one disease. Everybody’s cancer is different and the same cancer mutates as it goes along, so it’s very unlikely that there will ever be a single drug that cures cancer.
我们真正需要的是一个能输入患者基因组、肿瘤突变、患者病史及其他相关信息的程序,然后针对那种特定癌症推荐一种药物,或者可能是几种药物的组合,甚至设计出一种全新的药物。
What we really need is a program that takes in the patient’s genome, the tumor’s mutations, the patient’s medical history, and other relevant information, and then suggests a drug for that particular cancer. Or maybe a combination of drugs. Or even designs a new drug.
已经有一些公司开始这么做了,还有项目在整合患者数据和其他信息,因为如果没有肿瘤、药物和疗效的数据,我们根本没法开展这项工作。但归根结底,我们需要模拟活细胞如何运作、基因调控如何发生,因为正是这个过程出了差错,才会导致癌症。
There are already companies that are starting to do this and projects to pull together patient data and whatnot because without that data of the tumors, the drugs, and the outcomes, we can’t do this. But at the end of the day, it’s going to take modeling how living cells work, modeling how the gene regulation happens because it’s when that goes awry that you get cancer.
所以,这将需要大量数据,比如微阵列数据和基因测序数据。但同样,它也需要比我们今天拥有的更强大的学习算法,因为所有那五个问题在这里都会反复出现。
And so, that’s going to require a lot of data like microarray data and gene sequencing. But again, it’s also going to require more powerful learning algorithms than the ones that we have today because all of those five problems keep turning up here.
让我再回到推荐系统这个话题。如今每家公司都有自己的推荐系统。
Let me return once more to recommender systems. Today every company has its own recommender system.
那个模型只是它们根据掌握的数据给你勾勒的一小部分画像。所以 Netflix 根据你给电影的评分,建了一个你观影口味的模型;亚马逊根据你在网站上的行为,建了一个你购物偏好的模型;Facebook 建了一个你的模型,用来决定给你推荐哪些更新;Twitter 也有一个推文推荐模型,诸如此类。
That model is just a little sliver of you based on the data that they have. So Netflix has a model of your movie tastes based on your movie ratings. Amazon has a model of what you buy based on what you did on their website. Facebook has a model of you to choose which updates to show you, and Twitter has one for tweets, and so on.
但这并不是我作为一个消费者真正想要的东西。我想要的是从我生成的所有数据中学习,形成一个完整的 360 度模型,然后让这个模型在我人生的每个阶段帮助我做出必须做的决定。不只是挑选电影或书籍,还包括找工作、决定去哪里上大学、找房子,甚至寻找伴侣。
But this is not what I as a consumer really want to have. What I want to have is a single complete 360-degree model of me learning from all the data that I ever generate, and then to have this model help me with the decisions that I have to make at every stage of my life. Not just picking movies or books, but finding jobs, deciding where to go to college, finding a house, even finding a mate.
如今世界上大多数婚姻——或者说当今美国的婚姻吧——始于线上,而红娘们正在学习算法,根据人们的个人资料为他们挑选潜在人选。所以说,当今活着的孩子里,有些要不是因为机器学习,根本就不会出生。
Most marriages in the world today, or I should say in America today, start online, and the matchmakers are learning algorithms, picking a potential list for people based on their profiles. So there are actually children alive today who wouldn’t have been born if not for machine learning.
但这一流程的质量仍然很低,因为你仅凭肖像资料根本预测不了两个人是否合适。你能看到的他们生活细节越多,对他们的了解越深,就越有可能做出好的判断。
But the quality of the process is still very low because you can’t predict whether two people are going to be a good match just based on their profiles. The more of their life you can see, the better you know them, the better you’ll be able to do this.
我们需要把那些数据整合起来,这其中有很多有趣的问题,比如我会信任谁来干这件事?我会信任这些公司中的某一家吗?隐私方面的考量又是什么?诸如此类。
We’re going to need to have that data pulled together, and there are a lot of interesting issues such as who will I trust to do that? Will I trust one of these companies? What are the privacy considerations, and so forth?
所有公司都在争相实现这一目标。谷歌有 Google Now,微软有 Cortana。你会看到这些东西纷纷涌现,但即使你拥有所有这些数据,我们今天拥有的学习算法实际上也无法实现这一点。而万能算法则可以。我的发言到此结束,下面开始提问。
Companies are all in a race trying to do this. Google has Google Now and Microsoft has Cortana. You see all of these things coming out, but even if you have all that data, the learning algorithms that we have today would actually not allow us to do this. The master algorithm would. Let me conclude there and take questions.
在实证金融领域以及历史经验中,这个概念是这样的:你提出一个假设,然后你再
Question: In the field of empirical finance and historically, the concept is you have a hypothesis and then you
佩德罗·多明戈斯 华盛顿大学
Pedro Domingos University of Washington
用数据来检验它。我们当中一些人历来会利用这些输出来帮助做出投资决策。
test it using the data. Some of us historically have used the outputs of that to help make investment decisions.
有一种特别引人注目的做法,如果你愿意这么叫的话,那些深思熟虑的使用者会把它叫作“数据挖掘”(data mining)。
There is sort of an egregious stand, if you will, that people who are thoughtful users talk about which is called data mining.
数据挖掘在实证金融领域是个贬义词。它包含一系列问题:模型过度拟合、对异常值反应过度、研究不具全局代表性的数据集、以及处理随时间变化而呈现非平稳性的过程。
Data mining is derogatory in empirical finance. Among a number of things, it’s overfitting models, it’s reacting to outliers, it’s looking at datasets that are not representative of the big picture, having nonstationarity that does process changes over time.
那么,对于不熟悉机器学习概念以及它们如何运作的人来说,这在你的领域里问题有多大,又有哪些技术手段可以防止基于过度拟合的严重错误?
So not being as familiar with machine learning concepts and how that all works, how much of a problem is that in your world and what are the techniques used to prevent egregious errors based on overfitting?
佩德罗:是的,过拟合是机器学习中的核心问题。每当你采用一种我描述过的那种强大学习方法时,这就成为一个巨大的危险。你会在不存在模式的地方虚构出模式。因此,你需要找出避免虚构这些模式的方法。而每一种范式都有其不同的应对方式。
Pedro: Yes, overfitting is the central problem in machine learning. Every time that you have a powerful type of learning like all the ones that I have described, this becomes a huge danger. You hallucinate patterns where there aren’t any. And so you need to figure out ways to avoid hallucinating those patterns. And every one of these paradigms has their different ways of doing this.
最简单的方法是确保,直到你的模型能在未经训练的数据上做出正确预测之前,你都不该相信它。这是一条确实非常简单的规则,而违反它则是搞机器学习的人犯的头号错误。如果你不遵守,结果会一塌糊涂,因为你大可以死记硬背。
The simplest way is to make sure that you don’t believe your model until it makes correct predictions on data that it wasn’t trained on. This is a really simple rule and breaking it is the number one mistake made by people who do machine learning. If you do that, your results will be crap because you can just memorize.
想想简单的最近邻算法。你可以记住数据,它在过去总是完美准确,但问题是,当一位新病人出现时,或者遇到新情况时,你能在那时做出正确判断吗?
Think of the simple nearest neighbor algorithm. You can memorize the data and it’s always perfectly accurate on the past, but the problem is when a new patient comes along, or a new situation, will you do the right thing there?
因此,战胜这个问题的一种方法是使用某种保留测试,或者有时被称为交叉验证——具体取决于你的操作方式。每个学习算法通常至少有一个参数,你可以通过调整它来在过拟合和忽视实际存在的模式之间进行权衡。
And so, one way to beat this is with a type of holdout testing or sometimes cross-validation as it’s called, depending on how you do this. Every learning algorithm usually has at least one parameter that you can tweak to make a tradeoff between overfitting and being blind to patterns that are there.
你要能够看见确实存在的模式,而不凭空幻想出不存在的模式。在最近邻算法中,邻居数量是一个极其简单的参数。通常,人们不会只使用单一最近邻,而是使用 k-最近邻算法。例如,我找到与该患者最接近的 k 个最近邻,让他们投票,得票最多的结果胜出。如果我增大 k 值,过拟合的可能性就会降低。极限情况下,如果我把所有人都用上,那么我预测的就只是最常见的诊断结果。
You want to be able to see the patterns that are there without hallucinating patterns that aren’t. In the case of nearest neighbor, there’s a very simple parameter in the number of neighbors that you use. In general, people don’t just use the single nearest neighbor; they use the k-Nearest Neighbors. For example, I find the k-Nearest Neighbors to that patient and they vote, and the majority wins. If I increase k, I become less likely to overfit. At the limit, if I use everybody, then I just predict the most frequent diagnosis.
问题:在最近邻这样的案例中,最理想的样本量是多少?
Question: What is an optimal sample size for something like the nearest neighbor case?
佩德罗:嗯,这要看情况。我们现在是大数据的时代。事实上,在很多领域,你拥有的重采样数据多到根本用不完。但也有一些问题,你手头的重采样数据很少,尤其最棘手的是那些现象本身在不断变化的问题。所以,如果你收集大量数据,数据就已经过时了;但如果收集不到足够的数据,你又无法掌握这个现象的规律。
Pedro: Well, it depends. We are in the days of big data. In fact, there are many domains where you have so much r-sample data that you don’t even use all of it. There are, however, problems where you don’t have a lot of r-sample data and, in particular, the most difficult problems are the ones where the phenomenon is continually changing. So if you gather a lot of data, you are outdated, but if you don’t get a lot of data, you don’t learn the phenomenon.
这些问题很难,在极限情况下也许你无法解决它们,但解决它们的一个方法是,在多个时间尺度上工作,并引入其他知识。你要极力找出那些恒定不变的东西。
Those problems are hard and at the limit maybe you can’t solve them, but one way you solve them is by working at multiple time scales and bringing in other knowledge. You really try to find what the constants are.
很多事物在变化,但有些不会。具体到细胞生物学领域,细胞的工作方式实际上并未改变。你看到的某种具体癌症是新的,但我们在机器学习领域非常懂得如何应对这种情况。
A lot of things are changing but some aren’t. And in particular, in cellular biology, the way the cell works actually doesn’t change. The particular cancer that you’re seeing is new, but we know how to cope with that very well in machine learning.
提问者:我觉得在座很多人可能都看过《黑客帝国》《终结者》《机械姬》吧,
Question: I think a lot of the people in this room probably have seen The Matrix, The Terminator, Ex Machina,
等等。当计算机达到奇点,这种风险是真实存在的。在你们的世界里,正在形成哪些道德护栏?
etc. When computers reach singularity, this risk is real. What ethical guardrails are taking shape in your world?
华盛顿大学 佩德罗·多明戈斯
Pedro Domingos University of Washington
佩德罗:是的,所以计算机是否会达到奇点,这一点争议很大。大多数计算机科学家并不相信这会发生,但奇点意味着你拥有一个能够学习的学习算法……想象一下,一个学习算法可以制造出另一个学习算法。如果它制造出的学习算法比它自己更强,而那个更强的又能制造出更强大的,那么我们就有了失控的智能。
Pedro: Yes, so it is very controversial whether computers will reach the singularity or not. Most computer scientists do not believe that that will ever happen, but singularity means you have a learning algorithm that can learn . . . imagine a learning algorithm that makes another learning algorithm. If it makes a learning algorithm that is better than itself, and that one makes a better one, then we have a runaway intelligence.
这其实就是奇点这个概念所描述的情形。
This is actually what the concept of the singularity is.
不过,即使这一情况真的发生,也不意味着每一次都会比前一次更好。事实上,你所想象的带来指数级增长的那些奇点,实际上无法无限延伸。就像所有那些看似呈指数曲线的技术曲线一样,到了某个点上,你就会开始看到收益递减。它们起初增长得越来越快,但随后会逐渐放缓,智力也是如此。
Now, even if this happens, it doesn’t mean that each one is going to get better than the previous one. In fact, the singularities that you imagine get this exponential growth, can’t actually go to infinity. Like all these technology curves that look like exponential curves, at some point, you start to see diminishing returns. They start out growing faster and faster, but then they taper off, and it’s going to be the same thing with intelligence.
提问:您是否认为人工智能会发展到成为危险的地步?
Question: Do you see AI advancing to a point where it becomes a danger?
佩德罗:不,正是如此,但假设我们将拥有极其智能的机器。我们该为它们担忧吗?
Pedro: No, exactly but let’s suppose that we will have very intelligent machines. Should we worry about them?
现在,这些问题对埃隆·马斯克和斯蒂芬·霍金这样的人来说是个麻烦,他们曾说过,人工智能对人类构成生存威胁,这话上了不少新闻等等。
Right now, they are a problem to people like Elon Musk and Stephen Hawking who said AI is an existential danger to humanity, which has gotten a lot of press and whatnot.
我不认识任何真正的人工智能专家会把这些想法当真。他们不当真的原因是,AI 再聪明,也只不过是我们自身的延伸。人们想象中 AI 是那种会有不同于我们的目标、会与我们竞争或消灭我们的智能体,但既然设计它们的是我们,凭什么会发生这种事呢?
I don’t know anybody who is actually an AI expert that takes these ideas seriously. The reason they don’t take them seriously is that AI, no matter how smart, is still just an extension of us. People have this image of AIs as being agents that are going to have different goals from us and compete with us or wipe us out, but why would that happen if we are the ones designing them?
关于《终结者》、《机械姬》以及所有这些好莱坞电影,有一个共同点:里面的 AI 和机器人总是由人类乔装扮演的,因为只有这样拍,电影才有趣。
The thing about Terminator and Ex Machina and all these Hollywood movies is that in them, the AIs and the robots are always humans in disguise, because that’s how you make an interesting movie.
但真正的人工智能,真正的机器学习算法,外观上跟伪装成人类的机器人完全不同。关键在于,它们的目标是由我们设定的。只要目标由我们设定,它们越聪明就越好。因为如果它们被设定的目标是攻克癌症,你当然希望它越智能、越强大越好。
But real AIs, real machine learning algorithms look nothing like humans in disguise. In particular, their goals are set by us. As long as their goals are set by us, the more intelligence they have, the better. Because if their set goal is to cure cancer, you want it to be as intelligent and as powerful as possible.
事实上,我们真正需要担心的不是计算机变得太聪明,而是它们太愚蠢。人们担心计算机过于聪明会接管世界,但真正的问题在于它们过于愚蠢,而它们已经接管了世界。
In fact, what we really need to be afraid of is not computers getting too smart but rather them being too stupid. People worry that computers will get too smart and take over the world, but the real problem is that they’re too stupid and they’ve already taken over the world.
计算机已经在替你和我做所有这些决策了,比如谁该获得信贷、谁被标记为潜在恐怖分子。它们非常容易出错,因为只是依赖数据集和程序设定。
Computers are already making all these decisions about you and me, like who gets credit and who gets flagged as a potential terrorist. They’re very fallible because they just lean on datasets and their programming.
因此,人工智能(AI)的危险来自企业信息集成(EII)——它不够了解情况,无法做出正确的事:因为它缺乏常识,误解了你的话。看看迈达斯国王(King Midas)的遭遇吧,他曾希望自己触碰的一切都变成黄金。
So the dangers from AIs come from the EII (Enterprise Information Integration) not knowing enough to do the right thing: from it not having common sense and misinterpreting what you say. Look what happened to King Midas, who wanted everything he touched to turn into gold.
要避免这些风险,办法是让计算机更聪明,而不是更笨。所以那种认为为了安全就应该限制计算机智能的想法,完全搞反了方向。计算机已经在开飞机了,很快它们还会开汽车。安全的关键在于让它们更智能,而不是更笨。
The way to avoid these things is to make the computers more intelligent, not less. So this idea that we should limit the intelligence of computers in order to be safe is exactly backwards. Computers are already flying airplanes, and soon they’ll be driving cars. The way to be safe is to make them more intelligent, not less.
问题:所以也许是一个类似的问题,换个问法,人类经验中有哪些部分是我们现有的五种感官无法触及的?
Question: So maybe a similar question asked differently, what parts of human experience are not reachable with the five that we have right now?
佩德罗:这是个好问题,让我先铺垫一下。十年之后,世界上的人工智能会比现在多得多。但其中绝大多数人工智能看起来一点都不像人,我们对它也会毫无察觉。它只是默默地在经济的某个角落里干着自己的活儿。
Pedro: That’s a good question and let me preface my answer with the following. Ten years from now, there will be a lot more AI in the world than there is now. But the vast majority of that AI will look nothing like people, and we will be completely unaware of it. It will just be doing its job in some corner of the economy.
一个小比例可能看起来像人,因为,比如说,如果我有一个家用机器人,一个非常可爱的——
A small percentage may look like people because, for example, if I have a home robot, one that’s very cute and
佩德罗·多明戈斯 华盛顿大学
Pedro Domingos University of Washington
可爱且拟人化的机器人,可能比一大块冷冰冰的金属卖得更好。所以,这些算法本身离人类还差得远,但你可以想象,再经过几代发展,就会诞生出一些算法,它们在自己的用途上,能完成人类所做的那些事。
cuddly and humanlike is likely to sell better than one that’s just a big hunk of metal. So these algorithms by themselves are not that close to people, but you can imagine some generations forward having algorithms that for their purposes do the things that people do.
那么问题来了,这个算法是否真有意识?这其实已经是人们在问某些人工智能聊天机器人的问题。因为它们看起来像是有意识,人们就把它们当人一样对待。
Here the question would be, does that algorithm really have consciousness? This is actually something that people already ask about some of these bots. Because they look like they’re conscious, people treat them like they’re human.
而我认为,我们永远都不会真正知道这个答案。你怎么知道我有意识?你只是假设我有意识,因为我们足够相似——既然你有意识,或许我也有,对吧?
And I think ultimately, we’re never really going to know that answer. How do you know that I am conscious? You just assume that I am conscious because we’re similar enough that since you’re conscious, maybe I am, right?
你只是在做类比推理。
You’re just doing analogical reasoning.
人们有一种惊人的能力,能将人性特质投射到那些行为稍具人类特征的事物上,因此同样的事情也会发生在智能机器人身上。无论它们是否真的有意识,我们都会把它们当作有意识的存在来对待。
And people have this amazing ability to project human qualities onto things that behave even slightly humanly, so the same thing will happen with intelligent robots. We will treat them as being conscious, whether or not they
我觉得这是一个有趣的问题,它们是否确实具备意识,但我想我们永远无法知道这个问题的答案。
are. I think it’s an interesting question whether they truly are conscious, but I think that we will never know the answer to that question.
问题:你一开始谈到了神经科学之间的类比,而我们在神经科学中学到的很多内容都来自历史上的反常事件。比如,有个人脑袋里插进一根杆子之类的事,然后有人就发现了某种东西。
Question: You started off talking about the analogues between neuroscience and a lot of what we learned in neuro were from anomalous incidents in history. You know, a guy getting a pole stuck through his brain, and so on, and then someone discovers something.
关于人工智能和机器学习,有很多正面的轶事。对此,我想大家都很熟悉 Target 的那则广告——它通过一个少女的购物模式预测到她怀孕了。
There are a lot of positive anecdotes with AI and machine learning. To that, I think people are familiar with, the Target ad that predicted a teenage girl was pregnant from her buying patterns.
现在有一家初创公司做视觉识别,能识别出那些我们误以为是色情图像但实际不是的图片。有没有计算机反过来搞错、彻底弄反的例子?
There’s a startup now that is doing visual recognition, identifying images that we would mistake for pornography that, in fact, are not. Are there examples where the computer is getting the opposite, where they get it completely wrong?
佩德罗:哦,是的,例子多得数不清。
Pedro: Oh, yes, there are examples galore.
问题是:我们从中能学到什么?
Question: What are we learning from that?
佩德罗:这个问题与上一个问题紧密相关,虽然看起来不太像。当我们看到某个学习算法能做到像识别物体这样的事情时,我们会假定它和我们采用同样的方式。
Pedro: This question is closely related to the previous one although it doesn’t seem to be. When we see one of these learning algorithms do something like, for example, recognize objects, we assume that it’s doing it the same way that we do.
举个例子,计算机通过学习能区分猫和狗,我们以为它大概知道猫或狗是什么,但事实上它并不知道。它只是捕捉到了某种信号,让自己能够分辨出猫和狗。
So for example, a computer learns to tell cats from dogs, and we assume that it kind of knows what a cat or a dog is, but in reality, it doesn’t. It just picked up on some signal that allows it to distinguish cats from dogs.
我的一位同事确实做了下面这个实验。他训练了一个神经网络,也就是那种深度模型之一,让它可以区分狗和狼,准确性达到 90% 以上,所以做得相当出色。
One of my colleagues actually did the following experiment. He trained a neural network, one of these deep models, to discriminate dogs from wolves, and it was 90-something percent accurate, so he was doing a very good job.
但他们后来有一种方法,能深入网络内部去尝试弄明白它学到了什么,也就是它关注的是图像中的哪些部分。
But then they had this way of looking inside the network to try and figure out what it has learned, what parts of the image that it was keying in on.
那么,你觉得它靠什么来区分?也许是鼻子,也许是耳朵?它真正关注的是图像中长长的横向白色斑块。为什么?因为狼的图像大多在雪地里,而狗的图像不在雪地里。所以它真正学到的是如何识别雪。[笑声] 而机器学习算法有一种惊人的倾向,总是干出这种事儿。
Now, what do you think it was keying on? Maybe the snout, maybe the ears? What it was focusing on was long horizontal white patches in the image. Why was that? Because the images of the wolves were in snow for the most part, and the images of the dogs were not. So what it really learned was how to classify snow. [Laughter] And machine learning algorithms have an uncanny tendency to do stuff like this.
现在,有一种看法是:天啊,这算法太蠢了。但现实世界中,很多动物也会犯类似的错误。如果你把一块方形的牛皮贴在一头母牛的侧面摩擦,这头牛
Now, one way to look at this is to say, oh my god, the algorithm is so dumb. But a lot of animals in the real world make mistakes like this, too. If you take a cow and you put a square of calfskin rubbing against its side, the cow
华盛顿大学佩德罗·多明戈斯
Pedro Domingos University of Washington
把小牛留在身后。是不是很震惊?骗一头奶牛就这么容易,而且我知道像这样的例子还有很多。
leaves its calf behind. Shocking, right? It’s that easy to fool a cow, and I know there are more examples like this.
但我认为从进化角度看,没有人会跟奶牛开这种玩笑,【笑声】。认知是有代价的。奶牛把能量用在了肠道上,而我们把能量用在了大脑上。但即使是我们的大脑,消耗的能量也极为惊人,所以用最少的能量把事办妥就行。
But I would look at this as evolutionarily, nobody was playing pranks like that on the cows, [Laughter]. Cognition has a cost. Cows use their energy in their intestine, whereas we use it in our brains. But even our brains are extremely energy-intensive, so using the least amount of energy that will get the job done is fine.
现在,如果你总能在雪地里遇到狼,而不在雪地里遇到狗,那你就没问题了。问题当然在于,有一天你在后院看到一只狼,而它可能会要了你的命。
Now if you will always meet wolves in snow and dogs not in snow, then you’re fine. The problem, of course, is that one day you see a wolf in your backyard and maybe this will get you killed.
提问者:谢谢。在您之前,比尔·格利(Bill Gurley)发言时提到一个观点:人工智能领域所有的技术进步之类,可能会给个人带来颠覆性冲击,但在社会层面却是一个巨大的积极因素,能创造增长。他援引了这样的例子:100 年前,99% 的人从事农业,而今天只有 1% 的人务农。人们找到了其他更好的工作来做。
Question: Thanks. We heard in Bill Gurley’s talk before you the idea that all of the technological advancement in AI and so forth might cause disruption for individuals, but be a big positive that creates growth at the level of society. He cited the idea that 100 years ago, 99 percent of people worked in agriculture, and today, one percent do. People found jobs doing other things that are better.
您是否有同感,或者更担心这次“第四次工业革命”是一种不同的技术变革——如果我们能大量复制人类能做的工作,那么留给人类工人的岗位可能会更少?
Do you feel similarly or are you more concerned that this Fourth Industrial Revolution is a different kind of technological change that may leave fewer roles for human workers if we can recreate a lot of what humans do?
佩德罗:这其实是个眼下正激烈的争论。有一派人说,这只是自动化的下一个阶段。每一种文化里,我们都见过这样的故事,没什么好担心的。人们总是更容易看到消失的岗位,而不是新出现的岗位。
Pedro: This is actually a raging debate right now. There are the people who say that this is just the next stage of automation. We’ve seen this story before in every culture, and there is nothing to worry about. It’s always easier to see the jobs that disappear than the ones that appear.
自动化的最大影响是让东西变得更便宜。人们现在会用同样的钱去购买其他东西。或者,一些过去完全不可能的事情变得可能了,于是就会产生新的工作岗位。
The biggest effect of automation is that things become cheaper. People are now going to buy other things with the same money. Or, things that were completely impossible become possible, and you have new jobs.
举例来说,如今全世界有数百万的 App 设计师。这个职业十年前还不存在。19 世纪的农民如果担心农耕岗位会消失,他们绝对不可能想象到,未来某一天会出现 App 设计师这类职业,甚至连此刻坐在这个房间里的我们所有人,也是他们无法预见的。
For example, there are millions of app designers in the world today. That job didn’t exist ten years ago. Farmers in the 19th century worrying about the destruction of farmers’ jobs couldn’t possibly have imagined that one day there would be app designers and other things like that, and all of us in this room for that matter.
但现在,持反对意见的人担心“这次不一样”。通常的论点是:工业革命自动化了体力劳动,但如今我们在自动化脑力劳动,这样一来就没剩下什么留给人类了。
But now, the counter to that is people who worry that this time is different. And the argument usually runs that the Industrial Revolution automated manual work, but now we are automating intellectual work, and then there will be nothing left for the people.
那么我站在什么立场呢?我认为,我们确实需要区分短期到中期与长期。在短期内——我说的短期是指未来十年——我不认为这次会有什么不同,因为人工智能是一条很长的路。
Now, where do I fall on this? I think we really need to distinguish the short to medium term from the long term. In the short term, and by short term I mean the next decade, I don’t buy the idea that this time is different, because AI is a very long road.
我们已走出千里之遥,但前方还有百万里路要走,而他们思考的是如何走好后面的百万里。这百万里路在可预见的未来不会到来。在遥远的将来,也许我们会有比人类更擅长一切的电脑和机器人。
We’ve come a thousand miles but there are a million more to go, and they’re thinking about this time as going the other million miles. That million miles will not happen in the near future. In the distant future, maybe we will have computers and robots that do everything better than people.
我认为这种情况下的结果是,我们所有人都会变得独立富裕。那时,人们已经在谈论像全民基本收入这类事了。但就政治层面而言,眼下要推行这样的措施是不可能的,因为我们绝大多数人都在工作,并没有领取失业救济金。
I think what’s going to happen in that case is that we’re all going to be independently wealthy. And then, people are already talking about these things like a universal basic income. But politically right now, it’s impossible to do something like that because the great majority of us are producing and not receiving unemployment benefits.
但是,如果发展到超过半数人口失业,而民主制度仍在运行,我看不出人们有什么理由不会为自己投票争取极为慷慨的失业福利,并将其称为终身收入。
But if you get to the point where more than half of the people are unemployed, and the democracy is in place, I don’t see how people will not vote themselves very generous unemployment benefits and call it lifetime income.
他们会为这些事情找出各种冠冕堂皇的道德理由。我认为这迟早会发生。
They would have all sorts of very good moral justifications for that stuff. I think that will happen.
在短期内,我认为将会出现大量就业岗位被取代的情况。许多工作会消失,这确实令人担忧。例如,如果我是一名卡车司机,我现在就会认真考虑换一份工作,因为在未来几年内——卡车司机是美国最常见的职业——其中大量岗位将会消失。所以我们必须关注接下来会发生什么。
In the short term, I think what’s going to happen is there is going to be a lot of displacement. A lot of jobs will disappear, and this is a real concern. For example, if I was a truck driver, I would be really looking at getting another job right now because sometime within the next several years—and truck driver is the most frequent occupation in the U.S.— a lot of these jobs are going to disappear. So we need to worry about what happens
佩德罗·多明戈斯 华盛顿大学
Pedro Domingos University of Washington
那些失去工作的人,情况并非我们必须把他们重新培训成数据科学家。今天有趣的一点是,一方面有人没有工作,另一方面,所有这些公司都极度缺乏合格的人才。
with the people who lose those jobs. It’s not that we have to retrain those people to be data scientists. Again, the interesting thing that is happening today is that on the one hand, there are people who don’t have jobs, but on the other hand, there are all these companies who are desperately short of qualified people.
因此,我们需要培训这些人,让他们能胜任那些工作。比尔说我们需要更多计算机科学家和数据科学家,但我也认为,我们需要让人们有能力自我再培训,找到他们接下来要干的工作。一个失业的卡车司机不会变成数据科学家。
So we need to train these people to be qualified for those jobs. Bill was saying we need more computer scientists and more data scientists, but also, I think we need to empower people to retrain themselves and to find the next job that they’re going to do. The unemployed truck driver isn’t going to become a data scientist.
假设货运成本下降,运输成本就降低,商品价格也会随之下降,这样一来很多人手上就有了更多余钱。也许他们现在会买更好的房子,于是建筑工人就有了更多就业机会。所以,卡车司机可能转行去做建筑工,而建筑工作又是极难实现自动化的。
Say the cost of trucking goes down, the costs of transportation go down, the costs of goods goes down, so a lot of people have more money. Maybe now they buy better houses so there are more jobs for construction workers. So maybe the truck driver becomes a construction worker, and construction work is very difficult to automate.
我们在人工智能领域痛苦地学到的一个教训——每个人都应该意识到——是:30 年前我们以为最容易自动化的工种是蓝领职业,而认为那些需要受教育程度的白领工作会很难被自动化。
One of the lessons that we’ve learned in AI painfully and that everybody should be aware of is that we used to think 30 years ago that the easiest jobs to automate were going to be the blue collar ones. We thought that the white collar jobs, which require education, were going to be hard to automate.
这种看法完全颠倒了。最难自动化的其实是建筑工这类工作,因为它们需要灵巧度、走动、不摔倒,还要识别那些我们完全不假思索就能看见的东西。这些任务可是花了好几亿年才进化出来的。
That has it exactly backwards. The hardest jobs to automate are things like construction work because they require dexterity, moving around, not stumbling, seeing things that we take completely for granted. These are tasks that took hundreds of millions of years to evolve.
另一方面,医生、律师、分析师、工程师、科学家,这些行当之所以难,是因为我们并非天生就会做。我们必须上大学才能学会这些技能。但这同时也意味着,计算机能比我们学得更快更多——因为人类正特意朝着这个方向在进化它们。所以,已经有很多白领岗位在萎缩或消失,未来还会有更多。
On the other hand, doctors, lawyers, analysts, engineers, scientists, these are hard because we didn’t evolve to do them. We have to go to college to learn how to do them. But that also means that the computers can learn that much more than we can because we are evolving them for that purpose. So there are already a lot of white collar jobs that have shrunk or disappeared and more of them will.
我认为未来这十年会发生的情况是,某些岗位会消失,但大多数工作还是会
I think what’s going to happen in this ten-year future is that some jobs will disappear, but the majority of jobs will
不会。它们只是会发生变化。
not. They’re just going to change.
我做这份工作的方式将会改变。工作中我需要思考的是:我的哪些工作内容可以被机器学习算法自动化?机器能否通过观察我,就学会做我做的事情?如果我做的所有事都能被自动化,那我得赶紧再找一份工作。
The way I do my job will change, and what I need to think about in my job is: what in my job can be automated by machine learning algorithms? Could a machine learn to do what I do by observing me? If it is the case that everything I do can be automated, then I need to get another job quickly.
但对大多数人来说——尤其是大多数白领工作——实际情况要有趣得多:我工作的某些部分可以被自动化,但其他部分则不能。
But what will be the case for most people and in particular, most white collar jobs, is actually much more interesting. It’s that some parts of my job can be automated but others can’t.
所以我想做的是,把自己的工作自动化。保住饭碗的方法,就是亲手实现自动化,然后把时间花在只有我才能做的、更高价值的事情上——这些事建立在计算机已经帮我自动化掉的基础之上。人们常常把这件事说成是人类与机器之间的赛跑,但真正的赛跑,是有机器的人与没机器的人之间的赛跑。
So what I want to do is I want to automate my job. The way to keep my job safe is to automate it myself, and then I spend my time doing the higher value things that only I can do on top of what the computer has now automated. People often frame this in terms of a race between humans and machines, but the real race is between a human with a machine and a human without a machine.
一个很好的例子是“深蓝”击败加里·卡斯帕罗夫的时候。人们当时认为,计算机已经是世界象棋冠军了,故事到此结束。但实际上,当今世界上最顶尖的棋手并非计算机——因为在现阶段,人类和机器各有所长。
A very good example of this is when Deep Blue beat Garry Kasparov. People figured computers were now the world chess champions, end of story. But actually the best chess players in the world today are not computers because for the time being, humans and machines have different strengths.
你要想清楚的是,如何借助计算机把自己的工作做得比单靠自己更好,或者比单纯靠计算机更好?在我看来,对大多数职业来说,这才是人们应该专注思考的问题。
What you want to figure out is how do I use the computer to do my job better than I alone could do it or better than the computer could do it? And I think for most occupations, this is what people should be preoccupied with.
迈克尔:我们先休会吃午饭,不过我想问一件事。两年前在这届大会上,主题是预测。当时我们有好几位演讲嘉宾认为,计算机在围棋上战胜人类还需要五到十年。但就在过去几个月里,我们看到了 AlphaGo。你能解释一下 AlphaGo 是怎么工作的吗?他们这么快就在围棋领域取得这样的成果,你感到意外吗?
Michael: We’re going to break for lunch, but I do want to ask one thing. Two years ago in this conference the theme was prediction. And a number of our speakers felt that computers beating humans at Go would take five to ten years. But in just the last couple months, we saw AlphaGo. Can you explain how AlphaGo works? Were you surprised by how quickly they were able to do what they did in the game of Go?
华盛顿大学的佩德罗·多明戈斯
Pedro Domingos University of Washington
佩德罗:是的,这个问题问得好。确实,如果你一年前问我,计算机要花多长时间才能击败人类围棋选手,我会说不确定。我当时可能认为还需要很久。那么,DeepMind 究竟做了什么如此了不起的事?
Pedro: Yes, that’s a great question. So indeed, if you had asked me a year ago how long it would take for computers to beat humans at Go, I would have said I wasn’t sure. I would have thought probably a long time. So what did DeepMind do that was so amazing?
德米斯·哈萨比斯就是创办 DeepMind 的那个人,这家公司的早期成功来自玩电脑游戏——用我们刚才看到的那种深度学习算法玩雅达利(Atari)游戏。后来,德米斯决定探究一个问题:为什么围棋比国际象棋或跳棋难那么多。
Demis Hassabis is the guy who created DeepMind, and its initial success was in playing computer games, in playing Atari games, using these deep learning algorithms that we just saw. Now, what Demis decided to do at one point was ask why Go was so hard compared to chess or checkers.
为什么围棋在 30 年前就被解决了,而且完全没有用到机器学习,只是靠搜索?问题在于,围棋的可能性空间要大得多。而选出一手好棋这个问题,与其说是逻辑推理,不如说更像是一种视觉模式识别问题。
Why was chess solved 30 years ago with no machine learning involved, just using search? The problem is that within Go, the space of possibilities is way larger. And the problem with picking a good Go move is almost more like a visual pattern recognition problem.
大多数棋类程序的运作方式是它们拥有对棋盘的评估函数。它们会说这个棋盘位置不错,因为我有棋子优势、控制了中心区域等等。我可以列举出这些因素。回到 50 年代,人们尝试了这些方法,并给它们赋予权重。这对围棋行不通。你问一位围棋高手他们为什么走那一步,他们解释不出来。
The way most game players work is they have these evaluation functions of the board. They say this is a good board position because I have a piece advantage, control of the center, etc. I can enumerate those things. And going back to the ’50s, what people did was they tried these things and they put weights on them. This doesn’t work for Go. You ask a Go expert why they made that move and they can’t explain.
所以,围棋更像是一个模式识别问题。而模式识别正是深度学习所擅长的:视觉、语音等等。于是,DeepMind 把这些神经网络接入到经典的围棋搜索流程中,这与人们之前使用的方法又略有不同。有一种叫蒙特卡洛树搜索的方法,围棋玩家借助它从一塌糊涂的水平进步到与人类业余棋手相当的水平。这种方法早已存在,但模式识别却没有。当蒙特卡洛树搜索与神经网络通过识别进行的思考结合起来时,你就得到了 DeepMind 所取得的成果。
So Go is almost more of a pattern recognition problem. Pattern recognition is what deep learning is good at: vision, speech, and whatnot. So DeepMind took these neural networks and plugged them into a classic Goplaying search process, which again is a little different from the ones that people used before. There is a thing called Monte Carlo tree search that Go players use to go from being a complete disaster to being as good as a human amateur. That was already there, but the pattern recognition wasn’t. When you combine the Monte Carlo tree search with the neural network’s thought by recognition, you get what DeepMind produced.
另一个有趣的方面是,DeepMind 之所以能达到今天的高度,最初是通过从它所能找到的所有围棋棋局的大型数据库中学习。那台机器花了三个月时间,动用数千台服务器,仅仅是自己跟自己下棋。
The other interesting aspect of it is that DeepMind got to where it is by initially learning from a big database of all the Go games that it could find. That thing spent three months burning thousands of servers just playing against itself.
实际上,这是整个领域中历史最悠久的概念之一:自我对弈。机器学习的说法最早出现——至少据我所知——是在 20 世纪 50 年代 IBM(国际商业机器公司)研究院一位叫阿瑟·塞缪尔的研究员的论文里。他教计算机下跳棋,方法是让计算机自我对弈,直到它的棋力达到人类水平。
And in fact, this is one of the oldest ideas in the whole field: self-play. The first known occurrence, at least known to me, of the term machine learning is in a paper by a guy at IBM (International Business Machines) Research in the ’50s called Arthur Samuel. He taught a computer to play checkers by playing against itself until it was as good as a human being.
当时,IBM 总裁托马斯·J·沃森说过,这份报告发表后,IBM 的股票会上涨 15%。事实也确实如此,因为人们对计算机能做什么的看法发生了变化。如今谷歌这样的公司也是如此。真正推动这一切的,是全新想法、半新想法和非常古老想法的组合,再加上强大的计算能力。
At the time, Thomas J. Watson who was the President of IBM, said that when this paper is published IBM’s stock will go up by 15 percent. And it actually did because people’s perception of what a computer could do changed. This is what’s happening with companies like Google today. It’s a combination of very new ideas, semi-new ideas, and very old ideas that actually made this all happen, coupled with a lot of computing power.
麦克尔:我想就到这里吧。非常感谢你,佩德罗。
Michael: I think on that, we’ll break. Thank you very much, Pedro.
宾夕法尼亚大学的凯德·马西
Cade Massey University of Pennsylvania
凯德·马西是沃顿商学院运营、信息与决策系的实践教授。
Cade Massey is a Practice Professor in the Wharton School’s Operations, Information and Decisions Department.
他在芝加哥大学获得博士学位,先后在杜克大学和耶鲁大学任教,之后进入宾夕法尼亚大学。凯德的研究聚焦于不确定条件下的判断——人们如何预测未来、预测得有多好。他的研究取材于实验数据以及"现实世界"数据,例如员工股票期权、401(k) 储蓄计划、美国国家橄榄球联盟(NFL)选秀和研究生录取。这些研究促成了他与谷歌(Google)、默克(Merck)及多家职业体育俱乐部的长期合作。
He received his PhD from the University of Chicago and taught at Duke University and Yale University before moving to the University of Pennsylvania. Cade’s research focuses on judgment under uncertainty—how, and how well, people predict what will happen in the future. His work draws on experimental and “real world” data such as employee stock options, 401(k) savings, the National Football League draft, and graduate school admissions. His research has led to long-time collaborations with Google, Merck, and multiple professional sports franchises.
Cade 的研究发表在顶尖心理学和管理学期刊上,并曾被《纽约时报》、《华尔街日报》、《华盛顿邮报》、《经济学人》、《大西洋月刊》及全国公共广播电台报道。他教授 MBA 及高管 MBA 课程长达 15 年,因讲授谈判、影响力、组织行为学和人力资源等课程,获得杜克大学、耶鲁大学和宾夕法尼亚大学的教学奖项。Cade 是沃顿商学院人员分析计划的教员联席主任、《华尔街日报》SiriusXM 商业电台“Wharton Moneyball”节目的联席主持人,以及梅西-皮博迪 NFL 实力排名的共同创作者。
Cade’s research has been published in leading psychology and management journals and has been covered by The New York Times, The Wall Street Journal, The Washington Post, The Economist, The Atlantic, and National Public Radio. He has taught MBA and Executive MBA courses for 15 years, receiving teaching awards from Duke, Yale, and Penn for courses on negotiation, influence, organizational behavior, and human resources. Cade is faculty co-director of Wharton’s People Analytics Initiative, co-host of “Wharton Moneyball” on SiriusXM Business Radio, and co-creator of the Massey-Peabody NFL Power Rankings for The Wall Street Journal.
宾夕法尼亚大学的凯德·马西
Cade Massey University of Pennsylvania
迈克尔·莫布森:我很荣幸介绍下一位演讲者,凯德·马西。凯德是宾夕法尼亚大学沃顿商学院运营、信息与决策系的实践教授。
Michael Mauboussin: It’s my pleasure to introduce our next speaker, Cade Massey. Cade is a Practice Professor at the University of Pennsylvania’s Wharton School in the Operations, Information, and Decisions department.
凯德的研究聚焦于不确定性下的判断,涵盖过度自信、乐观主义、反应不足与反应过度。
Cade’s work focuses on judgment under uncertainty, including overconfidence, optimism, underreaction, and overreaction.
首先我得说,我很喜欢跟凯德聊天。如果你对商业、体育或投资感兴趣,他的作品能提供数不清的启发,教你如何更善于思考、更有效地做决策。我之所以欣赏凯德,是因为他同时游走于实验世界和现实世界。他还跟谷歌等全球顶尖公司以及众多体育俱乐部合作过。
First of all, I have to say that I love talking to Cade. If you’re interested in business, sports, or investing, his work provides countless lessons about how to think better and decide more effectively. What I love about Cade is that he straddles both the experimental world and the real world. He has also collaborated with some of the largest companies in the world, such as Google as well as numerous sports franchises.
从很多方面来说,凯德站在今天所有讨论的交汇点上。你会听到关于算法厌恶、体育市场中的低效现象,甚至是我们如何在很大程度上主观性的流程(例如研究生招生)中引入更严谨性的方法。
In many ways, Cade stands at the intersection of all of the discussions today. You’re going to hear about algorithm aversion, inefficiencies in sports markets, and even ways that we can introduce more rigor to certain processes that are largely subjective, such as graduate school admissions.
最后我要提的是,凯德作为学院联合主任,深度参与沃顿商学院的“人才分析计划”。过去几年我有幸参加了该计划的年度会议,这项工作的进展非常令人振奋。一句话概括就是“《点球成金》遇上人力资源(HR)”。
The last thing I’ll mention is that Cade is very involved, as a faculty co-director, in Wharton’s People Analytics Initiative. I’ve had the pleasure of participating in the annual conference over the last couple of years and the work is very exciting. Think “Moneyball meets HR (Human Resources).”
请大家和我一起欢迎 凯德·梅西 教授。
Please join me in welcoming Professor Cade Massey.
凯德·梅西:谢谢你,迈克尔。非常感谢,真的很荣幸能来到这里。我们这个会议已经办了三年,迈克尔在过去两年里每次都做出了巨大贡献,对此深表感谢。
Cade Massey: Thank you, Michael. Appreciate it. Thank you. I’m delighted to be here. We’ve had our conference for three years. Michael has made huge contributions in each of the last two, and it is much appreciated.
所以迈克尔,我认真对待了你提出的这个主题——犯过的错如何教会我们做对的事。今天下午的讨论,我想先讲一个关于犯错的小故事。我决定给它起个更有诗意的标题,还是呼应你提出的主旨:“接受错误,以减少错误”。这个说法不是我原创的,你马上就会知道它出自哪里。
So Michael, I took seriously your theme of what being wrong can teach us about being right. I’m going to open the discussion this afternoon with a little story about being wrong. I also decided to title it a little bit more poetically, and again consistent with your message: “Accepting Error to Make Less Error.” I didn’t coin that phrase. You’ll see where that comes from momentarily.
不过,我想先讲一个关于梅西-皮博迪实力排名(Massey-Peabody Power Rankings)的故事。这个项目是我和以前的学生鲁弗斯·皮博迪(Rufus Peabody)共同合作的,他是一名职业体育投注者。我们发布橄榄球排名,起初只针对 NFL(美国国家橄榄球联盟),现在已覆盖大学联赛。这个排名在《华尔街日报》上刊登,我想已经有六年了。
But I want to start with a story about the Massey-Peabody Power Rankings. This is a collaboration with a former student of mine, Rufus Peabody, who is a professional sports gambler. We publish the football rankings, starting out just with the NFL (National Football League), and now spanning to college. We’ve done it I think for six years now in the Wall Street Journal.
我在耶鲁时,《华尔街日报》问我是否愿意为它们设计一套实力排名体系,我说可以,条件是必须让我的前学生与我合作完成这件事。报社同意了,于是过去三年里我们一直在发布这套排名。
When I was at Yale, the Wall Street Journal asked if I would put together a power-ranking system for them, and I said I would as long as I could get my former student to actually do this with me. They were up for it, and we’ve been publishing for the last three years.
如今,这些东西在网上更容易找到。我来给你看这个项目里的几个片段。这有点像是个车库里的业余项目,是我们私下里搞的东西。鲁弗斯是个全职赌徒,他也会用到这些输入信息,这和他职业生涯做的事确实很吻合。
These days, it’s more easily found online. I’m going to show you various clips from this project. It’s been kind of a garage project. This is just something we do on the side. Rufus is a full-time gambler. He uses these inputs a little bit. It’s certainly consistent with what he does in his professional life.
但对我来说,这是一个多了解一点不确定条件下判断的机会,也给了我一个谈论不确定条件下判断的平台,因为人们通常对听足球比赛更感兴趣,而不是我们在校园里做的最新实验。今天你们有福了,两样都会听到。
But for me, it’s been a way to learn a little bit more about judgment under uncertainty, and it also gives me a platform to talk about judgment under uncertainty, because people are generally more interested in hearing about football than the latest experiment that we’ve run on campus. Lucky for you, you’re going to get both today.
体系是这样设置的:我们将球队从第 1 名到第 32 名进行排名。这些都是 NFL 球队。我们对排名进行量化,得到的是每支球队在中立场地对阵平均水准球队时,预计能赢或输多少分。在这个排名中,最高是丹佛野马队,分值为正 8 分,最低是杰克逊维尔美洲虎队,分值为负 8.39 分。
The system is set up so that we rank teams from 1 to 32. These are NFL teams. And we quantify the ranking, so this is the number of points we’d expect them to win by or lose by to the average team on a neutral field. In this ranking, that’s the Broncos at plus eight all the way down to the Jaguars at minus 8.39.
所以,如果这两支队在中立球场交手,我们会看好野马队,胜过美洲虎队,差距是
So if these guys were to play each other on a neutral field, we would favor the Broncos over the Jaguars by
宾夕法尼亚大学的凯德·梅西(Cade Massey)
Cade Massey University of Pennsylvania
大约 16.4 分。我们这样设计,是为了如果你愿意,可以用它来对真实的 NFL 比赛下注。然后,他作为一个职业赌徒,而我作为一个教授,我们确实追踪了自己的表现。
something like 16.4 points. And we built it that way so that you could use it to bet in actual NFL games, if you wanted to do that. And then, his being a professional gambler and my being a professor, we actually track our performance.
我们想看看自己做得怎么样,结果发现我们在这个角色上干得不错。我们有点惊讶于自己居然干得这么好。每周,我们都会指定哪些是我们的“大注”。这些是我们最有信心的赌注。其他赌注也是下注,但信心没那么足。然后,为了增加一点趣味,我们还会额外加几场比赛。
We want to see how we’re doing, and it has turned out that we’ve done well in this role. We kind of surprised ourselves in how well we’ve done. Each week, we designate which are our “Big Plays.” Those are the ones we’re most confident about. The other plays are still bets but they’re not as confident. Then, just to add a little flavor, we throw a few extra games in there.
我们追踪这些大注的表现。在 2014 年左右,我们对了 10.5 场,错了 5 场,平了 5.5 场。但我们在整个赛季中都持续追踪。所以我们始终确保自己知道人们在沟通什么,以及我们表现如何。
We track performance on these Big Plays. In 2014 or so, we had 10-and-a-half games right, five games wrong, and five-and-a-half games tied. But we track this as we go through the season. So we’re always making sure we know where people are communicating and how we’re doing.
然后,在年底,我们想追踪一下总体表现。我稍微在这里建立一点可信度,这是前四年的数据。我在 2014 年赛季中期为另一个演讲做过这个数据,当时博彩市场的盈亏平衡线是 52.37。那是 50% 加上一点抽水。你必须做得比那更好才能赚钱。
And then, at the end of the year, we want to track how we’ve done. I’m building up credibility a little bit here, this is over the first four years. I ran this midseason in 2014 for another talk and the breakeven line in the betting markets is 52.37. That’s 50 percent plus a little bit of the vig. You have to do better than that to make money.
在我们所有的赛季中,我们都超过了那条盈亏平衡线。我们的校准显示,大注的表现优于其他注。我可以告诉你们,我们在 2014 年末并没有保持得这么好。均值回归同样会伤害体育博彩者。这个数字后来下降了一些。2015 年,我们度过了非常好的一年。我们再次盈利。几年前我们开始涉足大学橄榄球,在那方面也做得非常好。
Across all of our seasons, we’ve been above that breakeven line. We’re calibrated in that our Big Plays do better than our Other Plays. I can tell you that we didn’t end up 2014 this well. Regression to the mean hurts sports gamblers as well. This came down some. In 2015, we had a very good year. We were again profitable. We started doing college football a few years ago and we’ve done really well in that as well.
我们来做一下校准练习,展示我们的优势在哪里,以及我们与市场的差异在哪里——从与市场完全一致开始,到与市场差异越来越大。这是与市场 3 个点的差异。这个扩展到 6 个点。这里的直方图只是收益的频率,所以你可以看到,我们经常与市场一致,实际上通常与市场非常接近。
Let’s do the calibration exercise that says here is the edge and here is our difference with the market—starting with perfect agreement with the market to increasing difference with the market. This is a three-point difference with the market. This goes out to six. The histogram here is just frequency of gains, so you see that we very often agree with the market and in fact generally stay pretty close to the market.
但是,我们想知道的是,我们越是不一致,是否赢的次数就越多?我们看到这条线在上升,所以我们的校准很好。这完全是你希望看到的结果。但这只是总体情况,当然也有例外。我告诉过你们这是一个关于犯错的故事,所以我接下来会讲犯错的部分。
But then, what we want to know is, the more we disagree, do we more often win? And we see this increasing line, so we’ve got good calibration. This is exactly what you’d hope for. But this is just an aggregate, and of course there are exceptions. I told you this was a story about being wrong, so I’ll tell you the wrong part of it.
几年前,德克萨斯大学在他们与俄克拉荷马大学的一场重大对抗赛中遇到了对手,而德克萨斯大学那一年成绩很差。前两年,俄克拉荷马大学平均以 59 比 19 击败了他们,所以德州球迷很受伤,也不抱太大期望。赛季初,这场比赛的未来盘口是平手盘,但到赛季第五周时,盘口变成了让 14 分。
A few years ago, Texas came into their big rivalry game against Oklahoma, and they were having a bad year at the University of Texas. The previous two years, Oklahoma had beaten them by an average score of 59-to-19, so Texas fans were hurting and not expecting much better. At the beginning of the year, the futures line on this game was a pick ‘em, but five weeks into the year, it was a 14-point line.
所以此时,俄克拉荷马大学被预期赢 14 分。我们的模型输出了它对这场比赛的预测。Rufus 把预测结果发给我,我们有一个大注押在德克萨斯大学身上,这没问题。模型运行了四年半,我们从未改变过任何东西。问题在于我是德克萨斯大学的球迷。过去两年我经历了那些 59 比 19 的比赛。我觉得自己比任何人都更了解德克萨斯大学橄榄球队,而长角牛队即使拿到 14 分也不该被看好。系统认为他们会输 9 到 10 分,而不是 14 分。这在我们这里算是一个大注。
So at this point, Oklahoma is expected to win by 14. Our model kicks out its predictions for the game. Rufus sends me the predictions, and we have a Big Play on Texas, which is fine. Four-and-a-half years into the model, we’ve never changed anything. Trouble is I’m a Texas fan. I’ve lived through those 59-to-19 games the last couple of years. I feel like I know more about the University of Texas football team than anybody else out there, and there is no way the Longhorns should be favored even getting 14 points in this game. The system thought that they would lose by 9 or 10 instead of 14. That’s a Big Play in our world.
我以为 Rufus 在跟我开玩笑。我真的以为他在逗我。我说,模型不可能看好德克萨斯大学,他说,绝对看好,这是一个大注。我说,好吧,那我们不下这个注了。(笑声)他说,我们从来没有推翻过模型,从来没有。(笑声)我说,这次我们就破例一次。
I thought Rufus was pulling my leg. I literally thought he was joking with me. I said, no way the model likes Texas and he says, absolutely, it’s a Big Play. I said, well, we’re not going to run it. [Laughter] And he’s like, we’ve never overridden the model, ever. [Laughter] I said, we’re going to do it this time.
于是,他说,好吧,如果我们要这么做,你就得跟我打赌。他和我刚做完一些软件工作,所以我们赌了一把,赌注就是那笔软件工作的账单,只是为了让事情更有趣。
And so, he said, okay, well, if we’re going to do that, you’re going to have to bet with me. He and I, we had just done some software work, and so we bet between us the bill for that software work just to make it interesting.
他说,这样赌盘口可以,但如果德克萨斯大学直接赢了呢?我说,呃,随便你
And he said, that’s fine against the line, but what if Texas wins outright? And I’m like, well, whatever you want
凯德·梅西 宾夕法尼亚大学
Cade Massey University of Pennsylvania
因为那不可能发生。而我恰恰是研究过度自信的人。
because that’s not happening. And I’m the person who studies overconfidence.
所以赌注是,如果德克萨斯大学直接赢,接下来一周我们就要叫 Peabody-Massey。(笑声)不再是 Massey-Peabody,而是 Peabody-Massey,持续一周。
So the bet was that if Texas won outright, that we would be Peabody-Massey for the next week. [Laughter] None of this Massey-Peabody stuff. It would be Peabody-Massey for a week.
你们可能猜到结果了。长角牛队赢了,而且大胜,对长角牛球迷来说是辉煌的一天。接下来的一周,我们所有的渠道——我们的网站、我们的 Twitter 账号、《华尔街日报》——都以 Peabody-Massey 的名义出现。
You might guess where this is going. The Longhorns win, they win big, it was a glorious day for Longhorn fans, and the next week, in all of our outlets—our website, our Twitter account, the Wall Street Journal—we went out as Peabody-Massey.
那是我们唯一一次推翻模型。我感觉这给我上了一课,我很高兴付出了这点小代价来学到不推翻模型的教训。更明智的做法来自 Hilly Einhorn 一篇论文的标题,我今天演讲的题目也正是来源于此:“接受错误,以减少错误”。
It remains the only time we’ve overridden the model. And I feel like this was a lesson learned, and I’m happy to have paid that small price to learn that lesson of not overriding the model. The wiser way to go is the title from this Hilly Einhorn paper, which is where I get the title for the talk today: “Accepting Error to Make Less Error.”
Hilly 是芝加哥大学的研究员。他是芝加哥大学最早的行为研究者之一。在 Dick Thaler 之前,在 Josh Klayman 之前,Hilly 就已经在那里与那些新古典经济学家争论关于人的理性问题。
Hilly was a researcher at the University of Chicago. He was one of the first behavioral researchers at the University of Chicago. Before Dick Thaler was there, before Josh Klayman was there, Hilly was there arguing with all of those new classical economists about the rationality of man.
我不认识 Hilly,但我听说他很好斗,是个厉害的辩论者。他有很多出色的论文,其中有一篇很短的论文,你们大家都会喜欢,叫做“接受错误,以减少错误”。
And I didn’t know Hilly, but I’m told that he was pugnacious and he was a good fighter. He has a ton of great papers, and he has this very short paper that you guys would all enjoy, “Accepting Error to Make Less Error.”
其中的核心思想是临床判断与精算判断之间的区别,以及精算判断的优点和至高无上地位。
And the idea is this distinction between clinical judgment and actuarial judgment, and the virtue and supremacy of actuarial judgment.
但这其中包含一个概念:为了拥有那个优越的系统,你必须接受一些错误。本质上,为了拥有一个优越的模型,你必须放弃追求完美。这是一个很好的想法,也是一篇优雅的论文,但它在实践中是什么样子呢?
But in there is this notion that you have to accept some error in order to have that superior system. In order to have a superior model, you have to give up on being perfect, essentially. So that’s a nice idea and it’s an elegant paper, but what does that look like in practice?
在实践中更难。这些年来我们稍微进步了一些,在这个领域取得了相对出色的表现,但这仍然是一个非常嘈杂的领域。
It is harder in practice. We’ve gotten a little better over the years, and we’ve had relatively outstanding performance in this world, but it’s still a very noisy world.
所以 Einhorn 的“接受错误以减少错误”这个想法的问题在于,你必须接受每一次偏差,把它们当作你所能做到的最好结果。因此,那个算法就是预测本身。不允许你对任何一场比赛偏离那个预测。不允许你去追逐那些你认为自己实际上可以改进的错误。
So the trouble with Einhorn’s idea of accepting error to make less errors is you have to accept every one of those deviations as being as good as you can do. And so, that algorithm is the prediction. You’re not allowed to deviate from that prediction for any given game. You’re not allowed to chase those errors that you think you could actually improve.
当你整个赛季都在实践这一点时,Einhorn 的想法虽然好,但执行起来很艰难。我感觉,在过去六年从事 Massey-Peabody 项目的过程中,我学到的这一点,比我在任何实验中或课堂上空谈学到的都要多。
As you go through the season living that, Einhorn’s idea is nice but it’s tough. And I feel like I have learned that more from having worked on this Massey-Peabody project for the last six years than I could have ever running experiments or just talking about things in class.
我猜这跟你们的世界非常相似。我觉得,最能理解这类事情的人是体育博彩者和金融市场从业者,因为你有一个模型,有一套交易策略,但你在任何一天都不一定知道那个模型是否仍然有效。当你看到模型出现偏差时,这始终是对模型是否仍然有效的一种信心考验。
I assume this is very close to your world. I think my sense is the people who get this kind of thing very well are sports gamblers and folks who are involved in financial markets because you’ve got some model, you’ve got some trading strategy, and you don’t always know on any given day whether that model still holds. It’s always this test of confidence of whether it holds as you see these deviations from the model.
Einhorn 的观点会强力介入,而我今天要讲的一切,都将围绕坚持使用那个算法的优点展开,更重要的是,探讨人们想要偏离算法的心理动机。
Einhorn is going to strongly come in, and everything I’m going to talk about today is going to talk about the virtue of staying with that algorithm, and importantly, the psychology of the desire to depart from the algorithm.
最终,我们想探讨的问题是:如果人们确实想偏离算法,我们能做些什么来缓和他们的立场?我们能做些什么让他们更容易接受算法?
Ultimately, we want to get to the question, if people do want to depart from algorithms, what can we do to make them soften their position? What can we do to make them more amenable to algorithms?
所以,我想在第二部分,也就是演讲的主体部分,稍微谈谈我们是否有任何研究可以为这些问题提供信息。
So what I want to do in part two and kind of in the body of the talk is to speak a little bit about whether we have any research that can inform these questions.
我与两位合著者有几位论文和一个正在进行中的项目。第一篇论文去年发表在
I’ve got a couple of papers and an ongoing project with two co-authors. The first paper was out last year in the
凯德·梅西 宾夕法尼亚大学
Cade Massey University of Pennsylvania
《实验心理学杂志》上。两篇论文都是与 Berkeley Dietvorst 和 Joe Simmons 合作完成的。Dietvorst 上周末刚博士毕业。他在芝加哥大学获得了一个助理教授的职位。我为他感到非常骄傲。
Journal of Experimental Psychology. Both papers are with Berkeley Dietvorst and Joe Simmons. Dietvorst just graduated last weekend from our PhD program. He has taken an assistant professor position at the University of Chicago. I’m very proud of him.
Joe Simmons 是我长期的朋友和合作者,他一直处于心理学可复制性危机的核心位置。所以如果你们读过任何相关文章,如果你们知道过去六七年里那场风暴,Joe 与 Uri Simonsohn 和 Leif Nelson 合著的那篇论文就是该领域的首批关键论文之一。
Joe Simmons, a longtime friend and collaborator of mine, has been right at the heart of the replicability crisis in psychology. So if you guys have read any of that stuff, if you know about the storm that has brewed there over the last six, seven years, Joe’s paper with Uri Simonsohn and Leif Nelson was one of the key initial papers in that work.
我们有了这篇初始论文,然后紧随其后,我们现在有一篇正在审稿中的论文,关于克服算法厌恶。我在这一部分想做的就是,基本上向你们介绍我们一直在使用的研究范式,以及两个实验。
So we’ve got this initial paper, and then we came in behind that and we’ve got a paper under review now that is about overcoming algorithm aversion. What I want to do in this section is give you basically the paradigm we’ve been using, but also two experiments.
第一篇论文有三个实验,第二篇有四个,但我只给你们每个挑一个,相当于这两篇论文中最好的部分,让你们了解一下我们在这项研究中所做的工作。
There are three experiments in the first paper, four in the second, but I’m just going to give you one from each, kind of the greatest hits from each of these two papers to give you a sense of what we’re doing in this research.
我们正在研究的核心观点是,基于证据的算法的预测优于人类的预测。所以算法比人类更好。这一点早已被证实。
The idea we’re investigating is that the forecasts of evidence-based algorithms outperform human forecasts. So algorithms are better than humans. This has been long established.
这可以追溯到心理学家 Paul Meehl。Robyn Dawes 在心理学界也因这一点而闻名。
It goes way back to the psychologist Paul Meehl. Robyn Dawes is famous for it in psychological circles as well.
在这些圈子里,使用算法是毋庸置疑的。人类没有考虑到算法所考虑的因素,这是他们表现不如算法的四个主要原因之一。
And in these circles, there is no question that you should be using algorithms. Humans don’t consider things that the algorithms do, which is one of the four main reasons they do worse than algorithms.
他们没有把该考虑的所有因素都纳入模型。他们纳入了一些不该纳入的因素。他们不知道如何权衡每个属性。他们不知道应该给这些因素赋予正确的权重。
They don’t include everything that they should in the model. They do include things that they shouldn’t. They don’t know how to weigh each attribute. They don’t know the right weights to put on these things.
最后,也许是其中最关键的一点是:无论模型里装的是什么,使用者并不是始终如一地运用它。他们早上用一组权重,下午又换另一组权重;或者周一用三个因子,周二又换三个不同的因子。无论他们潜意识里在用什么算法,他们并不坚持使用。这就是我们对人类为何比算法更差的理解。
And then finally and probably most importantly, whatever they have in the model, they’re not using it consistently. They’re applying it with one set of weights in the morning and a different set of weights in the afternoon, or with three factors on Monday and three different factors on Tuesday. They’re not consistent in using whatever algorithm they’re using implicitly. So this is what we understand about why humans are worse than algorithms.
我们对人们为何抗拒算法却知之甚少。这正是我们想知道的。
We haven’t known much about why they are resistant to algorithms. That’s what we wanted to know.
此前已有研究表明,人们确实会抵触算法。有人做过类似于赛马对比的实验:给人们使用算法的机会,但他们并不去利用这个机会。所以,我们有证据表明他们不用算法,但关于他们为什么不用的证据却不多。
There has been research on the fact that they do resist them. People have run horse races essentially, they have given people the opportunity to use them, people don’t avail themselves to that opportunity. So we do have evidence that they don’t use algorithms, but we haven’t had much evidence on why they don’t use these algorithms.
这是我们的出发点。让我来谈谈我个人的出发点。我做这个项目的动机,还是来自 NFL。几年前,我做过一些关于 NFL 选秀的研究,由此被拉去给橄榄球、棒球和篮球组织做咨询。不过,我的主要工作还是与 NFL 球队打交道。
That’s where we’re coming from. Let me tell you where I’m coming from on the project. My motivation on this project came again from the NFL. I did some research a few years ago on the NFL draft and because of that, I’ve gotten pulled into consulting to football, baseball, and basketball organizations. But my main work has been with the NFL organizations.
大约从 2005 年起,我就觉得自己在球队应如何在 NFL 选秀中配置选秀资本这个问题上掌握着真理。不是说我知道该选哪个四分卫,而是我对状元签的价值与第 32 顺位签的价值有很好的判断,并且清楚这对球队的选秀策略意味着什么。
And since about 2005, I feel like I have had truth on my side with how teams should allocate their draft capital in the NFL draft. It’s not that I know which quarterback to pick, but I do have a good sense of the value of the first pick versus the value of the 32nd pick and what that should mean for their draft strategies.
时至今日,我们已经有了非常扎实的成果,也与多支 NFL 球队建立了长期合作关系,但令人略感沮丧的是,我们在改变 NFL 的决策方式上进展甚微。尽管,用一句引号里的话来说,我掌握着真理,并且比任何人都做更好的回归分析,但我们并没有产生什么实质影响。
At this point, we have very robust results, and I’ve got long-term relationships with multiple NFL teams, but it’s a little depressing how little progress we’ve made with changing decision making in the NFL. Despite having, quote, truth on my side, and a better regression than the next guy, we don’t move the needle very much.
那段经历引出了这个项目。它让我意识到,我们缺少让人们接受算法的工具。尽管我并非在推销,但多少还是在兜售一种算法。更广义地说,我在兜售一种
That experience led to this project. That experience led me to realize we don’t have the tools for opening people up to algorithms. Even though I’m not selling, I am a little bit selling an algorithm. More generally, I’m selling a
宾夕法尼亚大学凯德·马西
Cade Massey University of Pennsylvania
思维方式,而这些人对此很抵触,我们需要更好的工具来赢得这场辩论。
way of thinking about things that these guys are averse to, and we needed better tools for winning that argument.
所以这就是我的出发点:不只是出于枯燥的学术原因,而是为了取得实际进展,去更好地理解这个问题。但我越做下去,就越意识到这远不止 NFL 球队那点事。这适用于我们的许多学生,他们走向量化领域,努力在有史以来未经量化证据决策的谈判桌上争得一席之地。
So that’s where I’m coming from, trying to better understand this not just for dry academic reasons, but to actually make progress. But the more I’ve done this, the more I realize this applies way beyond NFL organizations. This applies to many of our students who are going out and working in quantitative fields and trying to get a seat at the table where decisions haven’t historically been made based on quantitative evidence.
所以我们接下来要问的研究问题是:人们为什么选择人类而非算法?然后,我们如何才能让人们选择算法而非人类?
So the research questions we’re going to ask are, why do people choose humans over algorithms, and then, how can we get people to choose algorithms over humans?
我们在这个领域做了不少研究。如你们所见,我们已发表了几篇论文。不过我们仍觉得只是刚刚入门。今天我能自信地说出下面这些内容,而再往后的推论就还是猜测性的。
We’ve done a number of studies in this area. As you’ve seen, we have a couple of papers. We still feel like we’re just getting into it. We can say confidently what I’m going to say today, and then what we say beyond that is still speculative.
我们的研究还在继续,还有更多假设要去验证。但至少这两个,我感觉可以讲了。
We’re continuing this research and we have more hypotheses we want to pursue. These two, I feel like we can give at least.
第一个问题,我能举的例子现在大家都很熟悉了,尤其是像 Waze 这类工具。比如说你在开车。开车的时候你在多大程度上会运用算法?
So for this first question, the example I would give is one that is common to all of our experiences now, especially with tools like Waze. Say you’re driving. To what extent do you use algorithms when you drive?
如果你在行驶中决定改变常规路线,结果却把自己堵在了半路上,你会作何反应?
If you decide to change your normal route as you’re driving, and you place yourselves in the middle of a traffic jam, how do you respond?
你可能会不高兴,但我认为你对自己判断的信心不会丧失太多。解释那一次到底哪里出错,对你来说大概不难。
You may be unhappy, but I suggest that you don’t lose much confidence in your judgment. It’s probably not very hard for you to explain away what went wrong on that particular occasion.
另一方面,如果 Waze 或其他 GPS 导航叫你改道,结果你同样被堵在了车流中,归因方式就大不一样了。你简直想把导航给开除了。
What happens on the other hand if Waze or some other GPS (Global Positioning System) tells you to change route and you end up in the same tracking channel? Very often, it’s a very different attribution. You kind of want to fire the GPS.
事实上,这正好是我和我妻子的亲身经历。我第一次试图让她用 Waze 时,开局不顺,我们刚用就遇到了糟糕的体验。结果又花了大概三个月,她才最终相信 Waze 比我们自己的判断更靠谱。
In fact, this was exactly my experience with my wife where I first started trying to get her to use Waze. We had the misfortune of having a bad experience with Waze right up front and it took probably three more months before she finally was convinced that it was superior to our judgment.
所以,核心观点是:我们对自己给出的指引与对算法给出的指引,推理方式是不同的。具体来说,我们目睹算法出错时,比目睹个体出错时苛刻得多。那么,为什么人们选择人类而非算法呢?
So this is the idea that we reason about guidance from ourselves differently than we reason about guidance from algorithms. And in particular, when we see algorithms err, we are much harsher on them than we are on individuals when we see them err. So why is it that people choose humans over algorithms?
第一,他们知道,在我们研究的这些领域,算法几乎注定会犯错。这不是那种可以做到完美预测的领域。
One, they see that it is almost inevitable, and in the domains that we’re studying, it is inevitable that algorithms are going to make errors. We’re not looking at domains where perfect predictions are possible.
用花哨的术语说,这就是“偶发性不确定性”——某种不可化约的不确定性。如果你问别人:“抛这枚硬币会是正面还是反面?掷这枚骰子是一点还是六点?”没有人能确知结果。在这种情况下,算法必定会犯错。这是前提的一部分。我们认为这涵盖了许多领域。
In fancy terms, these are aleatory uncertainties. There is some kind of irreducible uncertainty. If you were going to ask somebody, “is that coin going to be a head or a tail if I flip it, is that dice going to be a one or a six if we roll it,” there is no way of knowing for sure. And in those cases, it’s inevitable that an algorithm will err. So that’s part of the setup. We think that captures many domains.
第二,人们对人类犯的错容忍度更高,而在目睹算法犯错后,他们会失去对算法的信心。但目睹人类犯错后,他们却未必会失去对人的信心。这就是我们的假设所在。明确地说:看到算法犯错 → 失去对算法的信心 → 转而选择人类。
Second is that people will be more tolerant of errors made by humans and that they’ll lose confidence in algorithms after seeing it. They will not necessarily lose confidence in humans after seeing it. So that’s where our hypotheses are. To be clear, see algorithm err, lose confidence in algorithm, choose human instead.
这里我要坦白一下。我们刚开始研究时,曾追过不少不同的假设。最终,在最为纯粹的意义上我们发现,是“看到犯错”这个因素在驱动一切。我想
Let me make a confession here. We pursued a number of different hypotheses when we first started studying this. We discovered in kind of the purest sense that it was this seeing error that is driving things, and I want to
宾夕法尼亚大学凯德·马西
Cade Massey University of Pennsylvania
给你们一些这方面的证据。我先讲几个实验细节。在研究 1 中,我们有 361 名受试者,都是实验室参与者。他们需要预测 MBA 学生的成功程度。
give you some evidence on that. So let me give you a couple of experimental details. In study 1, we have 361 subjects. These are lab participants. They are estimating the success of MBA students.
我们有 115 名真实的 MBA 学生,并且知道他们的长期表现——用多种方式衡量:他们去了多好的公司工作、赚了多少钱、毕业时同学的评分、GPA。
We have 115 real MBA students, and we know how they did long-term measured in a number of different ways: how fancy a company they went to work for, how much money they made, ratings by their peers on graduation, GPA (Grade Point Average).
我们把所有这些合并成“学生成功”指标,并要求实验室参与者据此进行预测:给定这些输入,你们觉得这些 MBA 学生长期会表现多好?
We combined all that into student success, and we ask the lab participants to forecast that essentially. How good given these inputs do you think these MBA students will be long-term?
所以问题就是:他们更愿意依赖自己的估计,还是依赖一个统计模型的估计。这个问题永远存在。你可以自己做判断,也可以依赖模型来做。
So the question is whether they want to rely on their own estimates or a statistical model’s estimates. That’s always going to be the question. You can do this yourself or you can lean on a model to do it.
模型当然就是算法。为了做这个,他们最终需要做出 10 次预测。我们会设置激励:预测越准,赚的钱越多或越少。
The model is the algorithm, obviously, and to do this, they’re going to ultimately have ten forecasts. We’re going to put in incentives where they can earn more money or less money depending on how well they do.
我们告诉他们,他们将根据真实 MBA 学生的申请材料来估计这些学生的实际百分位排名,一个统计模型也会估计这个表现,然后我们给他们一份变量说明。
We tell them that they will estimate the actual percentile ranks of real MBA students based on their application, a statistical model will also estimate this performance, and then we give them a description of variables.
这就是这些变量的大致样子。输入变量包括:本科学位、GMAT 成绩、申请文书——基本上就是申请材料。
These are what the variables look like. These are the inputs: undergraduate degree, GMATs [Graduate Management Admission Test], essays, these are basically application materials.
这些都是 MBA 招生部门已知的信息。从某种意义上说,我们在复刻招生决策。实验设计是:他们要做出 15 次无激励的预测,然后得到真实结果的反馈。
These are things that MBA admissions departments know. In some ways, we’re replicating the admissions decision. The design is that they’re going to make 15 forecasts with no incentives and they get feedback on what actually happens.
用佩德罗·多明戈斯的话说,这是训练集。他们先在一个集合上学习,然后在另一个不同的集合上接受测试。所以这是一个学习机会。我们在学习阶段操纵的是:他们自己是否做出预测?他们是否能看到算法的预测?
This is the training set, if we’re going to follow Pedro [Domingos]. They’re going to learn on one set and then they’re going to be tested on a different set. So this is the learning opportunity, and then what we manipulate during this learning opportunity is if they themselves make forecasts, and if they have access to the algorithm’s forecast?
我们把这两个因素交叉,得到四个实验条件。你会被分配到其中某一个条件。具体是:你自己做预测并得到反馈(是/否),以及你是否看到模型的预测并得到反馈(是/否)。
We’re going to cross both of those things, so we have four experimental conditions. You are in one of these four conditions. Are you making your own forecasts with feedback, yes or no, and are you seeing the model’s forecast with feedback, yes or no?
再次强调,这是训练阶段。你要经历 14 次这样的训练:要么做预测,要么不做;要么获得模型的预测,要么没有。
Again, this is the training phase. You’re in 14 of these, you’re either making forecasts or not, and you’re either getting the model’s forecast or not.
我们设置了一个纯人工条件:只自己做预测并得到反馈;一个纯模型条件:不做自己的预测,只观察算法;两个对角线条件:模型和人工一起;还有一个控制条件:完全不经历学习阶段。
We have a human condition which is just making your own forecast and getting feedback; we have a model condition which is not making your own forecast, just observing the algorithm; and then we have the two diagonals which is model and human; and then the control condition where you’re not going through the learning phase at all.
所有的操纵都发生在学习阶段。受试者在某一种条件下做 15 次试验,区别只在于他们自己是否做预测并得到反馈,以及对模型是否做同样的事。
All of the manipulation happens during this learning phase. The subjects are doing 15 trials in one of these conditions, and it’s just whether or not they’re making predictions and getting feedback and/or whether they’re doing the same thing with the model.
这是他们看到的刺激示例。这是一名 MBA 学生,下面是一些该学生的基本人口统计学信息。
This is an example of the stimuli they get. Here is an MBA. Here are some basic demographics on the MBA.
你觉得这个人在 MBA 生涯中表现如何?你可能会自己做出判断。
How well do you think this person is going to do in their MBA career? You might make this judgment yourself.
你给他打多少分的百分位?启动你自己的判断模型。你可能还会考虑你对这个模型的信心。你愿意有一个统计算法来支持你吗?我希望大部分人都愿意。
What percentile would you put them in? Start cranking it through your own model. You might also think about your confidence in that model. Would you like the support of a statistical algorithm? I hope most of you would.
宾夕法尼亚大学凯德·马西
Cade Massey University of Pennsylvania
在这个例子中,平均预测值是第 75 百分位,而模型的预测值是 28。但我要说的是,这个学生实际百分位的结果——答案是 2。这个例子可能不太有代表性。我们是自然取样,你会得到一些噪声非常大的结果。
And in this case, the average prediction was the 75 percentile, the model’s prediction was 28, and I mean, the students’ actual percentile, the answer essentially was two. So that may not be a super representative one. We naturally sampled this space. You’re going to get some very noisy outcomes.
在这个刺激示例中,他们看到的是人口统计学资料,自己决定要不要做预测,模型要不要做预测,然后得到这样的反馈。
Just in the example of the stimuli they go through, they get the profile of demographics, they make a prediction or not, the model makes a prediction or not, and then they get this feedback.
这个例子来自人类与模型条件实验。我想向各位展示的是,当它们处于这些不同条件下时,显然会发生什么情况。
This is an example from the human and model condition. And what I want to show you is what happens obviously when they are in these different conditions.
所以他们每猜中一个在五个百分点以内的估值,就能拿到 1 美元奖金,也就是说,他们在这上面押了些钱。这笔钱显然不算多,但对我们的实验对象来说,积少成多,而这正是他们的基本选择。
So they’re going to get paid a $1 bonus each time the estimate is within five percentile, so they’ve got some money riding on this. This isn’t a lot of money, obviously, but for our experimental subjects, it adds up, and this is their basic choice.
做出这 15 次预测后,他们面临一个选择:接下来还要再做 10 次类似预测,他们必须决定,是继续使用统计模型,还是依靠自己的判断?
After they have done the 15, they face this choice. They’re going to make 10 more of these forecasts, and they have to decide if they want to use the statistical model or if they want to use their own judgment?
而这对于我们来说确实是至关重要的问题。他们有足够机会了解这种环境,也有机会自行预测或观察模型的预测结果。他们在没有激励的情况下做了 15 次,仅仅是学习性试验,现在他们被要求,在金钱刺激下再做 10 次:你是想靠自己判断,还是用模型?
And this is really the most important question for us. They’ve had a chance to learn about this environment, they’ve had a chance to either exercise their own forecast or observe the model’s forecast. They’ve done this 15 times with no incentives, just a learning trial, and now they’re asked, okay, ten more times for money: do you want to use yourself or do you want to use the model?
这个模型用的是同一组数据。在这个案例中,我们没有做正确的事——没用一个数据集来建模、用另一个来测试。所以它没有更深刻的信息,但确实有相同的信息。在这个案例中,模型是在样本内工作。我们在第二项研究中修正了这个问题。
The model in this case is from the same data. In this case, we didn’t do the proper thing and model him on one dataset and test on another. So it doesn’t have deeper but it does have the same. The model is working insample in this case. We fixed that in the second study.
结果是,参与者的预测比模型多出 15% 的误差,这一误差幅度稳定地高于模型。如果他们采用模型的预测,本可多赚 29% 的钱。
The outcome is that the participant’s forecasts had 15 percent more error than the model, which is reliably more error than the model. They would have earned 29 percent more money had they used the model’s forecast.
以下是学生们的表现。这是模型的表现。关键在于选择使用模型的学生百分比。我按四个条件来细分结果。
This is how the students performed. This is how the model performed. The key bit is the percentage who chose to use the model. I’m breaking it down by the four conditions.
控制组的情况是他们完全不这样做。他们遇到的第一个问题是:好吧,你要做十个这样的预测,你想用自己的判断还是用模型?他们没有进行任何学习。
The control condition is where they didn’t do this at all. The first question they got was, okay, you’re doing ten of these forecasts, do you want to use your own or do you want to use the model? They didn’t have any learning.
结果如何?大约 65% 的情况下,他们都会选择使用这个模型。
What happens? About 65 percent of the time, they want to use the model.
这有点像你们的经历。你们更愿意用模型还是自己的判断?这点很重要,因为有时我们的论文被误解成在说所有人都讨厌算法。我们没说人们讨厌算法。事实上,他们刚来这里时挺天真的,大约三分之二的人其实更喜欢算法。
This is a little bit like your experience. Would you rather use the model or your own judgment? This is important because sometimes our paper gets misconstrued as saying that everybody hates algorithms. We don’t say that people hate algorithms. In fact, they come in naively here and about two-thirds would actually prefer the algorithm.
下一个是涉及人的条件(human condition),他们自己做判断,得到反馈,但从未看到模型做任何事。结果如何?同样,大约三分之二的人实际上更愿意使用模型。所以他们意识到自己在这一点上并不完美,宁愿借助一些统计工具的帮助。
The next one is the human condition, where they made their own judgments, they got feedback, and they never saw the model do anything. What happens? Again, about two-thirds would actually prefer to use the model. So they realized they are not perfect in this and would rather use some statistical help.
另外两个条件实际上展现了这个模型的表现,我刚才跟你们报告了它在这些条件下的表现,与其他模型相比结果如何。这个模型更好。这个模型明显更好。所有看过演示的人都亲眼看到它表现更佳。
The other two conditions actually saw the model perform, and I just reported to you how the model performed relative to how they perform. The model is better. The model is demonstrably better. Everybody who saw it saw it do better.
但这些人使用该模型的实际兴趣大幅下降。在仅看到模型表现的人群中,只有 26% 选择了使用它。他们没有谦逊地去做自己的预测。所以,这或许是最糟糕的情况。
But their actual interest in using the model drops dramatically. Only 26 percent of the people who only saw the model perform chose it. They didn’t have the humility of making their own predictions. So maybe this is the worst case.
宾夕法尼亚大学 凯德·马西
Cade Massey University of Pennsylvania
但事实上,当他们看到两者时,那仍然只是一个四分之一。而我实际上监督了整个过程。因为我们自然抽取了这 115 个样本,并且有真实的受试者做出真实的预测,所以研究中的每个人并非都被模型超越。
But in fact, when they saw both, it’s still just a quarter. And I actually oversee this. Because we’re naturally drawing these 115 and because we have real subjects making real predictions, it’s not the case that everybody in the study is outperformed by the model.
有些人确实纯属运气好,表现和模型一样出色。但即便是那些亲眼见过模型比自己表现更好的人——而这占我们员工的多数——在 69% 的情况下仍然倾向于用自己的判断,而不是模型的判断。
Some people actually do as well as the model just by chance, but even those who saw the model outperform themselves, and this is the majority of our people, still prefer to choose their own judgment 69 percent of the time over the model’s judgment.
你现在可以开始理解我们为何得出这个结论了。正是看到模型出错,才引发了这种逆向反应。这不是对算法天生的反感,而是在这样一个预测困难的世界里,错误是不可避免的。当人们看到模型犯错时,他们就会惩罚这个模型。
You can start to see where we’re getting our conclusion. It’s seeing the model err that leads to the inversion. It’s not an inborn aversion to the algorithm, it’s that in this world where prediction is tough, error is inevitable. When people see a model err, they punish the model.
我们还有一些附加的衡量指标,可以更深入地了解这里的情况。例如,我们询问他们对于模型的信心,结果发现,当他们看到自己的表现时,信心并没有太大变化。你也不会指望它会变。但每次他们看到模型的表现时,信心就会受到打击。
We have some additional measures to give some insight into what’s going on here. For example, we ask about confidence in the model, and we see that confidence doesn’t change much when they see their own selves perform. You wouldn’t expect it to. But the confidence gets hurt whenever they see the model perform.
因此,置信度完全随模型选择而定——换句话说,置信度调节了他们对模型的选择。你可能会问,那人类自身的信心呢?它起作用了吗?——它不起作用。
So confidence goes exactly as the choice of model goes, in other words, confidence mediates their choice in the model. You might ask, well, what about confidence in humans? Did that play a role? That doesn’t play a role.
在全部四种情境下,人类对自身的信心都同样坚定。无论他们是否亲眼见证了自己表现糟糕,都不影响自信。他们不会因此失去对自己的信心。回到那个 GPS 的例子。他们总有办法为自己开脱——那个导致堵车的错误,不会成为他们判断力的一贯缺陷。这种信心的差异,正是理解其中机制的关键线索之一。
Confidence in humans is constant across all four conditions. It doesn’t matter whether they saw themselves perform poorly or not. They don’t lose confidence in themselves. Back to the GPS example. Somehow they rationalize how their mistake leading to the traffic jam isn’t a permanent feature of their judgment. That distinction in confidence is one of the clues to what’s going on here.
我们在几种不同的条件下、几项不同的测试中进行了这项研究。我们完成了不同的任务,而且我们始终需要更多这样的任务——在实验环境中你能进行的任务数量毕竟有限——但我们尝试过预测各种经济指标,以及不同类型的学生表现。我们还操控了算法胜过人类的程度。
We have done this in a few different conditions, a few different tests. We have done different tasks. We always need more of these tasks. There are only so many of these tasks you can run in the experimental setting, but we have tried forecasting various economic indicators, different kinds of student performance. We have manipulated the extent to which algorithms outperform the humans.
你可能会合理地猜测,当算法仅比人类强一点点时,对算法产生一些抗拒或许更合理。于是我们设计了算法远胜于人类的情境,然而,即便面对明显的错误,这种对算法的抗拒仍然顽固存在。
You might reasonably guess that when the algorithm is only a little bit better than the humans, some algorithm aversion might be more rational. So we put them in situations where the algorithm is a lot better than the humans, and yet, this algorithm aversion in the face of error remains robust.
于是我们又做了一系列研究,选择的选项不再是模型与自己判断之间的比较(这种比较会引入大量自我中心偏差),而是模型与他人判断之间的比较。结果我们发现,算法厌恶的程度几乎没有减少,只是稍有缓解。自我中心偏差似乎确实有一定影响,但总体而言,更大的效应仍然是:人们对算法的厌恶超过了对人类判断本身的普遍偏好。
And then we run studies where the choice isn’t between a model and your own judgment, which introduces a lot of egocentric biases, but between a model and somebody else’s judgment. And we see almost as much algorithm aversion. It’s mitigated a little bit. There is, it seems a role for egocentric biases, but the bigger effect still is the aversion to algorithms over just human judgment in general.
所以我们做了一些探索性研究来探讨原因。这些发现我们已写入第一篇论文,做法很简单——直接问:“如果你要比较模型和人类的判断力,你会担心哪些问题?”你觉得呢?是人类更擅长,还是模型更出色?
So we have some exploratory measures on why this is. We reported these in our first paper and we ask, we literally just ask, what are some things that you might worry about if you’re trying to compare judgments of models and humans. What do you think? Do you think humans are better? Do you think models are better?
因此,对于这些因素中的一部分,他们认为模型优于人类。对于另一部分,他们认为人类优于模型。根据我们的受访者所说,模型在以下方面比人类更胜一筹:一致地权衡信息、恰当地权衡属性,以及避免明显的错误。这就是我们受访者的观点。
So for some of these factors, they believe models are better than humans. For some, they believe humans are better than models. Models are better than humans, according to our subjects, at weighing information consistently, weighing attributes appropriately, and avoiding obvious mistakes. This is what our participants say.
这在他们偏好的心理层面也占了一部分原因。
This is part of the psychology for their preference.
人们常说,在发掘被低估的人选、识别特殊情况、从错误中学习,并通过实践不断进步这几方面,人类比模型更在行。
They say that humans are better than models at finding underappreciated candidates, detecting exceptions, learning from mistakes, and getting better with practice.
所以对我们来说,这相当符合直觉。那就是变得更好的可能性。这关乎至少大多数算法的静态特征,以及外行人对它们的印象,也关乎人类的动态特征——在那里我们
So for us, this comports quite a bit with our intuition. It’s this possibility of getting better. It’s the static feature of at least most algorithms, and laypeople’s impressions of them, and the dynamic feature of humans where we
宾夕法尼亚大学的 凯德·马西(Cade Massey)
Cade Massey University of Pennsylvania
随着时间的推移,他们能够不断学习和进步。
can learn and improve over time.
还有一些其他有趣的细节。我完全不认为“避免明显错误”应该被归入那个类别。稍后我会在演讲中再给你举一个这方面的例子。
There are also some other interesting details. It isn’t clear to me at all that avoiding obvious mistakes should be in that category. I’ll give you another example of that later in the talk.
不过,察觉异常这一点——我认为也非常接近问题的核心。这又回到我给你看的那张图表,上面是所有 NFL 比赛的数据和那条贯穿其中的回归线,以及形形色色的例外情况。
But this idea of detecting exceptions I think is also very close to the heart of it. It goes back to me showing you that graph of all the NFL games and the line through it, all kinds of exceptions.
游戏的偏离情况与我们的策略大相径庭。如果你能发现那些例外,那当然很棒,你绝对能改进算法。问题在于,你相信自己能够发现那些例外。所以,我认为这恰恰是这类偏见的症结所在——人们更偏爱人类判断,而非模型。
Games deviate dramatically from our line. If you could detect those exceptions, fantastic, you could definitely improve on the algorithm. The trouble is believing you can detect those exceptions. And so, this I think is real close to the heart of the bias here, a preference for human judgment over the models.
好的,这是第一项研究,核心结论是:人们在目睹算法和人类犯错后,放弃算法的可能性远高于放弃人类。再次强调,尽管论文标题的前两个词是“算法厌恶”,但我们并非说人们憎恨算法。而是当他们看到算法出错时(在我们关注的许多领域,这几乎是不可避免的),人们对算法的惩罚会比对待人类更严厉。
Okay, that’s study one, and the main takeaway is that people are much more likely to abandon an algorithm than a human after seeing them err. And again, despite the first two words of the title, algorithm aversion, we’re not saying that people hate algorithms. It’s that when they see algorithms err, which in many domains that we’re interested in is inevitable, they punish algorithms more than they do humans.
第二份分析更多地探讨了这个问题:我们能做些什么?有没有办法缓解这种情况?
The second paper is more on this question of what can we do about it? Are there ways we can mitigate that?
我们集思广益了众多想法,开展了大量研究,并且可以肯定,我们只找到了其中一条答案。还有更多问题有待厘清,但我确实想分享其中一项成果。
We have brainstormed many ideas, we’ve run a number of studies, and we’re sure we only have one of the answers. There are more to be pinned down, but I do want to share what we have on one of those.
那么,我们怎么能让人们去使用算法呢?关键想法在于,人们希望对自己的预测保持一定程度的掌控。他们还没准备好将预测和决策权拱手让给这些黑箱算法。
So how can we get people to use algorithms? The idea here is that people want to retain some control over their forecasts. They are not ready to cede forecasts and decision making to these black box algorithms.
思路是,当人们能够修改算法的预测时,他们就会更愿意使用它。即便修改能力受到很大限制,也是如此。
The idea is that they’ll be more willing to use an algorithm when they can modify its forecast. Even when the ability to modify is quite constrained.
你可能会想,如果让你对那个算法拥有一丁点控制权,会有什么不同。
You might wonder what difference it would make if you were just given a modicum of control over that algorithm.
这会有影响吗,还是说你需要的控制权很多?你是只需要一点控制权,还是要很多才会开始把一部分决策权交给算法?你对其他人会如何回应算法有什么直觉?这就是我们这里要研究的问题。
Would it make a difference or do you need a lot? You just want a little control or would it be a lot before you would actually start ceding some decision rights to it? And what is your intuition for how other people will respond to algorithms? This is what we’re going to investigate here.
所以总体来看,实验设置还是一样的。我们会换用另一套数据集,解决我之前提到的那个样本内问题。这次会用一些线上参与者,而不是实验室参与者。
So broadly, the setup is going to be the same. We’re going to use a different dataset. We’re going to fix this insample problem I mentioned earlier. We’re going to use some online participants this time instead of lab participants.
可重复性危机的主要教训之一是,心理学研究的统计检验力不足,样本量实在太少了。与乔·西蒙斯共事后,我开始使用非常大的样本。现在我们一般每个实验条件争取达到 200 人,比如这是一个四组实验的设计。
One of the main learnings from the replicability crisis is that psych studies have been underpowered using far too small a sample. Doing work with Joe Simmons has lead me to use very large samples. Now we try to get 200 a cell generally, where this is going to be a four-cell experiment.
816 名线上参与者,这次的任务是估算学生在标准化数学考试中的成绩。我们有一个真实的数据集、真实的高中生,他们需要为这些估算做出决定,我们会给他们提供激励。但对于这些估算,他们是愿意依赖自己的判断,还是愿意依赖模型?大致上与之前的设计相同。
Eight hundred sixteen online participants, and the task in this case is to estimate students’ standardized math test performance. We have a real dataset, real high school students, and again, they have to decide for these estimates, and we’re going to incentivize them. But for these estimates, do they want to rely on their own judgment or do they want to rely on a model? Broadly the same design as we had before.
最后,也是最重要的,操纵变量的核心是他们能在多大程度上修改模型的预测。以下是他们看到的介绍:你将估算 20 名真实高中生在标准化数学考试中的实际百分位排名,这里是一些自变量,这里有一个模型,它用你拥有的相同信息估算了所有学生的百分位。哦,顺便说一句,这个模型的平均偏差是 17.5 个百分位。我们只是提前告诉他们模型的性能。
Finally and most importantly, the manipulation is how much they’re able to modify the model’s prediction. Here is the introduction they see: you will estimate the actual percentile ranks of 20 real high school students on a standardized math test, here are some independent variables, and here is a model that has estimated all students’ percentiles using the same information you have. Oh, and by the way, the model is wrong by 17 ½ percentiles on average. We’re just telling them upfront this is the model performance.
我们这里想说的是:看,这很难。我们通常会这样说:“一位有见地、深思熟虑的统计学家”或
What we’re saying here is look, it’s hard. We usually say something like “informed, thoughtful statistician” or
凯德·梅西
宾夕法尼亚大学
Cade Massey University of Pennsylvania
“一位有见地、深思熟虑的建模者。”我们稍微给模型造了势头。
“informed, thoughtful modeler.” We pimp the model a little bit.
但在这里,我们也想给他们一个关于模型性能的真实描述。所以我们说:这是一个困难的领域,充满噪声,模型不完美,它是个好模型,但并非完美,明白吗?这就是初始设定。我们再次进行了培训,然后给了他们选择。
But here, we also wanted to give them a truthful take on what the performance is. So we’re saying this is a hard domain, it’s noisy, the model is imperfect, it’s a good model but it’s imperfect, okay? That’s the setup. And we ran the training sessions again and then we gave them this choice.
这些是他们拥有的变量,具体是什么不太重要,但这正是他们需要处理的数据。再次强调,这是一个具有挑战性的任务,表面上看就很有挑战性。我们都觉得,为什么有人不直接选择用模型呢?结果发现,他们确实不太愿意用模型,但你可以看到这与研究一有很高的相似度。
So these are the variables they have, kind of doesn’t matter but just this is what they’re working with. Again, it’s a challenging task, it’s prima facie challenging. We all kind of thought, why wouldn’t anybody just go straight to the model? It turns out that they really don’t like to go to the model, but you can see that there is a lot of similarity to study one.
这些是激励措施。不过我们试着稍微加大力度。我们只能按照这个奖金计划支付四个条件组中的一个。接着,最有趣的部分是这四个实验条件组的设置。
These are the incentives. We tried to crank it up a little bit though. We were only able to pay one of the four conditions on this bonus schedule. And then, here is the interesting bit, the four experimental conditions.
他们被随机分配到这四个实验条件组中的一个。第一个条件是:他们不能改动模型。
They were randomly assigned to one of these four experimental conditions. One is they can’t change the model.
他们必须在自己的预测和模型之间做出选择,如果选了模型,就必须 100% 使用模型的预测。
They have to choose between their forecasts or the model, and if they choose the model, they have to go with the model 100 percent.
第二个条件是:如果他们选了模型,他们可以对模型调整最多 10 个百分位。所以如果模型预测值是 37.5,他们可以把它下调到 27.5 或上调到 47.5,或者调整 5 个点、调整 2 个点……
The second is that if they choose the model, they can adjust the model by up to ten percentile. So if the model predicts 37.5, they can move it down to 27.5 or up to 47.5, or adjust by five, adjust by two . . .
他们被随机分配到这两个条件组中。我们很好奇这会如何影响他们对模型的采纳程度。我们给了他们相当多的控制权、一点控制权、以及几乎毫无控制权,来看看这对他们使用模型的兴趣有什么影响。
They were randomly assigned to one of these two conditions. We were curious how that would impact the take-up of the model. We’re giving them a fair bit of control, a little bit of control, and almost no control over the model to see what the impact is on their interest in using the model.
我认为这是论文四项研究中最有趣的一项,也是最能捕捉其核心精神的一项。这是我们偶然发现的,当时我们开始观察他们的行为。我们一直在推进“他们喜欢控制权”这个想法,并想看看我们能把控制权降到多低还能行得通。
This I think is the most interesting of the paper’s four studies and kind of captures the spirit of it most closely. This is something that we happened into as we started seeing what they were doing. We were pushing the idea that they like control, and we wanted to see how little control we could get away with.
你可以对模型调整 10 个百分位。现在,你要做若干预测。你是想用模型并做调整,还是想用自己的判断?
You can adjust the model by ten percentiles. Now, you’re going to make a number of predictions. Do you want to use the model and adjust it or do you want to use your own judgment?
这是我们的发现。再次,我向你报告的是我之前在研究一中报告的同类数据:选择使用模型的人的比例,我会按不同条件组展示给你。对于不能改动模型的条件组,大约 47% 的人选择使用模型,这介于我们之前两个条件组之间。
This is what we found. Again, I’m reporting to you what I reported to you for study one, which is the percentage of people who chose to use the model, and I’m going to show it to you by these different conditions. And for the can’t-change condition, about 47 percent chose to use the model, so it’s kind of in-between our two conditions before.
请记住,我们告诉他们模型会有误差,所以他们并没有直接体验误差,但他们知道存在噪声。他们知道该模型有 17.5% 的误差。他们看着这个数据,大约一半的人说,是的,我用模型。大约一半的人想用他们自己的判断。
Remember, we told them that it errs, so they haven’t experienced the error firsthand, but they know that it’s noisy. They know that it has 17.5 percent error. They look at that and about half of them say, yes, I’ll use the model. About half of them want to use their own judgment.
如果允许他们对模型调整 10 个点,会发生什么?对模型的兴趣显著增加。
What happens when you allow them to adjust the model by ten? Significant uptake in their interest in the model.
现在,我们又回到了超过三分之二的参与者愿意使用模型的状态,只要他们能稍微调整一下。那么,如果你允许他们调整,但调整幅度更小呢?对于 5% 的调整限制,71% 的参与者选择使用模型。然后,最有趣的是,如果你把调整幅度一路降到只能调整 2 个点,会发生什么?68%。
Now we’re back up there above two-thirds of participants willing to use the model if they can tweak it a little bit. And then, what happens if you let them tweak it but don’t let them tweak it as much? For 5 percent, it’s 71 percent of participants. And then, most interesting, what happens if you crank it all the way down and you can only tweak it by two percent, what happens? Sixty-eight percent.
这让我们深感震撼,这也是第二篇论文的核心观点。没错,这有点直觉上的道理:人们会在能调整模型时更喜欢模型;但反直觉的是:你根本不需要让他们调整太多,就能让他们接受模型的帮助。所以,我们看到了大幅提升的采纳率和模型使用率,尽管我们给了他们严格的限制。
We were really struck by this and this is the main point of the second paper. Yes, it’s kind of intuitive, people will like the model more when they can move it, but it’s counterintuitive as you don’t have to let them move it very much to get them to accept the model’s help. So we get dramatic uptake and model usage even though we’ve tightly constrained them.
凯德·梅西
宾夕法尼亚大学
Cade Massey University of Pennsylvania
人们已经表现出无法持续改进模型的能力,所以我们喜欢把他们牢牢控制在模型附近:如果能在不牺牲采纳率的前提下做到这一点,那我们就是占了上风。
People have demonstrated an inability to improve the model consistently, so we like being able to keep them very close to the model, so if we can do that without costing us uptake, then we’re ahead of the game.
事实也确实如此。大约 50% 的人选择模型、50% 的人选自己判断,他们的平均误差是 22 个百分位。
And so, that’s in fact what we see. These people went about 50 percent with the model and about 50 percent on their own, and their average error was 22 percentile points.
如果可以调整 10 个点,那么超过三分之二的人使用模型。因为更多人用模型了,而且他们实际上并未调整太多,所以误差更小。
If you can adjust the model by ten, now we have more than two-thirds of the people using the model. Because more people are using the model and they’re not all actually adjusting it by that much, they have less error.
将模型调整限制在 5 个点,误差甚至更小。调整幅度限制在 2 个点,误差也更小。
Constrain model adjustments by five, you get even less error. Adjust the model by two, less error.
于是你得到了更好的表现。数据中有噪声,所以要是每个样本都能改善就好了,因为我们知道我们把人们控制在了更接近模型的位置。使用模型的人数大致相同。你预计误差会下降。数据里确实有噪声。
So you get this better performance. There is noise in here so it’d be nice if this were each one improving because we know that we’re keeping them closer to the model. About the same number of people are using the model. You would expect this to come down. It’s noisy.
我们是如实采样的,所以有时贝叶斯性能或模型的性能并不如它应有的那么好。但这里的要点是:通过给他们这些调整选项,在所有情况下,他们的表现都优于完全不给控制权的情况。
We sampled truthfully, and so, sometimes the Bayesian performance or the model’s performance doesn’t do as well as it should. But the idea here is that by giving them these adjustment choices, in all cases, they outperform not giving them the control at all.
如果你不给他们控制权,不给他们一点控制权的选项,他们就不会用模型,表现也不好。给他们一点控制空间,他们用模型的兴趣就大得多。当他们用模型时,他们不会大幅改动它,最终表现更好。
If you don’t give them control, if you don’t give them the option of some of the control, they don’t use the model and they don’t do as well. Give them a little room for control and they’re much more interested in using the model. When they use it, they don’t push it around that much and they end up performing better.
还有额外的下游效应。我们喜欢这一点,因为我们让他们表现更好了;但同样令我们深感震撼的是其他下游效应。所以我们收集了更多指标。
There are additional downstream consequences. We like this because we’ve made them perform better, but we’re also really struck by the other downstream consequences. So we collected additional measures.
我们了解到的是:当他们对模型有一定控制权时,他们对过程的满意度会显著提高。这一点并不太令人惊讶。
Things that we have learned are that they are dramatically more satisfied with the process when they can have some control over it. That’s not terribly surprising.
还有另外几个稍微更令人惊讶的发现。参与模型、对模型拥有一定控制权,会改变他们对自己的信念,也会改变他们对模型的信念。具体来说,他们对自己的能力变得不那么自信。
We have a couple other which are a little more surprising. Being involved with the model, having some control in the model changes their beliefs about themselves and it changes their beliefs about the model. In particular, they become less confident in their own ability.
他们基本上通过这种与模型的互动学到了一些谦逊,而且,相当深刻地,他们对模型的能力变得更加有信心了。
They learn some humility basically by engaging with the model in this way, and kind of profoundly, they get more confident in the model’s ability.
所以,如果你比较没有机会用模型的人——这些人是随机分配到这些条件组的——那些有机会用模型的人,实际上增加的是对模型能力的信心。
So if you compared people who didn’t have the chance to work with the model, these are people who were randomly assigned to these conditions, those who have had a chance to work with the model actually increase their confidence in the model’s ability.
然后最后一点是:他们选择只用模型的可能性是 2 倍。我们在下游任务中又给了他们一次选择机会,这次我们不再随机分配他们到不同条件组,而是让他们自己选择想进入哪个条件组。
And then, finally, they are three times more likely to choose model-only. We give them a downstream task where they again get to choose, and this time, instead of randomly assigning them to these different conditions, we let them choose which condition they want to be in.
再次说明,我们试图更贴近现实世界。我们试图更接近你的世界:在这个世界里,他们可以选择是否使用算法,是否对员工强加算法。如果他们之前有机会摆弄过模型、在早期轮次中有过一些控制权,那么他们更有可能选择只用模型的世界——而这正是预测的最优世界。
Again, we’re trying to get closer to the real world. We’re trying to get closer to your world where they get to choose whether or not to use algorithms or not, whether they impose algorithms on their employees. They’re much more likely to choose the model-only world, which is the optimal world for prediction here, if they’ve had a chance to play with the model and if they’ve had some control in the earlier rounds.
所以人们更愿意使用模型。你限制他们到多大程度——甚至我们那 2% 的限制——这都不重要。使用模型还会带来这些积极的下游效应。
So people are much more willing to use the model. It doesn’t matter how much you constrained them to the limit, to our two percent limit even, and there are these positive downstream consequences from using the model.
我们的主要结论是:在决定是否选择使用模型时,人们对他们能调整算法的程度并不敏感。
Our main takeaway is in deciding whether to choose to use the model, people were insensitive to the amount by
凯德·梅西
宾夕法尼亚大学
Cade Massey University of Pennsylvania
人们似乎希望对算法拥有一定控制权,但不一定需要更大的控制权。
which they could adjust the algorithm. It seems people want to have some control over the algorithm, not necessarily greater control.
我们最初的研究问题是:为什么人们会选择人类而非算法?以及我们如何能让人们选择算法而非人类?我们的答案是:他们对算法的错误容忍度更低。这就是原因。
The research questions we started out with were why do people choose humans over algorithms? And how can we get people to choose algorithms over humans? Our answers are they are less tolerant of algorithms’ mistakes. This is why.
那么,我们如何修改这一点?我们可以通过让他们修改算法来实现,哪怕只是一点点。所以这就是研究部分。我想分享这些内容。我尽量精简到只讲两个实验,让你感受一下那两篇论文的核心精神。
And then, how can we modify that? We can modify it by letting them modify the algorithm even if just a little bit. So those are the research parts. I wanted to share those, too. I tried to hew it down to just two experiments, give you a sense of the spirit of those two papers.
我想用一个真实世界的例子来结束,这个例子是我同时在做的:沃顿商学院的录取过程。并不是我们做了这项研究,然后我跑去应用它。我实际上是同时在推进这两件事。
I want to close with a real world example, and I was doing this example concurrently, which is the admissions process at Wharton. It’s not that we did this research and then I went out and applied it. I’m literally pursuing these two things concurrently.
你们中有多少人看过蒂娜·菲主演的那部电影?我想它叫《 Admissions 》(录取通知书)。片中她是在普林斯顿。就像你从背景里的橙色能看到的,她是普林斯顿大学的招生官。据我所知,这实际上非常真实地反映了招生工作的实际情况。
How many of you have seen this movie with Tina Fey? I think it’s called Admissions. She is at Princeton. As you can see from the orange in the background, she’s the Princeton admissions officer. From what I’m told, this is actually a very representative take on what admissions is like.
几年前,院长让我参与学院的 MBA 招生工作。所以,过去两年我一直和招生团队紧密合作,试图把我们这一行的最佳实践引入到招生办公室。普林斯顿的做法和其他所有常春藤盟校没什么两样,我们把它看作是招生 1.0 版本。
A couple of years ago, our dean asked me to get involved with our MBA admissions. And so, for the last two years, I have worked very closely with those guys trying to bring the best of our world into that office. What Princeton does is no different than what all the Ivy League schools do, and we think of it as kind of Admissions 1.0.
上世纪 70 年代,一些非常聪明的人决定要改进那些沿用了几百年、同时也是人们长久以来入学所依靠的传统体系,于是他们设计出了这套系统。这套系统不仅是这些学校的模式,也成了全美所有招生机构效仿的样板。它是一个非常个案化、费时费力、因人而异、每一个细节都很重要的过程。我们认为,有些地方可以做些改进。
Some very smart people back in the ’70s decided they were going to improve on the legacy systems that had been the way people had been admitted for centuries, and they designed this system. It’s been the model for not just these schools but for everybody in admissions around the country. And it’s a very case-by-case, laborious, individual, all-the-details-matter kind of process. And we think that there are some ways we can improve it.
所以我们一直在做的,就是我们谦虚地称之为 招生 2.0。招生 1.0,就是那部电影里描绘的普林斯顿模式——读申请材料、讨论、然后作出个人决定。
So what we’ve tried to do is what we so humbly call Admissions 2.0. Admissions 1.0, the Princeton model as depicted in that movie, is read the files, debate, and then make an individual decision.
我这么说的意思是,他们真的会围坐在会议桌旁整整一周,面前堆着一摞摞档案,然后逐个做出决定,要对 1000 个项目通过或不通过,从周一早上一上班就开始,一直干到周五下午全部完成。
What I mean by that is they are literally sitting around a conference table for a week with the stacks of files, and they are deciding one-by-one to go through 1,000 up-or-down, and they do this from first thing Monday morning until Friday afternoon when they’re done.
你可能会担心,他们在整整一天里的决策到底有多系统、有多一致,也可能会好奇:面对 500 个录取名额或 1000 个录取名额的招生组合,他们如何做出最优决策?如果像我们担心的那样、也是我们试图解决的那样——你是在依次逐个、逐案例地做决定,那又怎么才能做出整个组合的最优决策呢?
You might worry about how systematic or how consistent they’re being over that full day, and you might also wonder how do they make an optimal decision for the portfolio of whatever it is: 500 admitted students, 1,000 admitted students. How do you make the optimal portfolio decision if you’re making sequential one-by-one, case-by-case decisions, which is what we worried about and which is what we’ve tried to address.
当院长第一次邀请我参与时,他大概以为我会像之前跟几支体育队做的那样,搞点什么预测类的事情。但当我开始研究这问题时,我的兴趣与其说是改进预测,不如说是改进更广泛的决策流程。
When the Dean first asked me to get involved, I think he thought that I would do some kind of forecasting thing as I have begun doing with some sports teams. But when I started looking at this, my interest was less in improving the forecast than it was improving the broader decision process.
所以我们做的事情是,先在顶层做预测,然后关键环节是中间这个优化过程。我们试图构建一个投资组合,就像你们在七秒内完成优化那样,而不是花四天半时间逐一做决策。而一旦有了这个系统,我们就会根据每年学到的东西来评估并改进它。
So what we do is we do the forecasts up top, and then the key bit is this optimization in the middle. We’re trying to make a portfolio, we’re trying to optimize the portfolio as you guys would do in seven seconds as opposed to making case-by-case decisions for four-and-a-half days. And then we are trying, once we have this system, to evaluate and refine that system based on what we learn each year.
我想说,我们现在多少是公开谈论这件事的,否则我也不会坐在这里聊它,但我们目前并没有大力向外推这个说法。我们希望能最终做到。
I will say that we’re talking a little bit publicly here or else I wouldn’t be here talking about it, but we’re not pushing this story out there much right now. We hope to eventually.
宾夕法尼亚大学的凯德·马西
Cade Massey University of Pennsylvania
我们实际上已经按照这个模式录取了一个班级的学生。好吧,准确地说,我们已经有一整年的学生按照这个模式在学,而且刚刚录取了第二个班。一旦我们可以宣布,按照这个模式录取的学生已经顺利毕业,并且这套体系确实行得通,我们就会开始更公开地谈论它。
We’ve actually admitted one class on this model. Well, actually, we have a full year of students under this model, and we’ve just admitted our second class. We will begin talking about it more publicly once we can say we just graduated the kids that we admitted under this model and it’s actually a working system.
关键是这种组合方法。模型里的每一项输入都是主观的,每一个数据点都带有主观判断。所以,我们依然借助读者,依然依靠所有招生专家。但一旦数据进入模型,基本上我们只需转动曲柄,就能强制优化过程对所有人保持一致性。
The key part is this portfolio approach. Everything that goes into the model is subjective. Every single input is subjective. So we’re still using the readers, we’re still using all the admissions experts, but once we have it in the model, basically we turn the crank and force the optimization to be consistent across everybody.
这本质上就是一个最大化模型,在给定所有阅读过这些文件的读者的预测后,我们试图尽可能多地实现我们的目标。一切都是主观的,一切都由人决定,所有的聚合都通过模型问题系统地完成。
It’s literally just a maximization model where we’re trying to get as much of our objectives as possible, given the forecast of all the readers who have read these files. Everything is subjective, everything is human, all the aggregation happens systematically through the model question.
这将带我们进入比我当前打算讲述的更深入的细节。不过我们有自己的目标,这些目标并不存在多少争议。我们眼下能看到的是他们在校园里两年期间的表现。
This is going to take us into more detail than I’m going to actually talk about right now. But we have our objectives. They’re not very controversial. What we can see right now is what they do in their two years on campus.
显然,我们真正关心的是他们长期职业生涯的发展,但我们在校园里看到的远不止 GPA 这么简单。因此,我们设定了更广泛的目标集合,并会针对这些目标进行预测。
Obviously, what we really care about is long term what they do with their careers, but what we see on campus is broader than just GPA. So we have a broader set of objectives. And so, we’re going to forecast against those objectives.
我们将此作为一项预测任务向所有读者说明,当他们阅读文章和推荐信时也是如此。这确实是一项预测任务。所以请给出你们的预测,我们还有其他考量因素。
We framed it to all of our readers as if it’s a forecasting task when they read essays and letters of recommendation. It is a forecasting task. So give us your forecast, and then we have other considerations as well.
这是整个故事中重要的一环。这就是 Hilly Einhorn 所说的“接受错误才能减少错误”。我们在投资组合层面做的事情,就是对所有标的都一视同仁地配置权重,无论是你第一只读到的股票、第 67 只读到的,还是第 670 只读到的,得到的权重都一样。如果你今年提交了一份申请,明年提交一份完全相同的申请,你受到的对待也会一样。
This is an important part of this story. This is Hilly Einhorn’s “Accepting Error to Make Less Error.” We are doing something at the portfolio level where we’re applying all of our weights consistently, there is no difference whether you’re the first file read, the 67th file read, or the 670th file read, you’re getting the same weight. If you came in with an application this year, identically the same application the next year, you’ll be treated the same.
现在更加公平、一致了。一切都系统化了。
There is much more fairness, consistency. Everything is systematic now.
我们放弃的是这种纠结:“达拉斯高地公园的乔,和加州门洛帕克的金尼,到底谁更优秀?”我们会做出判断,但用的是模型来决策,而且方式非常一致。
What we give up is this whole, “Is Joe from Highland Park in Dallas better than Ginny from Menlo Park in California?” We make that decision but we do it in the model and we do it in a very consistent way.
我们换回的,是四天半的时间,可以用来更扎实地判断,这到底是不是一个正确的投资组合决策。现在我们能讨论的,不再只是“班级里女性比例应该定多少才对”这种问题——我们还可以做敏感性分析:如果把比例从某个值调到另一个值,结果会怎样?还有其他任何你关心的因素,都能一并跑一遍,看看影响是什么。而我们真正在乎的,又是什么呢?
What we get back is four-and-a-half days where we can be more robust on whether it’s the right portfolio decision to make. Now we can argue about what’s the right percentage of women to have in the class? Not just argue about it, we can run sensitivity analysis asking what would happen if we cranked the percentage from this to that? And any other considerations you want, we can run them all and ask what’s the impact? And what do we care about?
我们做着所有这类政策考量。一方面,我们只是在最大化这些预测,但同时也让它们受到所有其他政策考量的约束。正因为我们这样操作,现在才能讨论这些政策。我们能把时间花在讨论政策考量上,而不是逐一审查个案。
We have all these policy considerations. On the one hand, we’re just maximizing these forecasts but we’re also subjecting them to constraints on all other policy considerations. Because we do it this way, now we can talk about these policies. We can spend our time talking about the policy considerations instead of case-by-case.
心理学家们会告诉你——尤其是决策科学家们特别会告诉你——人类判断的最大缺陷,不在于没有考虑对的因素,甚至不在于权重给得不对。决策科学家会说,给他们单位权重就行。实际的权重并没有那么重要。
The psychologists will tell you and the decision scientists especially will tell you that the biggest flaw on human judgment isn’t not considering the right factors or even getting the weights right. The decision scientist will say just give them unit weights. The actual weights don’t matter that much.
最关键的,是成一个体系,日复一日、始终如一地运用同样的标尺。
What matters the most by far is being systematic, being consistent in applying those same weights day-in, day-
走出去了。这正是我们在这里尝试做的事情,也是我们一年半以来一直在探索的方向。那么,我们学到了什么?我主要想强调几件事。其中一些是关于基本决策流程的东西——把流程做对、把关系排好优先级、做个好的翻译者。我想强调两点。
out. That’s what we’ve tried to do here and that’s what we’re a year-and-a-half into exploring. So what have we learned? I want to emphasize mostly a couple of things. Some of these are basic decision process stuff, getting the process right, prioritizing relationships, being a good translator. I want to emphasize two
宾夕法尼亚大学的 凯德·马西
Cade Massey University of Pennsylvania
有四和五这两件事。
things, numbers four and five.
第四点:给决策者留有余地。正如我所说,我是在研究算法厌恶的同时做这件事的。当我们告诉招生办公室这些不会是约束性决定时,我们获得了更多的支持。每轮我们会给招生办公室 1000 个推荐名额,1200 个推荐名额给
Number four: cut decision-maker slack. As I said, I was doing this concurrently with the research on algorithm aversion. We got much more traction with the admissions office when we told them these aren’t going to be binding decisions. Each round, we’re going to make 1,000 recommendations to you, 1,200 recommendations to
你。用不用它们,全由你自己决定。对每位申请者,我们会根据模型给出一个“1”或“0”,即录取或不录取。顺便说一句,模型是你建的,数据是你输入的,我们只是机械执行,对吧?我们会给出 600 个、1200 个“1”和“0”,你自己决定怎么处理它们。
you. It’s up to you on whether you use them or not. For every applicant, we’ll give you a one or a zero, admit or not, according to the model. By the way, it’s your model and it’s your input, so we’re just turning the crank, right? We’re going to give you 600, 1,200 ones and zeros, you decide what you want to do with them.
事实上,我们意识到在优化过程中必须预留弹性空间,因此我们并不会真的分配全部 600 个名额。我们会分配其中一部分,因为我们知道需要为之后的主观调整留出余地。
And in fact, we have learned that in the optimization, we have to build in the slack, so we don’t actually allocate all 600 slots. We’ll allocate some fraction of that knowing that we want to give them some room for subjective changes to it after the fact.
这一点至关重要。你也可以把它看作一种策略。这是我们让他们参与进来的方式。就像那个“算法厌恶”实验——你说,你可以把这个模型调整 10 个单位,或者调整 5 个单位。如果我们给他们这种调节空间,他们对这个模型的兴趣会更大。但还有一个原因:另一个原因就是第五点,即监督你的算法。
That has been critical. And you could think of it as strategic. It’s our way of getting them to buy in. It’s like the algorithm aversion experiment where you say you can adjust this model by ten or you can adjust this model by five. They will be more interested in the model if we give them that slack. But there is another reason: the other reason is number five, the supervise-your-algorithm reason.
所以,我用一个关于再次犯错的简短故事来收尾。你看,迈克尔,我这是在呼应你的主题。
And so, I’ll close with a quick story on being wrong again. See, I’m book-ending it, Michael, with your theme.
对这些算法进行监督,确有其必要。我知道,跟你们做的某些事情相比,这个例子有点小儿科,但它很重要,影响着很多人的生活。算法有多危险,我深有体会。它就像一把电动工具,干活又快又省力,但你得小心。这把工具也能造成很大的破坏。
The need for supervision is real with these algorithms. And I know this is kind of a low-powered example relative to some of the things you guys do, but it’s an important example. It affects a lot of people’s lives. And I’ve been humbled by how dangerous algorithms are. It’s like a power tool that does quick efficient work but you’ve got to be careful. You can do a lot of damage with the power tool.
来举一个相当残酷的例子,这个例子我以后大概会学乖,不会再公开讲,但这次是第一回公开说。我们第一次这么做的时候,实际上是针对 MBA 第一轮录取面试人选做出决定。
So one really brutal example which I’ll probably learn not to give on the record, but I’ll give for the first time on the record. The first time we did this, this was literally making decisions on who to interview for MBA admissions round one.
历史上,招生人员有一套评分体系,1 分是最好的,4 分是最差的。你看一篇文书,1 分最好,4 分最差;你看一封推荐信,1 分最好,4 分最差。我们收集了所有这些数据,建立了一个模型,一直在跟招生人员互动,我们知道评分标准就是这样。可等到我们真正要运转这个模型时,我们做了什么?
Historically, the admissions folks had a ranking system, and one was the best you could get. You read an essay, a one is the best, four is the worst. You read a letter of recommendation, one is the best, four is the worst. So we collect all these data, we build a model, we’re interacting with them all the time, we know that this is the scale, and yet, when we go to turn the crank, what do we do?
我们会做到极致。我们走进会场,展示我们的成果,所有数据都摆在眼前,我们逐一讨论,直到会议进行到 15 到 20 分钟,才意识到我们还没有搞清楚这个扭曲到极致的最差投资组合。
We maximize. We get into the meeting, we’re presenting our results and we’ve got everything showing up, we’re talking it through, and we are 15, 20 minutes into the meeting before we realize that we haven’t figured out this distorted worst possible portfolio.
因为我们把这事儿搞反了方向。这令人警醒,也正因如此,摆弄电动工具时,你得戴上护目镜、戴好手套。它需要一点监管。
Because we ran the thing in the wrong direction. That’s humbling and that’s also why, playing with a power tool, you’ve got to have the safety glasses on, got to get the gloves on. It needs a little bit of supervision.
我们还有过几次类似的时刻,让你意识到不和算法打交道的妙处在于:你们这些人虽然会犯各种稀奇古怪的错误,但不太可能因为一个大错就伤及一大群人。而用算法的话,错误就不再是零星的了。你要是搞砸一个大问题,一次就能把一大票人全带沟里去。
We’ve had a few other moments like that where you realize the beauty of not working with algorithms is that for all you guys who are going to make idiosyncratic error, you’re unlikely to get one big thing wrong that hurts a lot of people. With an algorithm, error is not idiosyncratic anymore. You get one big thing wrong, you can take out a whole bunch of people in one go.
好了,我用 xkcd 网络漫画里的一幅小漫画来收尾。你们很多人可能都知道这个漫画。它巧妙地把算法按复杂程度排成了一个连续谱,漫画作者用 xkcd 的话说:“实际上,在复杂度的世界里,佩德罗·多明戈斯那些深度学习的东西根本不算什么,最复杂的算法是一个教堂团体花了 20 年搭起来的一张庞大 Excel 表格。”
Okay, so I’ll close with just a quick cartoon from “xkcd”, the online comic strip. A lot of you guys probably know these guys. It’s a clever little continuum of algorithms by degree of complexity, where our favorite cartoonist at xkcd says, “actually, in the world of complexity, forget anything Pedro [Domingos] does, his deep learning stuff, the most complicated algorithm is a sprawling Excel spreadsheet built up over 20 years by a church group in
宾夕法尼亚大学的凯德·马西
Cade Massey University of Pennsylvania
内布拉斯加州来协调他们的日程安排。”
Nebraska to coordinate their scheduling.”
我最喜欢这一点的是,它并没有说人们排斥算法。事实上,他们会使用算法,但他们需要在一定程度上参与算法过程。这或许不是最优解,但如果你希望他们对算法投入承诺,你可能愿意偏离最优方案。我敢打赌,那些内布拉斯加州的教友对那个算法是充满承诺的。好了,各位,谢谢,祝你们算法愉快。
What I love about this is it doesn’t say that people are averse to algorithms. In fact, they will use algorithms, but they need to have some involvement in the algorithm. This may not be optimal, but you may be willing to go away from optimal if they are going to be committed to the algorithm. And I bet those Nebraska churchgoers are committed to that algorithm. Okay, guys, thank you and happy algorithming.
提问者:关于您的那个实验——允许正负 2、5 和 10 个变动——您是否考虑过或研究过,反过来允许用户在试验中纳入条件信息,比如说,在某个时间段内重复 2 次、5 次、10 次?
Question: With respect to your experiment where you allow plus or minus two, five, and ten changes, have you thought about or looked at instead, allowing the user to incorporate conditioning information, you know, X times in the trial? So two, five, ten times over the course of the period?
我要说的核心是,在构建财务模型时如何思考这个问题。你喜欢做算法推演,但如果石油输出国组织(OPEC)开会了,或者爆发金融危机,或者发生恐怖袭击,你可能会判断,我的模型参数校准在当前时点很可能不适用。这会帮我避免那些真正糟糕的结局——那些一旦出错就会很要命的错误。我想,如果你允许人们有一定的控制权,你可能会看到更多的人愿意参与进来。
What I’m getting at is in thinking about it with respect to financial models. You like doing algorithms, but if there’s an OPEC (Organization of the Petroleum Exporting Countries) meeting or there’s a financial crisis or there’s a terrorist attack, you might evaluate that the calibration of my models are probably inappropriate for this point in time. And that would help me avoid those really bad outcomes that can come, the errors that can be really bad. I’d imagine that if you allowed people control, you might see more buy-in.
凯德:这个观察很到位。事实上,在我们那篇论文的第一项研究中,我们设置了两个不同条件。一个条件是,你可以对任一推荐结果做一定程度的修改。
Cade: It’s a great observation. In fact, in one of our studies, in the first study in that paper, we had two different conditions. One was you can modify any given recommendation by a certain amount.
另一个条件是,你可以跟任意多家这类公司签约,按你的方式来操作,所以这更接近你提到的欧佩克会议那种情况。我们得到的接种率也处在同等水平。他们喜欢这样。
The other condition was you can strike any number of them and do your own thing, so it’s closer to what you’re talking about with the OPEC meeting or whatever. And we get the same comparable levels of uptake. They like that.
我们当时只是想,进一步排列组合下去,跟另一个选项打交道更有趣。但作为一名建模者和预测者,我承认这确实至关重要,可这也恰恰直击问题的核心——你怎么知道什么时候破例是可行的,什么时候又不行?
We just thought for further permutation, it was more interesting to play with the other. But as a modeler and a forecaster, I agree that it’s critical but it’s also just right back to the heart of the problem—when do you know that exception is okay and when is it not okay?
提问:您如何看待心理学中所谓的“断腿”问题,即临床医生过于频繁地推翻公式的风险?
Question: How do you think about the “broken leg” problem from psychology, where you have the risk of clinicians overriding the formulas too frequently?
凯德:没错。假阳性实在太多了。我在 NFL 选秀研究上的合作者是我的导师迪克·泰勒,他在金融行业有些经验,我们跟橄榄球队打交道已经有十到十二年了。早年有一次,我们跟一位教练坐下来谈比赛日的决策。在这些会谈中,我们尽量保持谦逊,确实如此。
Cade: Right. There are way too many false positive. My collaborator on the NFL draft research is my advisor, Dick Thaler, who has some experience in the finance industry, and we have been working with the football teams for 10 or 12 years now. Early on, we had a sit-down with a coach and we were talking about game day decision making. We try to be humble in these meetings, we really do.
但就这么一个教练,在这么一种情况下,我们当时在讨论第四档进攻,我记得是讨论是否要尝试第四档,这个问题有海量数据支撑,而且已经被从各种角度反复研究过。我们正聊着这个,那个教练却问,那风的影响呢?
But this one coach in this one situation, we were talking about fourth down, I think, going for it on fourth down, and there’s just so much data on this and it’s been worked over in so many different ways. We’re talking about this and the coach asks, what about the wind?
你知道风向当然会起作用,这肯定有影响。但他就是接连抛出一串问题:要是刮风怎么办?左护锋又怎么考虑?总有别的因素让他惦记着,他就是不肯接受那个模型。所以每次我们跟他聊着聊着,一听到那种“万一起风了怎么办之类的问题”时,这简直就成了一个持续的笑料。(笑声)
And you know that the wind is going to matter, of course it matters. But it was just that he asked a series of these questions: What about the wind? What about the left guard? There’s always some other consideration he wants, he’s never going to accept the model. So it’s like an ongoing joke when we have these conversations with the “what about the wind type questions.” [Laughter]
提问者:我只是好奇,在招生过程中,教职员工是否觉得课堂的构成、学生的行为或能力方面有任何变化?
Question: I’m just curious on the admissions process if the faculty feels there has been any change in the composition of the class or the behavior or the capability?
凯德:是的,是的。所以我们招生办、招生主管、一直在这件事上做我搭档的招生副院长,她是从交易领域过来的,所以她在这些方面相当老练。然而,她还是有点担心。
Cade: Yes, yes. So our admissions, the head of admissions, the vice dean of admissions who’s been my partner in this all along, she came from the trading world, and so, she is pretty sophisticated on these fronts. And yet, she was a little worried.
宾夕法尼亚大学的凯德·马西
Cade Massey University of Pennsylvania
所以那年春天我们有个活动,那些学生是冬天录取的……我们为所有被录取的学生办了这个春季活动。他们来学校参观,我们差不多是想说服他们选择这所学校。他们有的会来我们学校,有的会去其他学校,她第一年去参加这个首次活动时还挺担心的。
So we had these events in the Spring, they were admitted over the Winter . . . we have this event in the Spring for all admittees. They come visit school, we kind of try to sell them on the school. They’re going to our school, they’re going to other schools, and she went to this first event the first year and she was worried.
她后来说,进门之前还挺担心,怕里面会跟《星球大战》里那个酒吧场景似的。[笑声] 你懂吧,就是满屋子奇形怪状的人?她以为我们这儿会那样——这想法可有点不着调了。她一半是拿我开涮,一半倒也是认真的。
She said afterwards she was worried that when she went in, it was going to be like going into the bar scene in Star Wars. [Laughter] You know, with all the freaks? She thought that’s what we were going to have, which is misplaced. She’s partly ribbing me but partly serious.
这种定位并不准确,因为我们所做的,本质上只是把他们过去三四年所做决策背后的判断,系统性地总结归纳了出来。我们真正做的,只是把他们过去几年中做得稍微不那么系统化的事情,用更系统的方式继续做下去。
It’s misplaced because all we’ve really done is codified the judgment that we pulled from what they’ve done the previous three or four years. All we’ve really done is do systematically what they’ve been doing a little less systematically in previous years.
而这让我困扰了一阵子,因为我来自决策的世界,几十年来我们一直痴迷于偏见问题。不过我们现在不是在消除偏见。我们稍后会谈到这类事情。我们将在未来应对那个挑战。
And this troubled me for a little while because I come from the decision-making world, and we have obsessed for decades now about bias. And we’re not fixing bias right now. We will move on to that kind of thing. We will tackle that challenge down the road.
但我们现在所做的,只是变得更加系统化,而我过去的领域并不聚焦于此。丹尼尔·卡尼曼终于开始谈论这个话题,迈克尔今年也在我们大会的开幕环节出席了。
But what we’ve done right now is we have just been more systematic, and my field was not focused on that in the past. Danny Kahneman has finally started talking about this some and Michael was there in the opening session of our conference this year.
我们这个领域可能过度担忧偏差(bias),而对噪音(noise)的重视却远远不够。这里恰恰就是一个为噪音操碎了心的绝佳案例。
It’s possible that our field is worried too much about bias and not enough about noise. And this is a great example of worrying a lot about noise.
这个系统里杂音实在太多了,但即便我们不做任何去偏(de-bias)处理,哪怕只是把他们一直在做的事情程序化,也能去掉大量噪音,而且所有参与方都觉得这样更好,因为这套系统更公平、更公正、更系统化。所以绕了一大圈,归根结底就是——他们基本还是同一批学生。目前我们没看到班级构成有任何变化。
There has just been too much noise in the system, but even if we don’t de-bias anything, even if we just codify what they’ve been doing, we’re wringing a lot of noise out of it and everybody involved feels better about it because it’s a more fair, just, systematic system. So that’s a long way of saying they’re the same students basically. We don’t see any changes in class composition right now.
提问者:你在给不同组别调整百分位的灵活度上赋予的权限不同,权限更大的那些组是真正利用了它,还是没怎么用?他们的利用模式跟增加幅度有关联吗?
Question: When you gave your different groups different amounts of latitude in adjusting percentiles, did the ones with more latitude use it or did they not use it much? Were there any patterns of use based on how much they add?
凯德:所以我们有三挡——十挡、五挡和两挡——你会发现用十挡的人比用两挡的多,因为他们额度大,能用的地方多,但即便如此,他们的使用程度也远未达到本可达到的水平。而且具体情况因人而异,说明他们在认真关注每个案子的细节。
Cade: So we had ten, five, and two, and you see ten’s use it more than two’s because they have more to use than two’s, but they don’t use it anywhere near as much as they could. And it varies by case, so they are paying attention to the details of the case.
具体平均数我记不太清了,但大概也就是 4% 左右,所以他们实际上真正动用的自由度只占全部权限的一小部分。这对我们来说是个相当明显的信号——我们可以把他们的权限收得更紧。最后我们确实把权限收紧到比他们实际使用的还要小。这些按理说应该是严格的约束条件,但他们似乎并不怎么在意。
I forget the exact averages but it’s going to be something like four [percent], so it’s really only a fraction of the discretion that they actually had, which was a pretty good clue to us that we could constrain them more. We ended up constraining them more than they were using. These should have been binding constraints, but they just didn’t seem to mind very much.
问题:你好,我得说演讲很精彩,但我有种发自内心的负面反应。
Question: Hi, I have to say great presentation, but I had a visceral negative reaction.
凯德:(笑声)太棒了。
Cade: [Laughter] Excellent.
提问:感觉你做的事就是在骗别人接受糟糕的算法。我想回到刚才那个橄榄球剧本的话题。作为风险管理者,我看到那个剧本时想说的是:好,把你基于失误差的条件结果给我看。把你预测的比分给我看——如果你把所有射门得分都用它们的期望值替换掉。我想把噪音剔除,直到得到一个没有太多噪音的算法。现在,这里面仍然有很多噪音,因为失误本身就很随机,射门得分也很随机……
Question: It feels like what you’re doing is tricking people into accepting bad algorithms. And I want to go back to the football script which is out there right now. When I see that as a risk manager, what I want to say is, okay, show me the results conditional on the turnover differential. Show me your prediction, the score, if you replace all the field goal results with their expected value. I want to take the noise out until I can get to an algorithm that doesn’t have a lot of noise. Now, there is still a lot of noise here because turnovers are pretty random and field goals are pretty random . . .
宾夕法尼亚大学 凯德·马西
Cade Massey University of Pennsylvania
凯德:是的。我对此深表理解。在我确信某个算法是最优选择之前,我不想用它来误导人。关于“误导”我还想说最后一点——我是认真的:我现在对数据的敬畏,比人生中任何时候都要更强。
Cade: Yes. I’m sympathetic to that. I don’t want to trick people into an algorithm until I’m very confident that it’s the best option. Let me say one last thing on tricking. I am serious when I say that I’m more humble now about data than I’ve ever been in my life.
而据我所知,那些最擅长模型和数据的人,其实恰恰是最谦逊的。他们最终也会被现实所教训。你和他们聊上一会儿,然后反过来,你会发现数据本身也有它的局限——你也会因此变得谦逊。
And I think the people I know who are best with models and data are actually the most humble. They end up getting humbled. You’re talking with them a while and then on the other side of that, you actually get humbled about what data can do.
所以我不会去骗人。我打心底认为,在这里或那里保持一点参与会更好。但在与机构合作时,我也在玩说服的游戏。所以,如果某种工具有用,我就会用它。
So I don’t want to trick people. I honestly think it’s better to have a little bit of involvement here and there. But I’m also in the persuasion game when I’m working with organizations. And so, if it’s helpful then I’ll use whatever tool is helpful.
这个橄榄球的例子很恰当,因为这个模型我花了五六年时间才搭建起来,而且我也不会放弃继续改进它。每个休赛期,我们都会对它做一点微调。
The football example is a good one because that model is one that I’ve built over five or six years now, and I’m not giving up on getting it better. Every offseason, we tweak it a little bit.
但这个模型是目前能找到的最好模型,人们不可能战胜它,就是不可能。想找出这个模型的例外情况,几率微乎其微。所以我很乐意推广它,因为我相信这是最合理的判断。不是说它会完美无误,因为这个世界本就充满艰难。
But that model is as good a model as there is out there and people aren’t going to beat it, they’re just not. The odds of someone identifying the exceptions to that model are just exceedingly rare. And so, I’m very happy to kind of push it because I believe it’s the best judgment. That’s not to say it’s going to be perfect because it’s a very hard world.
我想这其中的界限很微妙——我知道它还有改进空间,每个休赛期都想让它变得更好,但同时,它也已经近乎完美了。我可以从另一个足球领域给你举个例子——NFL 选秀,我早期的工作就是在那里完成的。要判断 NFL 选秀中哪个四分卫会比另一个四分卫更出色,真的非常困难。要预测这些大学新秀的场上表现,也真的很困难。
I guess it’s a fine line because I recognize that it can be improved and I want to work every offseason to make it better, but at the same time, it’s about as good as it gets. I could give you the example from another football domain—the NFL draft where my early work was done. It’s really hard to pick which quarterback is going to be better than the next quarterback in the NFL draft. It’s really hard to forecast the performance of these college kids coming out.
这并不容易。如果你看看随着时间的推移,这方面的改善有多明显,就会感到这项任务非常令人谦卑。这是一项极其困难的任务。拿球员从大学出来时的评估结果——比如依据他的选秀顺位——与某种长期表现指标之间的相关性来举例:首发场次、职业生涯收入,随便什么,只要是个相关性就行,好吗?我们可以选取很多不同的数字,只要给我一个相关性数据,然后看看它从 80 年代中期到本世纪头十年中期这 20 年间发生了怎样的变化。
It’s not easy, and if you look at how much this has improved over time, this is a very humbling thing about this task. It’s a very hard task. Take a correlation between how a player is evaluated coming out of college, say, by where he is drafted and some measure of his long-term performance: games started, career earnings, whatever, some correlation, okay? We could pick a lot of different numbers, just give me a correlation and ask how it’s changed from like the mid-80s to the mid-aughts, 20 years’ worth of work.
随着计算机和大数据的兴起,这种相关性发生了怎样的变化?有些相关性就是 0.3。1985 年是 0.3,到了 2005 年还是 0.3。这其中总存在一定程度的、无法再减的不确定性。
With the advent of computers, big data now, how has that correlation changed? Some correlation would be like 0.3. That 0.3 in 1985 is still 0.3 in 2005. There is just a degree of irreducible uncertainty.
所以,我们能学到的确实有限。但你和那些优秀的选秀专家聊这个,他们会说,没错没错没错,可我们在进步。我相信这一点,也愿意保持这种可能性,因为他们确实在进步。他们绝对在进步。但在我们变得更好之前,必须相当谦逊,坚持这个模式。
So there is a limit to how much we can learn. But you talk about that with good draft guys and they’ll say, yes, yes, yes, but we’re getting better. And I believe it and I want to hold open that possibility because they are getting better. They absolutely are getting better. But until we’re better, we need to be really humble and stay with the model.
问:我想问一下关于录取流程的事,我喜欢这个流程的一点是,它把讨论环节去掉了,从某种意义上说,只是汇总了个人的判断。
Question: Can I ask in terms of the admissions process, one of the things that I like about the process is what happens to the debate in that it removes the discussion element of it and just in a sense aggregates the judgments of the individuals.
但是,招生委员会对缺少这些数据会有什么反应?我的意思是,我猜你是说他们事后确实会讨论,但在我看来,在某些方面,讨论实际上可能会使讨论和判断产生偏差,假设人们在群体中更有影响力。如果你去掉这一点,似乎就能接近某种“群体智慧”。那么,对于没有那四天半时间来讨论每个候选人的情况,各方反应如何?
But what’s the reaction then from the Admissions Committee about the absence of the data? I mean, I guess you’re saying they do talk about it afterwards, but from my perspective, in some ways, talking actually can bias the discussion and judgments, assume people are more influential in a group. And if you remove that, it seems like you get close to kind of the wisdom of crowds. What’s the reaction been in terms of not having those four-and-a-half days to talk about each individual candidate?
凯德:我认为早期确实存在怀疑,但在经历了这一整套体系之后,已经得到了广泛的认同。
Cade: I think there was skepticism early and there has been broad buy-in having gone through this system,
宾夕法尼亚大学 凯德·梅西
Cade Massey University of Pennsylvania
即便第一年、第一轮走下来就做到了,原因在于我们并没有取消辩论,只是重新聚焦了辩论的方向。
even just having gone through it the first round the first year, and it’s because we didn’t remove debate. We just refocused the debate.
关于标准是什么、如何从一篇文章中判断这些标准,仍然存在争论;关于各项目标之间如何合理分配权重,存在争论;关于正确的组合策略、政策约束、生源来源以及这些要素的恰当搭配,我们也在不断争论,所有这些都在讨论范围之内。
There is still debate on what the criteria are, how you judge that criteria from an essay, there is debate on what the right weights are across our objectives, there is debate on what’s the right portfolio, the policy constraints, where we’re pulling students from, and what the right mix of those are, we debate all those things.
关于例外情况,确实存在争论。每当我们审视那一串由 1 和 0 构成的名单——1200 个 1 和 0——我们就会关注那些处于临界点上的人,那些如果你把权重往一边调,他们就会被剔除,如果你往另一边调,他们就会被纳入的人。所以对于处于边缘的候选对象,我们会更仔细地审视他们,然后展开讨论。我们把更多精力放在真正关键的地方,而不是把时间浪费在毫无产出的环节上。
There are debates on the exceptions. Whenever we look at that list of ones and zeros, 1,200 ones and zeros, we look at the people who are on the fence, we look at the people who if you shift the weights one way, they get out, if you shift the weight the other way, they get in. So for the marginal candidates, we go and look at them in more detail, and then they get debated. We get much more attention on the right places as opposed to spending all of this time where it’s just not productive.
提问者:我觉得这种思维方式确实挑战了人的决策习惯。我担心的是结果本身——我们该如何衡量结果?
Question: I think it’s great how this kind of defies one’s decision making. The concern I have is about the outcomes. How do we measure the outcome?
那么,用传统方式跑一小部分业务,同时在模型里也跑相同的一小部分,这样一来,你实际上就有了一个准自然实验——随着时间的推移,你能看到结果可能会发生变化,并且可以回过头去审视,这样是不是说得通?
Would it make sense to just run a sliver of things the traditional way against a sliver of things within the model and then you actually have somewhat of a natural experiment where over time, you can see that outcome might change, and you’ll be able to go back and look?
凯德:是的。好吧,现在我遇到了这位先生在这里指导的反方向问题——我要依赖一个我知道很嘈杂的判断?真的,我真的想这么做吗?
Cade: Yes. Well, now I’ve got the opposite problem of this gentleman’s instruction up here which is I’m going to rely on judgment that I know was noisy? Really, do I really want to do that?
总体而言,我完全认同这个问题背后的精神。我们希望能尽可能做一些试验。回到样本量的问题,在任何特定条件下都很难获得足够的数据来得出太多推论。所以,我们在试验中能做什么、不能做什么,一直让我们感到自己很渺小。
Broadly, the spirit of the question, I agree with entirely. We want to run some experiments if at all possible. Back to the sample size issue, it’s really hard to have enough in any given condition to draw much inference. And so, we’ve been humbled by what we’re able to do and not do experimentally.
不过,关于衡量学习成果这件事,我得说,我们现在确实没做好。而且我们也不知道该怎么做,我们还没有解决办法,也完全清楚这项任务的难度。这个过程带来的最大好处是,我们现在开始讨论这件事了,我们真的在辩论到底应该如何评估学生的表现。
I want to say though about measuring outcomes, we know we’re not getting that right now. And we don’t know, we don’t have the solution and we fully appreciate the difficulty of the task. The best thing to come of this process is that we’re having the conversation now, that we’re actually debating exactly how it is that we should be assessing our students’ performance.
我们之所以采取这一策略,部分灵感来源于“为美国而教”(Teach for America)。你可能不太了解,但“为美国而教”是我见过的最精明的招聘组织。他们在我们第一届人才分析大会上发表了演讲。
We were partly inspired to pursue this tack by Teach for America. You might not know it, but Teach for America are the most sophisticated hiring organization I’ve ever been around. They spoke at our first conference, the People Analytics Conference.
仔细想想,这其实挺有道理的,因为每年有五六万人申请同一份工作。这份工作差不多是同质化的——成千上万份申请,而他们干这一行已经 15 年了。那家公司里有极其聪明、精通量化分析的人,这 15 年来一直在不断优化他们的筛选流程。
If you think about it, it makes some sense because they have 50,000 or 60,000 applicants every year for the same job. It’s a homogenous job, more or less—tens of thousands of applications and they’ve been doing it now 15 years. They have really smart, quantitatively sophisticated people in that organization who have been refining their process for 15 years now.
他们在我们第一年的大会上说了件漂亮事。他们说,这事儿永远不会有尽头。它不是一个修复录取流程的项目,也不是一个修补招聘流程的项目。它是一个持续的过程。他们还提到,他们内部已经有这个模式了,而且这项工作永远不会结束。
They said this beautiful thing at our conference the first year. They said we’re never going to be done. It’s not a project to fix admissions. It’s not a project to fix recruiting. It’s an ongoing process. And they said they had this model inside, and they’re never going to be done.
所以,我和招生副院长玛丽艾伦·赖利·兰姆从一开始就是这么说的。这几乎等于给我们发了张许可证,可以说我们现在确实不知道该如何衡量结果,因为这事真的很难,但我们还是会开始尝试,并且会继续讨论这个问题。因为这个过程,我们现在正在展开以前从未有过的对话。但这是个非常大、也非常棘手的问题。
And so, Maryellen [Reilly Lamb] and I, the vice dean of admissions, have said that from day one. It kind of licenses us to say we don’t really know right now how we’re going to measure our outcomes because it’s really hard, but we’re going to start trying and we’re going to have the conversation. And we’re having the conversation now that we have never had before because of the process. But that’s a very big and very difficult question.
问题:谢谢。刚才第一个实验里,你谈到人们对事物有非常强烈的抵触情绪。
Question: Thank you. Just in the first experiment, you were talking about how people have a very adverse
宾夕法尼亚大学的凯德·梅西
Cade Massey University of Pennsylvania
对算法带来负面结果的反应。我对您大学里的实验很好奇——假设您以 GPA、诺贝尔奖、工资等指标为目标进行优化,这在某种意义上或许说得通,但您同时也可能突然造出更多连环杀手以及其他您可能并未筛选的特征。难道您不会说,“哦,这是个问题,我们干不了这事,因为后果是我们会多出进监狱的人,多到我们根本没法衡量”?
reaction to negative outcomes from algorithms. I’m curious with your college experiment, let’s say you’re solving for GPA and Nobel Prizes and wages and things, if that might make sense, but you also could suddenly get more axe murderers and other characteristics you may not be screening for. Wouldn’t you say, oh, this is a problem, we can’t do this because we have more people who might go to prison than we can measure.
凯德:(笑)我倒不是担心杀妻凶手什么的,那不是我操心的事。但总的来说,我有这样一种感觉——因为自己做过的研究,我比任何人都更清楚这一点——我很幸运能推动这套新招生系统上线,帮着把它运转起来,因为它得等上几年才会看到成效。
Cade: [Laughs] I’m not worried about axe murderers per say. That’s not something I’m worried about. But generally, this idea, I’m aware of something, more aware than anybody else because of the research I’ve done, I’m greatly privileged in getting this new system going in admissions, helping get it going because we won’t see the outcomes for a couple of years.
这项研究的一切都表明,如果人们得不到反馈、看不到算法的错误,反而会对算法越来越感兴趣。所以,我大概有两年窗口期,在最有利的环境下把事情推起来,然后,一旦我们开始去衡量,就会发现算法充满了噪声。
Everything about this research says that people are getting more interested in algorithms if they don’t get feedback, if they don’t see the algorithm error. And so, it’s like I’ve got this two-year window essentially to get things going in the most hospitable environment possible, and then, once we’ve started measuring, we’ll learn that the algorithm is noisy.
我们会发现,那些我们排名很高的人有时表现并不完美,这会让一些人感到担忧,我同意。我希望我们不会遇到杀人狂魔,但一旦出现对我们不利的糟糕结果,情况肯定会更具挑战性。而且这种情况会有很多,因为这是一项艰巨的任务。
We’ll learn that those people that we ranked so highly sometimes don’t turn out so perfectly and that will cause some people some concern, agreed. I hope we don’t turn up axe murderers, but it will definitely be more challenging once we have hard outcomes that go against us. And there will be plenty because it’s a hard task.
提问者:我很好奇,您是否和您的同事菲尔·泰特洛克有过交流?
Question: I’m curious if you spend any time with your colleague, Phil Tetlock?
Cade: Sure.
Cade: Sure.
问:我是说,很明显,他有自己的超级预测方法,手下有各种团队,其中一些团队年复一年地持续跑赢大盘。
Question: I mean, obviously, he’s got his superforecasting approach and he’s got all these myriad teams and some of those teams consistently outperform year after year.
我的问题是,你有没有考虑过尝试参与其中,并构建算法来组建你自己的团队?第二个问题是,你认为那些成功的团队在做的某些事情,本质上是不是具有算法特征,并且在某种程度上与你正在做的事情是一致的?
My question is, have you thought about trying to participate and build algorithms to field your own team? And the second question is, do you think that there are elements of what those successful teams are doing that are algorithm-like in nature and that kind of coincide with what you’re doing?
凯德:当然,我跟菲尔很熟,也很喜欢他,还有芭布[梅勒斯]也一样,他们是一支了不起的团队,我们确实会聊这些事。事实上,他们有个团队一直想利用招生数据作为某种研究的刺激素材,所以他们在跟那边一位叫莱尔·昂加尔的计算机科学人士积极合作。
Cade: Certainly, I know and I thoroughly enjoy Phil, and Barb [Mellers] as well, and that’s a phenomenal team and we do talk about these things. In fact, they’ve had a team interested in using data from admissions for one of their stimuli in some kind of study. So they’re actively engaged in a way with Lyle Ungar, a computer science guy there.
第二个问题:他们在做的事情中,任何与算法相关的内容都值得关注。但他们并没有非常明确地走那条路。可能有些人是单独用算法做预测的,但那并不是他们的招牌特色。
For the second question, anything in what they’re doing that’s algorithmic is interesting. They don’t very explicitly go down that road. There may be people who are individual forecasters who are working with algorithms, but that’s not part of their shtick.
他们发现了一些非凡的东西。他们找到了优秀预测者身上那些能将他们与糟糕预测者区分开来的特质。而同样意义深远的,或许是他们开发出了一些训练技巧,能够改善人们在不确定条件下的判断力。
They have discovered some phenomenal things. They’ve discovered some qualities in good forecasters that differentiate good forecasters from bad forecasters. And probably as profound, have developed some training techniques that improve people’s judgment under uncertainty.
再说回到我的专业领域,从保罗·米尔(Paul Meehl)到现在,这一直都在研究不确定性条件下的判断。严格来说米尔不属于我的领域,但自从丹尼尔·卡尼曼和阿莫斯·特沃斯基用非常严谨的方式研究过之后,还没有任何人真正在这方面做出过改进。而这两位后来陆续提出了一些方法,确实能帮助人们改善自己的判断——这正是我们应该融入实践的那类东西。
And again, my field has looked at judgment under uncertainty all the way back to [Paul] Meehl. Really Meehl’s not my field, but since Danny Kahneman and Amos Tversky did so pretty rigorously, no one has ever really improved upon it. And these guys have come along and come up with some techniques for actually improving their judgment. That’s the kind of thing that we need to be incorporating.
例如,我们可以用一些同样的技术来培训我们的招生人员,很可能就会得到更准确的预测。但我们很幸运,菲尔和芭布就在那里做这项工作。
We can train our admissions people, for example, using some of those same techniques, and we’d probably see better forecasts. But we’re lucky to have Phil and Barb doing that work right there.
迈克尔:好的,我们就到这里了。谢谢你,凯德。
Michael: Well, we’ll call it there. Thank you, Cade.
保罗·德波德斯塔 克利夫兰布朗队
Paul DePodesta Cleveland Browns
保罗·德波德斯塔的职业生涯就是评估、衡量和衡量人才价值,这一点在迈克尔·刘易斯的著作《点球成金:赢得不公平游戏的艺术》中有详实记载。《点球成金》方法论已成为商业领袖寻求新方法来改革僵化体系的一项主流策略。
Paul DePodesta has made a career of evaluating, measuring, and assigning value to talent, as documented in Michael Lewis’s book, Moneyball: The Art of Winning an Unfair Game. The Moneyball methodology has become a mainstay strategy for business leaders looking for new approaches for overhauling stagnant systems.
德波德斯塔曾任纽约大都会队球员发展与业余球探副总裁,他帮助该队自 2000 年以来首次打入 2015 年世界大赛。大都会队总经理桑迪·奥尔德森表示,保罗是大都会队成功的“巨大因素”。
Formerly the Vice President of Player Development and Amateur Scouting for the New York Mets, Paul helped lead the team to the 2015 World Series for the first time since 2000. Mets GM Sandy Alderson said Paul was a “huge factor” in the Mets’ success.
2016 年 1 月,德波德斯塔加入美国国家橄榄球联盟(NFL)的克利夫兰布朗队,担任首席战略官。在这个新职位上,他负责评估和实施最佳实践与策略,为布朗队提供所需的全面资源,以便为球员和球队做出最优决策。
In January 2016, Paul joined the NFL’s Cleveland Browns as Chief Strategy Officer. In this new role, he is responsible for assessing and implementing the best practices and strategies that will give the Browns the comprehensive resources needed to make optimal decisions for their players and team.
保罗还担任斯克里普斯转化科学研究所的生物信息学助理教授。
Paul is also an Assistant Professor of Bioinformatics at the Scripps Translational Science Institute.
注:无相关记录。
Note: No transcript available.
本文件由瑞士信贷编制,其中所表达的观点仅为瑞士信贷在撰写之日的观点,并可能随时变更。本文件仅为信息目的而编制,仅供接收者使用。本文件不构成瑞士信贷代表任何人购买或出售任何证券的要约或邀请。本材料中的任何内容均不构成投资、法律、会计或税务建议,也不表示任何投资或策略适合或适用于您的个人情况,亦不构成对您的个人建议。所提及投资的价格和价值以及可能产生的任何收入可能波动,可能下跌或上涨。任何过往业绩的引用均不代表未来表现。
This document was produced by and the opinions expressed are those of Credit Suisse as of the date of writing and are subject to change. It has been prepared solely for information purposes and for the use of the recipient. It does not constitute an offer or an invitation by or on behalf of Credit Suisse to any person to buy or sell any security. Nothing in this material constitutes investment, legal, accounting or tax advice, or a representation that any investment or strategy is suitable or appropriate to your individual circumstances, or otherwise constitutes a personal recommendation to you. The price and value of investments mentioned and any income that might accrue may fluctuate and may fall or rise. Any reference to past performance is not a guide to the future.
本出版物中包含的信息和分析系从被认为可靠的来源汇编或获得,但瑞士信贷对其准确性或完整性不作任何陈述,且对因使用本文件而产生的任何损失不承担任何责任。瑞士信贷集团旗下公司在向瑞士信贷客户提供本出版物之前,可能已根据其中包含的信息和分析采取了行动。新兴市场的投资具有投机性,且比成熟市场的投资波动性大得多。部分主要风险包括政治风险、经济风险、信用风险、货币风险和市场风险。外币投资受汇率波动影响。在进行任何交易之前,您应考虑该交易是否适合您的特定情况,并(在必要时与您的专业顾问一起)独立评估具体的财务风险以及法律、监管、信用、税务和会计后果。本文件在美国由美国注册经纪交易商瑞士信贷证券(美国)有限责任公司发行和分发;在加拿大由瑞士信贷证券(加拿大)公司发行和分发;在巴西由瑞士信贷投资银行(巴西)股份公司发行和分发。
The information and analysis contained in this publication have been compiled or arrived at from sources believed to be reliable but Credit Suisse does not make any representation as to their accuracy or completeness and does not accept liability for any loss arising from the use hereof. A Credit Suisse Group company may have acted upon the information and analysis contained in this publication before being made available to clients of Credit Suisse. Investments in emerging markets are speculative and considerably more volatile than investments in established markets. Some of the main risks are political risks, economic risks, credit risks, currency risks and market risks. Investments in foreign currencies are subject to exchange rate fluctuations. Before entering into any transaction, you should consider the suitability of the transaction to your particular circumstances and independently review (with your professional advisers as necessary) the specific financial risks as well as legal, regulatory, credit, tax and accounting consequences. This document is issued and distributed in the United States by Credit Suisse Securities (USA) LLC, a U.S. registered broker-dealer; in Canada by Credit Suisse Securities (Canada), Inc.; and in Brazil by Banco de Investimentos Credit Suisse (Brasil) S.A.
本文件在瑞士由瑞士银行瑞士信贷股份公司发行和分发。瑞士信贷受瑞士金融市场监管局(FINMA)授权和监管。本文件在欧洲(瑞士除外)由瑞士信贷(英国)有限公司和瑞士信贷证券(欧洲)有限公司(伦敦)发行和分发。瑞士信贷证券(欧洲)有限公司(伦敦)和瑞士信贷(英国)有限公司,由审慎监管局(PRA)授权并由金融市场行为监管局(FCA)和 PRA 监管,是瑞士信贷内部相互关联但独立的法律和受监管实体。英国金融服务管理局为私人客户提供的保护不适用于由英国境外人士提供的投资或服务,如果投资发行方未能履行其义务,金融服务补偿计划也将不适用。本文件在根西岛由瑞士信贷(根西岛)有限公司发行和分发,该公司是一家根据根西岛法律注册的独立法律实体,公司编号 15197,注册地址为 Helvetia Court, Les Echelons, South Esplanade, St Peter Port, Guernsey。瑞士信贷(根西岛)有限公司由瑞士信贷全资拥有,并受根西岛金融服务委员会监管。经要求可提供最新经审计账目的副本。本文件在泽西岛由瑞士信贷(根西岛)有限公司泽西岛分行发行和分发,该分行受泽西岛金融服务委员会监管。瑞士信贷(根西岛)有限公司泽西岛分行在泽西岛的业务地址为:TradeWind House, 22 Esplanade, St Helier, Jersey JE2 3QA。本文件在亚太地区由以下在相关司法管辖区获得适当授权的实体之一发行:在香港由持有香港证券及期货事务监察委员会牌照的公司瑞士信贷(香港)有限公司,或由香港金融管理局监管的授权机构瑞士信贷香港分行,以及根据香港法例第 571 章《证券及期货条例》监管的注册机构;在日本由瑞士信贷证券(日本)有限公司;在亚太其他地区由以下在相关司法管辖区获得适当授权的实体之一:瑞士信贷股票(澳大利亚)有限公司、瑞士信贷证券(泰国)有限公司、瑞士信贷证券(马来西亚)私人有限公司、瑞士信贷股份公司新加坡分行,以及在世界其他地区由上述实体的相关授权附属机构。
This document is distributed in Switzerland by Credit Suisse AG, a Swiss bank. Credit Suisse is authorized and regulated by the Swiss Financial Market Supervisory Authority (FINMA). This document is issued and distributed in Europe (except Switzerland) by Credit Suisse (UK) Limited and Credit Suisse Securities (Europe) Limited, London. Credit Suisse Securities (Europe) Limited, London and Credit Suisse (UK) Limited, authorised by the Prudential Regulation Authority (PRA) and regulated by the Financial Conduct Authority (FCA) and PRA, are associated but independent legal and regulated entities within Credit Suisse. The protections made available by the UK‘s Financial Services Authority for private customers do not apply to investments or services provided by a person outside the UK, nor will the Financial Services Compensation Scheme be available if the issuer of the investment fails to meet its obligations. This document is distributed in Guernsey by Credit Suisse (Guernsey) Limited, an independent legal entity registered in Guernsey under 15197, with its registered address at Helvetia Court, Les Echelons, South Esplanade, St Peter Port, Guernsey. Credit Suisse (Guernsey) Limited is wholly owned by Credit Suisse and is regulated by the Guernsey Financial Services Commission. Copies of the latest audited accounts are available on request. This document is distributed in Jersey by Credit Suisse (Guernsey) Limited, Jersey Branch, which is regulated by the Jersey Financial Services Commission. The business address of Credit Suisse (Guernsey) Limited, Jersey Branch, in Jersey is: TradeWind House, 22 Esplanade, St Helier, Jersey JE2 3QA. This document has been issued in Asia-Pacific by whichever of the following is the appropriately authorised entity of the relevant jurisdiction: in Hong Kong by Credit Suisse (Hong Kong) Limited, a corporation licensed with the Hong Kong Securities and Futures Commission or Credit Suisse Hong Kong branch, an Authorized Institution regulated by the Hong Kong Monetary Authority and a Registered Institution regulated by the Securities and Futures Ordinance (Chapter 571 of the Laws of Hong Kong); in Japan by Credit Suisse Securities (Japan) Limited; elsewhere in Asia/Pacific by whichever of the following is the appropriately authorized entity in the relevant jurisdiction: Credit Suisse Equities (Australia) Limited, Credit Suisse Securities (Thailand) Limited, Credit Suisse Securities (Malaysia) Sdn Bhd, Credit Suisse AG,Singapore Branch,and elsewhere in the world by the relevant authorized affiliate of the above.
未经作者和瑞士信贷的书面许可,不得以整体或部分形式复制本文件。
This document may not be reproduced either in whole, or in part, without the written permission of the authors and CREDIT SUISSE.