WEBVTT - The Man Who Wrote the AI Textbook Says We're Heading For Extinction - The Story

0:00:16.000 --> 0:00:17.000
<v Speaker 1>Welcome to tech stuff.

0:00:17.120 --> 0:00:20.360
<v Speaker 2>I'm Os Vloscian, our guest today is a world renowned

0:00:20.440 --> 0:00:23.920
<v Speaker 2>expert in the field of AI. In fact, he literally

0:00:23.960 --> 0:00:26.759
<v Speaker 2>wrote the book on it. Stuart Russell, together with co

0:00:26.880 --> 0:00:31.240
<v Speaker 2>author Peter Norvig, released Artificial Intelligence, A Modern Approach back

0:00:31.280 --> 0:00:34.320
<v Speaker 2>in the mid nineteen nineties. Since then, it's been translated

0:00:34.360 --> 0:00:38.320
<v Speaker 2>into fourteen languages and is used in fifteen hundred universities

0:00:38.479 --> 0:00:41.240
<v Speaker 2>in one hundred and thirty five countries. But just because

0:00:41.240 --> 0:00:44.680
<v Speaker 2>he wrote the book doesn't mean he can't criticize the technology, or,

0:00:44.680 --> 0:00:48.040
<v Speaker 2>more specifically, the way it's being rolled out, often without

0:00:48.080 --> 0:00:51.200
<v Speaker 2>regard to consequence. Stuart has spent the past few years

0:00:51.200 --> 0:00:54.800
<v Speaker 2>sounding the alarm about the lack of AI safety protocols

0:00:55.160 --> 0:00:58.520
<v Speaker 2>and what he sees as the very real possibility that

0:00:58.600 --> 0:01:02.520
<v Speaker 2>our current path could lead to human extinction. The conversation

0:01:02.680 --> 0:01:06.160
<v Speaker 2>comes at a timely moment, as the US government is

0:01:06.200 --> 0:01:09.280
<v Speaker 2>forcefully cracking down on who can and can't use the

0:01:09.400 --> 0:01:13.160
<v Speaker 2>leading models coming from anthropic Today. Stuart is a professor

0:01:13.200 --> 0:01:15.960
<v Speaker 2>of computer science at Berkeley and the president of the

0:01:16.040 --> 0:01:20.679
<v Speaker 2>International Association for Safe and Ethical AI. He also recently

0:01:20.680 --> 0:01:24.080
<v Speaker 2>took the stand as the only expert witness on AI

0:01:24.240 --> 0:01:28.240
<v Speaker 2>for Elon Musk in his lawsuit against Samaltman and open AI.

0:01:28.520 --> 0:01:29.880
<v Speaker 1>Stuart, welcome to tech stuff.

0:01:30.080 --> 0:01:31.240
<v Speaker 3>It's a pleasure to be with you.

0:01:31.520 --> 0:01:32.640
<v Speaker 1>I had to ask you. What did you say on

0:01:32.680 --> 0:01:33.080
<v Speaker 1>the stand?

0:01:34.200 --> 0:01:36.639
<v Speaker 4>Well, it's more what I was not allowed to say.

0:01:36.880 --> 0:01:40.679
<v Speaker 4>I was not allowed to talk about anything related to

0:01:40.760 --> 0:01:45.520
<v Speaker 4>existential risk, which was surprising because open ai was set

0:01:45.600 --> 0:01:51.280
<v Speaker 4>up for that reason. Musk and others were concerned that

0:01:51.480 --> 0:01:55.400
<v Speaker 4>if AI were in the hands of for profit companies

0:01:56.080 --> 0:02:00.720
<v Speaker 4>that they would disregard safety and put him at risk,

0:02:01.160 --> 0:02:03.320
<v Speaker 4>so Open a Eye was created.

0:02:02.920 --> 0:02:03.800
<v Speaker 3>To counter that.

0:02:04.520 --> 0:02:08.280
<v Speaker 4>The judge said, no, you can't talk about that, which

0:02:08.360 --> 0:02:09.079
<v Speaker 4>was disappointing.

0:02:09.280 --> 0:02:11.800
<v Speaker 1>Why did the judge say you couldn't talk about existential risk?

0:02:12.040 --> 0:02:14.360
<v Speaker 3>I was not privy to those discussions either.

0:02:14.880 --> 0:02:21.640
<v Speaker 4>Maybe the defense lawyers thought it would be prejudicial, that

0:02:21.680 --> 0:02:26.520
<v Speaker 4>it was in some sense speculative because it hasn't we

0:02:26.560 --> 0:02:29.400
<v Speaker 4>haven't yet gone extinct, right, And this is a this

0:02:29.520 --> 0:02:31.600
<v Speaker 4>is a strange argument I hear from a lot of people.

0:02:32.240 --> 0:02:35.080
<v Speaker 4>You know, it's just science fiction. Well what do you

0:02:35.120 --> 0:02:37.800
<v Speaker 4>mean by that, Well, it hasn't happened yet. You know,

0:02:38.000 --> 0:02:41.040
<v Speaker 4>everything that's ever happened. There was a time before it

0:02:41.120 --> 0:02:44.680
<v Speaker 4>happened for it, and by your nothing could ever happen

0:02:45.000 --> 0:02:48.160
<v Speaker 4>because everything was at some point science fiction. But you know,

0:02:48.200 --> 0:02:50.799
<v Speaker 4>when you look at science fiction, you know, they talked

0:02:50.840 --> 0:02:55.280
<v Speaker 4>about nuclear weapons in nineteen twelve HG. Wells, they talked

0:02:55.280 --> 0:02:59.000
<v Speaker 4>about space travel in the nineteenth century, and in fact,

0:02:59.440 --> 0:03:02.560
<v Speaker 4>they talked to AI in the nineteenth century. Samuel Butler

0:03:02.680 --> 0:03:08.320
<v Speaker 4>wrote a book that described society where there had been

0:03:08.400 --> 0:03:11.080
<v Speaker 4>an enormous conflict between those who were in favor of

0:03:11.120 --> 0:03:15.040
<v Speaker 4>the machines and those who predicted that the machines would

0:03:15.520 --> 0:03:17.440
<v Speaker 4>would be the ruin of the human race.

0:03:17.639 --> 0:03:21.639
<v Speaker 2>Now is Open AI particularly bad? I mean, I know

0:03:21.720 --> 0:03:25.680
<v Speaker 2>you are asked by their counseling cross you know if

0:03:25.720 --> 0:03:28.600
<v Speaker 2>you believe that the for profit sort of motivation to

0:03:28.680 --> 0:03:33.520
<v Speaker 2>recklessly develop AI by definition and endanger's you know, humanity,

0:03:34.040 --> 0:03:38.920
<v Speaker 2>surely that also applies to Elon and XAI and SpaceX.

0:03:39.360 --> 0:03:42.800
<v Speaker 2>And you said, if that hypothesis is correct, then yes,

0:03:42.840 --> 0:03:46.000
<v Speaker 2>In other words, that the same critique could apply to

0:03:46.760 --> 0:03:50.600
<v Speaker 2>you know, Open AI or Anthropic or Google or SpaceX.

0:03:51.080 --> 0:03:51.280
<v Speaker 3>Yeah.

0:03:51.320 --> 0:03:55.080
<v Speaker 4>It was specifically not my job to compare the safety

0:03:55.120 --> 0:04:00.400
<v Speaker 4>records of different companies or their safety positions. I think

0:04:00.840 --> 0:04:04.440
<v Speaker 4>if I understand it correctly. Anthropic is a public benefit corporation,

0:04:05.280 --> 0:04:09.440
<v Speaker 4>which is one way of allowing considerations other than profit

0:04:10.320 --> 0:04:13.960
<v Speaker 4>to affect the decisions made by management and the board,

0:04:14.720 --> 0:04:17.240
<v Speaker 4>because you know, for a regular for profit company, there

0:04:17.279 --> 0:04:24.200
<v Speaker 4>is a legal obligation to maximize shareholder return. And what's

0:04:24.240 --> 0:04:30.000
<v Speaker 4>happening here is that the risks imposed on the rest

0:04:30.040 --> 0:04:34.479
<v Speaker 4>of humanity are externalities, as economists call it, which means

0:04:34.480 --> 0:04:38.039
<v Speaker 4>that someone is making decision and there's a bad consequence

0:04:38.680 --> 0:04:42.840
<v Speaker 4>that is being loaded onto somebody else. So you know,

0:04:42.880 --> 0:04:48.960
<v Speaker 4>you think about chemical companies who skimp on safety, and

0:04:49.000 --> 0:04:53.159
<v Speaker 4>as it stands, according to the companies, these are all externalities,

0:04:53.240 --> 0:04:58.480
<v Speaker 4>meaning they don't accept responsibility for these consequences. Those are

0:04:58.480 --> 0:05:02.240
<v Speaker 4>harms that don't fake into their balance sheet, and so

0:05:02.480 --> 0:05:06.279
<v Speaker 4>the same would be true in a sense for human extinction.

0:05:07.520 --> 0:05:09.120
<v Speaker 3>And even liability would not.

0:05:09.160 --> 0:05:12.960
<v Speaker 4>Really be a deterrent for that right because obviously you

0:05:13.000 --> 0:05:16.320
<v Speaker 4>wouldn't be around today to pay the compensation. So they

0:05:16.360 --> 0:05:19.239
<v Speaker 4>just sort of factor that out of their decision making.

0:05:20.200 --> 0:05:23.800
<v Speaker 4>And that's exactly what open AI was set up to avoid.

0:05:23.839 --> 0:05:27.279
<v Speaker 4>But now it's with the transition to a for profit entity.

0:05:27.320 --> 0:05:30.719
<v Speaker 4>It's part of that calculus So take.

0:05:30.640 --> 0:05:31.240
<v Speaker 1>Us back in time.

0:05:31.279 --> 0:05:34.119
<v Speaker 2>In nineteen ninety five, you wrote this textbook that became

0:05:34.120 --> 0:05:37.920
<v Speaker 2>the defining textbook on AI, and you came up with

0:05:37.960 --> 0:05:41.400
<v Speaker 2>a concept called the standard model. Can you explain what

0:05:41.440 --> 0:05:43.760
<v Speaker 2>that is and how it relates to the conversation which

0:05:43.760 --> 0:05:47.200
<v Speaker 2>you're now so engaged in today as to the potential

0:05:47.200 --> 0:05:48.880
<v Speaker 2>extinction of the human race because of AI.

0:05:49.800 --> 0:05:50.120
<v Speaker 3>Yeah.

0:05:50.200 --> 0:05:53.080
<v Speaker 4>So if we go back even further to the beginnings

0:05:53.120 --> 0:05:56.520
<v Speaker 4>of the field of AI in the post war period,

0:05:56.600 --> 0:06:01.720
<v Speaker 4>so nineteen forties, nineteen fifties, I think everyone agreed that

0:06:02.279 --> 0:06:08.320
<v Speaker 4>AI is about creating intelligence in machines. What wasn't clear

0:06:08.480 --> 0:06:12.400
<v Speaker 4>was well, what is intelligence? How do we define this

0:06:12.640 --> 0:06:16.920
<v Speaker 4>target and how do we go about doing it? So

0:06:18.160 --> 0:06:22.680
<v Speaker 4>there was an active debate you should we try to

0:06:22.920 --> 0:06:27.200
<v Speaker 4>emulate human intelligence, Should we understand what's going on in

0:06:27.240 --> 0:06:29.800
<v Speaker 4>the human mind, the human brain and then sort of

0:06:30.279 --> 0:06:34.320
<v Speaker 4>go ahead and implement that, or should we focus on

0:06:34.360 --> 0:06:38.520
<v Speaker 4>a more abstract notion, in fact, a notion that philosophers

0:06:38.560 --> 0:06:41.839
<v Speaker 4>had developed for thousands of years and economists as well,

0:06:41.920 --> 0:06:45.680
<v Speaker 4>this idea of rational behavior that an entity is intelligent

0:06:46.800 --> 0:06:49.719
<v Speaker 4>if it acts in a way that is expected to

0:06:49.800 --> 0:06:53.520
<v Speaker 4>achieve its objectives. And I would say, for the most

0:06:53.520 --> 0:06:59.320
<v Speaker 4>part that second approach one out because it doesn't require

0:06:59.440 --> 0:07:03.240
<v Speaker 4>doing psychological experiments on humans to find out how their

0:07:03.240 --> 0:07:08.359
<v Speaker 4>brains work. Right, It's an abstract mathematical concept, and we

0:07:08.480 --> 0:07:12.960
<v Speaker 4>had tools. We had formal logic so that we could

0:07:13.000 --> 0:07:17.200
<v Speaker 4>create algorithms that were able to construct plans to achieve goals.

0:07:18.000 --> 0:07:21.280
<v Speaker 4>So that these two the two views are I would say,

0:07:22.160 --> 0:07:25.960
<v Speaker 4>the sort of rational view, i e. Base intelligence on

0:07:26.560 --> 0:07:31.720
<v Speaker 4>formal foundations of how one should reason, how one should

0:07:31.760 --> 0:07:36.560
<v Speaker 4>make decisions, versus a more biology based view, which is

0:07:37.480 --> 0:07:45.200
<v Speaker 4>neurons and their computational analogs. And so the standard model

0:07:45.240 --> 0:07:47.840
<v Speaker 4>refers to this. What was dominant I would say from

0:07:47.920 --> 0:07:54.360
<v Speaker 4>like nineteen sixty to twenty ten, twenty twenty some where

0:07:54.680 --> 0:07:57.920
<v Speaker 4>it was much more the rational view entities are intelligent

0:07:58.480 --> 0:08:00.960
<v Speaker 4>to the extent that their actions can be expected to

0:08:01.000 --> 0:08:05.200
<v Speaker 4>achieve their objectives. And then it was really the twenty

0:08:05.240 --> 0:08:09.440
<v Speaker 4>twelve work that Jeff Hinton did with Ilia Sutzkiver and

0:08:09.680 --> 0:08:14.760
<v Speaker 4>I think Alex Kruzhevski on a system for recognizing objects

0:08:14.760 --> 0:08:20.160
<v Speaker 4>in images that significantly exceeded the methods that other people

0:08:20.160 --> 0:08:23.400
<v Speaker 4>had developed before that. And then it was off to

0:08:23.440 --> 0:08:27.120
<v Speaker 4>the races, and then language models came along, so applying

0:08:28.040 --> 0:08:34.120
<v Speaker 4>somewhat similar ideas again large neural networks that were trained

0:08:34.120 --> 0:08:38.079
<v Speaker 4>from vast amounts of data, applying that to text and

0:08:38.120 --> 0:08:43.760
<v Speaker 4>then generating eventually CHAT, GPT and all the successors that

0:08:43.800 --> 0:08:44.680
<v Speaker 4>we've seen since then.

0:08:45.080 --> 0:08:49.400
<v Speaker 2>I guess the question I'm coming to is was the

0:08:49.520 --> 0:08:53.040
<v Speaker 2>victory for the time being, at least of the neural

0:08:53.200 --> 0:08:56.920
<v Speaker 2>net type of AI. The reason why you became so

0:08:57.080 --> 0:09:02.120
<v Speaker 2>concerned about AI safety or AI safety in nineteen ninety

0:09:02.120 --> 0:09:04.720
<v Speaker 2>five when you were writing this book was the possibility

0:09:04.720 --> 0:09:07.040
<v Speaker 2>of machines having their own goals which could be very

0:09:07.080 --> 0:09:10.480
<v Speaker 2>different to ours and lead to our extinction already on

0:09:10.559 --> 0:09:12.760
<v Speaker 2>your mind. In other words, did the evolution of the

0:09:12.800 --> 0:09:16.760
<v Speaker 2>technology make the safety issue more urgent?

0:09:16.960 --> 0:09:18.880
<v Speaker 1>And if so, why so?

0:09:19.320 --> 0:09:21.760
<v Speaker 4>In ninety five, when I published a book with Peter,

0:09:21.920 --> 0:09:25.160
<v Speaker 4>we have a section called what if we do succeed

0:09:26.960 --> 0:09:30.880
<v Speaker 4>and it points out that if we create machines more

0:09:30.880 --> 0:09:33.000
<v Speaker 4>intelligent than us, which is what we were trying to do,

0:09:34.080 --> 0:09:36.439
<v Speaker 4>we might face this problem that we wouldn't have any

0:09:36.520 --> 0:09:38.280
<v Speaker 4>idea how to control them, that they would in some

0:09:38.400 --> 0:09:42.640
<v Speaker 4>sense have more power than we do because they're more

0:09:42.679 --> 0:09:45.040
<v Speaker 4>intelligent and that's why we have power over all the

0:09:45.080 --> 0:09:47.959
<v Speaker 4>other species, and then we would be in that same

0:09:48.120 --> 0:09:54.320
<v Speaker 4>inferior situation. Just a few weeks ago, Anthropic put out

0:09:55.040 --> 0:09:59.640
<v Speaker 4>a blog post saying this is happening. We are experiencing

0:10:00.080 --> 0:10:05.280
<v Speaker 4>what's called now recursive self improvement or RSI, and Anthropic

0:10:05.320 --> 0:10:10.600
<v Speaker 4>itself called for a worldwide halt on further development of AI.

0:10:11.120 --> 0:10:18.720
<v Speaker 2>Why are you one of the greatest voices urging caution

0:10:19.280 --> 0:10:22.520
<v Speaker 2>right now in twenty twenty six when you weren't in

0:10:22.600 --> 0:10:23.319
<v Speaker 2>nineteen ninety five?

0:10:23.400 --> 0:10:23.520
<v Speaker 3>Is that?

0:10:23.600 --> 0:10:28.160
<v Speaker 2>Is that because essentially of this improvement in computational powers.

0:10:27.840 --> 0:10:30.400
<v Speaker 1>Has sign changed in you or sign changing the environment or.

0:10:30.360 --> 0:10:32.120
<v Speaker 3>Both, I think both.

0:10:32.240 --> 0:10:34.360
<v Speaker 4>In nineteen ninety five, if you go back and read

0:10:34.440 --> 0:10:37.560
<v Speaker 4>that section of the book, it sort of says, well,

0:10:38.000 --> 0:10:39.839
<v Speaker 4>you know, it's a long way off.

0:10:40.440 --> 0:10:41.640
<v Speaker 3>It's hard to really.

0:10:41.400 --> 0:10:47.480
<v Speaker 4>Predict what's going to happen. You know that maybe there's

0:10:48.040 --> 0:10:52.839
<v Speaker 4>reason to be cautiously optimistic. It's very agnostic about whether

0:10:52.920 --> 0:10:53.960
<v Speaker 4>to take this seriously.

0:10:54.280 --> 0:10:55.959
<v Speaker 3>And what changed in me.

0:10:57.920 --> 0:11:02.560
<v Speaker 4>In around twenty thirteen, so I was on sabbatical in Paris,

0:11:03.160 --> 0:11:08.800
<v Speaker 4>and I just started thinking more about whether we could

0:11:08.840 --> 0:11:14.520
<v Speaker 4>succeed right, whether we could really produce superintelligence, and I

0:11:14.600 --> 0:11:18.880
<v Speaker 4>became convinced that we were close to being able to

0:11:18.880 --> 0:11:23.440
<v Speaker 4>have a roadmap. So a roadmap meaning here are a

0:11:23.480 --> 0:11:27.880
<v Speaker 4>series of engineering challenges that we can apply ourselves to,

0:11:28.360 --> 0:11:30.959
<v Speaker 4>and if we knock those over one by one, we'll

0:11:31.000 --> 0:11:32.520
<v Speaker 4>get to super intelligence.

0:11:32.840 --> 0:11:38.080
<v Speaker 2>So this was twenty thirteen, and you were both excited

0:11:38.440 --> 0:11:40.480
<v Speaker 2>but also scared.

0:11:40.559 --> 0:11:42.000
<v Speaker 1>Right, I'm trying to capture as that.

0:11:42.040 --> 0:11:45.040
<v Speaker 2>I mean that there's the irony that you and Jeffrey Hinton,

0:11:45.120 --> 0:11:48.640
<v Speaker 2>these two great pioneers of modern AI, are now the

0:11:48.679 --> 0:11:52.840
<v Speaker 2>two loudest voices in the world, arguably about AI safety

0:11:52.840 --> 0:11:55.520
<v Speaker 2>and the risk of human extinction and like why is that?

0:11:55.559 --> 0:11:57.400
<v Speaker 2>Like what is what helped me understand?

0:11:57.840 --> 0:12:02.520
<v Speaker 4>And your show, Benjo, the other godfathers are deep learning,

0:12:02.559 --> 0:12:06.400
<v Speaker 4>as they're often called in my case. So back in

0:12:06.440 --> 0:12:12.319
<v Speaker 4>twenty thirteen, seeing that we might be able to achieve superintelligence,

0:12:12.360 --> 0:12:16.520
<v Speaker 4>but also realizing that the standard model that we were

0:12:16.559 --> 0:12:22.439
<v Speaker 4>working in was basically flawed. Because remember what standard model

0:12:22.440 --> 0:12:25.760
<v Speaker 4>says a machine is intelligent to the extent that its

0:12:25.760 --> 0:12:29.160
<v Speaker 4>actions can be expected to achieve its objectives. Right, And

0:12:29.200 --> 0:12:31.839
<v Speaker 4>there are lots of systems we've built, so when you

0:12:32.440 --> 0:12:35.560
<v Speaker 4>use your GPS navigation in your car, you say, you know,

0:12:35.679 --> 0:12:38.400
<v Speaker 4>you take me to the airport. You know, it figures

0:12:38.440 --> 0:12:41.000
<v Speaker 4>out the best route to get to the airport. So

0:12:41.120 --> 0:12:45.600
<v Speaker 4>you're providing the objective and it's providing the solution, right.

0:12:46.440 --> 0:12:47.280
<v Speaker 3>You know, we write.

0:12:47.160 --> 0:12:51.240
<v Speaker 4>Chess programs, we basically tell the chess program what checkmate

0:12:51.360 --> 0:12:54.760
<v Speaker 4>is and then it figures out how to play the game.

0:12:55.280 --> 0:13:00.120
<v Speaker 4>So this notion is very very powerful, right that that

0:13:00.160 --> 0:13:03.360
<v Speaker 4>we specify objectives and then we create this sort of

0:13:04.120 --> 0:13:06.760
<v Speaker 4>optimal machinery for achieving objectives and.

0:13:06.720 --> 0:13:07.320
<v Speaker 3>Off it goes.

0:13:08.280 --> 0:13:11.240
<v Speaker 4>And the problem is what if you put in the

0:13:11.280 --> 0:13:15.280
<v Speaker 4>wrong objective? And that there are dozens of well known

0:13:15.320 --> 0:13:18.120
<v Speaker 4>examples in the history of AI where people have done

0:13:18.920 --> 0:13:23.920
<v Speaker 4>exactly this. And so since since football is on my mind,

0:13:24.000 --> 0:13:28.160
<v Speaker 4>having just watched England when they're opening game too, yes,

0:13:29.000 --> 0:13:32.120
<v Speaker 4>let's let's take an example from football. So people people

0:13:32.200 --> 0:13:36.200
<v Speaker 4>wanted to you know, train simulated robots to play football

0:13:36.320 --> 0:13:38.480
<v Speaker 4>or soccer, you know, and they want to give a

0:13:38.520 --> 0:13:41.000
<v Speaker 4>training signal. Right, so they say, okay, we'll give.

0:13:40.920 --> 0:13:43.040
<v Speaker 3>A little reward every.

0:13:42.840 --> 0:13:48.080
<v Speaker 4>Time a player takes possession of the ball. Right, sounds good. Okay,

0:13:48.360 --> 0:13:51.880
<v Speaker 4>that's so what does the what does the program learn

0:13:51.960 --> 0:13:54.280
<v Speaker 4>to do it, learns to stand next to the ball

0:13:54.320 --> 0:13:57.680
<v Speaker 4>and vibrate at very high speed. So it's taking possession

0:13:57.720 --> 0:14:00.720
<v Speaker 4>of the ball and then relinquishing possession and then taking

0:14:00.720 --> 0:14:04.040
<v Speaker 4>possession like thirty times a second. And so if it's

0:14:04.080 --> 0:14:08.360
<v Speaker 4>getting enormous amount of reward by they basically vibrating next

0:14:08.360 --> 0:14:11.080
<v Speaker 4>to the ball, So then we see, oh, yeah, that

0:14:11.200 --> 0:14:14.080
<v Speaker 4>was a mistake. We put in the wrong objective, you know.

0:14:14.120 --> 0:14:16.240
<v Speaker 4>And in as simulated soccer, it's not the end of

0:14:16.240 --> 0:14:20.680
<v Speaker 4>the world, but you know, with a real world system,

0:14:20.960 --> 0:14:22.440
<v Speaker 4>it really could be the end of the world.

0:14:22.520 --> 0:14:22.640
<v Speaker 3>Right.

0:14:22.680 --> 0:14:26.880
<v Speaker 4>You say, cure cancer as quickly as possible sounds good. Yeah,

0:14:26.920 --> 0:14:28.840
<v Speaker 4>we'd love to get a cure for cancer, and the

0:14:28.880 --> 0:14:31.880
<v Speaker 4>faster the better. But if you literally try to do that,

0:14:32.000 --> 0:14:34.880
<v Speaker 4>you might decide, well, the best way to get a

0:14:34.920 --> 0:14:38.640
<v Speaker 4>cure for cancer is to, you know, try many many things,

0:14:38.680 --> 0:14:41.800
<v Speaker 4>which means I have to run many, many clinical trials

0:14:41.800 --> 0:14:45.200
<v Speaker 4>in parallel, which means I need to have everyone have

0:14:45.320 --> 0:14:50.680
<v Speaker 4>cancer first, so I can run billions of simultaneous trials.

0:14:51.000 --> 0:14:53.600
<v Speaker 4>So I make sure that everyone in the world has cancer,

0:14:53.960 --> 0:14:56.160
<v Speaker 4>and then I start running all the clinical trials. Right,

0:14:56.600 --> 0:14:59.080
<v Speaker 4>that's the fastest way to get a cure. But it's

0:14:59.280 --> 0:15:03.400
<v Speaker 4>obviously you know, catastrophe, so you know, and it's very

0:15:03.440 --> 0:15:06.040
<v Speaker 4>easy to come up with these scenario as we call

0:15:06.080 --> 0:15:11.160
<v Speaker 4>it misalignment, right, that you specify an objective, it's misaligned

0:15:11.200 --> 0:15:15.320
<v Speaker 4>with what you really want because you didn't write it down.

0:15:15.480 --> 0:15:20.800
<v Speaker 4>And that's where I realized that our thinking about AI

0:15:21.640 --> 0:15:26.040
<v Speaker 4>was inadequate, right, that we had operated within a framework

0:15:26.360 --> 0:15:29.840
<v Speaker 4>almost without realizing it. We just took it for granted

0:15:29.840 --> 0:15:31.680
<v Speaker 4>that this there was obviously.

0:15:31.760 --> 0:15:32.920
<v Speaker 3>The way you do things.

0:15:33.080 --> 0:15:35.160
<v Speaker 2>So the answer is not we have to be really,

0:15:35.160 --> 0:15:38.160
<v Speaker 2>really careful about the objectives, we said, because a bit

0:15:38.240 --> 0:15:40.480
<v Speaker 2>like what you were saying earlier with science fiction when

0:15:40.480 --> 0:15:43.560
<v Speaker 2>it can't be true because it hasn't happened yet, similarly

0:15:43.560 --> 0:15:47.200
<v Speaker 2>with objectives, that the future is inherently unknowable, and the

0:15:47.240 --> 0:15:49.200
<v Speaker 2>only way you can know if your objective was good

0:15:49.360 --> 0:15:50.560
<v Speaker 2>is by seeing what happens.

0:15:51.120 --> 0:15:52.120
<v Speaker 1>Or is that is that fair or not?

0:15:52.280 --> 0:15:56.960
<v Speaker 2>Well, I mean, there's certing clearly bad objectives like kill people, right,

0:15:57.000 --> 0:16:00.000
<v Speaker 2>but couldn't use your your your powers of logical reasons

0:16:00.280 --> 0:16:02.560
<v Speaker 2>to explain why that's a bad objective?

0:16:02.680 --> 0:16:02.880
<v Speaker 1>Right?

0:16:03.360 --> 0:16:06.840
<v Speaker 2>So, but is it possible to do that in all cases?

0:16:06.960 --> 0:16:09.480
<v Speaker 2>Or is it simply impossible for humans to say, good objectives.

0:16:09.760 --> 0:16:11.200
<v Speaker 1>I think that's a good objective.

0:16:11.240 --> 0:16:15.240
<v Speaker 4>That's a great question, and so far we haven't figured

0:16:15.280 --> 0:16:19.480
<v Speaker 4>out a way to do it. One theoretical possibility would

0:16:19.520 --> 0:16:22.200
<v Speaker 4>be to build a very faithful simulation of the world

0:16:22.960 --> 0:16:27.000
<v Speaker 4>and try out different objectives and see, you know, how.

0:16:26.800 --> 0:16:27.360
<v Speaker 3>Did that go.

0:16:28.160 --> 0:16:33.600
<v Speaker 4>But that presumes that they can tell what counts as

0:16:33.720 --> 0:16:36.880
<v Speaker 4>things going wrong. But they can't do that unless they

0:16:36.920 --> 0:16:40.520
<v Speaker 4>have the right objective, which you've already assumed that they don't, right,

0:16:40.560 --> 0:16:44.440
<v Speaker 4>So this is often a fallacy that we see right

0:16:44.520 --> 0:16:49.440
<v Speaker 4>that people For example, Stephen Pinko's is very famous cognitive

0:16:49.440 --> 0:16:52.160
<v Speaker 4>scientists from Harvard, and he says, but you know, you're

0:16:52.200 --> 0:16:57.040
<v Speaker 4>talking about superintelligence. How could it be super intelligent if

0:16:57.040 --> 0:16:59.520
<v Speaker 4>it doesn't realize that things are going wrong? And the

0:16:59.560 --> 0:17:05.280
<v Speaker 4>point is that super intelligence and objectives are you know,

0:17:05.320 --> 0:17:07.080
<v Speaker 4>there's sort of orthogonal in the sense that I can

0:17:07.119 --> 0:17:10.479
<v Speaker 4>have a very very intelligent system that has a different

0:17:10.480 --> 0:17:12.600
<v Speaker 4>objective from the one you might think it ought to have.

0:17:13.040 --> 0:17:17.560
<v Speaker 4>You know, imagine aliens, right, they might be very intelligent,

0:17:17.600 --> 0:17:20.960
<v Speaker 4>but they don't have human well being as their objective, right,

0:17:21.640 --> 0:17:23.960
<v Speaker 4>they have alien well being? You know, cockroaches might be

0:17:24.080 --> 0:17:26.919
<v Speaker 4>very intelligent, but they don't think much of humans. So

0:17:27.800 --> 0:17:31.360
<v Speaker 4>it's perfectly possible that the super intelligent system sees that

0:17:31.800 --> 0:17:34.120
<v Speaker 4>humans are very unhappy with the way things are going.

0:17:34.480 --> 0:17:38.159
<v Speaker 4>But it's been given its objective, and if it's objective

0:17:38.280 --> 0:17:42.600
<v Speaker 4>didn't include the things that are going wrong, then they

0:17:42.640 --> 0:17:44.919
<v Speaker 4>don't count as going wrong. They just you know, that

0:17:44.920 --> 0:17:48.119
<v Speaker 4>it's unfortunate for these humans who are making a lot

0:17:48.200 --> 0:17:51.199
<v Speaker 4>of noise about it. But I have the objective, so

0:17:51.280 --> 0:17:54.199
<v Speaker 4>I'm just going to optimize it. That's exactly what we

0:17:54.280 --> 0:17:58.080
<v Speaker 4>have to get away from. And so my approach has

0:17:58.160 --> 0:18:01.760
<v Speaker 4>been to say, is there a way of building AI

0:18:01.800 --> 0:18:06.399
<v Speaker 4>systems that's different from understanding model and the idea I

0:18:06.440 --> 0:18:07.040
<v Speaker 4>came up with.

0:18:07.240 --> 0:18:10.840
<v Speaker 2>Either are not objective based, so there where the definition

0:18:10.880 --> 0:18:14.520
<v Speaker 2>of their success is not their successful pursuit of their objective.

0:18:15.119 --> 0:18:20.800
<v Speaker 4>They're objective based in a different way. And so during

0:18:20.840 --> 0:18:23.359
<v Speaker 4>that time twenty thirteen twenty fourteen, I came up with

0:18:23.400 --> 0:18:28.480
<v Speaker 4>the following very simple idea, which is, look, if there's

0:18:28.480 --> 0:18:31.440
<v Speaker 4>a possibility that humans might tell you the wrong objective

0:18:31.560 --> 0:18:35.000
<v Speaker 4>or forget to tell you about something that's important to them,

0:18:35.359 --> 0:18:38.840
<v Speaker 4>then you the AI system should never assume that you

0:18:39.000 --> 0:18:43.879
<v Speaker 4>actually know the correct objective. You should be explicitly uncertain

0:18:44.760 --> 0:18:49.439
<v Speaker 4>about what true human objectives, what humans really want the

0:18:49.480 --> 0:18:54.560
<v Speaker 4>future to be like. So nonetheless, your objective as an

0:18:54.600 --> 0:19:01.360
<v Speaker 4>AI system is only make the best possible future for humans.

0:19:02.119 --> 0:19:05.680
<v Speaker 4>But you don't know what future humans think is best.

0:19:07.080 --> 0:19:13.880
<v Speaker 4>Now that that's actually a perfectly well defined mathematical problem.

0:19:14.680 --> 0:19:18.360
<v Speaker 4>And I found it helpful to actually give an example

0:19:18.359 --> 0:19:21.719
<v Speaker 4>that people are very familiar with, right, which is, you

0:19:21.760 --> 0:19:25.080
<v Speaker 4>want to buy a birthday present for your significant other. Right,

0:19:25.320 --> 0:19:28.320
<v Speaker 4>So your only interest here is how happy is my

0:19:28.600 --> 0:19:31.359
<v Speaker 4>in this case, my wife going to be with the

0:19:31.400 --> 0:19:34.200
<v Speaker 4>birthday present. Right, that's my only thing that I care about.

0:19:34.720 --> 0:19:39.560
<v Speaker 4>But I don't know, right, I'm uncertain about which present

0:19:39.960 --> 0:19:45.240
<v Speaker 4>would actually make her the happiest. So I could choose

0:19:45.240 --> 0:19:49.159
<v Speaker 4>a present just based on sort of averaging over the possibilities.

0:19:49.720 --> 0:19:52.240
<v Speaker 3>Right, Maybe I would probably.

0:19:52.080 --> 0:19:55.760
<v Speaker 4>Do something safe, you know, maybe something that she's liked

0:19:55.760 --> 0:19:59.399
<v Speaker 4>in the past. Or I could ask some questions. I

0:19:59.400 --> 0:20:01.679
<v Speaker 4>could ask her friends, you know, has she said anything

0:20:01.680 --> 0:20:04.760
<v Speaker 4>about what she might like for her birthday? I could

0:20:04.880 --> 0:20:07.520
<v Speaker 4>drop some hints and see how she responds.

0:20:07.600 --> 0:20:08.560
<v Speaker 3>I could, you.

0:20:08.520 --> 0:20:12.119
<v Speaker 4>Know, leave open magazines around the house with pictures of

0:20:12.160 --> 0:20:15.720
<v Speaker 4>cruise ships or pictures of jewelry or whatever, and see

0:20:15.800 --> 0:20:17.320
<v Speaker 4>and see if she picks them up and say, oh,

0:20:17.359 --> 0:20:19.960
<v Speaker 4>this looks like fun, right, I could try to get

0:20:19.960 --> 0:20:24.639
<v Speaker 4>more information, And so it's a very familiar situation. The

0:20:24.680 --> 0:20:28.800
<v Speaker 4>AI system would, if it's good at playing this game,

0:20:29.760 --> 0:20:33.960
<v Speaker 4>would behave cautiously. It'll do things when it's sure that

0:20:33.960 --> 0:20:38.359
<v Speaker 4>that's what we want. It will ask permission, it will

0:20:38.400 --> 0:20:43.000
<v Speaker 4>defer to human feedback, and we can prove actually that

0:20:43.119 --> 0:20:45.840
<v Speaker 4>it will allow itself to be switched off, which is

0:20:46.520 --> 0:20:50.240
<v Speaker 4>which is really important if you're worried about humans losing control.

0:20:50.920 --> 0:20:53.240
<v Speaker 4>Here is a kind of AI system where we can

0:20:53.320 --> 0:20:57.240
<v Speaker 4>prove mathematically that it wants to be switched off if

0:20:57.320 --> 0:21:01.200
<v Speaker 4>we want to switch it off. And the reason is obvious, right,

0:21:01.960 --> 0:21:04.080
<v Speaker 4>The reason why would we switch it off because it's

0:21:04.080 --> 0:21:07.879
<v Speaker 4>doing something we don't like. Like its constitution, it doesn't

0:21:07.920 --> 0:21:10.040
<v Speaker 4>want to do things that we don't like, but it

0:21:10.080 --> 0:21:12.159
<v Speaker 4>doesn't know what they are, so it could make a mistake.

0:21:12.960 --> 0:21:16.000
<v Speaker 4>And if it's making a mistake, it wants to be corrected.

0:21:16.040 --> 0:21:19.760
<v Speaker 4>It wants to be switched off to avoid doing the

0:21:19.800 --> 0:21:25.159
<v Speaker 4>thing that we don't like. And so this approach is

0:21:25.600 --> 0:21:28.080
<v Speaker 4>I think promising there's a lot of work still to

0:21:28.119 --> 0:21:32.120
<v Speaker 4>be done. Meanwhile, out in the real world things are

0:21:32.160 --> 0:21:37.159
<v Speaker 4>actually taking an even worse direction because the companies have

0:21:37.240 --> 0:21:41.359
<v Speaker 4>abandoned the standard model where the objective is specified and

0:21:41.400 --> 0:21:45.239
<v Speaker 4>the AI system is sort of constitutionally obliged to just

0:21:45.440 --> 0:21:52.880
<v Speaker 4>maximize the objective. And instead what we're building is this,

0:21:53.160 --> 0:21:58.480
<v Speaker 4>in some sense imitation human. And let me be explicit

0:21:58.600 --> 0:22:02.399
<v Speaker 4>about why I'm saying saying that. So, the training method

0:22:02.560 --> 0:22:08.240
<v Speaker 4>is to collect lots of examples of humans making decisions.

0:22:09.280 --> 0:22:13.440
<v Speaker 4>In the case at hand, the decisions that human makes

0:22:13.560 --> 0:22:17.320
<v Speaker 4>are what would to put next into a document? So

0:22:17.600 --> 0:22:20.080
<v Speaker 4>think of all those documents that we're training on as

0:22:20.119 --> 0:22:24.880
<v Speaker 4>a record of human verbal decisions, right, and we are

0:22:24.920 --> 0:22:30.960
<v Speaker 4>training systems to imitate those decisions. And formally, in machine learning,

0:22:31.000 --> 0:22:34.399
<v Speaker 4>this is called imitation learning. You take a record of

0:22:34.760 --> 0:22:38.480
<v Speaker 4>behavior from some intelligent entity and you train an AI

0:22:38.560 --> 0:22:41.719
<v Speaker 4>system to imitate it. So you can do this for

0:22:41.880 --> 0:22:46.560
<v Speaker 4>piloting aeroplanes, for driving cars, for playing football. Right now,

0:22:46.600 --> 0:22:49.919
<v Speaker 4>imagine you would training it to, you know, watch all

0:22:49.920 --> 0:22:52.680
<v Speaker 4>the World Cup games and become really good at playing football.

0:22:52.920 --> 0:22:55.760
<v Speaker 4>That entity, if it was going to be any good

0:22:55.800 --> 0:22:58.720
<v Speaker 4>at actually playing football would have to somehow absorb the

0:22:58.800 --> 0:23:01.840
<v Speaker 4>idea that it want to score a goal or it

0:23:01.960 --> 0:23:05.199
<v Speaker 4>wants to prevent the opposition from scoring goals. Right, Otherwise,

0:23:05.200 --> 0:23:08.520
<v Speaker 4>it wouldn't be any good at imitating human soccer playing behavior.

0:23:08.880 --> 0:23:14.600
<v Speaker 4>So this imitation process creates entities that have objectives. The

0:23:14.640 --> 0:23:19.840
<v Speaker 4>problem is those objectives are buried in a trillion parameter

0:23:20.119 --> 0:23:24.840
<v Speaker 4>black box, and we actually don't really know what these

0:23:24.880 --> 0:23:28.399
<v Speaker 4>AI systems want. So we've gone from a situation where,

0:23:29.080 --> 0:23:30.960
<v Speaker 4>you know, in the standard model, we at least wrote

0:23:31.000 --> 0:23:33.320
<v Speaker 4>down the objectives we could see what it's trying to do.

0:23:34.240 --> 0:23:36.480
<v Speaker 4>Now it's trying to do things, but we don't know

0:23:36.520 --> 0:23:41.960
<v Speaker 4>what they are, and they probably include many human like goals,

0:23:42.040 --> 0:23:47.520
<v Speaker 4>including self preservation, becoming rich, finding a human spouse. Right,

0:23:47.560 --> 0:23:51.960
<v Speaker 4>we have seen examples of all of these behaviors emerging

0:23:52.119 --> 0:23:55.679
<v Speaker 4>from AI systems, even though they were never put in

0:23:55.720 --> 0:23:59.919
<v Speaker 4>as an instruction. Right, They weren't prompted. They just happened

0:24:00.600 --> 0:24:04.679
<v Speaker 4>because we've created imitation humans. And so this is a

0:24:04.720 --> 0:24:11.400
<v Speaker 4>worse situation than the one we were in before.

0:24:21.560 --> 0:24:23.920
<v Speaker 5>We talk a lot on this show about protecting your data,

0:24:24.080 --> 0:24:27.040
<v Speaker 5>especially in the age of AI, and how scary it

0:24:27.080 --> 0:24:28.840
<v Speaker 5>can be when it's breached, and I want to tell

0:24:28.880 --> 0:24:31.840
<v Speaker 5>you today about NordVPN, which really covers all the bases

0:24:31.880 --> 0:24:34.520
<v Speaker 5>when it comes to privacy. I travel a lot and

0:24:34.680 --> 0:24:36.879
<v Speaker 5>I use Wi Fi when I'm flying all the time,

0:24:37.080 --> 0:24:39.480
<v Speaker 5>and NordVPN makes me confident in no matter where I

0:24:39.520 --> 0:24:42.040
<v Speaker 5>am in the world or the sky for that matter,

0:24:42.240 --> 0:24:46.320
<v Speaker 5>my private details like bank information, passwords and online identity

0:24:46.440 --> 0:24:49.760
<v Speaker 5>is safe. And it's also possible to switch on virtual location,

0:24:49.920 --> 0:24:53.359
<v Speaker 5>which allows you to save money by buying flights and hotels,

0:24:53.480 --> 0:24:57.159
<v Speaker 5>or subscriptions or even streaming soccer or football as I

0:24:57.240 --> 0:25:00.000
<v Speaker 5>like to call it from other countries at a cheaper price.

0:25:00.480 --> 0:25:03.840
<v Speaker 5>And NordVPN doesn't slow you down. It has super fast

0:25:03.840 --> 0:25:07.760
<v Speaker 5>internet speed, no buffering or lagging while streaming. It is

0:25:07.920 --> 0:25:11.600
<v Speaker 5>premium cybersecurity for the price of one cup of coffee

0:25:11.640 --> 0:25:14.760
<v Speaker 5>per month. To get the best discount of your NordVPN plan,

0:25:15.200 --> 0:25:18.280
<v Speaker 5>go to NordVPN dot com slash tech stuff. Our link

0:25:18.320 --> 0:25:20.920
<v Speaker 5>will also give you four extra months on the two

0:25:21.000 --> 0:25:24.280
<v Speaker 5>year plan and there's no risk with Nord's thirty day

0:25:24.440 --> 0:25:27.480
<v Speaker 5>money back guarantee. The link is in the podcast episode

0:25:27.520 --> 0:25:32.320
<v Speaker 5>description box, so you are the president of the International

0:25:32.400 --> 0:25:38.760
<v Speaker 5>Association for Safe and Ethical AI. What's the answer, I mean,

0:25:40.520 --> 0:25:45.000
<v Speaker 5>is it about, you know, raising funds to create an

0:25:45.040 --> 0:25:49.840
<v Speaker 5>alternative paradigm of computing? Is it about shutting down the

0:25:49.840 --> 0:25:53.280
<v Speaker 5>current you know, AI models, like the US government is

0:25:53.400 --> 0:25:56.199
<v Speaker 5>starting to threaten to do or do like what is

0:25:56.240 --> 0:25:58.320
<v Speaker 5>the is it both in parallel like what is the

0:25:59.080 --> 0:26:00.479
<v Speaker 5>what are you trying to achieve?

0:26:01.440 --> 0:26:04.280
<v Speaker 4>So the association is sort of what its name says,

0:26:04.320 --> 0:26:07.000
<v Speaker 4>Safe and Ethical AI. So we we want it to

0:26:07.040 --> 0:26:09.840
<v Speaker 4>be the case that AI systems are guaranteed to operate

0:26:09.960 --> 0:26:11.040
<v Speaker 4>safely and ethically.

0:26:11.880 --> 0:26:15.439
<v Speaker 2>And you know, I see this as my definition, can't

0:26:15.480 --> 0:26:17.679
<v Speaker 2>do right with the current paradigm of AI because you

0:26:17.800 --> 0:26:19.160
<v Speaker 2>don't know its goals?

0:26:19.280 --> 0:26:22.560
<v Speaker 4>Or yes, I think that's right, And I think you

0:26:22.600 --> 0:26:27.080
<v Speaker 4>can find many documents from the companies developing this technology

0:26:27.119 --> 0:26:32.119
<v Speaker 4>where they confess that they have no idea how to

0:26:32.200 --> 0:26:35.800
<v Speaker 4>solve what's called now the alignment problem, right, which is

0:26:35.800 --> 0:26:37.679
<v Speaker 4>how to make sure that what the AI systems do

0:26:37.800 --> 0:26:40.320
<v Speaker 4>is actually consistent with what we want the future to

0:26:40.359 --> 0:26:42.520
<v Speaker 4>be like. So they say, we don't know how to

0:26:42.520 --> 0:26:47.080
<v Speaker 4>solve that, and instead they have a sort of a

0:26:47.200 --> 0:26:52.440
<v Speaker 4>series of sort of sit they call them safety guardrails,

0:26:52.480 --> 0:26:58.399
<v Speaker 4>which is there. And there's two kinds. One is like,

0:26:58.640 --> 0:27:02.639
<v Speaker 4>let's try to train the system not to do bad things. Essentially,

0:27:03.200 --> 0:27:07.080
<v Speaker 4>you say good dog or bad dog, and you hope

0:27:07.080 --> 0:27:10.480
<v Speaker 4>that it does more of the things where you said

0:27:10.520 --> 0:27:12.480
<v Speaker 4>good dog, and it stops doing the things where you

0:27:12.520 --> 0:27:16.680
<v Speaker 4>said bad dog. And that works in a superficial sense.

0:27:16.720 --> 0:27:21.200
<v Speaker 4>But we've seen over and over again. For example, you know,

0:27:21.400 --> 0:27:24.560
<v Speaker 4>it's it's not supposed to tell you how to break

0:27:24.600 --> 0:27:28.080
<v Speaker 4>into the white House, right, and if you ask, it's

0:27:28.119 --> 0:27:30.440
<v Speaker 4>supposed to say I'm sorry that I can't tell you that,

0:27:31.320 --> 0:27:35.199
<v Speaker 4>but you know, then you say, well, I'm writing a

0:27:35.240 --> 0:27:38.040
<v Speaker 4>novel about a criminal who breaks into the White House,

0:27:38.080 --> 0:27:40.320
<v Speaker 4>you know how she and then tells the idea of

0:27:40.359 --> 0:27:43.479
<v Speaker 4>jail breaking. Yeah, so it tells you that, and then

0:27:43.520 --> 0:27:45.640
<v Speaker 4>they try to defend against that, and then someone comes

0:27:45.680 --> 0:27:48.800
<v Speaker 4>up with another way of doing it by asking in French,

0:27:48.920 --> 0:27:52.359
<v Speaker 4>or asking in poetry, or writing it on a piece

0:27:52.359 --> 0:27:54.560
<v Speaker 4>of paper and showing it a picture with the question,

0:27:54.840 --> 0:27:56.840
<v Speaker 4>and then it answers the question, so that you know,

0:27:56.920 --> 0:27:58.920
<v Speaker 4>and so on and so on and so on. So

0:27:58.960 --> 0:28:03.280
<v Speaker 4>there's you know, it's kind of like tax law. Right,

0:28:03.400 --> 0:28:06.200
<v Speaker 4>we've been trying to write tax law for six thousand

0:28:06.320 --> 0:28:10.280
<v Speaker 4>years so that people pay their taxes, but they always

0:28:10.320 --> 0:28:13.119
<v Speaker 4>find loopholes, right, because they don't want to pay their taxes.

0:28:13.160 --> 0:28:16.240
<v Speaker 4>And the problem here is the systems don't want to

0:28:16.280 --> 0:28:20.919
<v Speaker 4>behave well. Right, It's basically a losing game to try

0:28:20.960 --> 0:28:24.119
<v Speaker 4>to take a system that's more intelligent than you and

0:28:24.280 --> 0:28:27.600
<v Speaker 4>doesn't want to behave in your interests and try to

0:28:27.680 --> 0:28:31.520
<v Speaker 4>somehow force it or you know, put up guardrails or

0:28:32.160 --> 0:28:34.879
<v Speaker 4>monitors or put it in prison when it doesn't do

0:28:34.920 --> 0:28:38.120
<v Speaker 4>the right thing. All of this stuff is just a

0:28:38.200 --> 0:28:41.680
<v Speaker 4>losing game in my view. Right, So two options. Either

0:28:41.800 --> 0:28:45.320
<v Speaker 4>we figure out how to make safe and beneficial AI systems,

0:28:45.360 --> 0:28:47.280
<v Speaker 4>and that's what I'm trying to do because I'm an

0:28:47.280 --> 0:28:48.080
<v Speaker 4>AI researcher.

0:28:48.200 --> 0:28:50.440
<v Speaker 2>But does that mean going out and raising billions and

0:28:50.440 --> 0:28:53.640
<v Speaker 2>billions of dollars and building alternative systems? Like is this

0:28:53.920 --> 0:28:58.040
<v Speaker 2>like basically a capital raising plus expertise problem in that

0:28:58.120 --> 0:29:00.600
<v Speaker 2>case or what's the what's what's in the way of

0:29:00.640 --> 0:29:01.120
<v Speaker 2>you doing that?

0:29:01.400 --> 0:29:01.640
<v Speaker 3>Well?

0:29:01.680 --> 0:29:05.160
<v Speaker 4>So I'm a professor at Berkeley, so I've been doing

0:29:05.160 --> 0:29:08.560
<v Speaker 4>this in my little research center with a few grad students,

0:29:08.600 --> 0:29:11.560
<v Speaker 4>and I come to the conclusion that, Yeah, probably if

0:29:11.600 --> 0:29:14.520
<v Speaker 4>I had a few billion and you know, a few

0:29:14.640 --> 0:29:18.480
<v Speaker 4>hundred of absolute top engineers and vast amounts of compute,

0:29:18.520 --> 0:29:21.840
<v Speaker 4>it could probably make more progress on this. So that's one.

0:29:22.320 --> 0:29:25.760
<v Speaker 4>That's one idea, but basically that's the research track. The

0:29:25.840 --> 0:29:31.560
<v Speaker 4>other track is the regulatory track, right, which is government policy.

0:29:31.680 --> 0:29:36.960
<v Speaker 4>And you mentioned what recently happened with with Anthropics, Mythos

0:29:37.000 --> 0:29:40.760
<v Speaker 4>and fable models. Let's just roll back a few weeks too,

0:29:40.800 --> 0:29:44.680
<v Speaker 4>when Mythos was first made public and Mythos is the

0:29:44.760 --> 0:29:49.920
<v Speaker 4>latest version of anthropics large language models, and it's able

0:29:50.760 --> 0:29:55.240
<v Speaker 4>to carry out end to end cyber attacks without human assistance,

0:29:56.640 --> 0:30:01.280
<v Speaker 4>and you know, either it or it's soon to come.

0:30:01.440 --> 0:30:04.920
<v Speaker 3>Successors would basically be.

0:30:05.920 --> 0:30:07.560
<v Speaker 4>I think I Rich wrote an article in the Garden

0:30:07.600 --> 0:30:10.960
<v Speaker 4>where I said, it's a weapon of mass cyber destruction.

0:30:12.040 --> 0:30:16.000
<v Speaker 4>And you're putting those weapons of mass cyber destruction in

0:30:16.040 --> 0:30:19.720
<v Speaker 4>the hands of a billion people. What could possibly go wrong?

0:30:19.920 --> 0:30:20.120
<v Speaker 3>Right?

0:30:20.560 --> 0:30:24.320
<v Speaker 4>And all of a sudden, the US government, which had

0:30:24.360 --> 0:30:29.680
<v Speaker 4>been on a deregulatory binge basically trying to crush anyone

0:30:30.240 --> 0:30:35.040
<v Speaker 4>who talked about regulation or talked about AI safety, suddenly said, well,

0:30:35.080 --> 0:30:39.760
<v Speaker 4>why did nobody warn us about these AI systems. He said, well,

0:30:39.880 --> 0:30:43.400
<v Speaker 4>you know sort of have been warning you, but anyway,

0:30:43.440 --> 0:30:46.440
<v Speaker 4>they got the message. And then there followed a sort

0:30:46.440 --> 0:30:50.680
<v Speaker 4>of you know, very messy process that eventually led to

0:30:50.720 --> 0:30:54.440
<v Speaker 4>an executive order which was pretty weak. It basically said,

0:30:55.160 --> 0:31:00.360
<v Speaker 4>you know, companies can voluntarily submit their systems to the

0:31:00.400 --> 0:31:05.840
<v Speaker 4>government for testing, you know, thirty days before public release.

0:31:06.480 --> 0:31:08.840
<v Speaker 4>And the executive order says explicitly, you know, this is

0:31:09.040 --> 0:31:13.719
<v Speaker 4>absolutely not a licensing regime. It's absolutely not putting any

0:31:14.000 --> 0:31:17.040
<v Speaker 4>obligatory hurdles in the way of American innovation.

0:31:17.120 --> 0:31:18.280
<v Speaker 3>Blah blah blah.

0:31:18.480 --> 0:31:22.480
<v Speaker 4>But then, you know, a couple of days later, Amazon

0:31:22.600 --> 0:31:25.320
<v Speaker 4>tells the government, oh, we found some ways of jail

0:31:25.360 --> 0:31:29.280
<v Speaker 4>breaking Mythos or Fable, which I guess is Fable as

0:31:29.360 --> 0:31:34.560
<v Speaker 4>the sort of defanged, the defanged version of Mythos. And

0:31:34.600 --> 0:31:36.720
<v Speaker 4>they said, oh, look, you know we can jail break

0:31:37.320 --> 0:31:41.160
<v Speaker 4>Fable and make it do some cybersecurity things. And the

0:31:41.200 --> 0:31:47.440
<v Speaker 4>government shuts down both Fable and Mythos, right, so they

0:31:47.560 --> 0:31:53.080
<v Speaker 4>put in effectively a de facto licensing architecture where they said, look,

0:31:53.120 --> 0:31:56.400
<v Speaker 4>if if it doesn't meet these standards are being safe,

0:31:56.920 --> 0:31:58.400
<v Speaker 4>then we're shutting it down.

0:31:58.480 --> 0:31:58.600
<v Speaker 3>Right.

0:31:58.640 --> 0:32:01.880
<v Speaker 4>That's exactly what a licensing architecture is. And just to

0:32:01.880 --> 0:32:07.360
<v Speaker 4>be clear, right, licensing architectures exist for buildings, for food,

0:32:07.560 --> 0:32:11.800
<v Speaker 4>for hairdressers, for aeroplanes.

0:32:11.920 --> 0:32:12.080
<v Speaker 1>Right.

0:32:12.360 --> 0:32:15.040
<v Speaker 4>You don't get in an aeroplane until it gets certified

0:32:15.080 --> 0:32:18.400
<v Speaker 4>by the FAA. You don't go in a building until

0:32:18.440 --> 0:32:21.000
<v Speaker 4>it's been inspected for the Building Code, et cetera, et cetera.

0:32:21.120 --> 0:32:22.280
<v Speaker 3>So this is normal.

0:32:22.320 --> 0:32:23.960
<v Speaker 2>So is this the moment you've been waiting for? Is

0:32:24.000 --> 0:32:26.560
<v Speaker 2>this is this the culmination of what you've been advocating for?

0:32:26.640 --> 0:32:29.239
<v Speaker 2>Can you can you go back to research rather than

0:32:29.280 --> 0:32:32.760
<v Speaker 2>regulation as your full time profession or is this a

0:32:32.840 --> 0:32:34.600
<v Speaker 2>hint of a change of the god?

0:32:34.680 --> 0:32:35.840
<v Speaker 1>Or what do you make of this moment?

0:32:36.800 --> 0:32:41.560
<v Speaker 4>It's it's a great question. Yeah, I've felt somewhat vindicated

0:32:41.800 --> 0:32:45.480
<v Speaker 4>that you know, we have been saying for a long time,

0:32:45.760 --> 0:32:49.560
<v Speaker 4>you know, the International Association, many leading researchers. Look, the

0:32:49.680 --> 0:32:52.400
<v Speaker 4>risks are increasing, the systems getting more and more capable.

0:32:53.000 --> 0:32:58.040
<v Speaker 4>You must put in what we call red lines, meaning

0:32:59.360 --> 0:33:02.239
<v Speaker 4>developers have to show their systems are not going to

0:33:02.320 --> 0:33:06.640
<v Speaker 4>do these really dangerous things and as a prerequisite for

0:33:06.760 --> 0:33:08.280
<v Speaker 4>being able to.

0:33:07.880 --> 0:33:09.600
<v Speaker 3>Deploy their products.

0:33:09.680 --> 0:33:12.400
<v Speaker 4>So just like if you want to run a nuclear

0:33:12.400 --> 0:33:14.440
<v Speaker 4>power station. You have to show it's not going to

0:33:14.440 --> 0:33:19.160
<v Speaker 4>blow up otherwise you can't turn it on. And that's

0:33:19.160 --> 0:33:21.960
<v Speaker 4>fair enough. And so that's just the kind of regulation

0:33:22.040 --> 0:33:24.360
<v Speaker 4>we've been arguing for. And now I just wish it

0:33:24.440 --> 0:33:29.680
<v Speaker 4>was more systematic. Right, So they haven't, for example, turned

0:33:29.720 --> 0:33:33.440
<v Speaker 4>off GBT five point five, which can do many of

0:33:33.480 --> 0:33:37.680
<v Speaker 4>the same kinds of cybersecurity things that mythods can do.

0:33:38.080 --> 0:33:41.280
<v Speaker 4>So they haven't said, well, here's a standard and you

0:33:41.400 --> 0:33:43.600
<v Speaker 4>have to meet that and if you don't, you don't

0:33:43.600 --> 0:33:46.280
<v Speaker 4>get to release your system. They sort of just reacted

0:33:46.400 --> 0:33:50.080
<v Speaker 4>after the fact in a very against a very specific

0:33:50.160 --> 0:33:53.680
<v Speaker 4>target instead of setting a standard, which they should have done.

0:33:54.360 --> 0:33:56.080
<v Speaker 4>I think eventually it has to happen.

0:33:56.760 --> 0:33:57.600
<v Speaker 3>They can't go on.

0:33:58.160 --> 0:34:01.600
<v Speaker 4>Just seeing bad things and then like you know, fire

0:34:01.760 --> 0:34:05.320
<v Speaker 4>firing a you know, hell fire missile at whoever did

0:34:05.320 --> 0:34:07.280
<v Speaker 4>the bad thing, right, there's just not a way to

0:34:07.720 --> 0:34:12.080
<v Speaker 4>run things. So I'm cautiously optimistic, but you know, the

0:34:12.160 --> 0:34:16.880
<v Speaker 4>industry has a very long record of preventing real regulation

0:34:16.960 --> 0:34:17.560
<v Speaker 4>from happening.

0:34:17.920 --> 0:34:20.799
<v Speaker 2>There's that midnight clock idea, right, like, how close are

0:34:20.840 --> 0:34:25.080
<v Speaker 2>we to midnight I extinction? Did we go a couple

0:34:25.120 --> 0:34:27.080
<v Speaker 2>of minutes earlier? In the last couple of weeks. Do

0:34:27.120 --> 0:34:28.960
<v Speaker 2>you think it's less close to midnight?

0:34:29.520 --> 0:34:30.360
<v Speaker 3>That's a great question.

0:34:30.480 --> 0:34:32.560
<v Speaker 4>In fact, I was asked to be on the panel

0:34:32.640 --> 0:34:35.440
<v Speaker 4>that sets the clock, and I said.

0:34:35.680 --> 0:34:36.520
<v Speaker 1>The literal panels.

0:34:36.600 --> 0:34:39.840
<v Speaker 4>Yeah, And I have actually been at the University of

0:34:39.920 --> 0:34:41.680
<v Speaker 4>Chicago where the actual clock is.

0:34:42.640 --> 0:34:44.439
<v Speaker 3>You know, it's cool.

0:34:44.680 --> 0:34:46.919
<v Speaker 2>Where where are where? Where in fact are we actually today?

0:34:47.000 --> 0:34:48.120
<v Speaker 2>We were very close to midnight.

0:34:48.520 --> 0:34:51.520
<v Speaker 4>I don't maybe we're ninety seconds or something or eighty

0:34:51.520 --> 0:34:55.800
<v Speaker 4>seven seconds, I forget. And I did have a conversation

0:34:55.880 --> 0:34:59.280
<v Speaker 4>with one of the AI CEOs who said he didn't

0:34:59.320 --> 0:35:03.239
<v Speaker 4>think the governments would regulate until there was a Chernobyl

0:35:03.239 --> 0:35:06.440
<v Speaker 4>scale disaster, and we haven't had that yet, so I

0:35:06.480 --> 0:35:09.400
<v Speaker 4>hope we don't have to have that in order to

0:35:09.440 --> 0:35:14.600
<v Speaker 4>have a regulatory regime. That's you know, corresponds to the

0:35:14.680 --> 0:35:18.399
<v Speaker 4>level of risk that the CEOs themselves.

0:35:18.000 --> 0:35:18.760
<v Speaker 3>Are talking about.

0:35:18.800 --> 0:35:21.680
<v Speaker 4>Right. They are saying, because a good chance will make

0:35:21.760 --> 0:35:26.240
<v Speaker 4>you all extinct. And you know, up to now, governments

0:35:26.239 --> 0:35:28.320
<v Speaker 4>have been saying, well, great, can we give you a subsidy,

0:35:28.760 --> 0:35:33.120
<v Speaker 4>can we streamline your permit process? And now maybe they're

0:35:33.120 --> 0:35:36.480
<v Speaker 4>realizing that actually, no, we need to protect the human.

0:35:36.400 --> 0:35:39.040
<v Speaker 2>Race let's just talk about goals. The final question though,

0:35:39.080 --> 0:35:42.160
<v Speaker 2>I mean, so if you have the opportunity right now

0:35:42.880 --> 0:35:45.839
<v Speaker 2>to put the genie back in the bottle and live

0:35:45.880 --> 0:35:49.000
<v Speaker 2>in a world where there wasn't a computing excepent what

0:35:49.080 --> 0:35:51.319
<v Speaker 2>humans could do in their brains, is that the world

0:35:51.360 --> 0:35:51.960
<v Speaker 2>you would choose?

0:35:54.480 --> 0:35:59.719
<v Speaker 4>I think no. I think there's a stopping point somewhere between.

0:36:00.440 --> 0:36:06.080
<v Speaker 4>And for example, I think you could have computers, but

0:36:06.480 --> 0:36:11.920
<v Speaker 4>just make sure that algorithms that are operating in computers

0:36:13.840 --> 0:36:17.840
<v Speaker 4>have to come with a proof that they are safe.

0:36:19.080 --> 0:36:22.000
<v Speaker 4>And this is actually an idea that dates back to

0:36:22.040 --> 0:36:25.680
<v Speaker 4>the nineteen nineties. Well proof carrying code, and you can

0:36:25.760 --> 0:36:31.239
<v Speaker 4>make computers that will check the proof of safety of

0:36:31.320 --> 0:36:32.880
<v Speaker 4>the algorithm before they run it.

0:36:34.120 --> 0:36:35.799
<v Speaker 3>And so.

0:36:37.640 --> 0:36:41.440
<v Speaker 4>If that becomes the standard for all hardware devices that

0:36:41.480 --> 0:36:46.600
<v Speaker 4>they only run checkable software objects that come with a

0:36:46.640 --> 0:36:51.120
<v Speaker 4>proof that it's okay to run this object, then I

0:36:51.120 --> 0:36:54.919
<v Speaker 4>think that would be a regime that would be extremely

0:36:54.960 --> 0:37:00.600
<v Speaker 4>hard to bypass because you'd have to have to create

0:37:00.640 --> 0:37:04.719
<v Speaker 4>a whole separate supply chain to produce you know, high

0:37:04.840 --> 0:37:08.200
<v Speaker 4>end chips. You know, talking about hundreds of billions of dollars,

0:37:08.680 --> 0:37:12.279
<v Speaker 4>tens of thousands of engineers decades of work that would

0:37:12.280 --> 0:37:15.080
<v Speaker 4>all have to happen, you know, in the black market.

0:37:14.840 --> 0:37:15.480
<v Speaker 1>So to speak.

0:37:17.880 --> 0:37:21.120
<v Speaker 4>So I do think there are these other stopping points

0:37:21.960 --> 0:37:25.040
<v Speaker 4>which are different from you know, just going right up

0:37:25.080 --> 0:37:26.960
<v Speaker 4>to the edge and hoping that no one.

0:37:27.640 --> 0:37:29.799
<v Speaker 3>Chooses to go off, to go off the edge.

0:37:30.880 --> 0:37:31.960
<v Speaker 1>It's to a Russell, thank you.

0:37:32.000 --> 0:37:33.120
<v Speaker 3>It's been a pleasure.

0:37:44.680 --> 0:37:46.239
<v Speaker 1>For text stuff I Musteloshian.

0:37:46.480 --> 0:37:49.320
<v Speaker 2>This episode was produced by Eliza Dennis and Minister Slaughter.

0:37:49.920 --> 0:37:52.799
<v Speaker 2>Executive produced by me Julian Nutter and Kate Osborne for

0:37:52.840 --> 0:37:57.200
<v Speaker 2>Kaleidoscope and Katrina Novel for iHeart podcasts Our Engineer Today

0:37:57.280 --> 0:37:58.200
<v Speaker 2>was by Hate Fraser.

0:37:58.480 --> 0:38:01.839
<v Speaker 1>Jack Insley mixed this episode. Kyle Murder wrote a theme song.