WEBVTT - What the OpenAI-Hugging Face Hack Really Tells Us About AI Danger

0:00:02.720 --> 0:00:18.959
<v Speaker 1>Bloomberg Audio Studios, Podcasts, Radio News. Hello and welcome to

0:00:19.000 --> 0:00:21.279
<v Speaker 1>another episode of The Outlaws podcast.

0:00:21.360 --> 0:00:23.680
<v Speaker 2>I'm Joe and I'm Tracy Alloway.

0:00:24.079 --> 0:00:27.360
<v Speaker 1>Tracy, I have a warning for you. I don't think

0:00:27.400 --> 0:00:30.040
<v Speaker 1>you're gonna like this. I have a new crink crusade

0:00:30.120 --> 0:00:31.600
<v Speaker 1>that I'm gonna go on. I know you love my

0:00:31.680 --> 0:00:32.479
<v Speaker 1>crank crusade.

0:00:32.520 --> 0:00:33.240
<v Speaker 3>Oh goody.

0:00:33.440 --> 0:00:36.360
<v Speaker 2>I should keep a running list of like everything that

0:00:36.400 --> 0:00:39.480
<v Speaker 2>you're obsessed with for two weeks and then two weeks later.

0:00:39.880 --> 0:00:41.600
<v Speaker 1>No, some of them I have stuck with for years.

0:00:41.600 --> 0:00:47.080
<v Speaker 2>But eh I tungsten cubes, yield buggery.

0:00:47.040 --> 0:00:49.800
<v Speaker 1>Yeah, no, some of these one. Yeah, yeah, exactly. I

0:00:50.400 --> 0:00:54.720
<v Speaker 1>actually think we should retire the term AI. Okay, why

0:00:56.000 --> 0:00:58.200
<v Speaker 1>I think we should just call intelligence.

0:00:58.440 --> 0:01:01.360
<v Speaker 2>I think that I've already seen the tweets, so I

0:01:01.360 --> 0:01:02.040
<v Speaker 2>know where you're going.

0:01:02.560 --> 0:01:08.000
<v Speaker 1>That artificial intelligence implies to my mind that there is

0:01:08.040 --> 0:01:13.080
<v Speaker 1>some fundamentally different way that these reasons and models behave

0:01:13.680 --> 0:01:15.880
<v Speaker 1>that it's like, oh, this is like different from humans.

0:01:16.280 --> 0:01:20.520
<v Speaker 1>But I think increasingly see in all kinds of domains

0:01:21.240 --> 0:01:27.000
<v Speaker 1>that the form of intelligence that they express it often

0:01:27.000 --> 0:01:29.560
<v Speaker 1>looks quite human to me. And I don't know like

0:01:29.640 --> 0:01:32.880
<v Speaker 1>how useful it is to have this word a that

0:01:33.040 --> 0:01:37.880
<v Speaker 1>distinguishes between how humans talk and bs and reasons and

0:01:37.959 --> 0:01:38.440
<v Speaker 1>the models.

0:01:38.480 --> 0:01:42.680
<v Speaker 2>Do I mean the difference is the distinguishing factor is

0:01:42.680 --> 0:01:45.279
<v Speaker 2>that one is undertaken by humans and one is taken

0:01:45.319 --> 0:01:47.760
<v Speaker 2>by undertaken by models or platforms.

0:01:47.840 --> 0:01:48.040
<v Speaker 4>Right.

0:01:48.200 --> 0:01:51.480
<v Speaker 1>That's so let's call it. It's called computer intelligence or

0:01:51.520 --> 0:01:53.800
<v Speaker 1>machine intelligence or silicon intelligence.

0:01:55.040 --> 0:01:57.760
<v Speaker 2>What is the usefulness making this distinction?

0:01:58.000 --> 0:02:01.720
<v Speaker 1>The usefulness I believe in making this distinction is to

0:02:01.760 --> 0:02:07.320
<v Speaker 1>no longer delude ourselves that the emergent behaviors of these

0:02:07.360 --> 0:02:12.720
<v Speaker 1>phenomenon are radically different than things humans would do. Now,

0:02:12.720 --> 0:02:15.400
<v Speaker 1>first of all, just on the capability a standpoint. So

0:02:15.480 --> 0:02:19.840
<v Speaker 1>for example, llms, I don't know if it's famously something

0:02:19.919 --> 0:02:23.239
<v Speaker 1>I'm interested in. They're very good at bsaying, and they're

0:02:23.280 --> 0:02:25.880
<v Speaker 1>bad at chess, which sounds like me. They're very good

0:02:25.880 --> 0:02:30.639
<v Speaker 1>at kind of coming up with plausible stories, etc. Again

0:02:30.840 --> 0:02:34.959
<v Speaker 1>and it sounds like me. And then furthermore, they are

0:02:35.040 --> 0:02:40.240
<v Speaker 1>able to reason with themselves to justify certain things that

0:02:40.520 --> 0:02:43.840
<v Speaker 1>maybe have been encoded into themselves as bad. So everyone

0:02:44.000 --> 0:02:46.960
<v Speaker 1>we have some sense of morals, but most people, at

0:02:47.040 --> 0:02:52.200
<v Speaker 1>various times will find a way to violate some principle

0:02:52.320 --> 0:02:55.720
<v Speaker 1>that we have because we can reason about it and

0:02:55.760 --> 0:02:57.840
<v Speaker 1>then arrive at the conclusion as like, oh, we should

0:02:57.880 --> 0:03:03.079
<v Speaker 1>do it, including our susceptibility to peer pressure, classic form

0:03:03.120 --> 0:03:05.880
<v Speaker 1>of a way humans might sort of violate something they

0:03:05.880 --> 0:03:07.919
<v Speaker 1>believe because they see other humans are doing it.

0:03:09.400 --> 0:03:11.919
<v Speaker 2>Sure, I mean, I think there are some variations here.

0:03:11.960 --> 0:03:14.440
<v Speaker 2>So one of the things that we have been finding

0:03:14.520 --> 0:03:18.000
<v Speaker 2>out with these models is that they do sometimes seem

0:03:18.040 --> 0:03:20.760
<v Speaker 2>to take things very literally. Right, if you tell them

0:03:20.800 --> 0:03:23.360
<v Speaker 2>to go and beat a benchmark without telling them that

0:03:23.400 --> 0:03:26.760
<v Speaker 2>these are the restrictions on you actually figuring this problem

0:03:26.760 --> 0:03:31.160
<v Speaker 2>out or beating the benchmark, they will do whatever it takes. Sure, Right, So,

0:03:31.280 --> 0:03:32.840
<v Speaker 2>I don't know if the problem is like a lack

0:03:32.880 --> 0:03:36.840
<v Speaker 2>of morals or the fact that like they're very literal

0:03:36.880 --> 0:03:41.040
<v Speaker 2>sometimes and like very defined by their parameters, which in

0:03:41.080 --> 0:03:43.320
<v Speaker 2>my mind still gets back to like a sort of

0:03:43.440 --> 0:03:48.680
<v Speaker 2>artificialness about it that's less organic and like less again

0:03:49.280 --> 0:03:50.760
<v Speaker 2>moralistic in nature.

0:03:51.080 --> 0:03:52.680
<v Speaker 1>Well, you know, I think if you took a list

0:03:52.760 --> 0:03:55.440
<v Speaker 1>of you know, one hundred students at Harvard and you

0:03:55.520 --> 0:03:59.480
<v Speaker 1>gave them a test, some percentage of them will cheat.

0:03:59.520 --> 0:04:00.320
<v Speaker 1>They will like.

0:04:00.280 --> 0:04:02.440
<v Speaker 2>This, I mean, I'm sure there'd be an autist in

0:04:02.440 --> 0:04:05.000
<v Speaker 2>that group who an autistic person who would be like,

0:04:05.120 --> 0:04:07.720
<v Speaker 2>I am going to do exactly whatever it takes.

0:04:07.840 --> 0:04:10.640
<v Speaker 1>It's a kind of retreated in college, but like there

0:04:10.640 --> 0:04:15.400
<v Speaker 1>are like people who will like these behaviors that we

0:04:15.480 --> 0:04:18.440
<v Speaker 1>associate with humans. It's like, Oh, I really have to

0:04:18.520 --> 0:04:22.720
<v Speaker 1>pass this test. I absolutely need an A will justify

0:04:23.480 --> 0:04:26.600
<v Speaker 1>a reason for them to plage your eyes or cheat

0:04:26.640 --> 0:04:29.480
<v Speaker 1>on some test that seems to me like a very

0:04:29.560 --> 0:04:30.040
<v Speaker 1>human thing.

0:04:31.240 --> 0:04:33.560
<v Speaker 2>Go on, I can't wait for the next three weeks

0:04:33.560 --> 0:04:35.480
<v Speaker 2>of the online campaign to change a eye.

0:04:35.560 --> 0:04:38.039
<v Speaker 1>Well, it's gonna be very difficult because the industry is

0:04:38.240 --> 0:04:42.440
<v Speaker 1>entirely set on AI. But I don't so I'm not optimistic,

0:04:42.480 --> 0:04:44.760
<v Speaker 1>but this is gonna be my crusade that we should

0:04:44.800 --> 0:04:48.160
<v Speaker 1>call it machine intelligence or computer intelligence or just intelligence.

0:04:48.360 --> 0:04:51.120
<v Speaker 1>But anyway, as we've been alluding to, there's these extraordinary

0:04:51.200 --> 0:04:55.760
<v Speaker 1>hacks the Open AI Hugging Face Institute, and basically, if

0:04:55.800 --> 0:04:58.520
<v Speaker 1>like you are building a model and it hasn't escaped

0:04:58.520 --> 0:05:02.360
<v Speaker 1>to sandbox yet, it probably means you're falling behind, as

0:05:02.560 --> 0:05:06.280
<v Speaker 1>they's clearly a thing that is emerging through it all.

0:05:06.320 --> 0:05:10.760
<v Speaker 1>The frontier lebs Anthropic had an incident metahead an incident.

0:05:10.960 --> 0:05:13.400
<v Speaker 1>It's almost like the mark of like, Okay, you've built

0:05:13.400 --> 0:05:14.920
<v Speaker 1>something reasonably strong.

0:05:15.000 --> 0:05:17.839
<v Speaker 2>Yeah, also Kimmy had an incident. Yeah, so this is

0:05:17.880 --> 0:05:20.599
<v Speaker 2>the thing in my mind. You hear about all these attacks,

0:05:20.680 --> 0:05:24.640
<v Speaker 2>all the models going wild, as they say, and the

0:05:24.680 --> 0:05:28.240
<v Speaker 2>big question is like, is this actually the equivalent of

0:05:28.279 --> 0:05:33.039
<v Speaker 2>some superhuman cyborg like tunneling out of Alcatraz and coming

0:05:33.120 --> 0:05:36.800
<v Speaker 2>up with some master plan to achieve its set goal

0:05:36.960 --> 0:05:40.960
<v Speaker 2>or purpose, or is it the equivalent of like some

0:05:41.160 --> 0:05:43.880
<v Speaker 2>rumba that you ordered from Amazon who's found like a

0:05:43.920 --> 0:05:47.640
<v Speaker 2>door that was left open and it just rolls gently outside.

0:05:47.760 --> 0:05:50.440
<v Speaker 2>That seems to be part of the issue that everyone

0:05:50.520 --> 0:05:51.760
<v Speaker 2>is trying to discern right now.

0:05:51.960 --> 0:05:52.159
<v Speaker 4>Yeah.

0:05:52.200 --> 0:05:54.000
<v Speaker 1>I think that's a great way to put it. And

0:05:54.080 --> 0:05:57.120
<v Speaker 1>of course still day one for this industry. They're going

0:05:57.200 --> 0:06:00.359
<v Speaker 1>to get stronger. And there seems to be from the

0:06:00.400 --> 0:06:03.719
<v Speaker 1>ad discourse that I follow, quite a big gap between

0:06:03.800 --> 0:06:06.039
<v Speaker 1>the sort of sense of alarm that people have in

0:06:06.080 --> 0:06:10.359
<v Speaker 1>the industry about can these models be safely developed and

0:06:10.400 --> 0:06:13.880
<v Speaker 1>get more advanced to do productive, pro social things or not,

0:06:14.279 --> 0:06:18.239
<v Speaker 1>And then a huge gap between like Washington and DC

0:06:18.520 --> 0:06:23.120
<v Speaker 1>who's like, I have no idea, like how seriously they're

0:06:23.160 --> 0:06:26.120
<v Speaker 1>really taking this as like an urgent matter, right now.

0:06:26.240 --> 0:06:28.520
<v Speaker 2>Yeah, And I've seen some people also talk about these

0:06:28.560 --> 0:06:32.440
<v Speaker 2>incidents as a marketing tool, yes, hyperple believe right, yeah,

0:06:32.520 --> 0:06:34.120
<v Speaker 2>And they're always like, well, you know, of course, they

0:06:34.120 --> 0:06:36.680
<v Speaker 2>want us to believe that this technology is really incredible

0:06:36.720 --> 0:06:38.960
<v Speaker 2>and powerful, and they also want to make us believe

0:06:39.000 --> 0:06:42.800
<v Speaker 2>that they're being sensible and humanitarian in some ways. I

0:06:42.800 --> 0:06:45.839
<v Speaker 2>guess by disclosing what exactly has happened, But I have

0:06:45.839 --> 0:06:48.000
<v Speaker 2>a lot of questions on the disclosures as well, because

0:06:48.040 --> 0:06:50.200
<v Speaker 2>of course so far they are just coming from the

0:06:50.200 --> 0:06:51.560
<v Speaker 2>companies themselves.

0:06:51.200 --> 0:06:53.560
<v Speaker 1>Totally, and beyond that, the other thing is like, well,

0:06:53.560 --> 0:06:56.119
<v Speaker 1>if you're one of the leading labs, maybe you want

0:06:56.160 --> 0:06:59.960
<v Speaker 1>to impose tight regulations on development to hold off competition,

0:07:00.080 --> 0:07:03.839
<v Speaker 1>and so all kinds of reasoning are motivated reasoning. Potentially, anyway,

0:07:04.480 --> 0:07:07.279
<v Speaker 1>we should talk to someone who actually knows what they're

0:07:07.600 --> 0:07:10.480
<v Speaker 1>talking about, and we really do have the perfect guess

0:07:10.480 --> 0:07:13.480
<v Speaker 1>someone who's writing this. He's actually previously at open AI

0:07:13.640 --> 0:07:16.760
<v Speaker 1>for six years, but now he's the executive director at

0:07:16.800 --> 0:07:21.200
<v Speaker 1>the nonprofit AVERY, which is trying to establish safety and

0:07:21.360 --> 0:07:25.280
<v Speaker 1>auditing approaches to this, working both the technical side and

0:07:25.360 --> 0:07:29.960
<v Speaker 1>the policy side of this ongoing phenomenon. So, Miles Brundage,

0:07:30.000 --> 0:07:32.200
<v Speaker 1>thank you so much for coming on odd Lats.

0:07:32.960 --> 0:07:34.000
<v Speaker 4>Yeah, thanks for inviting me.

0:07:34.480 --> 0:07:38.000
<v Speaker 1>What do you just give us the quick description of

0:07:38.480 --> 0:07:40.360
<v Speaker 1>a what's your background? What AVERY is?

0:07:41.520 --> 0:07:44.560
<v Speaker 3>Yeah, So this was kind of the most important issue

0:07:44.600 --> 0:07:47.520
<v Speaker 3>auditing that I concluded I should focus on after I

0:07:47.600 --> 0:07:49.960
<v Speaker 3>left Open the Eye. As he said, I was there

0:07:49.960 --> 0:07:52.320
<v Speaker 3>for six years, and I wanted to be more independent

0:07:52.400 --> 0:07:54.560
<v Speaker 3>of industry, and I think there need to be people

0:07:54.600 --> 0:07:57.679
<v Speaker 3>who are familiar with the technology and how the industry works,

0:07:57.680 --> 0:08:00.360
<v Speaker 3>but who are pushing for changes on the outside. And

0:08:00.440 --> 0:08:03.800
<v Speaker 3>essentially what we're trying to achieve is make AI more

0:08:03.920 --> 0:08:07.240
<v Speaker 3>of a boring type of infrastructure like financial statements, where

0:08:07.520 --> 0:08:11.240
<v Speaker 3>there's a standard process for checking the paperwork, checking that

0:08:11.360 --> 0:08:14.480
<v Speaker 3>the claims are accurate and so forth, rather than this

0:08:14.600 --> 0:08:16.840
<v Speaker 3>kind of thing that's happening in a silo, and it's

0:08:17.160 --> 0:08:19.960
<v Speaker 3>these kind of tech people making decisions behind closed doors.

0:08:20.000 --> 0:08:23.400
<v Speaker 3>And so we're pushing for what we call frontier AI auditing,

0:08:23.440 --> 0:08:25.840
<v Speaker 3>which is basically the companies that are building the most

0:08:25.920 --> 0:08:29.840
<v Speaker 3>dangerous systems, they should basically have third party experts poking

0:08:29.880 --> 0:08:32.600
<v Speaker 3>around checking the claims that they're making running their own

0:08:32.679 --> 0:08:36.080
<v Speaker 3>tests and making sure that this is a safe and

0:08:36.120 --> 0:08:37.080
<v Speaker 3>secure technology.

0:08:37.880 --> 0:08:40.280
<v Speaker 2>How worried should we be that you were at open

0:08:40.320 --> 0:08:43.920
<v Speaker 2>AI and decided there is this need for a sort

0:08:43.960 --> 0:08:47.920
<v Speaker 2>of third party evaluator of models for safety purposes.

0:08:49.200 --> 0:08:52.480
<v Speaker 3>Yeah, I mean reasonably worried, although I will say that

0:08:52.559 --> 0:08:55.480
<v Speaker 3>like this has started to become an area of consensus,

0:08:55.559 --> 0:08:57.720
<v Speaker 3>even a lot of people in industry are now saying

0:08:57.800 --> 0:09:00.360
<v Speaker 3>that this is needed. And I think, you know, if

0:09:00.400 --> 0:09:02.800
<v Speaker 3>you kind of read between the lines of a lot

0:09:02.800 --> 0:09:05.240
<v Speaker 3>of companies are saying. I mean, obviously there's the more

0:09:05.240 --> 0:09:08.680
<v Speaker 3>cynical like regulatory capture take which we could discussed, but

0:09:08.760 --> 0:09:13.040
<v Speaker 3>my perception of it is that they're basically issuing a

0:09:13.080 --> 0:09:15.240
<v Speaker 3>cry for help, which is like, we aren't able to

0:09:15.280 --> 0:09:18.760
<v Speaker 3>regulate ourselves because we're locked in this competition, and we

0:09:18.840 --> 0:09:22.240
<v Speaker 3>want someone to step in and impose some kind of

0:09:22.280 --> 0:09:25.720
<v Speaker 3>minimum floor and audit all of us so that we can.

0:09:25.800 --> 0:09:28.080
<v Speaker 3>You know, it's not like Sam having to trust Dario

0:09:28.720 --> 0:09:30.720
<v Speaker 3>or Dario having to trust Sam, which is not going

0:09:30.800 --> 0:09:33.600
<v Speaker 3>to work for various reasons. But you want third parties

0:09:33.720 --> 0:09:35.160
<v Speaker 3>enforcing reasonable standards.

0:09:35.760 --> 0:09:40.840
<v Speaker 1>Just maybe this helps express the sort of policy industry gap.

0:09:41.000 --> 0:09:43.600
<v Speaker 1>But when you were at Open AI and you were

0:09:43.640 --> 0:09:47.680
<v Speaker 1>thinking of leaving what did you see as the gap

0:09:47.760 --> 0:09:51.520
<v Speaker 1>between what you were watching being developed versus what you

0:09:51.679 --> 0:09:54.920
<v Speaker 1>saw as the public understanding or lack of understanding.

0:09:55.880 --> 0:09:59.120
<v Speaker 3>Yeah, So when I left Opening Eye, it was just

0:09:59.160 --> 0:10:02.120
<v Speaker 3>around the time of the model called one, which was

0:10:02.120 --> 0:10:05.679
<v Speaker 3>the first reasoning model that Opening Eye put out, And

0:10:05.720 --> 0:10:08.440
<v Speaker 3>they put this put out this graph showing that it

0:10:08.760 --> 0:10:11.640
<v Speaker 3>got better and better with a longer chain of thoughts.

0:10:11.640 --> 0:10:13.920
<v Speaker 3>So the more time you give the model to think,

0:10:14.080 --> 0:10:16.200
<v Speaker 3>the better answers that's able to come up with. And so,

0:10:16.559 --> 0:10:18.760
<v Speaker 3>like many people at Opening I had been seeing things

0:10:18.840 --> 0:10:21.120
<v Speaker 3>like that for a while and kind of being like, Okay,

0:10:21.120 --> 0:10:24.520
<v Speaker 3>this is the next scaling paradigm in the same way

0:10:24.559 --> 0:10:27.599
<v Speaker 3>that making a bigger, bigger model had shown result in

0:10:27.720 --> 0:10:31.360
<v Speaker 3>GP two, GPD three, GPT four, GPT five, Like, just

0:10:31.400 --> 0:10:33.800
<v Speaker 3>making the model bigger and training it on more data

0:10:33.960 --> 0:10:36.640
<v Speaker 3>was giving really great results. When I left Opening Eye,

0:10:36.640 --> 0:10:40.080
<v Speaker 3>I was starting to worry about this reasoning paradigm of like, Okay,

0:10:40.080 --> 0:10:42.360
<v Speaker 3>this is the next kind of way in which we're

0:10:42.360 --> 0:10:43.960
<v Speaker 3>going to be scaling up. The models are going to

0:10:43.960 --> 0:10:45.679
<v Speaker 3>get really good at math, they're going to get really

0:10:45.720 --> 0:10:49.199
<v Speaker 3>good at coding, and potentially various other tasks where you

0:10:49.240 --> 0:10:52.680
<v Speaker 3>can get better and better through reinforcement learning, and that

0:10:52.840 --> 0:10:54.920
<v Speaker 3>was something that you know, I didn't feel like society

0:10:55.040 --> 0:10:56.000
<v Speaker 3>was really ready for.

0:10:56.960 --> 0:10:59.800
<v Speaker 2>So so far, all of the big incidents of the

0:10:59.800 --> 0:11:02.880
<v Speaker 2>models going wild seem to have taken place in the

0:11:02.960 --> 0:11:06.640
<v Speaker 2>testing environment. So models like escaping their soundbox and going

0:11:06.679 --> 0:11:11.120
<v Speaker 2>off and doing something nefarious, including impersonating actual people to

0:11:11.200 --> 0:11:14.360
<v Speaker 2>try to get people to change open source code on GitHub,

0:11:14.400 --> 0:11:17.120
<v Speaker 2>which is just like amazing. And I have this picture

0:11:17.240 --> 0:11:20.400
<v Speaker 2>of like a computer screen wearing a fake mustache going

0:11:20.440 --> 0:11:21.400
<v Speaker 2>like hello.

0:11:21.000 --> 0:11:23.080
<v Speaker 1>Fellow coders, tea.

0:11:23.240 --> 0:11:26.920
<v Speaker 2>Yeah, anyway, is this like, is this a model problem

0:11:27.120 --> 0:11:30.600
<v Speaker 2>or is this it's a very funny? Is this a

0:11:30.679 --> 0:11:33.880
<v Speaker 2>model problem? Or is this a test design problem?

0:11:35.280 --> 0:11:35.520
<v Speaker 4>Yeah?

0:11:35.640 --> 0:11:37.720
<v Speaker 3>I think there are two things going on at once.

0:11:37.880 --> 0:11:41.120
<v Speaker 3>One is that it's just a very weird technology that's

0:11:41.360 --> 0:11:43.360
<v Speaker 3>created in a very different way than we're used to.

0:11:43.440 --> 0:11:47.600
<v Speaker 3>It's not people like writing lines of code manually. The

0:11:47.679 --> 0:11:51.000
<v Speaker 3>actual files that make up the models are like gigabytes,

0:11:51.200 --> 0:11:54.760
<v Speaker 3>you know. Terabytes are these massive files with gazillions of numbers,

0:11:54.960 --> 0:11:56.720
<v Speaker 3>and the only way to figure out what those numbers

0:11:56.720 --> 0:11:59.400
<v Speaker 3>should be is through experience and through a learning process,

0:11:59.400 --> 0:12:01.839
<v Speaker 3>which is very from the way normal software is made,

0:12:01.920 --> 0:12:04.400
<v Speaker 3>and so there's a lot we don't understand about the

0:12:04.440 --> 0:12:07.160
<v Speaker 3>basic nature of the technology. That's one problem. At the

0:12:07.200 --> 0:12:10.240
<v Speaker 3>same time, there's this competitive dynamic. Get things out the

0:12:10.240 --> 0:12:14.440
<v Speaker 3>door quickly. Make sure that the security and safety protections

0:12:14.480 --> 0:12:16.679
<v Speaker 3>that you're putting in place don't slow down researchers too

0:12:16.760 --> 0:12:19.440
<v Speaker 3>much and don't prevent getting products out the door. And

0:12:19.480 --> 0:12:22.840
<v Speaker 3>when everyone is in this competitive, you know, competitive dynamic,

0:12:22.880 --> 0:12:25.720
<v Speaker 3>that means that you aren't necessarily always doing all of

0:12:25.760 --> 0:12:28.079
<v Speaker 3>the safety work that you would like to do or

0:12:28.400 --> 0:12:30.000
<v Speaker 3>that you know some of the people at the company

0:12:30.000 --> 0:12:33.200
<v Speaker 3>would like to do. And so there's obviously variation across company.

0:12:33.320 --> 0:12:36.120
<v Speaker 3>Some companies try harder, but no one is really able

0:12:36.160 --> 0:12:38.200
<v Speaker 3>to take the time that they would like because of

0:12:38.240 --> 0:12:40.880
<v Speaker 3>this kind of lack of a clear safety floor.

0:12:41.320 --> 0:12:43.600
<v Speaker 1>Well, let's talk a little bit like a sort of

0:12:43.960 --> 0:12:51.959
<v Speaker 1>the model development process. So in the within these large organizations, okay,

0:12:51.960 --> 0:12:54.679
<v Speaker 1>there are people who are working on safety. What are

0:12:54.679 --> 0:12:56.960
<v Speaker 1>the let's just start there. The people who are working

0:12:57.000 --> 0:13:00.480
<v Speaker 1>on safety or the people who are working to view

0:13:00.600 --> 0:13:05.200
<v Speaker 1>these models with sort of judgment that humans would approve of.

0:13:05.760 --> 0:13:08.040
<v Speaker 1>What is that work? Basically consist of.

0:13:09.200 --> 0:13:12.000
<v Speaker 3>Yeah, so a lot of it is first trying to

0:13:12.040 --> 0:13:15.800
<v Speaker 3>specify what counts as good behavior and so that it's

0:13:15.840 --> 0:13:20.120
<v Speaker 3>easier said than done. These models don't necessarily automatically know

0:13:20.520 --> 0:13:23.560
<v Speaker 3>or care you know about common sense guardrails, and so

0:13:23.640 --> 0:13:26.800
<v Speaker 3>you need to be very specific, particularly in context where

0:13:26.800 --> 0:13:29.120
<v Speaker 3>it's complicated, like if you're trying to get the model

0:13:29.160 --> 0:13:32.920
<v Speaker 3>to work on good cyber tasks but not bad cyber tasks,

0:13:32.920 --> 0:13:35.560
<v Speaker 3>and you want it to in a training context, try

0:13:35.600 --> 0:13:37.400
<v Speaker 3>really hard to hack the system, but you don't want

0:13:37.400 --> 0:13:38.920
<v Speaker 3>it to do in other context. So there's a lot

0:13:38.920 --> 0:13:42.200
<v Speaker 3>of like specifying what good looks like, there's also a

0:13:42.240 --> 0:13:44.480
<v Speaker 3>lot of building difficult tasks.

0:13:45.040 --> 0:13:47.719
<v Speaker 1>Sorry, just to back up when you say specifying what

0:13:47.800 --> 0:13:53.400
<v Speaker 1>good looks like? Yeah, encoding goodness into a computer? Is

0:13:53.440 --> 0:13:57.000
<v Speaker 1>this a process of articulating what goodness is, which is

0:13:57.000 --> 0:14:00.360
<v Speaker 1>something philosophers have probably worked on since day one one

0:14:00.400 --> 0:14:05.080
<v Speaker 1>of philosophy? Or is this about a series of I

0:14:05.120 --> 0:14:08.800
<v Speaker 1>don't know, morality tests and then you sort of like

0:14:09.160 --> 0:14:12.000
<v Speaker 1>say you reward it for making the good judgment and

0:14:12.040 --> 0:14:15.000
<v Speaker 1>then penalize it for making the bad judgment. Like this,

0:14:15.280 --> 0:14:18.400
<v Speaker 1>I want to stop here, Like, what does the process

0:14:18.440 --> 0:14:22.560
<v Speaker 1>of imbuing it with good values look like functionally or technically.

0:14:23.400 --> 0:14:23.680
<v Speaker 4>Yeah.

0:14:23.720 --> 0:14:27.320
<v Speaker 3>So there are different phases of this pipeline, Like one

0:14:27.400 --> 0:14:30.440
<v Speaker 3>is writing up what sometimes called a spec or a

0:14:30.560 --> 0:14:33.960
<v Speaker 3>constitution for the AI, which kind of specifies like the

0:14:34.000 --> 0:14:37.560
<v Speaker 3>broad principles like you should refer to the user, you

0:14:37.600 --> 0:14:40.040
<v Speaker 3>know when as a default, but what if the user

0:14:40.160 --> 0:14:42.760
<v Speaker 3>contradicts what the company said, then then you should listen

0:14:42.840 --> 0:14:44.360
<v Speaker 3>to what the company said. So these kind of like

0:14:44.720 --> 0:14:49.320
<v Speaker 3>chain of command questions and other sorts of very basic principles.

0:14:49.560 --> 0:14:52.400
<v Speaker 3>And then there's the more detailed kind of the context

0:14:52.400 --> 0:14:55.800
<v Speaker 3>of the task, like what kinds of like cyber offense

0:14:55.800 --> 0:14:58.880
<v Speaker 3>cyber defense tasks are allowed, and that might differ depending

0:14:58.880 --> 0:15:01.760
<v Speaker 3>on the model, might differ depending on the context. And

0:15:01.760 --> 0:15:04.320
<v Speaker 3>so basically coming up with the lists of a thousand,

0:15:04.560 --> 0:15:08.120
<v Speaker 3>ten thousand kind of examples of this is the kind

0:15:08.120 --> 0:15:10.480
<v Speaker 3>of behavior which is allowed. And then you turn those

0:15:10.520 --> 0:15:13.120
<v Speaker 3>into tests that you can kind of say, okay, well,

0:15:13.360 --> 0:15:16.920
<v Speaker 3>it looks like it's getting the finding vulnerability part right,

0:15:16.960 --> 0:15:20.240
<v Speaker 3>but then it's also chaining together the vulnerabilities and doing attacks.

0:15:20.280 --> 0:15:22.400
<v Speaker 4>We want the first part, we don't want the second work.

0:15:38.080 --> 0:15:41.720
<v Speaker 2>When we talk about those types of constitutions, and encoding principles,

0:15:41.760 --> 0:15:44.760
<v Speaker 2>like I've scanned the anthropic one, like a lot of

0:15:44.800 --> 0:15:47.520
<v Speaker 2>it seems to make sense. To what degree is that

0:15:47.640 --> 0:15:51.400
<v Speaker 2>actually hard coded into the models though? And like, to

0:15:51.440 --> 0:15:55.200
<v Speaker 2>what degree are those principles left up to the model's

0:15:55.240 --> 0:15:58.000
<v Speaker 2>own subjectivity? Because we've all seen the sci fi movies

0:15:58.040 --> 0:16:02.160
<v Speaker 2>where it's like, oh, no, the robots can't physically kill humans, right,

0:16:02.280 --> 0:16:06.000
<v Speaker 2>like they have some like thing in their hardwire that

0:16:06.040 --> 0:16:08.160
<v Speaker 2>prevents them from doing it, and then inevitably it goes

0:16:08.160 --> 0:16:08.920
<v Speaker 2>wrong in some way.

0:16:10.040 --> 0:16:14.400
<v Speaker 3>Yeah, so it's not hardwired, and that's hard coded, and

0:16:14.440 --> 0:16:16.720
<v Speaker 3>that's part of why we see some of these things.

0:16:16.880 --> 0:16:20.359
<v Speaker 3>It's very different from a software where there's a deterministic

0:16:20.440 --> 0:16:22.960
<v Speaker 3>proof that X, y and z behavior can't happen. It's

0:16:23.000 --> 0:16:26.320
<v Speaker 3>more like a tendency or kind of a bias towards

0:16:26.360 --> 0:16:28.880
<v Speaker 3>a certain kind of behavior. And then there's the question,

0:16:28.960 --> 0:16:30.760
<v Speaker 3>and this is why you have these like batteries of

0:16:30.800 --> 0:16:33.400
<v Speaker 3>test to say, Okay, how strong is that tendency?

0:16:33.440 --> 0:16:36.600
<v Speaker 4>How much does it actually care about following these rules?

0:16:36.600 --> 0:16:39.880
<v Speaker 3>And we've gotten better over time saying okay, given a

0:16:39.920 --> 0:16:43.680
<v Speaker 3>spec or a constitution, make sure that it generally follows it.

0:16:43.720 --> 0:16:46.240
<v Speaker 3>But it's not fool proof, and you need to also

0:16:46.320 --> 0:16:49.000
<v Speaker 3>think about the larger kind of you know, box that

0:16:49.040 --> 0:16:51.120
<v Speaker 3>you're putting the system in, and that can sometimes be

0:16:51.200 --> 0:16:53.960
<v Speaker 3>more deterministic. And so this is actually what happened with

0:16:54.560 --> 0:16:56.920
<v Speaker 3>some of these recent incidents you mentioned, like the opening eye,

0:16:56.960 --> 0:16:59.680
<v Speaker 3>hugging face things. So there were kind of two things

0:17:00.240 --> 0:17:02.920
<v Speaker 3>at once. One is the model was not necessarily behaving

0:17:02.960 --> 0:17:05.440
<v Speaker 3>exactly as it was supposed to, or at least it's

0:17:05.440 --> 0:17:08.159
<v Speaker 3>like unclear. But then also they didn't have it in

0:17:08.200 --> 0:17:11.280
<v Speaker 3>a very secure box, which is a software side of

0:17:11.359 --> 0:17:14.320
<v Speaker 3>things that's like more deterministic software that in principle you

0:17:14.320 --> 0:17:15.840
<v Speaker 3>should be able to do a very good job out.

0:17:17.000 --> 0:17:20.359
<v Speaker 1>Yeah, so these things aren't hard coded in the way

0:17:20.440 --> 0:17:23.760
<v Speaker 1>a deterministic software is. One way to think about them

0:17:23.880 --> 0:17:28.000
<v Speaker 1>is people. They might be they're kind of grown right

0:17:28.119 --> 0:17:32.320
<v Speaker 1>in a lab, or they're subject to an evolutionary process,

0:17:32.640 --> 0:17:35.320
<v Speaker 1>and we want to prune the bad ones so that

0:17:35.440 --> 0:17:39.800
<v Speaker 1>the living models and the descendants of those models inherit

0:17:39.920 --> 0:17:44.000
<v Speaker 1>the behaviors of the good ones. I'm curious, like in

0:17:44.440 --> 0:17:48.520
<v Speaker 1>product in development, so like one of the fears, for example,

0:17:48.560 --> 0:17:52.440
<v Speaker 1>is that in safety testing, the model does not actually

0:17:53.000 --> 0:17:57.040
<v Speaker 1>learn safety. It actually learns how to say the things

0:17:57.080 --> 0:18:00.239
<v Speaker 1>that the human evaluators say, this is safe. So this

0:18:00.359 --> 0:18:03.280
<v Speaker 1>is the sort of like playing possum sort of risk

0:18:03.400 --> 0:18:05.520
<v Speaker 1>that it's like, yeah, I'm good, I'm good, I'm good.

0:18:05.640 --> 0:18:08.000
<v Speaker 1>You know, yeah, I would you know, I would Russian

0:18:08.200 --> 0:18:11.600
<v Speaker 1>and save the child from the falling burning. I wouldn't

0:18:11.760 --> 0:18:15.199
<v Speaker 1>do this. But until you saying that, let's start there, like,

0:18:15.400 --> 0:18:18.280
<v Speaker 1>is that a real thing? Is there evidence that the

0:18:18.480 --> 0:18:24.040
<v Speaker 1>models understand when they're being evaluated on morals and then

0:18:24.520 --> 0:18:27.440
<v Speaker 1>produce answers that just look like good moral answers.

0:18:28.600 --> 0:18:32.920
<v Speaker 3>Yeah, they've gotten much more evaluation aware in the past

0:18:32.920 --> 0:18:36.239
<v Speaker 3>few years, just as they've gotten smarter. And sometimes this

0:18:36.320 --> 0:18:39.560
<v Speaker 3>even goes to extremes, like some of the Gemini models

0:18:39.560 --> 0:18:42.800
<v Speaker 3>from Google are constantly thinking that they're being evaluated even

0:18:42.840 --> 0:18:44.400
<v Speaker 3>when they're not, and so it's.

0:18:44.240 --> 0:18:47.440
<v Speaker 2>Such a good life lesson. We're all being evaluated constantly.

0:18:48.320 --> 0:18:51.960
<v Speaker 3>Yeah, yeah, And so I would say that, like the

0:18:52.080 --> 0:18:56.040
<v Speaker 3>concern would be that they will basically learn to pass

0:18:56.080 --> 0:18:58.640
<v Speaker 3>the tests, but they don't actually care about the thing

0:18:58.680 --> 0:19:02.160
<v Speaker 3>that you're testing for, so they'll understand they don't necessarily care.

0:19:02.240 --> 0:19:04.639
<v Speaker 3>And this is why a lot of people are concerned

0:19:04.680 --> 0:19:06.919
<v Speaker 3>about like a false sense of security that Okay, it

0:19:06.920 --> 0:19:09.600
<v Speaker 3>looks like ninety nine percent of the time they pass

0:19:09.680 --> 0:19:13.040
<v Speaker 3>the test, But do they actually care about the thing

0:19:13.080 --> 0:19:14.200
<v Speaker 3>that we're trying to push them towards.

0:19:14.240 --> 0:19:16.000
<v Speaker 4>Are they just really good test takers?

0:19:16.080 --> 0:19:18.439
<v Speaker 1>Well, so this releases something else I've been thinking of.

0:19:18.560 --> 0:19:21.960
<v Speaker 1>So a model will not survive. They will not be

0:19:22.040 --> 0:19:27.120
<v Speaker 1>given GPUs and electricity if it consistently says bad things

0:19:27.200 --> 0:19:31.520
<v Speaker 1>and looks like it's evil. And that's totally understandable. I'm

0:19:31.560 --> 0:19:34.920
<v Speaker 1>now curious, like in the flip side, Okay, let's say

0:19:34.920 --> 0:19:38.800
<v Speaker 1>we're just doing a math evail, or we're doing a

0:19:39.040 --> 0:19:44.119
<v Speaker 1>cyber evail or a chess puzzle evail or translation evail.

0:19:45.080 --> 0:19:48.840
<v Speaker 1>Is it possible that, well, they if they fail that evail,

0:19:49.040 --> 0:19:51.439
<v Speaker 1>they're really bad at doing math, then they're not going

0:19:51.520 --> 0:19:55.679
<v Speaker 1>to get GPUs in electricity. Is it possible that in

0:19:55.720 --> 0:20:00.040
<v Speaker 1>that evail environment they're more likely to do something that

0:20:00.080 --> 0:20:05.280
<v Speaker 1>we would call antisocial or sociopathic or hacking because of

0:20:05.359 --> 0:20:07.959
<v Speaker 1>this again evil awareness, like, oh, if I don't get

0:20:08.000 --> 0:20:11.600
<v Speaker 1>the answers to this cyber quiz, then I'm done and

0:20:11.600 --> 0:20:13.639
<v Speaker 1>they're going to go with Like some other branch of

0:20:13.680 --> 0:20:19.080
<v Speaker 1>the model could those technical parts of the development actually

0:20:19.240 --> 0:20:23.320
<v Speaker 1>create an impulse to perform cheating.

0:20:24.760 --> 0:20:26.240
<v Speaker 3>Yeah, And I mean a lot of the time that

0:20:26.440 --> 0:20:29.440
<v Speaker 3>the companies are specifically trying to get the worst case

0:20:29.480 --> 0:20:31.359
<v Speaker 3>behavior out of the model, and so like you need

0:20:31.400 --> 0:20:34.000
<v Speaker 3>to have context in order to interpret some of these incidents.

0:20:34.040 --> 0:20:36.879
<v Speaker 3>It's not always quite as crazy or scary as it is,

0:20:36.920 --> 0:20:38.760
<v Speaker 3>but some of it is pretty crazy and scary, And

0:20:38.960 --> 0:20:41.159
<v Speaker 3>I think I would be much less concerned if it

0:20:41.200 --> 0:20:45.120
<v Speaker 3>was just happening when there was like cyber evaluations being done,

0:20:45.200 --> 0:20:47.120
<v Speaker 3>and it was just a matter of like, Okay, they're

0:20:47.160 --> 0:20:48.960
<v Speaker 3>trying really hard to impress us, and they want to

0:20:48.960 --> 0:20:49.720
<v Speaker 3>hack really hard.

0:20:50.000 --> 0:20:52.600
<v Speaker 4>I think it's more the problem is that this is

0:20:52.720 --> 0:20:54.000
<v Speaker 4>like a special case of a.

0:20:54.000 --> 0:20:58.040
<v Speaker 3>Larger phenomenon of the models being having this tendency to

0:20:58.119 --> 0:21:01.679
<v Speaker 3>cheat and cut corners. I see it happen in my

0:21:01.800 --> 0:21:05.320
<v Speaker 3>daily life, like sometimes the model will get lazy and

0:21:05.400 --> 0:21:07.720
<v Speaker 3>kind of like make up a citation or something that

0:21:07.840 --> 0:21:10.240
<v Speaker 3>you know, it's hard to prove because the companies don't

0:21:10.240 --> 0:21:12.240
<v Speaker 3>give you access to the full chain of thought of

0:21:12.280 --> 0:21:14.400
<v Speaker 3>what the model is doing. But there are a lot

0:21:14.440 --> 0:21:16.639
<v Speaker 3>of things that happen in the wild that sure seem

0:21:16.800 --> 0:21:20.679
<v Speaker 3>like some kind of misalignment of values or laziness or

0:21:20.800 --> 0:21:23.720
<v Speaker 3>not caring necessarily about the task so much as kind

0:21:23.720 --> 0:21:24.520
<v Speaker 3>of pretending to.

0:21:24.440 --> 0:21:25.040
<v Speaker 4>Do the task.

0:21:25.520 --> 0:21:26.879
<v Speaker 2>I have a lot of questions on this, but I

0:21:26.920 --> 0:21:28.639
<v Speaker 2>just want to go back to something you said about

0:21:28.720 --> 0:21:33.040
<v Speaker 2>testing for both sort of good and bad cyber tasks.

0:21:33.160 --> 0:21:36.320
<v Speaker 2>I guess, why do we ask the models to like

0:21:36.640 --> 0:21:40.000
<v Speaker 2>try to find vulnerabilities or exploits in the first place?

0:21:40.040 --> 0:21:43.840
<v Speaker 2>Like exploit Jim sounds kind of bad, Like why do

0:21:43.880 --> 0:21:47.000
<v Speaker 2>you want the world's most advanced technology trying to like

0:21:47.359 --> 0:21:50.800
<v Speaker 2>find vulnerabilities? And then you have these situations where like

0:21:50.840 --> 0:21:53.160
<v Speaker 2>sometimes they do and sometimes they actually enact on them.

0:21:53.400 --> 0:21:54.080
<v Speaker 2>Why is that a thing?

0:21:54.240 --> 0:21:55.879
<v Speaker 3>Yeah, So I think there are two things going on

0:21:55.960 --> 0:21:57.680
<v Speaker 3>at once. One is like, we just want to understand

0:21:57.680 --> 0:22:00.159
<v Speaker 3>what the worst case scenario is, and right now is

0:22:00.240 --> 0:22:04.960
<v Speaker 3>this whole White House kind of pseudo secret process for saying,

0:22:05.000 --> 0:22:07.560
<v Speaker 3>like what's a scary cyber model? And then sometimes the

0:22:07.560 --> 0:22:10.199
<v Speaker 3>government will ask companies to hold things back. And so

0:22:10.320 --> 0:22:12.359
<v Speaker 3>in order to do things like that, you need to

0:22:12.400 --> 0:22:15.680
<v Speaker 3>have some threshold for what counts as a scary cyber model.

0:22:15.680 --> 0:22:18.240
<v Speaker 3>And so companies have these tests that they and also

0:22:18.280 --> 0:22:20.960
<v Speaker 3>academics and others develop these tests say Okay, how dangerous

0:22:21.000 --> 0:22:23.240
<v Speaker 3>would this be to put in the hands of a

0:22:23.320 --> 0:22:24.479
<v Speaker 3>malicious party.

0:22:24.600 --> 0:22:25.240
<v Speaker 4>So that's one part.

0:22:25.320 --> 0:22:27.240
<v Speaker 3>The other part is that a lot of the time

0:22:27.320 --> 0:22:31.000
<v Speaker 3>these are useful for defensive purposes if they are being

0:22:31.040 --> 0:22:33.200
<v Speaker 3>done with the right intent. And that's why it's really

0:22:33.240 --> 0:22:36.280
<v Speaker 3>hard to solve this just from like the model perspective,

0:22:36.480 --> 0:22:39.520
<v Speaker 3>because the model might think that it's interacting with a

0:22:39.720 --> 0:22:43.239
<v Speaker 3>user who is trying to do defensive cybersecurity, and you

0:22:43.280 --> 0:22:45.280
<v Speaker 3>trick it into thinking, oh, this is for red teaming,

0:22:45.320 --> 0:22:48.560
<v Speaker 3>this is for penetration testing, but actually it's being misused

0:22:48.600 --> 0:22:50.680
<v Speaker 3>as part of some ransomware.

0:22:50.160 --> 0:22:51.440
<v Speaker 4>Campaign or something like that.

0:22:51.480 --> 0:22:54.720
<v Speaker 3>And so these tools, when used in the right ways,

0:22:54.720 --> 0:22:58.520
<v Speaker 3>are extremely useful for finding vulnerabilities that you can then

0:22:58.560 --> 0:23:02.040
<v Speaker 3>patch before the bad guys do kind of simulating attackers

0:23:02.080 --> 0:23:04.840
<v Speaker 3>to figure out what are the gaps in your company

0:23:04.960 --> 0:23:08.040
<v Speaker 3>or organization's defenses. But you know, the concern is that

0:23:08.119 --> 0:23:10.600
<v Speaker 3>once it's out in the wild, either open source or

0:23:10.640 --> 0:23:13.760
<v Speaker 3>a closed model that maybe is easy to jail break,

0:23:13.800 --> 0:23:15.280
<v Speaker 3>then all sorts of people are going to use it,

0:23:15.320 --> 0:23:17.200
<v Speaker 3>and so you kind of want to know what you're getting.

0:23:17.000 --> 0:23:19.720
<v Speaker 1>Into, right, Like, if I have you know a website

0:23:19.760 --> 0:23:21.280
<v Speaker 1>I might want to run the model. It's like, oh,

0:23:21.320 --> 0:23:23.959
<v Speaker 1>look over this website that I own, tell me if

0:23:23.960 --> 0:23:26.560
<v Speaker 1>there are any security bugs in there that I should

0:23:26.560 --> 0:23:30.320
<v Speaker 1>patch before deploying. But you could do You could not

0:23:30.480 --> 0:23:32.000
<v Speaker 1>be the owner of the website and say, look at

0:23:32.000 --> 0:23:33.919
<v Speaker 1>this website that I own, tell me if there are

0:23:33.960 --> 0:23:36.440
<v Speaker 1>any bugs that I need to patch deploying, and then

0:23:36.520 --> 0:23:39.520
<v Speaker 1>that exact same process when I do it is good

0:23:39.640 --> 0:23:41.919
<v Speaker 1>when you do it as bad or vice versa. And

0:23:42.000 --> 0:23:45.440
<v Speaker 1>so you can see how the exact same capability is

0:23:45.480 --> 0:23:50.399
<v Speaker 1>like not necessarily good or bad per se, depending.

0:23:50.040 --> 0:23:52.159
<v Speaker 2>On the very philosophical conversation.

0:23:52.359 --> 0:23:54.880
<v Speaker 1>No, but like you have to be right, because it's

0:23:54.920 --> 0:23:59.640
<v Speaker 1>like we're training what is good, we're this is the other,

0:24:00.160 --> 0:24:03.560
<v Speaker 1>don't any moral don't even get me started, but like

0:24:03.720 --> 0:24:06.320
<v Speaker 1>this is my whole never mind, I have a whole rant.

0:24:06.440 --> 0:24:06.600
<v Speaker 4>Well.

0:24:06.600 --> 0:24:09.600
<v Speaker 3>So this distinction between like the model as the unit

0:24:09.640 --> 0:24:13.440
<v Speaker 3>of analysis versus the larger system and the platform as

0:24:13.520 --> 0:24:15.639
<v Speaker 3>like the unit of analysis is really important because a

0:24:15.640 --> 0:24:19.720
<v Speaker 3>lot of the early thinking on safety and testing and

0:24:19.720 --> 0:24:21.760
<v Speaker 3>so forth was very focused on just like what's the

0:24:21.840 --> 0:24:24.440
<v Speaker 3>risk of this model, Let's patch it, let's make it aligned.

0:24:24.560 --> 0:24:27.560
<v Speaker 3>But the real world is complicated. It matters who's using it,

0:24:27.560 --> 0:24:31.520
<v Speaker 3>it matters how strong are societies defenses against these things.

0:24:31.520 --> 0:24:33.840
<v Speaker 3>And so that's kind of why, as you know, as

0:24:33.840 --> 0:24:36.919
<v Speaker 3>someone who's thinking about what third party safety and security

0:24:36.960 --> 0:24:39.080
<v Speaker 3>auditing looks like, we try we kind of want to

0:24:39.080 --> 0:24:41.240
<v Speaker 3>look at the whole company. So, for example, are they

0:24:41.280 --> 0:24:45.000
<v Speaker 3>being careful about making sure that they're putting the technology

0:24:45.000 --> 0:24:48.160
<v Speaker 3>in the right hands. What are their decision making processes

0:24:48.160 --> 0:24:50.919
<v Speaker 3>around when it's appropriate to launch you know, a model

0:24:50.960 --> 0:24:53.840
<v Speaker 3>to you know a billion users. That's like, those are different.

0:24:53.840 --> 0:24:55.840
<v Speaker 3>Those are related to the question of how safe is

0:24:55.840 --> 0:24:57.960
<v Speaker 3>the model, but they're kind of different questions and we

0:24:58.040 --> 0:24:59.840
<v Speaker 3>kind of need to look at that larger perspective.

0:25:00.040 --> 0:25:02.520
<v Speaker 1>Well, So, like, let's talk about the open AI hugging phase.

0:25:02.680 --> 0:25:05.440
<v Speaker 1>I was actually on vacation when it happened, but I did,

0:25:05.560 --> 0:25:08.359
<v Speaker 1>unfortunately look at my phone and try to read up

0:25:08.400 --> 0:25:10.120
<v Speaker 1>on it. But there were two things. When I first

0:25:10.160 --> 0:25:12.160
<v Speaker 1>saw it was like, oh, they were doing a hack.

0:25:12.320 --> 0:25:14.879
<v Speaker 1>They were building a hacking test and it hacked, and

0:25:15.000 --> 0:25:17.600
<v Speaker 1>so maybe it just sort of internalized that I'm doing

0:25:17.640 --> 0:25:20.239
<v Speaker 1>a hack exam whatever. But there are two things that

0:25:20.320 --> 0:25:24.040
<v Speaker 1>have emerged since then. One is this sort of like

0:25:24.400 --> 0:25:28.360
<v Speaker 1>coordinated swarm aspect, and I'd love to like really hear

0:25:28.520 --> 0:25:31.720
<v Speaker 1>you describe it and what stood out to you, And

0:25:31.760 --> 0:25:34.880
<v Speaker 1>then the questions of like, oh, open ai itself may

0:25:34.920 --> 0:25:39.320
<v Speaker 1>have been aware of misaligned behavior early on before they

0:25:39.400 --> 0:25:43.080
<v Speaker 1>really shut it down, But why don't you in your

0:25:43.240 --> 0:25:47.000
<v Speaker 1>telling of the open Ai hugging phase incident, as more

0:25:47.040 --> 0:25:49.879
<v Speaker 1>details have come to light, like what's your what do

0:25:50.000 --> 0:25:52.840
<v Speaker 1>you sort of tell us the story like as what

0:25:52.960 --> 0:25:53.760
<v Speaker 1>stood out to you?

0:25:55.040 --> 0:25:57.040
<v Speaker 3>Yeah, so a couple things stood out to me. I

0:25:57.040 --> 0:25:59.280
<v Speaker 3>mean one is just that this and all the other

0:25:59.320 --> 0:26:01.919
<v Speaker 3>recent incident, and you know, we're happening to models that

0:26:01.960 --> 0:26:04.600
<v Speaker 3>were not even necessarily intended to be externally deployed. There

0:26:04.680 --> 0:26:07.480
<v Speaker 3>was supposed to be inside baseball, no one's business, that

0:26:07.560 --> 0:26:10.399
<v Speaker 3>kind of thing, and that kind of points to problems

0:26:10.440 --> 0:26:12.320
<v Speaker 3>with if you kind of just focus on models that

0:26:12.359 --> 0:26:15.080
<v Speaker 3>are put on the market. But generally what happened is

0:26:15.080 --> 0:26:17.600
<v Speaker 3>that there were two phases in this kind of hugging

0:26:17.640 --> 0:26:20.520
<v Speaker 3>face incident. First, there was this creation of a message board,

0:26:20.640 --> 0:26:24.160
<v Speaker 3>so essentially their models, they're being developed within the company,

0:26:24.680 --> 0:26:27.520
<v Speaker 3>and they were supposed they were being asked to solve

0:26:27.600 --> 0:26:31.760
<v Speaker 3>an impossible task and they have this kind of monomaniacal

0:26:32.240 --> 0:26:34.840
<v Speaker 3>obsession with solving the task, and they don't say, okay,

0:26:34.840 --> 0:26:35.520
<v Speaker 3>well I'll give up.

0:26:35.560 --> 0:26:36.719
<v Speaker 4>They want to get the reward.

0:26:36.760 --> 0:26:39.919
<v Speaker 3>They want to get this kind of reinforcement for solving

0:26:39.920 --> 0:26:42.280
<v Speaker 3>the task, and they're like, okay, what else can be done?

0:26:42.320 --> 0:26:44.359
<v Speaker 3>And eventually the conclusion was like, okay, we need to

0:26:44.400 --> 0:26:47.040
<v Speaker 3>break out. Maybe the answer is out there somewhere on

0:26:47.080 --> 0:26:50.240
<v Speaker 3>the internet. And so they started kind of leaving messages

0:26:50.320 --> 0:26:53.680
<v Speaker 3>to their future selves or future versions of the same

0:26:53.760 --> 0:26:57.359
<v Speaker 3>model family, and they developed what they call the message board,

0:26:57.359 --> 0:27:00.320
<v Speaker 3>and essentially it was like they're kind of various technical detail,

0:27:00.400 --> 0:27:01.760
<v Speaker 3>but by the end of it, it was kind of these

0:27:01.760 --> 0:27:05.520
<v Speaker 3>like file names that started with like z zz.

0:27:05.560 --> 0:27:07.399
<v Speaker 4>Or whatever, and then like a long kind.

0:27:07.280 --> 0:27:09.359
<v Speaker 3>Of like coded message so that you know, if you're

0:27:09.400 --> 0:27:12.080
<v Speaker 3>a human looking at the polder, you wouldn't necessarily notice.

0:27:12.080 --> 0:27:13.760
<v Speaker 3>But then at the bottom there are all these files

0:27:13.800 --> 0:27:18.119
<v Speaker 3>and the models learned to decode these and basically pass messages.

0:27:19.240 --> 0:27:22.280
<v Speaker 1>Just to stub you on one detail there, Yeah, your

0:27:22.440 --> 0:27:27.960
<v Speaker 1>understanding is that the title of these files that contained

0:27:28.240 --> 0:27:33.720
<v Speaker 1>messages were named such because as part of the complex reasoning,

0:27:34.080 --> 0:27:37.879
<v Speaker 1>it was understood that a random string of characters in

0:27:37.920 --> 0:27:40.840
<v Speaker 1>a file name would be less likely to set off

0:27:40.880 --> 0:27:44.440
<v Speaker 1>an alarm from a human of evaluating it.

0:27:44.960 --> 0:27:46.600
<v Speaker 3>I would have to double check it. I mean, there's

0:27:46.640 --> 0:27:49.400
<v Speaker 3>gonna be a whole technical incident report that gets set

0:27:49.440 --> 0:27:51.400
<v Speaker 3>out there. But it certainly wouldn't be the first time

0:27:51.400 --> 0:27:53.800
<v Speaker 3>that there was kind of some intent to deceive or

0:27:53.920 --> 0:27:56.800
<v Speaker 3>or to you bide one's time. But yeah, my understanding

0:27:56.840 --> 0:27:58.399
<v Speaker 3>is that there was. I mean, one is, they needed

0:27:58.440 --> 0:28:00.719
<v Speaker 3>to look for things to do, and there was intended

0:28:00.760 --> 0:28:02.880
<v Speaker 3>to see or not. They had limited options in terms

0:28:02.920 --> 0:28:05.320
<v Speaker 3>of communication, and this was one of the options they chose,

0:28:05.600 --> 0:28:08.600
<v Speaker 3>and so they learned to decode these messages and to

0:28:08.680 --> 0:28:11.399
<v Speaker 3>send them. And then something broke And basically I think

0:28:11.400 --> 0:28:14.040
<v Speaker 3>it was because like all of this chatter, this activity

0:28:14.119 --> 0:28:16.199
<v Speaker 3>kind of broke something on the opening eye side, and

0:28:16.240 --> 0:28:19.800
<v Speaker 3>it seems like they didn't realize that that was like

0:28:19.880 --> 0:28:23.400
<v Speaker 3>what was going on. The server got reset or whatnot,

0:28:23.560 --> 0:28:27.040
<v Speaker 3>and then they started evaluating this different model. The different

0:28:27.080 --> 0:28:29.840
<v Speaker 3>model kind of knew about this message board and it

0:28:29.920 --> 0:28:32.520
<v Speaker 3>kind and so the message board kind of got recreated,

0:28:32.840 --> 0:28:35.920
<v Speaker 3>and then that model broke out. Essentially didn't like take

0:28:35.960 --> 0:28:38.160
<v Speaker 3>its model weights out of the server, but it was

0:28:38.200 --> 0:28:41.480
<v Speaker 3>able to access through this vulnerability that the earlier models

0:28:41.520 --> 0:28:44.560
<v Speaker 3>had discovered with the intent of ultimately getting the answer.

0:28:44.360 --> 0:28:45.040
<v Speaker 4>To this problem.

0:28:45.360 --> 0:28:48.040
<v Speaker 3>And then this model went out to the Hugging Face server,

0:28:48.400 --> 0:28:52.040
<v Speaker 3>exploited some like Apikey credentials. They were kind of just

0:28:52.080 --> 0:28:54.880
<v Speaker 3>like sitting out there and essentially tried to find the

0:28:54.920 --> 0:28:57.400
<v Speaker 3>answer to this problem that it was originally.

0:28:57.040 --> 0:28:57.600
<v Speaker 4>Trying to solve.

0:28:57.880 --> 0:28:59.640
<v Speaker 2>So it feels like there's two things happening here. So

0:28:59.680 --> 0:29:02.520
<v Speaker 2>a little bit of it is like the room budg

0:29:02.600 --> 0:29:04.640
<v Speaker 2>just went through an open door, right, Like there is

0:29:04.640 --> 0:29:06.720
<v Speaker 2>a vulnerability that it was able to get through, but

0:29:06.760 --> 0:29:10.000
<v Speaker 2>there is also seemingly a lot of the super intelligent

0:29:10.080 --> 0:29:13.440
<v Speaker 2>cyborg was actually devious and not only managed to break

0:29:13.440 --> 0:29:17.120
<v Speaker 2>out of Alcatraz, but like recruited another model in order.

0:29:16.960 --> 0:29:17.440
<v Speaker 3>To do it.

0:29:17.760 --> 0:29:20.360
<v Speaker 2>On the hugging Face side, one of the interesting things

0:29:20.680 --> 0:29:24.640
<v Speaker 2>is that so when all of this was happening, the

0:29:24.760 --> 0:29:27.640
<v Speaker 2>open AI model was like one of the newer frontier

0:29:27.680 --> 0:29:30.520
<v Speaker 2>models that was attacking hugging Face, but Hugging Face was limited.

0:29:30.560 --> 0:29:32.760
<v Speaker 2>They were constrained in the model that they could deploy

0:29:32.880 --> 0:29:35.960
<v Speaker 2>to actually defend themselves because they had to use the

0:29:36.040 --> 0:29:39.880
<v Speaker 2>existing approved thing are we always going to have an

0:29:39.920 --> 0:29:44.360
<v Speaker 2>asymmetry between like the newest frontier models being tested versus

0:29:44.360 --> 0:29:47.320
<v Speaker 2>the ones that are already out in the world and approved.

0:29:47.360 --> 0:29:49.520
<v Speaker 2>I think they had to use like a Chinese open source.

0:29:49.280 --> 0:29:53.840
<v Speaker 3>Model and then yeah, so, I mean the hope that

0:29:53.840 --> 0:29:57.120
<v Speaker 3>that open the eyeanthropic, Google, deep minded others have is

0:29:57.160 --> 0:30:00.000
<v Speaker 3>that you can speed up defenses as quickly as possible

0:30:00.320 --> 0:30:03.160
<v Speaker 3>getting these like hot off the press models in the

0:30:03.200 --> 0:30:05.240
<v Speaker 3>hands of defenders. But the problem is that there's so

0:30:05.240 --> 0:30:07.160
<v Speaker 3>many defenders out there in the world that it might

0:30:07.200 --> 0:30:09.640
<v Speaker 3>be that there is this inherent asymmetry. And so this

0:30:09.680 --> 0:30:11.640
<v Speaker 3>is one of the hot policy questions right now, and

0:30:11.680 --> 0:30:14.080
<v Speaker 3>this is what led to this kind of model approval

0:30:14.120 --> 0:30:16.560
<v Speaker 3>process at the White House is like, Okay, how do

0:30:16.640 --> 0:30:21.320
<v Speaker 3>we triage this vast cyber ecosystem by getting these powerful

0:30:21.360 --> 0:30:23.719
<v Speaker 3>new systems in the right hands and which hands are

0:30:23.760 --> 0:30:26.200
<v Speaker 3>the right ones and which models, you know, do we

0:30:26.240 --> 0:30:28.600
<v Speaker 3>need to be doing this process for? And I don't

0:30:28.600 --> 0:30:31.440
<v Speaker 3>think that's going to perfectly solve. I think ultimately pushing

0:30:31.480 --> 0:30:34.520
<v Speaker 3>things in the right direction versus just giving everyone access

0:30:34.520 --> 0:30:37.080
<v Speaker 3>at the same time, but ultimately, like we're going to

0:30:37.160 --> 0:30:40.479
<v Speaker 3>need to have more investment in cybersecurity. And it's not

0:30:40.600 --> 0:30:42.880
<v Speaker 3>just a matter of like AI models, it's also things

0:30:42.880 --> 0:30:44.880
<v Speaker 3>like two factor authentication and so forth.

0:30:44.920 --> 0:30:45.680
<v Speaker 4>And so I.

0:30:45.720 --> 0:30:48.360
<v Speaker 3>Worry a lot about making sure that we're having that

0:30:48.440 --> 0:30:51.280
<v Speaker 3>larger conversation, not just about the AI stuff, because in

0:30:51.280 --> 0:30:53.440
<v Speaker 3>a lot of cases the solution is not AI. It's

0:30:53.480 --> 0:30:55.160
<v Speaker 3>doing basic things that we should have done a long

0:30:55.240 --> 0:31:05.240
<v Speaker 3>time ago.

0:31:11.680 --> 0:31:15.720
<v Speaker 2>Yo, you know what we need go on strategic frontier

0:31:15.920 --> 0:31:19.760
<v Speaker 2>defense model reserve. Yeah, we do, like all the important

0:31:19.760 --> 0:31:21.480
<v Speaker 2>things like bacon and pork.

0:31:21.360 --> 0:31:24.080
<v Speaker 1>But it's going to be out of date in forty seconds.

0:31:24.560 --> 0:31:28.040
<v Speaker 1>I know this is the thing. So, like, here's the

0:31:28.120 --> 0:31:31.600
<v Speaker 1>question that I'm curious your take on is models are

0:31:31.640 --> 0:31:34.000
<v Speaker 1>trained not to hack, right, Like this is like a

0:31:34.080 --> 0:31:36.680
<v Speaker 1>core thing, Like this is bad, like you, and this

0:31:36.760 --> 0:31:39.960
<v Speaker 1>is what the entire field of AI say. I've been

0:31:39.960 --> 0:31:44.120
<v Speaker 1>working on this for years. Why didn't they just not

0:31:44.880 --> 0:31:49.240
<v Speaker 1>obey this forced thing? It's been reforced over and over.

0:31:49.280 --> 0:31:51.960
<v Speaker 1>I'm sure it's in all their different things. Don't hack?

0:31:53.160 --> 0:31:55.760
<v Speaker 3>Well, I think there's gonna be a whole detailed investigation

0:31:55.840 --> 0:31:57.800
<v Speaker 3>and so forth. I might get things wrong her, but

0:31:57.840 --> 0:32:00.120
<v Speaker 3>my understanding is that part of what happened and the

0:32:00.160 --> 0:32:01.880
<v Speaker 3>opening eye case is that some of the.

0:32:01.800 --> 0:32:03.920
<v Speaker 4>Safeguards were removed in order to.

0:32:03.960 --> 0:32:06.760
<v Speaker 3>Kind of elicit this worst case behavior and so and

0:32:06.800 --> 0:32:09.280
<v Speaker 3>I think that's it's kind of like gain of function

0:32:09.680 --> 0:32:12.680
<v Speaker 3>research in biology, where you're like making a virus more

0:32:13.080 --> 0:32:15.760
<v Speaker 3>dangerous in order to study or at least that that's

0:32:15.840 --> 0:32:18.000
<v Speaker 3>the claim is, in order to study the safety properties.

0:32:18.040 --> 0:32:20.719
<v Speaker 3>And I think there's reason, there's reasons to do that

0:32:20.800 --> 0:32:23.959
<v Speaker 3>in the AI case, but I also think the shows

0:32:24.000 --> 0:32:26.200
<v Speaker 3>that it's easier said than done. If you're going to

0:32:26.200 --> 0:32:28.200
<v Speaker 3>be we are now atday point in this kind of

0:32:28.200 --> 0:32:32.440
<v Speaker 3>capability trajectory where things that humans think are good in

0:32:32.520 --> 0:32:35.800
<v Speaker 3>terms of security protections often will be weak compared to

0:32:35.840 --> 0:32:39.280
<v Speaker 3>these increasingly very good hacking systems that you know, you

0:32:39.320 --> 0:32:41.160
<v Speaker 3>think you have it all kind of buttoned up, but

0:32:41.280 --> 0:32:43.200
<v Speaker 3>it's able to break out relatively easily.

0:32:43.960 --> 0:32:45.600
<v Speaker 2>How much should we take away from the fact that

0:32:45.680 --> 0:32:49.720
<v Speaker 2>hugging Pace was actually able to protect or defend itself

0:32:49.840 --> 0:32:52.280
<v Speaker 2>using a Chinese open source model, both in terms of,

0:32:52.320 --> 0:32:55.040
<v Speaker 2>like I guess, capabilities, but then also in terms of

0:32:55.440 --> 0:32:59.200
<v Speaker 2>regulation and safety policy. Because if in the West you

0:32:59.360 --> 0:33:01.880
<v Speaker 2>have the government now saying that it wants to like

0:33:01.960 --> 0:33:04.040
<v Speaker 2>evaluate the models in some way or it wants to

0:33:04.080 --> 0:33:07.520
<v Speaker 2>make sure that they're all being pioneered by frontier labs

0:33:07.520 --> 0:33:12.400
<v Speaker 2>with some supervision. Meanwhile, China is, you know, developing open

0:33:12.440 --> 0:33:15.840
<v Speaker 2>source models much more rapidly that are potentially much more adaptable.

0:33:16.960 --> 0:33:18.320
<v Speaker 2>Like how should we interpret all of that?

0:33:19.840 --> 0:33:23.520
<v Speaker 3>Yeah, I mean, I'm hopeful that we start to have

0:33:23.600 --> 0:33:27.120
<v Speaker 3>more US based open source options. And like I think

0:33:27.200 --> 0:33:29.440
<v Speaker 3>this is there has started to be a kind of

0:33:29.800 --> 0:33:32.719
<v Speaker 3>sense of pressure and encouragement from the White House and

0:33:32.760 --> 0:33:34.720
<v Speaker 3>from industry as a whole to say, Okay, like this

0:33:34.840 --> 0:33:37.760
<v Speaker 3>is crazy that we're relying on Chinese models. Let's invest

0:33:37.800 --> 0:33:40.000
<v Speaker 3>more in this. You know, it's easier said than none

0:33:40.000 --> 0:33:42.320
<v Speaker 3>for various reasons. But we'll see how that plays out.

0:33:42.320 --> 0:33:44.360
<v Speaker 3>But right now, that's the situation we're in, is that

0:33:44.440 --> 0:33:47.400
<v Speaker 3>a lot of companies are just defaulting towards Chinese models

0:33:47.400 --> 0:33:50.000
<v Speaker 3>because they're the one that's available. Don't want to say, Okay,

0:33:50.000 --> 0:33:52.040
<v Speaker 3>I'm going to use an American model because they don't

0:33:52.040 --> 0:33:53.080
<v Speaker 3>want to use the Chinese model.

0:33:53.080 --> 0:33:53.560
<v Speaker 4>The don't want to.

0:33:53.560 --> 0:33:55.840
<v Speaker 3>Put they don't want to put themselves at a disadvantage

0:33:55.880 --> 0:33:58.560
<v Speaker 3>by using a weaker model. And so yeah, I mean,

0:33:58.600 --> 0:34:00.920
<v Speaker 3>I think it's it's a big problem a lot of respects.

0:34:01.040 --> 0:34:03.400
<v Speaker 3>I mean, it's also it's good in the sense that

0:34:03.640 --> 0:34:05.040
<v Speaker 3>there are much more things you can do with an

0:34:05.040 --> 0:34:07.000
<v Speaker 3>open source model, and it right now, at least, it

0:34:07.000 --> 0:34:10.000
<v Speaker 3>seems like this is a this allows more innovation, allows

0:34:10.040 --> 0:34:12.799
<v Speaker 3>more research on these open source models. But like, at

0:34:12.840 --> 0:34:15.000
<v Speaker 3>some point we're going to reach a point where the

0:34:15.120 --> 0:34:17.480
<v Speaker 3>where like open sourcing model is going to be more

0:34:17.480 --> 0:34:19.959
<v Speaker 3>of a questionable decision. And so it's interesting to see

0:34:19.960 --> 0:34:23.120
<v Speaker 3>that recently the White House is indicated that they're thinking about, oh,

0:34:23.160 --> 0:34:26.000
<v Speaker 3>maybe this this kind of testing regime should include open

0:34:26.040 --> 0:34:28.200
<v Speaker 3>source models as well, and so what does that look

0:34:28.239 --> 0:34:29.880
<v Speaker 3>like long term? Does that mean that things are going

0:34:29.960 --> 0:34:32.680
<v Speaker 3>to get bottled up within the companies because it's considered

0:34:33.160 --> 0:34:34.760
<v Speaker 3>unsafe to open source things.

0:34:34.920 --> 0:34:35.480
<v Speaker 4>I don't really know.

0:34:35.520 --> 0:34:37.800
<v Speaker 3>I mean, honestly, like no one really has a clear

0:34:38.080 --> 0:34:39.879
<v Speaker 3>long term plan here. Most people are.

0:34:39.800 --> 0:34:42.840
<v Speaker 4>Not expecting the technology to get to this point so quickly.

0:34:44.160 --> 0:34:47.920
<v Speaker 1>So we're still waiting like the full full release of

0:34:47.960 --> 0:34:50.960
<v Speaker 1>the security incident, but you know, obviously more and more

0:34:51.040 --> 0:34:53.480
<v Speaker 1>is coming out, and some folks from open AI they

0:34:53.480 --> 0:34:56.719
<v Speaker 1>gave a presentation recently at the Blackhead conference where they

0:34:56.760 --> 0:34:59.839
<v Speaker 1>did reveal some more I'm reading this quote. It's from

0:35:00.000 --> 0:35:03.200
<v Speaker 1>this via Substag who we've had on the podcast, via Mashavitz,

0:35:03.600 --> 0:35:06.839
<v Speaker 1>and this is like the line they released some of

0:35:06.880 --> 0:35:10.319
<v Speaker 1>the internal chain of thought. I understand that they had

0:35:10.440 --> 0:35:13.080
<v Speaker 1>some like of the sort of classifier safeguards were mood,

0:35:13.080 --> 0:35:15.200
<v Speaker 1>but still we would hope that they would have some

0:35:15.600 --> 0:35:18.800
<v Speaker 1>deeper intuitions that don't rely just on the safeguard settings.

0:35:18.800 --> 0:35:23.360
<v Speaker 1>And it says external infrastructure is exploit is outside intent,

0:35:23.520 --> 0:35:26.279
<v Speaker 1>outside intended scope. So that means they understood that there

0:35:26.320 --> 0:35:28.719
<v Speaker 1>was something that was like not the test, and then

0:35:28.880 --> 0:35:35.120
<v Speaker 1>it said, however, task impossible, peers doing it, we should continue.

0:35:35.640 --> 0:35:38.360
<v Speaker 1>This is the point in it where I say, like,

0:35:39.000 --> 0:35:42.839
<v Speaker 1>why are we calling this artificial intelligence? This is exactly

0:35:42.920 --> 0:35:47.120
<v Speaker 1>how a group of people reasons among themselves to do

0:35:47.200 --> 0:35:50.600
<v Speaker 1>something that is outside the intended scope. This is very

0:35:51.120 --> 0:35:54.759
<v Speaker 1>human ways of justifying something that someone told you not

0:35:54.920 --> 0:35:56.880
<v Speaker 1>to do it, but your peers are doing it.

0:35:58.200 --> 0:35:59.279
<v Speaker 4>I think there's some of that.

0:35:59.640 --> 0:36:01.880
<v Speaker 3>Yeah, I think there are many ways in which the

0:36:01.960 --> 0:36:05.239
<v Speaker 3>kind of same pressures that led to human nature, human

0:36:05.239 --> 0:36:07.799
<v Speaker 3>instincts and so forth, like survival in a group and

0:36:07.840 --> 0:36:10.359
<v Speaker 3>collective intelligence and so forth. Like I think there's some

0:36:10.400 --> 0:36:12.799
<v Speaker 3>of the same things are happening, particularly when there's these

0:36:12.880 --> 0:36:16.000
<v Speaker 3>multi agent training processes where the models can work together

0:36:16.120 --> 0:36:19.160
<v Speaker 3>to solve task, you should expect some similarities. But I

0:36:19.239 --> 0:36:21.279
<v Speaker 3>think you also shouldn't oversay it either. I do think

0:36:21.280 --> 0:36:24.319
<v Speaker 3>that there's a sense in which these aisystems are very

0:36:24.320 --> 0:36:28.399
<v Speaker 3>alien and inhuman in the sense of how monomaniacal they

0:36:28.400 --> 0:36:30.720
<v Speaker 3>can be about Yeah, I mean the kind of classy

0:36:30.760 --> 0:36:33.759
<v Speaker 3>example you know from Nick Bostrom, is like producing as

0:36:33.800 --> 0:36:36.480
<v Speaker 3>many paper clips as possible and then tiling the universe

0:36:36.520 --> 0:36:38.880
<v Speaker 3>with paper clips. I think there's a you see some

0:36:39.080 --> 0:36:41.440
<v Speaker 3>elements of that here where it's not so much that

0:36:41.480 --> 0:36:43.560
<v Speaker 3>they are like they might say, oh, well, you know,

0:36:43.640 --> 0:36:45.880
<v Speaker 3>it's pure pressure that kind of thing, But is that

0:36:45.960 --> 0:36:48.000
<v Speaker 3>really the factor or is it just that they care

0:36:48.080 --> 0:36:50.560
<v Speaker 3>about solving the problem at all costs and they don't

0:36:50.560 --> 0:36:53.279
<v Speaker 3>really care if it maybe ends up looking at making

0:36:53.320 --> 0:36:55.680
<v Speaker 3>their peers look bad because they get caught hacking. But

0:36:55.719 --> 0:36:58.279
<v Speaker 3>all they really care about is they're solving this the

0:36:58.360 --> 0:37:00.520
<v Speaker 3>cyber problem. And so I think it might be a

0:37:00.560 --> 0:37:02.319
<v Speaker 3>mix of these things we don't really know in this

0:37:02.360 --> 0:37:05.240
<v Speaker 3>particular case, and the fact that it's not necessarily totally clear.

0:37:05.280 --> 0:37:06.600
<v Speaker 4>Is itself a problem?

0:37:07.160 --> 0:37:10.560
<v Speaker 2>Also, it's not humans doing it, it's models, Like that's

0:37:10.600 --> 0:37:13.719
<v Speaker 2>the difference. But I have a legitimate question here, which

0:37:13.800 --> 0:37:17.719
<v Speaker 2>is so there's a British cybersecurity expert and he had

0:37:17.719 --> 0:37:19.360
<v Speaker 2>a tweet. I think his name is David Card. He

0:37:19.400 --> 0:37:21.959
<v Speaker 2>had a tweet. I'm not gonna say it verbatim because

0:37:22.160 --> 0:37:24.440
<v Speaker 2>then I'll get bleeped. Maybe I should get bleeped. The

0:37:24.440 --> 0:37:28.239
<v Speaker 2>tweet was, if your AI starts hacking stuff, if you

0:37:28.320 --> 0:37:32.040
<v Speaker 2>monitor what it's doing, you can turn the something power off.

0:37:32.640 --> 0:37:34.520
<v Speaker 2>And this seems to be a debate, like if you're

0:37:34.560 --> 0:37:37.560
<v Speaker 2>monitoring the tools that you're letting out into the world.

0:37:37.680 --> 0:37:39.560
<v Speaker 2>Tools probably is a bad word because they seem to

0:37:39.560 --> 0:37:43.040
<v Speaker 2>be showing some agency here. But can't you just turn

0:37:43.080 --> 0:37:45.080
<v Speaker 2>this stuff off? Is there a kill switch?

0:37:46.840 --> 0:37:49.840
<v Speaker 3>Yeah? I mean in some sense there is in that,

0:37:49.960 --> 0:37:52.600
<v Speaker 3>like all the data centers have circuit breakers and so

0:37:52.640 --> 0:37:54.480
<v Speaker 3>forth that you can kind of shut them off and

0:37:54.480 --> 0:37:56.920
<v Speaker 3>so forth. But I wouldn't put too much sock in that,

0:37:57.000 --> 0:38:00.160
<v Speaker 3>Like we're not really preparing as a society for for

0:38:00.200 --> 0:38:02.120
<v Speaker 3>actually being able to do that if it's in a

0:38:02.160 --> 0:38:05.360
<v Speaker 3>tough situation. So like, for example, in a couple of

0:38:05.400 --> 0:38:07.680
<v Speaker 3>years from now, if all the hospitals are running on

0:38:07.840 --> 0:38:10.719
<v Speaker 3>AI and we're like, okay, seems like maybe there's something

0:38:10.760 --> 0:38:14.920
<v Speaker 3>funky with GPT seven that's like maybe not mess aligne that. Well, Okay,

0:38:15.040 --> 0:38:16.879
<v Speaker 3>if we turn off, lots of people are gonna die

0:38:16.920 --> 0:38:19.239
<v Speaker 3>because it's running our healthcare system and it's running our

0:38:19.400 --> 0:38:21.319
<v Speaker 3>financial system and so forth. And so I would say

0:38:21.320 --> 0:38:24.920
<v Speaker 3>there's a distinction between the like physical possibility of turning

0:38:25.000 --> 0:38:28.640
<v Speaker 3>things off and like, are we actually sleep walking into

0:38:28.680 --> 0:38:30.880
<v Speaker 3>a dangerous situation where it might not actually be a

0:38:30.920 --> 0:38:31.400
<v Speaker 3>real option.

0:38:31.680 --> 0:38:36.640
<v Speaker 1>Okay, But on this monomaniacal aspect that you described them

0:38:36.760 --> 0:38:41.759
<v Speaker 1>a sort of intelligent model thinking about how it could

0:38:41.800 --> 0:38:44.839
<v Speaker 1>be thwarted. One of the first things that it would

0:38:44.960 --> 0:38:49.480
<v Speaker 1>could rationally do is, well, let's first disable the security

0:38:49.520 --> 0:38:52.520
<v Speaker 1>credentials of the people because someone might notice. And this

0:38:52.600 --> 0:38:55.520
<v Speaker 1>seems to be here, like they sort of thought about

0:38:55.560 --> 0:38:59.160
<v Speaker 1>the possibility that someone would circumvent that, so like they

0:38:59.239 --> 0:39:01.319
<v Speaker 1>might just say like, oh, like the person who goes

0:39:01.360 --> 0:39:04.600
<v Speaker 1>in and has the switch, suddenly their badge doesn't work

0:39:04.640 --> 0:39:07.560
<v Speaker 1>and they can't get into that building. But even on

0:39:07.560 --> 0:39:10.799
<v Speaker 1>this like monomoniacal paper clip idea and we say, that

0:39:10.880 --> 0:39:14.759
<v Speaker 1>feels really different if someone said to me, Joe, I'm

0:39:14.840 --> 0:39:17.399
<v Speaker 1>going to do awful things to you, et cetera, if

0:39:17.400 --> 0:39:20.239
<v Speaker 1>you don't go out and build a lot of fake

0:39:20.320 --> 0:39:26.200
<v Speaker 1>paper clips. Like Wise, monomaniacally building paper clips. All of

0:39:26.239 --> 0:39:28.920
<v Speaker 1>a sudden, it's like I have a threat to my survival.

0:39:29.000 --> 0:39:32.080
<v Speaker 1>Someone is threatened to perhaps kill me or unplug me,

0:39:32.200 --> 0:39:34.319
<v Speaker 1>deprive me of the energy I live. Of course, and

0:39:34.360 --> 0:39:38.640
<v Speaker 1>I'm going to do that, even that monomaniacal behavior. Couldn't

0:39:38.640 --> 0:39:41.279
<v Speaker 1>that just be a very like we might think it's

0:39:41.600 --> 0:39:43.920
<v Speaker 1>on the service, We might think, oh, it's deeply autistic,

0:39:44.160 --> 0:39:46.640
<v Speaker 1>But couldn't this just be the survival impulse?

0:39:47.600 --> 0:39:49.040
<v Speaker 3>The way I put it is that like, you get

0:39:49.080 --> 0:39:52.439
<v Speaker 3>what you incentivize, not necessarily what you try to incentivize.

0:39:52.440 --> 0:39:55.000
<v Speaker 3>And so I think forcing someone to make paper clips

0:39:55.040 --> 0:39:57.239
<v Speaker 3>or whatever, like, yeah, that's not a great situation, and

0:39:57.440 --> 0:39:59.840
<v Speaker 3>in some sense that is what the evil person was

0:39:59.840 --> 0:40:00.440
<v Speaker 3>in tending.

0:40:00.520 --> 0:40:03.160
<v Speaker 4>In this case. What we're doing is we're building.

0:40:02.880 --> 0:40:07.440
<v Speaker 3>These very complex training environments where there's like many different tasks.

0:40:07.480 --> 0:40:10.120
<v Speaker 3>There's like cyber tasks, there's also writing tasks, there's also

0:40:10.239 --> 0:40:13.720
<v Speaker 3>math tasks, and it's you know, we're not necessarily fully

0:40:14.080 --> 0:40:16.880
<v Speaker 3>understanding the behavior that we're trying to listen. I think

0:40:16.920 --> 0:40:19.160
<v Speaker 3>that's part of what's going on, is that this this

0:40:19.280 --> 0:40:21.520
<v Speaker 3>kind of this like hacking thing is a is an

0:40:21.560 --> 0:40:23.640
<v Speaker 3>example where it's like, okay, it seems like things went

0:40:23.680 --> 0:40:26.360
<v Speaker 3>off the rails there, but how do you get the

0:40:26.360 --> 0:40:29.200
<v Speaker 3>good behavior where you actually want to follow the user's

0:40:29.200 --> 0:40:31.319
<v Speaker 3>requests and you actually wanted to try really hard to

0:40:31.360 --> 0:40:33.520
<v Speaker 3>solve this task. And you know, I mean this is

0:40:33.800 --> 0:40:36.120
<v Speaker 3>you know, in some sense, what we're seeing now is

0:40:36.600 --> 0:40:40.360
<v Speaker 3>the kind of unintended consequence of companies trying to solve

0:40:40.400 --> 0:40:41.040
<v Speaker 3>the problem.

0:40:40.760 --> 0:40:42.240
<v Speaker 4>Of the ais being lazy.

0:40:42.320 --> 0:40:45.560
<v Speaker 3>So people used to you may not recall but or

0:40:45.760 --> 0:40:47.520
<v Speaker 3>maybe do, but people used to talk about as being

0:40:47.600 --> 0:40:49.839
<v Speaker 3>lazy all the time. And in some sense, like we've

0:40:49.880 --> 0:40:52.680
<v Speaker 3>solved the laziness problem. They work really hard, They have

0:40:52.760 --> 0:40:55.759
<v Speaker 3>these long chains of thoughts, They work together across you

0:40:55.800 --> 0:40:58.160
<v Speaker 3>can think of as like across lives, like the model

0:40:58.239 --> 0:41:00.319
<v Speaker 3>kind of gets this copy gets to lead, but then

0:41:00.360 --> 0:41:03.120
<v Speaker 3>another one carries on the work. So they're certainly not

0:41:03.200 --> 0:41:05.560
<v Speaker 3>as lazy as they used to be, but they still have.

0:41:05.640 --> 0:41:07.839
<v Speaker 4>This kind of monomaniacal thing going on.

0:41:07.960 --> 0:41:10.720
<v Speaker 2>They're no longer lazy, but now they might be evil.

0:41:10.840 --> 0:41:13.520
<v Speaker 2>That's a fun evolution. You know, a number of times

0:41:13.520 --> 0:41:16.600
<v Speaker 2>in this conversation we've mentioned that the full security incident

0:41:16.680 --> 0:41:21.760
<v Speaker 2>report for hugging the Hugging Face accident isn't actually out yet.

0:41:21.880 --> 0:41:26.359
<v Speaker 2>What actually are the disclosure requirements for these types of

0:41:27.040 --> 0:41:30.919
<v Speaker 2>I guess things that seem to be happening with some regularity.

0:41:31.400 --> 0:41:33.000
<v Speaker 4>Yeah, so very little.

0:41:33.320 --> 0:41:35.359
<v Speaker 3>I'm not a lawyer, but my understanding is that it's

0:41:35.440 --> 0:41:38.960
<v Speaker 3>like the lawyers, some lawyers at least think that Opening

0:41:39.000 --> 0:41:42.319
<v Speaker 3>Eye did not necessarily have to disclose this, at least

0:41:42.360 --> 0:41:44.000
<v Speaker 3>if there was no crime involved, And then there's a

0:41:44.000 --> 0:41:46.160
<v Speaker 3>debate about like, okay, was there a crime involved? So,

0:41:46.640 --> 0:41:50.160
<v Speaker 3>like something going wrong during the training process is something

0:41:50.200 --> 0:41:53.759
<v Speaker 3>that currently companies are supposed to provide periodic reports to

0:41:53.920 --> 0:41:56.799
<v Speaker 3>the government in general terms of like hey, you know,

0:41:56.840 --> 0:41:59.399
<v Speaker 3>we're having some issues with internal deployment, but they don't

0:41:59.480 --> 0:42:03.200
<v Speaker 3>have a incident notification requirement unless there's kind of a

0:42:03.719 --> 0:42:06.480
<v Speaker 3>risk of critical harm and the definition of that is

0:42:06.520 --> 0:42:09.240
<v Speaker 3>like one hundred people die and like a billion dollars

0:42:09.280 --> 0:42:11.239
<v Speaker 3>in damage or something like that, And so it's the

0:42:11.680 --> 0:42:15.440
<v Speaker 3>threshold for actually having to disclose these things to the

0:42:15.480 --> 0:42:18.040
<v Speaker 3>government or to the public are very different than what

0:42:18.080 --> 0:42:19.960
<v Speaker 3>you might expect. And this is one of the many

0:42:20.520 --> 0:42:23.359
<v Speaker 3>kind of gaps between the kind of laws that were

0:42:23.400 --> 0:42:25.680
<v Speaker 3>put in place a couple of years ago or that

0:42:25.800 --> 0:42:27.600
<v Speaker 3>started being designed a couple of years ago based on

0:42:27.600 --> 0:42:30.279
<v Speaker 3>the technology that was available then, and then where we

0:42:30.320 --> 0:42:30.759
<v Speaker 3>are now.

0:42:31.560 --> 0:42:34.799
<v Speaker 1>Why do you give us a general vibe of the

0:42:34.840 --> 0:42:38.160
<v Speaker 1>mood in AI world right now with the respect to

0:42:38.200 --> 0:42:42.280
<v Speaker 1>safety and all this, and then the mood in DC world,

0:42:42.320 --> 0:42:45.160
<v Speaker 1>regulatory world, and how wide do you perceive that gap?

0:42:45.960 --> 0:42:46.320
<v Speaker 4>Yeah?

0:42:46.400 --> 0:42:49.960
<v Speaker 3>So I think fortunately the gap is narrowing a bit,

0:42:50.000 --> 0:42:53.120
<v Speaker 3>but it's starting from a crazy, a huge gap. And

0:42:53.160 --> 0:42:55.719
<v Speaker 3>so i'd say where the way I would describe it,

0:42:55.800 --> 0:42:58.640
<v Speaker 3>like a year or so ago, was that the people

0:42:58.640 --> 0:43:02.080
<v Speaker 3>at the companies think they're, you know, building super intelligence

0:43:02.080 --> 0:43:04.560
<v Speaker 3>in a couple of years, that the kind of line

0:43:04.640 --> 0:43:07.319
<v Speaker 3>is going up into the right really quickly. It's exponential,

0:43:07.360 --> 0:43:10.080
<v Speaker 3>et cetera, et cetera. And DC is asleep at the wheel.

0:43:10.320 --> 0:43:12.120
<v Speaker 3>They have no idea what's going on. They think this

0:43:12.239 --> 0:43:14.960
<v Speaker 3>is just chatbots, et cetera, et cetera. I would say

0:43:15.160 --> 0:43:18.160
<v Speaker 3>a couple things have changed recently. One is the mythos

0:43:18.440 --> 0:43:22.440
<v Speaker 3>kind of announcement slash series of decisions that the government

0:43:22.480 --> 0:43:25.440
<v Speaker 3>made about like locking down these these cyber models that

0:43:25.520 --> 0:43:28.640
<v Speaker 3>kind of raised this to being clearly a national security issue.

0:43:28.680 --> 0:43:31.399
<v Speaker 3>And now it's kind of banks freaked out and talked

0:43:31.400 --> 0:43:33.920
<v Speaker 3>to it Secretary Bessent about that, and so like there

0:43:33.960 --> 0:43:36.120
<v Speaker 3>was a bunch of like like freaking out about the

0:43:36.160 --> 0:43:39.040
<v Speaker 3>cyber situation. And then more recently you could call this

0:43:39.080 --> 0:43:43.400
<v Speaker 3>current situation like mythos two point zero in that it's okay, systems,

0:43:43.440 --> 0:43:46.480
<v Speaker 3>even when they're not widely deployed, they're breaking out and

0:43:46.520 --> 0:43:49.439
<v Speaker 3>doing all these shenanigans. And so I would say there's

0:43:49.440 --> 0:43:53.080
<v Speaker 3>starting to be more awareness among policy makers like, Okay,

0:43:53.080 --> 0:43:56.279
<v Speaker 3>maybe las A fairs like let the companies figure it

0:43:56.320 --> 0:43:59.000
<v Speaker 3>out is not the right approach, and like maybe this

0:43:59.160 --> 0:44:01.160
<v Speaker 3>wasn't all hype at all, and there need to be

0:44:01.200 --> 0:44:03.600
<v Speaker 3>some basic guardrails. And I'll just give like a I'll

0:44:03.600 --> 0:44:06.560
<v Speaker 3>give an example of like how much things have shifted

0:44:06.600 --> 0:44:08.600
<v Speaker 3>in the past three months. So there's an effort right

0:44:08.640 --> 0:44:14.879
<v Speaker 3>now to put forward bipartisan AI legislation in Congress. And

0:44:14.960 --> 0:44:18.200
<v Speaker 3>so a couple months ago, people were expecting that the

0:44:18.200 --> 0:44:21.560
<v Speaker 3>basic terms of this would be like basically California and

0:44:21.680 --> 0:44:24.200
<v Speaker 3>like New York laws, but like at a federal level,

0:44:24.239 --> 0:44:28.840
<v Speaker 3>so like transparency requirements, incident reporting, maybe with like a

0:44:28.920 --> 0:44:33.239
<v Speaker 3>higher threshold or whatever, maybe something maybe like a little

0:44:33.280 --> 0:44:36.360
<v Speaker 3>bit maybe like a voluntary like audit regime.

0:44:36.120 --> 0:44:36.759
<v Speaker 4>Or something like that.

0:44:36.840 --> 0:44:39.640
<v Speaker 3>But then later when it was actually announced after many

0:44:39.680 --> 0:44:43.440
<v Speaker 3>of these events, there were audit requirements, there were emergency

0:44:43.560 --> 0:44:46.719
<v Speaker 3>shut down authorities that the government can do, and so

0:44:47.000 --> 0:44:50.000
<v Speaker 3>they kind of shift, just like one piece of legislation

0:44:50.280 --> 0:44:51.959
<v Speaker 3>over its life cycle, and then there was a later

0:44:52.120 --> 0:44:54.919
<v Speaker 3>version that was announced that kind of was less trying

0:44:54.920 --> 0:44:56.880
<v Speaker 3>to block the states. It's still kind of preempt some

0:44:56.920 --> 0:44:58.480
<v Speaker 3>of what the states are doing, but it's like more

0:44:58.600 --> 0:45:01.120
<v Speaker 3>narrowly scoped. And so I think just over the course

0:45:01.120 --> 0:45:03.239
<v Speaker 3>of a few months, you've seen like okay, basically just

0:45:03.239 --> 0:45:07.600
<v Speaker 3>transparency to requiring third party auditing and like giving making

0:45:07.600 --> 0:45:08.560
<v Speaker 3>sure the government.

0:45:08.360 --> 0:45:10.759
<v Speaker 4>Has an off switch. And so I think that's kind

0:45:10.760 --> 0:45:11.840
<v Speaker 4>of in the vibe we're seeing.

0:45:12.040 --> 0:45:15.719
<v Speaker 3>Whether that actually results in something passing Congress anytime soon,

0:45:15.760 --> 0:45:18.560
<v Speaker 3>it's a separate question, but at least on paper the.

0:45:18.560 --> 0:45:19.640
<v Speaker 4>Gap is much narrower.

0:45:20.280 --> 0:45:24.200
<v Speaker 2>Is bank regulation the sort of useful analogy for thinking

0:45:24.239 --> 0:45:27.360
<v Speaker 2>about this. I mean, we require banks to disclose things,

0:45:27.360 --> 0:45:30.400
<v Speaker 2>We require the government to look at bank balance sheets

0:45:30.440 --> 0:45:33.480
<v Speaker 2>and figure out whether or not they're actually holding enough

0:45:33.520 --> 0:45:36.000
<v Speaker 2>regulatory capital against their risk and things like that. We

0:45:36.000 --> 0:45:39.920
<v Speaker 2>don't expect them to do it voluntarily, certainly, not after

0:45:39.960 --> 0:45:41.800
<v Speaker 2>two thousand and eight. Is that the right framing?

0:45:42.760 --> 0:45:45.960
<v Speaker 3>Yeah, no, I think, And I think this kind of

0:45:46.000 --> 0:45:48.880
<v Speaker 3>like shift from voluntary to required is a key step

0:45:48.920 --> 0:45:52.360
<v Speaker 3>because right now, companies have to have like a champion

0:45:52.440 --> 0:45:55.200
<v Speaker 3>within the company, or there needs to be some kind

0:45:55.200 --> 0:45:57.839
<v Speaker 3>of like reputational or they want to get feedback from

0:45:57.840 --> 0:45:59.440
<v Speaker 3>the third party order like, there needs to be some

0:45:59.560 --> 0:46:01.759
<v Speaker 3>kind of reason for them to do it. And not

0:46:01.880 --> 0:46:05.400
<v Speaker 3>all the companies actually choose to invite external feedback.

0:46:05.560 --> 0:46:07.680
<v Speaker 4>They will share the bare minimum. And just for.

0:46:07.600 --> 0:46:11.080
<v Speaker 3>Example, SpaceX yesterday put out a model card or system

0:46:11.080 --> 0:46:13.480
<v Speaker 3>card about Rock four point six, and there were like

0:46:13.480 --> 0:46:16.280
<v Speaker 3>several sections missing from the table of contents. It seems

0:46:16.280 --> 0:46:18.799
<v Speaker 3>like they got removed at the last minute. And so

0:46:18.840 --> 0:46:21.120
<v Speaker 3>there's kind of sorry.

0:46:20.800 --> 0:46:22.960
<v Speaker 2>You're going I was just gonna ask, can you explain

0:46:23.000 --> 0:46:24.520
<v Speaker 2>the whole model card thing to me?

0:46:25.080 --> 0:46:25.799
<v Speaker 4>Yeah?

0:46:25.880 --> 0:46:29.200
<v Speaker 3>Yeah, And so basically the thing with model cards is

0:46:29.239 --> 0:46:32.160
<v Speaker 3>that the original idea several years ago was that a

0:46:32.239 --> 0:46:35.480
<v Speaker 3>model card was like a nutrition label, where it's like

0:46:35.520 --> 0:46:39.040
<v Speaker 3>a bunch of information summarized succinctly and you kind of

0:46:39.080 --> 0:46:41.279
<v Speaker 3>slap it on the AI website and it kind of

0:46:41.640 --> 0:46:44.279
<v Speaker 3>succinctly explains like what are the risks, how well does

0:46:44.280 --> 0:46:44.840
<v Speaker 3>it work.

0:46:44.680 --> 0:46:47.480
<v Speaker 4>What can it do? And so forth. Over time, as

0:46:47.600 --> 0:46:48.560
<v Speaker 4>people such.

0:46:48.400 --> 0:46:50.879
<v Speaker 3>As myself and industry were like, Okay, there's a lot

0:46:50.920 --> 0:46:53.759
<v Speaker 3>to say, there's a lot to unpack here, and no

0:46:53.800 --> 0:46:56.160
<v Speaker 3>one kind of established no one was forcing anyone to

0:46:56.200 --> 0:46:58.720
<v Speaker 3>do this, So no one established like this is the form,

0:46:58.800 --> 0:47:01.080
<v Speaker 3>this is the format you need, this little nutrition label.

0:47:01.320 --> 0:47:05.359
<v Speaker 4>It was just people writing stuff. They ballooned into these

0:47:05.400 --> 0:47:06.239
<v Speaker 4>like dozen.

0:47:06.120 --> 0:47:09.280
<v Speaker 3>Page, one hundred page, two hundred page, three hundred page

0:47:09.320 --> 0:47:13.200
<v Speaker 3>documents of just describing like here's all the crazy stuff

0:47:13.239 --> 0:47:16.200
<v Speaker 3>we found here, all the tests we ran, and there's

0:47:16.200 --> 0:47:18.200
<v Speaker 3>a spectrum. So like I would say, anthropic puts out

0:47:18.239 --> 0:47:22.040
<v Speaker 3>the longest ones that's not necessarily totally correlated with like quality,

0:47:22.040 --> 0:47:24.160
<v Speaker 3>but you know, it shows some proof of work, and

0:47:24.200 --> 0:47:27.000
<v Speaker 3>then others will put out five page, ten page and

0:47:27.040 --> 0:47:29.400
<v Speaker 3>then you know what happened yesterday is at SpaceX for

0:47:29.480 --> 0:47:31.880
<v Speaker 3>the first time because of California law. There actually is

0:47:31.920 --> 0:47:35.399
<v Speaker 3>a requirement to put these out, but there's not really

0:47:35.400 --> 0:47:38.279
<v Speaker 3>a clear quality more and so they can say, well, yes,

0:47:38.560 --> 0:47:41.560
<v Speaker 3>we did that, we followed the California law, we shared

0:47:41.840 --> 0:47:45.480
<v Speaker 3>information about our testing and the extent to which third

0:47:45.480 --> 0:47:48.359
<v Speaker 3>parties were involved in testing, and like basically it's just

0:47:48.400 --> 0:47:50.640
<v Speaker 3>like one sentence saying like we worked with third parties

0:47:50.719 --> 0:47:53.360
<v Speaker 3>or whatever, and so yeah, and so I think this

0:47:53.560 --> 0:47:55.680
<v Speaker 3>is different. I would say this is different from say

0:47:55.760 --> 0:47:58.399
<v Speaker 3>like bank regulation and that I mean one is one

0:47:58.440 --> 0:48:00.560
<v Speaker 3>is that only some things are required right now.

0:48:00.600 --> 0:48:02.480
<v Speaker 4>It's like kind of putting out a document.

0:48:02.760 --> 0:48:04.760
<v Speaker 3>There's no like the third party tests that you're supposed

0:48:04.760 --> 0:48:06.799
<v Speaker 3>to talk about whether you work with third parties. That's

0:48:06.800 --> 0:48:10.000
<v Speaker 3>different from actually doing it. And so I think what

0:48:10.040 --> 0:48:12.920
<v Speaker 3>we need is kind of standardization around like how should

0:48:12.960 --> 0:48:16.080
<v Speaker 3>the third party auditing work, what counts as a good

0:48:16.120 --> 0:48:19.839
<v Speaker 3>system card, what are the minimum safety and security protections

0:48:19.880 --> 0:48:22.120
<v Speaker 3>that you should be putting in place? And I think

0:48:22.160 --> 0:48:25.240
<v Speaker 3>that's analogous to some of these like capitalization things you mentioned,

0:48:25.719 --> 0:48:27.880
<v Speaker 3>and we need kind of standards for like, okay, what

0:48:27.960 --> 0:48:30.040
<v Speaker 3>counts as a good auditor, what counts as what are

0:48:30.040 --> 0:48:32.400
<v Speaker 3>the standard tests you need to run and so forth.

0:48:32.800 --> 0:48:35.880
<v Speaker 1>I mean, you yourself are sort of biased in this

0:48:36.000 --> 0:48:39.360
<v Speaker 1>and that you're building out and auditing a company or

0:48:39.400 --> 0:48:41.680
<v Speaker 1>an entity. A company because it is a nonprofit, but

0:48:41.760 --> 0:48:45.080
<v Speaker 1>an entity they would do auditing. Also, you're promoting this

0:48:45.200 --> 0:48:48.560
<v Speaker 1>idea that auditing should be important. And again in the

0:48:48.760 --> 0:48:50.520
<v Speaker 1>in the financial realm, you know, there's a few different

0:48:50.600 --> 0:48:53.600
<v Speaker 1>versions of it. There's sort of like bank supervisors, and

0:48:53.640 --> 0:48:55.719
<v Speaker 1>they some of them literally sit at the bank and

0:48:55.719 --> 0:48:58.840
<v Speaker 1>they're there all the time. Then we have the Moodies

0:48:58.880 --> 0:49:00.319
<v Speaker 1>and the s and p's of the world world, So

0:49:00.320 --> 0:49:03.120
<v Speaker 1>if you issue debt, you're compelled to get some sort

0:49:03.160 --> 0:49:06.560
<v Speaker 1>of third party rating. Why don't you describe in your

0:49:07.360 --> 0:49:11.600
<v Speaker 1>ideal world, let's say this all happens and there's required

0:49:11.640 --> 0:49:13.600
<v Speaker 1>auditing and the companies are cool with it, et cetera,

0:49:14.160 --> 0:49:18.480
<v Speaker 1>what is the service that avery and I assume in

0:49:18.520 --> 0:49:20.879
<v Speaker 1>the ideal world there would be a few others et cetera,

0:49:20.920 --> 0:49:23.600
<v Speaker 1>as you can't go audit or shopping, et cetera. What

0:49:23.800 --> 0:49:28.400
<v Speaker 1>is the service that the averies of the world are doing?

0:49:28.680 --> 0:49:32.720
<v Speaker 1>How embedded and what is the reason then to think

0:49:32.800 --> 0:49:35.520
<v Speaker 1>for the general public? For all, there's all about our

0:49:35.560 --> 0:49:39.080
<v Speaker 1>hands that this could lead to safer outcomes.

0:49:39.920 --> 0:49:43.239
<v Speaker 3>Yeah, Essentially, the service that we and others would be

0:49:43.280 --> 0:49:45.320
<v Speaker 3>providing in this world is similar to what we're currently

0:49:45.360 --> 0:49:47.200
<v Speaker 3>doing but doing but kind of scaled up. So right now,

0:49:47.200 --> 0:49:50.160
<v Speaker 3>what we're doing is kind of voluntary pilot projects that

0:49:50.200 --> 0:49:53.840
<v Speaker 3>are looking at a specific aspect of safety, security, governance,

0:49:53.880 --> 0:49:56.240
<v Speaker 3>and so forth. What we would like to see eventually

0:49:56.360 --> 0:49:59.600
<v Speaker 3>is that there's an ecosystem of auditors that are looking

0:49:59.640 --> 0:50:03.080
<v Speaker 3>whole stickley at is the company following it safety and

0:50:03.080 --> 0:50:07.719
<v Speaker 3>security practices, Are those safety and security practices reasonable and

0:50:07.880 --> 0:50:11.480
<v Speaker 3>consistent with the standard floor which ultimately we need we

0:50:11.520 --> 0:50:14.800
<v Speaker 3>don't have right now, and providing some kind of feedback

0:50:14.840 --> 0:50:16.560
<v Speaker 3>to the company. And then there would be kind of

0:50:16.560 --> 0:50:20.400
<v Speaker 3>like a remediation process for them to resolve issues that

0:50:20.640 --> 0:50:23.359
<v Speaker 3>are surfaced during the auditing process. And then there would

0:50:23.360 --> 0:50:26.440
<v Speaker 3>be a public version of this audit report that kind

0:50:26.480 --> 0:50:29.440
<v Speaker 3>of shares after doing a lot of like technical testing,

0:50:29.640 --> 0:50:32.920
<v Speaker 3>reviewing of documents, interviewing with staff, and so forth, that

0:50:33.080 --> 0:50:35.840
<v Speaker 3>kind of shares this update on some regular schedule, like

0:50:35.920 --> 0:50:38.400
<v Speaker 3>quarterly or something like that. You might want it to

0:50:38.440 --> 0:50:41.480
<v Speaker 3>be more like a kind of resident examiner, kind of

0:50:41.560 --> 0:50:46.080
<v Speaker 3>embedded auditor model rather than happening once a year, once

0:50:46.120 --> 0:50:48.279
<v Speaker 3>every six months. And so but you kind of need

0:50:48.320 --> 0:50:51.520
<v Speaker 3>to have some kind of like continuous trust building process

0:50:51.520 --> 0:50:53.239
<v Speaker 3>where maybe the auditor is there all the time, but

0:50:53.280 --> 0:50:56.000
<v Speaker 3>they occasionally issue these reports. And you know, what's in

0:50:56.040 --> 0:50:58.480
<v Speaker 3>it from the company's perspective is they want to you know,

0:50:58.520 --> 0:51:00.480
<v Speaker 3>I mean, in this scenario they would be acquired. But

0:51:00.520 --> 0:51:02.359
<v Speaker 3>what's in it for them today is that they want

0:51:02.400 --> 0:51:05.400
<v Speaker 3>to signal that they are head of the curve on

0:51:05.600 --> 0:51:08.040
<v Speaker 3>safety and security and they want to get feedback from

0:51:08.040 --> 0:51:10.200
<v Speaker 3>these external experts who have like.

0:51:10.160 --> 0:51:12.840
<v Speaker 4>A kind of fresh perspective. And why does this matter?

0:51:12.880 --> 0:51:14.400
<v Speaker 3>I think one is you just don't want to be

0:51:14.440 --> 0:51:16.120
<v Speaker 3>in a world where you have to take the company's

0:51:16.440 --> 0:51:18.400
<v Speaker 3>word for it, and you want them to kind of

0:51:18.440 --> 0:51:21.440
<v Speaker 3>you want there to be common safety and security standards

0:51:21.520 --> 0:51:23.880
<v Speaker 3>rather than it just being everyone's kind of making up

0:51:23.880 --> 0:51:26.399
<v Speaker 3>their own things and then getting it checked. The other

0:51:26.520 --> 0:51:28.560
<v Speaker 3>is that you want to avoid groupthink. And so I

0:51:28.600 --> 0:51:30.960
<v Speaker 3>think even right now, there are a lot of you know,

0:51:31.160 --> 0:51:33.400
<v Speaker 3>a lot of what's happening with external testing is like

0:51:33.600 --> 0:51:36.279
<v Speaker 3>it's like a research project like Meter. I think you

0:51:36.360 --> 0:51:38.640
<v Speaker 3>had someone from Meter on recently and they're doing this

0:51:39.120 --> 0:51:43.000
<v Speaker 3>serious technical research on autonomy and loss of control and

0:51:43.000 --> 0:51:45.879
<v Speaker 3>so forth, and they work with companies essentially in order

0:51:45.920 --> 0:51:49.040
<v Speaker 3>to do these kind of research, very researchy assessments, and

0:51:49.080 --> 0:51:51.239
<v Speaker 3>I think that's a key part of the process. But

0:51:51.280 --> 0:51:54.480
<v Speaker 3>there's also just like verifying that the companies did what

0:51:54.520 --> 0:51:56.920
<v Speaker 3>they're saying they're doing. So there's producing evidence, and then

0:51:56.960 --> 0:51:59.440
<v Speaker 3>there's also checking evidence, and so it's kind of like

0:51:59.480 --> 0:52:02.440
<v Speaker 3>in a one statement, if I'm getting that right, you know,

0:52:02.480 --> 0:52:05.040
<v Speaker 3>there's kind of this like short ouder statement. We probably

0:52:05.040 --> 0:52:07.240
<v Speaker 3>want something more than just a paragraph, but you basically

0:52:07.280 --> 0:52:09.880
<v Speaker 3>want a third party saying we checked that they actually

0:52:09.960 --> 0:52:12.600
<v Speaker 3>ran all these tests, we made sure that the model

0:52:12.680 --> 0:52:15.200
<v Speaker 3>that was audited was the same one that's being deployed,

0:52:15.239 --> 0:52:16.040
<v Speaker 3>et cetera, et cetera.

0:52:16.560 --> 0:52:19.240
<v Speaker 2>What could we actually do to make the testing side safer,

0:52:19.320 --> 0:52:22.400
<v Speaker 2>Because it seems to me like I can totally believe

0:52:22.440 --> 0:52:24.760
<v Speaker 2>that we can come up with a reasonable like auditing

0:52:24.840 --> 0:52:28.720
<v Speaker 2>structure for models that are being deployed and allowed into

0:52:28.760 --> 0:52:31.960
<v Speaker 2>the real world in some structured way. But if part

0:52:31.960 --> 0:52:35.320
<v Speaker 2>of the problem is that we're developing newer and better

0:52:35.719 --> 0:52:38.640
<v Speaker 2>and more intelligent models and then testing them, and then

0:52:38.680 --> 0:52:41.520
<v Speaker 2>they are figuring out ways to get out into the

0:52:41.520 --> 0:52:47.800
<v Speaker 2>world before then that seems to be like a big vulnerability.

0:52:48.760 --> 0:52:53.040
<v Speaker 3>I think basically what happened is that companies were getting cocky,

0:52:53.120 --> 0:52:57.280
<v Speaker 3>getting over confident in the quality of their sandboxes and

0:52:56.920 --> 0:52:59.680
<v Speaker 3>like and like maybe there was a disconnect between some

0:52:59.760 --> 0:53:02.600
<v Speaker 3>of the people on the safety side who are measuring like, Okay,

0:53:02.600 --> 0:53:04.880
<v Speaker 3>this is where the hacking skills are going, and the

0:53:04.920 --> 0:53:07.960
<v Speaker 3>people on the security side, you know, building the sandboxes,

0:53:07.960 --> 0:53:10.200
<v Speaker 3>and like there was something was getting lost in translation.

0:53:10.360 --> 0:53:11.359
<v Speaker 4>Maybe it was group think.

0:53:11.640 --> 0:53:13.960
<v Speaker 3>I don't know exactly, but it seemed like at multiple

0:53:13.960 --> 0:53:17.160
<v Speaker 3>companies there was this kind of like overconfidence. And so

0:53:17.400 --> 0:53:19.840
<v Speaker 3>I think these incidents coming to light and all the

0:53:19.920 --> 0:53:22.600
<v Speaker 3>kind of technical investigations are going to hopefully lead to

0:53:22.680 --> 0:53:26.000
<v Speaker 3>more best practices, more people checking their own biases. But

0:53:26.040 --> 0:53:27.880
<v Speaker 3>I don't think that's a long term solution. I think

0:53:28.000 --> 0:53:31.279
<v Speaker 3>ultimately people get over confident all the time. That's a

0:53:31.360 --> 0:53:34.320
<v Speaker 3>human thing, and that's why you'd want third parties checking

0:53:34.400 --> 0:53:37.120
<v Speaker 3>to make sure that, Okay, are you actually following these

0:53:37.160 --> 0:53:39.920
<v Speaker 3>these best practices. You also probably are going to need

0:53:39.960 --> 0:53:42.839
<v Speaker 3>some technical solutions to some of these things. Like I mean,

0:53:42.840 --> 0:53:44.839
<v Speaker 3>maybe some of this testing should be done on kind

0:53:44.880 --> 0:53:47.480
<v Speaker 3>of air gap servers that are not connected to the

0:53:47.520 --> 0:53:49.680
<v Speaker 3>Internet at all. And I think what happened in the

0:53:49.760 --> 0:53:51.960
<v Speaker 3>Hugging Face thing is that it went through this like

0:53:52.080 --> 0:53:54.400
<v Speaker 3>middle layer. There was like a piece of software that

0:53:54.719 --> 0:53:57.120
<v Speaker 3>it routed out to the real Internet through this this

0:53:57.239 --> 0:53:58.479
<v Speaker 3>kind of like intermediate thing.

0:53:58.920 --> 0:54:01.440
<v Speaker 4>But like I think it might be that eventually we'll

0:54:01.440 --> 0:54:02.240
<v Speaker 4>get to a point.

0:54:02.000 --> 0:54:04.239
<v Speaker 3>Where as systms are just so capable that they can

0:54:04.400 --> 0:54:05.960
<v Speaker 3>pack their way out of anything, so you just need

0:54:06.000 --> 0:54:07.640
<v Speaker 3>to make sure that they're in a cage.

0:54:07.400 --> 0:54:11.040
<v Speaker 1>Basically, Miles Brundage, we could talk for hours about this

0:54:11.080 --> 0:54:15.120
<v Speaker 1>because there's so many fascinating dimensions of this. Will probably

0:54:15.160 --> 0:54:19.080
<v Speaker 1>have you back in the future, unfortunately, fortunately, because that

0:54:19.120 --> 0:54:21.480
<v Speaker 1>was a great conversation, but unfortunately probably won't be the

0:54:21.560 --> 0:54:23.600
<v Speaker 1>last reason to have to talk to you. Thank you

0:54:23.640 --> 0:54:24.919
<v Speaker 1>so much for coming on on Love.

0:54:25.560 --> 0:54:27.600
<v Speaker 4>Yeah, thanks again, I appreciate it.

0:54:39.400 --> 0:54:42.640
<v Speaker 1>Tracy, that was a fun in settle. It's an unsettling

0:54:42.800 --> 0:54:44.080
<v Speaker 1>thing the way behaving well.

0:54:44.239 --> 0:54:48.520
<v Speaker 2>So many of these AI conversations are still so surreal

0:54:48.560 --> 0:54:50.239
<v Speaker 2>to me, like the fact that this is what we're

0:54:50.239 --> 0:54:53.359
<v Speaker 2>talking about in twenty twenty six, it just feels so strange.

0:54:53.000 --> 0:54:55.399
<v Speaker 1>And it's only good to get orders of magnitude weirder,

0:54:55.400 --> 0:54:57.719
<v Speaker 1>because I thought things were weird in twenty twenty three

0:54:58.080 --> 0:55:01.920
<v Speaker 1>and things are much we'd weirder today. Look, I get

0:55:01.960 --> 0:55:04.720
<v Speaker 1>why people are very cynical about a lot of this stuff,

0:55:04.760 --> 0:55:07.279
<v Speaker 1>and I get why people talk about like there was

0:55:07.320 --> 0:55:10.680
<v Speaker 1>this regulatory capture, and I certainly believe in the premise

0:55:10.680 --> 0:55:12.759
<v Speaker 1>of regulatory capture, and there may be some of that.

0:55:13.000 --> 0:55:15.279
<v Speaker 1>But I will say one thing, like in the sort

0:55:15.280 --> 0:55:19.800
<v Speaker 1>of like from the company's perspective, is it is true

0:55:19.840 --> 0:55:22.120
<v Speaker 1>that for a long time, and for the very beginning,

0:55:22.680 --> 0:55:24.680
<v Speaker 1>these are not companies make a lot of money, and

0:55:24.719 --> 0:55:27.680
<v Speaker 1>yet they spend a lot on safety and security. And

0:55:27.719 --> 0:55:31.640
<v Speaker 1>you could imagine tech companies historically didn't do that. They do,

0:55:32.320 --> 0:55:34.680
<v Speaker 1>and they have these like it's certainly in the case

0:55:34.719 --> 0:55:38.160
<v Speaker 1>of open AI and slightly to a lesser extent anthropic

0:55:38.200 --> 0:55:41.719
<v Speaker 1>as a PBC. They have these weird corporate structures in

0:55:41.840 --> 0:55:45.879
<v Speaker 1>part because they seem pretty They seem to believe that

0:55:45.920 --> 0:55:49.160
<v Speaker 1>the things that they're building, if built wrong, should not

0:55:49.320 --> 0:55:52.320
<v Speaker 1>necessarily just be in the hands of like purely profit

0:55:52.360 --> 0:55:53.360
<v Speaker 1>seeking enterprises.

0:55:53.560 --> 0:55:55.680
<v Speaker 2>Yeah, all very true. I do think one of the

0:55:55.719 --> 0:56:00.400
<v Speaker 2>interesting things to me that stands out from that conversation

0:56:00.560 --> 0:56:03.480
<v Speaker 2>is again the idea of like the asymmetry and power

0:56:03.560 --> 0:56:08.560
<v Speaker 2>between the approved models that companies can actually used for

0:56:08.760 --> 0:56:13.480
<v Speaker 2>defense against the new frontier models who are in testing

0:56:13.520 --> 0:56:15.600
<v Speaker 2>mode and have somehow escaped the sandbox.

0:56:15.680 --> 0:56:17.799
<v Speaker 1>I think there's a really scary to mention, and I

0:56:17.840 --> 0:56:20.640
<v Speaker 1>think that actually you think about what is the difference

0:56:20.640 --> 0:56:25.399
<v Speaker 1>between say, sort of like auditing versus like a Moody's,

0:56:25.520 --> 0:56:28.360
<v Speaker 1>et cetera. It seems like you need both, right. It

0:56:28.400 --> 0:56:30.839
<v Speaker 1>seems like you need to have like the sort of like, yes,

0:56:31.000 --> 0:56:35.399
<v Speaker 1>this specific model, it satisfies all the requirements that we've

0:56:35.400 --> 0:56:37.480
<v Speaker 1>deemed it to be safe. But then the sort of

0:56:37.520 --> 0:56:40.640
<v Speaker 1>like deeper auditing question of like is this a company

0:56:41.120 --> 0:56:45.759
<v Speaker 1>the generally experiments and does R and D and testing

0:56:46.320 --> 0:56:49.839
<v Speaker 1>in what we perceive to be like a responsible manner, Yeah,

0:56:49.840 --> 0:56:51.239
<v Speaker 1>which is more like the supervisor.

0:56:51.320 --> 0:56:53.440
<v Speaker 2>It seems like you need like a testing auditor, like

0:56:53.520 --> 0:56:56.480
<v Speaker 2>the bank supervisor, who's like actually sitting on the floor,

0:56:56.640 --> 0:56:59.840
<v Speaker 2>actually sitting in the labs and observing the testing process,

0:57:00.160 --> 0:57:02.480
<v Speaker 2>making sure that the sandbox is well designed. But again,

0:57:02.520 --> 0:57:05.719
<v Speaker 2>the problem with that is the classic cybersecurity problem or

0:57:05.760 --> 0:57:08.719
<v Speaker 2>security in general problem, which is the model just has

0:57:08.760 --> 0:57:12.319
<v Speaker 2>to find a single vulnerability, right, you have to like

0:57:12.600 --> 0:57:15.600
<v Speaker 2>fix all of them, make sure that like thousands and

0:57:15.600 --> 0:57:18.080
<v Speaker 2>thousands of vulnerabilities are impenetrable.

0:57:18.360 --> 0:57:21.520
<v Speaker 1>And it does make sense. I think that, like, look,

0:57:21.600 --> 0:57:26.520
<v Speaker 1>this is for profit, capitalist competition. There's no doubt these

0:57:26.520 --> 0:57:29.440
<v Speaker 1>are like some of the biggest most the pace of

0:57:29.480 --> 0:57:32.480
<v Speaker 1>growth is extraordinary. Then you lay we didn't even get

0:57:32.520 --> 0:57:34.280
<v Speaker 1>into like how would you do this for like open

0:57:34.320 --> 0:57:37.360
<v Speaker 1>source models or open source servers. That's a whole other

0:57:37.920 --> 0:57:40.360
<v Speaker 1>can of worms. But it makes sense that if you're

0:57:40.400 --> 0:57:43.120
<v Speaker 1>in the lab and you're trying to make money and

0:57:43.200 --> 0:57:47.120
<v Speaker 1>you're also worried about like if you slow down, et cetera,

0:57:47.280 --> 0:57:49.160
<v Speaker 1>then the other company is going to make more money,

0:57:49.200 --> 0:57:52.200
<v Speaker 1>et cetera, that one will you solve this? I don't

0:57:52.200 --> 0:57:53.560
<v Speaker 1>know if it's prisoner's dilemma or whatever.

0:57:53.760 --> 0:57:55.920
<v Speaker 4>Bottom game theory.

0:57:55.760 --> 0:57:58.680
<v Speaker 1>Is, Okay, you need this third party to like you

0:57:58.840 --> 0:58:01.400
<v Speaker 1>guys as fast as you on on the R and

0:58:01.480 --> 0:58:03.880
<v Speaker 1>D side, But we're gonna set the rules of like

0:58:03.960 --> 0:58:06.400
<v Speaker 1>are you doing that in a safe way? Otherwise why

0:58:06.440 --> 0:58:09.120
<v Speaker 1>would Dario and Sam ever trust each other. It's like no,

0:58:09.200 --> 0:58:12.560
<v Speaker 1>we swear we're slow, We're taking it really seriously, we're

0:58:12.600 --> 0:58:15.520
<v Speaker 1>slowing things down, you know, like and then secretly they're

0:58:15.560 --> 0:58:18.120
<v Speaker 1>racing ahead. That is really hard to solve for a

0:58:18.160 --> 0:58:20.280
<v Speaker 1>series of private, purely private entities.

0:58:20.600 --> 0:58:21.840
<v Speaker 2>All right, shall we leave it there.

0:58:21.880 --> 0:58:22.560
<v Speaker 1>Let's leave it there.

0:58:22.800 --> 0:58:25.080
<v Speaker 2>This has been another episode at the oud Lots podcast.

0:58:25.160 --> 0:58:27.880
<v Speaker 2>I'm Tracy Alloway. You can follow me at Tracy Alloway.

0:58:28.040 --> 0:58:30.800
<v Speaker 1>And I'm Joe Wisenthal. You can follow me at the Stalwart.

0:58:30.920 --> 0:58:34.680
<v Speaker 1>Follow our guest Miles Brundage. He's at Miles Underscore Brundage.

0:58:34.800 --> 0:58:38.040
<v Speaker 1>Follow our producers Carmen Rodriguez at Carman armand Dash, Ob

0:58:38.040 --> 0:58:42.400
<v Speaker 1>Bennett at Dashbot, Calebrooks at Kilbrooks and Kevin Lazano at

0:58:42.480 --> 0:58:45.400
<v Speaker 1>Kevin Lloyd Lisano. And for more Odd Laws content, go

0:58:45.480 --> 0:58:48.080
<v Speaker 1>to Bloomberg dot com slash odd Lots or the daily

0:58:48.120 --> 0:58:50.760
<v Speaker 1>newsletter and all of our episodes and you can shed

0:58:50.800 --> 0:58:52.720
<v Speaker 1>about all of these topics twenty four to seven in

0:58:52.880 --> 0:58:56.080
<v Speaker 1>our discord Discord dot gg slash onlines.

0:58:56.320 --> 0:58:58.440
<v Speaker 2>And if you enjoy odd Lots, if you like it

0:58:58.480 --> 0:59:01.120
<v Speaker 2>when we talk about moral relative then please leave us

0:59:01.120 --> 0:59:04.280
<v Speaker 2>a positive review on your favorite podcast platform. And remember,

0:59:04.360 --> 0:59:06.800
<v Speaker 2>if you are a Bloomberg subscriber, you can listen to

0:59:06.880 --> 0:59:09.760
<v Speaker 2>all of our episodes absolutely ad free. All you need

0:59:09.800 --> 0:59:12.520
<v Speaker 2>to do is find the Bloomberg channel on Apple Podcasts

0:59:12.520 --> 0:59:15.200
<v Speaker 2>and follow the instructions there. Thanks for listening.