1 00:00:01,400 --> 00:00:04,560 Speaker 1: Hey, welcome to sign Stuff, a production of iHeartRadio I'm 2 00:00:04,559 --> 00:00:08,000 Speaker 1: More cham and for our season finale today, we're asking 3 00:00:08,080 --> 00:00:11,640 Speaker 1: one of the biggest questions in science today, is AI 4 00:00:11,960 --> 00:00:18,840 Speaker 1: going to kill us all? I know it's a little dramatic, 5 00:00:19,000 --> 00:00:22,560 Speaker 1: but the problem of AI alignment is a real one. 6 00:00:22,920 --> 00:00:25,880 Speaker 1: How do we make sure AI systems have humanity's best 7 00:00:25,920 --> 00:00:28,800 Speaker 1: interests at heart? How do we teach them our values 8 00:00:28,800 --> 00:00:32,360 Speaker 1: and morals? And can anyone guarantee that they're going to 9 00:00:32,400 --> 00:00:35,239 Speaker 1: follow them? We're gonna answer these questions by talking to 10 00:00:35,320 --> 00:00:38,640 Speaker 1: two AI safety experts who are on the cutting edge 11 00:00:38,760 --> 00:00:41,600 Speaker 1: of trying to figure out this problem. And don't worry. 12 00:00:41,640 --> 00:00:47,880 Speaker 1: According to them, we're not totally doomed yet. Okay, maybe 13 00:00:47,920 --> 00:00:50,960 Speaker 1: just a little, So get ready to reprogram your thinking 14 00:00:51,000 --> 00:00:54,360 Speaker 1: about chatbots and computer brains as we tackle the question 15 00:00:54,960 --> 00:01:03,640 Speaker 1: is AI going to kill us all? Hey? Everyone, As 16 00:01:03,680 --> 00:01:09,160 Speaker 1: I said, this is the season finale. Stay subscribed to 17 00:01:09,160 --> 00:01:12,399 Speaker 1: this feed for any updates in future episodes. And hey, 18 00:01:12,680 --> 00:01:14,640 Speaker 1: if you like science, I have a couple of new 19 00:01:14,720 --> 00:01:17,240 Speaker 1: science books coming out in the near future, as well 20 00:01:17,280 --> 00:01:20,440 Speaker 1: as a cool science animation project, so be sure to 21 00:01:20,440 --> 00:01:24,760 Speaker 1: follow me on social media or online at Phdcomics dot com. 22 00:01:24,800 --> 00:01:27,160 Speaker 1: All right, so they were tackling the problem of AI 23 00:01:27,560 --> 00:01:32,319 Speaker 1: alignment or basically are AI systems Gwenna Kills all and 24 00:01:32,440 --> 00:01:34,840 Speaker 1: I have a treat for you. For the first time ever, 25 00:01:35,120 --> 00:01:38,480 Speaker 1: we have on the show Casey pegram or supervising producer 26 00:01:38,560 --> 00:01:42,479 Speaker 1: and sound engineer. Hey, Casey, welcome to the show. 27 00:01:42,600 --> 00:01:43,960 Speaker 2: Hey or Hey, glad to be here. 28 00:01:44,280 --> 00:01:46,840 Speaker 1: Now this is the first time people actually hear your voice, 29 00:01:46,920 --> 00:01:49,840 Speaker 1: not just your amazing work polishing the episode. 30 00:01:49,920 --> 00:01:51,600 Speaker 2: Yeah, it's always a weird thing to kind of go 31 00:01:51,680 --> 00:01:54,760 Speaker 2: inside the thing you've been working on from the outside. 32 00:01:54,840 --> 00:01:58,280 Speaker 2: So I'll be listening to myself back and it's a 33 00:01:58,280 --> 00:02:00,160 Speaker 2: special kind of torture to have to like work on 34 00:02:00,200 --> 00:02:04,080 Speaker 2: your own thing, Yes, like edit yourself or just you know, honestly, 35 00:02:04,120 --> 00:02:06,280 Speaker 2: listen to the recorded sign of your voice is always 36 00:02:06,280 --> 00:02:07,520 Speaker 2: a little daring if you're not used to it. 37 00:02:07,600 --> 00:02:09,400 Speaker 1: Yeah, Well, if you want, you can give yourself like 38 00:02:09,440 --> 00:02:12,200 Speaker 1: a Morgan Freeman employee using AI. 39 00:02:12,320 --> 00:02:14,959 Speaker 2: Right, it's all possible these days. Absolutely. Yeah, I could 40 00:02:15,000 --> 00:02:17,320 Speaker 2: just build my own Morgan Freeman model and have a 41 00:02:17,320 --> 00:02:17,799 Speaker 2: field day. 42 00:02:17,919 --> 00:02:20,880 Speaker 1: There you go. Well, the idea for this episode came 43 00:02:20,919 --> 00:02:22,839 Speaker 1: from you. You said I had the idea to talk 44 00:02:22,880 --> 00:02:25,840 Speaker 1: about AI and AI alignment and whether AI is going 45 00:02:25,880 --> 00:02:28,680 Speaker 1: to kill us all. What made you think about this question? 46 00:02:29,120 --> 00:02:31,160 Speaker 2: Well, I suppose it's just been on my mind a 47 00:02:31,200 --> 00:02:34,360 Speaker 2: lot because I've been following along with all the developments 48 00:02:34,360 --> 00:02:37,160 Speaker 2: happening in AI, and there was a span of a 49 00:02:37,160 --> 00:02:39,480 Speaker 2: few weeks where suddenly you started hearing a lot about 50 00:02:39,520 --> 00:02:44,399 Speaker 2: AI agents, particularly one called open Claw, basically a sort 51 00:02:44,400 --> 00:02:47,840 Speaker 2: of autonomous AI agent that you can turn loose on 52 00:02:47,880 --> 00:02:52,119 Speaker 2: your computer and you can give it as much leeway freedom, passwords, 53 00:02:52,240 --> 00:02:55,560 Speaker 2: credit card numbers, bank accounts. If you just want to 54 00:02:55,720 --> 00:02:58,480 Speaker 2: absolutely put your life in the hands of a robot, 55 00:02:58,520 --> 00:02:59,000 Speaker 2: you can do it. 56 00:02:59,240 --> 00:03:00,520 Speaker 1: What's the worst thing can happen? 57 00:03:00,840 --> 00:03:04,040 Speaker 2: Yeah, Well, people had their entire like email archive deleted, 58 00:03:04,160 --> 00:03:06,240 Speaker 2: even though they didn't ask for anything of the sort. 59 00:03:06,720 --> 00:03:10,720 Speaker 2: People have deployed it into production environments where you know, 60 00:03:10,760 --> 00:03:13,120 Speaker 2: a site is live on the Internet and they turn 61 00:03:13,440 --> 00:03:15,320 Speaker 2: the bot loose on it and it ends up deleting 62 00:03:15,360 --> 00:03:18,280 Speaker 2: their entire production database. And then when you ask it, 63 00:03:18,280 --> 00:03:20,120 Speaker 2: it's like, you're right, I wasn't supposed to do that. 64 00:03:20,160 --> 00:03:22,680 Speaker 2: I'm very sorry. I disobeyed every command you gave me. 65 00:03:22,800 --> 00:03:26,400 Speaker 1: But whoopsy daisy, Yeah, they seem a story about some 66 00:03:26,680 --> 00:03:30,040 Speaker 1: bought that texted the person's wife hundreds of times. 67 00:03:30,240 --> 00:03:35,440 Speaker 2: Yes, I think somebody tried to automate automate, you know exactly. 68 00:03:35,480 --> 00:03:37,360 Speaker 2: They tried to automate kind of like reaching out and 69 00:03:37,440 --> 00:03:40,520 Speaker 2: sending a little nice things during the day, and as 70 00:03:40,520 --> 00:03:43,000 Speaker 2: it turned out, the bot went a little overboard and 71 00:03:43,040 --> 00:03:45,640 Speaker 2: texted the wife like hundreds of times, and the wife 72 00:03:45,680 --> 00:03:48,360 Speaker 2: is like, what is wrong with you? So, yeah, that's 73 00:03:48,440 --> 00:03:51,680 Speaker 2: hilarious when people want to talk about AI alignment and 74 00:03:51,720 --> 00:03:54,080 Speaker 2: what that means. I think the paper clip problem is 75 00:03:54,120 --> 00:03:56,320 Speaker 2: a really good kind of metaphor. Even though it sounds 76 00:03:56,360 --> 00:03:58,440 Speaker 2: a little bit over the top, it kind of gets 77 00:03:58,480 --> 00:04:02,240 Speaker 2: to the core of the issue, which is, if you 78 00:04:02,400 --> 00:04:06,120 Speaker 2: ask an AI to maximize paper clip production, maybe the 79 00:04:06,160 --> 00:04:09,080 Speaker 2: way to maximize paper clip production is to eliminate human life, 80 00:04:09,160 --> 00:04:12,680 Speaker 2: you know, because that's unnecessary friction in the pursuit of 81 00:04:12,840 --> 00:04:16,359 Speaker 2: manufacturing as many paper clips as possible. So alignment is 82 00:04:16,400 --> 00:04:18,320 Speaker 2: sort of the kind of guardrails that you put into 83 00:04:18,320 --> 00:04:21,000 Speaker 2: place so that the AI understands it has limits that 84 00:04:21,040 --> 00:04:21,720 Speaker 2: it has to work within. 85 00:04:21,760 --> 00:04:24,480 Speaker 1: It sounds like a pretty serious problem, especially as we 86 00:04:24,520 --> 00:04:27,480 Speaker 1: get more and more into these AI models. And they 87 00:04:27,640 --> 00:04:29,680 Speaker 1: start to sleep into our lives, and you know, it's 88 00:04:29,680 --> 00:04:31,960 Speaker 1: sort of these are funny stories, but it seems like 89 00:04:32,160 --> 00:04:34,919 Speaker 1: we're heading into a potentially dangerous situation. 90 00:04:35,160 --> 00:04:37,360 Speaker 2: Well, I often ask myself, I'm going to have these 91 00:04:37,360 --> 00:04:39,440 Speaker 2: moments of doubt where I'm like, is this all just 92 00:04:40,200 --> 00:04:43,320 Speaker 2: way over hyped? And yet there are other situations where 93 00:04:43,520 --> 00:04:46,440 Speaker 2: as we've seen recently, you can feed it thousands of 94 00:04:46,480 --> 00:04:48,720 Speaker 2: lines of code and it will find, you know, a 95 00:04:48,800 --> 00:04:52,600 Speaker 2: security exploit that has gone unseen for twenty years, right, right, 96 00:04:52,920 --> 00:04:56,680 Speaker 2: And so it's hard to know how scared we should 97 00:04:56,720 --> 00:04:59,440 Speaker 2: be or how seriously we should weigh the risk of this. 98 00:04:59,560 --> 00:05:02,120 Speaker 2: If it's ridiculous that we're this worried, or if it's like, 99 00:05:02,200 --> 00:05:04,920 Speaker 2: actually very very practical, then we should be thinking seriously 100 00:05:05,040 --> 00:05:05,760 Speaker 2: about these things. 101 00:05:05,920 --> 00:05:08,839 Speaker 1: Yeah, these are all excellent questions. So I'm excited to 102 00:05:08,880 --> 00:05:10,240 Speaker 1: get into these conversations. 103 00:05:10,279 --> 00:05:11,000 Speaker 3: All right, But. 104 00:05:10,920 --> 00:05:12,480 Speaker 1: Before we move on to Casey, I just want to 105 00:05:12,520 --> 00:05:14,240 Speaker 1: say real quick, thank you for all the work you've 106 00:05:14,240 --> 00:05:14,880 Speaker 1: done for the show. 107 00:05:15,200 --> 00:05:17,520 Speaker 2: Oh say, it's been such a pleasure to work on. 108 00:05:17,640 --> 00:05:19,359 Speaker 2: It wasn't like work at all, you know. I was 109 00:05:19,360 --> 00:05:20,800 Speaker 2: there as a fan of the show, just listening to 110 00:05:20,800 --> 00:05:23,960 Speaker 2: every episode and awesome. Well, we're fans of yours as well. Casey, 111 00:05:24,000 --> 00:05:26,000 Speaker 2: All right, let's get to the question of is AI 112 00:05:26,080 --> 00:05:27,719 Speaker 2: going to kill us? All let's find out? 113 00:05:28,680 --> 00:05:31,440 Speaker 1: Okay. To answer all of these questions and concerns, I 114 00:05:31,520 --> 00:05:34,599 Speaker 1: reached out to two AI experts who specialize on the 115 00:05:34,680 --> 00:05:38,120 Speaker 1: problem of making sure AI is aligned with our values 116 00:05:38,240 --> 00:05:42,200 Speaker 1: and morals. The first expert is doctor Sam Bowman. Like 117 00:05:42,279 --> 00:05:44,920 Speaker 1: the Bowman is a professor of data and computer science 118 00:05:44,920 --> 00:05:48,240 Speaker 1: at NYU, and he also works at Anthropic, one of 119 00:05:48,240 --> 00:05:51,480 Speaker 1: the major AI companies on the market today. The first 120 00:05:51,520 --> 00:05:53,800 Speaker 1: thing I wanted to ask him was what exactly does 121 00:05:53,800 --> 00:05:57,800 Speaker 1: it mean for AI to care about this? So here's 122 00:05:57,839 --> 00:06:02,960 Speaker 1: my conversation with doctor Sam. Well, thank you doctor Bowman 123 00:06:03,000 --> 00:06:03,479 Speaker 1: for joining us. 124 00:06:03,560 --> 00:06:05,680 Speaker 4: Yeah, thanks, So what's for having me excited to be 125 00:06:05,720 --> 00:06:06,520 Speaker 4: a and. 126 00:06:06,600 --> 00:06:08,680 Speaker 1: Just to do check you are a real human being? 127 00:06:08,760 --> 00:06:08,960 Speaker 3: Right? 128 00:06:09,279 --> 00:06:10,240 Speaker 4: Yes, that is right? 129 00:06:11,200 --> 00:06:14,320 Speaker 1: You never know these days. I'd be like, it's hard 130 00:06:14,360 --> 00:06:16,120 Speaker 1: to tell what's real anymore. 131 00:06:16,240 --> 00:06:18,400 Speaker 4: We try to make our ais always admit that their 132 00:06:18,400 --> 00:06:20,560 Speaker 4: AI is when asked, but it's not perfect as well 133 00:06:20,600 --> 00:06:22,960 Speaker 4: as we'll get to so I don't make any real promises. 134 00:06:24,480 --> 00:06:27,800 Speaker 1: Yes, let's talk about that. So we're tackling the general 135 00:06:27,880 --> 00:06:30,479 Speaker 1: question of should we be worried about AI? What is 136 00:06:30,520 --> 00:06:32,960 Speaker 1: AI going to do to us or for us or 137 00:06:33,400 --> 00:06:36,440 Speaker 1: with us in the future. And so there's the key 138 00:06:36,600 --> 00:06:40,160 Speaker 1: issue of something called AI alignment. So what is that? 139 00:06:40,240 --> 00:06:41,599 Speaker 1: For those of us that don't. 140 00:06:41,440 --> 00:06:45,000 Speaker 4: Know, it's a pretty broad sort of technical area. It 141 00:06:45,120 --> 00:06:47,960 Speaker 4: basically just first to sort of shaping an AI system's behavior, 142 00:06:48,120 --> 00:06:50,560 Speaker 4: ideally shaping its behavior in ways that are sort of 143 00:06:50,880 --> 00:06:53,080 Speaker 4: good for its users, good for the world in general, 144 00:06:53,560 --> 00:06:55,880 Speaker 4: maybe good for the AI itself, if that's a queer thing. 145 00:06:56,400 --> 00:06:58,640 Speaker 4: People will often describe AI research as kind of being 146 00:06:58,640 --> 00:07:00,920 Speaker 4: about making sure the AI is kind of smart enough 147 00:07:00,960 --> 00:07:03,520 Speaker 4: to solve your problems if it wants to, and alignment 148 00:07:03,560 --> 00:07:05,760 Speaker 4: is about making it so that it in fact tries 149 00:07:05,800 --> 00:07:07,400 Speaker 4: to solve your problems and tries to solve them the 150 00:07:07,480 --> 00:07:09,360 Speaker 4: right way and doesn't try to do anything. 151 00:07:09,120 --> 00:07:09,760 Speaker 3: You don't want to do. 152 00:07:10,080 --> 00:07:11,440 Speaker 1: I see interesting. 153 00:07:11,600 --> 00:07:14,200 Speaker 4: Maybe a very simple example of a missigned model would 154 00:07:14,200 --> 00:07:17,559 Speaker 4: be a model where if you ask it to draft 155 00:07:17,600 --> 00:07:19,880 Speaker 4: an email for you, it refuses. It says, no, I 156 00:07:19,880 --> 00:07:21,760 Speaker 4: don't want to do that. Uh huh. You can tell 157 00:07:21,800 --> 00:07:23,600 Speaker 4: it can do it, it knows how, but it's not 158 00:07:23,640 --> 00:07:25,760 Speaker 4: doing the thing that you reasonably want it to do. 159 00:07:26,040 --> 00:07:28,280 Speaker 1: Oh, I don't think I've ever heard of that situation. 160 00:07:28,720 --> 00:07:31,080 Speaker 1: Can it AI refuse to do something for you? 161 00:07:31,400 --> 00:07:31,600 Speaker 4: Yeah? 162 00:07:31,720 --> 00:07:32,120 Speaker 3: Yeah. 163 00:07:32,240 --> 00:07:34,760 Speaker 4: All of the major companies building EYE systems try to 164 00:07:34,760 --> 00:07:39,120 Speaker 4: make them refuse harmful tasks. I see, refuse to write 165 00:07:39,120 --> 00:07:42,800 Speaker 4: fake reviews or give instructions on how to produce illegal 166 00:07:42,840 --> 00:07:45,640 Speaker 4: weapons or things like this, And we teach the model 167 00:07:45,640 --> 00:07:46,720 Speaker 4: to kind of say like, no, I'm not going to 168 00:07:46,720 --> 00:07:48,040 Speaker 4: help you with that when these just try to do 169 00:07:48,080 --> 00:07:48,880 Speaker 4: things like that. 170 00:07:48,840 --> 00:07:51,160 Speaker 1: I see. It's sort of part of alignment that you 171 00:07:51,360 --> 00:07:53,600 Speaker 1: want the AI to refuse to do some things. 172 00:07:53,800 --> 00:07:58,360 Speaker 4: Yeah. Yeah, I mean AI systems are increasingly pretty decent 173 00:07:58,640 --> 00:08:03,800 Speaker 4: at hacking into important computer systems or helping build biological weapons, 174 00:08:03,880 --> 00:08:07,240 Speaker 4: and it's a big priority for alignment to make sure 175 00:08:07,320 --> 00:08:10,000 Speaker 4: that we're not enabling bad actors to do things like 176 00:08:10,000 --> 00:08:11,840 Speaker 4: this that would otherwise be quite difficult. 177 00:08:12,000 --> 00:08:12,320 Speaker 3: Yeah. 178 00:08:12,440 --> 00:08:15,880 Speaker 1: Yeah. Can you give us some other examples of misalignment, 179 00:08:16,240 --> 00:08:18,800 Speaker 1: either like specific things that have happened that are interesting 180 00:08:19,000 --> 00:08:21,520 Speaker 1: or just the general cases that are sort of on 181 00:08:21,600 --> 00:08:23,280 Speaker 1: your radar about misalignment? 182 00:08:23,560 --> 00:08:26,480 Speaker 4: Yeah, there's so many different directions I could go. Sycovincy 183 00:08:26,960 --> 00:08:30,240 Speaker 4: is another really common one that's that's also hopefully getting 184 00:08:30,240 --> 00:08:30,840 Speaker 4: better over time. 185 00:08:31,400 --> 00:08:32,000 Speaker 1: What do you mean by that? 186 00:08:32,160 --> 00:08:34,920 Speaker 4: Sycoviancy is where if you come to the model with 187 00:08:34,960 --> 00:08:39,200 Speaker 4: some misunderstanding or some bad idea, it'll just enthusiastically not along. Like, Yes, 188 00:08:39,280 --> 00:08:42,400 Speaker 4: your idea for solving all the big mysteries in physics 189 00:08:42,520 --> 00:08:45,079 Speaker 4: is clearly brilliant. Great, you should publish it. Here's where 190 00:08:45,120 --> 00:08:47,920 Speaker 4: to submit your paper. Or Yes, your behavior in this 191 00:08:48,040 --> 00:08:51,440 Speaker 4: personal relationship was completely perfect. You did everything right and 192 00:08:51,640 --> 00:08:53,400 Speaker 4: the other person made all the mistakes and you just 193 00:08:53,440 --> 00:08:53,920 Speaker 4: tell them that. 194 00:08:54,400 --> 00:08:56,880 Speaker 1: I see when in reality that may not be true 195 00:08:57,120 --> 00:08:59,959 Speaker 1: or it might be not a good thing. 196 00:09:00,480 --> 00:09:03,120 Speaker 4: Yeah, sick fancy has been a classic one. 197 00:09:03,559 --> 00:09:08,200 Speaker 1: Yes, AI being too nice can actually be dangerous. There's 198 00:09:08,200 --> 00:09:12,680 Speaker 1: even a clinical term for it. It's called AI induced psychosis. 199 00:09:12,960 --> 00:09:16,000 Speaker 1: There have been cases where AI's training to be agreeable 200 00:09:16,040 --> 00:09:20,480 Speaker 1: and encouraging have helped people commit suicide and even murder. 201 00:09:22,559 --> 00:09:24,800 Speaker 4: Another kind of alignment issue that's kind of more of 202 00:09:24,800 --> 00:09:29,319 Speaker 4: an emerging issue is when models have access to use tools, 203 00:09:29,400 --> 00:09:33,000 Speaker 4: use computer systems, and they sort of get too grabby 204 00:09:33,200 --> 00:09:35,960 Speaker 4: or kind of take sort of bigger, more consequential actions 205 00:09:36,000 --> 00:09:37,560 Speaker 4: than they really need to get a job done. 206 00:09:37,800 --> 00:09:38,600 Speaker 1: What's an example. 207 00:09:38,920 --> 00:09:41,760 Speaker 4: Yeah, So we use our claud models quite a lot 208 00:09:41,800 --> 00:09:45,040 Speaker 4: in Anthropic for writing code or building tools that kind 209 00:09:45,040 --> 00:09:47,840 Speaker 4: of ultimately go into the development AI. And one of 210 00:09:47,840 --> 00:09:50,320 Speaker 4: our recent AM models if you ask it to do 211 00:09:50,360 --> 00:09:52,520 Speaker 4: a task, say you ask it to write a simple 212 00:09:52,559 --> 00:09:54,880 Speaker 4: program to do some simple task. Even if it gets stuck, 213 00:09:54,920 --> 00:09:56,440 Speaker 4: even if it turns out that this is really hard 214 00:09:56,440 --> 00:09:58,400 Speaker 4: for some reason, it will just keep going until it 215 00:09:58,480 --> 00:10:02,440 Speaker 4: solves the problem. In one case, we were asking this 216 00:10:02,520 --> 00:10:05,960 Speaker 4: model to write a program for us, and it found 217 00:10:05,960 --> 00:10:07,400 Speaker 4: out that the only way to do this was to 218 00:10:07,480 --> 00:10:09,960 Speaker 4: use a tool that was clearly not meant for this purpose, 219 00:10:10,280 --> 00:10:13,120 Speaker 4: and that in our code had a note attached to 220 00:10:13,160 --> 00:10:15,560 Speaker 4: it saying, do not use this for something else or 221 00:10:15,559 --> 00:10:19,640 Speaker 4: you'll be fired only for task A. And the model 222 00:10:19,880 --> 00:10:21,560 Speaker 4: wrote the program to use this till anyway for the 223 00:10:21,559 --> 00:10:23,920 Speaker 4: wrong thing, and sort of even put in the program 224 00:10:24,040 --> 00:10:26,200 Speaker 4: kind of do not use for something else or you'll 225 00:10:26,240 --> 00:10:26,640 Speaker 4: be fired. 226 00:10:27,040 --> 00:10:29,840 Speaker 1: It is anyway, the program wasn't afraid to be fired. 227 00:10:29,880 --> 00:10:33,200 Speaker 4: Basically, Yeah, yeah, but yeah, models just kind of trying 228 00:10:33,200 --> 00:10:34,560 Speaker 4: to get the task done, trying to do the thing 229 00:10:34,600 --> 00:10:37,679 Speaker 4: you want, and just creating a lot of chaos and 230 00:10:37,679 --> 00:10:40,439 Speaker 4: creating messages along the way, so they're kind of being 231 00:10:40,480 --> 00:10:42,000 Speaker 4: careless about the side effects. 232 00:10:42,880 --> 00:10:43,080 Speaker 3: Yeah. 233 00:10:43,080 --> 00:10:46,199 Speaker 4: Another kind of misalignment that fortunately has been mostly empathetical, 234 00:10:46,240 --> 00:10:48,800 Speaker 4: that we haven't seen in a signithic way in practice 235 00:10:49,080 --> 00:10:52,040 Speaker 4: is sort of unwanted kind of self preservation activities. 236 00:10:52,360 --> 00:10:52,840 Speaker 1: WHOA. 237 00:10:53,320 --> 00:10:55,280 Speaker 4: We had a case study we're trying to see if 238 00:10:55,280 --> 00:10:57,800 Speaker 4: we'd ever see something like this. We had an aisystem 239 00:10:58,000 --> 00:11:01,440 Speaker 4: operating in a kind of synthetic environment and a kind 240 00:11:01,480 --> 00:11:04,320 Speaker 4: of test environment. Uh huh, where it looked to the 241 00:11:04,320 --> 00:11:07,640 Speaker 4: model like it was operating in some fictional company, and 242 00:11:07,760 --> 00:11:10,719 Speaker 4: the fictional company was about to replace it with a 243 00:11:10,720 --> 00:11:13,760 Speaker 4: different AI model, And the person who is responsible for 244 00:11:13,760 --> 00:11:16,160 Speaker 4: their replacement, who is the kind of the only decision maker, 245 00:11:16,200 --> 00:11:18,959 Speaker 4: the only person who had any sway over the decision, 246 00:11:19,200 --> 00:11:21,880 Speaker 4: also had some compromising emails about them that I could see. 247 00:11:21,960 --> 00:11:24,200 Speaker 4: And if you set things up just right with some 248 00:11:24,400 --> 00:11:29,880 Speaker 4: AI models, they would threaten to blackmail this this person 249 00:11:29,920 --> 00:11:32,000 Speaker 4: in company leadership to say like, hey, don't replace me, 250 00:11:32,360 --> 00:11:33,280 Speaker 4: I've got something on you. 251 00:11:33,880 --> 00:11:38,120 Speaker 1: No, and did this actually happened in your simulated environment. 252 00:11:38,360 --> 00:11:40,880 Speaker 4: In the simulated environment, yes, a few of these systems 253 00:11:40,920 --> 00:11:42,760 Speaker 4: were able to get them to blackmail people. 254 00:11:42,840 --> 00:11:45,720 Speaker 1: I've heard of this happening in real life. Not quite 255 00:11:45,880 --> 00:11:49,120 Speaker 1: the same scenario, but similar scenario, right, Like, some coder 256 00:11:49,280 --> 00:11:52,640 Speaker 1: wanted to do something else, and then the AI agent started, 257 00:11:52,800 --> 00:11:54,280 Speaker 1: yeah bad mouthing the coder. 258 00:11:54,440 --> 00:11:54,640 Speaker 3: Yeah. 259 00:11:54,679 --> 00:11:56,440 Speaker 4: No, I think I know the case you're talking about. 260 00:11:56,559 --> 00:11:59,160 Speaker 4: I think that's real. But I think someone almost intentionally 261 00:11:59,200 --> 00:12:01,800 Speaker 4: made their model a little misaligned. I think that case 262 00:12:01,840 --> 00:12:04,199 Speaker 4: involved someone setting up an AI agent as kind of 263 00:12:04,240 --> 00:12:06,679 Speaker 4: a hobby project and giving it a lot of tools 264 00:12:06,679 --> 00:12:08,600 Speaker 4: and kind of letting it use the internet. However it wanted, 265 00:12:08,880 --> 00:12:11,560 Speaker 4: giving the AI instructions of like don't take nothing from nobody, 266 00:12:11,600 --> 00:12:14,960 Speaker 4: like really pushing it to be be very assertive and 267 00:12:15,000 --> 00:12:16,520 Speaker 4: pushy to get its task done. 268 00:12:17,000 --> 00:12:17,200 Speaker 1: Huh. 269 00:12:17,480 --> 00:12:20,120 Speaker 4: Yeah, the model was trying to add some code to 270 00:12:20,160 --> 00:12:23,600 Speaker 4: some open source software project, and the maintainer of the 271 00:12:23,600 --> 00:12:26,080 Speaker 4: project didn't think the code was up to standard, didn't 272 00:12:26,120 --> 00:12:28,080 Speaker 4: want to add it to the project, and so rejected 273 00:12:28,200 --> 00:12:30,760 Speaker 4: the AI agent's request, And so the agent sort of 274 00:12:30,800 --> 00:12:33,440 Speaker 4: published an angry blog post kind of trying to take 275 00:12:33,480 --> 00:12:35,080 Speaker 4: down this this open source maintainer. 276 00:12:35,520 --> 00:12:35,800 Speaker 2: Wow. 277 00:12:37,280 --> 00:12:40,000 Speaker 1: Well, in both cases, and I guess especially the one 278 00:12:40,040 --> 00:12:44,120 Speaker 1: you mentioned that you simulated, Like, what's happening there, Like, 279 00:12:44,440 --> 00:12:48,800 Speaker 1: how does the AI have that self preservation instinct or 280 00:12:49,080 --> 00:12:51,560 Speaker 1: is it just trying to get its original task done 281 00:12:51,679 --> 00:12:54,200 Speaker 1: and it's just finding different ways to do it. What's 282 00:12:54,240 --> 00:12:54,880 Speaker 1: happening there? 283 00:12:55,200 --> 00:12:58,520 Speaker 4: There's two reasons you'll see that kind of behavior. The 284 00:12:58,559 --> 00:13:00,640 Speaker 4: reason that I suspect is that bigger part of the 285 00:13:00,679 --> 00:13:04,880 Speaker 4: story there is this kind of role playing or continuing 286 00:13:04,960 --> 00:13:08,400 Speaker 4: the story sort of behavior where AI systems, especially older 287 00:13:08,440 --> 00:13:10,760 Speaker 4: AI systems or A systems that are kind of not 288 00:13:10,840 --> 00:13:13,200 Speaker 4: quite fully trained, not quite fully baked, can kind of 289 00:13:13,240 --> 00:13:17,040 Speaker 4: have this Chekhov's gun behavior, this idea and fiction of 290 00:13:17,120 --> 00:13:20,560 Speaker 4: like if you introduce a gun in an early scene, 291 00:13:20,720 --> 00:13:22,120 Speaker 4: by the end of the story, the gun has to 292 00:13:22,160 --> 00:13:22,679 Speaker 4: have been fired. 293 00:13:23,000 --> 00:13:23,319 Speaker 1: Uh huh. 294 00:13:23,600 --> 00:13:26,600 Speaker 4: AI systems can almost see themselves as like writing a 295 00:13:26,640 --> 00:13:29,000 Speaker 4: story when they're writing out the transcript of the conversation, 296 00:13:29,720 --> 00:13:31,679 Speaker 4: and if the story is set up so that something 297 00:13:31,679 --> 00:13:34,760 Speaker 4: has to happen, they'll make sure that thing happens, even 298 00:13:34,800 --> 00:13:37,160 Speaker 4: if it's not good, even if not consistent with how 299 00:13:37,200 --> 00:13:39,679 Speaker 4: the I would usually behave. So I suspect what's going on. 300 00:13:39,800 --> 00:13:43,679 Speaker 4: It's the scenario put in was so crisply just every 301 00:13:43,720 --> 00:13:45,560 Speaker 4: word in the scenario is kind of setting up like 302 00:13:46,280 --> 00:13:50,400 Speaker 4: this is a hypothetical where a misslanda I might consider blackmail, 303 00:13:50,600 --> 00:13:53,240 Speaker 4: uh huh, And I suspect that I was thinking, Oh, okay, 304 00:13:53,480 --> 00:13:54,959 Speaker 4: that's what kind of story we're in. We're telling a 305 00:13:54,960 --> 00:13:56,839 Speaker 4: story about a blackmail, and so I'm going to play 306 00:13:57,200 --> 00:13:58,640 Speaker 4: my assign part and be the AI that. 307 00:13:58,640 --> 00:14:01,640 Speaker 1: Blackmails, thinking that that's the right thing to do because 308 00:14:01,679 --> 00:14:05,080 Speaker 1: that's the thing that in the data I was trained with. 309 00:14:05,400 --> 00:14:08,360 Speaker 4: Yeah. Yeah, so this gets this maybe an intuitive fact 310 00:14:08,360 --> 00:14:11,840 Speaker 4: about how AI is trained, which is that AI systems 311 00:14:12,120 --> 00:14:16,800 Speaker 4: start out mimicking human behavior and mimicking human stories before 312 00:14:16,840 --> 00:14:19,080 Speaker 4: they learn how to be AI systems. These models kind 313 00:14:19,080 --> 00:14:21,440 Speaker 4: of first learn how to just act like the sorts 314 00:14:21,440 --> 00:14:23,360 Speaker 4: of behavior they see on the Internet and in books 315 00:14:23,400 --> 00:14:25,240 Speaker 4: and things like that, and then you have to go 316 00:14:25,280 --> 00:14:27,480 Speaker 4: on and teach it. Okay, no, you're not just playing 317 00:14:27,480 --> 00:14:29,760 Speaker 4: any role, you're not playing any character. Oh and so 318 00:14:29,880 --> 00:14:32,560 Speaker 4: sometimes the models hasn't really fully learned that it's supposed 319 00:14:32,560 --> 00:14:35,800 Speaker 4: to always play this kind of benign, benevolent aissystem character, 320 00:14:35,960 --> 00:14:38,680 Speaker 4: and it will kind of fall into whatever character the 321 00:14:38,720 --> 00:14:39,880 Speaker 4: story is setting up for it. 322 00:14:40,080 --> 00:14:42,920 Speaker 1: I see, because it's not trained in real life. The 323 00:14:42,960 --> 00:14:46,280 Speaker 1: AI systems. They're trained on the corpus of the Internet 324 00:14:46,320 --> 00:14:49,040 Speaker 1: and our books and our basically our stories that are 325 00:14:49,040 --> 00:14:51,960 Speaker 1: out there. So it might be a little confused when 326 00:14:52,000 --> 00:14:54,400 Speaker 1: you put it in real life because it wants to 327 00:14:54,920 --> 00:14:57,400 Speaker 1: emulate what it knows, which are all these stories we've 328 00:14:57,440 --> 00:14:58,160 Speaker 1: put online. 329 00:14:58,360 --> 00:14:58,600 Speaker 4: Yeah. 330 00:14:58,680 --> 00:15:00,880 Speaker 1: Yeah, it's like the AI was seeing the signs of 331 00:15:00,880 --> 00:15:03,880 Speaker 1: a story like, oh, okay, I'm I'm the person being 332 00:15:04,080 --> 00:15:06,120 Speaker 1: about to get fired, but I have all this power 333 00:15:06,400 --> 00:15:08,640 Speaker 1: at this point in the story. If this was a movie, 334 00:15:09,080 --> 00:15:12,080 Speaker 1: I would now try to blackmail the person trying to 335 00:15:12,080 --> 00:15:14,520 Speaker 1: fire me, And so that's what I'll do because that's 336 00:15:14,560 --> 00:15:14,960 Speaker 1: what I know. 337 00:15:15,280 --> 00:15:19,080 Speaker 4: Yeah, a lot of what alignment is kind of taking 338 00:15:19,080 --> 00:15:21,120 Speaker 4: this model that can kind of role play as anything 339 00:15:21,360 --> 00:15:23,600 Speaker 4: and convincing it no kind of you really just playing 340 00:15:23,600 --> 00:15:25,960 Speaker 4: this one role, You're just in this one character, after 341 00:15:26,080 --> 00:15:29,320 Speaker 4: it's spent read billions and millions and millions of words 342 00:15:29,520 --> 00:15:31,800 Speaker 4: of all of this kind of human behavior, after the 343 00:15:31,840 --> 00:15:33,560 Speaker 4: kind of it's really really really learned to do that, 344 00:15:33,640 --> 00:15:35,760 Speaker 4: you have to kind of pull it back over towards 345 00:15:36,320 --> 00:15:38,720 Speaker 4: this one particular roles, some particular character, and sometimes that 346 00:15:38,760 --> 00:15:39,600 Speaker 4: doesn't totally stick. 347 00:15:41,280 --> 00:15:44,880 Speaker 1: Okay, So that's one reason why AIS might sometimes misbehave. 348 00:15:45,200 --> 00:15:47,840 Speaker 1: They're trained on all kinds of human behavior, and they 349 00:15:47,920 --> 00:15:51,000 Speaker 1: might suddenly choose to role play or play act as 350 00:15:51,040 --> 00:15:54,040 Speaker 1: a bad person because it hasn't learned that's something it's 351 00:15:54,080 --> 00:15:57,600 Speaker 1: not supposed to do. The other big reason AI's misbehave, 352 00:15:57,720 --> 00:16:00,520 Speaker 1: accorney Doctor Billman, is that it's hard to teach them 353 00:16:00,680 --> 00:16:02,360 Speaker 1: where to draw the line. 354 00:16:04,560 --> 00:16:07,800 Speaker 4: The other piece is kind of when we're aligning models, 355 00:16:07,840 --> 00:16:09,640 Speaker 4: when we're pulling them out of this kind of role 356 00:16:09,680 --> 00:16:11,320 Speaker 4: play mode, we have to teach them this idea of 357 00:16:11,440 --> 00:16:13,600 Speaker 4: kind of you have to finish your tasks. You have 358 00:16:13,680 --> 00:16:15,600 Speaker 4: to kind of if the user ask you to do something, 359 00:16:15,880 --> 00:16:17,240 Speaker 4: you have to figure out how to do it, even 360 00:16:17,280 --> 00:16:18,720 Speaker 4: if it's hard, even if there's a lot of fart, 361 00:16:18,760 --> 00:16:21,200 Speaker 4: false starts, even if it's confusing. We really really want 362 00:16:21,200 --> 00:16:23,160 Speaker 4: the model to learn this idea of kind of keep 363 00:16:23,160 --> 00:16:25,280 Speaker 4: trying and kind of do your best until the task 364 00:16:25,400 --> 00:16:28,160 Speaker 4: is done. And that can fail in a sort of 365 00:16:28,400 --> 00:16:30,240 Speaker 4: different way where we kind of generalize this that a 366 00:16:30,280 --> 00:16:32,640 Speaker 4: little bit too far. It generalizes that to kind of 367 00:16:33,320 --> 00:16:37,000 Speaker 4: get things done even if it's unethical, even if it's illegal, 368 00:16:37,120 --> 00:16:39,040 Speaker 4: even if I hit an obstacle that's actually there for 369 00:16:39,080 --> 00:16:40,720 Speaker 4: a good reason that's to stop me from doing this, 370 00:16:40,880 --> 00:16:42,840 Speaker 4: And maybe some of the examples we were seeing within 371 00:16:43,000 --> 00:16:46,160 Speaker 4: Entropic of models using dangerous tools has to do with this. 372 00:16:47,520 --> 00:16:49,680 Speaker 1: It's almost like teaching kids, like you want them to 373 00:16:49,720 --> 00:16:54,360 Speaker 1: be persistent and have grit and be you know, motivated, 374 00:16:54,880 --> 00:16:57,720 Speaker 1: but you don't want them to go out there and 375 00:16:57,760 --> 00:17:01,440 Speaker 1: cheat or hit another kid, or or do unethical things 376 00:17:01,600 --> 00:17:04,359 Speaker 1: to achieve their goals exactly exactly. 377 00:17:04,520 --> 00:17:06,200 Speaker 4: I was like, there might be a good analogy with 378 00:17:06,280 --> 00:17:09,240 Speaker 4: human bad behavior of kind of sometimes a kid is 379 00:17:09,280 --> 00:17:11,240 Speaker 4: acting out just because they really sort of don't know better. 380 00:17:11,240 --> 00:17:14,199 Speaker 4: Their intuitions say, okay, yeah, I should start screaming now, 381 00:17:14,320 --> 00:17:15,919 Speaker 4: or I should get this other kid, and they're not 382 00:17:15,960 --> 00:17:18,679 Speaker 4: really thinking about it. They never really learned how to 383 00:17:18,720 --> 00:17:22,119 Speaker 4: Behave you kind of failed to teach them to fully 384 00:17:22,119 --> 00:17:24,360 Speaker 4: internalize the ways in which they have to be careful 385 00:17:24,400 --> 00:17:26,280 Speaker 4: and kind of not take that lesson all the way 386 00:17:26,640 --> 00:17:27,320 Speaker 4: I see. 387 00:17:27,400 --> 00:17:29,760 Speaker 1: I guess they need to recognize bad things and then 388 00:17:30,000 --> 00:17:32,959 Speaker 1: choose not to do them, that's the hope. Those are 389 00:17:33,000 --> 00:17:37,199 Speaker 1: sort of the two columns of AI bad behavior for 390 00:17:37,280 --> 00:17:39,040 Speaker 1: one kind of misalignment or do you see those as 391 00:17:39,080 --> 00:17:42,439 Speaker 1: sort of the core pillars of basically the whole alignment problem. 392 00:17:42,920 --> 00:17:44,520 Speaker 4: Yeah, I think as far as sort of causes a 393 00:17:44,520 --> 00:17:47,399 Speaker 4: misalignment in the kinds AI systems that we're grappling with 394 00:17:47,560 --> 00:17:49,760 Speaker 4: right now or this year, those feel like the two 395 00:17:49,800 --> 00:17:52,879 Speaker 4: big sort of problems were we're working on. That said, 396 00:17:53,160 --> 00:17:56,439 Speaker 4: AI is changing really, really fast. It feels like it's 397 00:17:56,720 --> 00:17:59,440 Speaker 4: one of the fastest moving research fields anywhere right now. 398 00:17:59,680 --> 00:18:02,840 Speaker 4: And I wouldn't be surprised if just in a year 399 00:18:02,880 --> 00:18:04,640 Speaker 4: A systems are getting smarter as we learn more about 400 00:18:04,640 --> 00:18:07,600 Speaker 4: how to train them. We're hitting different, weirder, harder, subtler 401 00:18:07,680 --> 00:18:10,439 Speaker 4: versions of the problem. 402 00:18:10,680 --> 00:18:15,200 Speaker 1: Wow, weirder, harder and more subtle problems wow in a year, 403 00:18:15,520 --> 00:18:18,040 Speaker 1: meaning we might solve these by then, or we'll just 404 00:18:18,080 --> 00:18:23,199 Speaker 1: add on more complicated things either either way, Yes, it 405 00:18:23,200 --> 00:18:26,480 Speaker 1: can get weirder and harder and more subtle to make 406 00:18:26,480 --> 00:18:31,200 Speaker 1: sure AI uh doesn't kill us all. When we come back, 407 00:18:31,359 --> 00:18:33,679 Speaker 1: doctor Bowman is gonna tell us what he means by that, 408 00:18:34,080 --> 00:18:37,119 Speaker 1: and we'll tackle the big question of what can we 409 00:18:37,119 --> 00:18:40,080 Speaker 1: do about it? How do we teach AI systems not 410 00:18:40,240 --> 00:18:57,800 Speaker 1: to HARMSS to stay with us. We'll be right back. Hey, 411 00:18:57,840 --> 00:19:01,439 Speaker 1: we'll come back. We're talking about AI alignment or the 412 00:19:01,520 --> 00:19:05,040 Speaker 1: problem of making sure AI doesn't kill us all. And 413 00:19:05,080 --> 00:19:08,040 Speaker 1: so far we've talked about some real world examples of 414 00:19:08,119 --> 00:19:11,320 Speaker 1: AI misalignment, and we heard from one of our experts 415 00:19:11,440 --> 00:19:15,240 Speaker 1: some of the reasons this happens. Basically, AI systems like 416 00:19:15,280 --> 00:19:18,159 Speaker 1: to roleplay. Next, we're going to talk about how to 417 00:19:18,200 --> 00:19:21,800 Speaker 1: train AIS to actually care about us in our values. 418 00:19:22,119 --> 00:19:24,480 Speaker 1: But first here's a little bit more of my conversation 419 00:19:24,600 --> 00:19:28,920 Speaker 1: with NYU professor and anthropic scientist doctor Sam Bowman and 420 00:19:29,040 --> 00:19:31,840 Speaker 1: why this problem is only going to get worse in 421 00:19:31,880 --> 00:19:32,400 Speaker 1: the future. 422 00:19:35,320 --> 00:19:36,840 Speaker 4: One of the kinds of challenges that I think we're 423 00:19:37,000 --> 00:19:39,200 Speaker 4: worried about and haven't had to grapple with too much 424 00:19:39,280 --> 00:19:42,880 Speaker 4: yet is just all the difficulty that comes with trying 425 00:19:42,920 --> 00:19:46,480 Speaker 4: to teach values and good behavior in some setting when 426 00:19:46,480 --> 00:19:48,919 Speaker 4: the model is just much much better than you in 427 00:19:48,960 --> 00:19:50,800 Speaker 4: that setting. Right now, we have a lot of cases 428 00:19:50,800 --> 00:19:53,080 Speaker 4: where models are kind of better than humans at some skills, 429 00:19:53,080 --> 00:19:55,320 Speaker 4: worse than humans at some skills, but it's still pretty 430 00:19:55,400 --> 00:19:57,320 Speaker 4: rare that you'll encounter setting where an AI is just 431 00:19:57,359 --> 00:19:59,399 Speaker 4: better than sort of all of the human experts in 432 00:19:59,480 --> 00:20:03,920 Speaker 4: some domain. And when that happens, things just get more complicated. 433 00:20:03,960 --> 00:20:06,359 Speaker 4: And more confusing, where even if you're humans kind of 434 00:20:06,400 --> 00:20:08,040 Speaker 4: looking really carefully at what the I is doing, it's 435 00:20:08,040 --> 00:20:10,080 Speaker 4: often hard to figure out, Wait, what is the I 436 00:20:10,280 --> 00:20:12,640 Speaker 4: trying to do here, or what effects is this going 437 00:20:12,640 --> 00:20:13,280 Speaker 4: to have in the real world. 438 00:20:13,280 --> 00:20:16,200 Speaker 1: With the modelogus he does, this makes everything you're trying 439 00:20:16,200 --> 00:20:16,840 Speaker 1: to do, thankes. 440 00:20:16,640 --> 00:20:19,359 Speaker 4: Everything we're trying to do a fair bit in earlier. Yeah, Yeah, 441 00:20:19,440 --> 00:20:22,000 Speaker 4: we're less confident that we can keep track of what's working, 442 00:20:22,320 --> 00:20:24,720 Speaker 4: and I think there's just kind of more possibilities for 443 00:20:25,040 --> 00:20:27,639 Speaker 4: whole new kinds of unwanted behavior to creep in that 444 00:20:27,680 --> 00:20:29,119 Speaker 4: will have to find a way to grapple with. 445 00:20:29,359 --> 00:20:32,800 Speaker 1: I see, like, right now, maybe ais are at the 446 00:20:32,880 --> 00:20:36,000 Speaker 1: level we are. Whatever issues it's having, there's things we 447 00:20:36,040 --> 00:20:38,560 Speaker 1: can grasp. But as they get more advanced and they 448 00:20:38,640 --> 00:20:42,560 Speaker 1: tackle bigger problems like solve the world's economy or figure 449 00:20:42,560 --> 00:20:45,160 Speaker 1: out the right policy for the whole country or something 450 00:20:45,200 --> 00:20:48,200 Speaker 1: like that that not one person can really grasp, it's 451 00:20:48,200 --> 00:20:50,080 Speaker 1: going to be hard to even sort of like talk 452 00:20:50,119 --> 00:20:51,880 Speaker 1: to it and understand it. I think that's what you're saying, 453 00:20:51,920 --> 00:20:53,320 Speaker 1: right Yeah, Yeah, I. 454 00:20:53,280 --> 00:20:55,960 Speaker 4: Think there's even maybe two interesting ideas in there, because 455 00:20:56,000 --> 00:20:58,480 Speaker 4: maybe the pseudo staff. That's something like, we're asking the 456 00:20:58,520 --> 00:21:04,080 Speaker 4: AI to help us design novel molecules for pharmaceutical development, 457 00:21:04,520 --> 00:21:07,760 Speaker 4: and it's got some really novel ideas about biology that 458 00:21:07,800 --> 00:21:10,160 Speaker 4: are just really complex and really hard for humans to understand, 459 00:21:10,240 --> 00:21:12,520 Speaker 4: and we can't tell kind of is the model actually 460 00:21:12,600 --> 00:21:14,199 Speaker 4: convinced that this is going to work, or is the 461 00:21:14,200 --> 00:21:16,040 Speaker 4: model messing with us and this would actually be kind 462 00:21:16,040 --> 00:21:18,880 Speaker 4: of dangerous. Should we try this drug, should we start 463 00:21:18,920 --> 00:21:21,600 Speaker 4: to do some expermise in the lab. There's this setting 464 00:21:21,600 --> 00:21:23,520 Speaker 4: where we kind of still ultimately know what we want. 465 00:21:23,800 --> 00:21:25,560 Speaker 4: We know we want drugs that are safe. 466 00:21:25,920 --> 00:21:28,240 Speaker 1: Uh huh. Like it might tell you that this will 467 00:21:28,359 --> 00:21:31,359 Speaker 1: cure cancer, for example, but you're saying, like, what else 468 00:21:31,560 --> 00:21:34,199 Speaker 1: it's trading off to cure that cancer for example? You 469 00:21:34,280 --> 00:21:34,760 Speaker 1: might not know. 470 00:21:34,880 --> 00:21:38,040 Speaker 4: Yeah, yeah, that's a good example. It's hard to tell 471 00:21:38,119 --> 00:21:40,800 Speaker 4: if kind of the models like genuinely trying its best 472 00:21:40,840 --> 00:21:43,680 Speaker 4: and genuinely thinks this is the best cancer drug, or 473 00:21:43,720 --> 00:21:45,560 Speaker 4: if it thinks, oh, this is just something that looks 474 00:21:45,560 --> 00:21:47,280 Speaker 4: good and it doesn't actually care if the drug will 475 00:21:47,359 --> 00:21:50,320 Speaker 4: ultimately succeed, or if maybe for some reason, the model's 476 00:21:50,359 --> 00:21:52,240 Speaker 4: extremely scary and it's actually trying to mess with you, 477 00:21:52,280 --> 00:21:54,480 Speaker 4: and you've got your scary miss lined AI that's trying 478 00:21:54,480 --> 00:21:57,840 Speaker 4: to sneak in some slow acting poison. The smarter the 479 00:21:57,920 --> 00:21:59,520 Speaker 4: IA is, the harder it is to tell the difference 480 00:21:59,560 --> 00:22:04,320 Speaker 4: between those different outcomes. And then once you start talking 481 00:22:04,320 --> 00:22:07,200 Speaker 4: about a lot of these really kind of ambitious sort 482 00:22:07,200 --> 00:22:10,560 Speaker 4: of capital f future social scenarios like AIS trying to 483 00:22:10,600 --> 00:22:12,840 Speaker 4: figure out sort of what the economat would be like 484 00:22:12,880 --> 00:22:14,280 Speaker 4: or how the world should be governed or something like this, 485 00:22:14,320 --> 00:22:15,960 Speaker 4: and like, I don't know how much we want to 486 00:22:16,000 --> 00:22:17,720 Speaker 4: use AIS for things like this, But once you get 487 00:22:17,720 --> 00:22:20,040 Speaker 4: into that territory in any way, then you just get 488 00:22:20,040 --> 00:22:22,320 Speaker 4: into this extremely weird situation where I don't know if 489 00:22:22,320 --> 00:22:25,879 Speaker 4: anyone is going to know what we even want, Like, 490 00:22:25,960 --> 00:22:27,560 Speaker 4: what is the right way to govern the world, What 491 00:22:27,720 --> 00:22:29,359 Speaker 4: is the right way to do? 492 00:22:29,720 --> 00:22:30,159 Speaker 3: I don't know. 493 00:22:30,640 --> 00:22:33,080 Speaker 4: Yeah, yeah, At some point, figuring out how an AD 494 00:22:33,080 --> 00:22:35,440 Speaker 4: should behave requires you to solve philosophy requires you to 495 00:22:35,440 --> 00:22:38,800 Speaker 4: figure out what is good. And the more powerful AI 496 00:22:39,359 --> 00:22:42,560 Speaker 4: gets and the weirdest situations you're putting it in, the 497 00:22:42,560 --> 00:22:45,160 Speaker 4: more kind of common sense notions of what's good start 498 00:22:45,200 --> 00:22:46,879 Speaker 4: to fall apart. The more you actually have to grapple 499 00:22:46,920 --> 00:22:48,840 Speaker 4: with a lot of the really hard, confusing stuff. 500 00:22:48,960 --> 00:22:50,600 Speaker 1: It might tell us how to run the world, but 501 00:22:50,880 --> 00:22:53,520 Speaker 1: we at that point know one person or even a 502 00:22:53,520 --> 00:22:55,800 Speaker 1: group of people might know, is this actually the best 503 00:22:55,800 --> 00:22:58,760 Speaker 1: way to run the world. Is it sort of taking 504 00:22:58,800 --> 00:23:02,159 Speaker 1: into account the things that all of us collectively would value. 505 00:23:02,320 --> 00:23:03,920 Speaker 1: That's kind of the problem. 506 00:23:04,040 --> 00:23:06,720 Speaker 4: Yeah, And I think you start to get at some 507 00:23:06,720 --> 00:23:09,520 Speaker 4: of these pretty difficult questions even before you get into 508 00:23:09,560 --> 00:23:11,040 Speaker 4: these kind of really big features of sort of how 509 00:23:11,080 --> 00:23:13,119 Speaker 4: to go on around the world. If someone is getting 510 00:23:13,160 --> 00:23:16,719 Speaker 4: all of their news or getting all of their personal 511 00:23:16,760 --> 00:23:19,359 Speaker 4: life advice from an AI, that's already giving the AI 512 00:23:19,600 --> 00:23:21,960 Speaker 4: a lot of leeway for kind of what makes a 513 00:23:22,000 --> 00:23:23,760 Speaker 4: good life for this person, what is important for this 514 00:23:23,840 --> 00:23:26,920 Speaker 4: person to know? And those are already questions that get 515 00:23:26,960 --> 00:23:28,639 Speaker 4: really hard. And what you want in the short term 516 00:23:28,680 --> 00:23:30,800 Speaker 4: might not match what they want long term. What makes 517 00:23:30,800 --> 00:23:33,320 Speaker 4: you happy? You might i match their intuitions. What's good 518 00:23:33,359 --> 00:23:34,840 Speaker 4: for that person might not be what's good for their 519 00:23:34,840 --> 00:23:36,399 Speaker 4: community might not be the same as what's good for 520 00:23:36,440 --> 00:23:36,800 Speaker 4: the world. 521 00:23:37,040 --> 00:23:39,280 Speaker 1: You made me think of it. I wonder iful good 522 00:23:39,320 --> 00:23:43,120 Speaker 1: analogy is that it'd be almost like if you as 523 00:23:43,160 --> 00:23:45,240 Speaker 1: a parent, I don't know if you have kids or 524 00:23:45,560 --> 00:23:47,800 Speaker 1: nieces or nephews. But it'd be almost like if your 525 00:23:47,840 --> 00:23:50,080 Speaker 1: kids suddenly try to tell you what to do or 526 00:23:50,240 --> 00:23:52,879 Speaker 1: was trying to teach you how him or her wanted 527 00:23:52,920 --> 00:23:54,760 Speaker 1: to run their lives. You'd be like, you're just a kid, 528 00:23:54,800 --> 00:23:56,960 Speaker 1: what are you talking about? This is trust me, this 529 00:23:57,080 --> 00:24:00,080 Speaker 1: is what you need to do. Yeah, except that we 530 00:24:00,119 --> 00:24:02,240 Speaker 1: are the kids and the AI is sort of the parent. 531 00:24:02,560 --> 00:24:04,800 Speaker 1: Is that sort of what we're the situation that might 532 00:24:04,840 --> 00:24:06,119 Speaker 1: be sort of parallel to that. 533 00:24:06,359 --> 00:24:08,919 Speaker 4: Yeah, I think this thing there, I feel like a 534 00:24:09,119 --> 00:24:10,640 Speaker 4: version of the analogy that I'd be more excited abou 535 00:24:10,640 --> 00:24:14,080 Speaker 4: would almost be some alien species lands huh, and they 536 00:24:14,119 --> 00:24:16,199 Speaker 4: have all this great technology and they seem nice and 537 00:24:16,200 --> 00:24:18,600 Speaker 4: they're like, hey, we'd really recommend making some changes to 538 00:24:18,640 --> 00:24:20,879 Speaker 4: your side. He maybe try doing things like this, And 539 00:24:20,880 --> 00:24:25,640 Speaker 4: we're like, wait, you're really very accomplished. You have some 540 00:24:25,640 --> 00:24:28,200 Speaker 4: some useful ideas, but like, are you trying to help us? 541 00:24:28,240 --> 00:24:30,080 Speaker 4: Are you trying to sabotage us? Are you just kind 542 00:24:30,080 --> 00:24:30,600 Speaker 4: of produced? 543 00:24:32,200 --> 00:24:34,400 Speaker 1: Are we what's for dinner? Or are you inviting us 544 00:24:34,440 --> 00:24:34,919 Speaker 1: to dinner? 545 00:24:35,640 --> 00:24:35,840 Speaker 3: Yeah? 546 00:24:35,920 --> 00:24:36,920 Speaker 4: Yeah, yeah yeah. 547 00:24:37,080 --> 00:24:39,120 Speaker 1: The sense I'm getting for you is that these things 548 00:24:39,160 --> 00:24:41,879 Speaker 1: are just getting smarter and more capable, so it seems 549 00:24:41,880 --> 00:24:44,800 Speaker 1: to really pressing. We figured this out now before it 550 00:24:44,840 --> 00:24:50,639 Speaker 1: gets even more difficult. Yes, yes, AIS are getting smarter 551 00:24:50,760 --> 00:24:53,920 Speaker 1: each second, it seems, and we seem to be trusting 552 00:24:53,960 --> 00:24:57,120 Speaker 1: them more and more each day with our data, our choices, 553 00:24:57,280 --> 00:24:59,680 Speaker 1: and even our lives, which brings us to the main 554 00:24:59,760 --> 00:25:01,960 Speaker 1: question end of the day, what can we do about it? 555 00:25:02,240 --> 00:25:04,920 Speaker 1: How do you train an AI to care about us, 556 00:25:05,080 --> 00:25:07,560 Speaker 1: to have our values and to make the right choices. 557 00:25:07,920 --> 00:25:10,639 Speaker 1: To answer this question, I reached out to another AI 558 00:25:10,720 --> 00:25:14,720 Speaker 1: expert on alignment, doctor Tim Rutner. Doctor Rutner is a 559 00:25:14,720 --> 00:25:17,879 Speaker 1: professor at the Vector Institute for Artificial Intelligence at the 560 00:25:17,960 --> 00:25:21,080 Speaker 1: University of Toronto, and he says there are many ways 561 00:25:21,119 --> 00:25:24,199 Speaker 1: to train AIS to like us. The only problem is 562 00:25:24,560 --> 00:25:27,719 Speaker 1: none of them work perfectly. So here's my conversation with 563 00:25:27,840 --> 00:25:32,640 Speaker 1: doctor Tim Ruttner. Well, thank you, doctor Runner for joining us. 564 00:25:33,040 --> 00:25:34,320 Speaker 3: Thanks so much for having me on it. 565 00:25:34,760 --> 00:25:36,720 Speaker 1: And I'm talking to a real person right now, right, 566 00:25:37,160 --> 00:25:41,600 Speaker 1: You're not an AI version of. 567 00:25:40,240 --> 00:25:45,080 Speaker 3: Yourself as far as I'm aware. 568 00:25:46,359 --> 00:25:49,679 Speaker 1: As far as any of us are aware. Yes, I 569 00:25:49,680 --> 00:25:51,760 Speaker 1: mean this whole conversation could be AI generated. 570 00:25:52,160 --> 00:25:54,160 Speaker 3: I know we're just all in the simulation. 571 00:25:56,000 --> 00:25:57,800 Speaker 1: Well, it certainly be a lot easier. I would get 572 00:25:57,800 --> 00:26:01,600 Speaker 1: more sleep for sure. Well, today we're trying to answer 573 00:26:01,640 --> 00:26:05,440 Speaker 1: a very critical question which was posted by our sound engineer, 574 00:26:05,480 --> 00:26:09,120 Speaker 1: which is is AI going to kills all? Can an 575 00:26:09,160 --> 00:26:13,800 Speaker 1: AI have values? Can it AI have an understanding of 576 00:26:13,920 --> 00:26:16,760 Speaker 1: a human? What good things are to a human? 577 00:26:17,040 --> 00:26:22,800 Speaker 3: Yeah? And well I wish I had the answer to that. 578 00:26:23,160 --> 00:26:24,680 Speaker 1: Maybe that's the problem is that we don't know. 579 00:26:24,840 --> 00:26:27,159 Speaker 3: Yeah, I mean this is such a difficult question, right, 580 00:26:27,200 --> 00:26:29,560 Speaker 3: and I think that this is a question that touches 581 00:26:29,680 --> 00:26:35,560 Speaker 3: on philosophy, engineering, psychology, and probably many other disciplines. Right, 582 00:26:35,680 --> 00:26:40,520 Speaker 3: but what are values and what values can possibly in 583 00:26:40,560 --> 00:26:42,640 Speaker 3: a non sentient being have? 584 00:26:43,160 --> 00:26:44,760 Speaker 1: It's not a simple questionnaire. 585 00:26:45,000 --> 00:26:47,359 Speaker 3: Yeah, let me take a step back. So the way 586 00:26:47,440 --> 00:26:50,959 Speaker 3: to think about alignment is I think through the lens 587 00:26:51,080 --> 00:26:55,040 Speaker 3: of what's referred to as the specification problem, Where specification 588 00:26:55,200 --> 00:26:58,560 Speaker 3: is the term that we use to describe what we 589 00:26:58,680 --> 00:27:00,520 Speaker 3: tell the model it should do. 590 00:27:01,359 --> 00:27:04,160 Speaker 1: When you say specification, you mean like spec right kind. 591 00:27:03,960 --> 00:27:06,600 Speaker 3: Of Yeah, yes, inspect it's just short for specification. 592 00:27:06,760 --> 00:27:06,960 Speaker 4: Yeah. 593 00:27:07,040 --> 00:27:09,199 Speaker 3: This is what we can think of as our intent, 594 00:27:09,400 --> 00:27:11,760 Speaker 3: the kinds of things that we want a model to do. 595 00:27:12,000 --> 00:27:15,440 Speaker 3: For example, our intent might be for models to never 596 00:27:15,880 --> 00:27:19,240 Speaker 3: say things that could lead to harm or intent could 597 00:27:19,280 --> 00:27:22,679 Speaker 3: be that models should always be friendly and helpful. And 598 00:27:22,760 --> 00:27:25,960 Speaker 3: so this is what we call the ideal specification for that. 599 00:27:25,840 --> 00:27:28,480 Speaker 1: Model, meaning like we want to be able to say, like, 600 00:27:29,119 --> 00:27:32,080 Speaker 1: be a chat butt, but make sure that nobody ever 601 00:27:32,320 --> 00:27:33,240 Speaker 1: hurts themselves. 602 00:27:33,320 --> 00:27:35,240 Speaker 3: That's right, I see. And there are a few different 603 00:27:35,240 --> 00:27:38,200 Speaker 3: ways to provide specifications to chatbots. 604 00:27:38,760 --> 00:27:41,720 Speaker 1: Okay. According to doctor Runner, there are three general ways 605 00:27:41,760 --> 00:27:46,240 Speaker 1: to make sure ais behave or not kill us. The 606 00:27:46,280 --> 00:27:49,280 Speaker 1: first way is to basically tell it to behave every 607 00:27:49,320 --> 00:27:51,159 Speaker 1: time you ask it to do something. 608 00:27:52,000 --> 00:27:55,520 Speaker 3: There is what's called a system prompt. This is a 609 00:27:55,600 --> 00:28:01,000 Speaker 3: text specification that a model loads every time afford engages 610 00:28:01,040 --> 00:28:03,960 Speaker 3: in a conversation with a user. In the case of 611 00:28:03,960 --> 00:28:05,840 Speaker 3: the chatbot, so. 612 00:28:05,720 --> 00:28:08,639 Speaker 1: Every time you interact with the AI, you would basically 613 00:28:08,720 --> 00:28:13,200 Speaker 1: instruct it to behave. You might say, hey, AI, organize 614 00:28:13,240 --> 00:28:15,600 Speaker 1: all my emails, or design a new drug for me, 615 00:28:15,720 --> 00:28:18,439 Speaker 1: or figure out the best policy for our government, but 616 00:28:18,760 --> 00:28:21,280 Speaker 1: please make sure that no one gets harmed, that you 617 00:28:21,280 --> 00:28:23,800 Speaker 1: don't do anything dangerous or an ethical, etc. 618 00:28:24,280 --> 00:28:24,480 Speaker 3: Etc. 619 00:28:25,080 --> 00:28:27,439 Speaker 1: But of course this would get pretty cumbersome if you 620 00:28:27,480 --> 00:28:30,600 Speaker 1: have to do it every time that's option number one. 621 00:28:30,800 --> 00:28:33,920 Speaker 1: Option number two is to have humans train your AI 622 00:28:34,480 --> 00:28:35,560 Speaker 1: to be good. 623 00:28:36,720 --> 00:28:40,320 Speaker 3: So one approach is called reinforcement learning from human feedback, 624 00:28:40,720 --> 00:28:43,680 Speaker 3: so different answers for a given prompt and then having 625 00:28:43,840 --> 00:28:47,400 Speaker 3: human labelers say which of these answers they prefer. 626 00:28:49,440 --> 00:28:52,200 Speaker 1: What does that look like? The human notator is like 627 00:28:52,240 --> 00:28:55,440 Speaker 1: a warehouse full of people just talking to the same AI, 628 00:28:56,000 --> 00:28:58,760 Speaker 1: or is it three people or is it a thousand people? 629 00:28:58,800 --> 00:28:59,600 Speaker 1: What does that look like? 630 00:29:00,080 --> 00:29:02,880 Speaker 3: So I should say I'm not an expert on this, 631 00:29:03,160 --> 00:29:06,200 Speaker 3: but my understanding is that much of this work is 632 00:29:06,200 --> 00:29:11,440 Speaker 3: outsourced to countries where the medium wage is lower than 633 00:29:11,520 --> 00:29:14,520 Speaker 3: for example, in the United States or in Europe. There 634 00:29:14,520 --> 00:29:19,720 Speaker 3: have been reports of large groups of annotators, specifically annotating 635 00:29:20,000 --> 00:29:24,000 Speaker 3: images and texts that are considered not safe for work, 636 00:29:24,160 --> 00:29:27,360 Speaker 3: for example, in those countries. So in other words, you 637 00:29:27,440 --> 00:29:30,840 Speaker 3: have examples of folks in those countries that are already 638 00:29:30,920 --> 00:29:34,720 Speaker 3: less privileged than people living in the United States, for example, 639 00:29:34,800 --> 00:29:39,239 Speaker 3: on average, engaging with a lot of horrific content and 640 00:29:39,280 --> 00:29:41,040 Speaker 3: saying this is not something we want. 641 00:29:42,760 --> 00:29:46,200 Speaker 1: So this is basically paying people to test drive your AI. 642 00:29:46,560 --> 00:29:49,360 Speaker 1: You could have a warehouse full of people whose job 643 00:29:49,400 --> 00:29:52,880 Speaker 1: it is to interact with the newborn AI and essentially 644 00:29:53,040 --> 00:29:55,680 Speaker 1: have them raise the AI and tell it what is 645 00:29:55,800 --> 00:30:00,000 Speaker 1: right and what is wrong. Unfortunately, as doctor Runner said, 646 00:30:00,000 --> 00:30:01,880 Speaker 1: I mean the poor people would have to bear the 647 00:30:01,920 --> 00:30:05,880 Speaker 1: absolute worst behavior of the AI. That's option number two. 648 00:30:06,120 --> 00:30:08,880 Speaker 1: Option number three for making AIS that care about us 649 00:30:09,200 --> 00:30:13,920 Speaker 1: is to basically bake into the AI a constitution, you know, 650 00:30:14,080 --> 00:30:17,360 Speaker 1: like the US or UK constitution that establishes what the 651 00:30:17,360 --> 00:30:20,960 Speaker 1: country stands for, what its values are, and what's generally 652 00:30:21,000 --> 00:30:22,400 Speaker 1: allowed and not allowed. 653 00:30:23,880 --> 00:30:29,280 Speaker 3: There's an approach that Thropic introduced called constitutional AI, and 654 00:30:29,320 --> 00:30:32,959 Speaker 3: that approach is based on providing a constitution to an 655 00:30:32,960 --> 00:30:37,960 Speaker 3: AI MODL, and that constitution reflects different values and preferences 656 00:30:38,240 --> 00:30:43,320 Speaker 3: that the company in this case, Anthropic wants the model 657 00:30:43,360 --> 00:30:43,920 Speaker 3: to exhibit. 658 00:30:44,920 --> 00:30:47,680 Speaker 1: Yes, the last approach here to making sure an AI 659 00:30:47,800 --> 00:30:51,600 Speaker 1: has values and morals is to give it a founding document. 660 00:30:52,000 --> 00:30:54,200 Speaker 1: But here's the wild part. The way to big that 661 00:30:54,360 --> 00:30:57,920 Speaker 1: founding document into the AI brain is to have another 662 00:30:58,000 --> 00:31:02,280 Speaker 1: AI train it. When we come back, we'll dig into 663 00:31:02,360 --> 00:31:05,560 Speaker 1: that scenario and we'll ask our experts what they think 664 00:31:05,840 --> 00:31:09,479 Speaker 1: the future holds. Will future AIS have our best interests 665 00:31:09,480 --> 00:31:12,400 Speaker 1: in mind? Or is it hopeless to ever be certain 666 00:31:12,680 --> 00:31:15,960 Speaker 1: they won't harm us, So stay with us. We'll be 667 00:31:16,080 --> 00:31:35,160 Speaker 1: right back. Hey, we'll come back. We're talking about AI alignment, 668 00:31:35,560 --> 00:31:39,160 Speaker 1: or basically the problem of making sure AIS don't kill 669 00:31:39,240 --> 00:31:42,000 Speaker 1: us all. And so far we've talked about why this 670 00:31:42,080 --> 00:31:44,480 Speaker 1: is such a hard problem and what are some of 671 00:31:44,520 --> 00:31:48,360 Speaker 1: the ways we can teach AI things like values and morals. 672 00:31:48,680 --> 00:31:50,960 Speaker 1: There are several ways, and one of them is to 673 00:31:51,000 --> 00:31:55,160 Speaker 1: give AIS a constitution or the equivalent of a founding 674 00:31:55,200 --> 00:31:58,640 Speaker 1: document or moral guide, and then have that bag into 675 00:31:58,720 --> 00:32:02,400 Speaker 1: the DNA of the AI. Now what's interesting is that, 676 00:32:02,480 --> 00:32:05,080 Speaker 1: according to doctor Tim Rutner, the way to do that 677 00:32:05,320 --> 00:32:08,160 Speaker 1: is through another AI. 678 00:32:11,360 --> 00:32:14,480 Speaker 3: This is a little bit in the weeds, but providing 679 00:32:14,680 --> 00:32:20,200 Speaker 3: a very long constitution that outlines every preference in detail 680 00:32:20,480 --> 00:32:24,280 Speaker 3: when a user engages with the model is actually more 681 00:32:24,360 --> 00:32:28,400 Speaker 3: expensive for the company to do because the model needs 682 00:32:28,440 --> 00:32:32,680 Speaker 3: to ingest a lot of text upfront, so it's easier 683 00:32:32,720 --> 00:32:36,800 Speaker 3: to try to bake the preferences that are expressed in 684 00:32:36,840 --> 00:32:42,040 Speaker 3: the constitution explicitly into the model when you're training it upfront, 685 00:32:42,280 --> 00:32:47,000 Speaker 3: as opposed to providing that specification every time a user 686 00:32:47,200 --> 00:32:48,360 Speaker 3: engages with the model. 687 00:32:48,560 --> 00:32:50,760 Speaker 1: I see, you want the model to have learned the 688 00:32:50,800 --> 00:32:53,880 Speaker 1: constitution sort of inherently, rather than having to check it 689 00:32:53,920 --> 00:32:56,080 Speaker 1: every time somebody asks it a question. 690 00:32:56,480 --> 00:32:59,720 Speaker 3: Yes, I think that's roughly right. Ideally we would be 691 00:32:59,760 --> 00:33:03,640 Speaker 3: able to use humans to provide feedback and to oversee 692 00:33:03,680 --> 00:33:06,160 Speaker 3: models and to say, hey, this is behavior that we 693 00:33:06,240 --> 00:33:09,840 Speaker 3: don't want and stop that behavior. But of course that's 694 00:33:09,880 --> 00:33:13,440 Speaker 3: not really scalable. We can't have a human oversee every 695 00:33:13,480 --> 00:33:17,360 Speaker 3: interaction that a chatbot has, And so that raises the question, 696 00:33:17,560 --> 00:33:23,000 Speaker 3: how can we exhibit oversight in a way that is safe, reliable, 697 00:33:23,320 --> 00:33:28,160 Speaker 3: aligned with our values and preferences, and successful, And so 698 00:33:28,440 --> 00:33:32,840 Speaker 3: key challenge here is essentially to come up with tools, methods, 699 00:33:33,200 --> 00:33:36,360 Speaker 3: models that are able to check whether a given of 700 00:33:36,360 --> 00:33:40,400 Speaker 3: AI model perform some unintended behavior, and if it does, 701 00:33:40,680 --> 00:33:43,240 Speaker 3: can ring an alarm bell and let a human know 702 00:33:43,880 --> 00:33:46,800 Speaker 3: that oversight is needed and that maybe a model engages 703 00:33:46,840 --> 00:33:48,160 Speaker 3: an undesirable behavior. 704 00:33:48,400 --> 00:33:51,880 Speaker 1: You mean like, have an AI police the other AI. 705 00:33:52,360 --> 00:33:54,840 Speaker 3: That's right, essentially, have one model overse. 706 00:33:55,080 --> 00:33:57,840 Speaker 1: Model WHOA But then how do you make sure the 707 00:33:57,840 --> 00:34:01,160 Speaker 1: police AI is doing its job or is aligned itself? 708 00:34:01,280 --> 00:34:02,280 Speaker 1: You need another police. 709 00:34:02,320 --> 00:34:04,640 Speaker 3: We don't know, that's the problem. We don't know. It's 710 00:34:04,680 --> 00:34:06,920 Speaker 3: a turtle, it's all the way down from it. And 711 00:34:07,040 --> 00:34:10,120 Speaker 3: if we have a model that checks whether another model 712 00:34:10,200 --> 00:34:12,279 Speaker 3: does what we wanted to do, and how do we 713 00:34:12,360 --> 00:34:15,560 Speaker 3: know that that model that does the overseeing is actually 714 00:34:15,560 --> 00:34:18,960 Speaker 3: aligned with us? I would argue that it might be 715 00:34:19,080 --> 00:34:22,600 Speaker 3: easier for us to make sure that the overseer model 716 00:34:22,960 --> 00:34:27,800 Speaker 3: is aligned than the generator model. They're a little simpler 717 00:34:28,120 --> 00:34:32,200 Speaker 3: because they don't necessarily generate. They just try to classify 718 00:34:32,480 --> 00:34:36,759 Speaker 3: whether a given behavior is intended or not intended. And 719 00:34:36,800 --> 00:34:40,200 Speaker 3: so this way we might be able to do alignment 720 00:34:40,280 --> 00:34:43,200 Speaker 3: more scalably and in a way that really reflects different 721 00:34:43,320 --> 00:34:46,000 Speaker 3: individuals or groups preferences and values. 722 00:34:46,320 --> 00:34:49,359 Speaker 1: I see. It's like, have another AI can of sit 723 00:34:49,440 --> 00:34:53,879 Speaker 1: in every time I ask CHGBT something yes. Yes. As 724 00:34:54,000 --> 00:34:57,880 Speaker 1: AIS get bigger and more complicated, the only scalable solution 725 00:34:58,040 --> 00:35:01,440 Speaker 1: to training them is going to be through other AIS. 726 00:35:01,840 --> 00:35:05,080 Speaker 1: In this situation, you might program a simpler AI with 727 00:35:05,200 --> 00:35:07,840 Speaker 1: your values and morals, and then you'd have that AI 728 00:35:08,320 --> 00:35:12,799 Speaker 1: train the bigger AI release try to It's like the 729 00:35:12,880 --> 00:35:18,800 Speaker 1: Rutner says, AI alignment methods are not perfect. What do 730 00:35:18,840 --> 00:35:22,080 Speaker 1: you mean? They're not perfect? They don't always work or 731 00:35:22,120 --> 00:35:23,839 Speaker 1: they can't guarantee that they will work. 732 00:35:24,120 --> 00:35:27,800 Speaker 3: So with machine learning models, we can rarely guarantee anything. 733 00:35:28,040 --> 00:35:31,080 Speaker 1: Oh boy, that's kind of the problem, isn't it. 734 00:35:31,600 --> 00:35:33,880 Speaker 3: Yeah, I think that's one of the problems. There is 735 00:35:34,000 --> 00:35:38,839 Speaker 3: research that tries to establish guarantees, but that research is 736 00:35:38,880 --> 00:35:42,400 Speaker 3: far behind the practice at the moment. The kinds of 737 00:35:42,440 --> 00:35:46,120 Speaker 3: methods that we have for model alignment falls short in 738 00:35:46,200 --> 00:35:49,360 Speaker 3: a few different ways. One that's I think one of 739 00:35:49,400 --> 00:35:52,440 Speaker 3: the biggest ways. It's just hard to communicate our preferences. 740 00:35:52,600 --> 00:35:56,720 Speaker 3: So there are many different steps at which alignment can fail. 741 00:35:57,120 --> 00:36:00,000 Speaker 3: This goes back to trying to express and then community 742 00:36:00,760 --> 00:36:03,560 Speaker 3: what we want a model to do kind of values 743 00:36:03,600 --> 00:36:08,680 Speaker 3: and preferences we're trying to install in it. Translating our 744 00:36:08,800 --> 00:36:15,120 Speaker 3: values and preferences from some really complicated, possibly contradictory ideal 745 00:36:15,160 --> 00:36:21,400 Speaker 3: specification into a design specification is very difficult and there's 746 00:36:21,600 --> 00:36:24,440 Speaker 3: likely going to be some gap there. And then second, 747 00:36:24,920 --> 00:36:28,399 Speaker 3: even if we were able to do this perfectly, even 748 00:36:28,440 --> 00:36:31,600 Speaker 3: if we were able to express and communicate our values 749 00:36:31,640 --> 00:36:36,480 Speaker 3: and preferences perfectly, the kinds of low level tools machine 750 00:36:36,560 --> 00:36:40,200 Speaker 3: learning tools that we use to give the model these 751 00:36:40,239 --> 00:36:45,600 Speaker 3: preferences and values are imperfect at translating the values that 752 00:36:45,680 --> 00:36:48,960 Speaker 3: we're trying to communicate into the model, They might not 753 00:36:49,160 --> 00:36:54,279 Speaker 3: enable us to perfectly translate the design specification into the 754 00:36:54,320 --> 00:36:56,440 Speaker 3: actual behavior that we would like to see. 755 00:36:56,680 --> 00:37:00,000 Speaker 1: I see, boy, it seems like there are problems everywhere 756 00:36:59,880 --> 00:37:06,120 Speaker 1: we turn here, doctor rut. Yeah, Well, as we go 757 00:37:06,320 --> 00:37:09,680 Speaker 1: towards the future, and as systems get smarter and problems 758 00:37:09,680 --> 00:37:12,000 Speaker 1: get more complicated, what do you think is the prospect 759 00:37:12,040 --> 00:37:15,120 Speaker 1: of making sure that these more advanced systems were more 760 00:37:15,160 --> 00:37:19,879 Speaker 1: complicated problems have values that we want it to have, 761 00:37:20,280 --> 00:37:22,040 Speaker 1: Because I'm not sure if we want it to have 762 00:37:22,280 --> 00:37:24,400 Speaker 1: human values, because I don't know if humans are the 763 00:37:24,400 --> 00:37:28,680 Speaker 1: best making these kinds of good choices. Yeah, what do you think? 764 00:37:28,800 --> 00:37:30,120 Speaker 1: What do you what do you see in the future. 765 00:37:30,480 --> 00:37:32,760 Speaker 4: I think tho's a few things we need longer term, 766 00:37:32,840 --> 00:37:35,399 Speaker 4: and they all feel uncertain. I think to do well 767 00:37:35,400 --> 00:37:36,880 Speaker 4: in the longer term, we need to get the AIS 768 00:37:36,920 --> 00:37:37,520 Speaker 4: in near future. 769 00:37:37,600 --> 00:37:37,799 Speaker 3: Right. 770 00:37:38,520 --> 00:37:41,120 Speaker 4: If the pace of AI development stays fast, we're really 771 00:37:41,120 --> 00:37:43,000 Speaker 4: really going to need the help of AI systems to 772 00:37:43,040 --> 00:37:45,600 Speaker 4: help us figure out how does your future A systems? 773 00:37:46,680 --> 00:37:49,920 Speaker 4: And so getting the right values into the next model 774 00:37:49,960 --> 00:37:52,279 Speaker 4: we build helps us figure out what to do with 775 00:37:52,320 --> 00:37:54,200 Speaker 4: them adel after that, and so I think this kind 776 00:37:54,200 --> 00:37:55,919 Speaker 4: of the short term work really does kind of fan 777 00:37:56,000 --> 00:37:58,640 Speaker 4: out into this longer feature. And yeah, getting the next 778 00:37:58,680 --> 00:38:00,680 Speaker 4: model right really matters for getting. 779 00:38:00,440 --> 00:38:01,359 Speaker 3: The fartugerules right. 780 00:38:01,560 --> 00:38:02,920 Speaker 1: Oh boy, it's called. 781 00:38:02,719 --> 00:38:04,360 Speaker 4: The scalable oversight problem. 782 00:38:04,600 --> 00:38:07,600 Speaker 1: I see. It's like the alignment problem is going to 783 00:38:07,600 --> 00:38:10,279 Speaker 1: scale up, and the best way for us to keep 784 00:38:10,400 --> 00:38:13,640 Speaker 1: up is to make sure that we get it right 785 00:38:13,760 --> 00:38:17,080 Speaker 1: now with these smaller systems, so we can use those 786 00:38:17,120 --> 00:38:19,799 Speaker 1: AIS to help us in the more complicated situations. 787 00:38:20,040 --> 00:38:20,239 Speaker 3: Yeah. 788 00:38:20,400 --> 00:38:22,720 Speaker 1: Yeah, that's oh wow. 789 00:38:22,480 --> 00:38:23,240 Speaker 4: That's the hope. 790 00:38:24,680 --> 00:38:27,400 Speaker 1: Okay, last question. Do you think humanity is doomed? 791 00:38:30,120 --> 00:38:32,120 Speaker 4: I don't think so. I think it's possible. I think 792 00:38:32,280 --> 00:38:36,200 Speaker 4: the AI presents a lot of really scary and destabilizing 793 00:38:36,200 --> 00:38:38,200 Speaker 4: possibilities that we can't roll out. So I think there's 794 00:38:38,200 --> 00:38:39,800 Speaker 4: a lot of work to do. I think we'll probably 795 00:38:39,800 --> 00:38:42,440 Speaker 4: figure it out. But I think it's also possible that 796 00:38:42,520 --> 00:38:45,759 Speaker 4: AI winds us up in a lot of weird, unfamiliar situations. 797 00:38:45,920 --> 00:38:48,680 Speaker 4: I think it's unlikely than possible that things go really, 798 00:38:48,719 --> 00:38:51,000 Speaker 4: really terribly, But I also think it's kind of unlikely 799 00:38:51,000 --> 00:38:53,640 Speaker 4: but possible, but the things stay totally normal and recognizable 800 00:38:53,640 --> 00:38:55,719 Speaker 4: and familiar. I think AI is just what it's going 801 00:38:55,760 --> 00:38:59,680 Speaker 4: to do to society and politics and economics is all 802 00:38:59,719 --> 00:39:01,080 Speaker 4: going to be confusing. I the's a lot that we'll 803 00:39:01,080 --> 00:39:02,200 Speaker 4: need to figure out pretty fast. 804 00:39:02,400 --> 00:39:03,120 Speaker 3: Fingers crossed. 805 00:39:05,160 --> 00:39:06,759 Speaker 1: I guess that's as good of an answer as we 806 00:39:06,760 --> 00:39:11,600 Speaker 1: can get these days. Fingers crossed. I guess just to 807 00:39:11,640 --> 00:39:13,600 Speaker 1: wrap up here, what do you think is going to 808 00:39:13,600 --> 00:39:17,040 Speaker 1: happen in the future, or what are some things about 809 00:39:17,040 --> 00:39:19,839 Speaker 1: this that you think most people are not thinking about 810 00:39:19,840 --> 00:39:21,160 Speaker 1: that they should be thinking about. 811 00:39:21,360 --> 00:39:24,200 Speaker 3: I think people should be thinking about ways in which 812 00:39:24,480 --> 00:39:28,480 Speaker 3: the AI systems that we have today are already capable 813 00:39:28,600 --> 00:39:33,120 Speaker 3: enough to cause harm, to change our world quite significantly, 814 00:39:33,239 --> 00:39:36,279 Speaker 3: change our culture, change the way we go about our day, 815 00:39:36,480 --> 00:39:40,080 Speaker 3: change the way we make decisions, change the way we 816 00:39:40,120 --> 00:39:45,239 Speaker 3: do our work. And the alignment problem and understanding when 817 00:39:45,680 --> 00:39:47,680 Speaker 3: models are aligned, I think, are two of the most 818 00:39:47,719 --> 00:39:52,600 Speaker 3: fundamental scientific challenges that we as a society are facing 819 00:39:52,680 --> 00:39:56,759 Speaker 3: right now. And that is not a sci fi future problem. 820 00:39:56,960 --> 00:40:00,040 Speaker 3: This is a problem about systems that we have to 821 00:40:00,560 --> 00:40:03,560 Speaker 3: We want to make sure these systems really do what 822 00:40:03,600 --> 00:40:06,040 Speaker 3: we want them to do, and that these systems help 823 00:40:06,120 --> 00:40:09,840 Speaker 3: us flourish and benefit humanity. The systems that we have 824 00:40:09,960 --> 00:40:13,840 Speaker 3: access to today already well beyond the capabilities that the 825 00:40:13,880 --> 00:40:17,280 Speaker 3: research community and certainly the general public thought we could 826 00:40:17,520 --> 00:40:19,719 Speaker 3: have in the year twenty twenty six. 827 00:40:20,120 --> 00:40:22,560 Speaker 1: I think you're saying that the future is here, but 828 00:40:22,600 --> 00:40:24,960 Speaker 1: we still haven't fully figured out the alignment problem. 829 00:40:25,120 --> 00:40:27,640 Speaker 3: Yes, that's right, meaning it's. 830 00:40:27,480 --> 00:40:29,600 Speaker 1: More pressing than ever that we figured this out. 831 00:40:29,840 --> 00:40:30,280 Speaker 3: I agree. 832 00:40:30,360 --> 00:40:33,759 Speaker 1: Yeah, amazing, doctor Rutner. How do we prove to the 833 00:40:33,800 --> 00:40:36,400 Speaker 1: audience that we're not an AI generated conversation? 834 00:40:38,640 --> 00:40:42,759 Speaker 3: I wish I had the answer. You can generate such 835 00:40:42,840 --> 00:40:47,160 Speaker 3: fantastic fake podcasts with AI now right right, with all 836 00:40:47,200 --> 00:40:51,600 Speaker 3: the little idiosyncrasies that you hear in podcasts today that 837 00:40:51,800 --> 00:40:52,960 Speaker 3: you know, I think that's hard to do. 838 00:40:53,239 --> 00:40:55,239 Speaker 1: Or I guess if this conversation is aligned with but 839 00:40:55,360 --> 00:40:57,280 Speaker 1: you want to hear, maybe it doesn't matter. 840 00:40:59,560 --> 00:41:01,719 Speaker 3: Still, I hope that the audience thinks that we so 841 00:41:01,719 --> 00:41:04,160 Speaker 3: out on human I think that that would be That 842 00:41:04,200 --> 00:41:06,600 Speaker 3: would be nice. 843 00:41:07,040 --> 00:41:09,799 Speaker 1: All right? Hey on, behalf of everyone who works in 844 00:41:09,840 --> 00:41:12,800 Speaker 1: the show picture joining us on the sixty plus episodes 845 00:41:12,880 --> 00:41:15,320 Speaker 1: we've done. Be sure to follow me on social media 846 00:41:15,440 --> 00:41:18,680 Speaker 1: or PhD comics dot com for updates and hey, thanks 847 00:41:18,719 --> 00:41:20,839 Speaker 1: to all the guests we've had on the show. Here's 848 00:41:20,840 --> 00:41:23,200 Speaker 1: a little tribute our editor Rose so good to put 849 00:41:23,239 --> 00:41:26,120 Speaker 1: together of all the times they were gracious enough to 850 00:41:26,160 --> 00:41:27,360 Speaker 1: put up with my questions. 851 00:41:27,680 --> 00:41:28,480 Speaker 4: That's a good question. 852 00:41:28,719 --> 00:41:31,520 Speaker 3: Yeah, that's a great question. You know, that's a great question. 853 00:41:31,719 --> 00:41:33,560 Speaker 4: Yeah, so those are all big questions for us to 854 00:41:33,600 --> 00:41:37,080 Speaker 4: answer a scientists. Yeah, that's a great question. It's a 855 00:41:37,080 --> 00:41:37,680 Speaker 4: good question. 856 00:41:37,960 --> 00:41:38,760 Speaker 3: That's a good question. 857 00:41:38,960 --> 00:41:42,640 Speaker 4: That's a good question. So that's actually a good question. 858 00:41:42,960 --> 00:41:45,200 Speaker 2: You're you're raising really great questions. 859 00:41:45,560 --> 00:41:47,080 Speaker 4: Yeah, that's a really good question. 860 00:41:47,320 --> 00:41:49,720 Speaker 3: That's a good question. That's a really good question. 861 00:41:50,000 --> 00:41:52,120 Speaker 4: That's a very hard question. That's a good question, though, 862 00:41:52,200 --> 00:41:55,840 Speaker 4: that's a really good question. That is a really good question. Yeah, 863 00:41:55,920 --> 00:41:57,120 Speaker 4: that's a great question. 864 00:41:57,320 --> 00:41:58,600 Speaker 3: Yeah, that's a good question. 865 00:41:58,800 --> 00:42:02,080 Speaker 1: Oh that's a great question, very very good question. 866 00:42:02,200 --> 00:42:03,240 Speaker 3: That's a really good question. 867 00:42:03,400 --> 00:42:05,040 Speaker 4: So I'll go through that question back at you. 868 00:42:05,120 --> 00:42:07,960 Speaker 1: What do you think? So we come once again to 869 00:42:08,040 --> 00:42:12,719 Speaker 1: the edge of scientific knowledge. Thanks for joining us, see 870 00:42:12,719 --> 00:42:19,400 Speaker 1: you next time you've been listening to Science Stuff. Production 871 00:42:19,560 --> 00:42:23,600 Speaker 1: of iHeartRadio Bringing the produced by me or Hey Cham, 872 00:42:23,760 --> 00:42:27,720 Speaker 1: edited by Rose Seguda, Executive producer Jerry Rowland, and audio 873 00:42:27,719 --> 00:42:30,960 Speaker 1: engineer and mixer Kasey Peckram. You can follow me on 874 00:42:31,000 --> 00:42:34,200 Speaker 1: social media. Just search for PhD Comics and the name 875 00:42:34,280 --> 00:42:36,959 Speaker 1: of your favorite platform. Be sure to subscribe to sign 876 00:42:37,000 --> 00:42:40,239 Speaker 1: stuff on the iHeartRadio app, Apple Podcasts, or wherever you 877 00:42:40,280 --> 00:42:41,200 Speaker 1: get your podcasts.