WEBVTT - How I AI: Which AI model should I use for which task?

0:00:00.760 --> 0:00:03.680
<v Speaker 1>If you've opened up an AI tool lately, instead at

0:00:03.720 --> 0:00:08.560
<v Speaker 1>a list of names like flash so one, it opens Fable, Mythos,

0:00:08.760 --> 0:00:11.239
<v Speaker 1>plus a pile of numbers that seem to go up

0:00:11.280 --> 0:00:15.120
<v Speaker 1>and down at random, You are not alone. The naming

0:00:15.320 --> 0:00:19.120
<v Speaker 1>within these platforms is genuinely confusing, and most of us

0:00:19.200 --> 0:00:22.919
<v Speaker 1>just pick one and hope for the best. Today, Neo

0:00:23.000 --> 0:00:26.760
<v Speaker 1>and I walk through the big four platforms Gemini, Chatch, FT, Cord,

0:00:26.800 --> 0:00:30.319
<v Speaker 1>and Copilot and sort out which model to actually use

0:00:30.680 --> 0:00:34.520
<v Speaker 1>and when. You will learn some tricks that cut through

0:00:34.800 --> 0:00:38.519
<v Speaker 1>all these confusing names and how to avoid the token

0:00:38.600 --> 0:00:42.680
<v Speaker 1>bills that have left some companies hundreds of thousands of

0:00:42.800 --> 0:00:47.479
<v Speaker 1>dollars out of pocket just for one employee using the

0:00:47.520 --> 0:00:51.159
<v Speaker 1>wrong models for the wrong tasks. By the end of

0:00:51.200 --> 0:00:55.080
<v Speaker 1>this episode, you will know exactly how much brain power

0:00:55.240 --> 0:00:59.240
<v Speaker 1>to throw at any given task for the AI tool

0:00:59.360 --> 0:01:00.440
<v Speaker 1>that you're using.

0:01:06.520 --> 0:01:10.240
<v Speaker 2>Welcome to how IAI with me, Doctor Amantha Imba and

0:01:10.480 --> 0:01:15.000
<v Speaker 2>Neo Applin, head of Inventium AI. Each episode we share

0:01:15.080 --> 0:01:17.800
<v Speaker 2>one practical way to use AI better at.

0:01:17.640 --> 0:01:21.640
<v Speaker 1>Work and in life. No fluff, no tech jargon, just

0:01:21.720 --> 0:01:25.640
<v Speaker 1>things you can use straight away. So Neo I feel

0:01:25.680 --> 0:01:27.440
<v Speaker 1>like in the news at the moment, there is a

0:01:27.480 --> 0:01:30.560
<v Speaker 1>lot of talk about new and different models launching, like

0:01:30.720 --> 0:01:33.800
<v Speaker 1>Mythos and Fable, and they've all got different names, and

0:01:33.840 --> 0:01:36.520
<v Speaker 1>then there's different numbers, and I feel like it can

0:01:36.560 --> 0:01:40.360
<v Speaker 1>be very confusing to the average person what is going

0:01:40.400 --> 0:01:42.440
<v Speaker 1>on with all these models you have?

0:01:42.480 --> 0:01:45.600
<v Speaker 3>These AI companies are horrendous when they're naming their models.

0:01:45.920 --> 0:01:49.840
<v Speaker 3>There's agents, there's GPTs, there's projects as models, there's fable.

0:01:49.880 --> 0:01:52.800
<v Speaker 3>There's too many different terms that mean very little to

0:01:52.880 --> 0:01:57.560
<v Speaker 3>the average person. But what it's about is how much

0:01:57.640 --> 0:02:01.080
<v Speaker 3>the model thinks and how deeply it thinks through your problem.

0:02:01.520 --> 0:02:03.520
<v Speaker 3>Is for the different models you've got, and so you've

0:02:03.520 --> 0:02:05.480
<v Speaker 3>got smaller light models that do maybe a little bit

0:02:05.520 --> 0:02:08.040
<v Speaker 3>of thinking, and you might have the heavier thinking models

0:02:08.200 --> 0:02:11.400
<v Speaker 3>that are smarter and will work through your problem more. So,

0:02:11.400 --> 0:02:14.200
<v Speaker 3>that's what it's about. It's about how much brain you

0:02:14.200 --> 0:02:19.000
<v Speaker 3>would like to put towards your problem. Now, the thinking

0:02:19.040 --> 0:02:22.600
<v Speaker 3>models will cost more for these companies to run. Therefore

0:02:22.760 --> 0:02:25.560
<v Speaker 3>they'll let you use it less each hour, each month,

0:02:25.600 --> 0:02:29.160
<v Speaker 3>each year, whatever, in a particular period of windows. And

0:02:29.240 --> 0:02:32.640
<v Speaker 3>so you'll find that if you use the deep thinking

0:02:32.919 --> 0:02:36.600
<v Speaker 3>big models, you'll get fewer uses each hour each day

0:02:36.880 --> 0:02:39.320
<v Speaker 3>than you would if you're using the small and light models.

0:02:39.320 --> 0:02:41.320
<v Speaker 3>So it's really about you're getting the right model for

0:02:41.360 --> 0:02:43.960
<v Speaker 3>the right task. And so if you've got something important,

0:02:44.080 --> 0:02:46.800
<v Speaker 3>use the thinking model. When you don't have anything particularly important,

0:02:46.919 --> 0:02:48.320
<v Speaker 3>use a smaller model.

0:02:48.680 --> 0:02:54.760
<v Speaker 1>So let's unpack the four main AI platforms in terms

0:02:54.760 --> 0:02:59.520
<v Speaker 1>of getting into Gemini, CHATCHABT, Claude and copile it to

0:02:59.639 --> 0:03:02.520
<v Speaker 1>understand and what are the models within each one and

0:03:02.560 --> 0:03:05.760
<v Speaker 1>what should we be using when? How does that sound

0:03:05.760 --> 0:03:06.160
<v Speaker 1>as a plan?

0:03:06.720 --> 0:03:07.720
<v Speaker 3>Love? It sounds great.

0:03:07.919 --> 0:03:11.680
<v Speaker 1>Let's start with Gemini, I think, and that's probably anecdotally

0:03:11.960 --> 0:03:13.920
<v Speaker 1>for us and the clients that we're working with with

0:03:14.080 --> 0:03:18.280
<v Speaker 1>inventium AI. We're not saying too many Gemini users. But

0:03:18.400 --> 0:03:20.400
<v Speaker 1>let's start there and cover off what should we be

0:03:20.520 --> 0:03:21.720
<v Speaker 1>using for what in Gemini.

0:03:22.720 --> 0:03:24.919
<v Speaker 3>Now, when I go into this, I'll tell you the names,

0:03:24.960 --> 0:03:27.640
<v Speaker 3>but also underneath the names when you've got there's a

0:03:27.720 --> 0:03:30.600
<v Speaker 3>model picker on each of these tools. Usually it's in

0:03:30.639 --> 0:03:33.880
<v Speaker 3>the prompting bar or around the prompting bar. When you

0:03:33.960 --> 0:03:36.560
<v Speaker 3>see the names, there's always a little bit of text

0:03:36.640 --> 0:03:38.640
<v Speaker 3>underneath it which gives you a clue as to what

0:03:38.720 --> 0:03:41.280
<v Speaker 3>it is. So, for example, here in Gemini on my

0:03:41.360 --> 0:03:44.400
<v Speaker 3>pick list today, I've got three point one flash light

0:03:44.720 --> 0:03:47.760
<v Speaker 3>and then underneath that it goes fastest dances. Underneath that

0:03:47.840 --> 0:03:50.160
<v Speaker 3>you've got three point five flash and it goes all

0:03:50.280 --> 0:03:53.400
<v Speaker 3>round help. And then the three point one pro says

0:03:53.480 --> 0:03:56.840
<v Speaker 3>advanced math and code. So it's giving you a clue

0:03:56.880 --> 0:03:59.360
<v Speaker 3>of which one to use when. So you can then

0:03:59.480 --> 0:04:01.320
<v Speaker 3>look at your go other numbers. Yeah, who gets but

0:04:01.400 --> 0:04:03.880
<v Speaker 3>the numbers. The numbers will always change, but it's a

0:04:03.920 --> 0:04:07.440
<v Speaker 3>little comments underneath that often mean more than the actual

0:04:07.520 --> 0:04:10.040
<v Speaker 3>name of it. So in Gemini they've got three point

0:04:10.040 --> 0:04:13.480
<v Speaker 3>one flash light, three point five flash and three point

0:04:13.480 --> 0:04:16.240
<v Speaker 3>one pro. Now the numbers can be confusing because you

0:04:16.360 --> 0:04:18.920
<v Speaker 3>look at it going three point one pro. Well, three

0:04:18.920 --> 0:04:22.400
<v Speaker 3>point five is better than three point one, so three

0:04:22.440 --> 0:04:25.040
<v Speaker 3>point five flash should be better than three point one pro.

0:04:25.320 --> 0:04:27.040
<v Speaker 3>It doesn't actually work that way. So that's why I'm

0:04:27.080 --> 0:04:30.760
<v Speaker 3>suggesting forget the numbers, go for the little text underneath it,

0:04:30.800 --> 0:04:33.240
<v Speaker 3>so that the pro one, the advanced method code, would

0:04:33.279 --> 0:04:35.760
<v Speaker 3>be the smartest model. And so when would you use that?

0:04:35.960 --> 0:04:38.000
<v Speaker 3>You'd use it on things that matter. So if you're

0:04:38.000 --> 0:04:41.200
<v Speaker 3>working out something like strategy, or you're doing problem solving

0:04:41.440 --> 0:04:43.680
<v Speaker 3>or Hey, I'm dealing with this really complex problem and

0:04:43.839 --> 0:04:45.800
<v Speaker 3>a market thing and a research thing, and how do

0:04:45.839 --> 0:04:48.480
<v Speaker 3>I put those things together? Then you'd use the pro

0:04:48.640 --> 0:04:52.200
<v Speaker 3>model here everyday use, I'd probably use flash And then

0:04:52.240 --> 0:04:54.360
<v Speaker 3>if you've got something just as quick fast light though,

0:04:54.400 --> 0:04:57.320
<v Speaker 3>summarize this paragraph into two words or whatever, I'd probably

0:04:57.400 --> 0:04:58.720
<v Speaker 3>use flash light for that.

0:04:59.279 --> 0:05:01.960
<v Speaker 1>And we'll get in to chat chapt next. But just

0:05:02.000 --> 0:05:07.599
<v Speaker 1>as a general question with the deeper versus quicker thinking,

0:05:08.080 --> 0:05:12.160
<v Speaker 1>something I wonder when let's say I'm working on refining

0:05:12.279 --> 0:05:15.919
<v Speaker 1>a really important piece of content, and I think to myself,

0:05:16.640 --> 0:05:20.800
<v Speaker 1>this is not a complex task like refining copy based

0:05:20.880 --> 0:05:23.520
<v Speaker 1>on the rules that I give it, but it's a

0:05:23.560 --> 0:05:27.320
<v Speaker 1>really important one to me. So in that kind of scenario,

0:05:27.800 --> 0:05:29.360
<v Speaker 1>which model should I be using.

0:05:30.760 --> 0:05:34.400
<v Speaker 3>If it's more language and just summarize these things, then

0:05:34.480 --> 0:05:37.840
<v Speaker 3>I would be fine with a smaller model, one of

0:05:37.920 --> 0:05:40.799
<v Speaker 3>those just everyday ones, which in this case in Gemini,

0:05:41.200 --> 0:05:43.080
<v Speaker 3>the every day one is probably the three point five

0:05:43.120 --> 0:05:45.120
<v Speaker 3>flash or you can get away with a flash light

0:05:45.200 --> 0:05:47.520
<v Speaker 3>if it's just words and how do I put these

0:05:47.560 --> 0:05:50.560
<v Speaker 3>phrases together? Things like that, But if you want it

0:05:50.600 --> 0:05:53.400
<v Speaker 3>to actually analyze things and think through things, if it's

0:05:53.440 --> 0:05:55.800
<v Speaker 3>important to get it right. In your case, it's important

0:05:55.800 --> 0:05:58.039
<v Speaker 3>to get it right. I would definitely be using the

0:05:58.160 --> 0:06:01.320
<v Speaker 3>everyday one here, that's the three point five flash rather

0:06:01.440 --> 0:06:03.800
<v Speaker 3>than light because it's actually think through how it does

0:06:03.880 --> 0:06:07.440
<v Speaker 3>its word summary, those kind of things. But often we

0:06:07.480 --> 0:06:09.880
<v Speaker 3>need to do things that really matter. If it really matters,

0:06:09.920 --> 0:06:12.200
<v Speaker 3>I'm often doing PRO. But if it's really matter, usually

0:06:12.200 --> 0:06:15.520
<v Speaker 3>it's about thinking through a problem really matters, rather than

0:06:15.680 --> 0:06:18.800
<v Speaker 3>just summarize the sentence. Well, I need to summarize it, right, Yeah,

0:06:18.960 --> 0:06:21.599
<v Speaker 3>probably everyday model is probably fine for those things.

0:06:21.839 --> 0:06:24.680
<v Speaker 1>Okay, let's move on to chat GPT. What are the

0:06:24.760 --> 0:06:28.680
<v Speaker 1>different model options in chat GPT and what should we

0:06:28.760 --> 0:06:30.680
<v Speaker 1>be using when they've improved.

0:06:30.720 --> 0:06:34.800
<v Speaker 3>There's an instant then thinking, then PRO. That makes a

0:06:34.800 --> 0:06:37.080
<v Speaker 3>lot more sense to me. Instant is don't think, just

0:06:37.240 --> 0:06:39.360
<v Speaker 3>shoot me the answer straight away. And I love that

0:06:39.800 --> 0:06:42.599
<v Speaker 3>and that's exactly what it does. Instant effectively, don't think,

0:06:42.839 --> 0:06:46.680
<v Speaker 3>just do so use your own brain, go ahead thinking,

0:06:46.760 --> 0:06:49.080
<v Speaker 3>we'll actually think through it. That's what I'm generally on

0:06:49.240 --> 0:06:52.880
<v Speaker 3>all of the time. And then PRO, if it matters,

0:06:53.000 --> 0:06:55.400
<v Speaker 3>that's the one you'd pick, So do not. You get

0:06:55.440 --> 0:06:57.760
<v Speaker 3>quite a few pro uses, so a lot of people

0:06:57.800 --> 0:07:01.200
<v Speaker 3>are too scared to use them until they actually need them,

0:07:01.240 --> 0:07:02.920
<v Speaker 3>in which case you've actually got a whole bunch of

0:07:02.920 --> 0:07:05.520
<v Speaker 3>potential uses that you've missed out on using. So do

0:07:05.640 --> 0:07:07.920
<v Speaker 3>not be afraid to use the higher models.

0:07:08.240 --> 0:07:12.600
<v Speaker 1>So let's look at Claude now, which I know is

0:07:13.240 --> 0:07:17.080
<v Speaker 1>the AI tool that we use most often, and that

0:07:17.160 --> 0:07:19.880
<v Speaker 1>certainly we're saying a lot of our clients moving a

0:07:19.920 --> 0:07:24.680
<v Speaker 1>lot more towards Claude. And this is where certainly the

0:07:24.760 --> 0:07:27.760
<v Speaker 1>term fable you might have heard in relation to Claude.

0:07:27.760 --> 0:07:30.240
<v Speaker 1>What's going on there and what models should we be using?

0:07:30.960 --> 0:07:33.840
<v Speaker 3>Okay, so Claude have four models at the moment, well

0:07:34.320 --> 0:07:38.119
<v Speaker 3>almost four. So there starts at Haiku, that's their small

0:07:38.160 --> 0:07:41.200
<v Speaker 3>fast light model, then it goes to Sonnet, then to opus,

0:07:41.360 --> 0:07:44.400
<v Speaker 3>and then more recently there is fable. So instead of

0:07:44.440 --> 0:07:48.400
<v Speaker 3>giving you the words like chachipt does what's chechibit? He

0:07:48.440 --> 0:07:50.560
<v Speaker 3>got instant thinking and pro it kind of does what

0:07:50.600 --> 0:07:52.400
<v Speaker 3>it says in the box. I kind of like that

0:07:52.680 --> 0:07:55.480
<v Speaker 3>now with Claude, think of it as poems. A haiku

0:07:55.560 --> 0:07:57.520
<v Speaker 3>is a Japanese poem which only has a couple of words.

0:07:57.520 --> 0:07:59.400
<v Speaker 3>If it's only that three lights it's a tiny poem

0:08:00.440 --> 0:08:02.680
<v Speaker 3>is a larger poem, and Opus is quite a large poem,

0:08:02.720 --> 0:08:05.200
<v Speaker 3>and a Fable is obviously a huge poem. So it's

0:08:05.200 --> 0:08:08.000
<v Speaker 3>about poem sizes. That's how they've named their models. So

0:08:08.000 --> 0:08:09.760
<v Speaker 3>if you kind of get that in your mind, and.

0:08:09.680 --> 0:08:13.239
<v Speaker 1>Can I just say neo, I've never heard anyone describe

0:08:13.280 --> 0:08:16.720
<v Speaker 1>it like that. To me, that is quite revelatory, So

0:08:16.840 --> 0:08:17.960
<v Speaker 1>thank you. Continue.

0:08:18.800 --> 0:08:21.720
<v Speaker 3>So, yeah, these are the poem sizes and effective. It's

0:08:21.720 --> 0:08:23.840
<v Speaker 3>how smart you want it to be and how much

0:08:23.920 --> 0:08:27.080
<v Speaker 3>thinking you want it to do as well. So, as

0:08:27.080 --> 0:08:29.320
<v Speaker 3>I said, we've got hik So and Opus. Don't worry

0:08:29.320 --> 0:08:31.240
<v Speaker 3>about the numbers, just go for the names here. The

0:08:31.280 --> 0:08:35.319
<v Speaker 3>numbers will always change. Now, just recently they had had

0:08:35.600 --> 0:08:39.120
<v Speaker 3>released Fable. If you've also heard of Mythos, Mythos and

0:08:39.160 --> 0:08:41.160
<v Speaker 3>Fable are pretty much the same model. This is the

0:08:41.160 --> 0:08:42.880
<v Speaker 3>way you saw in the news. It was like, oh,

0:08:42.920 --> 0:08:45.320
<v Speaker 3>we can't release it. It's too scary, it's too smart,

0:08:45.320 --> 0:08:47.440
<v Speaker 3>a little hack every system in the world and all

0:08:47.480 --> 0:08:51.600
<v Speaker 3>those kind of things. So Fable is Mythos. It's the

0:08:51.640 --> 0:08:54.800
<v Speaker 3>full blown model. But they've just said you can't talk

0:08:54.840 --> 0:08:57.600
<v Speaker 3>about certain things, so you can't talk about chemical stuff

0:08:57.640 --> 0:09:01.120
<v Speaker 3>because someone might create chemical warfware. You can't talk about

0:09:01.400 --> 0:09:04.240
<v Speaker 3>security stuff because you might try to hack something. And

0:09:04.280 --> 0:09:06.760
<v Speaker 3>so what they've tried to do is get Fable without

0:09:06.800 --> 0:09:11.920
<v Speaker 3>those problematic topics, and then they released it. They've released it,

0:09:12.040 --> 0:09:16.320
<v Speaker 3>and since then there's obviously politics involved, but the US

0:09:16.440 --> 0:09:19.439
<v Speaker 3>government has said, well, in a way, you could actually

0:09:19.480 --> 0:09:22.800
<v Speaker 3>get it to review your security on your code, and

0:09:22.840 --> 0:09:26.280
<v Speaker 3>so therefore, no, we're not happy with that, and because

0:09:26.320 --> 0:09:29.520
<v Speaker 3>of that there's been discussions, and then Fable has been

0:09:30.120 --> 0:09:32.880
<v Speaker 3>paused will put it that way, and so you can't

0:09:32.880 --> 0:09:35.680
<v Speaker 3>actually use Fable at the moment. So for most people,

0:09:36.400 --> 0:09:39.800
<v Speaker 3>do you need something which is as awesome as Fable, No,

0:09:39.840 --> 0:09:44.280
<v Speaker 3>I'd probably say no. If you're getting AI to build

0:09:44.320 --> 0:09:46.920
<v Speaker 3>your full software suite or whatever, Fable would be absolutely

0:09:47.000 --> 0:09:50.680
<v Speaker 3>amazing for that. Opus is great Fable from everything I'm

0:09:50.720 --> 0:09:54.120
<v Speaker 3>reading from people who've been using it, it's mind blowingly amazing.

0:09:54.760 --> 0:09:56.839
<v Speaker 3>But it's not available just yet. So what I'd be

0:09:56.920 --> 0:10:00.319
<v Speaker 3>doing is Haiku small fast light every day kind of small,

0:10:00.440 --> 0:10:03.360
<v Speaker 3>like summarize this kind of stuff. Haiku is great. I

0:10:03.440 --> 0:10:06.719
<v Speaker 3>work in Sonnet pretty much all the time, but then

0:10:06.760 --> 0:10:09.040
<v Speaker 3>I turn it up to Opus when I need to

0:10:09.480 --> 0:10:11.880
<v Speaker 3>because Opus is it does use more tokens, which we

0:10:11.920 --> 0:10:14.480
<v Speaker 3>talk about later. But Opus is very smart and does

0:10:14.559 --> 0:10:17.920
<v Speaker 3>an awesome job. But don't sit in the most thinkingest

0:10:18.000 --> 0:10:20.120
<v Speaker 3>model all the time because you're just going to waste

0:10:20.120 --> 0:10:22.000
<v Speaker 3>all your usage and then you won't have the great

0:10:22.040 --> 0:10:24.160
<v Speaker 3>usage ready for when you need to use it.

0:10:25.000 --> 0:10:29.080
<v Speaker 1>Okay, let's now talk about co pilot. What are the

0:10:29.120 --> 0:10:30.480
<v Speaker 1>different models.

0:10:30.040 --> 0:10:32.560
<v Speaker 3>There right now? I need to give you a little

0:10:32.559 --> 0:10:37.760
<v Speaker 3>bit of market knowledge here. So Microsoft have had invested

0:10:37.800 --> 0:10:40.719
<v Speaker 3>in chat ChiPT, so actually Copilot under the hood has

0:10:40.760 --> 0:10:43.240
<v Speaker 3>a chat chipitt Brain, not just that you can use

0:10:43.320 --> 0:10:48.800
<v Speaker 3>chat Chipt's models in Copilot, and they've also recently done

0:10:48.840 --> 0:10:53.920
<v Speaker 3>a deal with Anthropic so you can use Claude inside copilot.

0:10:54.360 --> 0:10:56.360
<v Speaker 3>So this is within the copilot model. Pickup, So in

0:10:56.400 --> 0:10:59.320
<v Speaker 3>copilot you've got auto and I like the little text

0:10:59.360 --> 0:11:02.000
<v Speaker 3>and decides how to think. It's got quick response. That's

0:11:02.040 --> 0:11:04.120
<v Speaker 3>it's a qui equivalent of haiku, so don't think, just

0:11:04.120 --> 0:11:06.520
<v Speaker 3>give me a response and then think deeper, which is

0:11:06.600 --> 0:11:09.760
<v Speaker 3>its equivalent of something like the thinking ish models. So

0:11:10.240 --> 0:11:12.719
<v Speaker 3>I'm generally on auto all the time and it will

0:11:12.760 --> 0:11:14.800
<v Speaker 3>decide how long to think. But if it's important I

0:11:14.840 --> 0:11:17.560
<v Speaker 3>will go think deeper because it will always do a

0:11:17.600 --> 0:11:20.000
<v Speaker 3>better job of thinking if I tell it to think

0:11:20.040 --> 0:11:24.239
<v Speaker 3>deeper than auto if it thinks so if it's important, strategy,

0:11:24.520 --> 0:11:26.520
<v Speaker 3>market research, all those kinds of things, or how do

0:11:26.559 --> 0:11:27.959
<v Speaker 3>I do this is really important, I'm going to get

0:11:27.960 --> 0:11:31.040
<v Speaker 3>it right. For a client, absolutely think deeper. But I

0:11:31.160 --> 0:11:33.320
<v Speaker 3>encourage you to do some tests and say, all right, oh,

0:11:33.400 --> 0:11:36.080
<v Speaker 3>now try the Opus model because you can use Opus

0:11:36.320 --> 0:11:40.719
<v Speaker 3>that's Claudes model inside Copilot, which is quite amazing. And

0:11:40.760 --> 0:11:44.680
<v Speaker 3>there's also a GPT from AI Open AI and so

0:11:44.720 --> 0:11:46.520
<v Speaker 3>you can actually use one of the GPT once you've

0:11:46.520 --> 0:11:48.160
<v Speaker 3>got the think deepers in quick responses and all the

0:11:48.200 --> 0:11:51.920
<v Speaker 3>same things, so you've got more choices in Copilot, which

0:11:51.960 --> 0:11:52.840
<v Speaker 3>is actually pretty cool.

0:11:54.240 --> 0:11:57.480
<v Speaker 1>Let's finish by talking about tokens because this is so

0:11:58.320 --> 0:12:02.440
<v Speaker 1>important because you know you might just be listening and going, well,

0:12:02.559 --> 0:12:07.239
<v Speaker 1>why wouldn't I just have the AI think deeply about everything?

0:12:07.440 --> 0:12:10.960
<v Speaker 1>I mean, what are the implications of that? And I'd say,

0:12:11.080 --> 0:12:14.080
<v Speaker 1>in if you're a solopreneur and you've just got your

0:12:14.120 --> 0:12:18.160
<v Speaker 1>own individual license, or maybe you're in a small business

0:12:18.480 --> 0:12:20.480
<v Speaker 1>where I think the pricing is a little bit different

0:12:20.480 --> 0:12:23.840
<v Speaker 1>depending on how things are structured in your company. It's

0:12:23.880 --> 0:12:26.600
<v Speaker 1>so different to if you are working in a large

0:12:26.800 --> 0:12:30.560
<v Speaker 1>corporate and you've got an enterprise plan because what we

0:12:30.640 --> 0:12:33.760
<v Speaker 1>know there is that often there's a per user charge,

0:12:34.040 --> 0:12:38.040
<v Speaker 1>but there is also a charge for tokens. So how

0:12:38.080 --> 0:12:40.079
<v Speaker 1>do we need to be thinking about this because quite frankly,

0:12:40.120 --> 0:12:43.760
<v Speaker 1>I know we've both heard some nightmare stories around token

0:12:43.840 --> 0:12:49.000
<v Speaker 1>usage going crazy and skyrocketing the prices that people are

0:12:49.000 --> 0:12:50.199
<v Speaker 1>paying for their AI tool.

0:12:51.800 --> 0:12:56.400
<v Speaker 3>Yeah, so it's all about II recommend use the right

0:12:56.440 --> 0:12:58.559
<v Speaker 3>model that you need for the task you're doing. So,

0:12:58.600 --> 0:13:01.960
<v Speaker 3>for example, I need to summarize twenty customer feedback things

0:13:02.000 --> 0:13:06.080
<v Speaker 3>into a couple of themes. An auto model, a regular

0:13:06.080 --> 0:13:08.760
<v Speaker 3>model is probably fine for that. Or you might need

0:13:08.840 --> 0:13:11.920
<v Speaker 3>to say, hey, those same twenty comments do they reveal

0:13:11.920 --> 0:13:14.600
<v Speaker 3>a product strategy risk or does that really impact what

0:13:14.640 --> 0:13:16.440
<v Speaker 3>I need to do for the next kind of sets

0:13:16.440 --> 0:13:19.400
<v Speaker 3>of investments that matters, in which case I probably use

0:13:19.440 --> 0:13:22.280
<v Speaker 3>a larger top tier, a thinking model and Opus model

0:13:22.320 --> 0:13:25.120
<v Speaker 3>for that. There's also a timeframe thing, So if I'm

0:13:25.240 --> 0:13:27.200
<v Speaker 3>using Opus, it's going to spend a lot of time

0:13:27.280 --> 0:13:28.920
<v Speaker 3>thinking through the thing. It will take me ages to

0:13:28.920 --> 0:13:30.520
<v Speaker 3>get that response back and if you're going to get

0:13:30.520 --> 0:13:33.120
<v Speaker 3>frustrated with it taking a lot of time, then you're

0:13:33.160 --> 0:13:35.199
<v Speaker 3>wasting a lot of tokens. Okay, but let's actually go

0:13:35.280 --> 0:13:38.400
<v Speaker 3>back to the token thing. So if you're in a

0:13:38.400 --> 0:13:41.679
<v Speaker 3>particular plan, and only some plans have this, So if

0:13:41.720 --> 0:13:44.079
<v Speaker 3>you're in say everyday claud plan, this is a thirty

0:13:44.080 --> 0:13:47.720
<v Speaker 3>bucks a month plan, then you will have a usage limit.

0:13:47.760 --> 0:13:50.040
<v Speaker 3>So you can actually go to your settings page and

0:13:50.160 --> 0:13:52.280
<v Speaker 3>under settings there's a thing called usage and you can

0:13:52.320 --> 0:13:54.400
<v Speaker 3>see that you've got a five hour windows, a little

0:13:54.400 --> 0:13:56.240
<v Speaker 3>bit of a bar how much you've used, and there's

0:13:56.240 --> 0:13:58.319
<v Speaker 3>a weekly limit window and there's a bit of a

0:13:58.360 --> 0:14:01.160
<v Speaker 3>bar there, and you can also top up and buy

0:14:01.200 --> 0:14:02.720
<v Speaker 3>more on top of that if you need. And what

0:14:02.720 --> 0:14:05.439
<v Speaker 3>they're basically saying is you can have a certain number

0:14:05.440 --> 0:14:09.320
<v Speaker 3>of uses of your different models over that five hour

0:14:09.360 --> 0:14:12.360
<v Speaker 3>window or one week window until we say you've used

0:14:12.400 --> 0:14:17.319
<v Speaker 3>it enough, you've overused what our models do now, Haiku

0:14:17.520 --> 0:14:20.560
<v Speaker 3>is about two to three times cheaper than Sonnet, and

0:14:20.640 --> 0:14:24.800
<v Speaker 3>son It is minimum two times cheaper than Opus, so

0:14:24.880 --> 0:14:28.840
<v Speaker 3>there's like a minimum six times differential in price on

0:14:28.920 --> 0:14:32.040
<v Speaker 3>these different models. So if you're using Opus for hey,

0:14:32.080 --> 0:14:36.320
<v Speaker 3>summarize this email, you are burning through your tokens and

0:14:36.360 --> 0:14:38.200
<v Speaker 3>you'll hit at the end of the five hour period

0:14:38.240 --> 0:14:40.360
<v Speaker 3>and AI is going to say tote sos. You're going

0:14:40.440 --> 0:14:42.320
<v Speaker 3>to have to come back in a couple of hours, right,

0:14:42.600 --> 0:14:44.400
<v Speaker 3>And that becomes a problem if you actually needed to

0:14:44.400 --> 0:14:46.600
<v Speaker 3>get some work done. So picking the right model means

0:14:46.600 --> 0:14:48.440
<v Speaker 3>you can actually get some work done, which is great.

0:14:49.400 --> 0:14:52.920
<v Speaker 3>Now if you're in a company plan, and some companies

0:14:52.920 --> 0:14:56.560
<v Speaker 3>have a plan where it's simply you can you've get

0:14:56.640 --> 0:14:59.120
<v Speaker 3>access to it. But each time you get AI to

0:14:59.200 --> 0:15:01.600
<v Speaker 3>run a query to get a prompt or together to think,

0:15:01.800 --> 0:15:04.800
<v Speaker 3>that goes kachin kachin, kachin kachin on your bank balance.

0:15:05.480 --> 0:15:08.440
<v Speaker 3>So if that's the case, then you might be costing

0:15:08.480 --> 0:15:11.200
<v Speaker 3>an awful lot of cash to your poor IT department.

0:15:12.280 --> 0:15:15.880
<v Speaker 3>And we hear stories where some companies have got like

0:15:15.920 --> 0:15:19.920
<v Speaker 3>a five hundred thousand dollar AI bill for that month

0:15:20.360 --> 0:15:22.600
<v Speaker 3>and things like that because people have just been using

0:15:22.680 --> 0:15:25.480
<v Speaker 3>these things on summarize this email or stuff they don't need,

0:15:25.560 --> 0:15:28.720
<v Speaker 3>or they're letting it go crazy. And so my encouragement

0:15:28.920 --> 0:15:30.720
<v Speaker 3>is pick the right model for the job. But also

0:15:30.880 --> 0:15:33.840
<v Speaker 3>if you do have, say Claude, check your usage limits

0:15:33.840 --> 0:15:36.440
<v Speaker 3>and see whether you're actually using it right, so you

0:15:36.480 --> 0:15:39.560
<v Speaker 3>can actually see whether you're over using any of these models.

0:15:39.720 --> 0:15:42.880
<v Speaker 1>And just for context, that five hundred thousand dollar AI

0:15:42.920 --> 0:15:44.680
<v Speaker 1>bill that you talk about, that was just from one

0:15:44.760 --> 0:15:48.520
<v Speaker 1>employee from memory that particular story that you're referring to,

0:15:49.080 --> 0:15:49.440
<v Speaker 1>and they.

0:15:49.320 --> 0:15:54.000
<v Speaker 3>Were also getting AI. This is APIs and chlor Cowork

0:15:54.760 --> 0:15:56.800
<v Speaker 3>and sorry clock code, I should say, and all those

0:15:56.840 --> 0:15:59.600
<v Speaker 3>kind of things. So they were going nuts, right, But

0:15:59.640 --> 0:16:01.960
<v Speaker 3>here's the they were going nuts but not watching what

0:16:02.000 --> 0:16:05.440
<v Speaker 3>they're doing, and that's coming to be a bigger problem

0:16:05.600 --> 0:16:08.360
<v Speaker 3>because like Claude has its own usage limits and a

0:16:08.440 --> 0:16:11.200
<v Speaker 3>lot of companies they're off this. You know, you can

0:16:11.400 --> 0:16:13.760
<v Speaker 3>be the five our limit you stop and instead it's

0:16:13.800 --> 0:16:16.080
<v Speaker 3>just a paper usage. There's another thing coming out in

0:16:16.080 --> 0:16:18.800
<v Speaker 3>a couple of months time, which is Microsoft Cowork, which

0:16:18.840 --> 0:16:21.000
<v Speaker 3>is the same as clor COK in Microsoft, they will

0:16:21.040 --> 0:16:23.720
<v Speaker 3>also have usage limits. So picking the right model is

0:16:23.800 --> 0:16:26.480
<v Speaker 3>good practice now for when all of these things go

0:16:26.520 --> 0:16:27.520
<v Speaker 3>to usage only.

0:16:28.240 --> 0:16:31.080
<v Speaker 1>Thank you so much, NEO. Hopefully you are now feeling

0:16:31.320 --> 0:16:34.840
<v Speaker 1>more confident and clear in terms of which model you

0:16:34.880 --> 0:16:38.200
<v Speaker 1>should be using for all the different tasks that you're

0:16:38.320 --> 0:16:41.000
<v Speaker 1>using AI for. Thank you so much for listening and

0:16:41.040 --> 0:16:45.000
<v Speaker 1>we will see you next week. How IAI was hosted

0:16:45.040 --> 0:16:48.120
<v Speaker 1>by me Amantha Imber and Neo Applan. A big thank

0:16:48.120 --> 0:16:50.640
<v Speaker 1>you to Martin Imber who does our sound editing, and

0:16:50.760 --> 0:16:54.160
<v Speaker 1>Jim Rubio for production support, and thank you to John

0:16:54.280 --> 0:16:56.240
<v Speaker 1>Kilby who composed the theme music