WEBVTT - Cerebras Built a New Way to Make Sand Think - The Story

0:00:16.160 --> 0:00:19.160
<v Speaker 1>Welcome to tech Stuff. I'm os Voloscian. There are many

0:00:19.160 --> 0:00:22.040
<v Speaker 1>ways to think about AI. The one that recently struck

0:00:22.079 --> 0:00:24.320
<v Speaker 1>me quite a lot comes from Demis Hassebus, founder of

0:00:24.400 --> 0:00:27.160
<v Speaker 1>Google's Deep Mind, who said, if you stop to think

0:00:27.200 --> 0:00:30.960
<v Speaker 1>about it, we've essentially found a way to make sand think.

0:00:31.600 --> 0:00:34.479
<v Speaker 1>He was talking about silicon, the element that is melted

0:00:34.520 --> 0:00:38.040
<v Speaker 1>and sliced into the chips that power AI, and today

0:00:38.360 --> 0:00:41.040
<v Speaker 1>I want to talk about those chips. Our guest today

0:00:41.120 --> 0:00:44.360
<v Speaker 1>is Steve Vassalo, a venture capitalist who was the first

0:00:44.520 --> 0:00:47.880
<v Speaker 1>backer of Cerebras, a chip company that went public earlier

0:00:47.880 --> 0:00:52.920
<v Speaker 1>this year in the largest ever IPO for a semiconductor company. Steve,

0:00:53.200 --> 0:00:54.160
<v Speaker 1>Welcome to tech Stuff.

0:00:54.560 --> 0:00:56.720
<v Speaker 2>Thanks Arles. It's great to be here, great.

0:00:56.480 --> 0:00:59.360
<v Speaker 1>To see I'm curious what's your reaction to that, To

0:00:59.440 --> 0:01:01.760
<v Speaker 1>that framing of demis making sand think.

0:01:02.800 --> 0:01:05.320
<v Speaker 3>I mean, it's such a great reminder of what it

0:01:05.400 --> 0:01:08.520
<v Speaker 3>is that we're doing and how fundamental it is, both

0:01:08.520 --> 0:01:11.080
<v Speaker 3>from a physics perspective and all the hard work that

0:01:11.280 --> 0:01:13.600
<v Speaker 3>is required to take a piece of sand and turn

0:01:13.640 --> 0:01:16.640
<v Speaker 3>it into ultimately intelligence. And then there's also something that

0:01:17.120 --> 0:01:19.600
<v Speaker 3>feels sort of I don't know, almost I would say

0:01:19.720 --> 0:01:21.520
<v Speaker 3>sort of you know, just sort of scary and fun

0:01:21.560 --> 0:01:24.360
<v Speaker 3>in a humanity sense, like, Wow, what are we building

0:01:24.400 --> 0:01:26.440
<v Speaker 3>right now? What are we doing to be able to

0:01:26.920 --> 0:01:30.520
<v Speaker 3>sort of turn this, you know, very basic material that

0:01:30.560 --> 0:01:34.360
<v Speaker 3>surrounds us everywhere into something that is super intelligent.

0:01:35.240 --> 0:01:37.880
<v Speaker 1>I want to ask you all about the story of

0:01:37.920 --> 0:01:41.200
<v Speaker 1>Cerebrest and and also frankly, what wayfer scale means. But

0:01:41.280 --> 0:01:43.440
<v Speaker 1>before we get there, can you just zoom out a

0:01:43.440 --> 0:01:46.399
<v Speaker 1>little bit and give us the wider context of what's

0:01:46.440 --> 0:01:49.160
<v Speaker 1>going on with chips right now? Because I think we

0:01:49.200 --> 0:01:51.440
<v Speaker 1>will see the headlines. I mean, I'm going to read

0:01:51.440 --> 0:01:53.800
<v Speaker 1>a few, but the New York Times recently had it's

0:01:53.840 --> 0:01:56.560
<v Speaker 1>not just in Nvidia. The AI boom has ignited Asia's

0:01:56.600 --> 0:01:59.400
<v Speaker 1>chip companies. Dull Street Journal had a couple of months

0:01:59.400 --> 0:02:02.600
<v Speaker 1>ago inside Apple's push to build an all American Chip

0:02:02.880 --> 0:02:06.960
<v Speaker 1>axios in July, Chip craze for the layman, what's going

0:02:06.960 --> 0:02:07.360
<v Speaker 1>on here?

0:02:07.760 --> 0:02:08.760
<v Speaker 2>Yeah? I love that question.

0:02:08.880 --> 0:02:13.320
<v Speaker 3>So the way I think about it is really sort

0:02:13.320 --> 0:02:16.440
<v Speaker 3>of starts with what I'll call workloads, And the workloads are,

0:02:16.480 --> 0:02:18.560
<v Speaker 3>you know, what is the work of computing, And if

0:02:18.600 --> 0:02:21.760
<v Speaker 3>you go back to the nineteen seventies and eighties, the

0:02:21.840 --> 0:02:25.520
<v Speaker 3>workloads at that time were very serialized. They were kind

0:02:25.520 --> 0:02:28.120
<v Speaker 3>of diverse if you think about the invention of things

0:02:28.160 --> 0:02:30.240
<v Speaker 3>like the spreadsheet, or you know, as a kid, I

0:02:30.240 --> 0:02:33.000
<v Speaker 3>wrote my college essays on word Star and an eighty

0:02:33.040 --> 0:02:34.320
<v Speaker 3>eighty six IBM computer.

0:02:35.000 --> 0:02:38.360
<v Speaker 1>Diverse in a sense that different computers for different applications,

0:02:38.639 --> 0:02:40.119
<v Speaker 1>or diverse in what sense.

0:02:40.160 --> 0:02:43.440
<v Speaker 3>Diverse in that you had one computing platform, which turned

0:02:43.440 --> 0:02:46.280
<v Speaker 3>out to be Intel's X eighty six platform that turned

0:02:46.320 --> 0:02:50.079
<v Speaker 3>out to be a great general purpose computing platform for

0:02:50.720 --> 0:02:53.639
<v Speaker 3>a wide variety of applications. And then sometime in the

0:02:53.680 --> 0:02:56.320
<v Speaker 3>kind of mid nineteen eighties, we began to see a

0:02:56.360 --> 0:02:59.720
<v Speaker 3>new workload emerge. And that workload really came from two

0:02:59.760 --> 0:03:02.760
<v Speaker 3>areas as one was from gaming and the rendering of

0:03:02.760 --> 0:03:06.760
<v Speaker 3>pixels and a need for paralyzation of the compute. And

0:03:06.919 --> 0:03:09.600
<v Speaker 3>then also in believe it or not, the CAD tools,

0:03:09.639 --> 0:03:13.440
<v Speaker 3>so you know, building new products, whether its futures exactly.

0:03:13.480 --> 0:03:14.480
<v Speaker 2>So computer aided.

0:03:14.280 --> 0:03:18.760
<v Speaker 3>Design pushed the perimeter of what was possible on traditional

0:03:18.800 --> 0:03:22.080
<v Speaker 3>computing CPU general purpose systems. So that was really the

0:03:22.120 --> 0:03:24.640
<v Speaker 3>dawn of the GPU, and really that was Nvidia. And

0:03:24.680 --> 0:03:26.680
<v Speaker 3>by the way, at the time of Nvidia's founding, there

0:03:26.720 --> 0:03:29.639
<v Speaker 3>were like we're seeing today thirty other companies working on

0:03:30.000 --> 0:03:33.360
<v Speaker 3>graphics processing units. And I say graphics processing units. We

0:03:33.680 --> 0:03:35.840
<v Speaker 3>many people, you know, see the term GPU. They don't

0:03:36.240 --> 0:03:39.160
<v Speaker 3>know or remember that, like graphics is the g It

0:03:39.240 --> 0:03:42.000
<v Speaker 3>was purpose built for that workload. And then the next

0:03:42.080 --> 0:03:45.800
<v Speaker 3>real shift was a few decades later, but was really

0:03:46.120 --> 0:03:49.280
<v Speaker 3>a constraint a workload driven by mobile devices, and the

0:03:49.320 --> 0:03:52.760
<v Speaker 3>constraints there are totally different, right. This is form factor,

0:03:53.000 --> 0:03:56.720
<v Speaker 3>this is power consumption, battery life of course related to that,

0:03:57.040 --> 0:03:59.400
<v Speaker 3>and so you've got a totally different workload. And then

0:03:59.400 --> 0:04:01.600
<v Speaker 3>what we began to see this is now zooming forward

0:04:01.600 --> 0:04:05.280
<v Speaker 3>to kind of the twenty fifteen timeframe, really starting in

0:04:05.320 --> 0:04:08.600
<v Speaker 3>the twenty twelve timeframes, but the workloads began to really

0:04:08.640 --> 0:04:11.640
<v Speaker 3>build in the mid twenty tens. And what you saw was,

0:04:11.800 --> 0:04:14.560
<v Speaker 3>oh wait a minute, statistical inference over large data sets.

0:04:14.840 --> 0:04:17.960
<v Speaker 3>Our portfolio companies were beginning to hire data science teams

0:04:18.400 --> 0:04:22.359
<v Speaker 3>business models that weren't possible before you had sort of

0:04:22.360 --> 0:04:25.919
<v Speaker 3>a real awareness of kind of the data underlying your This.

0:04:26.720 --> 0:04:30.159
<v Speaker 1>Was the era when phrases like big data es out

0:04:30.160 --> 0:04:33.719
<v Speaker 1>computing mark andresen saying software eats the world. Is this

0:04:33.760 --> 0:04:35.040
<v Speaker 1>the area referring to sort.

0:04:34.920 --> 0:04:37.000
<v Speaker 3>Of about five years later than software into the world,

0:04:37.000 --> 0:04:40.360
<v Speaker 3>but kind of roughly contemporaneous with that. But really what

0:04:40.400 --> 0:04:43.000
<v Speaker 3>you're seeing big data is absolutely right and in a

0:04:43.120 --> 0:04:47.680
<v Speaker 3>sense that, wow, there are signals inside of our businesses

0:04:47.720 --> 0:04:50.320
<v Speaker 3>that if we're more attuned to we could make better decisions.

0:04:50.720 --> 0:04:53.080
<v Speaker 3>And what was interesting at that time, and this kind

0:04:53.080 --> 0:04:55.400
<v Speaker 3>of ties back to the sort of the GPU wave

0:04:55.960 --> 0:04:59.960
<v Speaker 3>was those kinds of workloads. AI workloads are very data

0:05:00.120 --> 0:05:02.559
<v Speaker 3>io intensive, so that just means the data is moving

0:05:02.600 --> 0:05:06.440
<v Speaker 3>through the system with with a lot of bandwidth, memory

0:05:06.480 --> 0:05:08.960
<v Speaker 3>bandwidth and with high high frequency.

0:05:08.520 --> 0:05:11.000
<v Speaker 1>Back and forth, back and forth. We had somebody who

0:05:11.040 --> 0:05:13.760
<v Speaker 1>I'm sure you know, Nick McEwen from course on the

0:05:13.760 --> 0:05:16.359
<v Speaker 1>show recently, who was at Cisco in the days of

0:05:16.400 --> 0:05:18.400
<v Speaker 1>building the kind of you know dominat net Round.

0:05:18.440 --> 0:05:20.880
<v Speaker 3>He was building networks, networking systems in that company was

0:05:20.880 --> 0:05:21.640
<v Speaker 3>acquired by Cisco.

0:05:21.680 --> 0:05:22.120
<v Speaker 2>Exactly.

0:05:22.240 --> 0:05:25.839
<v Speaker 1>He was saying, today's data center, a single data center,

0:05:25.880 --> 0:05:28.760
<v Speaker 1>has more back and forth connections than the whole Internet

0:05:28.760 --> 0:05:30.120
<v Speaker 1>did in the nineties.

0:05:29.680 --> 0:05:31.560
<v Speaker 2>Exactly, exactly. In fact, was fun.

0:05:31.600 --> 0:05:35.520
<v Speaker 3>There was an announcement yesterday and videos new vera Rubin

0:05:36.160 --> 0:05:37.600
<v Speaker 3>and you looked at I mean I sort of laughed

0:05:37.600 --> 0:05:40.559
<v Speaker 3>when I looked at the photo of the launch because

0:05:40.560 --> 0:05:42.840
<v Speaker 3>you're staring at this thing and it was just basically

0:05:42.880 --> 0:05:48.240
<v Speaker 3>this medusa of basically cables and connectors behind all of that,

0:05:48.400 --> 0:05:51.560
<v Speaker 3>you know, literally probably miles of AMFONYL cables.

0:05:52.160 --> 0:05:54.240
<v Speaker 2>Was of course some graphics processing.

0:05:53.960 --> 0:05:56.400
<v Speaker 3>Units, but what they're having to do is connect all

0:05:56.400 --> 0:05:59.839
<v Speaker 3>that stuff back together. And so this ties to the

0:06:00.000 --> 0:06:02.440
<v Speaker 3>point around wafer scale, which is, if you are going

0:06:02.520 --> 0:06:06.920
<v Speaker 3>to build something starting from scratch, not using platforms that

0:06:07.000 --> 0:06:11.320
<v Speaker 3>already exist, not relying on CPUs or GPUs or ARM architectures,

0:06:11.600 --> 0:06:13.719
<v Speaker 3>and you were to say, what would you build that's

0:06:13.960 --> 0:06:15.719
<v Speaker 3>really designed for these kinds of workloads?

0:06:15.920 --> 0:06:17.520
<v Speaker 2>You build what we built at Cerebras.

0:06:17.800 --> 0:06:20.560
<v Speaker 3>And that's the journey that we started in twenty fifteen,

0:06:21.000 --> 0:06:23.240
<v Speaker 3>and really it's born of this notion of big, big

0:06:23.279 --> 0:06:27.120
<v Speaker 3>workloads deserve custom purpose built silicon, and that has been

0:06:27.200 --> 0:06:30.320
<v Speaker 3>really the story of every silicon wave over the last

0:06:30.320 --> 0:06:31.000
<v Speaker 3>five decades.

0:06:31.560 --> 0:06:35.040
<v Speaker 1>And talk about this concept of wafers scale because I can.

0:06:35.320 --> 0:06:38.799
<v Speaker 1>I mean, if you look at the Cerebras chip, it's

0:06:38.839 --> 0:06:44.159
<v Speaker 1>fifty times bigger than an Nvidia GPU. And as you

0:06:44.240 --> 0:06:46.640
<v Speaker 1>mentioned with with Jensen holding up with verrub and it

0:06:46.720 --> 0:06:49.040
<v Speaker 1>is this kind of like, you know, sort of different

0:06:49.160 --> 0:06:52.520
<v Speaker 1>silicon chips connected together with cables and stuff, whereas your

0:06:52.600 --> 0:06:55.360
<v Speaker 1>chip is is kind of a single huge piece of silicon.

0:06:55.520 --> 0:06:56.960
<v Speaker 1>Why do they do that and why do you keep

0:06:56.960 --> 0:06:57.919
<v Speaker 1>it as a single piece.

0:06:58.200 --> 0:07:00.599
<v Speaker 3>The way to think about sort of silicon, MANUF is

0:07:00.640 --> 0:07:03.800
<v Speaker 3>it's grown as you started this conversation with sand, you know,

0:07:03.880 --> 0:07:08.400
<v Speaker 3>it's grown from sand into single crystalline columns, and then

0:07:08.440 --> 0:07:11.720
<v Speaker 3>those columns are sliced very thinly that is then cut

0:07:11.760 --> 0:07:13.840
<v Speaker 3>to basically turn it into a nine inch roughly nine

0:07:13.880 --> 0:07:17.120
<v Speaker 3>inch square, and that nine inch square has trillions of

0:07:17.120 --> 0:07:20.880
<v Speaker 3>transistors on it. And the way Nvidia or ARM or

0:07:20.960 --> 0:07:22.840
<v Speaker 3>others build their chips is they might start with a

0:07:22.840 --> 0:07:25.600
<v Speaker 3>twelve in wafer, but then they cut it. They cut

0:07:25.600 --> 0:07:28.520
<v Speaker 3>that wafer into many smaller components that are, as you said,

0:07:28.560 --> 0:07:31.280
<v Speaker 3>sort of fifty or fifty eight times smaller than ours,

0:07:31.640 --> 0:07:33.679
<v Speaker 3>and then they cut it and then in many cases

0:07:33.680 --> 0:07:36.560
<v Speaker 3>they're actually literally putting them back together. And you ask

0:07:36.640 --> 0:07:39.640
<v Speaker 3>a question, which is why cut it? That's exactly the

0:07:39.720 --> 0:07:42.160
<v Speaker 3>question that we asked ourselves, which is you shouldn't cut it.

0:07:43.000 --> 0:07:44.120
<v Speaker 2>You should keep it all together.

0:07:44.160 --> 0:07:46.840
<v Speaker 3>Because as soon as you as soon as you take

0:07:46.880 --> 0:07:51.080
<v Speaker 3>communications from a chip off of the chip. So once

0:07:51.600 --> 0:07:54.200
<v Speaker 3>once a bit needs to move from one chip to another,

0:07:54.640 --> 0:07:58.680
<v Speaker 3>you basically have other components involved, and everything you add

0:07:58.760 --> 0:08:00.440
<v Speaker 3>as you would imagine with any sytem them. As soon

0:08:00.480 --> 0:08:03.840
<v Speaker 3>as you add another component between two pieces of silicon,

0:08:04.000 --> 0:08:07.200
<v Speaker 3>you've slowed it down, You've created new constraints. And so

0:08:07.360 --> 0:08:10.120
<v Speaker 3>the big idea here is don't do any of that.

0:08:10.400 --> 0:08:12.400
<v Speaker 3>I'm sure you're familiar. I mean, Elon, you know he's

0:08:12.440 --> 0:08:14.880
<v Speaker 3>working on other stuff in this domain, but you know

0:08:14.920 --> 0:08:17.040
<v Speaker 3>he has this, this sort of rule of thumb. When

0:08:17.080 --> 0:08:19.960
<v Speaker 3>you work at SpaceX or one of his companies, you

0:08:20.000 --> 0:08:22.160
<v Speaker 3>hear this all the time, which is the best part

0:08:22.240 --> 0:08:25.600
<v Speaker 3>is no part. And the point of that is you'd

0:08:25.720 --> 0:08:28.280
<v Speaker 3>like to reduce all of your components down to their

0:08:28.320 --> 0:08:30.640
<v Speaker 3>simplest sort of platonic ideal form.

0:08:30.920 --> 0:08:31.920
<v Speaker 2>And that is exactly what.

0:08:31.880 --> 0:08:34.120
<v Speaker 3>We decided to do with Cerebus, which is, instead of

0:08:34.280 --> 0:08:38.040
<v Speaker 3>taking a twelve inch column of silicon, slicing it into

0:08:38.200 --> 0:08:41.920
<v Speaker 3>hundreds of smaller pieces, then putting those on a motherboard

0:08:41.920 --> 0:08:42.800
<v Speaker 3>and reconnecting them.

0:08:42.679 --> 0:08:45.240
<v Speaker 2>With copper wires or copper traces.

0:08:45.000 --> 0:08:47.400
<v Speaker 3>Let's keep it all on one piece so that the

0:08:47.559 --> 0:08:51.600
<v Speaker 3>data can move as fast as possible throughout that single

0:08:51.640 --> 0:08:54.679
<v Speaker 3>piece of silicon, and we'll get a thousand x improvement

0:08:54.800 --> 0:08:58.160
<v Speaker 3>in performance. And that's from a memory bandwidth perspective. That's

0:08:58.160 --> 0:09:00.839
<v Speaker 3>exactly what we delivered, and so's I were now call

0:09:00.880 --> 0:09:03.480
<v Speaker 3>it roughly fifteen in some cases even more than that

0:09:03.760 --> 0:09:06.720
<v Speaker 3>times faster than any other inference platform on the planet.

0:09:07.280 --> 0:09:09.400
<v Speaker 1>What do you think the biggest misconception is about the

0:09:10.040 --> 0:09:12.520
<v Speaker 1>quote unquote chip craze today.

0:09:13.480 --> 0:09:16.400
<v Speaker 3>I would say the biggest misconception is how hard it

0:09:16.480 --> 0:09:20.360
<v Speaker 3>is to build a new platform to yield those systems,

0:09:20.640 --> 0:09:22.839
<v Speaker 3>and the yield it's basically for everyone you make, how

0:09:22.840 --> 0:09:25.960
<v Speaker 3>many of those can you use? And that that is

0:09:26.000 --> 0:09:28.760
<v Speaker 3>sort of one of those fundamental drivers in the semiconductor industry.

0:09:29.040 --> 0:09:32.240
<v Speaker 3>But in the case of large semiconductor systems powering it

0:09:32.640 --> 0:09:35.000
<v Speaker 3>so you know, most of these systems are now taking

0:09:35.040 --> 0:09:38.719
<v Speaker 3>tens of kilowatts. Back when I was building startups a

0:09:38.760 --> 0:09:41.640
<v Speaker 3>couple decades ago, you know, a rack was somewhere on

0:09:41.679 --> 0:09:43.600
<v Speaker 3>the other of two to three kilowatts. We now have

0:09:43.679 --> 0:09:47.400
<v Speaker 3>racks that are sixty kilowatts, and so you're trying to

0:09:47.440 --> 0:09:49.120
<v Speaker 3>power this system now. Of course, once you're powering it.

0:09:49.200 --> 0:09:50.360
<v Speaker 3>Guess what else you have to do. You have to

0:09:50.360 --> 0:09:52.319
<v Speaker 3>cool it. You have to get that all the all

0:09:52.360 --> 0:09:54.679
<v Speaker 3>the heat out of that system. And then of course,

0:09:54.800 --> 0:09:57.120
<v Speaker 3>once if that's that's now working, now you're going to

0:09:57.200 --> 0:09:59.240
<v Speaker 3>program this thing that has you know, two and a

0:09:59.240 --> 0:10:02.920
<v Speaker 3>half trillion trans and then get you know, modern models

0:10:02.960 --> 0:10:04.720
<v Speaker 3>to work on that and get them up and running

0:10:04.840 --> 0:10:07.040
<v Speaker 3>in a matter of days or weeks instead of months

0:10:07.160 --> 0:10:10.920
<v Speaker 3>or years. And so you're compounding many, many hard problems.

0:10:11.240 --> 0:10:15.600
<v Speaker 3>And in most businesses, when you compound hard problems, what

0:10:15.679 --> 0:10:18.040
<v Speaker 3>comes out the other end often doesn't work. And so,

0:10:18.640 --> 0:10:21.040
<v Speaker 3>you know, the single greatest challenge really is can you

0:10:21.120 --> 0:10:22.520
<v Speaker 3>actually make this thing at scale?

0:10:23.160 --> 0:10:23.319
<v Speaker 4>Yeah?

0:10:23.360 --> 0:10:25.920
<v Speaker 1>I think I think you said somewhere, you know, I

0:10:25.960 --> 0:10:28.520
<v Speaker 1>sometimes joke that Cerebras was five startups in one. You've

0:10:28.520 --> 0:10:31.520
<v Speaker 1>sold wafer scale, great nive power system, cool it, package it,

0:10:31.520 --> 0:10:34.080
<v Speaker 1>deploy code onto it, integrated with an existing frameworks, get

0:10:34.120 --> 0:10:36.840
<v Speaker 1>into data centers. Each one of those challenges was quite

0:10:36.880 --> 0:10:39.800
<v Speaker 1>literally its own company. So I mean, let's go back

0:10:39.840 --> 0:10:43.640
<v Speaker 1>to twenty fifteen. I mean, you had this insight, you know,

0:10:43.720 --> 0:10:46.520
<v Speaker 1>every new wave of computing demands and new type of

0:10:46.520 --> 0:10:49.760
<v Speaker 1>silicon right and We talked about the computing of the eighties,

0:10:49.800 --> 0:10:53.040
<v Speaker 1>one type of chips and the GPUs for gaming, which

0:10:53.280 --> 0:10:56.559
<v Speaker 1>ended up by sort of accident essentially powering the air revolution,

0:10:57.080 --> 0:10:59.960
<v Speaker 1>the arm chips that power cell phones, and then these

0:11:00.080 --> 0:11:04.240
<v Speaker 1>these new chips that cerebreast design for data centers. Essentially

0:11:04.480 --> 0:11:07.200
<v Speaker 1>exactly who were you in twenty fifteen, and we've got

0:11:07.200 --> 0:11:10.320
<v Speaker 1>the kind of we've got the global context, the industry context,

0:11:10.320 --> 0:11:14.080
<v Speaker 1>what's the what was the Steve context? And what made you?

0:11:14.720 --> 0:11:17.120
<v Speaker 1>I mean nowadays twenty twenty six. When we think about

0:11:17.360 --> 0:11:20.280
<v Speaker 1>venture capital and investing, you know, we often think about

0:11:20.320 --> 0:11:23.839
<v Speaker 1>like momentum based, like how do I get into anthropic

0:11:23.880 --> 0:11:26.800
<v Speaker 1>before it goes public right? Versus like how do I

0:11:26.880 --> 0:11:29.720
<v Speaker 1>make a ten year plus bet on something which which

0:11:29.920 --> 0:11:33.880
<v Speaker 1>you know is not only requires all these macro things

0:11:33.920 --> 0:11:36.600
<v Speaker 1>to break my way, but also has like just layer

0:11:36.679 --> 0:11:39.200
<v Speaker 1>upon layer, up on layer of complexity. And you know,

0:11:39.240 --> 0:11:41.199
<v Speaker 1>these these all these different elements we talked about in

0:11:41.280 --> 0:11:45.160
<v Speaker 1>terms of cooling and packaging and integration. What what made

0:11:45.200 --> 0:11:49.160
<v Speaker 1>you want to do this? And how much of your career,

0:11:49.240 --> 0:11:52.240
<v Speaker 1>shall we say, or your capital, your financial capital, your

0:11:52.240 --> 0:11:54.240
<v Speaker 1>relation capital did you stake on this?

0:11:54.760 --> 0:11:57.679
<v Speaker 3>I love that, Yeah, so I mean the quick bounce

0:11:57.720 --> 0:12:00.680
<v Speaker 3>on me is you know, I'm trained as a roboticist.

0:12:01.600 --> 0:12:03.880
<v Speaker 3>I lived for the first ten years of my career

0:12:03.960 --> 0:12:07.280
<v Speaker 3>at the intersection of electrical and mechanical systems. I worked

0:12:07.280 --> 0:12:09.160
<v Speaker 3>at a company called Ideo, designing products.

0:12:10.280 --> 0:12:10.840
<v Speaker 2>Some of those.

0:12:10.720 --> 0:12:13.400
<v Speaker 3>Products brought those two worlds together as well as sort

0:12:13.400 --> 0:12:16.080
<v Speaker 3>of the design world, sort of really answering the question

0:12:16.120 --> 0:12:19.000
<v Speaker 3>of once you've solved those hard technical problems at inter

0:12:19.000 --> 0:12:22.079
<v Speaker 3>section of electrical and mechanical systems, does anyone actually care?

0:12:22.120 --> 0:12:24.000
<v Speaker 3>Do people love it and want it? So I spent

0:12:24.040 --> 0:12:26.800
<v Speaker 3>really a decade building products. Came to Foundation actually as

0:12:26.800 --> 0:12:30.040
<v Speaker 3>an entrepreneur in residence on the path to starting another company.

0:12:30.040 --> 0:12:31.920
<v Speaker 3>And this is now in the spring of two thousand

0:12:31.960 --> 0:12:36.760
<v Speaker 3>and seven. And so I'm a builder by nature, if

0:12:36.800 --> 0:12:39.400
<v Speaker 3>I would say, sort of more of an accidental venture

0:12:39.400 --> 0:12:42.960
<v Speaker 3>capitalist than one who sort of, you know, thought to

0:12:43.000 --> 0:12:45.400
<v Speaker 3>do this from when I was a teenager. I didn't

0:12:45.400 --> 0:12:48.000
<v Speaker 3>I literally didn't know the term venture capital until I

0:12:48.040 --> 0:12:50.600
<v Speaker 3>moved to Silicon Valley back in the mid nineteen nineties.

0:12:51.040 --> 0:12:53.920
<v Speaker 3>And the specific context and I love the question around

0:12:53.960 --> 0:12:56.240
<v Speaker 3>sort of what was required to get an investment like this.

0:12:57.000 --> 0:12:59.640
<v Speaker 2>Approved as it were, more than a decade ago.

0:13:00.679 --> 0:13:02.840
<v Speaker 3>The answer to that question is is really sort of

0:13:02.880 --> 0:13:05.120
<v Speaker 3>central to how we work at Foundation Capital, which is

0:13:05.679 --> 0:13:08.880
<v Speaker 3>we go really deep in opportunity space, and we're not

0:13:08.880 --> 0:13:12.079
<v Speaker 3>trying to be everything to everyone, and when we go deep,

0:13:12.640 --> 0:13:15.120
<v Speaker 3>we try to kind of build what we call our

0:13:15.160 --> 0:13:17.559
<v Speaker 3>points of view, a sense of what would need to

0:13:17.600 --> 0:13:19.520
<v Speaker 3>be true for us to be excited about a new

0:13:19.520 --> 0:13:21.760
<v Speaker 3>investment in this area. What are the attributes of the

0:13:21.760 --> 0:13:25.120
<v Speaker 3>founders that would make them successful or not, What is

0:13:25.160 --> 0:13:27.720
<v Speaker 3>the timing of the market opportunity, because of course many

0:13:27.800 --> 0:13:30.560
<v Speaker 3>venture capital investments, you know, have great ideas, but are

0:13:30.559 --> 0:13:32.680
<v Speaker 3>maybe a decade too early or in some cases multiple

0:13:32.679 --> 0:13:35.640
<v Speaker 3>dedcades too early. And so for me, I had been

0:13:35.640 --> 0:13:39.160
<v Speaker 3>studying this space, and I'd actually met Andrew Feldman and

0:13:39.200 --> 0:13:42.480
<v Speaker 3>his co founder Gary Back no joke, in October of

0:13:42.480 --> 0:13:45.480
<v Speaker 3>two thousand and seven, seven years before we really re

0:13:45.559 --> 0:13:49.079
<v Speaker 3>engaged around the opportunity at Cerebras, and they were actually

0:13:49.080 --> 0:13:51.400
<v Speaker 3>at that time just starting their prior company, a company

0:13:51.440 --> 0:13:53.840
<v Speaker 3>called c Micro, which was kind of also in the

0:13:53.960 --> 0:13:57.840
<v Speaker 3>kind of data center think sort of warehouse scale computing concept,

0:13:57.960 --> 0:14:02.400
<v Speaker 3>and we didn't actually invest in that company. I hit

0:14:02.440 --> 0:14:05.640
<v Speaker 3>it off with them and they were acquired by AMD

0:14:05.760 --> 0:14:07.800
<v Speaker 3>maybe five years later, and a couple of years into that,

0:14:08.280 --> 0:14:10.920
<v Speaker 3>I reached back out to Andrew and I could tell immediately,

0:14:10.960 --> 0:14:13.440
<v Speaker 3>I mean, I had this intuition going into the conversation

0:14:13.679 --> 0:14:15.160
<v Speaker 3>that he was not going to stay long for AMD.

0:14:15.280 --> 0:14:18.800
<v Speaker 3>He had another startup in him. And so we basically,

0:14:18.800 --> 0:14:21.160
<v Speaker 3>over the course of the next two years riffed on

0:14:21.240 --> 0:14:24.160
<v Speaker 3>a whole bunch of ideas that at that time were

0:14:24.280 --> 0:14:26.080
<v Speaker 3>kind of around that big data concept that you and

0:14:26.080 --> 0:14:28.440
<v Speaker 3>I talked about earlier, and we were looking at a

0:14:28.440 --> 0:14:31.040
<v Speaker 3>whole bunch of companies were using Andrew as a sounding

0:14:31.080 --> 0:14:34.000
<v Speaker 3>board for the opportunity space. And then he and Gary

0:14:34.000 --> 0:14:37.240
<v Speaker 3>and then Sean and then Michael and JP started to

0:14:37.280 --> 0:14:40.600
<v Speaker 3>basically converge on this idea around Hey, wait a minute,

0:14:40.640 --> 0:14:43.280
<v Speaker 3>there's this new workload, and this new workload we got

0:14:43.320 --> 0:14:46.160
<v Speaker 3>to pay attention to, and this new workload while as

0:14:46.160 --> 0:14:47.520
<v Speaker 3>you said, sort of NVIDIA is a bit of a

0:14:47.520 --> 0:14:50.080
<v Speaker 3>happy accident that it happens to be better than CPU's,

0:14:50.360 --> 0:14:53.200
<v Speaker 3>but it's not what you build. And so the you know,

0:14:53.240 --> 0:14:55.840
<v Speaker 3>these five founders basically set out to go build something

0:14:55.920 --> 0:14:58.800
<v Speaker 3>that was purpose built, and I had this prepared mind

0:14:58.880 --> 0:15:00.880
<v Speaker 3>I'd been working with Andrew on kind of a range

0:15:00.880 --> 0:15:03.760
<v Speaker 3>of concepts for two years. So when he finally said

0:15:03.760 --> 0:15:06.280
<v Speaker 3>this is now kind of March timeframe of twenty sixteen,

0:15:06.440 --> 0:15:08.680
<v Speaker 3>it's like, I think we're ready. I told him, I said, look,

0:15:08.720 --> 0:15:10.640
<v Speaker 3>we want to be your first term sheet and so

0:15:10.680 --> 0:15:14.200
<v Speaker 3>we It was no joke on April first, so April

0:15:14.240 --> 0:15:16.760
<v Speaker 3>Fool's Day, I said, Andrew, you know, we presented it

0:15:16.840 --> 0:15:18.800
<v Speaker 3>with an offer, and over the course of the next

0:15:18.840 --> 0:15:22.320
<v Speaker 3>few weeks we basically converged on that initial financing, which

0:15:22.360 --> 0:15:25.119
<v Speaker 3>we which we did in partnership with Eric at Benchmark

0:15:25.160 --> 0:15:28.640
<v Speaker 3>and Pierre and Leora at Eclipse, And yeah, that was

0:15:28.680 --> 0:15:31.880
<v Speaker 3>the beginning of a cerebras back in May of twenty sixteen.

0:15:33.040 --> 0:15:36.520
<v Speaker 1>What was the closest moment between twenty sixteen and ringing

0:15:36.520 --> 0:15:39.400
<v Speaker 1>the bet On in the New York Sulck Exchange to failure?

0:15:39.840 --> 0:15:44.080
<v Speaker 3>Oh many, We had many, so many white knuckle moments.

0:15:44.400 --> 0:15:46.520
<v Speaker 3>We didn't ship the first product for three and a

0:15:46.520 --> 0:15:49.760
<v Speaker 3>half years from that May twenty sixteen. I'd say there

0:15:49.760 --> 0:15:52.760
<v Speaker 3>were lots of challenges around getting that first semiconductor system

0:15:52.800 --> 0:15:55.840
<v Speaker 3>to work. You know, we'd spent multiple millions building the

0:15:55.840 --> 0:15:58.600
<v Speaker 3>first one and no joke. The very first system we

0:15:58.680 --> 0:16:02.040
<v Speaker 3>plugged it in and it blew up. I mean we

0:16:02.440 --> 0:16:04.560
<v Speaker 3>joke about like we call those thermal events because you

0:16:04.560 --> 0:16:05.920
<v Speaker 3>never want to tell your landlord you've.

0:16:05.840 --> 0:16:08.440
<v Speaker 2>Had a fire. But we had an event.

0:16:08.880 --> 0:16:11.120
<v Speaker 3>We had a small thermal event around our first system.

0:16:11.440 --> 0:16:13.440
<v Speaker 3>And what you're trusting in those cases is, wow, that

0:16:13.520 --> 0:16:16.040
<v Speaker 3>was very expensive. That one system you could you could

0:16:16.040 --> 0:16:18.400
<v Speaker 3>as certain cost you know, millions of dollars with all

0:16:18.440 --> 0:16:20.640
<v Speaker 3>the R and D that was invested in it. And

0:16:20.720 --> 0:16:22.240
<v Speaker 3>then you've got to just trust that we're going to

0:16:22.320 --> 0:16:25.520
<v Speaker 3>come back, you know, the next day and and dig

0:16:25.600 --> 0:16:28.320
<v Speaker 3>through what went wrong and figure it out.

0:16:28.400 --> 0:16:30.680
<v Speaker 2>And and we did. And that team is just.

0:16:30.680 --> 0:16:34.560
<v Speaker 3>Born of so much grit around sort of getting a

0:16:34.560 --> 0:16:37.280
<v Speaker 3>system that again there were no there were no playbooks

0:16:37.280 --> 0:16:39.560
<v Speaker 3>for this. TSMC had never done anything like what we

0:16:39.560 --> 0:16:41.600
<v Speaker 3>were asking them to do. And then the other piece

0:16:41.600 --> 0:16:43.720
<v Speaker 3>I would say that was a very scary moment was

0:16:43.760 --> 0:16:45.800
<v Speaker 3>when we realized, you know, we were initially actually going

0:16:45.840 --> 0:16:47.840
<v Speaker 3>to build an air cold system because every other system

0:16:48.280 --> 0:16:51.080
<v Speaker 3>in the data center, you know, including all the GPU systems,

0:16:51.120 --> 0:16:53.400
<v Speaker 3>were all air cooled, and we realized to make our

0:16:53.400 --> 0:16:55.960
<v Speaker 3>system work, we would need to do we'd need to

0:16:56.000 --> 0:16:58.720
<v Speaker 3>water cool it, and there was lots of questions around

0:16:59.040 --> 0:17:01.640
<v Speaker 3>how that would work. APR kind of head of mechanical

0:17:01.640 --> 0:17:04.119
<v Speaker 3>engineering and system design, had never done this.

0:17:04.200 --> 0:17:06.960
<v Speaker 2>He kind of admended that to me later. And you know, keap.

0:17:06.800 --> 0:17:09.480
<v Speaker 3>Dissipation at that scale is a really hard problem. And

0:17:09.520 --> 0:17:11.680
<v Speaker 3>so we we built a lot of prototypes. We actually

0:17:11.720 --> 0:17:13.879
<v Speaker 3>ended up having to bring in several consultants and I

0:17:13.920 --> 0:17:16.879
<v Speaker 3>spent you know, probably you know it was not hundreds

0:17:16.880 --> 0:17:20.119
<v Speaker 3>of hours, but you know up in that range working

0:17:20.119 --> 0:17:22.359
<v Speaker 3>with the team on how do we actually solve this

0:17:22.640 --> 0:17:24.480
<v Speaker 3>really really hard problems because you'd come out of the

0:17:24.520 --> 0:17:27.240
<v Speaker 3>lab some days thinking, guys, we might just be trying

0:17:27.280 --> 0:17:29.400
<v Speaker 3>to beat the laws of physics, which is never never

0:17:29.440 --> 0:17:30.040
<v Speaker 3>a good idea.

0:17:31.200 --> 0:17:33.080
<v Speaker 1>So you, in a sense you had all of these

0:17:33.560 --> 0:17:37.679
<v Speaker 1>pod engineering problems and now engineering problems. Some of them

0:17:37.720 --> 0:17:39.400
<v Speaker 1>are solved, but no doubt are new ones. But also

0:17:39.440 --> 0:17:42.800
<v Speaker 1>of course you have you know, you have competition within Vidia.

0:17:42.960 --> 0:17:45.199
<v Speaker 1>Is that is that essentially? I mean, how does the

0:17:45.240 --> 0:17:47.480
<v Speaker 1>business look today? Is it? Is it a pitch battle

0:17:47.520 --> 0:17:48.040
<v Speaker 1>of Nvidia?

0:17:48.080 --> 0:17:48.159
<v Speaker 4>Like?

0:17:48.240 --> 0:17:51.080
<v Speaker 1>Is that the central kind of business challenge?

0:17:51.400 --> 0:17:51.640
<v Speaker 2>Yeah?

0:17:51.680 --> 0:17:54.159
<v Speaker 3>I mean they are the eight hundred bound guerilla no question,

0:17:54.800 --> 0:17:57.480
<v Speaker 3>and we are the insurgent and they've built you know,

0:17:57.520 --> 0:18:01.240
<v Speaker 3>they've built an extraordinary category. When when Nvidia moved into

0:18:01.320 --> 0:18:03.760
<v Speaker 3>the data center space, which was kind of in the

0:18:03.840 --> 0:18:06.480
<v Speaker 3>mid twenty tens and it's now you know, it's more

0:18:06.480 --> 0:18:08.840
<v Speaker 3>than half of their business, and so they are definitely

0:18:09.520 --> 0:18:12.400
<v Speaker 3>you know there, they are the player that we most

0:18:12.440 --> 0:18:15.479
<v Speaker 3>directly get compared with, which I'm fine with. And they

0:18:15.480 --> 0:18:18.560
<v Speaker 3>have a different architecture and you know, we fundamentally believe

0:18:18.600 --> 0:18:22.000
<v Speaker 3>we're doing we're doing things differently for again back to

0:18:22.040 --> 0:18:23.840
<v Speaker 3>the sort of you know, initial premise here, because we

0:18:23.880 --> 0:18:26.320
<v Speaker 3>think this workload deserves its own silicon.

0:18:27.320 --> 0:18:30.600
<v Speaker 1>So so displacing them with clients, I guess is one

0:18:30.760 --> 0:18:33.640
<v Speaker 1>challenge on the on the on the demand side, shall

0:18:33.680 --> 0:18:35.879
<v Speaker 1>we say, what about on the supply side? I mean

0:18:35.920 --> 0:18:39.520
<v Speaker 1>you mentioned that TSMC actually you know, fabricate these these

0:18:39.560 --> 0:18:42.560
<v Speaker 1>chips for you guys, Are you like literally fighting of

0:18:42.640 --> 0:18:45.280
<v Speaker 1>a factory flaw space in Taiwan with Nvidia to to

0:18:45.440 --> 0:18:48.560
<v Speaker 1>to see which orders for chips they'll fulfill faster? I mean,

0:18:48.720 --> 0:18:50.800
<v Speaker 1>how do you how do you navigate that relationship?

0:18:51.160 --> 0:18:54.800
<v Speaker 3>Yeah, So TSMC has been an extraordinary partner to us,

0:18:54.840 --> 0:18:57.480
<v Speaker 3>and they are to really everyone in the industry. And

0:18:57.520 --> 0:19:00.240
<v Speaker 3>what I appreciate most about them, and I think what

0:19:00.320 --> 0:19:03.480
<v Speaker 3>they are uniquely well suited to doing is working on

0:19:03.560 --> 0:19:07.399
<v Speaker 3>new ideas even before the market is sort of clearly

0:19:07.440 --> 0:19:09.159
<v Speaker 3>there for them. And they have been a great partner

0:19:09.200 --> 0:19:11.760
<v Speaker 3>to us in that regard, because again, if you're looking

0:19:11.760 --> 0:19:15.679
<v Speaker 3>at Cerebras back in twenty sixteen and your TSMC and

0:19:15.680 --> 0:19:18.399
<v Speaker 3>you've got you know, customers like Video or Apple and

0:19:19.440 --> 0:19:21.840
<v Speaker 3>many many others of scale that are you know, would

0:19:21.880 --> 0:19:24.480
<v Speaker 3>dwarf anything that we could ask of them. For certainly

0:19:24.560 --> 0:19:27.400
<v Speaker 3>several years, they had no business in working with us,

0:19:27.440 --> 0:19:30.200
<v Speaker 3>and yet they know that and they've proven this out

0:19:30.280 --> 0:19:33.760
<v Speaker 3>time and time again that by working with startups on

0:19:33.880 --> 0:19:36.639
<v Speaker 3>hard problems, they get better. And so you really want

0:19:36.920 --> 0:19:39.199
<v Speaker 3>you want partners like that. You want first customers that

0:19:39.240 --> 0:19:42.560
<v Speaker 3>are like that, that help you make your systems better,

0:19:42.600 --> 0:19:45.359
<v Speaker 3>that help you debug them, help you find problems that

0:19:45.359 --> 0:19:48.040
<v Speaker 3>you wouldn't have and resolve those problems you you wouldn't

0:19:48.040 --> 0:19:48.560
<v Speaker 3>have otherwise.

0:19:48.760 --> 0:19:50.400
<v Speaker 2>And so they've been a great partner to us.

0:19:50.400 --> 0:19:54.199
<v Speaker 3>But yeah, I think everyone in the world is asking

0:19:54.200 --> 0:19:56.600
<v Speaker 3>more of TSMC. They're having to make you know, investments

0:19:56.680 --> 0:20:00.600
<v Speaker 3>and you know, as is ASML and that the companies

0:20:00.600 --> 0:20:03.760
<v Speaker 3>that make their machines are also in high demand. And

0:20:03.800 --> 0:20:05.879
<v Speaker 3>I think the entire world is basically looking at this

0:20:05.880 --> 0:20:07.520
<v Speaker 3>market opportunity in front of us and saying, well, what

0:20:07.920 --> 0:20:09.679
<v Speaker 3>is the market for intelligence?

0:20:10.440 --> 0:20:12.400
<v Speaker 2>It's unbounded, right like you you.

0:20:12.359 --> 0:20:16.000
<v Speaker 3>Know, these are trillion dollar markets today that could be

0:20:16.040 --> 0:20:18.080
<v Speaker 3>one hundred trillion dollar markets in the future. And so

0:20:19.200 --> 0:20:22.320
<v Speaker 3>you know, working on for them, working on the fundamental

0:20:22.320 --> 0:20:24.480
<v Speaker 3>technology underlying that wave that intelligence.

0:20:24.720 --> 0:20:26.880
<v Speaker 2>They're in high demand, no question. But they've been great,

0:20:26.880 --> 0:20:27.840
<v Speaker 2>great partners.

0:20:27.520 --> 0:21:02.360
<v Speaker 4>To us.

0:20:48.560 --> 0:20:51.240
<v Speaker 1>On the demands side. As I will send it. In

0:20:51.240 --> 0:20:54.440
<v Speaker 1>twenty twenty five, between eighteen ninety percent of the revenues

0:20:54.600 --> 0:20:59.199
<v Speaker 1>came from setting chips to to UAE based entities g

0:20:59.320 --> 0:21:03.760
<v Speaker 1>foot two and the Muhammad Benzai a University of Artificial Intelligence.

0:21:04.280 --> 0:21:07.320
<v Speaker 1>I guess, first question what do they use those chips for?

0:21:07.720 --> 0:21:10.040
<v Speaker 1>And second question is how do you How are you

0:21:10.080 --> 0:21:13.119
<v Speaker 1>diversifying the business and are you focus on winning market

0:21:13.160 --> 0:21:17.400
<v Speaker 1>share from in video or just maintaining your position As

0:21:17.480 --> 0:21:21.160
<v Speaker 1>you put it, the kind of the market for intelligence grows.

0:21:21.920 --> 0:21:26.000
<v Speaker 3>Yeah, So I mean we got started really selling systems

0:21:26.080 --> 0:21:28.920
<v Speaker 3>to great partners early in the life cycle of a company.

0:21:29.240 --> 0:21:33.879
<v Speaker 3>We had partners with National Labs, we had pharmaceutical companies,

0:21:33.880 --> 0:21:36.840
<v Speaker 3>we had energy companies working with US buying systems, and

0:21:36.840 --> 0:21:39.159
<v Speaker 3>that was primarily in the training days of.

0:21:40.760 --> 0:21:43.920
<v Speaker 2>Three versus History. In twenty twenty four, we launched our

0:21:44.000 --> 0:21:44.840
<v Speaker 2>inference offering.

0:21:45.400 --> 0:21:47.240
<v Speaker 3>We built it, we kind of had a board meeting,

0:21:47.280 --> 0:21:51.680
<v Speaker 3>realized that as everyone was seeing, inference workloads were growing

0:21:52.000 --> 0:21:56.879
<v Speaker 3>very very quickly, beginning to sort of really dwarf the

0:21:56.920 --> 0:22:00.719
<v Speaker 3>training workloads, and so we ourselves realized that our architecture,

0:22:00.760 --> 0:22:03.320
<v Speaker 3>our way for scale engine was going to be really

0:22:03.320 --> 0:22:06.960
<v Speaker 3>appropriate for that. So we basically rotated towards an inference offering,

0:22:07.560 --> 0:22:11.040
<v Speaker 3>and as part of that, we developed a partnership with

0:22:11.640 --> 0:22:13.800
<v Speaker 3>G forty two, which is the entity you just described,

0:22:14.760 --> 0:22:18.199
<v Speaker 3>and they have been purchasing our systems and working with

0:22:18.240 --> 0:22:21.080
<v Speaker 3>our cloud infrastructure for their customers. So think of them

0:22:21.119 --> 0:22:23.640
<v Speaker 3>as sort of the AWS for the Middle East, and

0:22:24.119 --> 0:22:27.479
<v Speaker 3>they have been able to really work closely with their

0:22:27.520 --> 0:22:29.720
<v Speaker 3>own customers as well as some of our customers. In fact,

0:22:30.040 --> 0:22:32.880
<v Speaker 3>we have actually put a number of our customer workloads

0:22:32.920 --> 0:22:35.919
<v Speaker 3>back onto those systems that we that we put in

0:22:35.960 --> 0:22:37.440
<v Speaker 3>place now a couple of years ago.

0:22:37.920 --> 0:22:39.840
<v Speaker 2>And then your question around diversification.

0:22:39.960 --> 0:22:44.560
<v Speaker 3>So earlier this year we announced our deal with open ai,

0:22:45.240 --> 0:22:47.280
<v Speaker 3>which is more than a twenty billion dollar multi year

0:22:47.320 --> 0:22:49.720
<v Speaker 3>deal to build seven hundred and fifty megawatts of compute

0:22:50.720 --> 0:22:54.159
<v Speaker 3>and we were after announcing that we were only six

0:22:54.240 --> 0:22:56.720
<v Speaker 3>weeks to launching our first model with them. We love

0:22:56.760 --> 0:22:58.720
<v Speaker 3>working with them. I mean it's so great to see

0:22:58.720 --> 0:23:02.000
<v Speaker 3>too super high performance teams focused on, you know, making

0:23:02.000 --> 0:23:04.080
<v Speaker 3>the fastest inference on the planet, which is which is

0:23:04.080 --> 0:23:06.359
<v Speaker 3>what we launched and what we have continued to prove

0:23:06.400 --> 0:23:09.119
<v Speaker 3>out with you know, lots more to come there. So

0:23:09.640 --> 0:23:12.440
<v Speaker 3>open aye is probably the single largest sort of counterweight

0:23:12.480 --> 0:23:14.879
<v Speaker 3>to the question of you know, kind of customer concentration.

0:23:15.400 --> 0:23:17.400
<v Speaker 3>And then back in March of this year, we also

0:23:17.400 --> 0:23:20.960
<v Speaker 3>announced our partnership with AWS. So I think what you'll

0:23:20.960 --> 0:23:23.320
<v Speaker 3>see is over the coming quarters is you know, us

0:23:23.320 --> 0:23:26.800
<v Speaker 3>scaling those deployments with those customers. And then there's many

0:23:26.840 --> 0:23:29.399
<v Speaker 3>many more that are you know, kind of wanting to

0:23:29.400 --> 0:23:31.280
<v Speaker 3>work with us. And I think back in the early

0:23:31.320 --> 0:23:34.560
<v Speaker 3>two thousands, there are lots of companies building search products,

0:23:34.840 --> 0:23:36.560
<v Speaker 3>and Google happened to have the best one. It was

0:23:36.560 --> 0:23:39.400
<v Speaker 3>also one that was just ruthlessly focused on speed. In fact,

0:23:39.440 --> 0:23:41.360
<v Speaker 3>you remember the early days of Google, they would tell

0:23:41.359 --> 0:23:43.439
<v Speaker 3>you how many milliseconds it took to return the you know,

0:23:43.720 --> 0:23:46.000
<v Speaker 3>to return the query. And they did that because a

0:23:46.119 --> 0:23:48.160
<v Speaker 3>it was it was important to users. But if they

0:23:48.240 --> 0:23:50.159
<v Speaker 3>if they weren't fast, you know, their users were going

0:23:50.200 --> 0:23:52.760
<v Speaker 3>to go do other things. And so performance and inference

0:23:53.080 --> 0:23:56.399
<v Speaker 3>is absolutely fundamental. And so we while while we might

0:23:56.440 --> 0:23:58.600
<v Speaker 3>only be i'll call it sort of ten or fifteen

0:23:58.640 --> 0:24:00.560
<v Speaker 3>percent of the market today in terms of who wants

0:24:00.640 --> 0:24:03.040
<v Speaker 3>to have high performance, we think one hundred percent of

0:24:03.040 --> 0:24:05.320
<v Speaker 3>the market towards most towards high performance over time.

0:24:06.119 --> 0:24:10.960
<v Speaker 1>Obviously G forty two's connections with with with China cause

0:24:11.040 --> 0:24:13.639
<v Speaker 1>some consternation in the US, and I believe that was

0:24:13.720 --> 0:24:19.160
<v Speaker 1>part of the reason why this rebrass IPO was originally delayed.

0:24:19.480 --> 0:24:22.000
<v Speaker 1>So I'm curious for your kind of take on on

0:24:22.000 --> 0:24:24.080
<v Speaker 1>on the geopolitics of all this. And also, I mean,

0:24:24.480 --> 0:24:27.399
<v Speaker 1>you are a you know, a sort of roboticist and

0:24:27.480 --> 0:24:30.480
<v Speaker 1>an IDEO guy based in Silicon Valley with like a

0:24:30.520 --> 0:24:32.640
<v Speaker 1>real passion of engineering, and all of a sudden you're

0:24:33.000 --> 0:24:34.560
<v Speaker 1>you know, on the board and the lead investor in

0:24:34.600 --> 0:24:36.679
<v Speaker 1>this in this company that's in the middle of you know,

0:24:36.720 --> 0:24:40.080
<v Speaker 1>politics and geopolitics, how do you navigate that?

0:24:41.640 --> 0:24:44.600
<v Speaker 3>Yeah, so we I mean we see these you know,

0:24:44.680 --> 0:24:48.679
<v Speaker 3>these markets. Market again for intelligence is infinite, is unbounded,

0:24:49.240 --> 0:24:51.000
<v Speaker 3>and it's coming from everywhere in the world.

0:24:51.200 --> 0:24:52.760
<v Speaker 2>And and I.

0:24:52.760 --> 0:24:55.200
<v Speaker 3>Would also assert that the you know, the geopolitics of

0:24:55.240 --> 0:24:58.080
<v Speaker 3>the last four or five years have only made those

0:24:58.119 --> 0:25:04.520
<v Speaker 3>international parties more to get access to alternative systems, alternative

0:25:04.520 --> 0:25:07.040
<v Speaker 3>technologies than just the sort of you know, you know,

0:25:07.119 --> 0:25:09.080
<v Speaker 3>kind of think of it as the traditional hyper scalers

0:25:09.119 --> 0:25:11.680
<v Speaker 3>or Nvidia for that matter. I mean, we have prospective

0:25:11.880 --> 0:25:16.720
<v Speaker 3>customers in Europe and they're desperately seeking any alternative to

0:25:17.119 --> 0:25:21.080
<v Speaker 3>purchasing more in video gear. And so there's I think

0:25:21.480 --> 0:25:24.040
<v Speaker 3>a really sort of a pull around the world for

0:25:24.200 --> 0:25:28.239
<v Speaker 3>alternatives that are coming from parties who don't feel like

0:25:28.280 --> 0:25:29.639
<v Speaker 3>they've they've got eighty.

0:25:29.440 --> 0:25:30.720
<v Speaker 2>Or eighty five percent market share.

0:25:31.119 --> 0:25:33.679
<v Speaker 1>Why because if fear that they're basically not going to

0:25:33.720 --> 0:25:35.760
<v Speaker 1>get they're not going to be a big enough client

0:25:35.800 --> 0:25:37.680
<v Speaker 1>to get preferential.

0:25:37.400 --> 0:25:39.080
<v Speaker 2>Get the allocations they are.

0:25:39.280 --> 0:25:41.960
<v Speaker 3>Yeah, there's you know, these are you know, you're you're

0:25:42.320 --> 0:25:45.639
<v Speaker 3>you're seeing companies have to and and sovereign governments have

0:25:45.680 --> 0:25:48.400
<v Speaker 3>to come to here in Silicon Valley and I mean

0:25:48.760 --> 0:25:50.600
<v Speaker 3>Jensen's made more than a few trips to the Middle

0:25:50.640 --> 0:25:52.800
<v Speaker 3>East over the last few years. But you know before

0:25:52.840 --> 0:25:54.800
<v Speaker 3>that they were all coming here to kiss the ring

0:25:54.840 --> 0:25:55.520
<v Speaker 3>and get allocation.

0:25:56.800 --> 0:26:00.600
<v Speaker 1>Well knows you about the open a ideal. Obviously huge

0:26:00.600 --> 0:26:03.720
<v Speaker 1>announcement a few guys, and then the market was sort

0:26:03.720 --> 0:26:05.840
<v Speaker 1>of questioning, oh, is it going to be more expensive

0:26:05.880 --> 0:26:08.480
<v Speaker 1>than they thought to actually service this deal which is

0:26:08.520 --> 0:26:10.920
<v Speaker 1>not a not a problem, which is you know, confined

0:26:10.920 --> 0:26:13.760
<v Speaker 1>to cerebras. I know, Google it getting beaten up for

0:26:13.840 --> 0:26:16.520
<v Speaker 1>the kind of cost of deployment of that cloud computing.

0:26:16.600 --> 0:26:18.280
<v Speaker 1>But how do you think about that?

0:26:19.440 --> 0:26:21.679
<v Speaker 3>Yeah, So we it was interesting because when you know,

0:26:21.680 --> 0:26:25.040
<v Speaker 3>we did the IPO roadshow, we shared all of that

0:26:25.600 --> 0:26:28.640
<v Speaker 3>in conversations as well as in the S one, and

0:26:28.880 --> 0:26:30.640
<v Speaker 3>you know, there was this question mostly in my sense

0:26:30.760 --> 0:26:33.760
<v Speaker 3>was it was folks either not reading or perhaps misunderstanding,

0:26:33.840 --> 0:26:36.960
<v Speaker 3>or just looking for ways to sort of throw bricks.

0:26:36.680 --> 0:26:38.240
<v Speaker 2>At young companies.

0:26:38.480 --> 0:26:40.760
<v Speaker 3>So we basically had a home run up and down

0:26:40.800 --> 0:26:43.760
<v Speaker 3>the P and L beat really every metric. And the

0:26:43.880 --> 0:26:46.440
<v Speaker 3>question that folks were raising was around gross margin, which

0:26:46.480 --> 0:26:48.720
<v Speaker 3>was off by I think ten or fifteen percent, and

0:26:48.760 --> 0:26:51.840
<v Speaker 3>that was entirely because we were renting back capacity over

0:26:51.880 --> 0:26:54.080
<v Speaker 3>the next few quarters from G forty two, which is

0:26:54.119 --> 0:26:56.440
<v Speaker 3>the thing we talked about earlier. So we had you know,

0:26:56.640 --> 0:26:59.320
<v Speaker 3>they had deployed systems, we had customers who wanted more,

0:26:59.600 --> 0:27:01.679
<v Speaker 3>and so we ultimately, I mean, we put those systems

0:27:01.720 --> 0:27:03.800
<v Speaker 3>in place, and then we've developed an agreement with them

0:27:03.800 --> 0:27:05.919
<v Speaker 3>which was, hey, can we rent back that capacity for

0:27:05.960 --> 0:27:08.080
<v Speaker 3>these customers, which they were they were happy to do,

0:27:08.119 --> 0:27:09.920
<v Speaker 3>and you know, we had to take a gross margin

0:27:10.000 --> 0:27:12.080
<v Speaker 3>hit in the in the near term for that. So

0:27:12.400 --> 0:27:15.240
<v Speaker 3>this was not a surprise for any of our institutional investors.

0:27:15.240 --> 0:27:17.320
<v Speaker 3>But you know, folks are looking for, you know, whatever

0:27:17.440 --> 0:27:18.920
<v Speaker 3>might be a crack in the story for your first

0:27:18.960 --> 0:27:19.439
<v Speaker 3>earnings call.

0:27:19.480 --> 0:27:21.439
<v Speaker 2>But yeah, this was this was a big nothing burger

0:27:21.520 --> 0:27:22.080
<v Speaker 2>in my mind.

0:27:23.680 --> 0:27:25.080
<v Speaker 1>And what about I mean, how do you look at

0:27:25.119 --> 0:27:30.879
<v Speaker 1>the competitive dynamics beyond Nvidia obviously, Microsoft, Google, Amazon, Meta

0:27:31.160 --> 0:27:34.879
<v Speaker 1>all now rushing into the space of chip design and

0:27:35.040 --> 0:27:38.919
<v Speaker 1>and and you know for their specialized use cases in

0:27:38.920 --> 0:27:41.159
<v Speaker 1>many case which I mentioned on how computeeing. How do

0:27:41.200 --> 0:27:43.320
<v Speaker 1>you look at what they're doing, are you worried about it?

0:27:44.080 --> 0:27:46.280
<v Speaker 1>And how do you maintain the edge?

0:27:47.240 --> 0:27:50.840
<v Speaker 3>Yeah, so I think broadly the way you think about

0:27:50.840 --> 0:27:53.080
<v Speaker 3>these things is you know, large markets like this, and

0:27:53.119 --> 0:27:54.920
<v Speaker 3>as we said, these are the mother of all markets,

0:27:55.359 --> 0:27:57.560
<v Speaker 3>you're never going to have them all to yourself. You know.

0:27:57.600 --> 0:27:59.800
<v Speaker 3>If you're right about these things, they're going to attract competition.

0:27:59.800 --> 0:28:05.120
<v Speaker 3>That's just basics and you know economics. And so what

0:28:05.160 --> 0:28:09.680
<v Speaker 3>we observe is every one of the hyperscalers has already

0:28:09.680 --> 0:28:12.320
<v Speaker 3>developed or is it you know, in development on their

0:28:12.359 --> 0:28:15.160
<v Speaker 3>own silicon. I mean, obviously Google's had their TPUs in place.

0:28:15.560 --> 0:28:18.640
<v Speaker 3>They're on their fourth or fifth generation now, the Trainium

0:28:18.920 --> 0:28:22.800
<v Speaker 3>and Infernia platform at AWS, and again we're in partnership

0:28:22.800 --> 0:28:25.919
<v Speaker 3>with AWS. Fact you know, we've we've shared some of

0:28:25.960 --> 0:28:28.240
<v Speaker 3>the ways in which we're working together on a very

0:28:28.280 --> 0:28:30.800
<v Speaker 3>novel disaggregated solution where we can kind of get the

0:28:30.840 --> 0:28:32.720
<v Speaker 3>best of their hardware and the best of our hardware,

0:28:33.240 --> 0:28:35.119
<v Speaker 3>which is where I think it will be really interesting

0:28:35.119 --> 0:28:35.560
<v Speaker 3>things to sort of.

0:28:35.560 --> 0:28:36.399
<v Speaker 2>See that play forward.

0:28:36.560 --> 0:28:39.640
<v Speaker 3>Facebook of course, has been talking about their own platform

0:28:39.680 --> 0:28:41.600
<v Speaker 3>to build their own silicon, which you know they've been

0:28:41.600 --> 0:28:44.640
<v Speaker 3>doing in partnership with Broadcom for some time. And so

0:28:45.080 --> 0:28:47.200
<v Speaker 3>you know, we see this as any large market is

0:28:47.240 --> 0:28:50.760
<v Speaker 3>going to attract multiple players, and what I actually really

0:28:50.800 --> 0:28:52.960
<v Speaker 3>appreciate about it is everyone's at least looking at their

0:28:52.960 --> 0:28:54.880
<v Speaker 3>workloads and saying, what is it that we should build

0:28:54.880 --> 0:28:57.960
<v Speaker 3>here that's well suited to us, And what we're building

0:28:58.000 --> 0:29:01.080
<v Speaker 3>at THREEBRUS is very well suited these inference workloads. It's

0:29:01.080 --> 0:29:03.760
<v Speaker 3>also well suited to training, and it's you know, it's

0:29:03.800 --> 0:29:05.640
<v Speaker 3>something that took as I said, you know, it took

0:29:05.720 --> 0:29:09.040
<v Speaker 3>us multiple years in three generations before it really felt

0:29:09.080 --> 0:29:12.320
<v Speaker 3>like we had a system that was really singing. And

0:29:12.520 --> 0:29:15.320
<v Speaker 3>most of the other insurgents, you know, are still telling

0:29:15.320 --> 0:29:16.840
<v Speaker 3>you what's about to ship you know whatever in the

0:29:16.880 --> 0:29:18.160
<v Speaker 3>fall or in Q one of.

0:29:18.080 --> 0:29:18.800
<v Speaker 2>Twenty twenty seven.

0:29:18.800 --> 0:29:20.480
<v Speaker 3>But we, you know, we have systems out in the wild,

0:29:20.480 --> 0:29:23.360
<v Speaker 3>and we're scaling dramatically, and we're also working on our

0:29:23.360 --> 0:29:25.680
<v Speaker 3>next generation systems, which are which are super exciting.

0:29:25.760 --> 0:29:26.760
<v Speaker 2>So that's how I think about it.

0:29:26.800 --> 0:29:29.080
<v Speaker 3>I think there's going to be many many more, could

0:29:29.080 --> 0:29:32.880
<v Speaker 3>be more upstarts. You know this, this this market is massive,

0:29:32.920 --> 0:29:35.440
<v Speaker 3>and so we should assume there's gonna be many, many winners.

0:29:36.120 --> 0:29:40.560
<v Speaker 1>So you have this claim now, which which all vcs

0:29:40.600 --> 0:29:42.760
<v Speaker 1>would love to be able to make. That you essentially

0:29:43.640 --> 0:29:46.480
<v Speaker 1>soar around the corner, right, and so that must be

0:29:46.560 --> 0:29:51.560
<v Speaker 1>a fun moment for you. And I'm curious about two things.

0:29:52.320 --> 0:29:57.720
<v Speaker 1>One is, you know, ten years ago you saw this

0:29:57.920 --> 0:30:02.080
<v Speaker 1>kind of very strong demand see all for you big

0:30:02.120 --> 0:30:06.360
<v Speaker 1>data processing what became AI workloads and invested in and

0:30:06.400 --> 0:30:10.160
<v Speaker 1>built Cerebrass entergy for thease opportunity. What are you seeing

0:30:10.240 --> 0:30:14.480
<v Speaker 1>now ten years into the future in terms of demand

0:30:14.720 --> 0:30:18.400
<v Speaker 1>signals that you're investing in, either through your fund or

0:30:18.520 --> 0:30:19.640
<v Speaker 1>as part of Cerebress.

0:30:20.080 --> 0:30:23.240
<v Speaker 3>Well, when I think about what today we're excited about,

0:30:23.400 --> 0:30:26.120
<v Speaker 3>I mean what you see in our industry. I think

0:30:26.120 --> 0:30:28.240
<v Speaker 3>in venture capital and tech more broadly, and you've seen

0:30:28.240 --> 0:30:29.240
<v Speaker 3>this as much.

0:30:29.120 --> 0:30:29.480
<v Speaker 2>As I have.

0:30:29.680 --> 0:30:34.880
<v Speaker 3>Is every wave, every platform, new platform technology, whether you

0:30:34.920 --> 0:30:36.400
<v Speaker 3>know we were talking about sort of the X eighty

0:30:36.400 --> 0:30:41.400
<v Speaker 3>six platform or GPUs or mobile and now AI. There's

0:30:41.440 --> 0:30:43.640
<v Speaker 3>a sort of quite roughly a decade of building out

0:30:43.640 --> 0:30:47.680
<v Speaker 3>the infrastructure and then there's multiple decades of building out

0:30:47.680 --> 0:30:50.760
<v Speaker 3>the application layer on top of it. And so there's

0:30:50.800 --> 0:30:53.840
<v Speaker 3>no question and we're spending you know, we more than

0:30:53.880 --> 0:30:57.040
<v Speaker 3>one hundred companies broadly, and that you describe as sort

0:30:57.080 --> 0:31:00.160
<v Speaker 3>of AIRAI native, you know, across our portfolio where a

0:31:00.160 --> 0:31:02.600
<v Speaker 3>thirty one year old venture firm we've been investing in

0:31:02.640 --> 0:31:05.160
<v Speaker 3>AI really sort of the sort of fundamentals of it

0:31:05.160 --> 0:31:07.080
<v Speaker 3>since two thousand and eight two thousand and nine time frame,

0:31:07.640 --> 0:31:11.280
<v Speaker 3>and so we look at this next waves as involving

0:31:11.320 --> 0:31:13.520
<v Speaker 3>a lot of innovation at the application layer.

0:31:13.560 --> 0:31:15.640
<v Speaker 2>Now what's tricky about this, as I'm.

0:31:15.480 --> 0:31:17.880
<v Speaker 3>Sure you're aware of as well, is well which parts

0:31:17.920 --> 0:31:21.160
<v Speaker 3>of those application layer are going to get completely commoditized

0:31:21.200 --> 0:31:24.240
<v Speaker 3>by perhaps even the fundamental foundation models that you're building

0:31:24.240 --> 0:31:27.960
<v Speaker 3>on top of. And so we look for the sort

0:31:28.000 --> 0:31:30.680
<v Speaker 3>of thin, sort of proprietary data loops, you know, ways

0:31:30.680 --> 0:31:33.080
<v Speaker 3>in which there's a data flywheel where with every use

0:31:33.120 --> 0:31:36.400
<v Speaker 3>of the product, you know, your insight around your customer

0:31:36.520 --> 0:31:39.520
<v Speaker 3>and your value proposition gets deeper and richer. And we

0:31:39.560 --> 0:31:41.960
<v Speaker 3>look for sort of really interesting ways in which things

0:31:42.000 --> 0:31:45.320
<v Speaker 3>that were today delivered as a service, oftentimes literally as

0:31:45.560 --> 0:31:47.480
<v Speaker 3>maybe it's a software service, maybe it's a human service,

0:31:47.520 --> 0:31:50.920
<v Speaker 3>maybe it's labor, and we look for ways that that

0:31:51.040 --> 0:31:53.520
<v Speaker 3>is going to be automated over time. And so we've

0:31:53.560 --> 0:31:55.800
<v Speaker 3>made you know, a number of really interesting investments in

0:31:55.840 --> 0:31:57.600
<v Speaker 3>this area. And then when it comes to sort of

0:31:57.640 --> 0:32:01.280
<v Speaker 3>AI infrastructure, I mean we all will have to kind

0:32:01.280 --> 0:32:03.560
<v Speaker 3>of like point to the you know, the the analogies

0:32:03.600 --> 0:32:06.400
<v Speaker 3>to how are you know natural systems work, whether it's

0:32:06.400 --> 0:32:08.600
<v Speaker 3>a human brain or not. But you know, these systems

0:32:08.600 --> 0:32:13.960
<v Speaker 3>that we're building are far less sample efficient than the

0:32:13.960 --> 0:32:15.080
<v Speaker 3>way natural systems work.

0:32:15.120 --> 0:32:15.240
<v Speaker 4>You know.

0:32:15.280 --> 0:32:18.120
<v Speaker 3>You know, if you've got a twenty whatt brain operating

0:32:18.120 --> 0:32:20.440
<v Speaker 3>at two hundred hertz and it can do, you know,

0:32:20.560 --> 0:32:24.320
<v Speaker 3>sample efficiency is much much higher than what is used

0:32:24.320 --> 0:32:28.640
<v Speaker 3>to train a very intelligent next generation frontier model.

0:32:29.080 --> 0:32:32.160
<v Speaker 1>Some of thefsionist means that you can learn from fewer examples.

0:32:31.680 --> 0:32:32.400
<v Speaker 2>Yeah, exactly.

0:32:32.520 --> 0:32:34.680
<v Speaker 3>Yeah, and so I think there's going to be some

0:32:34.800 --> 0:32:38.280
<v Speaker 3>really interesting work in this area that then allows us

0:32:38.280 --> 0:32:41.160
<v Speaker 3>to build models that are more efficient both from an.

0:32:41.120 --> 0:32:43.240
<v Speaker 2>Energy perspective but from a data perspective.

0:32:43.640 --> 0:32:46.360
<v Speaker 3>So yeah, I'm very excited about, you know, some of

0:32:46.360 --> 0:32:47.640
<v Speaker 3>the work that's being done in this domain.

0:32:48.280 --> 0:32:51.080
<v Speaker 1>Steve has outa thank you, thanks so much, as it's

0:32:51.120 --> 0:33:11.920
<v Speaker 1>just so great to talk with you for tech stuff.

0:33:12.000 --> 0:33:15.880
<v Speaker 1>I'm Osvoloshin. This episode was produced by Eliza Dennis. It

0:33:15.960 --> 0:33:19.280
<v Speaker 1>was executive produced by me and Julian Nutta for Kaleidoscope

0:33:19.640 --> 0:33:23.960
<v Speaker 1>and Katrian norvelve iHeart Podcasts. Jack Qinsley mixed this episode

0:33:24.000 --> 0:33:25.800
<v Speaker 1>and Kyle Murdoch wrote Olph theme song,