WEBVTT - Sound (featuring audio wizard Mat Keselman)

0:00:07.800 --> 0:00:11.799
<v Speaker 1>How any extraordinaries. Today, we're pulling back the dKu curtin

0:00:11.840 --> 0:00:14.200
<v Speaker 1>a bit. We're gonna see how the sausage is made,

0:00:14.280 --> 0:00:17.720
<v Speaker 1>as it were. So Daniel and I write pretty darn

0:00:17.800 --> 0:00:21.440
<v Speaker 1>detailed outlines before we record an episode, but often something

0:00:21.480 --> 0:00:25.560
<v Speaker 1>happens that totally derails our outline conversation, and that's good.

0:00:26.040 --> 0:00:29.520
<v Speaker 1>Sometimes the derailment happens because something I thought was kind

0:00:29.520 --> 0:00:32.120
<v Speaker 1>of obvious turns out to not be so obvious, and

0:00:32.240 --> 0:00:35.000
<v Speaker 1>Daniel will ask clarifying questions that forced me to come

0:00:35.080 --> 0:00:40.000
<v Speaker 1>up with different, better explanations on the fly. Or Daniel

0:00:40.000 --> 0:00:42.080
<v Speaker 1>will be working through his outline and I'll think of

0:00:42.120 --> 0:00:45.680
<v Speaker 1>some tangentially related topic that I now really want to

0:00:45.720 --> 0:00:48.559
<v Speaker 1>know about, and Daniel will hop online to do a

0:00:48.640 --> 0:00:51.520
<v Speaker 1>quick bit of research, and then we'll record a completely

0:00:51.680 --> 0:00:55.760
<v Speaker 1>unplanned segment. In each of these cases, our answers are

0:00:55.920 --> 0:00:59.280
<v Speaker 1>way less polished than we'd like, Lots of false starts,

0:00:59.400 --> 0:01:03.520
<v Speaker 1>lots of basically lots of stuff that the audience definitely

0:01:03.760 --> 0:01:06.880
<v Speaker 1>doesn't want to have to sit through. And then Daniel

0:01:06.880 --> 0:01:09.720
<v Speaker 1>and I have unexpected noises that intrude upon the audio.

0:01:10.120 --> 0:01:13.759
<v Speaker 1>In today's episode, I learned that my microphone actually does

0:01:13.840 --> 0:01:16.199
<v Speaker 1>pick up the sounds my goats make when they're trying

0:01:16.240 --> 0:01:19.520
<v Speaker 1>to get my attention, and sometimes Daniel apparently records from

0:01:19.560 --> 0:01:22.959
<v Speaker 1>the inside of a grocery store. More on that later. Anyway,

0:01:23.280 --> 0:01:25.720
<v Speaker 1>these are only a handful of the many challenges that

0:01:25.760 --> 0:01:28.840
<v Speaker 1>we throw at Matt Kesselman, the audio wizard who makes

0:01:28.880 --> 0:01:32.520
<v Speaker 1>Daniel and Kelly's Extraordinary Universe sound so good. So today

0:01:32.520 --> 0:01:34.720
<v Speaker 1>we're bringing Matt on the show to talk about sound

0:01:35.000 --> 0:01:37.320
<v Speaker 1>and how he goes about fixing hours so that your

0:01:37.400 --> 0:01:40.479
<v Speaker 1>listening experience is so much better that it would be

0:01:40.520 --> 0:01:42.720
<v Speaker 1>if you were a fly on the wall while Daniel

0:01:42.720 --> 0:02:00.480
<v Speaker 1>and I record. Welcome to Daniel and Kelly's Edited Universe. Hi.

0:02:00.600 --> 0:02:03.120
<v Speaker 2>I'm Daniel. I'm a particle physicist and I like thinking

0:02:03.160 --> 0:02:06.000
<v Speaker 2>about aliens, and I've often been told I have a

0:02:06.120 --> 0:02:07.160
<v Speaker 2>face for radio.

0:02:13.080 --> 0:02:16.560
<v Speaker 1>Hi. I'm Kelly leader Smith. I study parasites in space.

0:02:17.600 --> 0:02:20.800
<v Speaker 1>I often make jokes about having a face for radio.

0:02:20.880 --> 0:02:22.960
<v Speaker 1>But I think actually Daniel and I are but quite

0:02:23.000 --> 0:02:27.160
<v Speaker 1>attractive human beings. Oh yes, and Daniel has a great

0:02:27.280 --> 0:02:29.960
<v Speaker 1>voice for radio, So let's focus on the positive.

0:02:29.840 --> 0:02:32.040
<v Speaker 2>As do you, Kelly. I think our voices work really

0:02:32.080 --> 0:02:32.640
<v Speaker 2>well together.

0:02:32.919 --> 0:02:34.919
<v Speaker 1>I appreciate it. But I will note that when we

0:02:35.320 --> 0:02:39.120
<v Speaker 1>transitioned from Jorge on the show to Kelly on the Show,

0:02:39.480 --> 0:02:42.880
<v Speaker 1>we did get a fair number of folks writing in

0:02:42.919 --> 0:02:45.799
<v Speaker 1>to complain about the sound of my voice. But here

0:02:45.880 --> 0:02:49.680
<v Speaker 1>I am over one hundred episodes later. Kiss my butt, loosers.

0:02:49.840 --> 0:02:53.160
<v Speaker 1>I'm not going anywhere hater's gonna hate.

0:02:53.520 --> 0:02:55.400
<v Speaker 2>But we also got a lot of comments from people

0:02:55.440 --> 0:02:57.280
<v Speaker 2>who love your voice. And I think it's just a

0:02:57.320 --> 0:03:00.240
<v Speaker 2>subjective thing, you know, And it's a big part of

0:03:00.320 --> 0:03:02.760
<v Speaker 2>finding a podcast you'd like listening to. Is are the

0:03:02.880 --> 0:03:05.800
<v Speaker 2>voices pleasant and soothing? And that's very personal?

0:03:06.000 --> 0:03:08.440
<v Speaker 1>You know, That's right. But the good news is you

0:03:08.520 --> 0:03:11.840
<v Speaker 1>and I have Matt Kesselman to help us whenever our

0:03:11.919 --> 0:03:14.200
<v Speaker 1>voices are sounding less than their best.

0:03:15.760 --> 0:03:17.440
<v Speaker 2>That's right. And I want my voice to be smooth,

0:03:17.480 --> 0:03:18.880
<v Speaker 2>because a lot of folks write in and tell me

0:03:19.120 --> 0:03:22.120
<v Speaker 2>they like to fall asleep to my voice, and hey,

0:03:22.160 --> 0:03:25.080
<v Speaker 2>you know, I don't want to disturb anybody's slumber. No.

0:03:25.280 --> 0:03:28.040
<v Speaker 1>I wonder what my Calkuli laughed as to their dreams.

0:03:28.560 --> 0:03:32.280
<v Speaker 1>I dream of witches. But today we decided that we

0:03:32.280 --> 0:03:34.760
<v Speaker 1>were going to invite Matt Kesselman, who is just an

0:03:34.800 --> 0:03:38.240
<v Speaker 1>absolutely amazing audio engineer who, thankfully for us, spends a

0:03:38.280 --> 0:03:40.280
<v Speaker 1>lot of time working on our show. And part of

0:03:40.280 --> 0:03:42.480
<v Speaker 1>why we decided to invite him is because anyone who

0:03:42.600 --> 0:03:46.080
<v Speaker 1>submits a question that ends up on our Listener Questions

0:03:46.120 --> 0:03:49.760
<v Speaker 1>episode gets a raw audio file of Daniel and I

0:03:49.880 --> 0:03:52.320
<v Speaker 1>having our conversation, and we send it to them before

0:03:52.440 --> 0:03:55.120
<v Speaker 1>it goes to Matt, our audio engineer, so that they

0:03:55.160 --> 0:03:57.400
<v Speaker 1>have a chance to record themselves asking the question and

0:03:57.440 --> 0:04:02.080
<v Speaker 1>record themselves responding to our response to their question, and

0:04:02.120 --> 0:04:03.480
<v Speaker 1>then we can send all of that to Matt and

0:04:03.520 --> 0:04:05.880
<v Speaker 1>he can edit it all at once. And people often

0:04:05.960 --> 0:04:10.000
<v Speaker 1>respond to us and say, wow, your audio engineer has

0:04:10.040 --> 0:04:10.560
<v Speaker 1>a lot.

0:04:10.400 --> 0:04:14.880
<v Speaker 2>Of work to do, but not just that. Also they

0:04:14.960 --> 0:04:18.520
<v Speaker 2>say that it's really interesting to hear the unedited version

0:04:18.560 --> 0:04:21.680
<v Speaker 2>of the conversation because they can hear our side conversations

0:04:21.760 --> 0:04:24.080
<v Speaker 2>or when we pause to do research, or when we

0:04:24.120 --> 0:04:26.760
<v Speaker 2>back up to say something again, so they get like

0:04:26.760 --> 0:04:29.320
<v Speaker 2>a little glimpse behind the curtain. And so we want

0:04:29.360 --> 0:04:31.080
<v Speaker 2>to provide that to all of you today.

0:04:31.240 --> 0:04:33.800
<v Speaker 1>That's right, and so in the second segment of the show,

0:04:33.839 --> 0:04:35.560
<v Speaker 1>Matt is going to come on and talk to you

0:04:35.640 --> 0:04:38.440
<v Speaker 1>about the magic that he works on our episode, and

0:04:38.480 --> 0:04:41.640
<v Speaker 1>he'll play a couple raw clips so you'll get a

0:04:41.640 --> 0:04:43.919
<v Speaker 1>peek behind the curtain and see how Daniel and I

0:04:44.440 --> 0:04:49.400
<v Speaker 1>sound before Matt steps in. And Matt sent a question

0:04:49.920 --> 0:04:53.039
<v Speaker 1>for us to share with the listeners, and the question

0:04:53.560 --> 0:04:57.719
<v Speaker 1>was what kind of sound do humans hear best? So

0:04:57.800 --> 0:05:00.720
<v Speaker 1>let's go ahead and hear what the extraordinaries to say.

0:05:01.040 --> 0:05:05.040
<v Speaker 3>Maybe like the middle sounds like not super high or

0:05:05.040 --> 0:05:05.560
<v Speaker 3>super low.

0:05:05.640 --> 0:05:07.280
<v Speaker 1>That's what I was gonna say. But the humans are

0:05:07.279 --> 0:05:07.760
<v Speaker 1>in the middle.

0:05:08.080 --> 0:05:11.880
<v Speaker 2>I don't know what middle of what though, depending on

0:05:11.920 --> 0:05:16.520
<v Speaker 2>their gender, Like women would hear babies crying the best,

0:05:16.520 --> 0:05:20.360
<v Speaker 2>because I've heard stories that they often are moiasily working

0:05:20.440 --> 0:05:24.360
<v Speaker 2>up by that related to our size, Larger animals can

0:05:24.480 --> 0:05:28.400
<v Speaker 2>hear sounds with longer wavelengths. Smaller animals can hear sounds

0:05:28.440 --> 0:05:32.720
<v Speaker 2>with shorter wavelength high beach that sounds. Females tend to

0:05:32.760 --> 0:05:37.800
<v Speaker 2>hear higher frequencies better, possibly to help them hear their

0:05:37.839 --> 0:05:41.200
<v Speaker 2>offspring when they're crying or something similar.

0:05:41.560 --> 0:05:45.000
<v Speaker 1>Sounds that already stand out from the background and from

0:05:45.080 --> 0:05:45.560
<v Speaker 1>each other.

0:05:45.920 --> 0:05:48.120
<v Speaker 2>As a new mom, I want to say a baby crying.

0:05:48.520 --> 0:05:51.440
<v Speaker 2>These are great guesses and hilarious stories. Thank you very

0:05:51.520 --> 0:05:54.200
<v Speaker 2>much exterdinaries, and they're basically what I would have guessed.

0:05:54.240 --> 0:05:56.120
<v Speaker 2>You know that there's some range we can hear and

0:05:56.160 --> 0:05:59.400
<v Speaker 2>obviously ranges we can't hear, and that's going to be

0:05:59.480 --> 0:06:00.880
<v Speaker 2>tuned to our survival.

0:06:01.200 --> 0:06:04.560
<v Speaker 1>Yes, biology is the answer, and that's great. That's how

0:06:04.600 --> 0:06:06.080
<v Speaker 1>we know that. It's an interesting question.

0:06:06.279 --> 0:06:09.440
<v Speaker 2>No, it's a harmonious blend of biology and physics. It's

0:06:09.480 --> 0:06:11.480
<v Speaker 2>biology taking advantage of.

0:06:11.520 --> 0:06:16.080
<v Speaker 1>Physics, biophysics biofirst, but yes, biophysics.

0:06:17.920 --> 0:06:20.840
<v Speaker 2>You know, I think if biophysics means biology, it's modifying physics.

0:06:20.880 --> 0:06:22.839
<v Speaker 2>So it's at its core physics.

0:06:23.480 --> 0:06:24.880
<v Speaker 1>Well, you know, you and I could be at this

0:06:24.960 --> 0:06:28.120
<v Speaker 1>all day, and so we thought before we invited Matt

0:06:28.160 --> 0:06:30.040
<v Speaker 1>onto the show, we would let Daniel do a little

0:06:30.040 --> 0:06:33.120
<v Speaker 1>bit of his physics magic and talk about the physics

0:06:33.200 --> 0:06:35.440
<v Speaker 1>of sound. So we're going to dig into that a

0:06:35.440 --> 0:06:39.480
<v Speaker 1>little bit, and then we are going to bring Matt on. So, Daniel,

0:06:39.720 --> 0:06:40.680
<v Speaker 1>what is sound?

0:06:41.360 --> 0:06:44.440
<v Speaker 2>Yeah, so sound is not a thing. It's like an

0:06:44.640 --> 0:06:48.400
<v Speaker 2>arrangement of stuff. It's a traveling pattern of pressure changes

0:06:48.920 --> 0:06:54.080
<v Speaker 2>in matter. So if something is vibrating like a guitar string,

0:06:54.560 --> 0:06:58.080
<v Speaker 2>it's going to push and pull on nearby molecules creating

0:06:58.200 --> 0:07:01.720
<v Speaker 2>regions of compression and the regions when things are less dense,

0:07:02.320 --> 0:07:05.240
<v Speaker 2>and so those pressure changes then move through the medium.

0:07:05.320 --> 0:07:08.520
<v Speaker 2>Like if you slap your hand on the table, you're

0:07:08.520 --> 0:07:12.520
<v Speaker 2>creating momentary high pressure in the molecules on that table,

0:07:12.560 --> 0:07:13.960
<v Speaker 2>and they push on the ones next to it, and

0:07:13.960 --> 0:07:15.680
<v Speaker 2>they push on the ones next to it, so that

0:07:15.680 --> 0:07:19.000
<v Speaker 2>pressure wave travels right. And so this is a wave

0:07:19.080 --> 0:07:22.240
<v Speaker 2>that needs a medium. It moves through a medium. It's

0:07:22.240 --> 0:07:27.160
<v Speaker 2>a rearrangement through time of the relationship between the molecules

0:07:27.200 --> 0:07:27.880
<v Speaker 2>in the medium.

0:07:28.280 --> 0:07:32.160
<v Speaker 1>So I'm feeling even more amazed now by the microphone

0:07:32.200 --> 0:07:37.120
<v Speaker 1>that's sitting on my desk. How does those like squishiness

0:07:37.360 --> 0:07:41.720
<v Speaker 1>of the air get transferred into something my microphone can

0:07:41.880 --> 0:07:42.640
<v Speaker 1>play back to you.

0:07:43.240 --> 0:07:46.800
<v Speaker 2>Yes, so a speaker and a microphone are basically the

0:07:46.800 --> 0:07:49.400
<v Speaker 2>inverse of each other, right, So let's do a speaker first.

0:07:49.920 --> 0:07:53.640
<v Speaker 2>A speaker creates the sound. So there's something inside the speaker,

0:07:53.800 --> 0:07:56.920
<v Speaker 2>like a cone that moves back and forth. It's literally

0:07:56.960 --> 0:08:00.960
<v Speaker 2>physically mechanically moving at the frequent you want to create.

0:08:01.200 --> 0:08:04.160
<v Speaker 2>So we think about sound in terms of frequencies the

0:08:04.160 --> 0:08:06.800
<v Speaker 2>same way we think about light in terms of frequencies

0:08:07.120 --> 0:08:09.240
<v Speaker 2>you can have like red light and blue light or

0:08:09.280 --> 0:08:12.600
<v Speaker 2>green light. They are all vibrations of the electromagnetic field

0:08:12.640 --> 0:08:15.920
<v Speaker 2>at different frequencies. Sound has frequencies also, you can have

0:08:16.000 --> 0:08:20.760
<v Speaker 2>low frequencies and high frequencies, Daniel frequencies and Kelly frequencies exactly.

0:08:21.360 --> 0:08:25.120
<v Speaker 2>And so a speaker can make these by shaking effectively

0:08:25.120 --> 0:08:28.480
<v Speaker 2>a drum or a cone at that frequency. So if

0:08:28.520 --> 0:08:31.720
<v Speaker 2>you shake the cone at four hundred and forty hertz,

0:08:31.760 --> 0:08:34.480
<v Speaker 2>it's going to create four hundred and forty hertz sound,

0:08:34.960 --> 0:08:37.600
<v Speaker 2>just the same way that like a radio antenna oscillates

0:08:37.640 --> 0:08:41.439
<v Speaker 2>electrons at a certain frequency and creates photons or electromagnetic

0:08:41.440 --> 0:08:44.880
<v Speaker 2>waves at that frequency. So you have mechanical vibration. In

0:08:44.920 --> 0:08:47.959
<v Speaker 2>the case of sound transmits pressure waves to the air.

0:08:48.559 --> 0:08:52.000
<v Speaker 2>The microphone is essentially inverse pressure waves in the air

0:08:52.400 --> 0:08:56.040
<v Speaker 2>vibrates something sensitive, So you're converting pressure waves in the

0:08:56.040 --> 0:08:59.400
<v Speaker 2>air to physical vibrations of something, and then you just

0:08:59.440 --> 0:09:02.880
<v Speaker 2>need a divide that's going to record those vibrations. For example,

0:09:02.960 --> 0:09:06.160
<v Speaker 2>you could transmit them onto vinyl, or you could digitize

0:09:06.160 --> 0:09:08.800
<v Speaker 2>them and store them on a computer. So a microphone

0:09:08.880 --> 0:09:10.960
<v Speaker 2>or a speaker essentially inverses of each other.

0:09:11.480 --> 0:09:14.440
<v Speaker 1>Okay, I'm still amazed that we figured out how to

0:09:14.440 --> 0:09:16.040
<v Speaker 1>do that. That's that is pretty cool.

0:09:16.120 --> 0:09:17.520
<v Speaker 2>It is pretty cool. And there's a lot of really

0:09:17.559 --> 0:09:20.880
<v Speaker 2>fascinating physics of sound, like how the speed of sound

0:09:21.320 --> 0:09:24.840
<v Speaker 2>depends on how stiff and dense the material is. So,

0:09:25.040 --> 0:09:27.679
<v Speaker 2>in air, sound travels about three hundred and forty meters

0:09:27.720 --> 0:09:31.520
<v Speaker 2>per second. In water, it's faster because water is denser

0:09:32.000 --> 0:09:34.760
<v Speaker 2>and so molecules are closer together they could push up

0:09:34.760 --> 0:09:38.040
<v Speaker 2>against each other easier. It's like fifteen hundred meters per second,

0:09:38.440 --> 0:09:41.160
<v Speaker 2>and stel it's several kilometers per second.

0:09:41.240 --> 0:09:41.520
<v Speaker 3>Wow.

0:09:41.559 --> 0:09:44.640
<v Speaker 2>So sound travels faster through the ground. So you know

0:09:44.679 --> 0:09:47.360
<v Speaker 2>those scenes where somebody's like putting their ear on the

0:09:47.360 --> 0:09:49.680
<v Speaker 2>ground to hear if the buffalo are coming, Like, there's

0:09:49.720 --> 0:09:50.600
<v Speaker 2>real physics there.

0:09:50.880 --> 0:09:51.840
<v Speaker 1>WHOA, that's real.

0:09:53.040 --> 0:09:57.120
<v Speaker 2>That's real exactly. And so the things to understand about

0:09:57.120 --> 0:09:59.920
<v Speaker 2>sound is that there's frequencies, right, so hind ones and

0:10:00.080 --> 0:10:03.439
<v Speaker 2>low ones. There's also amplitude, right, is it loud or

0:10:03.559 --> 0:10:06.200
<v Speaker 2>is it quiet? Those are the crucial things to understand.

0:10:06.200 --> 0:10:09.600
<v Speaker 2>But musical sounds or voices are much more complicated than

0:10:09.679 --> 0:10:12.800
<v Speaker 2>just like a single frequency. Right, If you ask somebody

0:10:12.840 --> 0:10:15.920
<v Speaker 2>with a violin to play a certain note that's going

0:10:15.960 --> 0:10:18.960
<v Speaker 2>to correspond to a certain frequency, but a violin doesn't

0:10:19.000 --> 0:10:22.760
<v Speaker 2>produce sound that exactly one frequency. So if you want

0:10:22.760 --> 0:10:25.280
<v Speaker 2>to play four hundred and forty hertz, then a violin's

0:10:25.360 --> 0:10:27.839
<v Speaker 2>also going to play a harmonic which means eight hundred

0:10:27.880 --> 0:10:31.320
<v Speaker 2>and eighty hertz, which sounds like another version of the

0:10:31.360 --> 0:10:33.959
<v Speaker 2>same note. You know how somebody plays like the low

0:10:34.000 --> 0:10:36.520
<v Speaker 2>e and a high e and a guitar. They sound

0:10:36.640 --> 0:10:40.679
<v Speaker 2>nice together because they beat together, they're not shifted, but

0:10:40.679 --> 0:10:43.760
<v Speaker 2>there are different notes, and so a violin doesn't just

0:10:43.840 --> 0:10:45.880
<v Speaker 2>play the four to forty. It also plays the eight

0:10:45.960 --> 0:10:49.160
<v Speaker 2>eighty and the thirteen twenty and the seventeen sixty. And

0:10:49.240 --> 0:10:51.840
<v Speaker 2>if you play a guitar or a trombone with the

0:10:51.880 --> 0:10:55.080
<v Speaker 2>same note, it has a different relationship with those harmonics.

0:10:55.480 --> 0:10:57.760
<v Speaker 2>So you can hear a violin is different from a

0:10:57.760 --> 0:11:04.599
<v Speaker 2>trombone even when they're playing the same note. Why is

0:11:04.640 --> 0:11:06.680
<v Speaker 2>that Because they have a different set of frequencies that

0:11:06.720 --> 0:11:09.640
<v Speaker 2>they're playing. It's not a single pure frequency. And this

0:11:09.679 --> 0:11:12.560
<v Speaker 2>is what we call timber. This is why, like my

0:11:12.720 --> 0:11:15.120
<v Speaker 2>voice sounds different from somebody else's voice, even if it

0:11:15.160 --> 0:11:17.360
<v Speaker 2>was at the same frequency, like if you took Kelly's

0:11:17.440 --> 0:11:20.240
<v Speaker 2>voice and you shifted it down to Daniel range, oh no,

0:11:20.280 --> 0:11:22.160
<v Speaker 2>it would not sound like Daniels.

0:11:23.480 --> 0:11:26.560
<v Speaker 1>And that this was all new to me like forty

0:11:26.600 --> 0:11:28.440
<v Speaker 1>minutes ago when we talked to Matt, and so like

0:11:28.480 --> 0:11:30.720
<v Speaker 1>you know, heads up, Matt is going to mention this

0:11:30.760 --> 0:11:32.640
<v Speaker 1>a little bit and how he can see that in

0:11:32.679 --> 0:11:35.480
<v Speaker 1>the software that he uses, and that totally blew my mind.

0:11:35.520 --> 0:11:38.240
<v Speaker 1>I didn't realize that you were getting like harmonics of

0:11:38.280 --> 0:11:40.520
<v Speaker 1>people's voices and anyway, very cool, but.

0:11:40.559 --> 0:11:43.280
<v Speaker 2>There's also lots of fascinating biology here, like how the

0:11:43.360 --> 0:11:46.440
<v Speaker 2>human voice is actually generated. Right, you don't have like

0:11:46.440 --> 0:11:49.240
<v Speaker 2>a little violin in your throat, but you do have

0:11:49.360 --> 0:11:52.760
<v Speaker 2>things that vibrate. You have the vocal chords, or sometimes

0:11:52.840 --> 0:11:56.120
<v Speaker 2>called the vocal folds. They're not like guitar strings. They're

0:11:56.160 --> 0:11:59.680
<v Speaker 2>more like vibrating valves in an airflow and faster vibration.

0:12:00.040 --> 0:12:03.000
<v Speaker 2>It's a higher pitch, and so like adult male voices

0:12:03.040 --> 0:12:06.400
<v Speaker 2>are lower because they have longer, thicker vocal folds that

0:12:06.520 --> 0:12:09.880
<v Speaker 2>vibrate more slowly. And so this is why also two

0:12:09.920 --> 0:12:12.120
<v Speaker 2>people sing the same note but sound.

0:12:11.840 --> 0:12:15.960
<v Speaker 1>Different, and birds are somehow able to make at least

0:12:16.040 --> 0:12:19.240
<v Speaker 1>two noises at the same time, some birds, not all birds,

0:12:19.280 --> 0:12:21.640
<v Speaker 1>which is just absolutely like if you you should really

0:12:21.679 --> 0:12:23.360
<v Speaker 1>pay attention to the next time you hear a bird

0:12:23.440 --> 0:12:25.360
<v Speaker 1>song and try to figure out, like, wait a minute,

0:12:25.360 --> 0:12:28.160
<v Speaker 1>am I actually hearing that bird make two different notes

0:12:28.600 --> 0:12:32.520
<v Speaker 1>at the same time? And it's able to do that,

0:12:32.559 --> 0:12:34.640
<v Speaker 1>and anyway, sound is incredible and.

0:12:34.600 --> 0:12:36.360
<v Speaker 2>Some humans can do that too, right, this is this

0:12:36.440 --> 0:12:39.040
<v Speaker 2>like Himalayan dual note singing thing.

0:12:42.880 --> 0:12:45.360
<v Speaker 1>Oh, I didn't know that. I know that birds have

0:12:45.400 --> 0:12:48.400
<v Speaker 1>a very different like voice box than we do, so

0:12:48.440 --> 0:12:50.760
<v Speaker 1>that they can do that, and so now I need

0:12:50.800 --> 0:12:52.920
<v Speaker 1>to know what happens with the Himalayan people.

0:12:53.280 --> 0:12:55.680
<v Speaker 2>It's very hard to do, but you can train yourself

0:12:55.720 --> 0:12:58.920
<v Speaker 2>apparently to do that. And sound is much more complicated

0:12:58.960 --> 0:13:02.079
<v Speaker 2>than just like generate sound. Here a sound, you can

0:13:02.120 --> 0:13:05.280
<v Speaker 2>have interference effects because sound is a wave, and so

0:13:05.400 --> 0:13:08.560
<v Speaker 2>places where the pressure waves are in the same direction

0:13:08.679 --> 0:13:12.439
<v Speaker 2>will be louder, that's constructive interference, and places where they're

0:13:12.480 --> 0:13:15.240
<v Speaker 2>pushing in opposite directions it'll be weaker. And this is

0:13:15.280 --> 0:13:20.000
<v Speaker 2>how like noise cancelation works. Right, you play the opposite sound,

0:13:20.440 --> 0:13:22.400
<v Speaker 2>the one that's beating down when the other sound is

0:13:22.440 --> 0:13:25.400
<v Speaker 2>beating up and it cancels it out. So for example,

0:13:25.480 --> 0:13:29.680
<v Speaker 2>my daughter has air pods in and I say, hey, Hazel,

0:13:29.679 --> 0:13:32.080
<v Speaker 2>will you take out the trash? And her AirPod plays

0:13:32.120 --> 0:13:34.199
<v Speaker 2>the opposite of Hey Hazel, will you take out the trash?

0:13:34.240 --> 0:13:37.040
<v Speaker 2>And she hears nothing right, which is why she doesn't

0:13:37.040 --> 0:13:40.360
<v Speaker 2>take up the trash. Perfectly hypothetical example.

0:13:41.200 --> 0:13:44.720
<v Speaker 1>However, I would like to note that Matt mentions later

0:13:44.760 --> 0:13:48.040
<v Speaker 1>in the episode that noise canceling headphones don't necessarily cancel

0:13:48.080 --> 0:13:50.280
<v Speaker 1>out people talking to you. Yeah, I don't want to

0:13:50.280 --> 0:13:53.960
<v Speaker 1>get Hazel in trouble, but I'm wondering if maybe Hazel's

0:13:54.040 --> 0:13:54.760
<v Speaker 1>just not listening.

0:13:54.880 --> 0:13:56.960
<v Speaker 2>Yeah, there may be a few effects going on here.

0:13:57.679 --> 0:14:01.000
<v Speaker 2>But this goes into like the design of concert right,

0:14:01.080 --> 0:14:03.600
<v Speaker 2>because sound reflects and it can cancel out. So you

0:14:03.640 --> 0:14:06.679
<v Speaker 2>can have like dead spots in a room where you

0:14:06.720 --> 0:14:09.320
<v Speaker 2>can't hear someone from across the room, or places where

0:14:09.360 --> 0:14:12.520
<v Speaker 2>the sound all balances really really well. And this makes

0:14:12.559 --> 0:14:14.800
<v Speaker 2>a big difference between like a bad concert hall and

0:14:14.840 --> 0:14:17.439
<v Speaker 2>a good concert hall where like every seat sounds good

0:14:17.600 --> 0:14:21.320
<v Speaker 2>or more seats sound really good. You could also have

0:14:21.400 --> 0:14:23.880
<v Speaker 2>beating sounds, like if you have two sounds, they're very

0:14:23.920 --> 0:14:26.560
<v Speaker 2>similar to each other very close, but not quite on

0:14:26.640 --> 0:14:28.880
<v Speaker 2>top of each other. Like if you play your guitar

0:14:28.960 --> 0:14:30.840
<v Speaker 2>tuner at e and then you play your guitar but

0:14:30.880 --> 0:14:40.280
<v Speaker 2>it's not quite in tune. You'll hear this, Like, those

0:14:40.280 --> 0:14:43.600
<v Speaker 2>are the waves drifting out of alignment, right, you'll hear

0:14:43.640 --> 0:14:47.200
<v Speaker 2>them because they're not beating in sync harmonic like for

0:14:47.200 --> 0:14:50.240
<v Speaker 2>forty and a eighty. They have the down moments at

0:14:50.240 --> 0:14:52.280
<v Speaker 2>the same time, and so they sound nice.

0:14:52.920 --> 0:14:54.640
<v Speaker 1>I miss that. Could you make that sound again?

0:14:57.360 --> 0:14:59.080
<v Speaker 2>Mac can make that your ring toney for you if

0:14:59.120 --> 0:14:59.440
<v Speaker 2>you like.

0:15:00.960 --> 0:15:01.680
<v Speaker 1>That would be great.

0:15:02.480 --> 0:15:04.120
<v Speaker 2>And then there's a lot of work that needs to

0:15:04.160 --> 0:15:06.640
<v Speaker 2>be done when you're like doing noise removal. And Matt's

0:15:06.640 --> 0:15:08.880
<v Speaker 2>gonna tell us all about that when he gets here.

0:15:09.080 --> 0:15:12.920
<v Speaker 1>All right, so let us delay no longer, actually quick

0:15:12.960 --> 0:15:16.240
<v Speaker 1>delay we're gonna do. We're gonna do a commercial break,

0:15:16.320 --> 0:15:18.520
<v Speaker 1>and then when we come back, we will bring Matt

0:15:18.560 --> 0:15:19.000
<v Speaker 1>on the show.

0:15:19.080 --> 0:15:22.160
<v Speaker 2>The Great Matt Kesselman.

0:15:41.760 --> 0:15:44.480
<v Speaker 1>Matt Kesselman is an audio engineer and producer with over

0:15:44.560 --> 0:15:48.560
<v Speaker 1>two decades of multidisciplinary experience in the fields of film production,

0:15:49.080 --> 0:15:52.840
<v Speaker 1>audio post production, and music. Matt has worked with artists

0:15:52.880 --> 0:15:57.520
<v Speaker 1>and broadcasters like Netflix, the BBC, Lionsgate, Walk Off the Earth,

0:15:57.600 --> 0:16:01.120
<v Speaker 1>Ill Scarlet and many others. He developed software tools for

0:16:01.240 --> 0:16:05.360
<v Speaker 1>audio engineers through his company Oslot Audio. He brings over

0:16:05.440 --> 0:16:08.160
<v Speaker 1>a decade of teaching and communication experience to the table

0:16:08.200 --> 0:16:11.080
<v Speaker 1>as a college professor at Metalworks Institute, where he is

0:16:11.160 --> 0:16:14.280
<v Speaker 1>currently and before that at Durham College. He's also the

0:16:14.320 --> 0:16:17.880
<v Speaker 1>amazing audio engineer and sound designer behind dk EU. He's

0:16:17.920 --> 0:16:21.040
<v Speaker 1>the guy who makes us sound not horrible. Thanks for

0:16:21.120 --> 0:16:21.920
<v Speaker 1>joining us, Matt.

0:16:22.080 --> 0:16:24.440
<v Speaker 2>He's also the guy who has a shocking amount of

0:16:24.440 --> 0:16:26.000
<v Speaker 2>blackmail material on both of us.

0:16:26.040 --> 0:16:30.080
<v Speaker 3>That's right, that's all true. I'm a little nervous. I

0:16:30.120 --> 0:16:33.120
<v Speaker 3>feel like am I the first guest that doesn't have

0:16:33.160 --> 0:16:36.840
<v Speaker 3>a PhD? Oh not even in like chemistry.

0:16:37.480 --> 0:16:39.600
<v Speaker 2>Wow, I didn't realize we were gate keeping the guests.

0:16:39.640 --> 0:16:41.040
<v Speaker 1>That's right, Oh my gosh.

0:16:41.200 --> 0:16:43.920
<v Speaker 3>I mean you folks try to find the secrets to

0:16:44.000 --> 0:16:47.080
<v Speaker 3>life and how to better all of humankind. And I

0:16:47.200 --> 0:16:50.720
<v Speaker 3>once produced a song called Candy g String. So we're

0:16:50.720 --> 0:16:53.840
<v Speaker 3>a little different. So I'll try to cold man of

0:16:53.840 --> 0:16:55.120
<v Speaker 3>the bargain and make this entertaining.

0:16:55.240 --> 0:16:58.240
<v Speaker 2>Well, that's what makes us a harmonious collaboration. Everybody brings

0:16:58.240 --> 0:16:58.960
<v Speaker 2>their strengths.

0:16:59.200 --> 0:16:59.680
<v Speaker 3>That's right.

0:17:00.640 --> 0:17:03.200
<v Speaker 2>Well, maybe we should start actually by telling the story

0:17:03.280 --> 0:17:05.639
<v Speaker 2>of how we got connected to you, Matt, how you

0:17:05.800 --> 0:17:08.280
<v Speaker 2>ended up being the audio engineer for the pod.

0:17:09.280 --> 0:17:12.960
<v Speaker 3>Yeah, how indeed. Well, I was a fan for a

0:17:13.000 --> 0:17:17.400
<v Speaker 3>long time, and I think, like all the other Extraordinaries,

0:17:17.600 --> 0:17:19.880
<v Speaker 3>I was listening before we were even called the Extraordinaries.

0:17:19.920 --> 0:17:22.000
<v Speaker 3>But I always have questions and I always want to

0:17:22.000 --> 0:17:26.640
<v Speaker 3>know how things work. And a lot of popular science.

0:17:27.160 --> 0:17:30.199
<v Speaker 3>I can only consume popular science because I am not

0:17:30.359 --> 0:17:32.159
<v Speaker 3>smart enough, but a lot of it, you know, I

0:17:32.160 --> 0:17:34.560
<v Speaker 3>don't believe. That leaves me wanting and I feel like

0:17:34.560 --> 0:17:37.560
<v Speaker 3>there's something missing, and then there's conflicting things in different

0:17:37.600 --> 0:17:40.320
<v Speaker 3>popular science publications. And then I found you guys, and

0:17:41.119 --> 0:17:43.520
<v Speaker 3>you break it down and that was really enjoyable. And

0:17:43.520 --> 0:17:45.760
<v Speaker 3>then I thought, why don't I just tell them that

0:17:46.000 --> 0:17:48.480
<v Speaker 3>I can make them sound even better?

0:17:50.440 --> 0:17:52.560
<v Speaker 2>Yeah, we got an email out of the blue from

0:17:52.600 --> 0:17:56.280
<v Speaker 2>you saying, hey, I have notes on your audio engineering,

0:17:56.880 --> 0:17:59.280
<v Speaker 2>and we dug into it with you, and then eventually

0:17:59.280 --> 0:18:01.520
<v Speaker 2>we were like, hey, why do you come make the

0:18:01.560 --> 0:18:02.200
<v Speaker 2>show better?

0:18:02.560 --> 0:18:06.600
<v Speaker 1>And we haven't looked back yet, and so one of

0:18:06.640 --> 0:18:08.600
<v Speaker 1>the reasons we're super excited to have you on the

0:18:08.600 --> 0:18:11.200
<v Speaker 1>show is because after Daniel and I do a listener

0:18:11.280 --> 0:18:15.600
<v Speaker 1>questions episode, we always send the raw audio file export

0:18:15.640 --> 0:18:17.840
<v Speaker 1>it from riverside out to the listeners so that they

0:18:17.880 --> 0:18:20.000
<v Speaker 1>have plenty of time to listen to our conversation and

0:18:20.040 --> 0:18:24.399
<v Speaker 1>send us an audio file responding. And almost every single

0:18:24.400 --> 0:18:29.680
<v Speaker 1>one of them is like, Wow, this sounds awful. Your

0:18:29.840 --> 0:18:35.119
<v Speaker 1>audio guy must do so much work, and so we

0:18:35.240 --> 0:18:36.720
<v Speaker 1>thought that it would be great to have you on

0:18:36.760 --> 0:18:39.200
<v Speaker 1>the show to like pull back the curtain and talk

0:18:39.240 --> 0:18:42.440
<v Speaker 1>about like when Kelly forgets to turn off her air conditioner,

0:18:42.480 --> 0:18:44.679
<v Speaker 1>how do you remove that sound, and like all of

0:18:44.720 --> 0:18:46.240
<v Speaker 1>the work that you have to do to make us

0:18:46.320 --> 0:18:49.240
<v Speaker 1>sound good. Where does that start?

0:18:49.440 --> 0:18:53.200
<v Speaker 3>Okay? Well, first of all, I want to point out

0:18:53.240 --> 0:18:56.399
<v Speaker 3>that there are different kinds of podcasts that have different needs. Right,

0:18:56.440 --> 0:19:01.119
<v Speaker 3>everybody's familiar with the comedians just talking with no editing,

0:19:01.160 --> 0:19:03.359
<v Speaker 3>with no nothing, and they leave all the mistakes in

0:19:03.400 --> 0:19:06.600
<v Speaker 3>and that's part of the charm. They're also narrative podcast

0:19:06.680 --> 0:19:08.440
<v Speaker 3>that tell a story kind of like a movie, and

0:19:08.560 --> 0:19:10.920
<v Speaker 3>our very own network has a few examples of that,

0:19:11.040 --> 0:19:14.880
<v Speaker 3>like Red Elvis and the buzz Aldrin Show. I don't

0:19:14.920 --> 0:19:17.320
<v Speaker 3>know if you've heard that one, but then there's also

0:19:17.359 --> 0:19:22.600
<v Speaker 3>the informative kind where you're basically enjoying a lecture. And

0:19:22.920 --> 0:19:26.000
<v Speaker 3>I feel like this is kind of like a lecture

0:19:26.440 --> 0:19:29.680
<v Speaker 3>mixed in with some lighthearted comedy. And the thing about

0:19:29.720 --> 0:19:32.400
<v Speaker 3>lectures is that they take a lot of time to prepare,

0:19:33.000 --> 0:19:36.080
<v Speaker 3>and you two do something that very few people can do,

0:19:36.119 --> 0:19:39.000
<v Speaker 3>and that is that you basically deliver two lectures a week,

0:19:39.440 --> 0:19:45.160
<v Speaker 3>often on topics that you're not specialized in. And it's

0:19:45.200 --> 0:19:48.160
<v Speaker 3>no surprise that sometimes you need to write on the fly.

0:19:48.920 --> 0:19:52.159
<v Speaker 3>I mean, otherwise, it's just not possible for any person

0:19:52.240 --> 0:19:55.439
<v Speaker 3>to deliver that much information without having to stop and

0:19:55.480 --> 0:19:58.280
<v Speaker 3>sort of rewind and think back on it. And I

0:19:58.280 --> 0:20:00.639
<v Speaker 3>think we all sort of treat it that way, where

0:20:01.000 --> 0:20:03.119
<v Speaker 3>you have your outline and you have your idea, and

0:20:03.160 --> 0:20:05.560
<v Speaker 3>sometimes you come up with a new idea while we record,

0:20:06.160 --> 0:20:09.520
<v Speaker 3>and then you try it out a few times until

0:20:09.520 --> 0:20:11.240
<v Speaker 3>you get it right, and when you send it off

0:20:11.240 --> 0:20:14.840
<v Speaker 3>and you hope that I put it together, I feel like.

0:20:14.800 --> 0:20:17.919
<v Speaker 1>You really you get me, Matt, you get me. And

0:20:17.960 --> 0:20:20.880
<v Speaker 1>I want to start describing the podcast now as lectures

0:20:20.920 --> 0:20:23.880
<v Speaker 1>with light comedy, because I feel like that's actually that's

0:20:23.920 --> 0:20:25.560
<v Speaker 1>a pretty good description.

0:20:25.320 --> 0:20:28.199
<v Speaker 3>That's really what it is, right, It's like, yeah, weird

0:20:28.400 --> 0:20:31.520
<v Speaker 3>things that are very hard to understand, sometimes broken down

0:20:31.560 --> 0:20:37.600
<v Speaker 3>into small, bite sized, understandable pieces with some pro lapse jokes,

0:20:40.440 --> 0:20:42.640
<v Speaker 3>which is my favorite type of entertainment.

0:20:43.720 --> 0:20:46.680
<v Speaker 2>We need constant goat updates. You know, people are curious.

0:20:47.359 --> 0:20:49.760
<v Speaker 1>Well, she's kept her outsides on the inside recently, so

0:20:49.840 --> 0:20:51.480
<v Speaker 1>that's nice, wonderful. Yeah.

0:20:51.520 --> 0:20:54.080
<v Speaker 2>Well, from my perspective, I lean heavily into this editing

0:20:54.520 --> 0:20:56.879
<v Speaker 2>because I know that we can go back and fix

0:20:56.920 --> 0:20:59.879
<v Speaker 2>something or say something a better way I speak to

0:21:00.040 --> 0:21:02.520
<v Speaker 2>friendly than I do in person, or if I'm like

0:21:02.600 --> 0:21:05.080
<v Speaker 2>lecturing to a class, I'm lecturing to a class. I

0:21:05.080 --> 0:21:07.240
<v Speaker 2>don't often back up and say things another way and

0:21:07.359 --> 0:21:09.560
<v Speaker 2>just go with it. So I think the listeners are

0:21:09.560 --> 0:21:13.159
<v Speaker 2>hearing a very unusual slice of our speaking, you know,

0:21:13.240 --> 0:21:17.000
<v Speaker 2>the intended for editing but unedited version, which is not

0:21:17.119 --> 0:21:19.040
<v Speaker 2>how like we have a conversation.

0:21:18.640 --> 0:21:22.320
<v Speaker 3>Normally, Right, And if the listeners want some inside baseball,

0:21:22.359 --> 0:21:25.240
<v Speaker 3>I always send you the finished episode or a draft

0:21:25.240 --> 0:21:28.000
<v Speaker 3>of the episode, and then you listen intently and go, oh,

0:21:28.040 --> 0:21:31.240
<v Speaker 3>actually that detail is technically right, but there's a better

0:21:31.280 --> 0:21:33.000
<v Speaker 3>way to say it, And then you send me a

0:21:33.040 --> 0:21:36.040
<v Speaker 3>pickup because you're perfectionists and you want everything to be

0:21:36.480 --> 0:21:39.560
<v Speaker 3>correct and true, and that's very admirable as well, and you're.

0:21:39.400 --> 0:21:41.600
<v Speaker 1>Always a really good sport when we're like, oh, I

0:21:41.880 --> 0:21:44.520
<v Speaker 1>actually thank you so much for editing all of that,

0:21:44.600 --> 0:21:45.840
<v Speaker 1>but I'd like to do it again.

0:21:46.560 --> 0:21:48.840
<v Speaker 2>So let's make this concrete. Can you show us some

0:21:48.880 --> 0:21:51.560
<v Speaker 2>examples of what you get and then what you send

0:21:51.600 --> 0:21:55.000
<v Speaker 2>us back. Show us the raw, unedited Daniel and Kelly

0:21:55.040 --> 0:21:58.320
<v Speaker 2>craziness and then the smooth Matt output.

0:21:58.640 --> 0:22:02.320
<v Speaker 3>Sure, I have two separate kinds of examples. Here's a

0:22:02.400 --> 0:22:06.400
<v Speaker 3>couple where the noise reduction was already done, and we're

0:22:06.440 --> 0:22:09.760
<v Speaker 3>going to focus on your writing on the fly, if

0:22:09.840 --> 0:22:13.000
<v Speaker 3>you will, so take a close listen to. We'll start

0:22:13.040 --> 0:22:18.760
<v Speaker 3>with with Kelly trying to communicate this idea, backing up

0:22:18.760 --> 0:22:21.000
<v Speaker 3>a couple of times, a couple of uhs and ums,

0:22:21.200 --> 0:22:24.560
<v Speaker 3>and then after we'll listen to the cut up version.

0:22:24.640 --> 0:22:26.800
<v Speaker 1>And I have no idea what file met pick. So

0:22:27.640 --> 0:22:31.840
<v Speaker 1>I'm on the edge of my seat, and we don't

0:22:31.840 --> 0:22:34.320
<v Speaker 1>really have a good framework yet for predicting. Like if

0:22:34.359 --> 0:22:38.439
<v Speaker 1>you took you know X, and say you decided you

0:22:38.440 --> 0:22:41.800
<v Speaker 1>were going to take what's a say you were going

0:22:41.840 --> 0:22:43.560
<v Speaker 1>to take lions and domesticate them.

0:22:44.640 --> 0:22:50.200
<v Speaker 3>Okay, so there's that. Here's here's the edited version.

0:22:51.840 --> 0:22:53.560
<v Speaker 1>Yeah right, And we don't really have a good framework

0:22:53.640 --> 0:22:56.200
<v Speaker 1>yet for predicting, like say you were going to take

0:22:56.280 --> 0:22:59.000
<v Speaker 1>lions and domesticate them. I'm being ridiculous, right, But like

0:22:59.040 --> 0:22:59.879
<v Speaker 1>say you decided.

0:22:59.600 --> 0:23:01.760
<v Speaker 2>To Really that sounds awesome. I would love to have

0:23:01.800 --> 0:23:06.600
<v Speaker 2>a house lion. Hey, whatever happened to the house lion

0:23:06.680 --> 0:23:07.200
<v Speaker 2>I requested?

0:23:07.240 --> 0:23:09.199
<v Speaker 1>Anyway, I need a few more years.

0:23:09.960 --> 0:23:12.800
<v Speaker 3>Yeah. So, I mean you can see that you got

0:23:12.840 --> 0:23:15.920
<v Speaker 3>this idea of comparing it to lines and it came

0:23:15.960 --> 0:23:18.639
<v Speaker 3>on the fly, so you just needed to work through it,

0:23:19.359 --> 0:23:22.080
<v Speaker 3>which is great because I think everybody would rather hear

0:23:22.680 --> 0:23:25.000
<v Speaker 3>what you intend for them to hear and not for

0:23:25.080 --> 0:23:27.359
<v Speaker 3>you to sort of back away and just say something

0:23:27.400 --> 0:23:28.359
<v Speaker 3>that's easier to say.

0:23:28.720 --> 0:23:31.200
<v Speaker 2>And I think that these on the fly moments are

0:23:31.240 --> 0:23:35.000
<v Speaker 2>what make our conversations really fun because it is a conversation.

0:23:35.119 --> 0:23:37.119
<v Speaker 2>You're thinking on the fly, you're coming up with new ideas,

0:23:37.160 --> 0:23:40.400
<v Speaker 2>you're responding to questions. Also, we could never pull off

0:23:40.400 --> 0:23:42.840
<v Speaker 2>a scripted podcast twice a week. That's just so much

0:23:42.880 --> 0:23:46.800
<v Speaker 2>more work. And on top of that, I'm a terrible actor,

0:23:47.240 --> 0:23:49.320
<v Speaker 2>So if I had to read a scriptive podcast, it

0:23:49.359 --> 0:23:52.159
<v Speaker 2>would sound really stilted. The only way for this to

0:23:52.200 --> 0:23:53.960
<v Speaker 2>work is for it to be a natural conversation.

0:23:54.320 --> 0:23:57.119
<v Speaker 1>I massively appreciate that I can like stumble and start

0:23:57.160 --> 0:24:00.280
<v Speaker 1>over many, many, many times and know Mad is going

0:24:00.320 --> 0:24:01.879
<v Speaker 1>to listen to all of that and pick out the

0:24:01.880 --> 0:24:04.840
<v Speaker 1>bits that make sense. And you're amazing.

0:24:05.119 --> 0:24:07.720
<v Speaker 3>Thank you. Sometimes you do put a lot of faith

0:24:07.760 --> 0:24:11.800
<v Speaker 3>in me. And just to be clear, what I mean

0:24:11.800 --> 0:24:13.800
<v Speaker 3>by that is that you'll have a great first half

0:24:13.840 --> 0:24:16.480
<v Speaker 3>of an idea, and then you might write on the

0:24:16.480 --> 0:24:18.959
<v Speaker 3>fly and work out the second half of the idea,

0:24:19.280 --> 0:24:22.640
<v Speaker 3>and then when I put them together, sometimes they could

0:24:22.720 --> 0:24:25.959
<v Speaker 3>sound a little bit disjointed, kind of like a milder

0:24:26.040 --> 0:24:29.560
<v Speaker 3>version of this real segment that really aired on TV.

0:24:30.440 --> 0:24:31.879
<v Speaker 2>An idea, but what we do, we would not be

0:24:31.920 --> 0:24:34.600
<v Speaker 2>good at what we do, would we We would be sloppy?

0:24:35.560 --> 0:24:36.600
<v Speaker 2>You calling it sloppy.

0:24:37.400 --> 0:24:39.520
<v Speaker 3>I don't want your recordings to sound like that, even

0:24:39.560 --> 0:24:44.160
<v Speaker 3>though it's completely normal for speech cadence to change over time,

0:24:44.200 --> 0:24:46.840
<v Speaker 3>but when you put it all together and remove the

0:24:46.880 --> 0:24:49.439
<v Speaker 3>gaps in the middle, it becomes more noticeable. So sometimes

0:24:49.480 --> 0:24:52.199
<v Speaker 3>I have to do things like tune your voice on

0:24:52.440 --> 0:24:56.959
<v Speaker 3>certain words, or look around and find other times when

0:24:56.960 --> 0:24:59.720
<v Speaker 3>you said that word in the episode or in other episodes. Wow,

0:25:00.160 --> 0:25:03.480
<v Speaker 3>use it to help build one cohesive sentence. And just

0:25:03.480 --> 0:25:06.760
<v Speaker 3>to be clear, I never change what you've said. I

0:25:06.840 --> 0:25:09.200
<v Speaker 3>might just change how it sounds, but the ideas are

0:25:09.359 --> 0:25:10.080
<v Speaker 3>purely yours.

0:25:10.440 --> 0:25:12.600
<v Speaker 1>How often does it do we do that to you?

0:25:12.920 --> 0:25:14.760
<v Speaker 1>I'm feeling really bad right now.

0:25:15.160 --> 0:25:17.959
<v Speaker 3>It's not something you do to me, but it probably

0:25:18.000 --> 0:25:19.240
<v Speaker 3>happens more often than you think.

0:25:19.359 --> 0:25:21.600
<v Speaker 2>Oh no, no.

0:25:21.840 --> 0:25:25.160
<v Speaker 3>Which is completely normal. You're focused on educating the masses.

0:25:25.400 --> 0:25:28.280
<v Speaker 3>I'm focused on making it sound smooth without anybody noticing

0:25:28.400 --> 0:25:29.360
<v Speaker 3>anything was done at all.

0:25:29.760 --> 0:25:32.199
<v Speaker 2>Well, thank you for making this sound so good and

0:25:32.240 --> 0:25:34.840
<v Speaker 2>making it sound like you haven't done anything. That's the

0:25:34.840 --> 0:25:35.720
<v Speaker 2>most flattering part.

0:25:35.760 --> 0:25:37.760
<v Speaker 3>You want to hear one of you, Daniel, Yes and no,

0:25:39.320 --> 0:25:41.880
<v Speaker 3>let's do it. Okay, here we go before and after,

0:25:41.920 --> 0:25:42.800
<v Speaker 3>For Daniel.

0:25:43.600 --> 0:25:46.040
<v Speaker 2>Had that one piece of data argue that the Earth

0:25:46.080 --> 0:25:49.240
<v Speaker 2>could be a flat disc that was perfectly that was

0:25:49.240 --> 0:25:55.280
<v Speaker 2>perfectly aligned to make a circular shadow. Yes, oh boy, that's.

0:25:55.200 --> 0:26:00.760
<v Speaker 3>Real for some reason, like were you at because you

0:26:00.800 --> 0:26:02.960
<v Speaker 3>hear that in the Listen again.

0:26:03.200 --> 0:26:05.640
<v Speaker 2>And that one piece of data argue that the Earth

0:26:05.680 --> 0:26:08.760
<v Speaker 2>could be a flat to hear that where Yeah, I

0:26:08.800 --> 0:26:10.719
<v Speaker 2>was not buying a can of olives or something at

0:26:10.760 --> 0:26:13.240
<v Speaker 2>the same time as recording the podcast. I've no idea

0:26:13.280 --> 0:26:17.399
<v Speaker 2>where that comes from. So here's the after Yeah, that

0:26:17.440 --> 0:26:19.840
<v Speaker 2>one piece of data argue that the Earth could be

0:26:19.880 --> 0:26:22.399
<v Speaker 2>a flat disc that was perfectly aligned to make a

0:26:22.400 --> 0:26:27.439
<v Speaker 2>circular shadow. Yes, I sound great afterwards, thank you.

0:26:28.520 --> 0:26:31.840
<v Speaker 3>Yeah, and again, like you said, you don't talk like

0:26:31.880 --> 0:26:34.760
<v Speaker 3>that in person. It's just that you're trying to communicate

0:26:34.920 --> 0:26:37.399
<v Speaker 3>things that are new to you in some way and

0:26:37.440 --> 0:26:39.440
<v Speaker 3>in a way that fits for broadcast.

0:26:39.760 --> 0:26:42.840
<v Speaker 1>Okay, So when I listened to the audio files after

0:26:42.880 --> 0:26:44.879
<v Speaker 1>I send them to you, I'm like, oh, shoot, I

0:26:44.960 --> 0:26:48.000
<v Speaker 1>left my air conditioner on. Or Daniel was trying to

0:26:48.040 --> 0:26:51.480
<v Speaker 1>explain general relativity at the grocery store, like how do

0:26:51.680 --> 0:26:55.800
<v Speaker 1>you remove like the beeps and the background noises? Like

0:26:55.880 --> 0:26:56.720
<v Speaker 1>how hard is that?

0:26:57.480 --> 0:27:00.680
<v Speaker 3>So there are many different types of noises in many

0:27:00.720 --> 0:27:05.120
<v Speaker 3>different solutions. Some are automated, some are not. So if

0:27:05.160 --> 0:27:10.479
<v Speaker 3>it's here, I have some examples. Here is some unprocessed Daniel,

0:27:10.560 --> 0:27:12.280
<v Speaker 3>let me know if you can hear that background noise.

0:27:13.880 --> 0:27:18.120
<v Speaker 2>So Bell's paradox involves length contraction, the fact that when

0:27:18.400 --> 0:27:22.159
<v Speaker 2>things move fast, they look short. So relativity tells us

0:27:22.160 --> 0:27:26.960
<v Speaker 2>two things, moving clocks run slow and moving objects look short.

0:27:28.760 --> 0:27:31.240
<v Speaker 3>So that kind of noise is probably the easiest because

0:27:31.280 --> 0:27:31.960
<v Speaker 3>it's consistent.

0:27:32.359 --> 0:27:34.679
<v Speaker 2>It sounds like I'm crinkling paper or something.

0:27:35.040 --> 0:27:37.399
<v Speaker 3>That's a different thing. That's mouth clicks, and that's a

0:27:37.440 --> 0:27:39.640
<v Speaker 3>normal thing that happens when you put a microphone really

0:27:39.640 --> 0:27:42.879
<v Speaker 3>close to the human mouth. It's weird how it doesn't

0:27:42.880 --> 0:27:45.040
<v Speaker 3>happen in person, but it does on micro For that

0:27:45.160 --> 0:27:49.360
<v Speaker 3>kind of noise, it's because it's consistent. We have software

0:27:49.400 --> 0:27:51.840
<v Speaker 3>that we basically tell the software look at this, as

0:27:51.880 --> 0:27:54.200
<v Speaker 3>long as there's no talking, look at this piece of noise,

0:27:54.520 --> 0:27:57.000
<v Speaker 3>identify it as noise, and from now on remove any

0:27:57.080 --> 0:28:00.760
<v Speaker 3>time you encounter that. And they be come more and

0:28:00.800 --> 0:28:04.080
<v Speaker 3>more sophisticated over the years to the point where they're

0:28:04.440 --> 0:28:07.679
<v Speaker 3>quite transparent at it. It's still in many cases if

0:28:07.720 --> 0:28:10.320
<v Speaker 3>you go too hard, things start to sound watery and

0:28:10.359 --> 0:28:12.200
<v Speaker 3>sort of weird, so you have to be careful with that.

0:28:12.440 --> 0:28:14.560
<v Speaker 3>But dig into that a little bit more in detail,

0:28:14.640 --> 0:28:17.400
<v Speaker 3>because what you're saying is it's consistent, so you can

0:28:17.720 --> 0:28:20.280
<v Speaker 3>measure it when the person's not speaking, and then you'll

0:28:20.240 --> 0:28:23.440
<v Speaker 3>have a template to remove it. But still it's varying.

0:28:23.640 --> 0:28:27.280
<v Speaker 3>So is an average removal. It's a statistical removal because

0:28:27.320 --> 0:28:29.800
<v Speaker 3>it can't remove the exact noise moment to moment, because

0:28:29.840 --> 0:28:30.640
<v Speaker 3>there is variation.

0:28:30.800 --> 0:28:31.239
<v Speaker 2>We heard it.

0:28:31.920 --> 0:28:37.760
<v Speaker 3>There is variation, but it's fairly predictable. Like an air conditioner,

0:28:37.840 --> 0:28:40.560
<v Speaker 3>unless the air conditioner is changing temperatures, is going to

0:28:40.720 --> 0:28:45.280
<v Speaker 3>give or take make the same droney noise throughout the recording.

0:28:46.040 --> 0:28:48.479
<v Speaker 3>But yes, sometimes I do catch that throughout the recording.

0:28:48.520 --> 0:28:52.360
<v Speaker 3>Sometimes the air conditioner change gears and then it's no

0:28:52.440 --> 0:28:54.800
<v Speaker 3>longer being recognized and I have to compensate for that.

0:28:55.400 --> 0:29:00.200
<v Speaker 3>There are settings where it's supposed to follow the noise itself.

0:29:01.080 --> 0:29:03.120
<v Speaker 3>I don't find those to be very good yet, but

0:29:03.440 --> 0:29:05.440
<v Speaker 3>tuch changes all the time. Maybe tomorrow there'll be a

0:29:05.440 --> 0:29:09.120
<v Speaker 3>good one. And is this noise removable? Because it's at

0:29:09.160 --> 0:29:12.360
<v Speaker 3>a different frequency than the voice, so you can identify

0:29:12.400 --> 0:29:14.240
<v Speaker 3>it when I'm not speaking, and then you know exactly

0:29:14.280 --> 0:29:15.880
<v Speaker 3>how to remove it. And you can remove it without

0:29:15.920 --> 0:29:18.880
<v Speaker 3>removing the voice or does it overlap of the voice

0:29:19.240 --> 0:29:22.520
<v Speaker 3>and so you are affecting the actual voice as well?

0:29:22.720 --> 0:29:25.080
<v Speaker 3>Oh okay, that's a great question, you really want to

0:29:25.080 --> 0:29:27.560
<v Speaker 3>look under the hood. So the way it really works

0:29:27.680 --> 0:29:32.400
<v Speaker 3>is that the recording is divided into smaller bands of frequencies,

0:29:32.440 --> 0:29:35.080
<v Speaker 3>and each frequency is given what's called a gate or

0:29:35.080 --> 0:29:39.360
<v Speaker 3>an expander, where if a lot of volume hits that area,

0:29:39.880 --> 0:29:43.200
<v Speaker 3>it's allowed to pass through, so speech, for example. But

0:29:43.480 --> 0:29:47.840
<v Speaker 3>if there's a quieter moment, it enhances the quiet moment

0:29:47.880 --> 0:29:52.360
<v Speaker 3>and makes it even quieter. So if you really listen closely,

0:29:52.960 --> 0:29:55.680
<v Speaker 3>you may be able to notice that there are moments

0:29:55.680 --> 0:29:59.760
<v Speaker 3>of noise when you talk. But because you're talking and

0:29:59.760 --> 0:30:03.600
<v Speaker 3>then noise is only present as you're talking, it's not perceptible.

0:30:04.720 --> 0:30:06.520
<v Speaker 3>A lot of this is about magic tricks and not

0:30:06.640 --> 0:30:08.200
<v Speaker 3>actually about doing what you think we're doing.

0:30:08.480 --> 0:30:12.520
<v Speaker 2>By magic tricks, you mean distracting people and understanding how works.

0:30:12.600 --> 0:30:15.120
<v Speaker 3>Yeah, I mean I mean real real magic, not magic

0:30:15.200 --> 0:30:17.360
<v Speaker 3>magic were slide of hand and stuff like that.

0:30:19.920 --> 0:30:22.120
<v Speaker 2>I love what you call that real magic. That's awesome.

0:30:22.320 --> 0:30:25.320
<v Speaker 1>Okay, So then how would you remove something like the

0:30:25.480 --> 0:30:26.280
<v Speaker 1>grocery store.

0:30:26.840 --> 0:30:28.959
<v Speaker 2>I was not at a grocery store. I want that

0:30:29.040 --> 0:30:29.560
<v Speaker 2>on their front.

0:30:29.680 --> 0:30:30.320
<v Speaker 3>I think you were.

0:30:30.400 --> 0:30:30.880
<v Speaker 1>I saw it.

0:30:30.960 --> 0:30:35.640
<v Speaker 3>I saw it. So for a noise like that or

0:30:35.760 --> 0:30:42.680
<v Speaker 3>like goats, Kelly oh yeah, their goats. Yeah.

0:30:42.720 --> 0:30:43.840
<v Speaker 1>Oh, I'm so sorry.

0:30:43.840 --> 0:30:46.320
<v Speaker 3>Oh no, it's totally it's not your fault that they're goats.

0:30:46.640 --> 0:30:49.560
<v Speaker 3>They do goaty things, uh, for things like that. For

0:30:49.560 --> 0:30:53.960
<v Speaker 3>for noises that are not consistent or predictable, which, by

0:30:54.000 --> 0:30:56.440
<v Speaker 3>the way, let me back up one second. Noise canceling

0:30:56.440 --> 0:31:00.800
<v Speaker 3>headphones work on the same principle, So noise can headphones

0:31:01.160 --> 0:31:05.920
<v Speaker 3>can cancel predictable noise once they recognize that noise. However,

0:31:06.000 --> 0:31:08.600
<v Speaker 3>if somebody's talking to you or there's a dog barking,

0:31:08.800 --> 0:31:11.120
<v Speaker 3>you may notice that your noise canceling headphones don't work

0:31:11.160 --> 0:31:14.120
<v Speaker 3>for this very same reason. So if there is a

0:31:14.120 --> 0:31:17.080
<v Speaker 3>dog barking or a goat. Luckily, we live at a

0:31:17.160 --> 0:31:18.840
<v Speaker 3>time where I can just open it up in a

0:31:18.880 --> 0:31:21.640
<v Speaker 3>spectrol graph and I can actually see all the sound.

0:31:21.680 --> 0:31:25.080
<v Speaker 3>It looks like splotches, and I can see the goat

0:31:25.320 --> 0:31:27.520
<v Speaker 3>bleed to an extent, and I can dry it out

0:31:27.600 --> 0:31:28.280
<v Speaker 3>like an photoshop.

0:31:28.400 --> 0:31:33.719
<v Speaker 2>WHOA, So what you're looking at is the amplitude versus frequency.

0:31:34.160 --> 0:31:38.320
<v Speaker 3>Yeah, time is the x frequency is the y axis,

0:31:38.400 --> 0:31:43.480
<v Speaker 3>and then amplitude is color. Oh bright orange means it's

0:31:43.600 --> 0:31:46.640
<v Speaker 3>very loud. So I can actually see all the harmonics

0:31:46.680 --> 0:31:49.440
<v Speaker 3>of the goat or a door being slammed, or a

0:31:49.480 --> 0:31:52.120
<v Speaker 3>plane flying by, And sometimes it's hard to draw them

0:31:52.120 --> 0:31:54.160
<v Speaker 3>out because there's other stuff happening there. But in most

0:31:54.200 --> 0:31:57.400
<v Speaker 3>cases it just takes time. You have to stop, capture

0:31:57.400 --> 0:32:00.600
<v Speaker 3>that piece and follow the goat and remove it.

0:32:01.880 --> 0:32:04.480
<v Speaker 1>How does Daniel's voice show up different than mine? Go

0:32:04.680 --> 0:32:08.000
<v Speaker 1>So Daniel has a very deep voice. What does his

0:32:08.120 --> 0:32:10.360
<v Speaker 1>look like relative to mine?

0:32:10.560 --> 0:32:13.000
<v Speaker 3>Well, too bad. This is not a visual show. So

0:32:13.240 --> 0:32:17.360
<v Speaker 3>what you would see for any human is just a

0:32:17.440 --> 0:32:21.880
<v Speaker 3>series of lines on top of each other most sounds, really,

0:32:21.960 --> 0:32:25.120
<v Speaker 3>and the lowest line is the actual note that Daniel

0:32:25.360 --> 0:32:28.600
<v Speaker 3>or Kelly are saying. And then all the lines the

0:32:28.600 --> 0:32:31.480
<v Speaker 3>notes above that are harmonics that we don't perceive as

0:32:31.560 --> 0:32:35.000
<v Speaker 3>individual notes, but they're the sort of fingerprint that make

0:32:35.840 --> 0:32:37.880
<v Speaker 3>Daniel and Kelly sound different. Even if you sing the

0:32:37.880 --> 0:32:40.480
<v Speaker 3>same note, for example, is just a different combination of

0:32:40.480 --> 0:32:42.880
<v Speaker 3>those harmonics, huh.

0:32:42.560 --> 0:32:44.360
<v Speaker 2>Like, for the same reason that like, a guitar and

0:32:44.400 --> 0:32:47.360
<v Speaker 2>a violin sound different when they're playing the same note,

0:32:47.400 --> 0:32:50.760
<v Speaker 2>because there's a range of frequencies that gives you the

0:32:50.960 --> 0:32:53.200
<v Speaker 2>mental fingerprint to identify it exactly.

0:32:53.240 --> 0:32:55.640
<v Speaker 1>That's all it is, all right, Well, let's take a break,

0:32:55.680 --> 0:32:57.720
<v Speaker 1>and when we get back, we're going to chat with

0:32:57.800 --> 0:33:01.240
<v Speaker 1>Matt about how software is making his life easier or

0:33:01.360 --> 0:33:01.840
<v Speaker 1>maybe not.

0:33:02.120 --> 0:33:03.640
<v Speaker 2>And whether or not he's actually an AI.

0:33:03.960 --> 0:33:25.840
<v Speaker 1>And we're back and we've got Mat in the hot

0:33:25.880 --> 0:33:28.280
<v Speaker 1>seat today we are quizzing him about how he makes

0:33:28.360 --> 0:33:31.720
<v Speaker 1>us sound not horrible, because we sound pretty bad in

0:33:31.760 --> 0:33:34.680
<v Speaker 1>the raw audio that we send him at So AI

0:33:34.840 --> 0:33:38.360
<v Speaker 1>is like everywhere now, so could you like tell AI? Hey?

0:33:38.960 --> 0:33:42.520
<v Speaker 1>Here is an interview from two people for Starters. Cut

0:33:42.520 --> 0:33:45.000
<v Speaker 1>out every time they say and remove all the background

0:33:45.080 --> 0:33:48.400
<v Speaker 1>noise like are we there where that could happen? Or yeah,

0:33:48.400 --> 0:33:49.280
<v Speaker 1>where are we with AI?

0:33:49.360 --> 0:33:55.120
<v Speaker 3>Now? Uh, We're not there. We might be there. I

0:33:55.160 --> 0:33:57.360
<v Speaker 3>have no idea what's going to happen. Maybe next month

0:33:57.840 --> 0:34:00.280
<v Speaker 3>none of us will have jobs. I know this. People

0:34:00.280 --> 0:34:02.600
<v Speaker 3>have very strong beliefs about what's going to happen. I

0:34:02.640 --> 0:34:05.720
<v Speaker 3>think nobody knows, and we may hit a wall or

0:34:06.080 --> 0:34:08.839
<v Speaker 3>it may never stop. See. The thing is, there are

0:34:08.960 --> 0:34:11.319
<v Speaker 3>tools that can do things like that, but then they

0:34:11.360 --> 0:34:14.080
<v Speaker 3>require adult supervision where you check whether it did and

0:34:14.080 --> 0:34:16.600
<v Speaker 3>then you find that it deleted a very important question

0:34:16.880 --> 0:34:20.399
<v Speaker 3>or you know, things like that. So it doesn't help

0:34:20.520 --> 0:34:23.040
<v Speaker 3>me It's still faster for me to do it all

0:34:23.120 --> 0:34:25.840
<v Speaker 3>myself because I work pretty fast, and you know, I

0:34:25.880 --> 0:34:28.200
<v Speaker 3>don't have to go back and try to find what's missing.

0:34:29.239 --> 0:34:32.600
<v Speaker 3>The tools I use the most with AI are for

0:34:33.400 --> 0:34:36.800
<v Speaker 3>noise cleanup. That's also another option if I don't draw

0:34:37.160 --> 0:34:40.560
<v Speaker 3>it out with a pencil tool. Is There are AIS

0:34:40.600 --> 0:34:44.080
<v Speaker 3>that are pretty good at being able to tell what

0:34:44.200 --> 0:34:47.279
<v Speaker 3>noise is, and they sort of split into two categories.

0:34:47.320 --> 0:34:50.759
<v Speaker 3>One is it removes the noise on its own using

0:34:50.760 --> 0:34:54.640
<v Speaker 3>all kinds of cool phasing algorithms, and then there's one

0:34:54.719 --> 0:34:57.840
<v Speaker 3>that recognizes your voice and then resynthesizes your voice completely

0:34:57.840 --> 0:35:04.720
<v Speaker 3>from scratch. Wow wow, which is a little weird, scary scary. Yeah.

0:35:04.760 --> 0:35:06.680
<v Speaker 2>Well, I had a listener wright to me once that

0:35:06.719 --> 0:35:10.200
<v Speaker 2>they heard an AI generated podcast and that one of

0:35:10.239 --> 0:35:13.120
<v Speaker 2>the hosts on that podcast sounded a lot like me,

0:35:13.560 --> 0:35:15.320
<v Speaker 2>and I was like what, So then I listened to

0:35:15.400 --> 0:35:16.920
<v Speaker 2>it and I was like, that does kind of sound

0:35:16.960 --> 0:35:19.600
<v Speaker 2>like me, And then I realized there are hundreds of

0:35:19.719 --> 0:35:22.839
<v Speaker 2>hours of my voice out there for anybody to train

0:35:22.920 --> 0:35:25.359
<v Speaker 2>an AI to make an AI Daniel.

0:35:25.440 --> 0:35:29.439
<v Speaker 3>And you only need five seconds now. And by the way, PSA,

0:35:29.520 --> 0:35:32.160
<v Speaker 3>I don't know about your bank but my bank has

0:35:32.200 --> 0:35:35.440
<v Speaker 3>this voice print identification if you call them, Yeah, I

0:35:35.480 --> 0:35:38.840
<v Speaker 3>asked them to turn that off. Yeah, because it's incredibly easy,

0:35:38.920 --> 0:35:42.080
<v Speaker 3>especially for anyone that has their voice out there, for

0:35:42.120 --> 0:35:44.720
<v Speaker 3>anybody to fake your voice. So that is no longer

0:35:44.760 --> 0:35:46.760
<v Speaker 3>safe if you have that enabled at your bank.

0:35:46.840 --> 0:35:48.840
<v Speaker 2>And I'm not endorsing it, but I did find this

0:35:48.920 --> 0:35:51.080
<v Speaker 2>AI generated podcast to be pretty good. And I did

0:35:51.080 --> 0:35:54.360
<v Speaker 2>a little experiment where I took one of Katrina's latest

0:35:54.400 --> 0:35:56.560
<v Speaker 2>papers and I ran it through and I had produced

0:35:56.560 --> 0:35:59.120
<v Speaker 2>a fifteen minute podcast about it. Then I sent it

0:35:59.160 --> 0:36:01.600
<v Speaker 2>to her and didn't tell her it was AI generated.

0:36:01.680 --> 0:36:04.320
<v Speaker 2>I said, hey, I heard this podcast about your latest paper.

0:36:04.640 --> 0:36:06.920
<v Speaker 2>And she listened to it. She was really excited and

0:36:06.960 --> 0:36:08.080
<v Speaker 2>she shared it with her laugh.

0:36:10.400 --> 0:36:11.480
<v Speaker 3>Huh, that is weird.

0:36:11.640 --> 0:36:14.000
<v Speaker 2>It was surprisingly good. It's pretty compelling stuff.

0:36:14.080 --> 0:36:16.640
<v Speaker 1>Unfortunately, Well, so, I think you should be flattered that

0:36:16.800 --> 0:36:18.600
<v Speaker 1>out of all the voices they could have picked, they

0:36:18.640 --> 0:36:21.560
<v Speaker 1>decided to scrape your voice. Daniels a good voice.

0:36:21.800 --> 0:36:24.920
<v Speaker 2>I don't know that they did, you know, just generic

0:36:25.040 --> 0:36:29.040
<v Speaker 2>podcast dude with a deep voice, I guess. But no

0:36:29.120 --> 0:36:32.080
<v Speaker 2>AI is out there making goat jokes, Kelly, So you

0:36:32.080 --> 0:36:32.640
<v Speaker 2>are unique?

0:36:32.840 --> 0:36:35.680
<v Speaker 1>Yes, Oh that's great. I'm so glad, So Matt. Even

0:36:35.719 --> 0:36:40.520
<v Speaker 1>with the help of software, we send you about an hour,

0:36:40.719 --> 0:36:43.600
<v Speaker 1>an hour or twenty minutes of audio, it gets compressed

0:36:43.640 --> 0:36:47.840
<v Speaker 1>to about an hour of podcast. How many hours go

0:36:47.960 --> 0:36:50.320
<v Speaker 1>into editing? Uh?

0:36:50.560 --> 0:36:55.920
<v Speaker 3>Anywhere from five to ten wow, depending depending on physics

0:36:56.000 --> 0:36:59.759
<v Speaker 3>or bio subject matter. Yes, yes, Well, look if it's

0:37:00.239 --> 0:37:05.440
<v Speaker 3>if it's about particle accelerators, Daniel already has all that

0:37:05.520 --> 0:37:09.160
<v Speaker 3>stuff memorized and it's easy to do. Or if we're

0:37:09.160 --> 0:37:13.440
<v Speaker 3>talking about killie fish or something like that. But if

0:37:13.480 --> 0:37:15.920
<v Speaker 3>somebody just throws you a curveball question that you had

0:37:15.920 --> 0:37:17.839
<v Speaker 3>to look up, then it could take a bit longer.

0:37:17.840 --> 0:37:21.000
<v Speaker 3>If there's a guest, it usually takes longer because you

0:37:21.040 --> 0:37:23.239
<v Speaker 3>have a rapport and a banter that you're used to,

0:37:23.320 --> 0:37:25.239
<v Speaker 3>and then if somebody else comes in, they may not

0:37:25.640 --> 0:37:27.680
<v Speaker 3>have that same vibe and it takes a bit of time.

0:37:28.440 --> 0:37:30.239
<v Speaker 3>And I want to point out that five to ten

0:37:30.280 --> 0:37:33.000
<v Speaker 3>hours may sound like a lot, but it's not a

0:37:33.040 --> 0:37:35.640
<v Speaker 3>lot or unusual at all if you think about it.

0:37:35.640 --> 0:37:38.040
<v Speaker 3>It's usually roughly about an hour and a half of

0:37:38.280 --> 0:37:40.920
<v Speaker 3>raw audio that I get. I have to listen to

0:37:40.920 --> 0:37:43.040
<v Speaker 3>it at least one time, so that's already an hour

0:37:43.080 --> 0:37:46.160
<v Speaker 3>and a half. Then I need to do some noise reduction,

0:37:46.560 --> 0:37:50.719
<v Speaker 3>which can take thirty to fifty minutes, depending on the

0:37:50.760 --> 0:37:53.399
<v Speaker 3>type of noise there is. And then from there, every

0:37:53.560 --> 0:37:56.440
<v Speaker 3>edit that I make, I have to first hear it, stop,

0:37:56.640 --> 0:37:58.319
<v Speaker 3>listen to it one more time to make sure there

0:37:58.360 --> 0:38:01.120
<v Speaker 3>is a problem, do the edit, which takes between a

0:38:01.120 --> 0:38:03.560
<v Speaker 3>few seconds and a few minutes depending on the edit,

0:38:04.000 --> 0:38:07.160
<v Speaker 3>and then listen back. And these little things all add up,

0:38:07.200 --> 0:38:10.080
<v Speaker 3>and before you know it, you're easily crossing the five

0:38:10.120 --> 0:38:12.880
<v Speaker 3>hour mark. By the way, one other thing that I

0:38:12.920 --> 0:38:17.160
<v Speaker 3>forgot to mention is there's the editing sort of to

0:38:17.200 --> 0:38:19.439
<v Speaker 3>help everything feel smooth and makes sense. But then there's

0:38:19.480 --> 0:38:25.239
<v Speaker 3>also the issue of comedic timing over web conferencing. There's

0:38:25.239 --> 0:38:28.000
<v Speaker 3>a delay, whether you like it or not, and it's

0:38:28.040 --> 0:38:31.080
<v Speaker 3>not something as easy as shifting one voice back. Because

0:38:31.560 --> 0:38:35.479
<v Speaker 3>Kelly would say something funny, Daniel would accidentally talk over

0:38:35.560 --> 0:38:39.040
<v Speaker 3>Kelly because he's been waiting because there's a delay, then

0:38:39.120 --> 0:38:43.200
<v Speaker 3>realize that was funny, laugh at that. Kelly also hears

0:38:43.239 --> 0:38:47.120
<v Speaker 3>a delay, so she's already trying to damage control Daniel

0:38:47.200 --> 0:38:52.799
<v Speaker 3>not laughing, and it becomes this mess where I'm like,

0:38:52.840 --> 0:38:55.279
<v Speaker 3>there's a very funny joke here, and both laughed at

0:38:55.280 --> 0:38:58.799
<v Speaker 3>the joke, but because of that delay, it sounds a

0:38:58.800 --> 0:39:01.719
<v Speaker 3>little bit train wreckage. So that's another thing that.

0:39:01.680 --> 0:39:03.880
<v Speaker 2>I know, Wow, what a review?

0:39:04.120 --> 0:39:06.399
<v Speaker 3>Well, because I know that in person you both would

0:39:06.560 --> 0:39:08.440
<v Speaker 3>have this banter and laugh and it would be so

0:39:08.520 --> 0:39:11.640
<v Speaker 3>I'm sort of time traveling in a way to put

0:39:11.760 --> 0:39:14.920
<v Speaker 3>you where you wanted to be, at least where I

0:39:14.960 --> 0:39:17.440
<v Speaker 3>perceive you want it to be, so that that's also

0:39:17.480 --> 0:39:17.960
<v Speaker 3>part of it.

0:39:19.480 --> 0:39:21.520
<v Speaker 1>Thank you for making us funny.

0:39:21.920 --> 0:39:24.200
<v Speaker 3>No, you're the funny ones. I just undo the damage

0:39:24.239 --> 0:39:25.200
<v Speaker 3>the internet causes.

0:39:26.440 --> 0:39:29.560
<v Speaker 1>So why does my voice sound so much better in

0:39:29.640 --> 0:39:32.319
<v Speaker 1>my head? So like when I listened to the file

0:39:32.400 --> 0:39:35.040
<v Speaker 1>you send me, I think, wow, this sounds way better

0:39:35.360 --> 0:39:38.520
<v Speaker 1>than when I actually was like talking to Daniel in

0:39:38.560 --> 0:39:41.000
<v Speaker 1>this version, Daniel left at my joke right away, but

0:39:41.160 --> 0:39:44.759
<v Speaker 1>I still I still cringe because I don't love the

0:39:44.800 --> 0:39:47.360
<v Speaker 1>sound of my voice when I hear it back to me.

0:39:47.480 --> 0:39:49.799
<v Speaker 1>But when I'm saying it, it doesn't sound so bad.

0:39:49.840 --> 0:39:51.719
<v Speaker 1>So why is there that disconnect?

0:39:51.880 --> 0:39:53.400
<v Speaker 2>I think you have a great voice, Kelly. I love

0:39:53.440 --> 0:39:53.839
<v Speaker 2>hearing it.

0:39:53.920 --> 0:39:55.680
<v Speaker 3>Thanks Daniel the great great voice.

0:39:55.760 --> 0:39:56.279
<v Speaker 1>Thanks guys.

0:39:56.360 --> 0:39:59.000
<v Speaker 3>Otherwise people wouldn't be listening as interesting as you are.

0:39:59.040 --> 0:40:00.920
<v Speaker 3>If you had bad voice, as people would move on.

0:40:02.000 --> 0:40:06.359
<v Speaker 3>It's a matter of perspective. So we hear everything else

0:40:06.400 --> 0:40:10.800
<v Speaker 3>and everyone else through our ear canals. Right, there's air outside,

0:40:10.880 --> 0:40:12.959
<v Speaker 3>and that air vibrates and that goes into our ear canals.

0:40:13.000 --> 0:40:16.359
<v Speaker 3>Will hear it for ourselves. A big chunk of our

0:40:16.480 --> 0:40:20.440
<v Speaker 3>voice travels through conduction into our ear from the inside,

0:40:20.480 --> 0:40:21.960
<v Speaker 3>and that's what people call the inner ear. There's no

0:40:21.960 --> 0:40:23.680
<v Speaker 3>actual inner ear. I used to think there was when

0:40:23.719 --> 0:40:25.560
<v Speaker 3>I was a kid, like a third ear in your throat.

0:40:25.880 --> 0:40:30.200
<v Speaker 3>There isn't one. You're basically hearing sort of a muffled

0:40:30.239 --> 0:40:33.720
<v Speaker 3>version of yourself because it's traveling through soft tissue and bone,

0:40:34.239 --> 0:40:36.640
<v Speaker 3>and then you're also hearing a little bit of yourself

0:40:36.680 --> 0:40:38.759
<v Speaker 3>through the air going back into your ear, so it's

0:40:38.800 --> 0:40:40.920
<v Speaker 3>a different sort of like I don't know. When I

0:40:40.960 --> 0:40:43.960
<v Speaker 3>look in the mirror, I'm gorgeous. When I look at

0:40:44.000 --> 0:40:47.680
<v Speaker 3>a photo, I look like a crescent moon. I noticed

0:40:47.719 --> 0:40:51.600
<v Speaker 3>that my face is curved, and I never noticed that

0:40:51.640 --> 0:40:54.600
<v Speaker 3>in the mirror because it's just a perspective I'm not

0:40:54.719 --> 0:40:56.520
<v Speaker 3>used to. And I think that's exactly what happens with

0:40:56.600 --> 0:40:58.960
<v Speaker 3>your voices. You're used to hearing it a certain way,

0:40:59.280 --> 0:41:01.960
<v Speaker 3>and no matter how the recording sounds, you'll think it

0:41:02.000 --> 0:41:04.000
<v Speaker 3>sounds worse because you're just not used to it.

0:41:04.400 --> 0:41:07.560
<v Speaker 2>So somebody was to put their cheek up against Kelly

0:41:07.600 --> 0:41:10.080
<v Speaker 2>while she was speaking, so that her voice traveled like

0:41:10.160 --> 0:41:13.560
<v Speaker 2>through their soft tissue, would they hear Kelly's internal version

0:41:13.560 --> 0:41:14.200
<v Speaker 2>of her voice?

0:41:14.600 --> 0:41:20.439
<v Speaker 3>If you were somehow able to model Kelly's innards from

0:41:20.520 --> 0:41:23.080
<v Speaker 3>her throat all the way up to the top of

0:41:23.120 --> 0:41:27.399
<v Speaker 3>her skull and put that between the other person, yes,

0:41:27.719 --> 0:41:28.800
<v Speaker 3>it would be very strange.

0:41:28.800 --> 0:41:31.560
<v Speaker 1>But yes, I'd just like to put on record that

0:41:31.600 --> 0:41:33.840
<v Speaker 1>when I see you, I never think Crescent.

0:41:33.520 --> 0:41:36.560
<v Speaker 3>Moved, but I will now I have to tell myself

0:41:36.600 --> 0:41:39.480
<v Speaker 3>the same thing is that that tiny lack of symmetry

0:41:39.480 --> 0:41:42.879
<v Speaker 3>which every human has is just accentuated because I never

0:41:42.920 --> 0:41:46.279
<v Speaker 3>see it from that angle. And I feel like most

0:41:46.280 --> 0:41:48.719
<v Speaker 3>people when they look at a photo compared to the mirror,

0:41:48.719 --> 0:41:50.600
<v Speaker 3>they're like, oh, what is that? What happened?

0:41:50.719 --> 0:41:51.600
<v Speaker 1>Yeah?

0:41:51.680 --> 0:41:54.160
<v Speaker 2>So you mentioned earlier about real magic and how big

0:41:54.200 --> 0:41:56.480
<v Speaker 2>part of this is knowing how it's going to be

0:41:56.560 --> 0:41:59.960
<v Speaker 2>received by people. Tell us more about that, about how

0:42:00.320 --> 0:42:02.719
<v Speaker 2>understanding the way sound is actually perceived and the way

0:42:02.719 --> 0:42:06.040
<v Speaker 2>people pay attention changes how you do your editing and

0:42:06.080 --> 0:42:08.520
<v Speaker 2>tell us about psychoacoustics.

0:42:09.120 --> 0:42:12.239
<v Speaker 3>So, Daniel, you're gonna be proud of me. Before I

0:42:12.400 --> 0:42:15.600
<v Speaker 3>even knew who you were. I've talked to my students

0:42:15.600 --> 0:42:16.600
<v Speaker 3>about aliens a.

0:42:16.560 --> 0:42:19.720
<v Speaker 2>Lot plus ten points.

0:42:21.280 --> 0:42:23.520
<v Speaker 3>Yeah. The reason I do that is I always in

0:42:23.760 --> 0:42:26.440
<v Speaker 3>the first class. I tell them this is a class

0:42:26.920 --> 0:42:33.200
<v Speaker 3>for human beings in Earth's atmosphere, because, first of all,

0:42:33.280 --> 0:42:36.920
<v Speaker 3>sound behaves differently through different materials, and second, humans have

0:42:37.040 --> 0:42:40.400
<v Speaker 3>evolved to process their senses in ways that are useful

0:42:40.400 --> 0:42:43.879
<v Speaker 3>to them, which is another way of saying that what

0:42:44.080 --> 0:42:47.640
<v Speaker 3>we hear and see and feel is not reality, but

0:42:48.040 --> 0:42:52.640
<v Speaker 3>our useful version of reality, and sound is no exception

0:42:52.760 --> 0:42:56.040
<v Speaker 3>to that. I think we may not realize how much

0:42:56.080 --> 0:42:58.719
<v Speaker 3>of a role it plays. My theory is because there

0:42:58.760 --> 0:43:06.640
<v Speaker 3>are no earlids go on. With sight, you have these

0:43:06.680 --> 0:43:09.480
<v Speaker 3>eyelids where you can close them and not see and

0:43:09.560 --> 0:43:11.480
<v Speaker 3>get an idea of what it's like to not have

0:43:11.560 --> 0:43:13.960
<v Speaker 3>that sense. You also have to directly look at the

0:43:14.000 --> 0:43:17.120
<v Speaker 3>thing that you're looking at to perceive it, and with sound,

0:43:17.320 --> 0:43:20.319
<v Speaker 3>you hear everything around you. You don't have to look

0:43:20.360 --> 0:43:23.440
<v Speaker 3>at it with your ear, and you've never closed your ears.

0:43:24.239 --> 0:43:28.080
<v Speaker 3>Your earballs are always exposed. So it's a sense that's

0:43:28.640 --> 0:43:30.960
<v Speaker 3>so enveloping and so much part of our lives that

0:43:31.000 --> 0:43:33.320
<v Speaker 3>we sort of take it for granted and not understand

0:43:33.360 --> 0:43:38.120
<v Speaker 3>how much information we're constantly getting from our surroundings through sound.

0:43:38.080 --> 0:43:40.279
<v Speaker 2>And we don't even have the language for it, right,

0:43:40.360 --> 0:43:43.640
<v Speaker 2>you have to adapt the language from eyes. Like you said,

0:43:43.800 --> 0:43:46.120
<v Speaker 2>look at it with your ear instead of hear at

0:43:46.160 --> 0:43:48.680
<v Speaker 2>it with your ear, because that makes no sense, right.

0:43:49.280 --> 0:43:52.400
<v Speaker 3>Yeah, So we think there's just this baseline of nothing,

0:43:52.480 --> 0:43:55.640
<v Speaker 3>but in reality, you know the size of the room

0:43:55.640 --> 0:43:57.319
<v Speaker 3>you're in and what the walls are made out of

0:43:57.600 --> 0:44:01.120
<v Speaker 3>before you look and tap them. Walls, you know from

0:44:01.160 --> 0:44:02.920
<v Speaker 3>how sound reflects in that room.

0:44:03.000 --> 0:44:04.759
<v Speaker 2>I have a question about that because I've been told

0:44:04.760 --> 0:44:08.000
<v Speaker 2>many times to record in a room that doesn't have

0:44:08.120 --> 0:44:11.920
<v Speaker 2>auditory reflections, right, like surround yourself with laundry or something,

0:44:12.320 --> 0:44:14.520
<v Speaker 2>because nobody wants to hear the room that you're in.

0:44:15.360 --> 0:44:18.279
<v Speaker 2>So isn't that weird for people to hear audio that

0:44:18.320 --> 0:44:21.520
<v Speaker 2>sounds sort of like disembodied or de rumified.

0:44:22.719 --> 0:44:26.280
<v Speaker 3>No, because it's far from disembodied and de rumified unless

0:44:26.320 --> 0:44:28.279
<v Speaker 3>you've been to an anichoid chamber. You have never heard

0:44:28.320 --> 0:44:31.200
<v Speaker 3>silence in your life. It's a have you been to one?

0:44:31.280 --> 0:44:31.360
<v Speaker 1>No?

0:44:31.520 --> 0:44:34.240
<v Speaker 3>I highly recommend it. It's weird.

0:44:34.480 --> 0:44:36.400
<v Speaker 2>Sounds like a torture device, you know what.

0:44:36.520 --> 0:44:39.080
<v Speaker 3>People say that there's a rumor that no one has

0:44:39.120 --> 0:44:41.600
<v Speaker 3>ever lasted more than forty five minutes. Not true. I've

0:44:41.600 --> 0:44:44.640
<v Speaker 3>been there for hours. But you can hear blood rushing

0:44:44.640 --> 0:44:48.880
<v Speaker 3>through your ears. You can hear everything. And not only that,

0:44:49.040 --> 0:44:51.200
<v Speaker 3>if you're there with a buddy and they turn around

0:44:51.400 --> 0:44:53.320
<v Speaker 3>and talk and speak the other way, you can barely

0:44:53.360 --> 0:44:56.799
<v Speaker 3>hear them. We don't realize, wow, or we take for

0:44:56.840 --> 0:44:59.359
<v Speaker 3>granted how much of what we hear is actually reflections

0:44:59.719 --> 0:45:02.799
<v Speaker 3>of the source and not the source itself. So I

0:45:02.800 --> 0:45:07.480
<v Speaker 3>mean an example of some important psychoacoustics that engineers have

0:45:07.520 --> 0:45:10.000
<v Speaker 3>to keep in mind. And also the answer to the

0:45:10.040 --> 0:45:14.760
<v Speaker 3>listener question, what do you think what is the type

0:45:14.760 --> 0:45:16.319
<v Speaker 3>of sound or sound that humans hear?

0:45:16.360 --> 0:45:19.120
<v Speaker 2>Best? I would have guessed that there's something similar to

0:45:19.280 --> 0:45:22.040
<v Speaker 2>our visual arrange, where there's a whole spectrum of frequencies

0:45:22.440 --> 0:45:24.839
<v Speaker 2>and we can see something in the middle. Where in

0:45:24.880 --> 0:45:28.200
<v Speaker 2>the middle basically means where we can hear, which makes

0:45:28.239 --> 0:45:29.080
<v Speaker 2>it kind of a non.

0:45:28.920 --> 0:45:33.840
<v Speaker 3>Answer, right, Well, that's actually that's a good comparison with sound.

0:45:34.000 --> 0:45:36.200
<v Speaker 3>You can see red just as well as you can

0:45:36.239 --> 0:45:40.479
<v Speaker 3>see blue with sight, but with sound it doesn't work

0:45:40.520 --> 0:45:44.120
<v Speaker 3>that way. With sound, we can hear frequencies from twenty

0:45:44.200 --> 0:45:47.440
<v Speaker 3>vibrations per second twenty hurts to twenty thousand vibrations per

0:45:47.480 --> 0:45:49.919
<v Speaker 3>second twenty killer hurts, but out of those we don't

0:45:49.960 --> 0:45:52.680
<v Speaker 3>hear them all equally. The frequency range we hear best

0:45:53.320 --> 0:45:55.680
<v Speaker 3>is between one and five killer hurts, and that is

0:45:55.719 --> 0:46:01.040
<v Speaker 3>the range of leaves rustling and breaking.

0:46:00.880 --> 0:46:02.960
<v Speaker 2>And jaguar is creeping up on you so that.

0:46:02.960 --> 0:46:05.360
<v Speaker 3>Predators don't creep up on you. It's also the range

0:46:05.360 --> 0:46:08.680
<v Speaker 3>of a baby crying. You can hear a baby cry

0:46:08.800 --> 0:46:12.440
<v Speaker 3>from way farther away than I don't know a car

0:46:12.640 --> 0:46:14.960
<v Speaker 3>engine reving at the same distance, same volume.

0:46:15.360 --> 0:46:17.600
<v Speaker 2>It's also the sound of a candy bar wrapper, isn't

0:46:17.600 --> 0:46:18.440
<v Speaker 2>it That is?

0:46:19.120 --> 0:46:23.480
<v Speaker 3>Yes, yes, it is, and we've been using that. For example.

0:46:23.520 --> 0:46:26.680
<v Speaker 3>Police sirens are tuned to that frequency range, and that's

0:46:26.680 --> 0:46:28.759
<v Speaker 3>why you may have found yourself sitting in your car

0:46:28.840 --> 0:46:30.920
<v Speaker 3>or walking hearing a siren looking around being like, I

0:46:30.920 --> 0:46:32.879
<v Speaker 3>don't see the police anywhere. Where are they? It's because

0:46:32.920 --> 0:46:36.759
<v Speaker 3>we're so good at hearing those frequencies that we can

0:46:36.800 --> 0:46:39.120
<v Speaker 3>hear them from so far away that we can't even

0:46:39.160 --> 0:46:40.640
<v Speaker 3>see the source of the sound yet.

0:46:40.680 --> 0:46:40.960
<v Speaker 1>Wow.

0:46:41.000 --> 0:46:43.160
<v Speaker 2>So then how does all of this shape your editing?

0:46:43.840 --> 0:46:45.920
<v Speaker 2>Are you trying to push my voice up to police

0:46:45.960 --> 0:46:46.880
<v Speaker 2>siren frequencies?

0:46:48.080 --> 0:46:52.680
<v Speaker 3>Well, I definitely have to make sure that intelligibility is

0:46:52.719 --> 0:46:55.960
<v Speaker 3>there and that is in those frequencies. But there's a

0:46:56.040 --> 0:46:59.600
<v Speaker 3>challenge here because we hear those frequencies so well, they

0:46:59.640 --> 0:47:04.359
<v Speaker 3>become annoying very quickly too, because we're acutely tuned to them.

0:47:04.880 --> 0:47:07.839
<v Speaker 3>So there's a balancing act where if there's not enough

0:47:08.440 --> 0:47:10.520
<v Speaker 3>it sounds I can show you. Do you want me

0:47:10.560 --> 0:47:11.239
<v Speaker 3>to show you? Yeah?

0:47:11.320 --> 0:47:11.960
<v Speaker 1>Yeah, please?

0:47:12.160 --> 0:47:15.160
<v Speaker 3>So here let me play process Daniel here.

0:47:15.520 --> 0:47:19.759
<v Speaker 2>So Bell's paradox involves length contraction, the fact that when

0:47:20.040 --> 0:47:23.799
<v Speaker 2>things move fast, they look short. So relativity tells us

0:47:23.800 --> 0:47:28.600
<v Speaker 2>two things, moving clocks run slow and moving objects look short.

0:47:29.320 --> 0:47:31.400
<v Speaker 3>So now I'm going to take out that frequency range

0:47:31.800 --> 0:47:33.920
<v Speaker 3>that I was just talking about, and you'll notice so

0:47:33.960 --> 0:47:37.560
<v Speaker 3>you can still hear Daniel. He's loud, but the part

0:47:37.600 --> 0:47:40.279
<v Speaker 3>that makes him human loud, I mean loud is the

0:47:40.400 --> 0:47:42.760
<v Speaker 3>volume I made him. Not because Daniel is a loud person,

0:47:43.560 --> 0:47:46.920
<v Speaker 3>but the intelligibility is just not there. Take a listen.

0:47:47.600 --> 0:47:51.839
<v Speaker 2>So Bell's paradox involves length contraction, the fact that when

0:47:52.120 --> 0:47:55.840
<v Speaker 2>things move fast, they look short. So relativity tells us

0:47:55.880 --> 0:48:00.680
<v Speaker 2>two things, moving clocks run slow and moving objects look short.

0:48:01.239 --> 0:48:02.319
<v Speaker 1>Huh.

0:48:02.360 --> 0:48:04.799
<v Speaker 3>And now I'll exaggerate those frequencies and you'll see that

0:48:04.800 --> 0:48:06.040
<v Speaker 3>it'll kind of hurt your ears.

0:48:06.480 --> 0:48:10.719
<v Speaker 2>So Bell's paradox involves length contraction, the fact that when

0:48:11.000 --> 0:48:13.120
<v Speaker 2>things move fast they look short.

0:48:13.400 --> 0:48:15.760
<v Speaker 3>Ooh right, So I have to find a happy medium,

0:48:15.800 --> 0:48:18.200
<v Speaker 3>and that is sometimes something that has to be done temporally.

0:48:18.280 --> 0:48:21.040
<v Speaker 3>So I have to make sure that some words don't

0:48:21.080 --> 0:48:24.680
<v Speaker 3>have too much of those frequencies and others do. That's

0:48:24.719 --> 0:48:27.440
<v Speaker 3>just one of many ways in which I have to

0:48:27.600 --> 0:48:31.359
<v Speaker 3>look at the waveforms and understand the science and then

0:48:31.480 --> 0:48:36.120
<v Speaker 3>also understand the psychoacoustics, the psychology of how humans perceive

0:48:36.200 --> 0:48:36.640
<v Speaker 3>this stuff.

0:48:36.680 --> 0:48:39.080
<v Speaker 2>Well, you're doing a great job, because I sometimes meet

0:48:39.120 --> 0:48:41.160
<v Speaker 2>listeners in real life and they'll say to me, oh

0:48:41.160 --> 0:48:43.239
<v Speaker 2>my gosh, it's so weird. You sound just like you

0:48:43.320 --> 0:48:46.880
<v Speaker 2>do on the podcast. So you are reproducing this faithfully.

0:48:47.239 --> 0:48:48.399
<v Speaker 3>I did it, woo.

0:48:48.640 --> 0:48:51.360
<v Speaker 1>So how much physics did you have to learn to

0:48:51.480 --> 0:48:52.760
<v Speaker 1>be this good? At your job.

0:48:53.080 --> 0:48:56.360
<v Speaker 3>I see what you're saying. Yes, audio engineers are real engineers.

0:48:57.239 --> 0:49:02.000
<v Speaker 3>We're like the dentists of the audio profession.

0:49:02.120 --> 0:49:03.680
<v Speaker 2>She's just trying to sus that if you're on the

0:49:03.680 --> 0:49:06.640
<v Speaker 2>physics camp or the biology camp, she's wondering if you're biased.

0:49:07.000 --> 0:49:09.239
<v Speaker 3>See, I'm stuck right in the middle between physics and

0:49:09.280 --> 0:49:10.920
<v Speaker 3>biology exactly.

0:49:11.160 --> 0:49:12.040
<v Speaker 1>It's a good place to be.

0:49:12.880 --> 0:49:15.120
<v Speaker 3>Yeah, it's it's a very fun place to be because

0:49:15.120 --> 0:49:17.160
<v Speaker 3>who cares about physics if it's not for how it

0:49:17.200 --> 0:49:19.040
<v Speaker 3>affects sentient beings?

0:49:20.200 --> 0:49:21.040
<v Speaker 1>Amen?

0:49:21.320 --> 0:49:24.440
<v Speaker 2>Amen, hmmm, I have thoughts of that.

0:49:24.719 --> 0:49:26.440
<v Speaker 1>I'm so glad we got you. You know, Matt, you

0:49:26.440 --> 0:49:28.000
<v Speaker 1>should be on the show every week.

0:49:28.200 --> 0:49:30.719
<v Speaker 2>I see, Kelly just wants to invite people on the

0:49:30.719 --> 0:49:33.560
<v Speaker 2>show who will disagree with me gang up on me.

0:49:35.239 --> 0:49:39.520
<v Speaker 3>So let me tell you another psycho acoustic phenomenon called masking.

0:49:40.320 --> 0:49:44.120
<v Speaker 3>The way this works and this probably like we can't

0:49:44.160 --> 0:49:47.680
<v Speaker 3>ask evolution, right, but it looks like we've developed this

0:49:47.800 --> 0:49:53.239
<v Speaker 3>property to u curb over stimulation. What happens is if

0:49:53.239 --> 0:49:55.279
<v Speaker 3>there are certain and it's very predictable, if there are

0:49:55.280 --> 0:49:59.600
<v Speaker 3>certain frequencies present, we are deaf to certain other frequencies

0:50:00.480 --> 0:50:03.279
<v Speaker 3>in those moments. It's just like a blind spot in

0:50:03.320 --> 0:50:05.480
<v Speaker 3>the car, Like, no matter what mirror you look at,

0:50:05.520 --> 0:50:07.200
<v Speaker 3>you will not see the car right next to you

0:50:07.239 --> 0:50:08.839
<v Speaker 3>if they're in your blind spot. So we have these

0:50:08.880 --> 0:50:12.800
<v Speaker 3>blind spots that are constantly, very quickly changing as different

0:50:12.840 --> 0:50:18.520
<v Speaker 3>sounds are present, and very smart programmers mapped out the

0:50:18.640 --> 0:50:23.120
<v Speaker 3>exact way we have these blind spots and figured out,

0:50:23.160 --> 0:50:25.080
<v Speaker 3>you know what, we can save a lot of space

0:50:25.080 --> 0:50:27.799
<v Speaker 3>and audio files if we just throw that stuff away.

0:50:27.960 --> 0:50:31.759
<v Speaker 3>Oh oh wow. And that is how the MP three

0:50:31.840 --> 0:50:34.759
<v Speaker 3>file format was born. It does throw away a lot

0:50:34.800 --> 0:50:37.920
<v Speaker 3>of things, but it throws away only things you would

0:50:37.920 --> 0:50:40.560
<v Speaker 3>not perceive anyway. And so if you've ever had this

0:50:41.320 --> 0:50:43.919
<v Speaker 3>debate with some or just heard someone claim that MP

0:50:43.960 --> 0:50:47.600
<v Speaker 3>three's just don't sound like waves, man, there's just something missing.

0:50:48.120 --> 0:50:50.560
<v Speaker 3>Not necessarily, if it was a well encoded MP three,

0:50:51.360 --> 0:50:56.000
<v Speaker 3>you will never hear what is missing. Wow, Because you're human.

0:50:56.680 --> 0:51:01.360
<v Speaker 3>Our imperfect hearing has really helped us on file sizes. However,

0:51:01.480 --> 0:51:04.120
<v Speaker 3>if an alien came and you went, here's a wavefile,

0:51:04.200 --> 0:51:06.200
<v Speaker 3>like here's a recording, and then here's my MP three

0:51:06.239 --> 0:51:09.600
<v Speaker 3>of that recording to that alien, it might sound completely

0:51:09.640 --> 0:51:13.720
<v Speaker 3>different and upsetting because they don't have our same perception.

0:51:14.080 --> 0:51:16.600
<v Speaker 2>Those aliens who listen to the podcast might be really

0:51:16.640 --> 0:51:20.960
<v Speaker 2>sensitive to those frequencies that our audience, our human audience,

0:51:21.040 --> 0:51:21.520
<v Speaker 2>can't hear.

0:51:22.040 --> 0:51:23.759
<v Speaker 3>As far as we know, when we listen to an

0:51:23.840 --> 0:51:28.160
<v Speaker 3>MP three, our pets are suffering. They don't seem to

0:51:28.200 --> 0:51:31.239
<v Speaker 3>be complaining, but they're definitely not hearing what we hear

0:51:31.280 --> 0:51:33.200
<v Speaker 3>when we listen to the original recording the way that

0:51:33.280 --> 0:51:33.719
<v Speaker 3>humans do.

0:51:33.840 --> 0:51:36.680
<v Speaker 2>So this is like the moral equivalent of masking out

0:51:36.680 --> 0:51:40.239
<v Speaker 2>the ultraviolet off of a painting, and a bee might

0:51:40.239 --> 0:51:42.719
<v Speaker 2>see a painting very very differently, whereas a human would

0:51:42.719 --> 0:51:44.000
<v Speaker 2>be totally unaware.

0:51:44.760 --> 0:51:47.720
<v Speaker 3>Kind of It's like if humans, when you see yellow,

0:51:47.960 --> 0:51:49.800
<v Speaker 3>you're temporarily blind to purple.

0:51:51.320 --> 0:51:55.360
<v Speaker 1>Weird, Why why why does that happen?

0:51:56.280 --> 0:51:58.959
<v Speaker 3>I don't know. I haven't asked God yet, but I

0:51:59.160 --> 0:52:02.480
<v Speaker 3>think probably because of the overstimulation thing. I think if

0:52:02.520 --> 0:52:07.000
<v Speaker 3>we were to hear everything all at once, our brains

0:52:07.000 --> 0:52:09.280
<v Speaker 3>would explode. I think it's just too much information.

0:52:09.880 --> 0:52:13.400
<v Speaker 1>So people who get over stimulated, do they hear some

0:52:13.520 --> 0:52:15.040
<v Speaker 1>of that or is that just a different thing?

0:52:15.680 --> 0:52:19.560
<v Speaker 3>You know, that's a great question. That might be part

0:52:19.600 --> 0:52:21.200
<v Speaker 3>of it, That might be part of it that they

0:52:21.400 --> 0:52:24.040
<v Speaker 3>don't experience masking the same way. But the interesting is

0:52:24.040 --> 0:52:27.080
<v Speaker 3>that it's very predictable, this masking and how it happens.

0:52:27.080 --> 0:52:30.360
<v Speaker 3>It's very specific, and it's the same for children and

0:52:30.400 --> 0:52:33.239
<v Speaker 3>for women and for men and for everybody. I don't

0:52:33.280 --> 0:52:35.719
<v Speaker 3>know that there's a group of people to whom MP

0:52:35.800 --> 0:52:40.560
<v Speaker 3>three's don't sound right objectively, like a blind test. Yeah,

0:52:40.920 --> 0:52:45.200
<v Speaker 3>so yeah, I don't know. Maybe some people do experience

0:52:45.280 --> 0:52:46.759
<v Speaker 3>less masking. Maybe that's part of it.

0:52:47.600 --> 0:52:50.480
<v Speaker 1>So, Matt, I imagine you have loads of fun stories,

0:52:50.520 --> 0:52:52.840
<v Speaker 1>in part because like Daniel and I always hit record

0:52:52.880 --> 0:52:55.320
<v Speaker 1>and then we're like, we totally forgot. We hit record,

0:52:55.360 --> 0:52:58.440
<v Speaker 1>and we have conversations about the most personal things happening

0:52:58.440 --> 0:53:00.680
<v Speaker 1>in our lives, and then we send it to you,

0:53:01.160 --> 0:53:04.120
<v Speaker 1>and so you it's a very asymmetrical relationship. I think

0:53:04.160 --> 0:53:06.040
<v Speaker 1>you know a lot about the things that are going

0:53:06.040 --> 0:53:08.040
<v Speaker 1>wrong in my life. I don't know nearly as much

0:53:08.040 --> 0:53:09.520
<v Speaker 1>about what's going on in your life.

0:53:09.600 --> 0:53:10.960
<v Speaker 3>I can make you a list if you want to

0:53:11.200 --> 0:53:12.760
<v Speaker 3>even the playing field.

0:53:13.800 --> 0:53:16.520
<v Speaker 1>I mean that lay it on me. But but so

0:53:16.600 --> 0:53:18.800
<v Speaker 1>I imagine that one you know a lot of secrets

0:53:18.800 --> 0:53:22.120
<v Speaker 1>from people and tell us some weird, funny stories from

0:53:22.160 --> 0:53:24.280
<v Speaker 1>from life as an audio engineer.

0:53:24.840 --> 0:53:30.280
<v Speaker 3>Uh, there are many inappropriate stories. Uh there's okay, here's

0:53:30.360 --> 0:53:35.120
<v Speaker 3>one that is right in the middle between science and

0:53:35.120 --> 0:53:40.279
<v Speaker 3>and some biology and perception, so mostly physics. So I

0:53:40.440 --> 0:53:43.480
<v Speaker 3>was in the Santa Monica, which is uh is it

0:53:43.600 --> 0:53:45.279
<v Speaker 3>a part of LA or is it outside of La

0:53:45.719 --> 0:53:48.840
<v Speaker 3>boy Daniel, California.

0:53:48.840 --> 0:53:50.799
<v Speaker 2>You should know this. I don't know. It's well, you know,

0:53:50.840 --> 0:53:53.040
<v Speaker 2>there's the affective LA which is a big blob of

0:53:53.120 --> 0:53:56.480
<v Speaker 2>urbanized area, then there's La County, then there's LA City,

0:53:56.520 --> 0:53:58.560
<v Speaker 2>and I don't know the distinctions between all of them.

0:53:58.840 --> 0:54:01.359
<v Speaker 3>Yeah, I don't understand how that's a municipality because when

0:54:01.400 --> 0:54:04.040
<v Speaker 3>I was working there, they were shuttle us around from

0:54:04.200 --> 0:54:07.200
<v Speaker 3>location to location, and the names of the neighborhoods and

0:54:07.840 --> 0:54:11.120
<v Speaker 3>how things fit into the municipalities didn't make any sense. Anyway,

0:54:11.160 --> 0:54:17.000
<v Speaker 3>I was in Santa Monica filming with Smokey Robinson, Wow

0:54:17.840 --> 0:54:21.040
<v Speaker 3>for a very interesting documentary by the way, called Paid

0:54:21.040 --> 0:54:25.239
<v Speaker 3>in Full on the BBC and CBC wherever you are,

0:54:25.440 --> 0:54:30.480
<v Speaker 3>about the unfair treatment of black musicians throughout history. In

0:54:30.520 --> 0:54:34.960
<v Speaker 3>any case, I was there with Smokey Robinson and they

0:54:34.960 --> 0:54:37.840
<v Speaker 3>rented a special villa just for him. You know. The

0:54:37.920 --> 0:54:42.360
<v Speaker 3>VIP treatment, and there's always this air of when a

0:54:42.480 --> 0:54:47.160
<v Speaker 3>really important person comes in, no one stops the recording

0:54:47.239 --> 0:54:51.040
<v Speaker 3>for any reason. Right, do your job be quiet. We

0:54:51.120 --> 0:54:53.000
<v Speaker 3>have a very limited amount of time with this person.

0:54:53.520 --> 0:54:55.759
<v Speaker 3>So I'm sitting there, I'm set up, I have my

0:54:55.760 --> 0:54:58.399
<v Speaker 3>computer open, I've miked him up, I have a mic

0:54:58.480 --> 0:55:00.520
<v Speaker 3>on him, I have a mic above him, and I

0:55:00.560 --> 0:55:03.399
<v Speaker 3>have a spectrum analyzer open, just because I didn't really

0:55:03.440 --> 0:55:06.120
<v Speaker 3>need it, but it shows you all the frequencies in

0:55:06.160 --> 0:55:09.680
<v Speaker 3>real time. And suddenly during the interview, I see these

0:55:09.840 --> 0:55:14.600
<v Speaker 3>sort of spikes at like five hurts, which is way

0:55:14.640 --> 0:55:18.080
<v Speaker 3>lower than what humans can hear. And I'm realizing that

0:55:18.800 --> 0:55:21.480
<v Speaker 3>even though the microphone is suspended and they're supposed to

0:55:21.600 --> 0:55:24.880
<v Speaker 3>not pick anything up through the floor, I've created a

0:55:24.920 --> 0:55:29.200
<v Speaker 3>seismometer by accident. Oh and I can see that there's

0:55:29.200 --> 0:55:30.000
<v Speaker 3>an earthquake coming.

0:55:30.200 --> 0:55:31.600
<v Speaker 2>Oh wow wow.

0:55:31.800 --> 0:55:35.480
<v Speaker 3>And so in the middle of this interview, I go, hey,

0:55:35.600 --> 0:55:37.400
<v Speaker 3>excuse me, and the director looks at me, like what

0:55:37.440 --> 0:55:42.640
<v Speaker 3>are you doing, And I'm like, I'm very very sorry.

0:55:43.040 --> 0:55:45.839
<v Speaker 3>I think there's an earthquake coming and we should get

0:55:45.880 --> 0:55:48.319
<v Speaker 3>ready to evacuate or move away from the walls or

0:55:48.680 --> 0:55:50.600
<v Speaker 3>get too though I don't remember what you're supposed to do.

0:55:50.640 --> 0:55:53.240
<v Speaker 3>But we're in California. I expected everyone to know except

0:55:53.280 --> 0:55:56.040
<v Speaker 3>for me, and they're like, what are you talking about?

0:55:56.040 --> 0:56:00.200
<v Speaker 3>There's nothing happening, you've this is Smokey Robinson. What are

0:56:00.200 --> 0:56:02.800
<v Speaker 3>you doing? I mean, everybody was nice, we're all friends.

0:56:02.800 --> 0:56:06.960
<v Speaker 3>But as that's happening, the lights start to shake. Oh wow,

0:56:07.160 --> 0:56:10.080
<v Speaker 3>and there's a rumble. And luckily it wasn't a big

0:56:10.080 --> 0:56:13.439
<v Speaker 3>earthquake or anything. Everything was fine. But I predicted on an

0:56:13.440 --> 0:56:16.360
<v Speaker 3>earthquake using a microphone and a computer screen.

0:56:16.800 --> 0:56:17.520
<v Speaker 1>That's amazing.

0:56:17.560 --> 0:56:19.600
<v Speaker 3>And I interrupted an important recording.

0:56:19.440 --> 0:56:22.000
<v Speaker 1>And that's magic, right, it's Smokey Robinson is now like

0:56:22.040 --> 0:56:23.960
<v Speaker 1>Matt Kesselman is magic. Everybody knows.

0:56:24.400 --> 0:56:26.359
<v Speaker 3>He was like, that's okay, baby, he's got a very

0:56:26.440 --> 0:56:27.239
<v Speaker 3>very soft voice.

0:56:27.280 --> 0:56:27.799
<v Speaker 2>I love him.

0:56:29.080 --> 0:56:31.640
<v Speaker 1>Well, that is a great story. I don't have any

0:56:31.840 --> 0:56:36.640
<v Speaker 1>comparable stories for warning famous people about impending natural disasters.

0:56:37.160 --> 0:56:38.760
<v Speaker 3>It's a very niche story.

0:56:38.760 --> 0:56:42.759
<v Speaker 2>Market all right, well, thank you for coming on the show,

0:56:42.840 --> 0:56:45.120
<v Speaker 2>and we're telling us all about what happens in the

0:56:45.120 --> 0:56:48.160
<v Speaker 2>background to make us sound so good. Before you go,

0:56:48.280 --> 0:56:50.359
<v Speaker 2>tell people where they can find out more about you

0:56:50.440 --> 0:56:51.120
<v Speaker 2>and what's.

0:56:51.000 --> 0:56:54.759
<v Speaker 3>Up to you. Can find me for everything at Matt

0:56:55.160 --> 0:56:59.200
<v Speaker 3>Kesselman dot com. That's m A T K E S

0:56:59.280 --> 0:57:02.880
<v Speaker 3>E l m N dot com. And if you don't

0:57:03.000 --> 0:57:06.520
<v Speaker 3>like to spell things that are complicated, I just bought

0:57:06.680 --> 0:57:10.040
<v Speaker 3>Thegoshfather dot com and dot cost to my website too,

0:57:10.800 --> 0:57:12.680
<v Speaker 3>so you can use that instead.

0:57:13.480 --> 0:57:16.200
<v Speaker 1>Matt, I appreciate you so much for so many reasons,

0:57:16.240 --> 0:57:18.040
<v Speaker 1>and now I appreciate that you came on the show.

0:57:18.160 --> 0:57:18.760
<v Speaker 1>You're the best.

0:57:18.840 --> 0:57:21.120
<v Speaker 2>There are lots of reasons why the show sounds good

0:57:21.200 --> 0:57:23.120
<v Speaker 2>and it's fun to listen to and easy to listen to,

0:57:23.200 --> 0:57:25.600
<v Speaker 2>and many of those are because of Matt. Thank you

0:57:25.680 --> 0:57:26.120
<v Speaker 2>very much.

0:57:26.320 --> 0:57:27.960
<v Speaker 3>Amen, thank you for having me. This was a lot

0:57:27.960 --> 0:57:28.280
<v Speaker 3>of fun.

0:57:35.120 --> 0:57:37.560
<v Speaker 2>Thanks everybody for listening. Please go and do us a

0:57:37.560 --> 0:57:40.840
<v Speaker 2>favor and rate the show on whatever podcast app you're using.

0:57:40.920 --> 0:57:42.520
<v Speaker 2>It really helps people find us.

0:57:43.040 --> 0:57:46.960
<v Speaker 1>Daniel and Kelly's Extraordinary Universe is edited by the amazing

0:57:47.000 --> 0:57:47.720
<v Speaker 1>Matt Kesselman.

0:57:47.960 --> 0:57:51.200
<v Speaker 2>He really is a wizard. You can also find us

0:57:51.280 --> 0:57:56.400
<v Speaker 2>online on Blue Sky, Instagram, and x d K Universe.

0:57:56.440 --> 0:57:57.680
<v Speaker 2>Come engage with us.

0:57:57.920 --> 0:58:00.280
<v Speaker 1>You can email us I had questions at Dan and

0:58:00.400 --> 0:58:03.120
<v Speaker 1>Kelly dot org. We really do want to hear from you,

0:58:03.280 --> 0:58:03.520
<v Speaker 1>and you.

0:58:03.520 --> 0:58:07.520
<v Speaker 2>Can find our website www dot danieland Kelly dot org,

0:58:07.800 --> 0:58:10.920
<v Speaker 2>where you'll also find an invitation to join our discord

0:58:11.000 --> 0:58:14.040
<v Speaker 2>where everybody comes and talks about the amazing.

0:58:13.720 --> 0:58:17.960
<v Speaker 1>Universe, and we also have the most amazing moderators. This

0:58:18.400 --> 0:58:21.040
<v Speaker 1>is an iHeart podcast. Thanks for joining us.