00:00:02 Speaker 1: Bloomberg Audio Studios, Podcasts, radio News. 00:00:18 Speaker 2: Hello and welcome to another episode of the Odd Lots Podcast. 00:00:21 Speaker 3: I'm Joe Wisenthal and I'm Tracy Alloway. 00:00:24 Speaker 4: Racy, I have a question. 00:00:25 Speaker 2: I know you spent a lot of your youth overseas my youth. Did you ever go to many baseball games as a kid. 00:00:31 Speaker 3: Yeah, so I was in Chicago for a few years, so I went to the Cups games, and then I was in Japan, and the baseball scene in Japan is amazing, like the best vibes of a live sports event that I've ever ever witnessed. Her encounter. 00:00:46 Speaker 2: I was just talking to someone about this last night. I've always wanted to go to a Japanese baseball record. That's like, if we ever do a live show in Tokyo, let's try to schedule it during baseball season, because it is like I sort of think that's a bucket list thing. But anyway, the reason I asked this question is I have this really vague memory as a child going to see Detroit Tigers games with my grandfather, like probably when I was like maybe these memories are probably from when I was younger than like six or five, but there used to be a non trivial number of people who would go to the games and they would keep score, and they would write down every single a bat and they the outcome of every single. 00:01:25 Speaker 3: Your I've seen that, not in person, but like maybe movies or something like that. 00:01:29 Speaker 2: It doesn't really happen anymore, like I never see it. There are may be a few, like old timers who still has a hobby or habit do that, but it was like a non insignificant number of people. And it's interesting to me. You know, I've been to a couple of soccer games this year. There's no equivalent way you could do that, right, because it's like baseball is filled with all of these discrete events. The picture, who is the picture? Who is the batter? 00:01:52 Speaker 5: Hit? 00:01:53 Speaker 2: Not a hit, single, not a strikeout, walk, et cetera. Like what would even be the equivalent to soccer? 00:01:58 Speaker 3: So this has been a long running debate in soccer. And I remember when moneyball came out, and you know, sports analytics became a big thing, especially for baseball because as you point out, it's these sort of discrete events that have a lot of statistics embedded in them. A lot of people were saying that soccer. You could never use data analytics in the same way for soccer, Like it's too chaotic, it's too fluid, there's too many variables, there's not enough goals. That complaint comes up a lot whenever we talk about are we going to say soccer or football? By the way, you know what, let's just say soccer. Okay, all right, I'll try. 00:02:31 Speaker 2: We can say football. I actually don't feel strongly about this one. 00:02:34 Speaker 3: But that said, we do see soccer analytics on the rise, right, Like now we get all these stories about like tiny clubs that are using data to source you know, specific players in very moneyball style. Always we obviously have prediction markets. Yeah, people are doing a lot of sports betting, and so all the analytics seem to be like becoming more important. And I will just say, I've read this crazy stat right before we came on from the law firm Morgan Lewis, and they were saying, in the twenty twenty six FIFA World Cup match based data is going to be something like so it's one hundred and four matches generating more than ninety petabytes of data, yeah, which is a forty five fold increase over the volume produced in the last World Cup in twenty twenty two, you know, stunning. 00:03:24 Speaker 2: So this stories is a question, and I know we're going to get to this in the conversation, and it almost is like a philosophical question, which is, okay, we see the game of soccer is like very fluid, right, there's no we just said there's a fewer discrete events, like in theory, is something that's fluid a series of like microscopic discrete events. You know, could you get a million frames per second and actually turn it into discrete events? Like that is sort of like an interesting question to mean. A parallel that I think of in this conversation is chess is like all discrete events, and that it's seen it's like a very difficult thing to crack. But then like over twenty five years ago. 00:04:04 Speaker 3: Now it's like, yeah, of course. 00:04:05 Speaker 2: They cracked it. But then you would think, okay, well, like what is the opposite of chess, which would be language, And you say, okay, you can't. That's fluid, that's all over the place, And yet computers seemed to be understanding language pretty good. 00:04:17 Speaker 3: My inner luod Iite says, there must still be a secret, like unmodel leable you can't talk when it comes to like football, the beautiful game. But you're right that technology may prove me wrong very quickly. 00:04:29 Speaker 2: Well, as you mentioned, you know, soccer analytics growing, and obviously the interest now is for obvious reasons. There's all the betting on the World Cup. People talk about the x G of like I don't know a situation or a player the expected goals. 00:04:44 Speaker 3: Now we have body posing analytics as well, which. 00:04:47 Speaker 2: Yeah, and when our producer Kale introduced me, I hadn't seen them before earlier this year, like the momentum charts that show like you know how dominant a team is in given moment, you see it going waves and stuff like that. So there clearly is a lot of work being done. I don't know how much it works. I don't know how much like momentum charts consistently predict who is going to be the winner. But just this question of like the ability to model deeply fluid things the beautiful game, like it's an art, right, Like can computers actually model this? And then like if they can, what does that say about the ability to model a bunch of other things? Strikes me as just like a very relevant question right now, because it's the World Cup, but also relevant question in general about market for markets and finance exactly. So I'm curious to know like how it works, about how the moneyball revolution, where we are in the arc with soccer anyway, I'm really excited to say we really do have two perfect guests because they do sit right in this space and also at the intersection of everything that we're talking about. We're gonna be speaking with yours Becker's. He is a soccer analytics consultant, the professional soccer Analytics consultant for the decade. As well as Mike Tracy, he's the head of risk at Apex Fintech Solutions. He used to be a volatility arbitrage trader at Peak six and he also does soccer analytics for the club FC. So literally the two perfect guests to talk about some of these questions that are arising right now. So Mike and yours, thank you both so much for coming on the podcast. 00:06:21 Speaker 6: Thank you for having us is a great introduction. Yeah, thank you. 00:06:24 Speaker 2: You know, Mike, let me start with you question. A lot of people I think, have, for good reason an intuitive feel that like sports analytics and trading markets are adjacent ideas that like, we know this, we know that the a lot of the prop shops and market makers do both and so forth. And how would you articulate based on your career the sort of overlap between the skills and techniques you've developed as an arbitrage trader, a volatility trader, and the overlap of skills that they're like, Okay, this thing that we'll get into that we call soccer analytics. 00:07:04 Speaker 6: Well, soccer is unique because I'd say every aspect of the game is a distribution. When you look at on pitch, the performance by a team, when you look at performance by an individual player, when you look at the seasonal outcomes, and this unique component that is promotion relegation. So your finances are highly variant year over year. And so there's a lot of overlap between volatility trading and working in soccer in that you are making highly levered bets on often imperfect information. Then that's not necessarily predictive like other sports such as baseball. 00:07:51 Speaker 3: Can I ask a very basic question, which is what is the point of soccer analytics? So if we go back to the market's analogy, we talk about price disco right, like you're trying to find the right price for a particular asset. With soccer, are you trying to price players, improve the training, make more successful predictive bets? I imagine it's a bunch of different things. 00:08:13 Speaker 6: Yeah, i'd say everything. I think you said what are the things you're trying to do? And I think from an investor operator perspective, it's what are you trying to not do? You're not trying to get relegated and you're not trying to spend thirty million pounds on a player that is going to be terrible and you'll have to get rid of in two years yours? 00:08:34 Speaker 2: What don't you talk about from your perspective? I mentioned your soccer analytics pro actually are a producer. By the way, just put in the chat the Bloomberg style guide is football, which is injury. Maybe we'll just switch to Yeah, let's stay in a company style. You're a football analytics pro tracy as like, what is the point? Is it more on the team side of things in terms of identifying talent, et cetera, or is it more on I don't know, I guess like the sort of the best side. What is the problem you and your professional capacity are trying to. 00:09:04 Speaker 5: Solve, So it highly depends on the organization. Right, some organizations might want to like mic sts, make good hires, make sure that they don't lose millions of pounds. Some organizations might use soccer analytics to try and improve their strategy, but they might look at like in game data and try to figure out if there's some optimization that they can make based on like where the passes should go or where players should be, what play styles make more sense, or what yields more expected goals if you will, And then I guess the third one is entertainment. There are a lot of apps out there, websites out there that provide fans with some sort of information, and even back in the let's say in the nineties, they would have the overlays on the TV where they show number of corners, number of yellow cards, and possession percentage. It doesn't mean anything, but it's it's interesting. So you can only imagine that you can go much deeper with that and people will still be engaged. 00:10:03 Speaker 2: Well, if you just say more on tho you said it doesn't mean anything, Is that something that has always been understood or is this something in twenty twenty six, we could say that a few of these stats that they like to put on the TV were just extremely like low signal data point. 00:10:21 Speaker 5: Yes, so when I say they don't mean anything, I don't think those were necessarily predictive of the outcome of the match. Got corners might because it's an indicator of which team it has the overhand in a game, but I'm quite sure most of the others, like possession percentage, don't mean all that much. I'm trying to look at game outcomes. 00:10:41 Speaker 3: Wait, so we touched on this in the intro. But like the perceived wisdom in the sort of early two thousands was that soccer was far too complicated to be given the baseball moneyball style treatment. What actually changed to get us to this point where we're not just talking about things like expect goals, but we're also talking about like body movements and posturing and things like that. 00:11:05 Speaker 6: You know, baseball had moneyball in two thousand and three, and soccer had two things in the twenty ten. So in twenty thirteen, Chris Anderson and David Sally published this book called The Numbers Game that distilled soccer down to more of a weakest length game. And then the other thing that happened was I would say we had a bit of this revolution on Twitter, which you kind of spoke about recently, Joe, where ideas were incubated and it I give Michael Kayley a lot of credit for this, He's he's been on the pod. But yeah, every Saturday in the Premier League you would have six matches and then you would wait thirty minutes and Kaylee would just post a stream of every expected goals chart and it started to stimulate this discussion of does this represent what should have happened, what should have been the outcome? I think it started to lead us down this path of where is x G flawed? It only registers when you have a shot, and so there is much more to the game than simply which shots occur. There's possession and the threat of each possession. There's match momentum, and there's changing styles. So it's evolved from on ball data to tracking data to now we're going to that deeper layer of body pods and what movements can prevent or create opportunities to score. 00:12:38 Speaker 5: I think what goes heading out with that is when I started ten years ago, soccer football was perceived as two complex There were twenty two players, there was a ball. People are doing interesting analytics research and applied research already in basketball, and what was said as well, there's five players on each team, so it's a lot less complex. There's more games, and we have more data, and there's more scoring, so you could derive metrics a lot easier. 00:13:08 Speaker 6: And I think the. 00:13:10 Speaker 5: Evolution in just how AI was applied in data availability. So going from like Mike said, only on ball events where we know which player made the pass, which player made the shot, but we always have to say, well, we don't know where all the other players are because those are not recorded. And now we have the ability to do make highly complicated artificial intelligent neural nets with this positional tracking data where at ten trains per second or twenty five points per second, we know where all the players are, and we know where the ball is. And then additionally, at some point we will get the full body posts we know where like the whole skeleton of all players are. Basically like that sort of went hand in hand and it felt like it was inevitable. But also the game is just more complex, so it needed more compute, it needed more data, needed more knowledge. 00:14:13 Speaker 2: Tracy Mike mentioned that you know, just looking at a shot on goals, so again to tell you so much. For example, it doesn't tell you if the referees are going to revisit a call made a minute before the shot halfway out on the other half of the screen and take away the goal. 00:14:29 Speaker 3: This was going to be my question, which we were talking about before the podcast recording, But like, how do you factor in let's say, the occasional randomness of the game and perhaps some erratic trying to be diplomatic here, erratic decision making by referees and footballing bodies. 00:14:48 Speaker 6: I think it's control the controllables. That's a fixed parameter within the game. Is that uncertainty that you can't control and it's hard to predict. There is human error, obviously, and it's impossible to isolate when that's going to happen in a match, and how it will you just sort of play the game. 00:15:10 Speaker 5: You also don't look at individual events per city. You might from an analysis perspective, but if you're building a basic predictive model for football outcomes, you might just take all the scoreboard results from the last ten years and then all those things cancel out, Like if one team has a red guard somewhere you wouldn't even know from these models, but you can build some interesting, let's say rudimentary predictive models with just the outcomes. 00:15:36 Speaker 3: Of games out of curiosity. Has var changed football analytics at all or presented new opportunities for data. 00:15:44 Speaker 6: It presents new opportunities for data because that is an event that happens. If a player is fractionally offside because his hand is ahead of the last defender. That doesn't take away from the fact that he still got into a good opportunity to position, possessed the ball, turned and struck the ball in the net. And often that data is nullified because of the var But in theory, it's something that you should consider in your data set. Within the raw data itself, there's tons of data that that you could add or sensor out. You know, yours mentioned red cards, and a lot of our models we censor that out of our data set. There's a there's an infamous match from three years ago between Chelsea and Tottenham where Tottenham went down to nine men and their response was to play a very high line, meaning they put all of their defenders up near midfield and tried to catch Chelsea off sides, and as one would expect, Chelsea proceeded to score or multiple goals later in the match. And one player in particular, Nicholas Jackson, had three goals in that one match, and when you look at his seasonal outputs for that entire season, that three goals was probably around twenty percent of his total goals. So when something like that happens, when you have an irregular game state, it is wise to sort of manipulate your data and remove that to give you a full picture of what does this game look like at an equal game state. 00:17:35 Speaker 2: Oh, that's interesting. So it's not like that game in particular was a rich fountain of data. It's important to sort of like recognize that the data from this particular game is not going to be particularly predictive about other games. And therefore, you a guy who scores three goals in that match is probably not going to continue to score three goals the game for the rest of it. 00:17:59 Speaker 6: There's a rich fountain of data. Okay, So at the end of the season, when you see people analyzing the player, they often analyze their season and what do they do per ninety minutes of football, And so you have this highly skewed data set by some really poor data. Yeah, and so there's a lot of data mining involved in the process of building a model when you evaluate a team and individual players. 00:18:29 Speaker 2: So here's a question I have, and you're talking about using neural networks, and it's like, eventually, like you know, we'll have the compute to like have the position of every player's body flash maybe you know, hundreds of images, frames per second and so forth, and then you have feeded all into a model, and then we learned something about who's more likely to win. But one of the things that happens in a lot of other domains, and here I'm thinking about like chess or go, for example, is that you can have these models that are extraordinary, but they don't speak English, or they don't speak any human language, and so the transmission of like what was learned from these events is something usable by say a coach who's thinking about strategy or a general manager who's thinking about player selection. Talk to us about, like, when you think about these machine learning models that can't really communicate their findings in any way in language that humans can understand, how you sort of bridge that gap to where this is useful information for a team or a manager. 00:19:34 Speaker 5: This is generally be understood I think in the last couple of years or the last eight years, as the biggest promm you can build these models, these neural nets already exist. The main thing is the translation, like you said, from model outputs to coach, and I think in the last couple of years most most teams have found is that they need an expert analyst and data analysts to do this conversion or the translations, where the coach isn't fed the data directly. Sure the coach is fed just the information that the analyst finds from the data, and that could be in the end. That generally still boils down to having video clips, so you can use the data to find video, and then you can show the video to the coach, which then from bottom up you have this. 00:20:19 Speaker 2: Approach that makes sense. But the part about okay, we're going to use the data to find video. So all of this makes sense. You have the data special ist who translates, you have the video so that there's something tangible. But talk to us about that specific step where the data analyst sees some sort of model output and then is able to use that to find the relevant clip. To show the coach something potentially instructive, because that seems like the hard part to me. 00:20:45 Speaker 5: Yeah, So the model output could be many things. It could be outputs from a classification model that says, in this given ten second or fifteen or twenty second window, this team played in this sort of build up, or they have this type of structure. It could also be a little bit more advanced where you can simply say, well, in this instance, we had a high probability of conceding a goal. And that could be just from an expected goal shot if you're looking at only a cent data. But you could also have model outputs from something we call an expected possession value model or an expected tread model, where you measure the chance that the teams going to score and let's say the next thirty seconds or the next possession, and then you can find this fight there and either measure when your team is likely to concede or measure when the other team is likely to score, and you can use those kind of signals to boil it down to video. 00:21:33 Speaker 3: Yeah, this is something I wanted to ask. Actually, so you mentioned speed just then, like, what is the actual latency that we're talking about since we're using all these market Yeah, jeez, Like, are we talking about an insight that's actionable within seconds, like you're going to sub a player on after your model splits something out like live during the game. Or is it more realistically that you're reviewing the model and the analytics a d a game and sort of tweaking. I guess when things have calmed down. 00:22:05 Speaker 5: Apparently most of this does not happen live, So most of it happens either pre match or postmatch. 00:22:11 Speaker 6: But you can still do it. 00:22:12 Speaker 5: There's there's enough live data to make these instances the worthwhile. 00:22:16 Speaker 6: Yeah, I'd say most those types of adjustments based on data typically happen at halftime or if you're in the World Cup, during a higher dration break. And in the derivatives world, I always think of expected outcome versus realized outcome. So what is the market implying? What do you expect? And then what is actually happening? And that is sort of the nexus of how teams prepare for matches. They come up with their own expected outcome, How is our opposition going to play? How are we going to play? And then that live in game is your realized outcome? And so you're receiving that data and it's being transmitted to analysts who can then communicate it down to the bench to discuss with the manager, who can then make those changes at that halftime. 00:23:10 Speaker 3: It's interesting the game of two halves, Joe's now the game of four quarters, offering up more opportunities to make model based adjustments. 00:23:18 Speaker 4: That's right, we thought about it. 00:23:19 Speaker 2: We have We used to have two discrete events in a game, the first half of a second, and now we have at least four. So I guess that, yeah, that that creates more data as well as maybe more adjustment as well as more ad revenue. You know, obviously in baseball at least according to Michael lewis right that there was this period where you know, you had the old time scouts and they're like, oh, this guy has good hustle, right, or this guy has a good heart. There was like no data behind any of it. He may just sort of had the you know, the swagger of someone who looked like maybe a star player, and then it's like, oh no, but he look at his like on base person, that's his vorp or whatever, and then they get promoted. But it was like a culture thing has there been a similar cultural clash within sort of soccer scouting, where what the data says about what constitutes a player does not map to traditional intuitions. 00:24:18 Speaker 6: I think a lot of clubs look at it from two perspectives. I don't think there's this old school scouts versus the data guy of mentality. I think it's a very collaborative. I think what a lot you know, what a lot of clubs do is they have an individual scout who will go watch a player and give his assessment in his rating, and then they will have their own internal model with data, and they will look at the delta's between those two different models, and if something seems off, you often have collaboration between the data person and the scout and they figure out who's right, who's wrong. And then to your point about hustle, let's call it. Yeah, there are Swiss army knife ways to quantify that in soccer in certain instances, if you look at game state, say the game state. You know when game state is essentially zero plus one minus one plus two minus two plus three minus three, and so say you're in a plus three minus three game state, and win probability for one team is ninety eight percent. You can manipulate the data and only look at how players are performing in that game. State you're out of the game, are you still competing? Do you still care? Are you in the right position now? That may not align with the old school scout saying, you know, this guy has grit, But again there's there's Swiss army knife ways to give some type of indication, you know, to say, soccer data creates questions, It doesn't give us answers. 00:26:04 Speaker 3: Wait, just to better understand this, can you give us like an analytics framework if you were trying to judge the best This is the loaded question, but it comes up on every discussion. But you're trying to judge the best soccer player, either of all time or currently, Like, what would the analytics framework for that actually look like? 00:26:22 Speaker 6: Jude Bellingham? 00:26:24 Speaker 2: Okay, yeah, yeah, Like really seriously, this is a great question, Like, walk us through what the math says about Jude Bellingham and how you would derive that. 00:26:32 Speaker 6: Well, Jude Bellingham can play four or five different positions. Right, most players they have the number nine and their role is number nine. But if you think of the game as this book with different chapters within the story, Jude Bellingham can perform whatever task he needs to perform at every single chapter throughout the book, and he does it at the highest level. He could play any position on the pitch besides goalkeeper. 00:27:04 Speaker 2: And I just started like, how is this established? Like someone could say, oh, this guy, they say this is baseball too, he's a good all around player. Whatever we can see, But like, what is the data that actually establishes that Jude Bellingham can play at high levels in a wide range of position. How do you derive that conclusion quantitatively? 00:27:26 Speaker 6: So from a data lens, we think of it in two fronts, in possession and out of possession. Okay, So in possession is how you're progressing the ball into threatening areas and obviously creating high probability opportunities. Then out of possession is how are you preventing a team from moving the ball into threatening areas. Since it's a very tricky question sports, because you're trying to quantify the value of an event that does not happen. So let's say Joe, you have the ball out on the wing, Tracy, you're right in front of the box. Jude Bellingham, he would be both. He was always moving into that passing lane is right in between you guys at the right moment, and we have the ability to quantify the value of those movements and the closure of these lanes as players move out on the pitch, and then you'll obviously get to the next phase where it's how do they do this with their feet? 00:28:41 Speaker 2: We started talking about this sort of translation from what the model says to a coach or something like that, and that still seem as tricky for better someone who's betting on sports, that might be totally irrelevant, like they're just like here, the model says this, this team is better. The line doesn't match up with this. They're for going to bet on this team. I don't know why them a model says, My model says this team is better, but it does. So the translation is unnecessary if you're just betting, like are we at the stage or are we getting close to the stage where you could, for example, feed a model the first five minutes of a game scoreless zero zero, and to all the players like, oh, this looks this looks like a competitive match. It's going to be good both sides. But models are able to detect something that we can't articulate that says, oh no, this team is playing and even though it looks like a tie game and a competitive one, actually for reasons that we can't put into english, put into language, this team looks like they're going to win the game seven. They have a seventy percent chance. Is that a thing or is that a phenomenon or is that a realistic thing to expect? 00:29:43 Speaker 5: And there are game with probability models, yeah, which start with just a pre game cheap strength, so both teamch have some value. Perhaps you can think of an ELO rating, okay, and they boiled onto a win, a draw, and a lost percentage, and then those percentages can in game be updated, but they won't swing all that much because you have this prior information. So maybe after the first minute, one team has some egg, or maybe after the first five minutes, let's say some team has created some high probability chances and the team might be the underdog team. This was highly unexpected, I guess since since the ELW rating set that the other team would would be he the over end. So you can you can slightly update your beliefs there. I don't think it's like changing the needle or moving the needle all that much. Given that it's just five minutes of information, but you can definitely update your beliefs throughout the game given chances created, expected possession value as we just talked about, or momentum or something in between, depending on the data you have available to you. 00:30:46 Speaker 3: Mike, I wanted to ask you, given your involvement with Austin FC, we know that there are obviously differences between Major League Soccer and European leagues, and some of those are, you know, things like salary caps, designated players, roster rules, lack of relegation, lack of Yeah, that's a big one. Does that actually make soccer analytics and MLS like a little bit cleaner in some ways? In the sense that, like in the European leagues, the money is again trying to be diplomatic here, but it's very free flowing. There's a bit of rule stretching going on at times when it comes to salary restrictions and things like that. Like compare MLS analytics versus European analytics for US. 00:31:32 Speaker 6: So, in European analytics you have essentially your recruitment and then first team analysis. In MLS you have recruitment first team analysis, but you have this third vector that I call portfolio management. Right, this cap structure in the MLS is put in place with the intention of creating parity, and as you mentioned, it's highly complicated. The best way for you to understand it would be like me saying, Tracy, I'm going to give you ten million dollars. You can spend two million dollars on Navidia, seven million dollars on Walmart, and one million dollars on a speculative biotech stock. And so you have to think of each player from a relative value perspective based on where they slot in your cap structure. 00:32:26 Speaker 4: Has that worked? 00:32:27 Speaker 2: I mean with MLS. So it's like, I understand, you have this new league that you know, you don't want some really rich team to win all the time, and then I don't know, only Miami or wins, and then fans lose interest in the rest of the country. I'm just I don't know what the actual I think my understanding is Austin in particular is a very big fan base. But has that worked in practice to sort of create an equal level of like fan affinity that's geographically distributed. 00:32:54 Speaker 6: Well at Austin We're we're eling nine months into our project here, and we just had some turn and over with our sporting department, and so I would say that we have yet to integrate and prove this portfolio management theory and the impact on success in points in the table. But as a market participant, when I read these rules, I like, my brain immediately goes to the markets and portfolio allocation. 00:33:25 Speaker 3: Okay, so we have to watch Austin FC as a test case for the portfolio management thesis in football. 00:33:31 Speaker 2: Boy, Just to be clear, so this portfolio management thesis, would you obviously have a strong intuition for having traded volatility for a long time. This is the framework that you're bringing to your work in Austin. 00:33:46 Speaker 6: Yeah. Correct. Every player has a designated slot. So a player that could come in as a DP could be terrible. But if he's going to. 00:33:58 Speaker 4: Do oh yeah, okay. 00:34:02 Speaker 6: But that same player, if he is going to be a senior minimum salary player, could be in the top percentile of talent for that specific slot. And then within the league itself, you know, you have different roster construction strategies. You can either have three designated players or you could opt for two designated players and for U twenty two players and The way the league works is they have a They have a salary cap, but it's more of a salary cap charge. So whilst Lionel Messi may be making over twenty million dollars in salary, his salary cap charge as a designated player is going to be much lower than that. It could be seven hundred and fifty thousand. 00:34:57 Speaker 2: Wait, I don't understand that. 00:34:58 Speaker 6: Yeah, it's confusing. There's a salary cap charge based on these roster designations, so every slot has a dollar charge associated with it. So you have to work within that framework of what is the charge for each player and your finite amount of kapala you're allowed to spet. 00:35:19 Speaker 3: So the designated players are the ones you're allowed to pay more. 00:35:23 Speaker 5: Got it? 00:35:24 Speaker 3: Because this is the legacy of like La Galaxy getting David Beckham and wanting to pay him right, lots and lots and lots and lots of money. 00:35:32 Speaker 4: This is helpful, Okay. 00:35:33 Speaker 3: So one of the criticisms of modern football, I guess, is that it's dominated by the wealthiest clubs, like whoever has the most money can buy the best players, certainly in Europe, and so the big just kind of get bigger, And I could certainly see an argument where if analytics becomes more important to actually playing the game, then whoever has the most resources, the most compute, the most engineers I guess nowadays is going to have an edge here. But on the other hand, we have had technology before that has a sort of democratizing effect. 00:36:09 Speaker 1: Yeah, we have open source. 00:36:10 Speaker 3: Models, all of that, and I think there have been some instances of the smaller clubs actually like using this technology to perform better. But which way are we going to go? Is it the big get bigger or maybe we see some smaller clubs level the playing field here. 00:36:26 Speaker 6: I hope it's the latter. 00:36:27 Speaker 5: Yeah, same, So you mentioned open source models actually build open source software to help these smaller clubs. I guess it's also helping the bigger clubs, but it allows them to build these crossnural nets and these expected possession telling models, and then also load these tracting data districting data, which has been a big task just in general because of the data structure and the data size, and the software that I that I work on helps clubs any club basically get get started. So I hope by doing that it will help the smaller clubs, help data providers I guess provide better insights to level the playing fields. But I'm not sure. I'm not sure that it is working out necessarily because getting the knowledge, like we talked about before, actually I'm distilling the knowledge from the data. I still I think. 00:37:19 Speaker 6: One of the problem. 00:37:20 Speaker 3: Next. Oh yeah, so this reminds me. I wanted to ask as well, who actually owns the data here? Where's the data come from? 00:37:26 Speaker 5: It depends, okay, and I think some of it is a gray area. There are a lot of data providers, some have licenses, some don't. Yeah, it's a big question market in some instances. 00:37:37 Speaker 6: I'll tell you a story about when I first started building models for AFC Bourne, myth back in the Premier League in twenty sixteen. I had no idea where I could find the data, and so I went on upwork and I posted an ad and I just said I need a developer who has worked with European football betters, and a gentleman named Dmitri and Ukraine replied to me and he said, yeah, I've worked with many professional betters. And I said, give me all the data that you can find. And so it was very skunk works, but Dmitri was able to find a way for me to get access to all of the on ball data across the world at the time to start building my models. 00:38:29 Speaker 2: What is the data? So it's like, okay, you find some data provider or Dmitri and Ukraine collect data. Is this numerical data? Is this a series of frames? Like what can you when you say okay, you need to go out and get the data. What form are you getting it? 00:38:46 Speaker 1: What does that mean? 00:38:47 Speaker 6: Like? 00:38:47 Speaker 2: What does it look like? 00:38:48 Speaker 6: So it's varied throughout the years, but the most prevalent generic data out there is on ball data. And on ball data if you're just put on your Excel hat and think about what it looks like in an Excel spreadsheet, has every single event tagged, so you have it just look goes down a series of events past past, past drible shot, goal, past past drible tackle. It has the event, it has the player or players involved, It has the X Y coordinate on a pitch, and it has the time. And so when you have these raw data points, you can conditionally put together a mosaic of what is happening on the pitch because you can measure the events, the speed, and the location. Tracking data is a bit more nuanced. I'll let yours touch on that. 00:39:55 Speaker 5: So you can still imagine it an Excel spreadsheet, but I don't think your Excel spreadsheet would like it very much if you try to load in this data. It's basically a player identify a team identifier, and then X and y coordinates for all players at ten or twenty five frames per second, so that means I guess twenty five rows for a single frame. We never really touch, let's say the raw data in the sense that we don't get the pictures of the game and then try to figure out ourselves where the what the coordinates are or what the coordinates are. So there are a lot of data providers out there that either put cameras in the stadium or use the broadcast footage to extract this data, so they will have different models. One of the models would be first identify where all the pitch markings are, so they have a way to understand like where all the players are relative to the pitch markings. Then they know where all the players are and you can convert all of that into coordinates. They know where the goals are obviously, and then you'll get that in a single file for a single game. Some providers might give you one file per minute, they might give you one file per half, and then that's just the raw data with the identifiers in the cord that you might get an additional file that has all the metadata that says this identified blongs to this player, this identified blongs to this team. If you're lucky and you're tracking data, you get an identify that says which team is actually on the ball, because that's highly relevant, but sometimes it's not included and you have to figure it out yourself, like calculating conditions to the ball for each player. And then obviously you have skeletal data, which is I guess twenty seven times more dense, or maybe even more, because you have twenty seven coordinates, one for each body post point or body points, so you might have one for your left shoulder, for the tip of your nose, for your left ear, for your right ear, for your right foot, for your ankle, and that gets into the millions and millions of data points. So when you talked about in the introduction about these petabytes of data, I assume it's going to be mostly skeletal data because that data is incredibly rich. 00:41:53 Speaker 3: How do players actually feel about all of those because you know, if people were watching me do my job and monitoring like my next movie, they are all right, but no one's modeling like what I'm doing with my hands or my feet at all hours of this particular recording, Like I would have mixed feelings about it, right, Like, do you have any color on how players actually feel about I guess the rise of a statistical analysis in football. 00:42:23 Speaker 6: I think they find it useful. I think something that yours has worked on this specifically is is eyesight. What what can they see and what they cannot see? And so when you look at the data and you think about opportunity costs decision making, and you make a suboptimal decision, when you go and talk to the player, they might just simply tell you I couldn't see it, and then you move on to you move on to the next. So it's it's very complex. I find most players to embrace the data. There's curiosity around it, but you know, they likewise know the limitations. 00:43:02 Speaker 2: You know, we started to talk about o soccer's fluid. It's a beautiful game, it's an art. I totally agree, but like people actually did say this about chess thirty or forty years ago, and there was actually some people who held out the belief computers will never be able to beat humans at chess because of it as an art. And it almost seems hilarious that like this view held on for as long as it did. But it turns out that no, chess is just a calculation problem, and when you have enough data to compute, you could solve the game. Kind of is soccer in the end? Like, is it just a series of lots and lots of discrete events that our eyes are not capable of, But at the end of the day, with sufficient compute and data collection, is it just like chess? And that it's just a lot of micro binary decisions that can all be summed up. And is this therefore how life is? 00:43:54 Speaker 5: When you work with this data as long as I have, I mean with this tracking data specifically, in the beginning, it seems very overwhelming because it's just an infinite stream of coordinates, and so I strove at that a little bit. In the beginning, I was wondering what I should do with this? How are you going to model any of this? But I totally agree with you. You can discrictize this so you can discertise it into let's say, very minor events where we use the on ball event data, and if we align that with the positional tracking data, we might know every single moment where a player makes you pass, and then you might know the next moment where a player makes a reception, so that could be discretized into one event that might last two and a half seconds or three seconds. The next step would then be the player receives the ball, they make an onball action that also lasts two and a half seconds. That would then be your next discratized event. And then you have all these small events which are micro movements by players counter movements by defenders. And if you go these let's say sequences one at a time, you can still make aggregated metrics from this using all the tracking data that you have at your disposal, but it's actually still understandable for you. So you might say, well, this player made this many dribbles and it gained the team this much in terms of added value to scoring a goal. And you can do the reverse for defenders, where you can say, well, this defender was always close to the ball, so he was helping not concede a goal. And if you go back to the analysis part where you want your video analysts to look at this. They also do this just discriptized. In discriptizing it like this will help significantly. 00:45:33 Speaker 3: So I have one more question, which is it is obviously World Cup season, which means it is also a cell side analysts publishing World Cup prediction notes and research season. And in my experience, they tend not to be very good. Like often they will publish that England or the USA are going to win whatever World Cup and. 00:45:54 Speaker 2: As the numerista added, just predictor. 00:45:56 Speaker 3: Pan that not to my knowledge, but you know, they often get it wrong. If we think about statistical modeling data, I mean, the cell side firms they should be pretty good at this, and yet do you have a take on why they seem to struggle with soccer predictions every four years? 00:46:15 Speaker 6: I like, honestly, for them, I have no idea what data they're using, whether they're using an Elo model, whether Nomura has on ball event data tracking data. There's different ways that you can come up with these predictions. But I think, you know, I like to stay stay in your lane, and I think cell side analysts should stick to cell side analyzing. 00:46:40 Speaker 5: Simply. The age Old's problem. If you have a good model, you wouldn't publish it. You would just beat the bookies, right. 00:46:47 Speaker 2: Yeah, there's often criticism of people who publish things for a living. Mike and yours, thank you so much for coming on odd Laws. Learned a lot there and I really appreciate your time and enjoy the rest of the World Cup. 00:46:57 Speaker 4: Thank you, Thank you, Tracy. 00:47:11 Speaker 2: I have to admit I find it a little depressing that probably most things in life are probably just computation. 00:47:18 Speaker 6: Thanks. 00:47:18 Speaker 1: You know what I'm saying. 00:47:19 Speaker 2: It's like a lot of thing that like that there is something called art and beauty and intuition and something. It's probably just computers all the way down. Binary events that can be chunked and analyzed by microchips. 00:47:30 Speaker 3: Oh you know the question I would have asked, Yeah, but we ran out of time was the idea of like good Art's law, which is like once you have a measure, like the measure. Yeah, because you see this criticism of sports analytics, coaches like players start focusing on their stats. Coaches are focusing on the player stats, and then they get the players with the good stats, and then they just focus on improving their stats more. But the stats don't necessarily translate into like wins all the time. 00:47:58 Speaker 1: Yeah, you know. 00:47:59 Speaker 2: I was thinking something that Mike said in the beginning, which is that like a particularly like an English Premier League team, there's multiple things they could be optimizing for. So they could be optimizing for profit, it could be optimizing for avoiding relegation, and they could be optimizing for avoiding relegation. They could be optimizing for wins. Those are all distinct things. But you see this when a sport gets over optimized. It didn't come up with a good example. Is like basketball. When I was younger, it was like the game was fun because there were lots of slam dunks, and then everyone realized that three point attempts were not being taken enough. So suddenly the game is dominated by three pointers, which may be a better way to play, but it's not necessarily a more fun fan experience than watching a dunk. So you think, like, okay, could the game get better formally, but it becomes less entertaining. There's a lot of criticism. 00:48:49 Speaker 3: I think that's the possibility because people are already talking about convergence and like actual play style. 00:48:55 Speaker 2: Totally, and you see this, like you know, there was a lot of criticism in like, people are very critical of how Paraguay played right, like play to just survive into penalty kicks, then hope that the variance of the penalty gap period allows you to beat France. But it's like, no, that's like the game theory optimal or just the game optimal play if you're considered to be the weaker team. So it does feel like there's all different things, you know. Again, going just to the core of the question, stat's like, what are you solving for? Solving for winning is very different from solving for profit, is very different from solving for a gambler, And there are different answers to each one. 00:49:32 Speaker 3: Yeah, you win, but no one's paying for it or happy about it. 00:49:35 Speaker 2: Plausible, although for now people are paying crazy amounts of money still to go see a soccer games. 00:49:41 Speaker 3: So all right, shall we leave it there. 00:49:42 Speaker 5: Let's leave it there. 00:49:43 Speaker 3: This has been another episode of the All Thoughts podcast. I'm Tracy Alloway. You can follow me at Tracy Alloway. 00:49:48 Speaker 2: And I'm Joe Wisenthal. You could follow me at the Stalwart. Follow our producers Carmen and Rodriguez at Carman armand dash Ol Bennett at Dashbock, Kilbrooks at Kilbrooks and Kevin Lozano at Kevin Lloyd Lozano. From our Odd Lots content, go to Bloomberg dot com slash odd Lots or the daily newsletter and all of our episodes and you can shut about all of these topics twenty four to seven in our discord Discord dot gg slash odlines. 00:50:11 Speaker 3: And if you enjoy Oddlots, if you like it when we talk about stochastic soccer modeling, then please leave us a positive review on your favorite podcast platform. And remember, if you are a Bloomberg subscriber, you can listen to all of our episodes absolutely ad free. All you need to do is find the Bloomberg channel on Apple Podcasts and follow the instructions there. Thanks for listening 00:51:00 Speaker 4: And