How do organizations like ESPN and NFL analytics assign winning odds to games? How do Sports Analysts determine who is the best player on the court or field? In this episode, we sit down with Dr. Paul Sabin, a Senior Fellow in the Wharton Sports Analytics and Business Initiative and a Lecturer in the Statistics and Data Science Department. Dr. Sabin discusses his professional experience at ESPN, SummerSports, and Sabin Analytics, focusing on the invaluable player and team metrics he helped create which are now used in tournaments like March Madness. Additionally, Dr. Sabin dives into his research on Bayesian-driven models in soccer and American football, explaining how he approaches common sports analytics problems from the viewpoint of a statistician. The episode also reflects on the future of metrics in sports with additional player tracking data in an increasingly data-driven sports world.
00:00 Dr. Paul Sabin
The research that we've been doing is looking at expected goals and using player tracking data, via a data provider that we have a partnership with, that allows us to look at where every player is at all moments on the pitch, and not just when there's a shot, but we want to now calculate the probability that there is a shot.
00:25 Sheyan Lalmohammed
Hey everyone, this is Sheyan and you're listening to “Ever Thought About…?”, created by undergraduates at the University of Pennsylvania. We hope to bring you exciting episodes about the diverse research undertaken around campus. Sit down with us as we chat with Penn faculty about their research careers, academic journeys, and work they've dedicated their lives to. In this episode, we're incredibly fortunate to be joined by Dr. Paul Sabin from the Department of Statistics and Data Science at the Wharton School, where he serves as a lecturer and senior fellow in the Wharton Sports Analytics and Business Initiative. His work focuses on Bayesian statistics, machine learning, and sports analytics, and he teaches courses across the undergraduate, MBA, master's, and PhD levels. Dr. Sabin’s research has been published in the Journal of Quantitative Analysis in Sports, and he is an active contributor to the broader sports analytics community. In addition to his academic work, he has held several prominent industry roles, including Vice President of Football Analytics at SumerSports and Senior Sports Data Scientist at ESPN, where he developed advanced player evaluation metrics and predictive models across professional and collegiate sports. He's also the president of Sabin Analytics, consulting with professional teams and athletic departments on data-driven strategy and decision-making. He received his PhD in statistics from Virginia Tech. Welcome, Dr. Sabin.
01:39 Dr. Sabin
Thanks for having me.
01:41 Sheyan
Okay, so for the first part of this, we really just want to trace your own academic journey and how you came to be in the position that you are today. What originally drew you to apply high-level statistical modeling to the world of sports, and how has the industry changed since you first entered it?
01:56 Dr. Sabin
I was always interested in sports. As a young kid, I used to watch ESPN every day, every morning. It would be on for an hour, SportsCenter, and they’d go through the highlights and my mom would walk by an hour later and ask me if I had already seen what I was just watching and the answer was usually yes. And I was always attracted to sort of the box score statistics that they would flash at the end of each of the highlights. And when I was trying to figure out when I was an undergrad what I was going to study, I kept drawing back to the fact that I always kind of loved numbers and I always thought about sports from a sort of numbers perspective, even as a young child. So Statistics was always really interesting to me because it was applicable across many different disciplines. Now, I really didn't believe that I could get a career working in sports until it actually happened. In fact, I met with Ted Leonsis in 2013. I was in the summer between my first and second year of my master's degree. And if you don't know who Ted Leonsis is, he is the principal owner for the Washington Wizards and the Washington Capitals, so in the NBA and NHL. And he told me, and I had a family connection, which is how he was nice enough to talk to me. He told me that I should never work in sports, I wouldn't make any money, and I was much better off working in a finance role or something where quantitative analysis was more widely accepted and paid well. And it wasn't until I got the offer to work at ESPN that I really believed I could make a career out of it, and I've been fortunate to really do that ever since.
03:35 Sheyan
That's amazing. Thank you for that explanation. Of course, a significant portion of your career was at ESPN. During your time at ESPN, you were responsible for developing very influential metrics in sports, such as the Basketball Power Index and the College Football Playoff Predictor. What was the biggest challenge in creating a model that was both statistically rigorous and digestible for a mass audience like ESPN has?
03:58 Dr. Sabin
The hard thing about working at ESPN and doing sports analytics there, even compared to someone who works for a team, is that our audience was everybody. When you work for a team, you are mostly trying to do things to help your team win, and your number one goal is to convince the decision makers that you're working with that what you are doing will be helpful in the broader goal, which is, for an NFL team to win the Super Bowl or, to win the NBA championship, whatever sport you're in. At ESPN, our purpose was to, ESPN's motto was to serve sports fans anytime, anywhere. And we do that by telling stories. Well, in order to do that, we would want to tell stories. We wanted to do the complicated things that teams were doing in the NBA and in the NFL and have similar levels of rigor, but bring it to the masses. Well, the problem is the masses includes “Joe” from Alabama, who doesn't love numbers and hates math and just believes in the sheer willpower and physicality of football, to someone else who is a professional gambler and very critical of all models that have to do with sports because that's their livelihood. And we would get criticized from both ends. And typically that meant I felt like we were doing a good job of doing something that was sophisticated enough that we'd get some feedback or even criticism from a professional gambler, because we weren't trying to gamble. And also people yelling at us because they hated analytics in sports, and it was all about the heart and all about the intangibles. And so that was always the hardest thing, was making everyone happy, which is an impossible task. So ultimately, we would always just try to educate people, provide useful information, to help them feel like they were learning about sports. And now on ESPN, you hear terms like EPA, which stands for Expected Points Added, Mina Kimes and other journalists will use that term just as they're speaking. And that's something that we were educating talent and people that were on air at ESPN 10 years before any of them actually started using it in their analysis.
06:11 Sheyan
Yeah, that really is interesting. But obviously, sports analytics for you goes beyond the approach of a media-based company. You have experience working with sports teams specifically using analytics. At SumerSports, you transition from media-facing analytics to advising NFL organizations on roster construction and game strategy. How does the objective of your work as a statistician change when you're optimizing for a team's wins versus a media company's engagement or accuracy?
06:36 Dr. Sabin
Honestly, the main difference is that at ESPN, our models were used with broad strokes, and we were trying to apply it to every game for every team. And if you are working for a specific team, you are trying to help your specific team win. And one of the big differences is that you really are able to have a lot more information when you work for a team just about your team. But you don't have that same level of information potentially about players on other teams. What I'm talking about is like you have really detailed information about the injuries, how much the players are practicing and training and every little minutia that you want for the players on your team, you might know a lot about. And so, but then you're in this data hole about players from other teams. And teams themselves, they incorporate human elements, scouting. And that was, that's typically one of the big differences. At ESPN, we didn't typically use scouting because scouts work for the team. And so incorporating sort of like the coach and the scouting department’s, you know, philosophy and trying to make what they do something that is digestible, like turn it into data that then can be used and added to your models. That's the challenge. And good teams realize that, like a true Bayesian would, they would try to combine subjective information with the objective data. At ESPN, we were still Bayesians, but we didn't have sort of that same level of scouting department per se that an NFL team would have, or an NBA team or a Major League Baseball team.
08:12 Sheyan
That is very cool, very different sort of approach that you had there. Now I want to switch gears to some of your experience in research. You've done research in both an industry and academic capacity. I want to first focus on your time here at Penn. As a senior fellow at the Wharton Sports Analytics and Business Initiative, what are the primary research questions you and your students are currently tackling?
08:35 Dr. Sabin
What a lot of my research recently has been about is in the world of soccer – football, international football. And one of the more recent papers that I've worked on with students was around the use of “expected goals”. And expected goals is this metric that has been in soccer now for a while, and you might see it. It's even in the FIFA game, right? If you're playing at halftime, my son, when we play the video game FIFA, you know, it'll flash up the expected goals at halftime of how many expected goals, “xG” as it's abbreviated, for each player. And what an expected goal is, right? It's the probability of converting a shot into a goal. Every time there's a shot, what's the probability? We build a model. And people in soccer have been doing this for a long time because goals don't happen that often, but shots do. And so they sum up essentially the expected goals across the game. And that's a better sort of estimate of which team played better. Now, the research that we've been doing is looking at expected goals and using player tracking data, via a data provider that we have a partnership with, that allows us to look at where every player is at all moments on the pitch, and not just when there's a shot, but we want to now calculate the probability that there is a shot. Expected goals have been used just assuming that shots happen and no one's really modeled the probability that the shot actually occurs. And one of the interesting things in soccer analytics over the past decade or so is that people that work in soccer analytics have shown that if a player scores more goals than their expected goals would imply – we call that “goals over expected” – that is not something that is very repeatable, meaning if one player scores more than expected in a season, next season they're likely to regress to 0 goals above expected. And very few players in the world, I'm talking maybe a handful, have ever shown to be consistently able to perform better than expected goals, you know, Messi being the prime example. Now what we are doing is saying, okay, well, expected goals is only looking at shots. When you think about a statistical probability equation, it's really the probability of a goal conditional on a shot occurring. Well, we've been trying to calculate the actual probability that a shot occurs. So the true probability of a goal is actually the probability of a goal given a shot times the probability of a shot occurring. And we've shown that in doing so, that creating shots is a skill that is repeatable. So players like Erling Haaland, who is one of the best players in the world, he set the record in the Premier League recently, a couple years ago, for the most goals scored in a season. And he did that because he's incredible at creating shots when most players wouldn't be able to create a shot in that opportunity. And that year, he just happened to have a positive goals over expected, right, a lucky year. But he also is very good at shots over expected. So the combination of, he had a lucky year of converting shots into goals, but he's also very good at shooting and getting shots off. Our research is really focused on sort of correcting that bias that we're assuming that all players can create shots equally, and we know that's not true, and that's something that we have shown. And so, you know, if I owned a soccer team, I would pay much more attention to the players that are good at creating shots and less about making the shots that they create.
12:02 Sheyan
That's very interesting. I'm very interested in soccer, and that's something new that I've never heard before. And it seems like the work that you and your team are doing is very impactful. I want to focus on some of your academic work. Recently, you wrote a paper on hierarchical Bayesian plus-minus models for American football, which addresses one of the hardest problems in the field, which is finding the value of an individual player in a highly interdependent sport. What were the most significant sources you identified that traditional stats often miss? And how did you approach the problem?
12:33 Dr. Sabin
Yeah, so I mean, adjusted plus-minus models in sports have long been used most traditionally in basketball. In basketball, you have five players on each team on the court at a time. You have typically over100 possessions per team per game. And so you have a decent amount of data – and mixing. People substitute on and off. Even in basketball, you have this problem where certain players are on certain lineups. So LeBron James isn't going to play with the worst player on the court ever, or worst player on his team ever, because if he's not playing, then that's when the worst player comes on. They don't ever share the court together, and that causes some problems. So even in basketball, it's a hard problem, and people have been doing this for a long time, and they use methods like regularization and ridge regression type things to help with that. So a football's like a whole other level because the offensive line, those five players, unless there's an injury, like they're together, they're one unit and they don't really mix and match very much, which makes a model which is trying to separate and assign credit and blame to players very, very difficult. So what my solution was for this and this, you know, in the years since I've written this paper and had it published, one of the great things about research and putting it out there is I get feedback and then I look at things and if I had an updated data set, I would love to kind of do a follow-up on this one is that, we used additional information. And the reason why it's a Bayesian model is because we built priors distributions for each individual player based on additional statistics that we were able to gather. We had a company which is called Sports Information Solutions, and they collect what's called event data. So they'll look at a play and they might assign and say, this player, like this left tackle, he missed his block or he was beaten on his block. And so I could actually infer some prior distribution or I can build some sort of prior that says “this left tackle compared to all left tackles was worse at block percentage”. Okay, let's just make up a term and a sort of standardized scale compared to other left tackles. And that will then start to say, okay, if this player's bad at block percentage, then maybe he deserves more of the “blame” than this other player on his offensive line who's really good at block percentage compared to others of the same position. So by doing that sort of, we are able to disentangle a little bit some of the problems with players always being interdependent on playing. But of course, it's not a perfect – all models are wrong and have flaws, and I perfectly can acknowledge that. Some of the interesting findings from that, though, I mean, one of the things that came out of this paper was sort of an empirical estimate of positional value. And of course, and to no one's surprise (I split this up by both passing and running plays), the quarterback is by far the most impactful player on a passing play. And passing plays are way more impactful than running plays in football. They have a higher expected points added. They also have a higher variance. But interestingly enough, the quarterback is also the most impactful player on running plays. And it's not whether or not they're a good runner, like Jalen Hurts for the Eagles. He's a good runner. But even if someone is not a good runner, let's go back a few years to like Tom Brady, not known for running the ball. He was a very impactful player on running plays because the defense was worried about him passing. So just the fact that he was good at passing made the defense scared of him passing, which allowed the running backs that he would play with to do much better. That, you know, I think follows intuition and makes sense. One of the other things that we estimated with some of these positional values, and there's a lot of variance around them, but we're somewhat not traditionally what maybe a traditional football person would think. One of them is the offensive line. Most people consider the left tackle or whoever the – if it's a right-handed quarterback, it's the left tackle; if it's the left-handed quarterback, the right tackle – whoever's protecting the blind side of the quarterback to be the most important offensive lineman. And scouts think this because they're the most athletic, they're the tallest or the quickest. That's where they play. They're playing against the edge rushers, which are usually the most athletic defensive players on the field. But we found that more consistently than not, the center, when the center went out of the game or was injured, had a bigger impact on the team, on the offensive team, than if the left tackle was hurt. And so yes, the left tackle may be the most athletically gifted of them, but a center does something that's intangible, which is he is sort of the quarterback of the offensive line. He calls out the blocking schemes. So if we, we're here in Philadelphia, Jason Kelce, right? Like we can compare the Eagles offensive line performance with him in there and with him not. And yes, he's not going to be as athletic as the left tackle, but he probably understood and helped his teammates get in the right position more often than not. And that's a skill that you can really only measure in sort of these kind of plus-minus models. In the NBA, this is akin to a center who takes up so much space in the interior of the paint that no one even goes near him to shoot. So he doesn't have a lot of blocks, he doesn't have a lot of steals, but just his presence causes an absence of scoring. And we can observe the same thing, like the absence of a good center causes the offensive line to perform much worse because people are maybe more out of position. So that was one of the more interesting findings that we had.
18:13 Sheyan
It's really cool, and honestly, a really, really interesting paper. Going forward, I want to talk a little bit about your research at ESPN, and specifically some of the work that you did with metrics. The Strength of Record metric you helped create at ESPN has become a staple for the NCAA selection committee, for tournaments like the one we are in now with March Madness. Could you talk through the ideation and execution behind the SOR, and why is it a better evaluator of team deservingness than traditional simple win-loss records?
18:45 Dr. Sabin
One of the things that we were really proud of at ESPN was our team shaping narrative sometimes in sports. Nowadays, when you turn on any debate, I mean, ESPN's notorious now for all the debate shows, but if you turn on any sort of debate about who deserves to make the college football playoff or who deserves to make the NCAA tournament, often now people will separate this discussion to, well, which team is better and which team is more deserving? And we actually, at ESPN, our analytics team, we invented sort of that framing. We were the ones that brought it to the people on the air. And when I joined ESPN in 2015, they had what was called initially the Basketball Power Index for a year or two. And it was sort of this hybrid metric that tried to combine how deserving the team was and how good they were. And by doing that, it didn't really do either very well. And so when I came aboard, we separated and I was in charge of making two distinct metrics. So the current Basketball Power Index, which has now been around for 10, 11 years, I made and it was purely based to try to predict which team was more likely to win and the probabilities – it was built to optimize the probabilities of which team would win. So it actually does very, very well. It's comparable to what the Ken Palmer rankings would be. It's a little bit better at predicting games. No, it's not going to be Vegas or anything, but that's its purpose, is to predict games. The Strength of Record then, by separating it, we said, okay, what it's really asking the question is, what is the probability of, we usually define something like an average, we said like a typical 25th best team, or, you know, after strength of record came out, someone else made what's called WAB – Wins Above the Bubble. It's a similar concept. It's like someone who's roughly a bubble team or, you know, 25th ranked team, how many wins would they be expected to have against your schedule? And so it actually uses BPI, because BPI is an estimate of how good teams are. And then out of BPI, we can get the probability of a given team winning a game, you know, given who their opponent is, how many days rest they had, you know, whether they're traveling to elevation and they're not used to it, etc. So we essentially insert like a fake team into everyone's schedule and we say, okay, if it wasn't this team, but instead this fake team that's sort of like a typical, mid-ranked NCAA tournament team where to play against your schedule, what is their expected number of wins? And what you get out of this is essentially when statistics we talk about IID, which means “independent and identically distributed”. Well, this isn't identically distributed. Every game is a different probability, but they are independent observations. That's how we treat them. And so essentially, it's a Poisson binomial distribution, is what we call it. And essentially, you can calculate the probability that a team, if they have 31 games and their expected for each game has certain probability and they win 28 out of the 31, we can calculate the probability that some fake team wins 28 or more games against the schedule. And essentially, this is what Strength of Record is. So we do that for every single team's schedule, and the teams that have the highest wins above what a typical fake team would have, will have the highest Strength of Record. So like right now, that's Michigan #1, Arizona #2, and Duke at #3, UConn at #4. These are, you know, typically when people think best teams, that's actually how they're thinking, because most people look at wins and losses and compare it to Strength of Schedule. And that's what Strength of Record is. BPI is oddly actually very similar this year, except for that we don't think UConn is quite as good as their resume might suggest.
22:34 Sheyan
Very interesting. And I know you've mentioned this quite a few times while we've been having this conversation, but you often advocate for Bayesian approaches. For those less familiar with the methodology, what is a Bayesian framework? Why is a Bayesian framework superior to other methods when dealing with the inherent uncertainty of live sports and small sample sizes?
22:54 Dr. Sabin
Yeah, there's several reasons. You know, Bayesian statistics at large is essentially the idea that you start with some sort of prior belief about your parameters or your model or, you know, whatever it is, and you use data, then update those beliefs. It really is a great way to marry human knowledge with objective data. And in sports, this is really important because sample sizes are really tiny. Think about the NFL season, there's 32 teams and only 17 games each. So that's not a lot of information to really go off. The ratio of games to teams is not even 2 to 1. In other sports, it's larger, like in baseball, you have 162 games and only 30 teams. Okay, so it's over 5 to 1. But even that isn't very large. Anytime you're trying to fit a model to separate which teams are better, inherently, sports are noisy. So just because one team is better and they play two games, it doesn't mean they even win both of them or one of them, right? Just due to randomness and unlucky things that can happen in sports. And so Bayesian methods allow us to sort of infuse some of our prior beliefs or other information that might help us distinguish which teams are good. So I think about college football. You know, in college football, there's 130-something teams, and each team only plays 12 games. And if you just fit a typical regression model trying to estimate which teams are the best, you would get really odd results. You might get things like James Madison last year at #1 because they won so many games. You never saw James Madison play Alabama or an SEC team. Well, a Bayesian model would say, hey, I know that these teams from the SEC have better players going into the year, so I'm going to assume that they're better to start. As the teams play games, if James Madison is just so otherworldly that they are so much better than what we thought, then sure, they can catch up. But that knowledge of like, hey, SEC teams are probably better than James Madison, even if they lose to each other a few times, or like to other SEC teams, they're still probably better, helps us with prediction. And if you're trying to gamble or anything like that, not that I advocate for it, you're going to be really poor if you don't use a Bayesian model because it helps, essentially it helps us infuse additional data using some of our subject matter expertise.
25:14 Sheyan
Interesting. And I kind of want to look at the future of statistics in the world of sports analytics. With the advent of advanced tracking data like Next Gen Stats in the NFL, what is the next frontier of data collection? How is this changing how statisticians are looking at or working with the data that they have available?
25:33 Dr. Sabin
When I started at ESPN in 2015, it was all about player tracking data. That was the next new big thing. And all the major sports leagues in the world have player tracking data. And what I mean by player tracking data is you know where every player is and usually the ball on the surface, field or court or pitch, multiple times per second. And, the advanced teams have been using this data for years and been doing, you said, Next Gen Stats. That's the NFL's version of player tracking. The next frontier, which is just kind of getting going, is usually what we call “pose data” or, you know, three-dimensional, it's really like three-dimensional tracking, where we aren't just looking at where the center of mass is of each player. We know where their elbow, where their wrist, where their knees, where their head are in a three-dimensional space. This is something that the NBA is just kind of getting now into their framework, that Major League Baseball, like most situations in sports analytics, was the first to really get this kind of data. NBA typically follows next, and so far that's kind of as far as it goes. But I would expect in the future and NFL teams, soccer teams, NHL teams to get this kind of data. It's been really interesting for baseball because it has allowed really new research, especially into pitching mechanics, injury prevention, and things like that. And so in baseball, it's really useful for pitchers and batters, the swinging dynamics. In basketball, you know, people have ideas about maybe optimally changing the way someone shoots the ball to help them have a better quick release or whatever it may be, that data might be helpful. In the NBA before, we only knew if someone was near the shooter. So, hey, they were a foot away from the shooter, we would think it was contested. But if their back was to the shooter, that actually really isn't a contested shot. Well, now with this new data, we know if their wrist is in the shooter's face, that's a much more difficult shot than if their butt is next to them and they're facing away.
27:36 Sheyan
That's super interesting. Well, Dr. Sabin, I want to thank you for taking the time to join us today. I know all of our listeners really appreciate you taking the time out to talk about sports analytics. And thank you so much.
27:48 Dr. Sabin
I really appreciate coming on here. Anytime.