A Premier League data scientist answers popular questions
We asked readers of the Football Extra newsletter to send in questions to be answered by a Premier League data scientist. The result was an inbox full of superb queries (and a moan or two about xG). Topics covered include new ways to measure attacking threat, whether football intelligence can be quanitied, the sheer amount of data teams have, how to get into data science and what can be learny from other sports. We summarised the questions and put them to Alex Marin Felices, a football data scientist with experience at Nottingham Forest, Olympiacos and Rio Ave. Alex is also the author of The xG Football Club , where he writes about football analytics and data science. How much influence does the data team wield... Does the data team at a club get to influence tactics or team selection? And do the coaches and data team ever disagree? Alex's answer: Potentially they influence tactics and selections, yes; but usually not in the sense of a model producing a starting XI and telling the manager who has to play. A data team can help answer much more specific questions. Imagine a team has three possible wingers. The opposition full-back might be very aggressive when pressing but leave space behind them. One winger might be particularly effective attacking that space, while another is better receiving the ball to feet against deeper defences. You can analyse those kinds of stylistic match-ups. The same applies defensively: perhaps an opponent creates a very high proportion of their chances through crosses, struggles when pressed in a particular area, or relies heavily on one player progressing the ball. But analytics is only one layer of evidence. Coaches also have training performance, fitness, tactical understanding, player confidence and a lot of contextual information that may never appear in the data. That also means analysts and coaches can disagree, and I don't think that's necessarily a bad thing. A model might highlight a player as a strong option, while the coaching staff see something in training or video that makes them less convinced. Or the opposite can happen. The interesting part is understanding why those perspectives differ. Sometimes the model is missing context; sometimes intuition is being influenced by a small number of memorable moments; and sometimes the two sides are simply answering slightly different questions. The role of analytics is not to replace the decision-maker, but to provide another source of evidence. In fact, if the data simply confirms what everyone already thinks every time, it probably isn't adding very much. The route into data science Do you have to be a huge football fan already to work in football data science and how can people make start in this field at any age? Alex's answer: A lot of people working in football analytics originally came from statistics, computer science, engineering or completely different industries. Coming from outside can actually be valuable because you sometimes question assumptions that people within football have accepted for years. I wouldn't say football knowledge is unimportant, but it's something you can definitely develop as you go. The technical side alone isn't enough: you can build an incredibly sophisticated model and still produce something that isn't particularly useful if you don't understand the football problem a coach, scout or analyst is actually trying to solve. For me, the ideal combination is having strong technical skills while continuously developing your understanding of football, and also learning how to communicate insights effectively to non-technical stakeholders. That's also why multidisciplinary teams are so useful. You don't necessarily want ten people who all see football in exactly the same way. Someone with 20 years of coaching experience and someone coming from machine learning may look at the same problem completely differently, and the interesting part is often where those perspectives meet. For someone trying to get started, I would focus on building those skills together rather than waiting until you feel like an expert in football or data science. There is now plenty of public football data available, so a good starting point is to take a real football question you are genuinely curious about, analyse it, and try to explain the result clearly. Building projects is especially useful because it forces you to go through the full process: defining a useful question, working with imperfect data, choosing the right method, interpreting the result and communicating it. And I think sharing that work publicly can help a lot as well. A small number of thoughtful projects that show how you think are often more valuable than simply listing lots of technical skills. Learning from others and the future Can football learn anything from the analysis of other sports and - if we fast forward 10 years - might we see a team totally run by data? Alex's answer: A lot of ideas used in football analytics actually have roots in other sports. Baseball has a much longer history of using statistical models for player evaluation and recruitment. Basketball has pushed a lot of work around tracking, spacing, shot selection and player interactions, while sports such as ice hockey have also developed models for valuing actions and estimating how players contribute to creating or preventing scoring opportunities. Football can learn a lot from all of them, and in many areas I think we're still behind some other sports analytically. But there is an important reason for that: football is a particularly difficult sport to model. It is very continuous, low-scoring and highly interactive. Twenty-two players are constantly moving, and the value of one player's action depends heavily on what everyone else is doing around them. In baseball, for example, many situations can be broken down into relatively discrete interactions. Football is much less tidy. That makes simply importing a model from another sport difficult. The underlying ideas might transfer, but they normally have to be adapted to the characteristics of football. And that's one reason I don't think we'll reach a point where a club is completely run by models, with every signing, team selection or substitution decided automatically. There is simply too much context that either isn't in the data or is extremely difficult to quantify. A coach knows what happened in training that week, whether a player understood a tactical instruction, how someone is adapting to a new environment, whether confidence is low, or whether there are personal or physical circumstances affecting them. A model only knows the information it has been given. I think models will become involved in more and more decisions, and they will probably become much better at showing us the likely consequences of different options. But ultimately, I think there should still be a human making the decision. Partly that's because humans can incorporate context that models cannot, but also because football decisions have consequences. Someone ultimately has to take responsibility for choosing the player, making the substitution or spending the transfer fee. Data can make that person much better informed, but I don't think it should remove the decision-maker from the process. Thoughts on xG... Isn't xG flawed because it only measures shots, not attacking threat? Also you have written about measure possibilities rather than the current method of measuring outcomes. How would this change analytics? Alex's answer: This is one of the limitations of xG, but I wouldn't necessarily call it a flaw. xG is designed to use the context around a shot to estimate its chance quality: given where and how the shot was taken, how likely was it to become a goal? If nobody gets a touch on the cross, there is no shot, so there is nothing for an xG model to evaluate. But clearly the team has still created something dangerous. That's why there are other models designed to value what happens before the shot. Expected Threat (xT), for example, measures whether moving the ball from one area to another increases the likelihood of eventually scoring. Other possession-value models, such as VAEP, try to assign value to every pass, carry or other on-ball action. The second part of the question is also interesting. If one striker makes a brilliant run and just fails to finish a difficult chance, while another striker never makes the run at all, traditional event data can fail to recognise the advantage created by the first player, because there is no action at all recorded for the second. That's one reason tracking data is so exciting. Once you know where every player is at any given moment, you can start evaluating the run itself: did the striker recognise the space, did they arrive at the right moment, and could they realistically reach the pass? The same idea applies to passing. Traditional statistics might reward a safe five-metre pass because it is completed, while punishing a much more difficult pass that could have created a huge opportunity if it came off. What we really want to know is not only whether the action succeeded, but whether the player chose the best option available at that moment. That means modelling the alternatives: what else could they have done, how difficult were those options, and what would the likely outcome have been? So increasingly, football analytics is moving away from simply measuring outcomes. We want to understand what could have happened, whether the player made a good decision, and then separately whether they executed it well. What's on the data analyst's horizon? What will be the next xG, do you think, i.e. a statistic that is used a lot in match reports, on TV etc as a way of giving fans more information? Alex's answer: I'm not convinced there will be another single statistic that becomes as universal as xG. Part of the genius of xG is its relative simplicity. It takes quite a complicated question (how good was this scoring opportunity?) and gives you a number that most supporters can understand relatively quickly. Even then, though, xG is still often misunderstood or used to answer questions it was never really designed for. One possible next introduction is Expected Threat (xT). Rather than only looking at shots, xT tries to measure how much moving the ball from one position to another increases a team's attacking threat. Fifa brought a related idea to a much wider audience during the 2026 World Cup through its Match Momentum graphic, which tried to show how the balance of attacking threat changed throughout a match. That is already more complex than xG, and I think that points towards the broader challenge. A lot of the interesting questions we're tackling now are even harder to summarise in a single number. How much space did a player create? Did a defender prevent an opportunity just by their positioning? Did a midfielder choose the best passing option? Was a striker's run valuable even if the pass never arrived? Tracking data increasingly allows us to investigate those questions because we know where all 22 players are rather than just recording events involving the ball. So, if I had to guess, the next major step won't necessarily be one famous metric. It will be better ways of measuring off-ball behaviour, decision-making and the opportunities available to players. The challenge will be turning those complex models into something simple enough that supporters can understand and use them in the same way they now use xG. Are the other 21 players just walking? Can data spot players who are superb yet might not touch the ball as much as some similar players? Also, football intelligence is used by pundits a lot. Is this something that can be measured in any way? Alex's answer: I think these two questions are very closely related, because a lot of what we describe as "football intelligence" happens away from the ball. Traditional event data is naturally biased towards things that happen to the ball. We record passes, tackles, shots and carries very well, but football is full of important actions that never produce an obvious event. Only one player can have the ball, so what are the other 21 players doing? A centre-back might close a passing lane so effectively that the opposition never attempts the pass. A midfielder might position themselves in a way that forces the attack into a less dangerous area. A forward might make a run that drags two defenders away and creates space for a team-mate. Those are all examples of things we might describe as intelligent football actions, but they are difficult to measure if we only look at what happens to the ball. Tracking data helps enormously because we can observe where every player is and how those positions change over time. That allows us to start asking questions such as: did a player move into the right space? Did their movement create space for someone else? Did a defender remove a dangerous passing option? Did a player choose the best option available to them? We are getting much better at measuring these things, but I don't think "football intelligence" will ever be captured perfectly by one statistic. It is probably better thought of as a combination of positioning, movement, awareness and decision-making. We're gonna need a bigger server How much information do data scientists have on players from Europe and beyond? With this, is it limited to the bog, rich clubs or do ambitious lower-league clubs also have data teams? Alex's answer: There is a huge amount of football data available now, particularly across professional leagues. Clubs can buy data from specialist providers covering hundreds of competitions around the world. At the most basic level, that includes things like appearances, minutes, goals and assists, but more detailed datasets record almost every on-ball action a player makes: passes, shots, carries, pressures, duels and where on the pitch those actions happened. Increasingly, there is also tracking and physical data. That can give you information about things like a player's speed, movement, positioning and the runs they make. So if you're sitting in England analysing a player elsewhere in Europe or beyond, you can potentially know an enormous amount about them without having watched them live. A lot of that underlying data is available commercially, so different clubs may start with the same information. The competitive advantage often comes from what they do with it: the models they build, how they combine different sources, the questions they ask and, most importantly, how well the information is integrated into recruitment and decision-making. That's also why I wouldn't say data analytics is only for the richest clubs anymore. The biggest clubs can obviously invest much more in people, infrastructure and more expensive datasets, but ambitious clubs with smaller budgets can still do very good work. In some ways, when you can't simply outspend everybody else, finding an informational advantage can become even more important. So there are differences in resources, but having more data doesn't automatically mean making better decisions. How you use it matters much more. Is data a game-changer? Has the increased use of data analysis within football changed tactics and the options players take? Alex's answer: I think it has influenced football, although it's difficult to separate the effect of data from broader tactical trends. The decline of long-range shooting is a good example. Data makes it very obvious that, on average, shots from 25 or 30 metres are much less valuable than shots closer to goal. But coaches were also independently moving towards possession styles designed to create better chances. So did analytics cause that tactical change, or did analytics simply provide evidence for something good coaches were already discovering? Probably a bit of both. Where data can be particularly influential is when it repeatedly challenges intuition. If analysis consistently shows that a certain type of action produces very little value, or that another situation is much more dangerous than people assumed, eventually that information becomes part of how teams think. Set-pieces are probably one of the clearest areas where this has happened. Clubs can analyse thousands of corners, free kicks and throw-ins and identify patterns that would be very difficult to recognise simply from watching matches. I don't think data creates tactics on its own. But it can give coaches evidence that reinforces an idea, challenges one, or highlights something worth investigating.
News Source : Yahoo Sports and Read the full article →




