Showing posts with label coefficient of correlation. Show all posts
Showing posts with label coefficient of correlation. Show all posts

Monday, March 7, 2011

The Rising Cost of the Starting XI in the English Premier League

Somewhere in there is the most expensive starting XI, in terms of M£XI, in the history of the Premier League - Chelsea's 2006-2007 squad.  Too bad they finished second to Manchester United who were a full 1.5x lower than them on the M£XI metric.

In my recent series of posts on the correlation between squad transfer costs and table position in the English Premier League, it was shown that:
  1. Transfers are an ever increasing source of a team's players, with trainees constituting less than 20% of a team's squad
  2. Such squad transfer costs continue to escalate, with the average increase at £1.55M per year (or ~1.4% per year at the current average squade cost).  The actual average transfer cost inflation rate via the TPI database is often in the double digits percentage-wise. The average squad cost has now eclipsed more than £110M. 
  3. Squad transfer costs are highly correlated to squad wage costs, begging the question whether its really wages or transfers that is driving the high correlation between financially rich clubs and finishing in the top quarter of the table.
  4. A model was created to identify over and under performers, and was applied to teams and managers.
It's clear that if English soccer were an arms race, the result has been that those with the biggest budgets have won the war of championships and berths into UEFA competitions.

However, this explains the average behavior over the long term.  In the short term, players and managers have off years, players get injured, and in general things don't go according to plan.  Ultimately, this impacts who the manager can put on the pitch, which is quantified in Pay As You Play with the £XI metric.  This metric measures the cost of the starting XI on the pitch at a match.  If squad cost (Sq£) is a measure of cost of the weapons in a manager's arsenal (pun totally intended...), then the related £XI can be thought of the cost of the weapons he was able to bring to the battle.  Concurently, if MSq£ measures the relative cost of a manager's weapons against the league average's, a similar metric in M£XI can be used to measure the relative costs of a manager's talent on the pitch to the league average.

It's this metric - M£XI - that is best correlated to table position in the short term, and it will be the subject of several forthcoming posts.  Combined with the MSq£ this will provide a powerful forward- and backward-looking model at the impacts of squad transfer costs in the English Premier League.  After all, one must first have the tools at their disposal, and then deploy as many of them as possible on the pitch, to have a chance at success.

Note: In the interest of keeping this series of posts a bit shorter than the MSq£ posts, I will not be repeating my detailed explanation of the statistical theory that is being reused.  Newer readers, or those wishing a deeper discussion on the statistical theory, can see this post for a discussion of regression theory, this post for a discussion of prediction intervals, and this post for a discussion of ranking methdology based upon identifying over and under performance and then "shrinking variation before shifting the mean."

The Escalating Cost of Talent on English Premier League Pitches

While the average squad cost in terms of transfer expenditures has been going up over time, so too has the cost of the average starting XI that make it on to the pitch.  The table below shows how this cost has escalated over time.  Readers familiar with a similar graph in one of the MSq£ posts will notice a slight difference, as the graph below does not include the first three years of Premier League data.  This was done to eliminate the bias introduced in the data from those three seasons when the average squad and starting XI costs were diluted by the presence of two additional teams (click image to enlarge).


The elimination of the first three years actually worsens the fit of the regression lines compared to those for MSq£, but the coefficient in front of the x-term that indicates the typical annual rise in starting XI cost is far more accurate than the similar term found in the MSq£ graph.  Ultimately, the coefficient in the graph above is compared to the rise in MSq£ recounted at the outset of this post - £1.55M per year or 1.4% compared to 2010-11 squad costs.  In the case of M£XI, it is rising at a rate of £585.5k (or £0.5855M) per year, which in 2009-2010 £XI costs equates to a 1.1% increase.  While not statistically significant, this is lower than the rate of increase on the squad cost front.

Part of the explanation comes in the lower half of the graph, which re-plots the average utilization rate discussed in earlier posts.  With the removal of the first three seasons of data, the rate of decline in utilization is 0.33% per year, or approximately 1% every three seasons.  This declining utilization means a lower percentage of the average Premier League's cost, in terms of the Sq£ measure, is making it on to the pitch each passing year.  Thus while an increasing amount of money is spent on squad costs every year, a lower percentage of that squad cost shows up on the pitch and thus the growth in squad costs is higher than the starting XI cost.

Note: To the statistically inclined, the relationship between the average utilization rate and season is indeed statistically significant.  For the fourteen samples shown in the graph, an R-squared value of 0.4575 or higher would indicate statistical significance.

The Impact on Utilization on Starting XI Cost for the Big Six

While the first graph in this post explains what the league average has done over time, of greatest importance is to observe what the most successful clubs have done over the last several years.  Just like the MSq£ series, I have focused on the Big Six clubs in the post-Abramovich era - Arsenal, Chelsea, Liverpool, Manchester City, Manchester United, and Tottenham Hotspur.  The graph below shows their utilization rate by season (click on graph to enlarge)



Closely related to this graph is one that plots the difference between each squad's utilization rate and the season's average utilization rate.  This gives us an idea of how effectively the team used their purchased talent versus the average club.  Click on the graph below to enlarge.


There are a few things to take away from trends shown in the graphs above:
  • Manchester United is the only club that has a statistically significant higher utilization rate than the average team each season.  They finished with an average of 6.25% greater utilization than the league average each season, and only had two season out of seven that were below the average - the first (-0.5%) and last (-0.7%) in the series.  The first year of there latest championship three-peat saw them reach a peak of +15.8% vs. the league average.  During the seven season run, Manchester United has not finished lower than second amongst the big six in utilization, and hold the top three spots in the Big Six for utilization versus the league average over all seven seasons.  Say what you like about Alex Ferguson, but he not only knows how to buy talent he also knows how to get more of it on the pitch.
  • At the other end of the spectrum is Tottenham Hotspur.  They have finished to the positive side of the league average utilization only once - +5.9% in 2006-2007.  All other seasons they have been on the negative side, and while just barely not statistically significant they have averaged -3.2% over the seven seasons.
  • Arsenal, Chelsea, and Manchester City all have been on a general downward trend in terms of utilization since the 2005-2006 season.
  • Liverpool has generally bounced around the league average for utilization rate, but have been on a general upward trend since their nadir in the 2004-2005 season. Ironically, they reached their peak over the seven seasons in Rafael Benitez's final season as manager.
The graphs above confirm the general trend that teams who spend a good bit on transfer budgets for their squad get the talent on the pitch.  Besides being concerned about purchased championships, another side effect of huge spending is a "hoarding effect" - one where teams buy up players but where they don't see significant playing time due to limited space on the pitch.  While not the most economically efficient manner to win, it could nonetheless be a useful strategy to keep good players from signing with competitors.  This concern arises out of the general trend in the first graph, where over time transfer budgets have gone up by less of the costly talent shows up on the pitch.

Luckily, the statistics don't bear such a strategy out.  If one tests each season's utilization rates vs. MSq£'s, one will find that only four out of the seventeen seasons in the Premier League have had Pearson correlation coefficients that are significant - 94/95, 96/97, 99/00, and 05/06.  Each time, the coefficient was positive and the slope of the linear relationship was at least 3.5%, which means rather than horde talent the teams with higher squad transfer costs were putting more of it on the pitch.  The other 13 seasons showed no relationship between the two variables.  For the time being the main concern should remain focused on the ability of teams to buy better talent at higher costs and thus put more expensive talent on the pitch.

So what does all of this translate to in terms of an advantage in starting XI cost for the Big Six?  In general, their advantage versus the rest of the league has grown when compared to the similar advantages they enjoyed in MSq£.  See the graph below for a plot of the Big Six's M£XI for the last seven seasons (click on graph to enlarge).

Recall this graph that was part of my original post in the MSq£ series.  While most of the shapes of the lines are the same between the two graphs, the magnitude of the absolute values is certainly greater on the graph above.  Notice that Chelsea's M£XI reached a peak of 4.77 in the 2006/2007 season - this is a full 0.25 multiple above their similar peak in the MSq£ metric.  A similar shift is seen in Manchester United's M£XI data, but notice they actually close the gap to less than 0.25 with Chelsea's drop in utilization rate in the last few seasons.  Overall the order of the bottom four teams does not change much from their MSq£ order, although Tottenham's anemic utilization rate leads them to switch positions with Liverpool.  Thus, increasing one's squad costs relative to the competition seems to pay even greater dividends when it comes to starting XI transfer cost advantages.

Conclusions

The cost of talent on the pitch is certainly going up each year, although at a slower rate than the overall cost of the squad.  This is due to the steadily decreasing utilization rate, which is dropping by about 1% every three years.  While many of the same trends observed in the Big Six's MSq£ were maintained when looking at M£XI, the disparity was a bit larger and utilization rates provided a few subtle differences.

In the end, how does M£XI impact team performance in the table?  And if a relationship does exist, how can it be used to evaluate how well managers and teams have done given the cost of the players they could get on to the pitch?  These topics will be discussed in the second and third posts in this series.

Thursday, February 24, 2011

Comparing Econometric Models of the English Premier League: Reconciling the TPI and Soccernomics Data Sets

Note: This is a re-post from analysis I did back in January 2011 for the Transfer Price Index blog. I am posting it here to complete my series of posts on squad transfer costs, and to set up a forthcoming series of posts on the impact of starting XI transfer costs on table position

I’ve participated in many discussions since my original post on the relationship between a squad’s current transfer cost and their table position. Much of it has been centered on the debate over the predictive power of Soccernomics‘ wage data versus my analysis using current transfer costs. Many readers on The Tomkins Times have come to the same general conclusions as me: each analysis has its valid points and different uses, and the two are not necessarily in conflict with each other.

I’ve also had the pleasure of discussing the two studies with none other than Stefan Szymanski. I plan on keeping much of our conversation private, but you can get a sense of his respect for the overall Transfer Price Index approach and the differences in the two data sets via his review of Pay As You Play. Stefan’s review is a positive one, summarized best in the following observation.

“[I]n a fascinating new book Paul Tomkins, Graeme Riley and Gary Fulcher have developed a method of converting transfer fee data into a squad valuation… With every squad member given a value, this can then be used to compare spending to performance in the league. It is a true labour of love, collecting all the transfer fee values for Premier League clubs going back to the beginning of the 1990s.”
Szymanski closes out his review with this glowing recommendation:

“The book is a treasure trove of interesting financial facts and would make a great gift for any football statto…”
What’s interesting is how much correlation there is between the Soccernomics wage data and the TPI’s cost of the starting XI. Stefan’s metrics in the column are both relative measures (RW for wages and R£XI for relative starting squad cost), and he observes they show 90% correlation to each other. Unfortunately, the Evening Standard did not include the very compelling graph Stefan generated as part of his review of Pay As You Play. Luckily, Stefan has supplied us with that graph and it is reproduced below.


The graph clearly demonstrates the correlation between the two metrics, the weakness of the models at either end of the table, and the strength of the model in the middle of the table. Stefan’s observation of over predicting the resources needed for top table positions has been invaluable in explaining the discrepancy between regression predictions and historical data related to Champions League qualification that will be discussed in an upcoming post.

Stefan’s review rightfully points out the reliability of the publicly audited wage data versus the TPI’s privately compiled transfer data. At the same time, I would stand by the TPI as the most comprehensive and meticulously compiled set of transfer data within the English Premier League era. It was indeed a “labour of love” for the authors, a labour that continues to pay dividends in our financial understanding of the league.

Beyond the quality of the data and its impact on any resultant statistical analysis, Stefan’s data set has a bit of an advantage over the TPI. The Soccernomics wage data looks at overall team wages, thus taking into account the total cost of operating the squad in current British pounds. Combine this with the fact that wages are a dynamic measure adjusted over time by team and player, while the TPI is a static inflation of a one-time transfer fee, and we see why wages may be a better predictor of actual team success. It’s also no surprise that the £XI metric correlates very well with that wage data, as it takes into account all the players who have made it on the pitch and how much time they spent on it. There’s no dead weight contributing nothing to the team’s performance on the pitch, good or bad.

Indeed, analysis by Graeme Riley and me has proven this point statistically. Graeme looked at the squad and XI transfer cost order versus table position, while I looked at the multiple of the average squad and XI transfer costs. Both Graeme and I calculated these for each team, and then quantified the correlation of each metric to finish position for each individual season via the square of the Pearson product moment correlation coefficient (the commonly seen R² value in a regression plot). In Graeme’s analysis, the order of £XI had a higher R² value than the order of Sq£ in 16 out of the 18 seasons. In my analysis, M£XI had a higher R² value than MSq£ in 14 out of the 18 seasons. In the final comparison, I looked at the average and standard deviation of the R² values for each metric – order of £XI, order of Sq£, M£XI, and MSq£ – to determine which provides the best, most consistent prediction of table position over the 18 seasons. The M£XI had the lowest overall standard deviation (14.7%) and highest overall average (45.4%), indicating it provided the best fit versus table position (although it is far lower than the R² values in the long-term analysis in my original post and Soccernomics). Ultimately, this confirms my preference for relative measures, especially multiples of averages, and why I prefer to look at long term averages rather than individual seasons.

On the other hand, the TPI data I used in my original analysis only considered the impact of the total cost of transfers on team performance, and neglected those of the free variety as well as trainees. It also doesn’t look to utilization rate. It essentially looks at a reduced data set from the full squad or starting XI, and the graph below quantifies how much of a reduced data set non-free transfers represent over the history of the Premier League.


The graph above shows the cumulative percentage of three types of players within the league each year as categorized within the TPI – trainees, free transfers, and the rest of the players. The vast majority of this final category consists of transfers with confirmed fees, while the rest of it consists of a small number of players whose transfer fees couldn’t be confirmed. The graph is cumulative, so to understand the percentage of free transfers for any single year one must identify the free transfer value on the graph and then subtract the corresponding trainee value from it. As an example, the cumulative percentage (represented by the upper value of the red zone) in 2001-02 is approximately 30% while the league share of trainees is about 20%. This means that free transfers made up about 10% of the league in 2001-02.

What is clearly seen via the graph is that transfers have consistently accounted for nearly 70% of the Premier League’s players since its inception. That’s not to say 70% of the players transfer teams each year, but rather that at some point in their past they were purchased by the team they played for that season. What has changed over the league’s eighteen years is the number of trainees within it. This number has plummeted from nearly 30% of league player classifications in 1992/93 to below 20% by last season. Much of this change has happened due to an increasing number of free transfers, which were given official UEFA sanction with 1995's Bosman ruling. Free transfers have gone from only 2% of league player classification in 1995 to nearly 10% last season. Overall, transfers of any variety came to represent 80% of league players by the 2009/2010 season. In many regards, the Premier League is a microcosm of the increasingly globalized world it operates within: greater international ownership and investment, greater employee mobility, fewer employees staying with a single firm from “graduation” to retirement, and increased dominance by a few brands within the marketplace.

What this all means is that any analysis of league performance on a squad basis that uses the TPI is going to miss nearly 30% of the players in the league. Given that fact and the reasonably good R-squared value my regression analysis achieved, I would consider the relationship to be a reasonably strong one. Ultimately a study by Stefan Szymanski, similar to this one where he statistically examined the causality of the wage/performance correlation, would be fascinating. We might then determine whether it was transfer fees, wages, or table position that drove the relationship with the other two. That is a very advanced analysis best left to a statistician of Stefan’s caliber.

At the end of the day, what Stefan’s analysis, my analysis, and the overall TPI database prove is that one must pay, and pay big, to compete for the top few spots in the Premier League. One must pay dearly for the right to even negotiate wages with 70% of their players that end up on their squad, and then they must be willing to pay dearly again to keep the talent to challenge for a top spot. Each metric, whether it’s based upon £XI or MSq£, has its use in quantifying the roll of ever increasing transfer budgets in a club’s success. Generally, I concur with Paul Tomkins’ assessment that “Sq£ is the only predictive tool, but £XI is surely the better retrospective analyzer.”

To a certain degree this all makes sense, as we want a somewhat meritocratic system where excellence is financially rewarded. It all gives us pause, however, when the same teams can dominate everyone else each year by outspending their rivals, sometimes even with money that had no origination in the soccer world in which each team operates.

Saturday, April 17, 2010

Explaining regression through an enhanced Soccernomics analysis


The full results of the Soccernomics pay-for-play regression, including the actual equation and the distribution that accompanies it.

In this post I will attempt to tackle the much-abused and little understood topic of regression theory by using one of the more famous models in the soccer community: Soccernomics' infamous Figure 3.1 showing the pay-to-win regression of the top two English soccer leagues. This post will focus on the technical aspects of regression theory, walking through the step-by-step process likely used by the authors of the book. Occasionally I will go beyond what the authors showed, just to provide something more than a regurgitation of their study and hopefully provide greater insight into the general theory behind the analysis.

Before beginning, I would like to point out that the original inspiration behind this post was one via a request from one of my first tweeps. It just goes to show you that if you reach out to me on Twitter or via the open thread on the blog, I will respond with the requested analysis. I hope you enjoy this post, jblock49!

Background

In their seminal work Soccernomics, authors Simon Kuper and Stefan Szymanski lay out a very intuitive yet startling correlation: to finish higher in the tables of the top two English soccer leagues, one must spend more money than their opponents. Their analysis of the data, the results of which are shown in the graph below, shows that the team payroll as a function of a multiple of the leagues' average payroll explains a whopping 88.7% of the variation in finishing position within the tables.

Figure 1: Regression analysis from Soccernomics

The authors' use of transformed data (see "log" denotations on each axis), and their lack of discussion around the uncertainty inherent to any regression analysis, provides fertile ground for a case study in regression theory. Too often we equate regression analysis with dropping two data sets into Excel, plotting them with a fitted line, and hoping that the R-squared value comes out good enough to justify a relationship. What many people don't realize is that there are many more requirements of a good regression study whose conclusions can be accepted. I will explain those assumptions here.

The prerequisite: a normally distributed response variable

In the regression world, there are two types of variables. If we can imagine an equation in the form of y= mx+b, the following variables are named:
  • y = response variables
  • m = regressor coefficient
  • x = regressor variables
  • b = regression constant
In this case, y and x are data sets used to develop m and b and provide the regression equation we are used to seeing. Before beginning any regression analysis, we'd like to see a normal distribution to the data set y. In the case of the Soccernomics study, the regressor was the multiple of league average pay while the response was finishing position.

The trick with any analysis of league finishes is that the variable of interest is not the actual finish position, but rather how you finish relative to everyone else. The logic is similar to that used for wages - you don't need to spend a certain amount to win, just more than your opponents. That's where the first transform of the original data found in Soccernomics comes in. Instead of looking at the raw finish position, the authors looked at a relative finish position that provided a rough indication of how frequently another team would finish ahead of another. They did this by using the transformed data set of:

p/(45-p)

It appears they used 45 as a normalizing value due to the combined number of positions available to the teams in the top two leagues including relegation and promotion. This transform meant teams at the top of the table would have low values, while those at the bottom would have high values. However, this transform presented some challenges you can see in Figure 2 - the transformed data wasn't normal as indicated by p-value <>

Figure 2: Graphical Summary of p/(45-p)

The authors were on the right track, but they needed to perform an additional transform of the data to make it normal. This is when they would have turned to their commercial statistics program and asked it to run several simulations of common transforms (logarithmic, natural log, etc) to help identify a transform that provided a normal distribution. They settled on the natural log transform, and used a -1 coefficient in front of it. This is because a natural log transform would have taken the teams with high finish positions that now had low numbers in the p/(45-p) transform and given them a larger negative number after the natural log transformation. Using a -1 coefficient to invert the data set makes sense - high pay scales might lead to higher finish position, but not the other way around - and it doesn't affect the normality of the data set. The authors' suspicions were correct, and they were rewarded with a normal data set. See Figure 3 below, where the p-value is > 0.05 and thus we accept the assumption that the data is normally distributed.

Figure 3: Graphical Summary of -ln[p/(45-p)]

Transforming the Regressor

To preserve any chance of a linear relationship between the response variable and the regressor, the authors then likely set out transform the wage data by the similar method. A transformation using a natural log function produces a normal data set shown in Figure 4.

Figure 4: Graphical Summary of ln(wage multiple)

Prior to regression: determining if correlation exists

Prior to beginning regression modeling, the authors would have performed a correlation study and this is where we start getting into the claims of the book. Completing a correlation study is the first step because it helps understand the total amount of variation explained by the relationship of the data, and it provides a good statistical test as to whether the value is high enough to justify a regression analysis and equation. No more guessing at Excel-based R-squared values!

Correlation is measured via a correlation coefficient. This coefficient is calculated per the formula below, and it is essentially trying to measure the scatter around the mean of the two data sets (i.e. x(i) - x(bar), etc.) in relation to the overall scatter of the data (s(x), representing the sample standard deviation).

A graphical representation of this equation can be found in Figure 5.

Figure 5: Graphical representation of correlation measurement

Once a correlation coefficient has been calculated, it can be compared to values assigned to different risk levels based upon the number of samples in the data set. In the case of the English league data set of 58 samples, the authors would need a correlation coefficient of between 0.2948 to 0.3218 or greater to conclude there was less than a 1% risk of incorrectly concluding that a significant enough relationship exists between the two variables to proceed with a regression study. Instead of using the lookup tables, I used a statistical software package. Using the author's data and running it through a correlation study yields the results in Figure 6.

Figure 6: Correlation study of -ln[p/(45-p)] and ln (wage multiple)

The results of the study clearly indicate a low risk of assuming a correlation exists (p-value = 0.00), and that the relationship between the two variables explains 94% of the variation in their behavior. Now, this is a little bit different than the claim in Soccernomics, which was 92% of the variation in league position being explained by team expenditure. I triple checked the data that I copied from Figure 3.2 in the book, and could find no errors. I don't know if it is a typographical error in the book, or the result of some other analysis. Nonetheless, there seems to be a strong relationship between the two variables that warrants a regression analysis.

Checking the Regression Results

The step of making the regression equation and plot at this point is a formality for most of us. It should be noted that most regression algorithms are based upon the least squares method, which means that it uses multiple equations to describe the behavior of the system and finds the one regression equation that minimizes the sum of the squares of the errors made between the regression equation and the original data.

The trick isn't in the regression equation itself, but in the results that it produces. In general, regression analyses must have:
  • A normally distributed response variable data set
  • A statistically significant correlation coefficient
  • Have their residuals meet five basic requirements
Residuals are the difference between each response variable data point and the corresponding predicted response value from the regression equation. Residuals represent the error in the statistical model. For a regression equation to be accepted as statistically valid, the following five requirements must be met:
  • Residuals are normally distributed with a mean of zero
  • Residuals are random and show no pattern
  • Residuals have constant variance
  • Residuals are independent of the values of the regressor variables
  • Residuals are independent of each other
Meeting these requirements ensures that the relationship between the two data sets is real, and not the effect of an unseen factor, confounding variable, or test procedure.

When a regression analysis is carried out on the transformed finishing position and wage data, Figure 7 is produced. The two plots on the left suggest that the data is normal, and the normality test in Figure 8 confirms they are (barely) normally distributed with a p-value of 0.063. The two figures on the right measure the other four characteristics. There does seem to be some trouble at either end of the data sets. The upper right graph shows some increasing spread (non-constant variance) as one goes to either end of the data set. The graph in the lower right indicates a consistent under or over prediction in the model when looking at the ends of the data sets. I don't know if it would have been enough to conclude that the regression analysis was invalid, but it would definitely have caused me to look at the reasons why the ends are so skewed.

Figure 7: Four-in-one plot for regression residuals

Figure 8: Graphical summary of residuals

The Model's Results

Now that all of the assumptions have been checked, it is time to move on to analyzing the model's results. We are often used to seeing regression simply as a line and an R-squared value, but there is much more going on behind the scenes. As we are all aware of regression graph in Figure 3.1 in Soccernomics, I have instead focused on the the statistical analysis found in Figure 9 below.

Figure 9: Output from regression analysis

The first set of data to focus on is the R-Sq and R-Sq(adj). Notice that the R-Sq value is different that the correlation coefficient we calculated earlier. The R-Sq value is measuring the proportion of variation that is explained by the regression model that has been generated, and not the overall variation explained between the two data sets. R-sq is simply the SS (sum of squares) value from the "Regression" row in the Analysis of Variance subsection divided by the SS value from the "Total" row. Hence, 88.3% of the variation being explained by the regression model.

The R-Sq(adj) term is a modified form of R-Sq, which takes into account the number of terms in the model. In this case, the linear regression performed only has one term. If one were to try and predict the behavior by a cubic regression (i.e. an equation of y=ax^3+bx^2+cx+d), the equation would have three terms. The R-Sq adjusted is a way of telling whether or not the terms you are adding by going to more complex regression equations are actually improving the fit - higher R-Sq(adj) values means better regression bang-for-the-buck. In the case of the Soccernomics study no other regressions were run beyond the linear one.

With such a high R-Sq value, we can also look at the statistical tests of the predictors. Both the constant and the regressor have p-values <>

Finally, we can move on to the equation itself which was not included in the book. The equation relating average finish position to wages is:

-ln[p/(45-p] = 0.5465 +1.487*ln(wage multiplier)

For all you EPL fans out there, this is the equation that matters. If you ever were able to get you hands on a list of team payrolls at the beginning of a season, you would be able to project the average outcome of the season if it were played repeatedly at those price ratios for several years in a row.

Statistics are always about distributions

Finally, I'd like to close the discussion around regression by extending the author's analysis a bit further. For simplicity's sake, they presented what I would call a simplified model of regression. It works, and it gets the point across. But to statisticians, it's all a bit too simple. Statistics are always about distributions - nearly every test involves calculating means and variances or standard distributions. Sometimes these quantities are used as checks and pre-requisites to beginning tests, other times they are critical elements within the tests. More importantly, the output of the tests always has a distribution associated with it.

In the case of regression, the equation the line represents is associated with the mean values for the relationship. In reality, the likely distribution of the response data set at any one regressor can be calculated. The distribution of the predicted response will vary along the length of the regression line and across the data set, and is dependent upon the underlying data sets used in the regression. Thus, what we're really calculating when using the regression is the likely average of individual outcomes that could occur over time at that regressor test point that has been selected. In reality, there are a range of outcomes. This range of outcomes is called a prediction interval (PI), and the tighter band inside of it is the confidence interval (CI) for the predicted mean value.

Figure 10 is the same as the figure at the beginning of this blog post. It is the same regression analysis performed throughout this study, but I have turned on the 95% CI and PI lines. These lines represent the range of values of where we are 95% certain individual occurences of future observed data and its means will fall for each of the regressor values on the x-axis.

Figure 10: Regression analysis with PI and CI included.

As one can see, much of the observed variation from the data sets falls within the CI lines. This graph gives a much more complete picture of the expected behavior given the data, and becomes far more useful than the standard single regression equation line when one wants to understand whether or not a single season's finish in the English leagues is expected or an anomoly.

Conclusion

This was, believe it or not, a brief treatment of regression theory. I hope it has provided you a much better understanding of all the calculations and checks that must go on when performing regression studies. The next time someone shows you an Excel graph with a line and an R-squared value on it, ask them if they have checked their residuals and the p-values associated with the terms in the regression equation. Ask them if they know what the correlation coefficient for they data is, or if they are sure the response variable data is normally distributed. Until they can show you the data confirming all the critical checks of a regression analysis, it's just a pretty picture.

LinkWithin

Related Posts Plugin for WordPress, Blogger...