My interest in this blog is, one, to rate and rank teams, but, ultimately, to be able accurately quantify college football teams so that I can more accurately forecast game outcomes. While I'm revving up for another season, I thought it might be interesting to take a closer look at the industry standard in college football forecasting-the Vegas line. And since I'm writing about it now, you can guess I found something that at least I consider interesting.
I'll start with a quick note of the Vegas line. The line is not created to forecast results--its sole existential purpose is to split bets 50/50 above and below. If too many bets are made above or below the line then the line is adjusted. Therefore, the line is a product of the interaction of two forecasting methods. The first method uses a single model-part statistical, part qualitative-that attempts to predict the public attitude. The second method employs market forces, allowing the public to aggregate information and, thus, move the line up or down according to public sentiment. The public responds to the line and the line responds to the public. The Efficient Market Hypothesis tells us that if the Vegas casinos provide an open market, all available information should be aggregated in adjusting the line and it should be impossible to consistently outperform the line without special insider information (which can be purchased from your neighborhood crooked NBA ref). If someone can find a model that can consistently outperform the Vegas line (after it has been adjusted to bettor response) they can establish that the line does not satisfy the EMH-and they can make themselves millionaires. I will not, here, provide any evidence that the line does not satisfy the EMH.
Now, to the numbers. In 2007, the Vegas line and the actual game outcome (both in terms of point differentials) had a correlation of r=.4368. This is relatively high; as I mentioned before, this is the industry standard, but it is not overwhelming. For a little interpretation, if we were to guess the point differential using the line, we would, on average, be about 18% closer than if we just guessed that every game would end in a tie. And that's the industry standard.
The line is 12.25 points off from the actual point differential on average. But as you can see in the graph, the distribution is skewed--the average is pulled up by a few cases where the Vegas gamblers really missed the boat.
My first theory was the the Vegas line would have a tougher job accurately predicting the point differential in higher scoring games or games with a larger expected point differential. But with a correlation of .0635 of the total (total) combined scoring and the absolute difference between the line and actual outcome (difference). There was a slight increase in difference as the total score increases, but when we consider that the total has to be large in many cases for the difference to be large, we have to rule this out as a viable theory. So, is the line less accurate when one team is definitely better than the other (which leads to quirky 4th quarters with backups and such)? The answer is, again, a resounding no. In fact, if anything, the trend runs in the opposite direction.
Does the line give preference to favorites or underdogs? If you were to put one dollar on the underdog in the 688 games in 2007, you would have gone home a winner 346 times (50.3%), raking up a $4 profit. More impressive than a split that is almost exactly 50/50 is the fact that the mean and standard deviation of outcomes on both sides are almost exactly the same--in other words, the line is right in the middle of its own error distribution.
These null results were to be expected and they fit nicely with the efficient market hypothesis--the actually outcomes are normally distributed around the line. But one other null result was not expected. The line does not become a more accurate predictor of outcomes as the season progresses. One would think, as the season progresses, we get a larger data set that we can use to make more accurate predictions, but instead the predictions don't get more accurate. My only explanation is that injuries through the season cause enough fluctuations to offset the increased sample size-but I still find it surprising that the average error doesn't have more of a downward trend as the season progresses.
BPR | A system for ranking teams based only one wins and losses and strength of schedule. See BPR for an explanation. |
EPA (Expected Points Added) | Expected points are the points a team can "expect" to score based on the distance to the end zone and down and distance needed for a first down, with an adjustment for the amount of time remaining in some situations. Expected points for every situation is estimated using seven years of historical data. The expected points considers both the average points the offense scores in each scenario and the average number of points the other team scores on their ensuing possession. The Expected Points Added is the change in expected points before and after a play. |
EP3 (Effective Points Per Possession) | Effective Points Per Possession is based on the same logic as the EPA, except it focuses on the expected points added at the beginning and end of an offensive drive. In other words, the EP3 for a single drive is equal to the sum of the expected points added for every offensive play in a drive (EP3 does not include punts and field goal attempts). We can also think of the EP3 as points scored+expected points from a field goal+the value of field position change on the opponent's next possession. |
Adjusted for Competition | We attempt to adjust some statistics to compensate for differences in strength of schedule. While the exact approach varies some from stat to stat the basic concept is the same. We use an algorithm to estimate scores for all teams on both sides of the ball (e.g., offense and defense) that best predict real results. For example, we give every team an offensive and defensive yards per carry score. Subtracting the offensive score from the defensive score for two opposing teams will estimate the yards per carry if the two teams were to play. Generally, the defensive scores average to zero while offensive scores average to the national average, e.g., yards per carry, so we call the offensive score "adjusted for competition" and roughly reflects what the team would do against average competition |
Impact | see Adjusted for Competition. Impact scores are generally used to evaluate defenses. The value roughly reflects how much better or worse a team can expect to do against this opponent than against the average opponent. |
[-] About this table
Includes the
top 180 QBs by total plays
Total <=0 | Percent of plays that are negative or no gain |
Total >=10 | Percent of plays that gain 10 or more yards |
Total >=25 | Percent of plays that gain 25 or more yards |
10 to 0 | Ratio of Total >=10 to Total <=0 |
Includes the
top 240 RBs by total plays
Total <=0 | Percent of plays that are negative or no gain |
Total >=10 | Percent of plays that gain 10 or more yards |
Total >=25 | Percent of plays that gain 25 or more yards |
10 to 0 | Ratio of Total >=10 to Total <=0 |
Includes the
top 300 Receivers by total plays
Total <=0 | Percent of plays that are negative or no gain |
Total >=10 | Percent of plays that gain 10 or more yards |
Total >=25 | Percent of plays that gain 25 or more yards |
10 to 0 | Ratio of Total >=10 to Total <=0 |
Includes
the
top 180 players by pass attempts)
3rdLComp% |
Completion % on 3rd and long (7+
yards) |
SitComp% |
Standardized completion % for
down and distance. Completion % by down and distance are weighted by
the national average of pass plays by down and distance. |
Pass <=0 | Percent of pass plays that are negative or no gain |
Pass >=10 | Percent of pass plays that gain 10 or more yards |
Pass >=25 | Percent of pass plays that gain 25 or more yards |
10 to 0 | Ratio of Pass >=10 to Pass<=0 |
%Sacks |
Ratio of sacks to pass plays |
Bad INTs |
Interceptions on 1st or 2nd down
early before the last minute of the half |
Includes the top 240 players by carries
YPC1stD |
Yards per carry on 1st down |
CPCs |
Conversions (1st down/TD) per
carry in short yardage situations - the team 3 or fewer yards for a 1st
down or touchdown |
%Team Run |
Player's carries as a percent of team's carries |
%Team RunS |
Player's carries as a percent of team's carries in short
yardage situations |
Run <=0 |
Percent of running plays that
are negative or no gain |
Run >=10 |
Percent of running plays that
gain 10 or more yards |
Run >=25 | Percent of running plays that gain 25 or more yards |
10 to 0 | Ratio of Run >=10 to Run <=0 |
Includes the top 300 players by targets
Conv/T 3rd | Conversions per target on 3rd Downs |
Conv/T PZ | Touchdowns per target inside the 10 yardline |
%Team PZ | Percent of team's targets inside the 10 yardline |
Rec <=0 | Percent of targets that go for negative yards or no net gain |
Rec >=10 | Percent of targets that go for 10+ yards |
Rec >=25 | Percent of targets that go for 25+ yards |
10 to 0 | Ratio of Rec>=0 to Rec<=0 |
Includes the top 300 players by targets
xxxx | xxxx |
...
Includes players with a significant number of attempts
NEPA | "Net Expected Points Added": (expected points after play - expected points before play)-(opponent's expected points after play - opponent's expected points before play). Uses the expected points for the current possession and the opponent's next possession based on down, distance and spot |
NEPA/PP | Average NEPA per play |
Max/Min | Single game high and low |
Includes players with a significant number of attempts
NEPA | "Net Expected Points Added": (expected points after play - expected points before play)-(opponent's expected points after play - opponent's expected points before play). Uses the expected points for the current possession and the opponent's next possession based on down, distance and spot |
NEPA/PP | Average NEPA per play |
Max/Min | Single game high and low |
Adjusted | Reports the per game EPA adjusted for the strength of schedule. |
Defensive Possession Stats
Points/Poss | Offensive points per possession |
EP3 | Effective Points per Possession |
EP3+ | Effective Points per Possession impact |
Plays/Poss | Plays per possession |
Yards/Poss | Yards per possession |
Start Spot | Average starting field position |
Time of Poss | Average time of possession (in seconds) |
TD/Poss | Touchdowns per possession |
TO/Poss | Turnovers per possession |
FGA/Poss | Attempted field goals per possession |
%RZ | Red zone trips per possession |
Points/RZ | Average points per red zone trip. Field Goals are included using expected points, not actual points. |
TD/RZ | Touchdowns per red zone trip |
FGA/RZ | Field goal attempt per red zone trip |
Downs/RZ | Turnover on downs per red zone trip |
Defensive Play-by-Play Stats
EPA/Pass | Expected Points Added per pass attempt |
EPA/Rush | Expected Points Added per rush attempt |
EPA/Pass+ | Expected Points Added per pass attempt impact |
EPA/Rush+ | Expected Points Added per rush attempt impact |
Yards/Pass | Yards per pass |
Yards/Rush | Yards per rush |
Yards/Pass+ | Yards per pass impact |
Yards/Rush+ | Yards per rush impact |
Exp/Pass | Explosive plays (25+ yards) per pass |
Exp/Rush | Explosive plays (25+ yards) per rush |
Exp/Pass+ | Explosive plays (25+ yards) per pass impact |
Exp/Rush+ | Explosive plays (25+ yards) per rush impact |
Comp% | Completion percentage |
Comp%+ | Completion percentage impact |
Yards/Comp | Yards per completion |
Sack/Pass | Sacks per pass |
Sack/Pass+ | Sacks per pass impact |
Sack/Pass* | Sacks per pass on passing downs |
INT/Pass | Interceptions per pass |
Neg/Rush | Negative plays (<=0) per rush |
Neg/Run+ | Negative plays (<=0) per rush impact |
Run Short | % Runs in short yardage situations |
Convert% | 3rd/4th down conversions |
Conv%* | 3rd/4th down conversions versus average by distance |
Conv%+ | 3rd/4th down conversions versus average by distance impact |
Offensive Play-by-Play Stats
Plays | Number of offensive plays |
%Pass | Percent pass plays |
EPA/Pass | Expected Points Added per pass attempt |
EPA/Rush | Expected Points Added per rush attempt |
EPA/Pass+ | Expected Points Added per pass attempt adjusted for competition |
EPA/Rush+ | Expected Points Added per rush attempt adjusted for competition |
Yards/Pass | Yards per pass |
Yards/Rush | Yards per rush |
Yards/Pass+ | Yards per pass adjusted for competition |
Yards/Rush+ | Yards per rush adjusted for competition |
Exp Pass | Explosive plays (25+ yards) per pass |
Exp Run | Explosive plays (25+ yards) per rush |
Exp Pass+ | Explosive plays (25+ yards) per pass adjusted for competition |
Exp Run+ | Explosive plays (25+ yards) per rush adjusted for competition |
Comp% | Completion percentage |
Comp%+ | Completion percentage adjusted for competition |
Sack/Pass | Sacks per pass |
Sack/Pass+ | Sacks per pass adjusted for competition |
Sack/Pass* | Sacks per pass on passing downs |
Int/Pass | Interceptions per pass |
Neg/Run | Negative plays (<=0) per rush |
Neg/Run+ | Negative plays (<=0) per rush adjusted for competition |
Run Short | % Runs in short yardage situations |
Convert% | 3rd/4th down conversions |
Conv%* | 3rd/4th down conversions versus average by distance |
Conv%+ | 3rd/4th down conversions versus average by distance adjusted for competition |
Offensive Possession Stats
Points/Poss | Offensive points per possession |
EP3 | Effective Points per Possession |
EP3+ | Effective Points per Possession adjusted for competition |
Plays/Poss | Plays per possession |
Yards/Poss | Yards per possession |
Start Spot | Average starting field position |
Time of Poss | Average time of possession (in seconds) |
TD/Poss | Touchdowns per possession |
TO/Poss | Turnovers per possession |
FGA/Poss | Attempted field goals per possession |
Poss/Game | Possessions per game |
%RZ | Red zone trips per possession |
Points/RZ | Average points per red zone trip. Field Goals are included using expected points, not actual points. |
TD/RZ | Touchdowns per red zone trip |
FGA/RZ | Field goal attempt per red zone trip |
Downs/RZ | Turnover on downs per red zone trip |
PPP | Points per Possession |
aPPP | Points per Possession allowed |
PPE | Points per Exchange (PPP-aPPP) |
EP3+ | Expected Points per Possession |
aEP3+ | Expected Points per Possession allowed |
EP2E+ | Expected Points per Exchange |
EPA/Pass+ | Expected Points Added per Pass |
EPA/Rush+ | Expected Points Added per Rush |
aEPA/Pass+ | Expected Points Allowed per Pass |
aEPA/Rush+ | Expected Points Allowed per Rush |
Exp/Pass | Explosive Plays per Pass |
Exp/Rush | Explosive Plays per Rush |
aExp/Pass | Explosive Plays per Pass allowed |
aExp/Rush | Explosive Plays per Rush allowed |
BPR | A method for ranking conferences based only on their wins and losses and the strength of schedule. See BPR for an explanation. |
Power | A composite measure that is the best predictor of future game outcomes, averaged across all teams in the conference |
P-Top | The power ranking of the top teams in the conference |
P-Mid | The power ranking of the middling teams in the conference |
P-Bot | The power ranking of the worst teams in the conference |
SOS-Und | Strength of Schedule - Undefeated. Focuses on the difficulty of going undefeated, averaged across teams in the conference |
SOS-BE | Strength of Schedule - Bowl Eligible. Focuses on the difficulty of becoming bowl eligible, averaged across teams in the conference |
Hybrid | A composite measure that quantifies human polls, applied to converences |
Player Game Log
Use the yellow, red and green cells to filter values. Yellow cells filter for exact matches, green cells for greater values and red cells for lesser values. By default, the table is filtered to only the top 200 defense-independent performances (oEPA). The table includes the 5,000 most important performances (positive and negative) by EPA.
Use the yellow, red and green cells to filter values. Yellow cells filter for exact matches, green cells for greater values and red cells for lesser values. By default, the table is filtered to only the top 200 defense-independent performances (oEPA). The table includes the 5,000 most important performances (positive and negative) by EPA.
EPA | Expected points added (see glossary) |
oEPA | Defense-independent performance |
Team Game Log
Use the yellow, red and green cells to filter values. Yellow cells filter for exact matches, green cells for greater values and red cells for lesser values.
Use the yellow, red and green cells to filter values. Yellow cells filter for exact matches, green cells for greater values and red cells for lesser values.
EP3 | Effective points per possession (see glossary) |
oEP3 | Defense-independent offensive performance |
dEP3 | Offense-independent defensive performance |
EPA | Expected points added (see glossary) |
oEPA | Defense-independent offensive performance |
dEPA | Offense-independent defensive performance |
EPAp | Expected points added per play |
No comments:
Post a Comment