Friday, February 4, 2022

SCHEDULING RESOURCES

 I have two scheduling resources available for teams, both in the form of Excel workbooks:

Team Histories and Simulated 2022 Ranks

This is a resource that lets you review, for each team:

Its RPI rating and rank history, plus its Massey rank history;

Its history in terms of its rank as a contributor to opponents’ strengths of schedule;

Its history in terms of the average ranks of its opponents.

It also includes a trended RPI rating and rank I have assigned each team going into the 2022 season, based on its rank history, as well as its expected rank as a contributor to opponents’ strengths of schedule.

The workbook has a User Guide that explains in detail what is in the workbook and how to use it as part of your scheduling process.

This is a very large workbook, with a page for each team as well as additional resource pages.  Because of its size, it is not practical to use it on line.  Rather, you will need to download it using the following link: Team Histories and Simulated 2022 Ranks.  I have found that it takes the following steps to download it:

When you have clicked on the link, a page will come up in your browser with a rotating circle.  You may have to wait a while because of the size of the workbook.

A black page will come up with a rotating circle.  On the upper right, there will be a download icon.  Click on the icon (even if the circle still is rotating).  Again, you may have to wait a while.

A black page may come up again with a rotating circle.  You also may see some of the text of the User Guide page of the workbook.  Wait until a download box appears at the bottom left.  Click on the carat mark to open the box.  Then click on Open.

At this point, the Excel file should open.  Use Save As to save the file to your computer with whatever name you want to assign it.

If you have trouble downloading the file, email me at cpthomas@q.com and I will email a copy to you.

Team Scheduling Tool

This is a Tool custom set up for each team that wants one.  It lets the team try out different schedules to see where the team is likely to end up in terms of its RPI rating and rank and also in terms of several factors related to NCAA Tournament at large selections and seeds.  It contains a User Guide that explains what is in the Tool and how to use it.

I also can set the Tool up for use by a conference to see how scheduling by the individual teams affects the conference teams’ RPI ratings and ranks and NCAA Tournament prospects.

I will set up a tool for any team that wants one.  If you want one, simply email me at cpthomas@q.com.  In order to set up the tool, it will work best for me and for you if you will send me your conference’s entire in-conference schedule as well as your proposed non-conference schedule (I do not need game dates, but I do need game locations).  Once I have the needed information, I will set up the Tool and send it to you.

 

 

Friday, January 7, 2022

GRADING THE COMMITTEE ON ITS NCAA TOURNAMENT SEED AND HOME FIELD DECISIONS

Each year, in forming the NCAA Tournament bracket, the Women’s Soccer Committee selects 4 #1 seeds, 4 #2s, 4 #3s, and 4 #4s and places them in the bracket accordingly.

If we think of the #1 seeds as 1.1, 1.2, 1.3, and 1.4, it seems clear the Committee places 1.1 at the top left of the bracket.  After that, although it is not always clear, it appears to place 1.2 at the top right, 1.3 at the bottom right, and 1.4 at the bottom left.  (Over the last 11 seasons, but excluding 2020, the 1.1 top left have won 44 Tournament games, the 1.2 top right 41, the 1.3 bottom right 41, and the 1.4 bottom left 36.  It is possible the 1.2 and 1.3 positions are reversed from what I have indicated.)

As for the #2, #3, and #4 seeds, it is not clear to me how the Committee places them.  For purposes of this article, it does not matter.

In addition, in the Tournament first round, there are 16 games between unseeded opponents.  In these games, the Committee appears to give home field to the teams it considers stronger.

When the Committee makes these decisions, it is supposed to be free from bias, both in terms of the regions teams come from and the conferences they play in.  But is it really free from bias?  When I looked at the results of the 2021 Tournament, it occurred to me that this might be a good question to research.  So I did.

My approach was to see how many rounds regions’ and conferences’ teams actually won in each of the last ten Tournaments as compared to how many the Committee, based on its seed and host decisions, thought they would win.  In terms of how many games the Committee decisions say teams should win, here are the numbers:

#1.1  6 games (champion)

#1.2  5 games (runner up)

#1.3 and 1.3  4 games (semi-finalists)

#2s  3 games (quarterfinalists)

Unseeded pairs, 1st round host  1 game (get to second round)

Unseeded pairs, 1st round visitor  0 games (lose in first round)

Unseeded, playing seed in 1st round   0 games (lose in first round)

Regions

This year, I am changing my approach to regions by using four strictly geographic regions based on where a state’s teams play the greatest numbers of their games.  (Given the frequency of teams shifting conferences, this seems a better long term approach for studying regions than using conferences as the building blocks for regions.)

Middle: Illinois, Indiana, Iowa, Michigan, Minnesota, Missouri, Nebraska, North Dakota, Ohio, South Dakota, Wisconsin (64 teams)

North: Connecticut, Delaware, Maine, Maryland, Massachusetts, New Jersey, New York, Pennsylvania, Rhode Island, Vermont, Washington DC (79 teams)

South: Alabama, Arkansas, Florida, Georgia, Kansas, Kentucky, Louisiana, Mississippi, North Carolina, Oklahoma, South Carolina, Tennessee, Texas, Virginia, West Virginia (139 teams)

West: Arizona, California, Colorado, Hawaii, Idaho, Montana, Nevada, New Mexico, Oregon, Utah, Washington, Wyoming (62 teams).

For regions, I look at each region team’s bracket position to see how many games the positioning says it should win.  Then I look at actual results to see how many games it actually won.  From the games it actually won, I subtract the number the Committee positioning says it should have won.  If it did better than expected, it will have a positive number result.  If it did as expected, it will have a 0 result.  If it did poorer than expected, it will have a negative number result.

If two teams from the same region play each other, then either each will perform as expected and both will get 0 results from the game.  Or, neither will perform as expected and one will get a +1 and the other a -1 result, with these numbers canceling each other out from a region perspective.  Thus when I add up all the numbers for a region’s teams, the total will reflect how the region did in games against other regions.

The following table summarizes how the regions actually did as compared to how their Tournament positioning says they should have done:

The table gives one kind of look at the data.  The totals suggest that the Committee decisions, for whatever reason, have been wrongly biased in favor of the South region and against the North and Middle regions and, to a lesser extent, the West region.

A different kind of look considers how the numbers have trended over time:


In this chart, 2011 is on the left and 2021 on the right.  The circle markers are the region data points, connected by the solid lines.  The dotted lines are the trends (straight line trends).  The trend lines suggest that the Committee has slightly underrated the Middle region (blue) over time, but been pretty consistent about it.  In the past the Committee slightly overrated the North region (red) but that has trended downward and now is right about where it should be.  The Committee started the decade rating the South (grey) about right, but is trending towards significantly overrating it.  And conversely, the Committee started the decade overrating the West (gold), but is trending towards fairly significantly underrating it.

You can decide for yourself how seriously to take what the table and chart show.  The numbers, however, are correct.

Conferences

For conferences, the process is the same.  A problem, however, is that teams have changed conferences during the 2011 to 2021 period.  I did not want to be moving teams from one conference to another, so I have treated teams as though they were in their current conference throughout the period.


Here, the totals make it look like the Committee has been biased in favor of the American, Big 12, Pac 12, and SEC conferences in particular and against the Big 10 and Colonial, and especially the West Coast conference in particular.

The trends, however, give a different look:


The chart covers only the conferences that have had + or - results in at least half of the years: the Power 5 plus the American, Big East, and West Coast.  I have eliminated the lines connecting the data points to make it easier to see the trend lines.

Since the colors are a little hard to read, I will start at the top left of the chart and move down through the trend lines and what they suggest:

  • For the ACC (blue), the Committee has gone from significantly underrating it to significantly overrating it.
  • For the Big 10 ( black), the Committee has slightly underrated it pretty consistently.
  • For the Pac 12 (gold), the Committee started out rating it about right and is trending towards slightly overrating it.
  • For the Big East (green), the Committee consistently has rated it about right.
  • For the American (grey), the Committee started out rating it about right and is trending towards slightly overrating it.
  • For the West Coast (dark blue), the Committee has gone from overrating it to significantly underrating it.
  • For the SEC (dark green), the Committee has gone from underrating it to rating it about right.
  • For the Big 12 (purple-brown), the Committee has gone from underrating it to slightly underrating it.
Here again, you can decide for yourself how seriously to take what the table and chart show.  The numbers, however, are correct.

Conclusion

From both region and conference perspectives, the data suggest the possibility that, in seeding and in assigning home field to first round unseeded pairings, the Committee may have region and conference biases that are wrongly influencing its decisions.  Although this is just a possibility, perhaps it is something the Committee should bear in mind in its future bracket formation decisions.



Thursday, December 30, 2021

NCAA TOURNAMENT: ANALYZING THE COMMITTEE SEED AND AT LARGE DECISIONS

Each year, I analyze the Committee seed and at large selection decisions in relation to the Committee’s historic decision patterns.  This year, I have set up a series of tables for the analysis.  I will start with the #1 seed table, giving a detailed explanation of its data.  Then, I will show the tables for each of the #2, #3, and #4 seeds and for the at large selections, with a few comments.

For each table, you will need to scroll to the right to see the entire table.

#1 Seeds

Here is the #1 seed table:


First, some general background to help with the table.

For at large selections, the NCAA has set specific data factors from the season that the Committee must use and is limited to using.  The NCAA leaves it up to each Committee member to decide how much weight to assign to each factor.  The Committee also uses the factors in the seeding process, but for seeding they are not mandatory and the Committee is not limited to the factors in evaluating teams.

I break the NCAA factors down into 13 individual factors:
  • RPI (adjusted)
  • RPI Rank
  • Non-Conference RPI (adjusted)
  • NCRPI Rank
  • Top 50 Results (my modification of an NCAA factor)
  • Top 50 Results Rank
  • Conference Standing (I use average of regular season standing and conference tournament finishing position)
  • Conference RPI
  • Conference RPI Rank
  • Head to Head Results Against Top 60 Opponents (my modification of an NCAA factor)
  • Results Against Common Opponents with Other Top 60 Teams (my modification of an NCAA factor)
  • Common Opponent Results Rank (my modification of an NCAA factor)
  • Poor Results (my modification of an NCAA factor)
There are NCAA data available that allow an evaluation of a team for each factor.  For some of the factors, the NCAA has a scoring system.  For example, an NCAA formula assigns a value for the RPI.  Where the NCAA does not have a scoring system for a factor, I have created my own scoring system.

In addition, to mimic how a Committee member might think, I pair each factor with each other factor.  For each factor pair, I have a scoring system that gives each factor a 50% weight.

In addition, I have one other factor, Number of Games Against Top 60 Opponents.  This is not an NCAA mandated factor but rather is one I use as an aid to teams wishing to do their non-conference scheduling with a view towards the NCAA Tournament.

Altogether this produces a total of 92 individual and paired factors.

For each Top 60 team for each year since 2007 (the first year I began collecting data), my computer program has computed a score for each of the 92 factors.  (Hereafter, I will refer to my computer program and process as my "system.")  My system then compares the scores of all of the Top 60 teams since 2007 to the Committee seeding and at large selection decisions.  From this comparison, the system identifies two scores for each factor:

A "Yes" score, which means that for a particular Committee decision, if a team has had the Yes score or better for that factor, the team always has gotten a favorable decision from the Committee.  For example, if a team has had an RPI Rank of 1, it always has gotten a #1 seed.  I call such a Yes score the Yes standard for that factor.  Thus <=1 is the RPI Rank Yes standard for a #1 seed.

A "No" score, which means that for a particular Committee decision, if a team has had the No score or poorer for that factor, the team never has gotten a favorable decision from the Committee.  For example, if a team has had an RPI Rank of 8 or poorer, it never has gotten a #1 seed.  Thus >=8 is the RPI Rank No standard for a #1 seed.

Note:  As the data turn out, there are some factors, for some Committee decisions, that do not have a Yes standard or that do not have a No standard.  For example, a team’s Conference Standing does not have a Yes standard for any Committee decision.  In other words, your conference standing, all by itself and without reference to what conference you are in, will not assure you of any seed position or of an at large selection.

On completion of the regular season including conference tournaments, my system tallies up all of the Yes and No standards a team has met, for each Committee decision -- #1, 2, 3, and 4 seeds and at large selections.  For each team, for each Committee decision, there are four possible outcomes:

  • The team meets one or more Yes standards and no No standards.  This means that if the Committee follows its historic pattern, the team will get a Yes decision from the Committee.
  • The team meets no Yes standards and one or more No standards.  This means that if the Committee follows its historic pattern, the team will get a No decision from the Committee.
  • The team meets no Yes standards and no No standards.  This means that either a Yes or a No decision from the Committee will be consistent with its historic pattern.
  • The team meets one or more Yes standards and one or more No standards.  This means that the team has a profile the Committee has not seen historically.  Whatever decision the Committee makes cannot be fully consistent with its historic pattern.
With that background, the above table for #1 seeds shows data related to each of the teams with RPI ranks #1 through 7.  This is the candidate group for #1 seeds, since the No standard for a #1 seed is >=8.

Committee Decision: Green means the Committee gave the team a #1 seed.  Red means it did not.

RPI Rank 

Top 50 Results Rank:  I have included this in the table because historically the factor pair of RPI Rank and Top 50 Results Rank has proved to be a good predictor of Committee decisions, especially for at large selections.

Yes Standards Met and No Standards Met:  This is the number of Yes and No standards for #1 seeds that the team has met.

Yes Standard: If a team has met one or more Yes standards but does not get a Yes decision from the Committee, it is useful to know what the Yes standard is, so I list it here.  For example, Florida State had an RPI Rank of #1.  Suppose the Committee had not given it a #1 seed.  Then in this column you would have seen 2 RPI Rank (the 2 preceding RPI Rank simply is a number I have assigned to that standard).

Yes Value:  If I have listed a Yes Standard, I will state the standard score here.  For example, for a #1 seed, the RPI Rank Yes standard is <=1, which means that teams with RPI ranks of #1 always have gotten #1 seeds.  If the Committee had not given Florida State a #1 seed, you would have seen <=1 in this column.

Yes Actual:  If I have listed a Yes standard, I also state the team’s actual score for the standard.  So, if the Committee had not given Florida State a #1 seed, you would have seen 1 in this column representing its RPI rank.

No Standard.  If a team has met one or more No standards but gets a Yes decision from the Committee, it is useful to know what the No standard is, so I list it here.  For example, the RPI Rank No standard for a #1 seed is >=8.  If the Committee had given a #1 seed to the #8 RPI team, I would have listed 2 RPI Rank here.

No Value. If I have listed a No standard, I will state the standard score here.    If the Committee had given the RPI #8 team a #1 seed, you would have seen >=8 in this column.

No Actual:  If I have listed a No standard, I also state the team’s actual score for the standard.  So, if the Committee had given the #8 team a #1 seed, you would have seen 8 in this column.

Teams Affected:  If the Committee has given a No to a team that has met a Yes standard or a Yes to a team that has met a No standard, this column will give an indication of how significant a change the Committee has made from its historic pattern for that factor.  For example, since 2007 through 2019 there were 13 teams with #1 RPI ranks and they all received #1 seeds.  If the Committee had not given Florida State a #1 seed this year, then you would have seen 13 in the Teams Affected column.  This would mean that there are 13 teams that, based on past history, we would have thought assured of #1 seeds but that, based on the Committee decision this year, no longer could be considered as having been assured of #1 seeds.  The lower the Teams Affected number, the smaller the Committee change from its historic pattern.  If the Teams Affected number is 0, it means the team has a score for the factor and Committee decision that is just next to the historic standard and that the Committee has not seen before so that the Committee decision simply represents a refinement of the previous standard.

As a further note about Teams Affected, to give the numbers some context:

  • The #1 seed candidate range is teams with RPI Ranks of #7 or better, so the number of candidates for #1 seeds since 2007 has been 13 x 7 = 91.  So when you are looking at a Teams Affected number for #1 seeds, it is that number of teams out of a total of 91. 

  •  The #2 seed candidate range is RPI Ranks of #14 or better, so the number of #2 seed candidates has been 13 x 14, less the 52 teams that got #1 seeds, which amounts to a #2 seed candidate pool of 130 teams. 

  •  The #3 seed candidate range is RPI Ranks of #23 or better, so the number of #3 seed candidates has been 13 x 23, less the 104 teams that got #1 and 2 seeds, which amounts to a #3 seed candidate pool of 195 teams. 

  • The #4 seed candidate range is RPI Ranks of #26 or better, so the number of #4 seed candidates has been 13 x 26, less the 156 teams that got #1 through 3 seeds, which amounts to a #4 seed candidate pool of 182 teams. 

  • The At Large candidate range is RPI Ranks of #57 or better, less the 208 seeded teams, the Automatic Qualifiers in the Top 57, and teams in the Top 57 that failed to meet the 0.500 minimum winning record requirement.  Altogether since 2007, this has amounted to 447 teams. 

 Round Eliminated:  This column shows the NCAA Tournament round this year in which the team was eliminated.  It lets you look at the position the NCAA assigned the team and see how the team did in relation to that assigned position.  For example, Virginia’s #1 seed means that according to the Committee it should have made it at least to the semifinals, but instead it made it only to the 3rd round.  This lets you evaluate how the Committee decisions worked out. (For first round matchups between unseeded teams, I treat the home team as the stronger team according to the Committee.)

With the above explanation, I leave it to you to go up to the #1 Seeds table and see how the Committee decisions match up with its historic pattern.  My comment is that the Committee did not deviate from its historic pattern except when it had to with Duke and Virginia due to their meeting both Yes and No standards and that for those teams, the Committee deviation was small.

#2 Seeds

Here is the #2 seed table:

You can review the table and reach your own conclusions.  My comment is that in giving UCLA a #2 seed, the Committee deviated from its historic pattern.  The deviation was not large but also was not insignificant.  As an alternative, Tennessee would have been an easy #2 seed.

#3 Seeds

Here is the #3 seed table:


My comment is that there is nothing of major import in the Committee decisions.

#4 Seeds

Here is the #4 seed table:


There is only one significant Committee deviation from its historic patterns here, and it is the #4 seed given to BYU.  This was a pretty large deviation.  Ironically, BYU reached the championship game.

 At Large

Here is the At Large table:


My comment is that the only deviation here from Committee historic patterns is in its giving St. Johns an at large position rather than West Virginia, Colorado, Oregon, or Houston.  In looking at St. Johns’ Teams Affected numbers, however, the deviation was extremely small.

Summary and Two Additional Pieces of Information

My evaluation of the Committee decisions in relation to historic Committee patterns suggests that:

1.  The Committee At Large selections were quite consistent with historic patterns and, where it varied with St. Johns, the variation was very small.

2.  The Committee seeds were largely consistent with historic patterns.  Where the Committee varied from historic patterns, most of the variations were small.  The greatest variation was BYU getting seeded, which is ironic since BYU made it to the championship game.

In addition to the above, I have looked at two other aspects of the Committee decisions.

First, I looked at geographic regions based on the states where teams are located.  As it turned out this year, during the season teams from states in the West played roughly 90% of their games against other teams from the West and only 10% against teams from other regions.  This created a big problem for the RPI since 10% is not enough games for the RPI to properly rank teams from a region in relation to teams from other regions.  Because of this, I wanted to see if teams from the West performed differently in the Tournament than the Committee had evaluated them.  The following table addresses this question:


In this table, the Wins Difference column shows, for each region, the difference between (1) the number of games the Committee seeds and bracket placements for unseeded teams indicated teams should win and (2) the number of games teams actually won.  The numbers in this table represent how teams from a region did against teams from other regions, since all within-region games cancel each other out.  Although the numbers in the table are not large, they suggest that the Committee may have over-evaluated teams from the South and under-evaluated teams from the other regions and particularly the West.

Second, I took a similar look, but this time by conference.


This suggests that the Committee may have undervalued the West Coast Conference and, to a much lesser extent, the Big 10 and overvalued the ACC and Pac 12.

Since both of these tables are based on only one year’s results, I do not take them too seriously.  They suggest, however, that it might be worthwhile to do a study that considers more years of NCAA Tournaments, to see if there are any Committee region- or conference-based overvaluation and undervaluation patterns.

Sunday, November 7, 2021

END OF REGULAR SEASON NCAA TOURNAMENT BRACKET SIMULATION 11.7.21

[Note:  In the below table, Xavier should be a #4 seed rather than Auburn.  Also, Santa Clara should be a 5 AQ rather than a 6.  And Further Note:  On relooking at why I originally had Auburn as a #4 seed, I realized the system I use has it as a clear Yes for a #4 seed, so it indeed should be a #4 seed and not Xavier.]

Here are my simulated NCAA Tournament seeds and at large selections, based on my more complex system described here.  The table also includes the Automatic Qualifiers.

1 = #1 seed, 2 = #2 seed, 3 = #3 seed, 4 = #4 seed, 5 = unseeded automatic qualifier, and 6 = unseeded at large selection.

Below the table, I will indicate next teams in line for seeds and at large selections.


Other possible seeds:

#1 seed:  Virginia, Arkansas

#2 seed: Georgetown

#3 seed: Princeton

#4 seed: Mississippi, Hofstra, Harvard

At Large:

Last in: NC State, Santa Clara, Providence, Wisconsin, West Virginia, Oregon

Next in line: Alabama, Colorado, Houston, St. Johns, Indiana

Monday, November 1, 2021

2021 RPI: 10.31.21 RPI RATINGS (ACTUAL CURRENT AND SIMULATED END OF SEASON), AND SIMULATED NCAA TOURNAMENT AT LARGE SELECTIONS AND SEEDS

  This week’s reports use actual results of games played through Sunday, October 31.  They are:

1.  Actual current RPI ratings and ranks, showing which teams are in the current ranges for potential seeds and at large selections;

2.  Simulated RPI ratings and ranks based on the actual results of games played and simulated results of games not yet played.  The simulated results are based on opponents’ actual current RPI ratings.  This report includes simulated NCAA Tournament at large selections and seeds based on the simple system described here.

3.  Simulated NCAA Tournament bracket based on the simulated RPI ratings and ranks, using the more complex system described here.

For the tables below, you may need to scroll to the right to see the entire table. 

1.  Actual current RPI ratings and ranks, showing which teams are in the current ranges for seeds and at large selections:

Here is a link to an Excel workbook that shows current RPI and other information for all teams.

In addition, here is a table from the workbook.  On the left, it shows which teams, based on past history, are in the current seed and at large selection ranges as of this stage of the season.  It also includes the next group of teams, that appear to be close but out of the range for an at large selection.  If you look at the At Large Bubble column, the highest ranked teams at the top of the list that are not color coded likely are assured of getting at large selections even if not conference automatic qualifiers, based on past history.


2.  Simulated RPI ratings and ranks based on the actual results of games played and simulated results of games not yet played:

The simulated results of games not yet played are based on opponents’ actual current RPI ratings, as adjusted for home field advantage.  This report includes simple-system simulated NCAA Tournament at large selections and seeds.  [NOTE:  I have more confidence in the more complex system simulated Tournament selections and seeds described in part 3 below.]

Here is a link to an Excel workbook that shows the information for all teams.

In addition, below is a table from the workbook that shows simulated simple-system NCAA Tournament at large selections and seeds.

I have put the table in RPI rank order this week for a particular reason.  The table includes (1) the Top 57 RPI teams, since all at large selections since 2007 have come from the Top 57, plus (2) the Automatic Qualifiers.  The simple system selects unseeded Automatic Qualifiers based on a formula that combines team RPI ranks and their ranks based on their good results against Top 50 opponents, with those two factors weighted equally.  The Top 50 results ranks are based on a scoring system that is very highly skewed towards good results (wins or ties) against very highly ranked opponents.  The simple system uses this dual factor because on average, over the years since 2007 (excluding 2020), the dual factor rank correctly matches all but 2 per year of the Committee unseeded at large selections.

In the At Large Selections column, the color coding shows the teams to which the simple system assigns at large selections.  If a team is neither an Automatic Qualifier nor color coded in the At Large Selections column, it is a team that is among the Top 57 but that the simple system does not assign an at large selection.  In the Seed columns, the color coding shows the teams that the system assigns seeds (with the seeds based on RPI ranks).

The table shows an interesting anomaly:  The simple system assigns Harvard a #3 seed (as the #11 RPI team -- which has dropped to #12 as of November 2), but does not assign it an at large selection since it falls too far down on the dual factor list due to having no good Top 50 results (wins or ties).  In response to this and because it will have educational value, here is a detailed analysis of Harvard’s record.

First, I looked to see whether the differences between Harvard’s RPI rating and its opponents’ ratings seemed appropriate based on game results.  One of the things my system does is compute the RPI rating difference between opponents as adjusted for home field advantage.  Then, based on that difference and historic data, it computes the likelihood of the higher rated team winning, tieing, and losing the game.  If I then combine all of those likelihoods for a team’s schedule, I can tell what the ratings say the team’s win-tie-loss results should be for those games if the ratings for the teams are correct in relation to each other.  I did this first for Harvard’s non-conference games and the RPI ratings say that Harvard’s record over the course of those games should have been 7 wins, 1 tie, and 0 losses (or possibly 7-0-1).  In fact, its actual record was 7-1-0, exactly what it should have been if it and its opponents are rated correctly in relation to each other.  I did this next for Harvard’s conference games, where the ratings say its record (so far) should be 5-0-1, whereas it actually is 4-0-2.  What this suggests is that Harvard is rated appropriately in relation to the non-conference teams it played but is overrated in relation to its Ivy League partners Brown and Princeton.

Second, I looked at Harvard’s results against opponents it had in common with other current Top 60 teams.  This showed the following, based on current RPI ranks as of November 2:

Harvard: win home v #60 St Johns who had a tie home v #30 Butler and a win home v #34 Providence

Harvard: win away v #106 Northeastern who had a win away v #25 Hofstra

Harvard: win home v #91 Kansas who had a win home v #45 West Virginia

Harvard: win home v #111 Penn who had a win home v #54 Rice

Harvard: loss home v #13 Brown who had a loss away v #25 Hofstra, a loss away v #15 Notre Dame, and a loss away v #34 Providence

Harvard: loss home v #18 Princeton, who had a tie away v #14 Georgetown and a loss home v #25 Hofstra

Harvard: win home v #128 Dartmouth who had a tie away v #14 Georgetown

Looking at the common opponent results as a whole, they suggest a Harvard rank in the vicinity of #25 Hofstra to #34 Providence, possibly closer to the Hofstra side of that range.

Putting these two detailed looks at Harvard together, it appears that Harvard’s current RPI rating and rank are too high.  Realistically, it probably should be outside the range for a seed.  On the other hand, it seems well inside the range for an at large selection.

The above analysis can give some insight into the process the Committee must go through.

Interestingly, my more complex system, as shown in the last table below, does not seed Harvard but gives it an at large selection.  An at large selection certainly would be consistent with the history of Committee decisions, which always have given at large selections to teams with RPI ranks of #30 or better.

3.  Simulated NCAA Tournament bracket based on the simulated RPI ratings and ranks, using the more complex system:

Finally, below is a table that shows the simulated more-complex-system NCAA Tournament at large selections and seeds.

Of the at large teams, the last teams in are Butler, Wisconsin, Houston, and West Virginia.  The next teams in line are Oregon, Alabama, and Colorado.

1 = #1 seed, 2 = #2 seed, 3 = #3 seed, 4 = #4 seed, 5 = unseeded automatic qualifier, and 6 = unseeded at large selection.



Monday, October 25, 2021

2021 RPI: 10.24.21 RPI RATINGS (ACTUAL CURRENT AND SIMULATED END OF SEASON), AND SIMULATED NCAA TOURNAMENT AT LARGE SELECTIONS AND SEEDS

 This week’s reports use actual results of games played through Sunday, October 24.  They are:

1.  Actual current RPI ratings and ranks, showing which teams are in the current ranges for potential seeds and at large selections;

2.  Simulated RPI ratings and ranks based on the actual results of games played and simulated results of games not yet played.  The simulated results are based on opponents’ actual current RPI ratings.  This report includes simulated NCAA Tournament at large selections and seeds based on the simple system described here.

3.  Simulated NCAA Tournament bracket based on the simulated RPI ratings and ranks, using the more complex system described here.

For the tables below, you may need to scroll to the right to see the entire table. 

1.  Actual current RPI ratings and ranks, showing which teams are in the current ranges for seeds and at large selections:

Here is a link to an Excel workbook that shows RPI and other information for all teams.

In addition, here is a table from the workbook.  On the left, it shows which teams, based on past history, are in the current seed and at large selection ranges as of this stage of the season.  It also includes the next group of teams, that appear to be close but out of the range for an at large selection.  If you look at the At Large Bubble column, the highest ranked teams at the top of the list that are not color coded likely are assured of getting at large selections even if not conference automatic qualifiers, based on past history.

NOTE:  If you closely compare my ratings to those the NCAA has published, you will notice some minor differences.  This is because the October 24 game between Alcorn State and Grambling somehow dropped out of the NCAA data base.  Hopefully, that game will find its way back in.  In addition, the NCAA still has not adjusted the ranges for its two penalty adjustment tiers.  This has no significant effect, but also is causing some rating and ranking differences from mine, at the poorer end of the RPI.

Also, if you compare my ratings and ranks to the AllWhiteKit ratings, you may notice some very small differences.  The rating differences are due to our systems following different rounding conventions.  (For these differences, AllWhiteKit always will have a team rated 0.0001 higher than my rating.)  The ranking differences are due to the fact that when the AllWhiteKit system has teams with equal ratings when rounded to four decimal places, it puts them in alphabetical order.  My system puts teams in order based on calculations to 15 decimal places.  Ordinarily, these differences are inconsequential.


2.  Simulated RPI ratings and ranks based on the actual results of games played and simulated results of games not yet played:

The simulated results of games not yet played are based on opponents’ actual current RPI ratings, as adjusted for home field advantage.  This report includes simple-system simulated NCAA Tournament at large selections and seeds.

Here is a link to an Excel workbook that shows the information for all teams.  (NOTE:  Due to a programming error (by me), the originally linked workbook had the wrong information.  The currently linked workbook has the right information.)

In addition, here is a table from the workbook that shows simulated simple-system NCAA Tournament at large selections and seeds, plus the next teams in the RPI rankings down to #80 (some of which would not meet the NCAA Tournament 0.500 winning percentage requirement for at large selection).  (NOTE:  This is a corrected version of what I posted yesteray.)


3.  Simulated NCAA Tournament bracket based on the simulated RPI ratings and ranks, using the more complex system:

Finally, below is a table that shows the simulated more-complex-system NCAA Tournament at large selections and seeds.  It is worth noting that since I started keeping data in 2007, no team ranked poorer than #57 (using the current RPI formula) has gotten an at large selection.

Of the at large teams, the last teams in are Washington State, Santa Clara, and South Carolina.  The next teams in line are Colorado, Butler, and Michigan State, followed by Oregon State, Georgia, Indiana, and Clemson.

1 = #1 seed, 2 = #2 seed, 3 = #3 seed, 4 = #4 seed, 5 = unseeded automatic qualifier, and 6 = unseeded at large selection.



Tuesday, October 19, 2021

2021 RPI: 10.17.21 RPI RATINGS (ACTUAL CURRENT AND SIMULATED END OF SEASON), AND SIMULATED NCAA TOURNAMENT AT LARGE SELECTIONS AND SEEDS

This week’s reports use actual results of games played through Sunday, October 17.  They are:

1.  Actual current RPI ratings and ranks, showing which teams are in the current ranges for potential seeds and at large selections;

2.  Simulated RPI ratings and ranks based on the actual results of games played and simulated results of games not yet played.  The simulated results are based on opponents’ actual current RPI ratings.  This report includes simulated NCAA Tournament at large selections and seeds based on the simple system described here.

3.  Simulated NCAA Tournament bracket based on the simulated RPI ratings and ranks, using the more complex system described here.

For the tables below, you may need to scroll to the right to see the entire table. 

1.  Actual current RPI ratings and ranks, showing which teams are in the current ranges for seeds and at large selections:

Here is a link to an Excel workbook that shows RPI and other information for all teams.

In addition, here is a table from the workbook.  On the left, it shows which teams, based on past history, are in the current seed and at large selection ranges as of this stage of the season.  If you look at the At Large Bubble column, the highest ranked teams at the top of the list that are not color coded likely are assured of getting at large selections, based on past history.

NOTE:  If you compare these ranks to those the NCAA has published, you will note that I have Duke and Arkansas ranked #2 and #3 respectively, whereas the NCAA has them in the reverse order.  Those two teams have nearly identical ratings and the difference in their order is due only to the NCAA and I using different rounding conventions.  This is an unusual occurence in this area of the ratings and I expect it will disappear in next week’s ratings and ranks.

In addition, if you get into comparing my ratings to the NCAA’s, you might notice that I have #70 Minnesota with an RPI rating of 0.5676 whereas the NCAA has them at 0.5670.  The reason for this is that Minnesota has accrued an RPI penalty for a tie against Illinois-Chicago.  There are two tiers of penalties, the higher penalties being for poor results against the bottom 40 teams and the lower penalties for results against the next-to-bottom 40.  As new schools sponsor soccer each year, it is necessary to adjusted the penalty tiers in the RPI calculation system to match the total number of teams sponsoring soccer.  I have done that, but this year the NCAA has not yet done it.  As a result, the NCAA is treating Minnesota as having accrued a penalty as though its poor result was against a bottom 40 team whereas it actually accrued the result against a next-to-bottom 40 team.  The NCAA has been aware for a couple of weeks that it needs to adjusted its penalty tiers and has said they will do it, but they have not done it yet.  There are 15 other teams, farther down in the rankings, that this likewise affects.  Ordinarily, it would not be a significant issue, but since Minnesota still is within at large selection range as of this stage of the season, it actually may become important that the NCAA make the needed correction.


2.  Simulated RPI ratings and ranks based on the actual results of games played and simulated results of games not yet played:

The simulated results of games not yet played are based on opponents’ actual current RPI ratings, as adjusted for home field advantage.  This report includes simple-system simulated NCAA Tournament at large selections and seeds.

Here is a link to an Excel workbook that shows the information for all teams.

In addition, here is a table from the workbook that shows simulated simple-system NCAA Tournament at large selections and seeds, plus the next ten teams:


3.  Simulated NCAA Tournament bracket based on the simulated RPI ratings and ranks, using the more complex system:

Finally, below is a table that shows the simulated more-complex-system NCAA Tournament at large selections and seeds, plus the next 9 teams.  It is worth noting that since I started keeping data in 2007, no team ranked poorer than #57 (using the current RPI formula) has gotten an at large selection.

1 = #1 seed, 2 = #2 seed, 3 = #3 seed, 4 = #4 seed, 5 = unseeded automatic qualifier, and 6 = unseeded at large selection.