Friday, August 19, 2022

2022 PRE-SEASON ASSIGNED RPI RANKS AND RESULTING SIMULATED END-OF-SEASON RPI RANKS

Before each season, I prepare an Excel workbook that, among other things, assigns pre-season RPI ratings and ranks for all teams based on their historic rank trends.  You can use this link to download the Team Histories and Simulated 2022 Ranks Data Only  workbook.  (One purpose of the workbook is to be a data source for coaches doing non-conference scheduling.)  Once downloaded, the workbook will open to its User Guide page where you can find an explanation of how I assign pre-season ratings and ranks and of what their limitations are.

Once I have all team schedules, I apply the assigned pre-season RPI ratings to each game to project a game result on the assumption the result will be consistent with the two opponent ratings as adjusted for home field advantage.  (Home field advantage is worth 0.0148 in RPI rating points.)  If the location-adjusted rating difference between the two opponents is 0.0133 or less, then the projected game result is a tie.  This is because at that rating difference level, each team has a statistical win likelihood of less than 50%, so it is better to project a tie than a win-loss.  (The higher rated team is more likely to lose or tie the game than win it.)

My system also is set up to populate the conference tournament brackets based on the projected results of the in-conference games, as well as to project each conference tournament game result as described above.

With projected results for the entire season, including conference tournaments, my system then calculates final RPI ratings and ranks for all teams.  One might think that the final ratings and ranks would be about the same as the assigned pre-season ratings and ranks given that all game results are consistent with the assigned pre-season ratings.  In some cases, that is what happens.  In many cases, however, there are significant differences.  This is due in part to how team schedules interact with the RPI formula -- and can give insight into how important it is to do non-conference scheduling with a view to how the RPI formula works, if your team has NCAA Tournament aspirations.  It also is due, however, to the RPI’s patterns of discriminating against (or in favor of) some conferences and geographic regions.

With that background, the following table shows team assigned pre-season RPI ranks followed by their resulting RPI ranks after going through the game results assigning process described above.  You might want to bear in mind that since 2007 no team with an end-of-season RPI rank of 58 or poorer has received an NCAA Tournament at large selection, so it is likely that the end-of-season pool of candidates for an at large selection is, at best, the Top 60.

Team for Simulation Rating Purposes

Assigned Pre Season ARPI Rank

Resulting Final Simulated RPI Rank

Duke

1

1

FloridaState

2

2

ArkansasU

3

4

NorthCarolinaU

4

5

VirginiaU

5

7

Stanford

6

10

Pepperdine

7

17

BYU

8

24

TennesseeU

9

12

SouthernCalifornia

10

3

MichiganU

11

8

Rutgers

12

11

UCLA

13

22

TCU

14

14

NotreDame

15

19

Brown

16

6

Clemson

17

13

WashingtonState

18

21

Georgetown

19

15

SantaClara

20

33

PennState

21

16

Purdue

22

20

VirginiaTech

23

27

Xavier

24

9

Gonzaga

25

54

Princeton

26

18

LSU

27

26

SouthCarolinaU

28

23

ColoradoU

29

45

NCState

30

25

TexasA&M

31

38

OregonU

32

34

TexasU

33

49

Auburn

34

28

WakeForest

35

40

Pittsburgh

36

31

Hofstra

37

29

CaliforniaU

38

48

NebraskaU

39

50

MississippiU

40

30

WisconsinU

41

51

Butler

42

39

Milwaukee

43

84

ArizonaState

44

41

WestVirginiaU

45

78

IndianaU

46

100

Louisville

47

69

WashingtonU

48

60

Vanderbilt

49

36

OhioState

50

55

IowaU

51

117

OklahomaState

52

52

Memphis

53

61

MichiganState

54

58

NorthwesternU

55

68

Harvard

56

32

OregonState

57

96

GeorgiaU

58

73

SouthFlorida

59

93

UtahU

60

87

AlabamaU

61

106

MinnesotaU

62

107

Providence

63

37

SMU

64

46

UCF

65

53

TexasTech

66

82

StJohns

67

57

MississippiState

68

115

LoyolaChicago

69

43

UNCWilmington

70

44

StMarys

71

101

SanFrancisco

72

104

GrandCanyon

73

35

NewMexicoU

74

70

BostonCollege

75

98

SouthernMississippi

76

80

Baylor

77

177

Houston

78

95

Samford

79

42

BowlingGreen

80

47

FloridaU

81

171

MissouriU

82

168

KansasU

83

167

Columbia

84

76

Denver

85

99

SouthDakotaState

86

63

CalPoly

87

122

IllinoisU

88

182

PortlandU

89

128

OldDominion

90

66

ArizonaU

91

161

LongBeachState

92

94

UCIrvine

93

140

Pacific

94

198

OklahomaU

95

136

Cincinnati

96

166

EastCarolina

97

116

VCU

98

56

JamesMadison

99

74

StLouis

100

62

CalStateFullerton

101

143

KentuckyU

102

154

UtahValley

103

102

NorthTexas

104

72

UtahState

105

90

UCSantaBarbara

106

131

SouthAlabama

107

67

Seattle

108

103

PennsylvaniaU

109

114

KentState

110

64

MontanaU

111

137

DePaul

112

141

Marquette

113

149

FresnoState

114

156

Rice

115

65

KansasState

116

199

IowaState

117

217

MiamiFL

118

180

ConnecticutU

119

146

MarylandU

120

192

BoiseState

121

173

Creighton

122

145

SanDiegoU

123

203

Dayton

124

81

SouthDakotaU

125

113

Lipscomb

126

89

Buffalo

127

97

SIUEdwardsville

128

112

UTSA

129

126

SanDiegoState

130

174

Campbell

131

75

UNOmaha

132

201

NorthernColorado

133

159

Northeastern

134

88

Liberty

135

59

UAB

136

139

WesternKentucky

137

108

Monmouth

138

71

CaliforniaBaptist

139

125

Toledo

140

127

FloridaAtlantic

141

176

CentralMichigan

142

133

UCDavis

143

248

Tulsa

144

205

ArkansasState

145

77

UNLV

146

221

Charlotte

147

158

WeberState

148

170

Dartmouth

149

130

Elon

150

109

Drake

151

121

MassachusettsU

152

105

ColoradoState

153

212

OhioU

154

129

Navy

155

91

BallState

156

186

Quinnipiac

157

85

TennesseeTech

158

92

NorthFlorida

159

124

ColoradoCollege

160

227

BostonU

161

120

AirForce

162

223

Towson

163

134

Furman

164

118

Oakland

165

183

Bucknell

166

178

MiddleTennessee

167

196

Villanova

168

209

SanJoseState

169

264

UCSanDiego

170

226

Temple

171

244

NewMexicoState

172

181

William&Mary

173

153

IdahoU

174

252

HighPoint

175

142

LouisianaTech

176

163

MiamiOH

177

179

Niagara

178

110

FloridaGulfCoast

179

132

GeorgiaSouthern

180

164

WyomingU

181

287

CentralArkansas

182

135

NorthwesternState

183

79

Syracuse

184

185

UCRiverside

185

256

GeorgiaState

186

169

WesternIllinois

187

187

LouisianaMonroe

188

162

MurrayState

189

144

WesternMichigan

190

219

NorthernArizona

191

245

WesternCarolina

192

86

CentralConnecticut

193

111

HawaiiU

194

263

Hartford

195

138

Radford

196

152

VermontU

197

83

Drexel

198

190

SetonHall

199

200

Evansville

200

148

DelawareU

201

193

TennesseeMartin

202

202

KennesawState

203

184

Mercer

204

151

StephenFAustin

205

165

LoyolaMD

206

195

RhodeIslandU

207

155

NorthernKentucky

208

210

Valparaiso

209

204

UTEP

210

268

NevadaU

211

319

IndianaState

212

253

Fairfield

213

150

Richmond

214

197

IllinoisState

215

220

UMassLowell

216

119

Binghamton

217

175

EasternKentucky

218

206

LaSalle

219

224

UNCGreensboro

220

194

Lamar

221

123

GeorgeMason

222

216

Colgate

223

215

MissouriState

224

207

Lehigh

225

239

McNeeseState

226

157

UtahTech

227

228

LoyolaMarymount

228

279

Stonehill

229

160

StonyBrook

230

225

Albany

231

147

UALR

232

208

TexasState

233

189

GreenBay

234

218

SacramentoState

235

309

NewHampshireU

236

172

LouisianaLafayette

237

240

EastTennesseeState

238

230

FairleighDickinson

239

191

StJosephs

240

269

Army

241

267

EasternWashington

242

302

Cornell

243

234

Belmont

244

257

SoutheastMissouriState

245

266

CollegeofCharleston

246

231

SELouisiana

247

213

StBonaventure

248

236

CoastalCarolina

249

255

CalStateNorthridge

250

270

Yale

251

229

Siena

252

211

Chattanooga

253

233

UNI

254

265

Citadel

255

222

UMKC

256

262

Iona

257

250

AppalachianState

258

258

IncarnateWord

259

188

GeorgeWashington

260

278

IPFW

261

246

Marist

262

247

IUPUI

263

254

CalStateBakersfield

264

310

PortlandState

265

318

Bellarmine

266

273

Troy

267

242

HoustonBaptist

268

241

NJIT

269

259

OralRoberts

270

272

American

271

283

SamHoustonState

272

237

Duquesne

273

280

WrightState

274

274

EasternMichigan

275

297

Marshall

276

260

Fordham

277

299

Stetson

278

261

JacksonvilleState

279

235

Longwood

280

214

AbileneChristian

281

238

Lafayette

282

317

FIU

283

315

SacredHeart

284

249

Morehead

285

298

MaineU

286

271

Davidson

287

275

NorthDakotaU

288

316

ClevelandState

289

286

NorthernIllinois

290

296

JacksonvilleU

291

291

StThomas

292

308

AustinPeay

293

282

TexasCommerce

294

251

IllinoisChicago

295

314

NorthDakotaState

296

306

Manhattan

297

277

Bryant

298

281

Howard

299

276

SouthernUtah

300

290

Wofford

301

288

Akron

302

300

CharlestonSouthern

303

285

Rider

304

307

EasternIllinois

305

326

YoungstownState

306

325

UMBC

307

311

PrairieViewA&M

308

243

Detroit

309

313

TexasCorpusChristi

310

284

UNCAsheville

311

292

GardnerWebb

312

305

TexasRGV

313

295

Winthrop

314

294

Merrimack

315.5

293

StFrancisBrooklyn

315.5

301

Wagner

317

340

Queens

318.5

304

SouthernIndiana

318.5

332

HolyCross

320

337

Grambling

321

232

RobertMorris

322

324

IdahoState

323

320

MountStMary

324

303

NorthAlabama

325

323

Tarleton

326

312

VMI

327

338

TexasSouthern

328

289

LongIsland

329

336

USCUpstate

330

328

Canisius

331

335

StFrancis

332

333

AlabamaA&M

333

327

StPeters

334

342

Presbyterian

335

322

ChicagoState

336

341

SIUCarbondale

337

329

AlabamaState

338

330

Lindenwood

339

334

NichollsState

340

321

JacksonState

341

339

SouthernU

342

345

DelawareState

343

346

ArkansasPineBluff

344

344

Hampton

345

331

AlcornState

346

348

SouthCarolinaState

347

347

MississippiValley

348

343

 


Monday, April 25, 2022

THE OVERTIME RULE CHANGE: EFFECTS ON THE RPI AND THE NCAA TOURNAMENT BRACKET

Under the recent NCAA soccer rules change, there no longer will be overtime games during the regular season.  All games tied at the end of regular time will be ties.  In addition, for conference tournaments and the NCAA tournament, games tied at the end of regulation will have two 10-minute overtimes followed by kicks from the mark if still tied.  Thus no more games decided by "golden goal."

How Many Games Should We Now Expect to End as Ties?

Since 2007, 10.8% of Division I women’s soccer regular season games have ended in ties. This includes, as ties, conference tournament games decided by kicks from the mark, which the NCAA treats as ties for statistics purposes.  According to the NCAA description of the rule change, since 2013 47% of tied games going to overtime ended up still tied at the end of overtime.  Putting the numbers together, this suggests that roughly 23% of games were tied at the end of regular time.  Thus a good guess is that under the new rule 23% of games will be ties.

It is possible that teams will make tactical adjustments that will affect the percentage of tied games.  It seems unlikely to me that this would have a big effect on the expected 23% of games tied.

How Will This Affect Individual Team Ratings and Ranks?

The change can affect all three elements of a team’s RPI.  Obviously, the change will affect the team winning percentage (Element 1), if the team would have won or lost a game in overtime but instead ends up with a tie.  But in addition, the change will affect the team’s opponents’ winning percentages (Element 2) if those opponents would have won or lost games in overtime but instead end up with ties.  And it likewise will affect the team’s opponents’ opponents’ winning percentages (Element 3).  Thus the change will affect both teams’ winning percentages and their strengths of schedule (Elements 2 and 3).

For illustration, I ran a test for the 2019 season, using the NCAA’s data to change games won or lost during overtime to ties.  Here is the first of several tables based on that test.  It covers the top 10 automatic qualifiers (conference champions) based on RPI ratings and ranks as they would have been under the new rule.  It compares those ratings and ranks to the actual ratings and ranks under the old rule.  The ranks are in the highlighted columns on the right.


As you can see if you look at the top 3 of Stanford, North Carolina, and South Carolina, they have no changes in their own win-loss-tie records yet their RPI ratings changed.  The rating changes are due to their opponents and/or their opponents’ opponents having changes in their win-loss-tie records.  The rating changes do not affect the Stanford and North Carolina ranks.  South Carolina, however, drops from #5 to #6.  This is due to the change in its rating (resulting from changes in its strength of schedule) combined with changes in the rating of the team that moves up to #5 (which happens to be Arkansas) and also is affected by how close those two teams are in the ratings.  This illustrates that there are a lot of moving parts that contribute in varying degrees, depending on the data, to the effect of the rule change on a particular team.

Looking at 2019 from high above, the maximum rank change for any team is 52 positions.  This is for Oral Roberts, which has two games it won in overtime converted to ties.  Oral Roberts is in the middle of the Division I rankings where team’ ratings tend to be much close together than at the top and bottom of the rankings, so that a relatively small rating change can result in a relatively large rank change.  The smallest rank change is 0 positions, as the numbers for Stanford and North Carolina show.  The average rank change is 10.7 positions and the median is 8.

How Will This Affect Conference Ratings and Ranks?

Continuing with the 2019 season as an example, the following table shows conference teams’ average RPI ratings and ranks and how the conferences rank in relation to each other under the two systems.  Although the average ranks of teams in conferences change, there are not major changes in how the conferences rank in relation to each other.

How Will This Affect the NCAA Tournament Bracket?

To show how the change will affect the NCAA Tournament bracket, I will go through a series of similar tables, again for the 2019 season.  After the first table, I will explain what it shows.


This table shows all of the 2019 automatic qualifiers.  That is what the lime green highlighting signifies.  You can look at each one to see how its record changes under the new rule and also how its RPI rank changes.

I will not explain the entire table yet, but an important column is the NCAA Tournament Seed or Selection column.  That column uses a number code for the NCAA Tournament decision the Committee made for that team:

1  #1 seed

2  #2 seed

3  #3 seed

4  #4 seed

5  Unseeded automatic qualifier

6  Top 60 team in the actual 2019 RPI ranks that got an at large selection

7  Top 60 team in the actual 2019 RPI ranks that did not get an at large selection

To consider likely effects of the new rule, it helps to know that historically #1 seeds always have come from teams ranked #7 or better, #2s from teams ranked #14 or better, #3s from teams ranked #23 or better, and #4s from teams ranked #26 or better.  Further, teams ranked #30 or better (that are not automatic qualifiers) always have gotten at large selections.  And teams that are ranked #58 or poorer never have gotten at large selections.  What you want to be looking for are teams that have moved in and out of these ranges as a result of the rule change.

Here is the next table:


This is all of the teams, not automatic qualifiers, that are in the Top 30 under the new rule.  That is what the dark green highlighting signifies.  You will see that in the NCAA Tournament Seed or Selection column, I have highlighted two cells in dark green.  Each cell contains the number 7, which means that the team was in the RPI Top 60 but the Committee did not give it an at large selection.  The cells are for Florida Atlantic and Yale.  If you compare the actual RPI ranks to the OT Only in Tournament Games ranks for these teams, you will see that Florida Atlantic, due to the rule change, moves from #32 to #20 in RPI rank and Yale moves from #37 to #25.  Thus both teams moved from outside the #30 or better historically "protected" area to inside the protected area.  Because of this, it is reasonable to believe that both Florida Atlantic and Yale would have gotten at large selections, which means that two teams that actually did get at large selections would not have gotten them.  (The two teams that move out of the Top 30 are Florida and Louisville, which means they move from protected to the bubble.)

You also will see, in the NCAA Tournament Seed or Selection column, that I have grey highlighting for a number of cells.  These are for teams where the change in its RPI rank due to the new rule suggests that the team might have -- or definitely would have, according to the seed ranges -- gotten a different seed result from the Committee.

If you look at Santa Clara in the table, you will see something interesting.  It had no overtime games decided by a golden goal, so its record is unchanged.  Its rank, however, moves up from #29 to #17, a pretty significant improvement.  Since its own record is unchanged, this means all of the rank change is due to a change in the Santa Clara strength of schedule and/or to other teams moving to poorer ratings as a result of the rule change.  On looking at the changes for the teams Santa Clara played in 2019, they have 13 golden goal wins converted to ties and 22 golden goal losses converted to ties, the net effect of which would be an improvement in Santa Clara’s strength of schedule.  Santa Clara thus gives a good illustration of how a team can have no change in its own record but nevertheless move significantly in the rankings due to all of the other moving parts in the rating system.

Here is the last table.  It includes data for all of the teams (not automatic qualifiers)  ranked between #31 (of the new rule ranks) and #57 of either of the actual 2019 ranks or the new ranks.


A key column in this table is the RPI Rank and Top 50 Results Rank Rank column.  This column shows the ranks of teams based on combining their RPI ranks and their Top 50 Results Ranks, each weighted at 50%.  Top 50 Results Ranks are based on my own scoring system for good results (wins or ties) against Top 50 opponents -- it essentially is a measure of how high in the rankings a team has shown it is able to be successfully compete.  This is a key column because historically, if I use this column to predict Committee at large selections from among the teams ranked #31 through #57, on average it matches all but two of the Committee selections each year.

In the table, the blue highlighting indicates that, using team RPI Rank and Top 50 Results Rank Ranks, these are the teams that would get at large positions.  In the NCAA Seed or Selection column, the blue highlighting for Tennessee means it is a team (the only team) that did not actually get an at large selection but that would get one under the new rule using the Top 50 Results Rank Ranks.

Teams highlighted in orange are ones that, using RPI Rank and Top 50 Results Rank Ranks, actually did get at large selections but now likely would not -- TCU, Washington State, and Utah (replaced by Florida Atlantic and Yale who moved into the Top 30 and by Tennessee with the better RPI Rank and Top 50 Results Rank Rank).

There are couple of other interesting pieces of information in the table.

Wake Forest moves up from actual #61 (outside the bubble) to #54 (inside the bubble).  The red highlighting in the NCAA Tournament Disqualified column, however, indicates it actually had a winning percentage below 0.500 (which it likewise has under the new rule), so it is not eligible for an at large position.

Mississippi and Georgia both were within the actual top 57 and thus could have been considered bubble teams.  Under the new rule, however, they both fall out of the bubble group.  As the NCAA seed or selection column indicates, neither got an at large selection, so the rule change does not affect their NCAA Tournament status.

Denver and Boston College move up from actual #66 and #77 to #57 and #56 respectively so that both now can be considered in the bubble.  According to the RPI Rank and Top 50 Results Rank Ranks, however, they do not get at large positions, so the rule change does not affect their likely NCAA Tournament status.

Looking at 2019 as a whole, under the new rule it appears there would have been some NCAA Tournament seed changes (although nothing suggests changes in the four #1 seeds).  As I look at the at large selections, it appears that the last of the bubble teams likely would have been Notre Dame, Georgetown, Iowa, TCU, Washington State, and Utah, which are, in order, the bottom teams in the RPI Rank and Top 50 Results Rank Ranks.  All of those teams got at large selections in 2019, but most likely three of them do not get at large selections under the new rule.

Additional Comments

Looking at the above information for the 2019 season, the changes for particular teams and conferences appear relatively random.  If the test were to include more seasons, perhaps patterns would appear, but I do not see them at this point.

The test does seem, however, to give a picture of how big the effect of the rule change is likely to be on the NCAA Tournament bracket.  In essence, it likely will mean a few differences in seeds and a few different teams getting at large selections than would have been the case under the old rule, with the differences being relatively random.

This assumes that the Committee will not make significant adjustments in how it makes its bracket decisions, in response to the new rule.  Practically speaking, I think this is a pretty good assumption.  The Committee, as will be the case for all of us, will be looking at data very similar to the data it has had in the past except that there will be more ties.  It is hard to imagine anything changing in the Committee thought process simply because there are more ties.

There is one possible way I can think of in which the change might affect a particular class of teams.  It is possible that teams with better bench depth have an advantage in overtime games.  If this is the case, then eliminating regular season overtime games may work against these teams.  I suspect, however, that any change would be subtle, as teams with better bench depth may be able to make tactical adjustments to offset the loss of any advantage in overtime games.

The bottom line is that the rule change will result in some differences in RPI ratings and ranks and in the Committee’s NCAA Tournament bracket at large selections and seeds.  Overall, however, it seems like so far as the RPI and Committee decisions are concerned, seasons will not look significantly different than they have in the past.

Friday, February 4, 2022

SCHEDULING RESOURCES

 I have two scheduling resources available for teams, both in the form of Excel workbooks:

Team Histories and Simulated 2022 Ranks

This is a resource that lets you review, for each team:

Its RPI rating and rank history, plus its Massey rank history;

Its history in terms of its rank as a contributor to opponents’ strengths of schedule;

Its history in terms of the average ranks of its opponents.

It also includes a trended RPI rating and rank I have assigned each team going into the 2022 season, based on its rank history, as well as its expected rank as a contributor to opponents’ strengths of schedule.

The workbook has a User Guide that explains in detail what is in the workbook and how to use it as part of your scheduling process.

This is a very large workbook, with a page for each team as well as additional resource pages.  Because of its size, it is not practical to use it on line.  Rather, you will need to download it using the following link: Team Histories and Simulated 2022 Ranks.  I have found that it takes the following steps to download it:

When you have clicked on the link, a page will come up in your browser with a rotating circle.  You may have to wait a while because of the size of the workbook.

A black page will come up with a rotating circle.  On the upper right, there will be a download icon.  Click on the icon (even if the circle still is rotating).  Again, you may have to wait a while.

A black page may come up again with a rotating circle.  You also may see some of the text of the User Guide page of the workbook.  Wait until a download box appears at the bottom left.  Click on the carat mark to open the box.  Then click on Open.

At this point, the Excel file should open.  Use Save As to save the file to your computer with whatever name you want to assign it.

If you have trouble downloading the file, email me at cpthomas@q.com and I will email a copy to you.

Team Scheduling Tool

This is a Tool custom set up for each team that wants one.  It lets the team try out different schedules to see where the team is likely to end up in terms of its RPI rating and rank and also in terms of several factors related to NCAA Tournament at large selections and seeds.  It contains a User Guide that explains what is in the Tool and how to use it.

I also can set the Tool up for use by a conference to see how scheduling by the individual teams affects the conference teams’ RPI ratings and ranks and NCAA Tournament prospects.

I will set up a tool for any team that wants one.  If you want one, simply email me at cpthomas@q.com.  In order to set up the tool, it will work best for me and for you if you will send me your conference’s entire in-conference schedule as well as your proposed non-conference schedule (I do not need game dates, but I do need game locations).  Once I have the needed information, I will set up the Tool and send it to you.

 

 

Friday, January 7, 2022

GRADING THE COMMITTEE ON ITS NCAA TOURNAMENT SEED AND HOME FIELD DECISIONS

Each year, in forming the NCAA Tournament bracket, the Women’s Soccer Committee selects 4 #1 seeds, 4 #2s, 4 #3s, and 4 #4s and places them in the bracket accordingly.

If we think of the #1 seeds as 1.1, 1.2, 1.3, and 1.4, it seems clear the Committee places 1.1 at the top left of the bracket.  After that, although it is not always clear, it appears to place 1.2 at the top right, 1.3 at the bottom right, and 1.4 at the bottom left.  (Over the last 11 seasons, but excluding 2020, the 1.1 top left have won 44 Tournament games, the 1.2 top right 41, the 1.3 bottom right 41, and the 1.4 bottom left 36.  It is possible the 1.2 and 1.3 positions are reversed from what I have indicated.)

As for the #2, #3, and #4 seeds, it is not clear to me how the Committee places them.  For purposes of this article, it does not matter.

In addition, in the Tournament first round, there are 16 games between unseeded opponents.  In these games, the Committee appears to give home field to the teams it considers stronger.

When the Committee makes these decisions, it is supposed to be free from bias, both in terms of the regions teams come from and the conferences they play in.  But is it really free from bias?  When I looked at the results of the 2021 Tournament, it occurred to me that this might be a good question to research.  So I did.

My approach was to see how many rounds regions’ and conferences’ teams actually won in each of the last ten Tournaments as compared to how many the Committee, based on its seed and host decisions, thought they would win.  In terms of how many games the Committee decisions say teams should win, here are the numbers:

#1.1  6 games (champion)

#1.2  5 games (runner up)

#1.3 and 1.3  4 games (semi-finalists)

#2s  3 games (quarterfinalists)

Unseeded pairs, 1st round host  1 game (get to second round)

Unseeded pairs, 1st round visitor  0 games (lose in first round)

Unseeded, playing seed in 1st round   0 games (lose in first round)

Regions

This year, I am changing my approach to regions by using four strictly geographic regions based on where a state’s teams play the greatest numbers of their games.  (Given the frequency of teams shifting conferences, this seems a better long term approach for studying regions than using conferences as the building blocks for regions.)

Middle: Illinois, Indiana, Iowa, Michigan, Minnesota, Missouri, Nebraska, North Dakota, Ohio, South Dakota, Wisconsin (64 teams)

North: Connecticut, Delaware, Maine, Maryland, Massachusetts, New Jersey, New York, Pennsylvania, Rhode Island, Vermont, Washington DC (79 teams)

South: Alabama, Arkansas, Florida, Georgia, Kansas, Kentucky, Louisiana, Mississippi, North Carolina, Oklahoma, South Carolina, Tennessee, Texas, Virginia, West Virginia (139 teams)

West: Arizona, California, Colorado, Hawaii, Idaho, Montana, Nevada, New Mexico, Oregon, Utah, Washington, Wyoming (62 teams).

For regions, I look at each region team’s bracket position to see how many games the positioning says it should win.  Then I look at actual results to see how many games it actually won.  From the games it actually won, I subtract the number the Committee positioning says it should have won.  If it did better than expected, it will have a positive number result.  If it did as expected, it will have a 0 result.  If it did poorer than expected, it will have a negative number result.

If two teams from the same region play each other, then either each will perform as expected and both will get 0 results from the game.  Or, neither will perform as expected and one will get a +1 and the other a -1 result, with these numbers canceling each other out from a region perspective.  Thus when I add up all the numbers for a region’s teams, the total will reflect how the region did in games against other regions.

The following table summarizes how the regions actually did as compared to how their Tournament positioning says they should have done:

The table gives one kind of look at the data.  The totals suggest that the Committee decisions, for whatever reason, have been wrongly biased in favor of the South region and against the North and Middle regions and, to a lesser extent, the West region.

A different kind of look considers how the numbers have trended over time:


In this chart, 2011 is on the left and 2021 on the right.  The circle markers are the region data points, connected by the solid lines.  The dotted lines are the trends (straight line trends).  The trend lines suggest that the Committee has slightly underrated the Middle region (blue) over time, but been pretty consistent about it.  In the past the Committee slightly overrated the North region (red) but that has trended downward and now is right about where it should be.  The Committee started the decade rating the South (grey) about right, but is trending towards significantly overrating it.  And conversely, the Committee started the decade overrating the West (gold), but is trending towards fairly significantly underrating it.

You can decide for yourself how seriously to take what the table and chart show.  The numbers, however, are correct.

Conferences

For conferences, the process is the same.  A problem, however, is that teams have changed conferences during the 2011 to 2021 period.  I did not want to be moving teams from one conference to another, so I have treated teams as though they were in their current conference throughout the period.


Here, the totals make it look like the Committee has been biased in favor of the American, Big 12, Pac 12, and SEC conferences in particular and against the Big 10 and Colonial, and especially the West Coast conference in particular.

The trends, however, give a different look:


The chart covers only the conferences that have had + or - results in at least half of the years: the Power 5 plus the American, Big East, and West Coast.  I have eliminated the lines connecting the data points to make it easier to see the trend lines.

Since the colors are a little hard to read, I will start at the top left of the chart and move down through the trend lines and what they suggest:

  • For the ACC (blue), the Committee has gone from significantly underrating it to significantly overrating it.
  • For the Big 10 ( black), the Committee has slightly underrated it pretty consistently.
  • For the Pac 12 (gold), the Committee started out rating it about right and is trending towards slightly overrating it.
  • For the Big East (green), the Committee consistently has rated it about right.
  • For the American (grey), the Committee started out rating it about right and is trending towards slightly overrating it.
  • For the West Coast (dark blue), the Committee has gone from overrating it to significantly underrating it.
  • For the SEC (dark green), the Committee has gone from underrating it to rating it about right.
  • For the Big 12 (purple-brown), the Committee has gone from underrating it to slightly underrating it.
Here again, you can decide for yourself how seriously to take what the table and chart show.  The numbers, however, are correct.

Conclusion

From both region and conference perspectives, the data suggest the possibility that, in seeding and in assigning home field to first round unseeded pairings, the Committee may have region and conference biases that are wrongly influencing its decisions.  Although this is just a possibility, perhaps it is something the Committee should bear in mind in its future bracket formation decisions.



Thursday, December 30, 2021

NCAA TOURNAMENT: ANALYZING THE COMMITTEE SEED AND AT LARGE DECISIONS

Each year, I analyze the Committee seed and at large selection decisions in relation to the Committee’s historic decision patterns.  This year, I have set up a series of tables for the analysis.  I will start with the #1 seed table, giving a detailed explanation of its data.  Then, I will show the tables for each of the #2, #3, and #4 seeds and for the at large selections, with a few comments.

For each table, you will need to scroll to the right to see the entire table.

#1 Seeds

Here is the #1 seed table:


First, some general background to help with the table.

For at large selections, the NCAA has set specific data factors from the season that the Committee must use and is limited to using.  The NCAA leaves it up to each Committee member to decide how much weight to assign to each factor.  The Committee also uses the factors in the seeding process, but for seeding they are not mandatory and the Committee is not limited to the factors in evaluating teams.

I break the NCAA factors down into 13 individual factors:
  • RPI (adjusted)
  • RPI Rank
  • Non-Conference RPI (adjusted)
  • NCRPI Rank
  • Top 50 Results (my modification of an NCAA factor)
  • Top 50 Results Rank
  • Conference Standing (I use average of regular season standing and conference tournament finishing position)
  • Conference RPI
  • Conference RPI Rank
  • Head to Head Results Against Top 60 Opponents (my modification of an NCAA factor)
  • Results Against Common Opponents with Other Top 60 Teams (my modification of an NCAA factor)
  • Common Opponent Results Rank (my modification of an NCAA factor)
  • Poor Results (my modification of an NCAA factor)
There are NCAA data available that allow an evaluation of a team for each factor.  For some of the factors, the NCAA has a scoring system.  For example, an NCAA formula assigns a value for the RPI.  Where the NCAA does not have a scoring system for a factor, I have created my own scoring system.

In addition, to mimic how a Committee member might think, I pair each factor with each other factor.  For each factor pair, I have a scoring system that gives each factor a 50% weight.

In addition, I have one other factor, Number of Games Against Top 60 Opponents.  This is not an NCAA mandated factor but rather is one I use as an aid to teams wishing to do their non-conference scheduling with a view towards the NCAA Tournament.

Altogether this produces a total of 92 individual and paired factors.

For each Top 60 team for each year since 2007 (the first year I began collecting data), my computer program has computed a score for each of the 92 factors.  (Hereafter, I will refer to my computer program and process as my "system.")  My system then compares the scores of all of the Top 60 teams since 2007 to the Committee seeding and at large selection decisions.  From this comparison, the system identifies two scores for each factor:

A "Yes" score, which means that for a particular Committee decision, if a team has had the Yes score or better for that factor, the team always has gotten a favorable decision from the Committee.  For example, if a team has had an RPI Rank of 1, it always has gotten a #1 seed.  I call such a Yes score the Yes standard for that factor.  Thus <=1 is the RPI Rank Yes standard for a #1 seed.

A "No" score, which means that for a particular Committee decision, if a team has had the No score or poorer for that factor, the team never has gotten a favorable decision from the Committee.  For example, if a team has had an RPI Rank of 8 or poorer, it never has gotten a #1 seed.  Thus >=8 is the RPI Rank No standard for a #1 seed.

Note:  As the data turn out, there are some factors, for some Committee decisions, that do not have a Yes standard or that do not have a No standard.  For example, a team’s Conference Standing does not have a Yes standard for any Committee decision.  In other words, your conference standing, all by itself and without reference to what conference you are in, will not assure you of any seed position or of an at large selection.

On completion of the regular season including conference tournaments, my system tallies up all of the Yes and No standards a team has met, for each Committee decision -- #1, 2, 3, and 4 seeds and at large selections.  For each team, for each Committee decision, there are four possible outcomes:

  • The team meets one or more Yes standards and no No standards.  This means that if the Committee follows its historic pattern, the team will get a Yes decision from the Committee.
  • The team meets no Yes standards and one or more No standards.  This means that if the Committee follows its historic pattern, the team will get a No decision from the Committee.
  • The team meets no Yes standards and no No standards.  This means that either a Yes or a No decision from the Committee will be consistent with its historic pattern.
  • The team meets one or more Yes standards and one or more No standards.  This means that the team has a profile the Committee has not seen historically.  Whatever decision the Committee makes cannot be fully consistent with its historic pattern.
With that background, the above table for #1 seeds shows data related to each of the teams with RPI ranks #1 through 7.  This is the candidate group for #1 seeds, since the No standard for a #1 seed is >=8.

Committee Decision: Green means the Committee gave the team a #1 seed.  Red means it did not.

RPI Rank 

Top 50 Results Rank:  I have included this in the table because historically the factor pair of RPI Rank and Top 50 Results Rank has proved to be a good predictor of Committee decisions, especially for at large selections.

Yes Standards Met and No Standards Met:  This is the number of Yes and No standards for #1 seeds that the team has met.

Yes Standard: If a team has met one or more Yes standards but does not get a Yes decision from the Committee, it is useful to know what the Yes standard is, so I list it here.  For example, Florida State had an RPI Rank of #1.  Suppose the Committee had not given it a #1 seed.  Then in this column you would have seen 2 RPI Rank (the 2 preceding RPI Rank simply is a number I have assigned to that standard).

Yes Value:  If I have listed a Yes Standard, I will state the standard score here.  For example, for a #1 seed, the RPI Rank Yes standard is <=1, which means that teams with RPI ranks of #1 always have gotten #1 seeds.  If the Committee had not given Florida State a #1 seed, you would have seen <=1 in this column.

Yes Actual:  If I have listed a Yes standard, I also state the team’s actual score for the standard.  So, if the Committee had not given Florida State a #1 seed, you would have seen 1 in this column representing its RPI rank.

No Standard.  If a team has met one or more No standards but gets a Yes decision from the Committee, it is useful to know what the No standard is, so I list it here.  For example, the RPI Rank No standard for a #1 seed is >=8.  If the Committee had given a #1 seed to the #8 RPI team, I would have listed 2 RPI Rank here.

No Value. If I have listed a No standard, I will state the standard score here.    If the Committee had given the RPI #8 team a #1 seed, you would have seen >=8 in this column.

No Actual:  If I have listed a No standard, I also state the team’s actual score for the standard.  So, if the Committee had given the #8 team a #1 seed, you would have seen 8 in this column.

Teams Affected:  If the Committee has given a No to a team that has met a Yes standard or a Yes to a team that has met a No standard, this column will give an indication of how significant a change the Committee has made from its historic pattern for that factor.  For example, since 2007 through 2019 there were 13 teams with #1 RPI ranks and they all received #1 seeds.  If the Committee had not given Florida State a #1 seed this year, then you would have seen 13 in the Teams Affected column.  This would mean that there are 13 teams that, based on past history, we would have thought assured of #1 seeds but that, based on the Committee decision this year, no longer could be considered as having been assured of #1 seeds.  The lower the Teams Affected number, the smaller the Committee change from its historic pattern.  If the Teams Affected number is 0, it means the team has a score for the factor and Committee decision that is just next to the historic standard and that the Committee has not seen before so that the Committee decision simply represents a refinement of the previous standard.

As a further note about Teams Affected, to give the numbers some context:

  • The #1 seed candidate range is teams with RPI Ranks of #7 or better, so the number of candidates for #1 seeds since 2007 has been 13 x 7 = 91.  So when you are looking at a Teams Affected number for #1 seeds, it is that number of teams out of a total of 91. 

  •  The #2 seed candidate range is RPI Ranks of #14 or better, so the number of #2 seed candidates has been 13 x 14, less the 52 teams that got #1 seeds, which amounts to a #2 seed candidate pool of 130 teams. 

  •  The #3 seed candidate range is RPI Ranks of #23 or better, so the number of #3 seed candidates has been 13 x 23, less the 104 teams that got #1 and 2 seeds, which amounts to a #3 seed candidate pool of 195 teams. 

  • The #4 seed candidate range is RPI Ranks of #26 or better, so the number of #4 seed candidates has been 13 x 26, less the 156 teams that got #1 through 3 seeds, which amounts to a #4 seed candidate pool of 182 teams. 

  • The At Large candidate range is RPI Ranks of #57 or better, less the 208 seeded teams, the Automatic Qualifiers in the Top 57, and teams in the Top 57 that failed to meet the 0.500 minimum winning record requirement.  Altogether since 2007, this has amounted to 447 teams. 

 Round Eliminated:  This column shows the NCAA Tournament round this year in which the team was eliminated.  It lets you look at the position the NCAA assigned the team and see how the team did in relation to that assigned position.  For example, Virginia’s #1 seed means that according to the Committee it should have made it at least to the semifinals, but instead it made it only to the 3rd round.  This lets you evaluate how the Committee decisions worked out. (For first round matchups between unseeded teams, I treat the home team as the stronger team according to the Committee.)

With the above explanation, I leave it to you to go up to the #1 Seeds table and see how the Committee decisions match up with its historic pattern.  My comment is that the Committee did not deviate from its historic pattern except when it had to with Duke and Virginia due to their meeting both Yes and No standards and that for those teams, the Committee deviation was small.

#2 Seeds

Here is the #2 seed table:

You can review the table and reach your own conclusions.  My comment is that in giving UCLA a #2 seed, the Committee deviated from its historic pattern.  The deviation was not large but also was not insignificant.  As an alternative, Tennessee would have been an easy #2 seed.

#3 Seeds

Here is the #3 seed table:


My comment is that there is nothing of major import in the Committee decisions.

#4 Seeds

Here is the #4 seed table:


There is only one significant Committee deviation from its historic patterns here, and it is the #4 seed given to BYU.  This was a pretty large deviation.  Ironically, BYU reached the championship game.

 At Large

Here is the At Large table:


My comment is that the only deviation here from Committee historic patterns is in its giving St. Johns an at large position rather than West Virginia, Colorado, Oregon, or Houston.  In looking at St. Johns’ Teams Affected numbers, however, the deviation was extremely small.

Summary and Two Additional Pieces of Information

My evaluation of the Committee decisions in relation to historic Committee patterns suggests that:

1.  The Committee At Large selections were quite consistent with historic patterns and, where it varied with St. Johns, the variation was very small.

2.  The Committee seeds were largely consistent with historic patterns.  Where the Committee varied from historic patterns, most of the variations were small.  The greatest variation was BYU getting seeded, which is ironic since BYU made it to the championship game.

In addition to the above, I have looked at two other aspects of the Committee decisions.

First, I looked at geographic regions based on the states where teams are located.  As it turned out this year, during the season teams from states in the West played roughly 90% of their games against other teams from the West and only 10% against teams from other regions.  This created a big problem for the RPI since 10% is not enough games for the RPI to properly rank teams from a region in relation to teams from other regions.  Because of this, I wanted to see if teams from the West performed differently in the Tournament than the Committee had evaluated them.  The following table addresses this question:


In this table, the Wins Difference column shows, for each region, the difference between (1) the number of games the Committee seeds and bracket placements for unseeded teams indicated teams should win and (2) the number of games teams actually won.  The numbers in this table represent how teams from a region did against teams from other regions, since all within-region games cancel each other out.  Although the numbers in the table are not large, they suggest that the Committee may have over-evaluated teams from the South and under-evaluated teams from the other regions and particularly the West.

Second, I took a similar look, but this time by conference.


This suggests that the Committee may have undervalued the West Coast Conference and, to a much lesser extent, the Big 10 and overvalued the ACC and Pac 12.

Since both of these tables are based on only one year’s results, I do not take them too seriously.  They suggest, however, that it might be worthwhile to do a study that considers more years of NCAA Tournaments, to see if there are any Committee region- or conference-based overvaluation and undervaluation patterns.