Archived referee methods
How foul patterns are compared across officials' games, with season adjustments, sample thresholds, and tests against random assignments. Calls are recorded at crew level.
Keep in mind This is archived research on foul frequency in games each official worked. It cannot identify who made a call or whether the call was correct.
On this page
What an official's row represents
Three officials work each game, but the play-by-play does not identify who made a call. Each foul is attributed to the crew.
a game credits all three officials equally
⇒ an official's rate includes calls made by their crewmates
⇒ it does not isolate that official's individual effectOfficials work with different partners over time. This can reduce the influence of one crewmate, but it does not establish random assignments or remove all confounding. The figures describe games an official worked, not calls they made.
The research samples
The analysis uses cached ESPN play-by-play and box scores, so each test can use the same game records.
The three game counts reflect different input requirements and extraction dates. Each analysis reports its own denominator.
Timing covers more games than the foul mix because a play stream survives in games whose box score does not. The folklore chapter covers more than either because it was rebuilt later, on a filter that admits the games where ESPN lists a standby fourth official alongside the three who worked, a case the earlier extracts dropped. Playoff games are counted only there, and only for the questions that are about the postseason.
Comparing foul shares within each season
Officiating changes with the rulebook. The league called a different game in 2015-16 than in 2025-26, so comparing an official’s raw rate to a pooled average would credit them with the era they happened to work in. Every figure is therefore a deviation from the league’s own average in the same season, and on shares rather than counts wherever pace could otherwise masquerade as a tendency.
deviation = official's share of foul type T
− the league's share of T, that season
emphasised when |z| ≥ 2, at that official's own sample size
published only for officials with ≥ 200 gamesThe bar cuts both ways and is meant to. At |z| ≥ 2, about 3.4 of 74 officials clear it from noise alone, so a bold cell does not establish an individual tendency. The page does not lead with a name on that basis. Muted cells are shown rather than hidden, because a table of only the significant ones invites the reader to find a pattern that was selected for them.
Why each row uses the latest 200 games
Careers in this data run from 200 games to more than 600, and a z-score bar moves with sample size: an identical quirk that clears |z| ≥ 2 at n = 700 is out of reach at n = 200. Worse, whistles measurably change. A pre-registered drift test split every official with ≥ 350 games into their most recent 200 games and everything earlier: 37 of 312 cells sat beyond |zΔ| ≥ 2 (11.9%), where chance produces about 4.6%. This suggests that career averages can obscure changes over time.
So since 2026-08-24 the table scores every official on their most recent 200 games, the publication bar, so every published row is a full window, the same n and the same bolding bar on every line. At n = 200 the bar is harder to clear, so the table bolds 83 type cells where the career basis bolded 119. The full-span figures ship alongside in the same artifact for anyone comparing.
The per-season split was measured in the same pre-registration and, against expectation, cleared its declared bars (11.8% of official-season cells beyond |z| ≥ 2; 75.2% within-official sign agreement). It still is not shown in the browse table because adding a season axis would make it much larger. The published artifact retains those results.
Comparing officials with random assignments
Counting how many officials clear a bar is not enough to say officials differ, because some always will. The verdict comes instead from a single question asked once: is the spread between officials wider than the spread you get by dealing the same games out at random?
observed: the spread of per-official means
null: redraw each official's games at random from the
same seasons, holding games-per-season fixed
(2,000 times)
verdict: how often the null's spread reaches the observed oneHolding games-per-season fixed is what stops an era doing the work: two officials who worked different decades cannot be made to differ by the league’s foul rate changing between them. And because it is one test rather than one per official, the test does not select an individual official from many comparisons. Other questions on the page still need their own treatment of multiple testing.
Unusual official-player records and chance
The folklore chapter compares official-player records with the extremes expected from checking many pairs. An unusual record alone does not establish bias.
pairs examined 689 minimum shared playoff games 10 most extreme p from PURE NOISE 0.00145 the most famous pair actually 0.0016 cleared p < 0.01 7 (chance predicts 6.9) cleared p < 0.05 27 (chance predicts 34.5)
Across 689 pairs, the expected minimum p-value under the chance calculation is about 0.00145. This is a reference for the scale of extremes, not a corrected significance threshold. The featured pair’s result is less extreme than that reference.
One-sided and two-sided p-values are not interchangeable. The noise floor is the expected minimum of a two-sided sweep, so every pair compared against it is quoted two-sided too. A test fails if that ever stops being true.
Preregistered questions and decision rules
Testing many questions and reporting only favorable results can exaggerate the evidence. The questions and decision rules were committed before the analyses ran.
Two consequences are visible on the surface. The Q4 “clutch” question was gated behind a coarser per-quarter test, so the narrow window was only allowed to proceed if the broad test passed. It did not, and that result is published. And the five famous claims were named in writing before the postseason was even fetched, reducing the risk of choosing claims after seeing their results.
The protocol required publishing null results. The reported tests include player foul rates, player win records, star foul trouble, crowd effects, and make-up calls, including those that did not meet their declared evidence thresholds.
Full limitations
- —It cannot attribute a call to an individual. Each figure describes the games an official worked with two crewmates.
- —It cannot judge whether a call was correct. The figures measure call frequency, not accuracy.
- —It cannot see before 2015-16. Named officials are available further back, but the play-by-play detail these measures need is not.
- —It cannot test the playoff legends properly. A pair shares a handful of postseason games in a lifetime; both eras of the most famous claim fall below the minimum this page requires before it will judge a pair at all.
- —It cannot undo how a claim was found. A record the public discovered by scanning outcomes can only be confirmed on games nobody had seen when they found it, and there are rarely enough of those.
- —It cannot separate officiating from changes in how teams play while leading or trailing. The score-state gradient is an association, not a causal estimate.