Skip to content
Leaderboard

How ratings work

Your rating comes from your speaker scores and your team results at every university BP tab we have, compared with the rest of each field, and is shown on a 66 to 90 scale.

What your rating is

Your rank comes first. We sort every rated debater from strongest to weakest, and your rating is your rank put on a scale you already know from tabs: 66 at the bottom, 90 at the top, about 75.5 in the middle.

So a rating of 80 does not mean you average 80 speaks. It means your results put you ahead of about 85% of rated debaters in that period. Because the number follows your rank, it can move when other people's results change, even if you haven't debated.

What counts

  • University BP tabs. High school, middle school and World Schools tabs are kept apart and never count.
  • Tournaments where you spoke in at least 3 inrounds. Fewer than that and the tab is skipped for you.
  • Your speaker score in each inround, and your team's result in each room.
  • How far you went in outrounds, if you broke.
  • Tournaments with a field-strength score of at least −1. That leaves out only the weakest fields, mostly novice tabs. A tab with no strength score yet still counts.

Each board covers a period. On the last 6 months board, tabs up to about 6 months old count in full, then less in a straight line until they drop out at 12 months: a 9-month-old tab counts half. The longer boards (a year, 5 years, the pandemic years) use every tab in the period.

On every board the speaks half of your rating also leans toward recent tabs. Speaks count half as much for every 18 months that pass, so speaks from 3 years ago count a quarter as much as this month's.

What happens after a tournament

Ratings don't change round by round. When new tabs come in, we rebuild the whole board, and each tournament you debated feeds two halves of your rating.

Your speaks, against the tab's field

Each speaker score is compared with every other speaker score at the same tab. A 78 where judges hand out 73s is worth more than a 78 where they hand out 78s. We average how far above or below the field you were, tab by tab.

Your results, against who you beat

In every room your team finishes 1st to 4th. We record the gap in team points between your team and each of the other three, then solve for the set of ratings that best explains every gap at once (a Massey rating). Beating a strong team moves you more than beating a weak one.

Outrounds

Breaking already shows in your inround results, so it earns nothing extra on its own. Each round you get through after the break does, and more when the teams in those rooms were strong. Your best outround runs count most: each next-best run counts 80% of the one before.

Weighting and the blend

Each tournament is weighted by how recent it is (above) and by how strong its field was: from 0.8 times for a weak field to 1.25 times for the strongest. Room results at tournaments with fewer than 80 teams count a little less, down to 0.55 times for the smallest. The speaks half and the results half are each turned into a position in the field, then averaged 50/50. That average sets your rank.

From rank to rating

Your position in the field runs from 1 (first) to 0 (last): position = 1 − (rank − 1) ÷ (field size − 1). Each band of positions maps in a straight line onto a band of ratings, and the result is rounded to the nearest half point.

Position in the field and the rating it maps to
PositionRating
0 to 0.5066.0 to 75.5
0.50 to 0.7075.5 to 78.0
0.70 to 0.9078.0 to 80.5
0.90 to 0.9780.5 to 82.5
0.97 to 0.9982.5 to 84.5
0.99 to 184.5 to 90.0

Worked example

Made-up figures, real formula.

You are ranked 50th of 2,000 debaters. Your position is 1 − 49 ÷ 1,999 = 0.975. That sits in the 0.97 to 0.99 band, which runs from 82.5 to 84.5, so your rating is 82.5 + (0.975 − 0.97) ÷ 0.02 × 2 = 83.0.

After the next rebuild you are 40th. Your position is 1 − 39 ÷ 1,999 = 0.980, and your rating is 82.5 + (0.980 − 0.97) ÷ 0.02 × 2 = 83.5 after rounding. Ten places near the top moved you half a point.

When two debaters are within 2 percentile points of each other, the one we are surer about ranks higher. That is usually the one with more tabs, so a two-tab hot streak doesn't sit above a long record on thin evidence. Nobody's number is marked down for it.

Tiers

Tiers are bands of the rating. They show as row headers on the leaderboard. The shares below follow from the mapping above, after rounding.

Debater tiers
TierRatingShare of the field
Champion82.5 and upabout the top 4%
Outrounds80.0 to 82.0top 4% to 16%
Bubble77.5 to 79.5top 16% to 36%
Developing73.0 to 77.0top 36% to 64%
Novice72.5 and belowthe rest, about 36%

Why you need 2 or more tabs

One tournament is one weekend. A strong draw, a friendly panel or a partner on form can carry it. To be ranked you need at least 2 tabs on the last-6-months board and at least 3 on every longer board, where one old tab says less about who was active in the period.

Even then, a short record is pulled toward the field average, and the pull shrinks with each tab you add. How strong the pull is gets measured from the whole field at every rebuild: how much one debater's results vary from tab to tab, against how much debaters differ from each other.

Debaters whose rating would put them in the top 100 but who are short of the minimum are listed under “Strong results, too few tabs” at the bottom of the leaderboard. If you know a tab one of them debated at that we don't have, import it or send it to hello@buildacase.ca.

How judges are rated

A judge's rating comes from what they were trusted to judge. At each tournament we take the deepest seat they earned and give it points:

Points for the deepest seat at one tournament
Deepest seatPoints
Grand final chair, or Chief Adjudicator12
Grand final panel, or semifinal chair10
Semifinal panel, or quarterfinal chair9
Quarterfinal panel, or octofinal chair8
Octofinal panel6
A bubble room in the inrounds5
Inrounds only0

Those points are multiplied by how hard the seat was to reach: how few teams made that round out of the whole field, and how few of the judges there were trusted with it. A stronger judging pool adds a little, and WUDC gets a small extra weight of 1.12.

Only tournaments with a field-strength score of at least 0.3, or a named major (WUDC, EUDC, NAUDC, Cambridge, Oxford and others), count. Then we add up a judge's best 8 tournaments, best first: the best counts in full, the next 70%, then 45%, 25%, and a small tail after that. Judges who are also in the top 20% of debaters get a bonus of up to 10 points. A judge needs 2 qualifying tournaments to be ranked.

The total sets the rank, and the rank sets the number: a straight line from 0.5 for last to 10 for first, rounded to the nearest half. The tiers:

Judge tiers
TierRatingShare of judges
CA9.5 and upabout the top 8%
Outround Chair8.5 to 9.0top 8% to 18%
Outround panellist7.0 to 8.0top 18% to 34%
Inround chair4.0 to 6.5top 34% to 66%
Inround panellist2.0 to 3.5top 66% to 87%
Trainee1.5 and belowthe rest, about 13%

What this can't see: a CA testing a judge in a weak room rather than rewarding them, a judge who left before the break, or the private feedback scores a CA team keeps. It reads the pattern across many tournaments instead of any single seat.

Countries, field strength and arrows

Countries

Each country is ranked on the average standing of its 10 best-rated debaters, so 10 strong debaters count for more than 30 average ones. A country needs 3 rated debaters to appear.

Field strength

Field strength answers one question: how hard was it to break here? We find the teams still fighting for the break before the last two inrounds, look at the ratings of the teams they met in those rounds, and take the middle value. Round robins use the average of the whole field, because every team meets every other team. “Top 5%” means a stronger field than 95% of tournaments in the period.

Arrows and up-and-coming

On the last-6-months board, the small up and down arrows beside a rank compare it with a saved weekly snapshot, and the date is in each arrow's tooltip. Moves of 1 or 2 places are not shown. “Up and coming” lists people now in the top 300 who have climbed at least 25 places in the last 6 months and held or improved in at least 7 of every 10 weekly snapshots.

Technical notes

The full write-ups, for anyone who wants to check the maths.

How the debater rating is built

Why Massey, not Plackett-Luce

Plackett-Luce (PL) is the textbook model for "rank these four teams" data. It's a generative model: the 1st-place team is drawn from a softmax over team strengths, then 2nd from the remaining three, and so on. We tried it. Two consecutive PL runs on identical code and data produced completely different top-5 rankings, with zero overlap. The solver was finding different local optima or hitting the iteration cap before converging. Outcome rating ranges varied by 0.44 points between runs. Massey on the same data produced stable, differentiated rankings. The deeper reason PL struggled: it assumes judges decide by sequential elimination. BP judges don't work like that. They think in margins. "Team A was clearly first, Team D clearly last, B and C were close." The deliberation assigns quality scores. It doesn't run a sequential pick. Massey fits that process. It estimates latent quality scores which explain observed margins, rather than modelling a selection procedure the judges never ran. What matters is matching the process which generated the data. Defending a model in the abstract is a separate exercise. We archived PL and stayed with Massey.

How well it predicts

We ran a held-out validation: trained the ratings on BP rounds before 2024, then predicted 2024 and 2025 rounds the model had never seen. It beat a random baseline by a clear margin, and the gap between in-sample and held-out accuracy was small, so the model isn't just memorising the rounds it was built on. We check it two ways, because they answer different questions. One asks "is the rating signal real on rooms where we already know the teams?" The other asks "how well does it predict across the whole recent field, including the many rooms where every speaker is too new to have a rating yet?" The first is the cleaner test of signal. The second is the honest test of deployment. It clears both.

How the judge rating is built

Why judging is so hard to rate

Debaters leave measurable signal every round: speaker scores, team points, wins and losses against specific opponents. Judges leave almost none. There's no objective measure of how good a call was. The only reliable signal is what the CA did with you: which room you got, whether you chaired or wung, whether you broke, and how deep you went. Even that signal is messy. CAs allocate for reasons that aren't all about quality. Sometimes they panel a judge in a weak room to test them, not reward them. Sometimes a judge only sits four rounds because they showed up to help a friend who's CA-ing and weren't free for the rest, so they don't break despite being trusted. Allocation runs on personal trust, so a CA's friends end up on their panels, and that's often because those people are genuinely good, not because of favouritism. Some judges are quiet workhorses who deliver a correct call in a mid room every single time and never need testing or a marquee seat. At some IVs the final has to be chaired by someone from the host institution, so a GF chair is sometimes a local rule rather than a global signal. And CAs don't usually judge deep rooms at their own tournaments, because they're running the tab. So the signal is noisy. We use it anyway, because it's the only consistent one there is, and we lean on the pattern across many tournaments rather than any single seat.

What we actually score

Each tournament, we look at the deepest seat a judge earned. A grand final chair sits at the top, then a grand final panel or a semi chair, on down through the quarters and octos to the early outrounds. Being Chief Adjudicator scores at the top too: you don't get handed a tournament's adjudication unless you're trusted to chair its biggest room. Chairing counts for more than sitting as a wing, because leading the deliberation is its own kind of trust. A seat in a bubble room counts as well, even though it's a prelim, because CAs hand-pick who they trust with the rounds that decide who breaks. Judging only prelims, with no bubble room and no outround, doesn't score on its own.

Stronger tournaments count for more

A grand final chair at WUDC is not the same as a grand final chair at a small novice IV. So every seat is weighted by the tournament: how strong the field was, and how prestigious the tournament is. Depth at a major outweighs the same depth at a weaker tab. The scarcity of the seat matters too. At a tournament where only a handful of judges make outrounds, breaking is worth more than at one where half the pool does. That keeps a WUDC grand final scarce even though plenty of people break the tournament overall.

From seats to one number out of 10

We add up a judge's best tournaments, leaning on their strongest results rather than counting every weekend equally. One big tournament can't be diluted by a busy season of small ones, and one great weekend can't carry a thin record on its own. Judges who are also strong debaters get a small bonus, since the ability to win rounds tracks the ability to call them. The total maps onto a scale out of 10, the way debaters already read adjudicator feedback: a trusted wing sits in the middle, a chair higher, and the chairs of finals approach 10. Judges on similar records get similar scores on purpose, because a band of equally-trusted judges genuinely is equally good.

What it doesn't do

It can't tell a CA who tested you from a CA who trusted you. It can't see that you would have chaired the semi if you'd been free for rounds 7 to 9. It can't see internal feedback scores either: Tabbycat stores the CA team's private adjudicator ratings, but the API doesn't expose them. If it did, they'd be a strong signal, though one that shifts between CA teams. What it can do is read the pattern across many tournaments and many CA teams. One good tournament could be luck or a friendly CA. Ten of them, at fields of varying strength, with bubble rooms and outround seats, almost certainly isn't. One weekend is a data point. A career is a signal.

How tournament strength is scored

The question we're answering

Tournament strength should answer one specific question: how hard is it to break at this tournament if you're aiming for it? This isn't the same as "how impressive is the field on paper." A 400-team WUDC has the most elite debaters in absolute count, but power-pairing isolates the bubble from those elites for most of the tournament. By the time you're fighting for the last break spot, you're likely in a room with other 14-point teams, not with the world champions running away at the top. Sharper, smaller tournaments like LSE Open or Doxbridge concentrate the elite in a narrow break field. There's nowhere to hide. After two or three rounds you're in their rooms whether you like it or not. The bubble path is genuinely harder even though the absolute count of WUDC-level debaters is lower.

What the bubble is, exactly

A team is on the bubble entering the last two rounds if their break/no-break status still depends on those last two rounds. In BP you can earn 0 to 6 team points in two rounds (two 4ths to two 1sts). So if break_cut is, say, 17 points and your team has 11 to 16 points after round 7 of a 9-round tournament, your break fate is undecided. You could still make it. You could still miss it. That's the bubble. Teams who have already clinched the break (already at break_cut or above) aren't on the bubble. Teams who are mathematically out (below break_cut minus 6) aren't either. The bubble is the contested middle.

How we measure path difficulty

For each bubble team, we look at the opponent teams they played in their last two prelim rounds. That's up to 6 opponent teams (3 per room in BP). For each opponent, we look up the speakers' debater ratings. We take the rolling average of each opponent speaker's ratings from the 3 tournaments before this one and the 3 after. A 6-tournament window centered on the tab. This catches "current form at the time" rather than career averages which include speakers' future achievements. We average across all opponent speakers a bubble team faced in their last two rounds. That's the team's path-strength number. We then take the median across all bubble teams at the tournament. That's the tournament's pathRoll. A higher pathRoll means the bubble had to play harder opponents to break. That's the headline difficulty.

Why we use a rolling 6-tournament window, not lifetime averages

A speaker's "rating" should reflect their level at the time they competed at this tournament, not their level five years later when they've made the WUDC final. Using career averages would inflate the strength of any tournament which happened to feature someone who got famous afterwards. Using only past observations would miss late-blooming speakers whose strength wasn't visible yet. Three before and three after the tab in question is a fair compromise. It captures recent form without leaning too far into the future.

Why round robins use a different formula

At a round robin like HWS RR, every team plays every other team across the prelims. There's no power-pairing. So every bubble team's opponents are identical: the rest of the field. That means the "bubble path" question collapses. The path is the field. So for round robins we use fieldStrength: the mean rolling rating of all speakers in the field. Same idea, simpler math, honest about what the metric is measuring. At HWS RR specifically, this means difficulty is essentially the average strength of the 16 best teams in the world. Those tabs sit at the top of the difficulty ranking on this metric, where they belong.

What this changes vs the old formula

The old formula used five signals about who was in the field: top-decile attendance count, recent-WUDC elite count, AC quality, etc. All measured field composition. None measured what the bubble actually played against. With the new formula, WUDC drops sharply in the rankings. The deepest field in absolute terms doesn't matter when the bubble is well-protected from the elite by power-pairing. Doxbridge, LSE Open, KCL Open, and similar sharp invitationals rise. HWS Round Robin tabs leap to the top because every team plays the entire elite field, with no protection.

What it doesn't do, honestly

This metric only works on tournaments where we have pairings data. For power-paired tabs, we need to know who faced whom in the last two rounds. Tabs which died with their Heroku dynos before we could fetch them are missing entirely. We don't adjust for motion quality, how the tournament was run, or how good the judging was beyond what it implies for speaker ratings. The metric is one question. "How hard was it to break here?" That's all. It doesn't pretend to be anything broader. The bubble itself is defined mechanically, not from final standings. If a tournament's break math is unusual (very small fields, unusual break sizes), the bubble might be sparse or empty. We need at least 2 bubble teams with enough rated opponents to compute a meaningful median. Tabs without enough signal get no strength score rather than a noisy one.

How this compares to other rating approaches

Why this section exists

People ask why we don't use Elo. Or TrueSkill. Or the discrete-class system going around the Facebook groups. Fair questions. This section is for the people who care about the math choice, not just the ranking output. The short version: every method that isn't broken produces a roughly similar top of the leaderboard. The arguments are at the margin. Reading this section helps if you want to understand where the margin is and what each method actually does. Reading it isn't required to use the leaderboard.

Elo

Elo is the chess rating you've heard of. Each player has a number. After every game, the winner gains a few points and the loser loses the same amount. The amount you gain or lose depends on how unexpected the result was. Beat someone way above you and you jump. Lose to someone way below and you tank. It works well in chess because chess is 1v1, the matchings are sequential, and the only signal is win or lose. BP isn't any of those things. Each round has four teams of two debaters, ranked 1st through 4th, with separate speaker scores carrying their own information. You can hack Elo to handle ranked outcomes (treat the four-way result as six pairwise wins and losses), but you're throwing away the speaker score signal and ignoring strength of schedule. We tried it on a sample. The ratings drift toward whoever showed up at lots of low-bar tabs because every "win" against a soft opponent still costs them points and rewards you. Massey handles strength of schedule directly. Elo doesn't.

Glicko-2 and TrueSkill

Same family. Both are Bayesian extensions of Elo. You track not just a player's rating but also an uncertainty around that rating. New players have wide uncertainty bands. Established players have narrow ones. After each match the system updates both numbers. Useful if you want confidence intervals on the leaderboard. We don't currently surface them. Instead we use Bayesian shrinkage on the Massey side: low-tab debaters get pulled toward zero automatically. Same idea, different exposure. If we ever want to publish "this debater is somewhere between top 10% and top 25% with 90% confidence", Glicko-2 would be the natural fit. For now, shrinkage covers it.

Plackett-Luce

The textbook model for "rank these four teams" data. It's a generative model: the 1st-place team is drawn from a softmax over team strengths, then 2nd from the remaining three, and so on. We tried it. Two consecutive runs on identical code and data produced completely different top-5 rankings, with zero overlap. The solver was finding different local optima or hitting the iteration cap before converging. Massey on the same data produced stable rankings. The deeper reason it struggled: PL assumes judges decide by sequential elimination. BP judges don't work like that. They think in margins. "Team A was clearly first, Team D clearly last, B and C were close." The deliberation assigns quality scores. It doesn't run a sequential pick. Massey fits that process. PL doesn't. We archived it.

Discrete-class scoring systems

In 2024 someone proposed a class-based scoring system for debating CVs. Every tournament gets sorted into one of eleven discrete classes (S-WUDC at the top, S-E at the bottom). Each (class, achievement) pair has a hand-set point value. Your career score is the average of your top-N achievements, with ghost copies padding the lower slots. We ported it for a side-by-side comparison and then killed the surface. The math itself is reasonable and reads like a debating resume, which is part of its appeal. "I broke at WUDC, SF at EUDC, top-20 at Cambridge IV" maps cleanly to discrete classes and points. The problems are mostly philosophical. Class boundaries are arbitrary (why exactly 72 WUDC-breaker speakers for S-AAA+, why not 75?). The point values are someone's opinion. Two tournaments inside the same class can have very different actual strength, which our continuous strengthZ catches and the discrete system flattens. And running it alongside our composite created the obvious question: "which one is the real score?" That question doesn't have a clean answer because the two systems are measuring different things. If we ever resurface the class view, it'll be on a page like this one as an educational comparison. Not on individual profiles.

So which is right?

None of them are right. They're different ways of summarising the same underlying performance. The interesting question isn't "which algorithm". It's "which question are we asking". Our composite asks "how good were they, on average, recently". The discrete-class CV asks "how impressive is their achievement list". Elo would ask "who would beat whom in a head-to-head". Plackett-Luce asks "what generative model best explains the ranks we observed". Different questions get different answers. Picking a method without thinking about the question is just picking a method. The reassuring part: the same 30 or so people end up at the top across every reasonable method. Whatever you build, the WUDC champions and finalists of the last decade come out near the top. The choice of algorithm shifts ranks by 1 to 3 spots in the middle of the leaderboard. The fights at the margin are the only ones the algorithm choice actually decides. And those fights are usually inside the uncertainty band of whatever method you used. That doesn't mean any method works. A bad method can produce nonsense (we have one in our git history: Plackett-Luce's non-converging top-5s). It does mean that once you're past the "this method is broken" floor, the disagreements between methods at the top are smaller than people instinctively believe.

Why we wrote this

The methodology page exists so anyone can check the math. The math itself isn't where the interesting choices are. The interesting choices are upstream: which tournaments count, how to handle recency, what to do with the 1-tab-elite, whether to display percentile or raw rating, whether to publish provisional names at all. Those decisions move the leaderboard more than the choice between Massey and Plackett-Luce ever could. When you read these ratings, treat the algorithm as a detail. The interesting questions are about the questions.

How the leaderboards relate

Shared data, shared recency

The three leaderboards share source data but use independent strength signals, because a tournament's debater difficulty and judge difficulty aren't the same thing. A tournament has TWO strength scores: • Debater-bubble strength. How hard the bubble teams' last-2-rounds path was. Used as the weighting multiplier in the debater rating computation. Shown in the tournaments pane on the leaderboard. • Judge-field strength. Mean rating-percentile of every judge in the adjudication pool. Used as the weighting multiplier in the judge rating computation. Not currently surfaced as its own ranking, but it's why the judge ladder reflects judging-field quality rather than debater-bubble difficulty. These can diverge meaningfully. WUDC has elite debaters AND elite judges, so both scores rank high. A smaller regional tab might have a brutal debate bubble and a regional adjudication pool: high debater strength, lower judge strength. The methodology measures each honestly without forcing them to track each other. When we recompute the leaderboards, debater ratings › tournament debater-strength › updated debater ratings (one iteration; converges within a single pass since the rolling speaker ratings don't depend on tournament strength). Judge ratings › tournament judge-strength › updated judge ratings (we iterate this 2-3 times for convergence since judge field strength uses prior judge ratings as input).

Correct or hide your name

Email hello@buildacase.ca to fix a wrong name or result. To hide your name, see leaderboard privacy.