Interrater Reliability

A shared label in a shared field doesn't mean two people are making the same judgment when they choose it. Aggregating the labels without checking that is a real, measurable mistake, not a caveat.

Title card reading 'Interrater Reliability' in green type on a large cream circle, framed by abstract organic shapes in terracotta, gold, dusty blue, sage, and cream on a textured cream background

A pipeline report says 40% of open deals sit in "proposal" stage. One rep only moves a deal to "proposal" once a signed statement of work has actually gone out. Another rep moves a deal to "proposal" the moment pricing comes up on a call. Both of them filled in the CRM field correctly, by their own definition of what the field means. The 40% is an accurate sum of individually accurate entries, and it still doesn't mean what it looks like it means, because "proposal" was never one thing to begin with.

A shared label doesn't mean a shared judgment

Two people can use the exact same word for the exact same field and still be measuring different things, if the rule each of them uses to decide when the word applies is different. This isn't a data-entry error. Neither rep did anything wrong by their own standard, and neither field is technically false. What's missing is the thing that would make the two entries comparable in the first place: some check that both reps, looking at the same deal, would have picked the same stage.

That check doesn't happen automatically just because the field is a dropdown with a fixed set of options. A shared set of labels constrains what anyone can type. It says nothing about whether two different people would type the same one for the same underlying situation. Treat every "proposal" entry across the whole pipeline as equivalent without checking that, and the resulting percentage is really measuring something closer to which rep happened to touch which deal, not how far along the deals actually are.

Where kappa gets its name

Measurement science has worked out ways to check this rather than just assume it, and one of the most widely used comes from Jacob Cohen. In 1960, Cohen published "A Coefficient of Agreement for Nominal Scales" (Educational and Psychological Measurement), introducing what's now called Cohen's kappa: a statistic for how much two raters actually agree when independently categorizing the same set of cases into a fixed set of labels, corrected for the agreement they'd get from chance alone. That correction is the whole point, and the chance baseline isn't a fixed fifty-fifty guess, it's computed from how often each rater actually uses each label, their own marginal rates. If both raters lean heavily toward the same one or two labels already, chance alone predicts a high agreement rate before either of them has judged anything, which means a raw 90% agreement number can still mean almost nothing was confirmed beyond what their habits already implied. Kappa answers the narrower question that matters: agreement above and beyond what each rater's own labeling habits would already produce.

Cohen's original kappa treats every mismatch the same, "qualified" recorded against "proposal" counts as a disagreement exactly the same size as "qualified" against "closed," and a CRM stage field isn't actually that kind of category set. Stages run in a funnel order, so a one-stage miss and a four-stage miss aren't obviously the same mistake. Unweighted kappa still answers a real, narrower question, whether reps land on the exact same stage, but where the distance between stages is supposed to matter, Cohen's own follow-up, "Weighted Kappa: Nominal Scale Agreement with Provision for Scaled Disagreement or Partial Credit" (Psychological Bulletin, 1968), extends the same chance-correction logic to give partial credit for a near miss instead of scoring every disagreement identically.

The honest limit: kappa needs something a routine sync was never built to produce

This isn't an argument that CRM data is untrustworthy, or that stage fields should be abandoned. A field correctly recording what a rep entered is a real, checkable fact, exactly what a system that syncs deal records is supposed to capture, and it not being reliable in the interrater sense doesn't make it inaccurate in the recording sense. Those are different properties, and only one of them is what a sync job promises.

The harder property is one a routine sync was never positioned to produce on its own. Measuring interrater reliability the way Cohen described it requires multiple raters independently judging the same cases, two reps staging the identical sample of deals without seeing each other's answer, so the agreement between them can actually be computed. A live pipeline doesn't generate that by default: each deal has exactly one rep entering exactly one stage, once, and there's no second, independent judgment on the same deal to compare it against. Knowing reliability is a real question doesn't hand you the answer. Answering it takes a deliberate check that ordinary CRM usage doesn't produce as a side effect.

Where this fits at Brief

Brief's own description of its agents is explicit that Pipeline Agent belongs to the group dealing in observed fact: a deal is at a stage or it is not, so the agent keeps that record true rather than second-guessing it. That's the correct job for what a stage field actually is, a real entry someone made. This post's point sits one layer above that job, not against it: whether the stage a given rep entered means the same thing as the same stage entered by a different rep is a separate, harder measurement question, one a fact-keeping sync was never built to answer and doesn't need to in order to do its own job well.

Pull up your team's pipeline by stage. If two different reps had independently looked at the same batch of deals, is there any reason to expect they'd have sorted them the same way?

Frequently asked questions

What is interrater reliability? The general question of how consistently different people agree when each independently makes the same kind of judgment on the same cases. It covers more than one way of measuring that consistency; Cohen's kappa is the specific coefficient for two raters sorting cases into a fixed set of labels, correcting for how much agreement their own labeling habits would produce by chance alone, with a weighted version for cases where the labels themselves have an order.

Where does Cohen's kappa come from? Jacob Cohen's 1960 paper "A Coefficient of Agreement for Nominal Scales," which introduced the statistic, subtracting out the agreement expected from chance so that what's left actually reflects real agreement between raters on a categorical judgment. Interrater reliability as a general concept doesn't trace to one single paper the way this specific coefficient does.

Why does a shared dropdown in a CRM not guarantee reliable data? Because a fixed list of options only constrains what any one person can enter, not whether different people apply the same rule to choose between them. Two reps can each fill in a field correctly by their own definition of the label and still mean different things by it, which a shared list of choices can't catch on its own.

Does this mean Pipeline Agent's synced deal data is wrong? No, and it isn't the question Pipeline Agent is answering in the first place. Its documented job stops at a single right value: what a rep actually typed into the field. Whether typing the same word means the same thing coming from two different reps is a question sitting one level up, outside what any record-keeping sync claims to check.

GET TLDR FROM:
← Back to Blog