The Pilot: Kappa and Alpha

S20 · Chapter 10 · MC 451 Research Methods in Mass Media

Dr. Alex Leith

Today’s Agenda

  • What a pilot test is, and why coders work alone
  • Why raw agreement flatters a codebook
  • Kappa, alpha, and reading the number
  • Diagnosing a failure and re-piloting

Putting the Codebook on Trial

Lab session

  • Tuesday you drew a sample, today the codebook earns its status
  • We run a pilot and turn agreement into a number
    • We watch a real codebook fail, and diagnose why
    • You leave with a coefficient for your own instrument

This session can send you back to Chapter 8, and that is fine.

What a Pilot Test Is

  • A trial run before any real coding begins
    • Take a subset of the sample, a hundred messages
    • Two people code it from the same codebook
    • Then their codes are compared
  • The question: do the same rules produce the same classifications?

Independently, and Why

  • Genuine isolation: no conferring, no comparing
    • Not suspicion: the pilot measures the codebook, not the coders
  • Discuss a hard message and you measure memory
    • Only independent coding tests the written rules
  • Reliability is consistency, whoever holds the codebook

The Obvious Number

  • The two coders agreed on 74 of 100
    • Seventy-four percent sounds respectable
  • It is also misleading
    • Seeing why is the key idea of the chapter
  • Some agreement happens for no good reason at all

Your Turn

  • Is 74 percent good enough to start on? Commit to an answer
  • Two coders flipping coins on two categories agree how often?
  • What if both use one category 90 percent of the time?
  • What number would satisfy you, and where did it come from?

Why Raw Agreement Misleads

  • Cohen’s 1960 paper used two clinicians sorting patients
    • Each assigns categories blindly, at the observed rates
    • Coordinating on nothing, they still agree often
  • That agreement reflects the number of categories
    • It reflects no judgment at all
  • Raw agreement counts coincidence and insight alike

Cohen’s Kappa

kappa = (observed agreement - expected agreement) / (1 - expected agreement)

Kappa starts from the agreement you observed, subtracts what chance alone would have produced, and credits the codebook only with the difference.

  • No better than chance: the numerator is zero, so kappa is zero
  • Perfect agreement: kappa is one
  • Built for two coders and nominal categories

Krippendorff’s Alpha

  • The general instrument, when kappa will not fit
    • Any number of coders, not just two
    • Any level of measurement, nominal through ratio
    • Datasets with missing codes
  • For a two-coder nominal pilot they agree
    • Kappa is the more familiar choice

Running the Check

v2v::reliability(pilot$coder_a, pilot$coder_b, method = "kappa")

reliability() takes the two coders’ columns of codes and returns the coefficient, with method = "kappa" or method = "alpha" selecting which one.

The $ pulls a single named column out of a table.

The Result

Inter-coder reliability: Cohen's kappa

  Messages coded   100
  Cohen's kappa    0.236
  • Observed agreement was 0.74, the respectable-looking number
  • Chance alone predicts about 0.66, given the flag rates
  • The formula returns 0.236

The raw agreement was almost all coincidence.

Two Rules, One Name

  • The codebook said “high-energy” and never defined it
    • Coder A read it as carrying a Twitch emote
    • Coder B read it as shouting, so all capitals
  • Both readings are reasonable
    • Both coders believed they coded one variable
    • Emotes and all-caps overlap enough to agree by accident

Reading the Number

Labels have long been attached to ranges (Landis & Koch, 1977):

  • 0.0 to 0.2: slight
  • 0.2 to 0.4: fair
  • 0.4 to 0.6: moderate
  • 0.6 to 0.8: substantial
  • Above 0.8: almost perfect

By these labels 0.236 is “fair.” Published work treats below 0.70 as too low to proceed on, with 0.80 a comfortable floor.

Diagnose, Revise, Re-Pilot

  • A kappa of 0.236 is an instruction, not a result
  • The two coders finally meet
    • They walk the messages they classified differently
    • The disagreements are the data
  • Repair the rule in writing, then pilot again

Fresh messages, independent coding, another coefficient.

Common Errors Today

  • object 'pilot' not found: the codes are not in one table
  • Columns of different lengths: someone skipped a message
  • Kappa of exactly 1: you passed the same column twice
  • Kappa near 0 with high agreement: the loose-definition signature
  • NA in the result: missing codes, so use alpha

Run ?v2v::common_errors for the ones specific to this data.

Checkpoint

You should now have:

  • A hundred-message pilot, coded independently by two people
  • A v2v::reliability() call that runs and returns a number
  • A written note on whether it clears 0.70
  • The list of messages the coders split on, your revision list
  • If you failed, a revised rule for each disagreement

Before Next Time

  • Revise against your disagreement list, then re-pilot
  • Read Chapter 11, Wrangling the Data
  • Next week is R for real
    • Unreadable timestamps, missing variables, two tables apart
  • Due next week: Sampling Plan and Pilot, and Data Wrangling

Questions?