The Pilot: Kappa and Alpha
S20 · Chapter 10 · MC 451 Research Methods in Mass Media
Today’s Agenda
- What a pilot test is, and why coders work alone
- Why raw agreement flatters a codebook
- Kappa, alpha, and reading the number
- Diagnosing a failure and re-piloting
Putting the Codebook on Trial
Lab session
- Tuesday you drew a sample, today the codebook earns its status
- We run a pilot and turn agreement into a number
- We watch a real codebook fail, and diagnose why
- You leave with a coefficient for your own instrument
This session can send you back to Chapter 8, and that is fine.
What a Pilot Test Is
- A trial run before any real coding begins
- Take a subset of the sample, a hundred messages
- Two people code it from the same codebook
- Then their codes are compared
- The question: do the same rules produce the same classifications?
Independently, and Why
- Genuine isolation: no conferring, no comparing
- Not suspicion: the pilot measures the codebook, not the coders
- Discuss a hard message and you measure memory
- Only independent coding tests the written rules
- Reliability is consistency, whoever holds the codebook
The Obvious Number
- The two coders agreed on 74 of 100
- Seventy-four percent sounds respectable
- It is also misleading
- Seeing why is the key idea of the chapter
- Some agreement happens for no good reason at all
Your Turn
- Is 74 percent good enough to start on? Commit to an answer
- Two coders flipping coins on two categories agree how often?
- What if both use one category 90 percent of the time?
- What number would satisfy you, and where did it come from?
Why Raw Agreement Misleads
- Cohen’s 1960 paper used two clinicians sorting patients
- Each assigns categories blindly, at the observed rates
- Coordinating on nothing, they still agree often
- That agreement reflects the number of categories
- It reflects no judgment at all
- Raw agreement counts coincidence and insight alike
Cohen’s Kappa
kappa = (observed agreement - expected agreement) / (1 - expected agreement)
Kappa starts from the agreement you observed, subtracts what chance alone would have produced, and credits the codebook only with the difference.
- No better than chance: the numerator is zero, so kappa is zero
- Perfect agreement: kappa is one
- Built for two coders and nominal categories
Krippendorff’s Alpha
- The general instrument, when kappa will not fit
- Any number of coders, not just two
- Any level of measurement, nominal through ratio
- Datasets with missing codes
- For a two-coder nominal pilot they agree
- Kappa is the more familiar choice
Running the Check
v2v::reliability(pilot$coder_a, pilot$coder_b, method = "kappa")
reliability() takes the two coders’ columns of codes and returns the coefficient, with method = "kappa" or method = "alpha" selecting which one.
The $ pulls a single named column out of a table.
The Result
Inter-coder reliability: Cohen's kappa
Messages coded 100
Cohen's kappa 0.236
- Observed agreement was 0.74, the respectable-looking number
- Chance alone predicts about 0.66, given the flag rates
- The formula returns 0.236
The raw agreement was almost all coincidence.
Two Rules, One Name
- The codebook said “high-energy” and never defined it
- Coder A read it as carrying a Twitch emote
- Coder B read it as shouting, so all capitals
- Both readings are reasonable
- Both coders believed they coded one variable
- Emotes and all-caps overlap enough to agree by accident
Reading the Number
Labels have long been attached to ranges (Landis & Koch, 1977):
- 0.0 to 0.2: slight
- 0.2 to 0.4: fair
- 0.4 to 0.6: moderate
- 0.6 to 0.8: substantial
- Above 0.8: almost perfect
By these labels 0.236 is “fair.” Published work treats below 0.70 as too low to proceed on, with 0.80 a comfortable floor.
Diagnose, Revise, Re-Pilot
- A kappa of 0.236 is an instruction, not a result
- The two coders finally meet
- They walk the messages they classified differently
- The disagreements are the data
- Repair the rule in writing, then pilot again
Fresh messages, independent coding, another coefficient.
Common Errors Today
object 'pilot' not found: the codes are not in one table
- Columns of different lengths: someone skipped a message
- Kappa of exactly 1: you passed the same column twice
- Kappa near 0 with high agreement: the loose-definition signature
NA in the result: missing codes, so use alpha
Run ?v2v::common_errors for the ones specific to this data.
Checkpoint
You should now have:
- A hundred-message pilot, coded independently by two people
- A
v2v::reliability() call that runs and returns a number
- A written note on whether it clears 0.70
- The list of messages the coders split on, your revision list
- If you failed, a revised rule for each disagreement
Before Next Time
- Revise against your disagreement list, then re-pilot
- Read Chapter 11, Wrangling the Data
- Next week is R for real
- Unreadable timestamps, missing variables, two tables apart
- Due next week: Sampling Plan and Pilot, and Data Wrangling