Measurement, Reliability, and the Codebook
Week 8 · Chapter 8 · MC 501 Research Methods for Mass Communications
Today’s Agenda
- NOIR, and what each level permits
- Reliability and validity, kept apart
- The four-step reliability protocol
- The codebook, and its five parts
Tonight
Week 8 · Chapter 8
- This looks like vocabulary and is a set of decisions
- The level decides which tests you may run in November
- By then recoding two thousand messages is not an option
- Then the codebook
- The instrument the back half of this course runs on
- Due tonight: Definitions Practice
NOIR
- Every operationalized variable has a level
- It determines what you are later allowed to do
- Nominal, ordinal, interval, ratio
stream_log spans all four
- Channel, title, game, viewers, date
Levels are not a property the data hands you. They are a property of your decision.
Nominal
- Categories with no inherent order
game: Fortnite, Just Chatting, Art, Hearthstone
- No arithmetic relates them
channel is nominal, and so is a raw title
- You can count and name the most common
You cannot average it. The mean of Fortnite and Art is not a category.
Ordinal
- Ordered categories, unequal distances
- Bin titles by length: short, medium, long
- Long outranks short
- The gaps are not guaranteed equal
- Short-to-medium need not match medium-to-long
- You can rank and take a median, but not a true mean
Interval
- Equal distances, arbitrary zero
date counts milliseconds since January 1, 1970
- Minute-to-minute distance is constant across the column
- 1970 does not mean “no time”, it is a convention
- Subtract two timestamps for a duration, never a ratio
Ratio
- Equal intervals and a true zero
viewers is ratio: zero means nobody is watching
- So 28,000 really is twice the audience of 14,000
- The full range of arithmetic is available
- Most counts are ratio: viewers, characters, messages
One Column, Three Levels
- Raw string: nominal
- Character count: ratio, since a title can have zero
- Binned short, medium, long: ordinal
One column, three levels, depending entirely on how you operationalize it.
The Cost of Choosing Low
- Nominal where ratio was available narrows the study
- And it does so before you code a single case
- Nobody notices at the time
- It surfaces in November as a test you may not run
- On data you already spent forty hours coding
Two Kinds of Variable
- Some arrive ready-made
viewers is a number, game is a category
- The work is classification: pick the level and move on
- Others do not exist until you build them
- Message target is not a column in
chat_log
- It exists only once a coder reads each message
A coder can only do that consistently from written instructions.
Reliability
- Consistency: the same result under the same conditions
- A scale reading 150, then 162, then 147 is worthless
- Here it means inter-coder reliability
- Two trained coders, same messages, applied independently
- Kappa and alpha correct for chance agreement
Validity
- Accuracy: does it capture what it claims?
- Face validity: does it look like the right instrument?
- Content validity: does it cover the whole concept?
- Construct validity: does it behave as theory predicts?
Validity is a claim you argue for, not a statistic you compute.
The Two Are Independent
- Reliable and invalid: a scale ten pounds light
- The same every time, and wrong every time
- Valid and unreliable: accurate on average
- So inconsistent a single reading is unusable
- Reliability is necessary, not sufficient
You get reliability first, because the codebook produces it.
The Protocol, First Half
A complete protocol has four steps, and the order is not optional.
- The second coder receives the codebook alone, with no walk-through, and codes a training set of 20 to 30 units independently.
- The two coders compare and discuss every disagreement. No production data is coded at this stage.
The urge to explain what you meant is what the protocol forbids.
The Protocol, Second Half
- The codebook is revised, adding a decision rule for every disagreement.
- The second coder codes a fresh reliability sample, separate from the training set. That agreement is the kappa you report.
- Reporting the training set’s agreement is a serious error
- It is contaminated by the reconciliation conversation
- Report more than one statistic (Lombard et al., 2002)
Why the Protocol Exists
“Conclusions from such data can be trusted only after demonstrating their reliability.”
Hayes and Krippendorff (2007, p. 77)
- The four steps are the demonstration that sentence demands
- Nothing downstream outranks this dependency
- A finding on an untested codebook is a number without a warrant
Never Considered Valid
Lombard et al. (2002)
“It is widely acknowledged that intercoder reliability is a critical component of content analysis and (although it does not ensure validity) when it is not established, the data and interpretations of the data can never be considered valid.”
Lombard et al. (2002, p. 589)
- Your study has one coder. Now what?
- Is that sentence a standard or a threat?
- What can you still claim honestly?
The Researcher as Coder
“Coding must be done independently and without consultation or guidance. If possible, the researcher should not be a coder.”
Lombard et al. (2002, p. 601)
- You are both. Name the specific risk
- What would catch it if it happened?
- Their rule of thumb for a reliability sample is 30 units
The Threshold and the Sample
- The standard for kappa is 0.70 or higher (Landis & Koch, 1977)
- Below it the codebook needs revision, not a caveat
- At least 50 coded units in the reliability sample
- For rare codes, 100 or more are preferred
v2v::reliability() computes kappa with labels
Two coders agreeing on 40 of 50: does that clear the bar?
The Codebook Is the Instrument
- Not paperwork assembled to satisfy a requirement
- The useful analogy is source code
- If coding were a program, the codebook is the program
- It specifies every operation and every conditional
- Two coders running it on one input agree
- The heart of a content analysis (Neuendorf, 2017)
Five Parts of a Codebook
- Unit of analysis: exactly what one coded case is
- Variables and categories: both definitions, and each value
- Decision rules: what to do when a case will not sort
- Examples: two or three prototypes per category
- Special cases: recurring complications, thin at first
Exhaustive and Mutually Exclusive
- Exhaustive: every case can be coded
- The catch-all category is what guarantees it
- Mutually exclusive: every case fits exactly one
- Overlap destroys reliability, because coders split forever
- The fix: sharper definitions, or a precedence rule
More than one case in ten in the catch-all means the scheme is incomplete.
Draft Codebook, Unit and Target
Unit of analysis. A single chat message, defined as one row of chat_log: one sender, one message string, one timestamp. Each message is coded independently of those around it, except where a decision rule says otherwise.
Variable 1: message target. Categories: directed at streamer, directed at another viewer, broadcast to the room, unclassifiable. Level: nominal.
Draft Codebook, Two More
Variable 2: message length. Conceptual: the verbal extent of the message. Operational: the count of characters, including spaces and emote tokens. Derived rather than judged. Level: ratio.
Variable 3: contains emote. Operational: yes if the message contains at least one token from the project’s emote list. Level: nominal.
Two manifest measures, one latent. The mix is deliberate.
The Five Decision Rules
- Precedence: explicit address wins over performance
- Emote-only: broadcast, unless it carries an at-mention
- Copypasta: coded by content, like any other message
- Non-English: code target if structure determines it
- Bot accounts: unclassifiable, noted for possible exclusion
Before Week 9
- Read Chapter 9, and bring a laptop with R installed
- Re-read Munafo et al. (2017) on reporting and usability
- Due next week: Extended Codebook and Reliability Protocol
- Planned sample size, thresholds, and revision triggers
- Have your three variables written out with rules
- Write a journal entry, 450 to 500 words