Measurement, Reliability, Validity
S15 · Chapter 8 · MC 451 Research Methods in Mass Media
Today’s Agenda
- NOIR, the four levels of measurement
- What the level lets you conclude later
- Reliability, and how chance inflates it
- Validity, and why it is a separate question
A Decision, Not Vocabulary
- Today looks like vocabulary
- It is a decision with November consequences
- The level decides which tests you may run
- Choose narrowly now and the paper cannot say some things
- By then recoding two thousand messages is not an option
- You leave with a level on every variable
NOIR
- Every operationalized variable has a level
- That level decides what you may later do with it
- Four of them: nominal, ordinal, interval, ratio
- The
stream_log table spans all four
- Its columns: channel, title, game, viewers, date
Nominal
- Categories with no inherent order
game: Fortnite, Just Chatting, Art, Hearthstone
- No arithmetic relates them
channel is nominal, and so is a raw title
- One string is not ranked above another
- You can count and name the most common, never average
Ordinal
- Ordered categories, uneven distances
- Bin titles by length: short, medium, long
- Long outranks short
- The gaps are not guaranteed equal
- Short-to-medium need not match medium-to-long
- You can rank and find the median, but not a true mean
Interval
- Equal distances, arbitrary zero
date is milliseconds since January 1, 1970
- Minute-to-minute distance is constant across the column
- 1970 does not mean “no time”, it is a convention
- Subtract two timestamps for a duration, never a ratio
Ratio
- Equal intervals and a true zero
viewers is ratio: zero means nobody is watching
- So 28,000 really is twice the audience of 14,000
- The full range of arithmetic is permitted
- Most counts are ratio: viewers, characters, messages
One Column, Three Levels
Stream title, measured three ways:
- Raw string: nominal
- Character count: ratio, since a title can have zero
- Binned short, medium, long: ordinal
Level of measurement is not a property the data hands you. It is a property of the decision you make.
Your Turn
- Take one variable from your Definitions Practice
- What level did you land on, and could it be higher?
- What would you have to record instead?
- What would you lose by binning it into categories?
Levels Govern the Statistics
- Nominal gets a chi-square test for association
- Ratio can be averaged, correlated, and modeled
- Choosing nominal where ratio was available narrows you
- Nobody catches it at the time
- It shows up in November as a test you may not run
- On data you already spent forty hours coding
You make the choice before coding a single case.
Two Kinds of Variable
- Some arrive ready-made
viewers is already a number, game already a category
- The work is classification: pick the level and move on
- Others do not exist until you build them
- Message target is not a column in
chat_log
- It exists only after a human reads each message
For the second kind, operationalization is most of the work.
Reliability
- Reliability is consistency: same conditions, same result
- A scale reading 150, then 162, then 147 is worthless
- Whatever you actually weigh
- Here it means inter-coder reliability
- Two trained coders, same messages, independently
- Without it, nothing downstream repairs the codebook
Correcting for Chance
- Agreement statistics correct for luck
- Two coders agree sometimes by chance alone
- Cohen’s kappa and Krippendorff’s alpha
- Raw percent agreement flatters you
- Especially when one category is common
A codebook’s job is to make independent coders agree.
Validity
- Validity is accuracy: does it capture what it claims?
- Face validity: does it look like the right instrument?
- Content validity: does it cover the whole concept?
- Construct validity: the most demanding of the three
- Does it behave as theory says it should?
They Are Independent
- Reliable and invalid: a scale ten pounds light
- Same answer every time, wrong every time
- Valid and unreliable: accurate on average
- So inconsistent that a single reading is unusable
- Figure 8.2 draws four targets, tight-on-center to scattered
Consistency is the easiest thing to mistake for accuracy.
The Principle to Hold
- Reliability is necessary, not sufficient
- An inconsistent measure cannot be capturing anything
- So unreliability rules validity out
- A consistent measure can still measure the wrong thing
- You need both, and reliability comes first
A codebook is the instrument that produces it.
Before Next Time
- Due this week: Definitions Practice
- Include the level of measurement for every variable
- Bring your three variables, categories, and edge-case log
- We meet as a studio, and you leave with a draft codebook
- Re-read the model codebook in Chapter 8
- Have your field notes open and searchable