Measurement, Reliability, Validity

S15 · Chapter 8 · MC 451 Research Methods in Mass Media

Dr. Alex Leith

Today’s Agenda

  • NOIR, the four levels of measurement
  • What the level lets you conclude later
  • Reliability, and how chance inflates it
  • Validity, and why it is a separate question

A Decision, Not Vocabulary

  • Today looks like vocabulary
    • It is a decision with November consequences
  • The level decides which tests you may run
    • Choose narrowly now and the paper cannot say some things
    • By then recoding two thousand messages is not an option
  • You leave with a level on every variable

NOIR

  • Every operationalized variable has a level
    • That level decides what you may later do with it
  • Four of them: nominal, ordinal, interval, ratio
  • The stream_log table spans all four
    • Its columns: channel, title, game, viewers, date

Nominal

  • Categories with no inherent order
    • game: Fortnite, Just Chatting, Art, Hearthstone
    • No arithmetic relates them
  • channel is nominal, and so is a raw title
    • One string is not ranked above another
  • You can count and name the most common, never average

Ordinal

  • Ordered categories, uneven distances
    • Bin titles by length: short, medium, long
    • Long outranks short
  • The gaps are not guaranteed equal
    • Short-to-medium need not match medium-to-long
  • You can rank and find the median, but not a true mean

Interval

  • Equal distances, arbitrary zero
    • date is milliseconds since January 1, 1970
    • Minute-to-minute distance is constant across the column
  • 1970 does not mean “no time”, it is a convention
  • Subtract two timestamps for a duration, never a ratio

Ratio

  • Equal intervals and a true zero
    • viewers is ratio: zero means nobody is watching
    • So 28,000 really is twice the audience of 14,000
  • The full range of arithmetic is permitted
  • Most counts are ratio: viewers, characters, messages

One Column, Three Levels

Stream title, measured three ways:

  • Raw string: nominal
  • Character count: ratio, since a title can have zero
  • Binned short, medium, long: ordinal

Level of measurement is not a property the data hands you. It is a property of the decision you make.

Your Turn

  • Take one variable from your Definitions Practice
  • What level did you land on, and could it be higher?
  • What would you have to record instead?
  • What would you lose by binning it into categories?

Levels Govern the Statistics

  • Nominal gets a chi-square test for association
  • Ratio can be averaged, correlated, and modeled
  • Choosing nominal where ratio was available narrows you
    • Nobody catches it at the time
    • It shows up in November as a test you may not run
    • On data you already spent forty hours coding

You make the choice before coding a single case.

Two Kinds of Variable

  • Some arrive ready-made
    • viewers is already a number, game already a category
    • The work is classification: pick the level and move on
  • Others do not exist until you build them
    • Message target is not a column in chat_log
    • It exists only after a human reads each message

For the second kind, operationalization is most of the work.

Reliability

  • Reliability is consistency: same conditions, same result
    • A scale reading 150, then 162, then 147 is worthless
    • Whatever you actually weigh
  • Here it means inter-coder reliability
    • Two trained coders, same messages, independently
  • Without it, nothing downstream repairs the codebook

Correcting for Chance

  • Agreement statistics correct for luck
    • Two coders agree sometimes by chance alone
  • Cohen’s kappa and Krippendorff’s alpha
    • The two you will meet
  • Raw percent agreement flatters you
    • Especially when one category is common

A codebook’s job is to make independent coders agree.

Validity

  • Validity is accuracy: does it capture what it claims?
  • Face validity: does it look like the right instrument?
  • Content validity: does it cover the whole concept?
  • Construct validity: the most demanding of the three
    • Does it behave as theory says it should?

They Are Independent

  • Reliable and invalid: a scale ten pounds light
    • Same answer every time, wrong every time
  • Valid and unreliable: accurate on average
    • So inconsistent that a single reading is unusable
  • Figure 8.2 draws four targets, tight-on-center to scattered

Consistency is the easiest thing to mistake for accuracy.

The Principle to Hold

  • Reliability is necessary, not sufficient
    • An inconsistent measure cannot be capturing anything
    • So unreliability rules validity out
  • A consistent measure can still measure the wrong thing
  • You need both, and reliability comes first

A codebook is the instrument that produces it.

Before Next Time

  • Due this week: Definitions Practice
    • Include the level of measurement for every variable
  • Bring your three variables, categories, and edge-case log
    • We meet as a studio, and you leave with a draft codebook
  • Re-read the model codebook in Chapter 8
  • Have your field notes open and searchable

Questions?