The Sample and the Pilot

Week 10 · Chapter 10 · MC 501 Research Methods for Mass Communications

Dr. Alex Leith

Today’s Agenda

  • Why unbiased is not the same as representative
  • Stratified sampling, and where power meets the data
  • The pilot, and why independence carries it
  • Kappa, alpha, and what a failure instructs

Two Jobs This Week

Chapter 10

  • Draw the sample: the messages you will actually code
  • Test the codebook: the pilot that makes it an instrument
    • It has standing to send you back to Chapter 8
    • Being sent back is the pilot working
  • Both are due together as the Sampling Plan and Pilot

Why You Sample

  • A hundred hours for 35,267 messages, at ten seconds each
    • The full 2018 collection would outlast your degree
  • A census is right when the population is small
    • The stream fixture ships whole, at 32,000 snapshots
  • Chat is the opposite, and the ordinary condition

A sample makes the work finite without making conclusions worthless.

Unbiased Is Not Representative

  • Simple random sampling gives every message equal chance
    • Its virtue is real: no researcher tilt is built in
  • Chat volume is severely right-skewed
    • The busiest channel logged over a million that week
    • The median about 1,440, the quietest one
  • Forty-three of the fifty are gaming channels

Draw at random and you get giant gaming channels.

Stratified Sampling

  • Divide into subgroups, then draw within each
    • Every stratum gets a place, in the proportion you decide
  • The strata must matter for the question
    • Here the stratifying variable is the channel
  • Gaming status is a property of the channel
    • So every non-gaming channel is in by construction

The Corpus Is a Stratified Design

  • Its fifty were not the fifty busiest
    • Eight were fixed anchors chosen for variety
    • Bob Ross among them, the deliberate non-gaming case
  • The other 42 came from ten volume bands
    • Every band was sampled

The skew was not sampled away. It was sampled across, on purpose.

Small Strata, Take All of It

  • Ask 30 of a channel that produced 9, and there is no 30
    • The rule the tools follow: take everything it has
  • Nineteen of the fifty hold fewer than a thousand
  • A stratum gives its full weight, or its full self

Document this in your sampling plan. It is a decision, not an accident.

Drawing the Sample

set.seed(409)

coding_sample <- v2v::sample_messages(
  chat,
  by = channel,
  n  = 30
)

by = channel makes each channel a stratum and n = 30 asks for thirty from each, aiming at roughly 1,500 messages across fifty channels.

set.seed() Is Not Optional

  • A draw is random: run it twice, get two samples
    • set.seed() fixes the generator’s starting point
  • Without it your sample is unreproducible
    • Nobody can check your results, including you
  • The number is arbitrary, and 409 means nothing

What matters is that one is set and recorded, so the draw is frozen.

Where Power Meets the Data

  • Sampling is not only about capturing variation
    • Your prospectus specified a minimum n
    • This is the week that number meets your dataset
  • Need 128 and the data gives 10,000: stratification solves it
  • Need 128 and the data gives 80: the study is underpowered

Then you change the design, or justify the limitation explicitly.

Lakens on What Honesty Demands

“Do not try to spin a story where it looks like a study was highly informative when it was not. Instead, transparently evaluate how informative the study was given effect sizes that were of interest, and make sure that the conclusions follow from the data.”

Lakens (2022, p. 11)

Purposive sampling improves construct validity and does nothing for power. Conflating them gives a study that is coherent and indefensible.

The Pilot Test

  • A trial run before the real coding begins
    • A hundred messages is a reasonable size
    • Two people, the same codebook, then a comparison
  • The question it answers
    • Applied independently, do these rules give the same codes?

Reliability is the property that the codebook answers the same whoever holds it.

Independence Carries the Logic

  • No conferring, no comparing, no checking as they go
    • Not suspicion: the pilot measures the codebook
  • Discuss a hard message and you measure their memory
    • Not the clarity of the written rules
  • The rules must carry a classification alone

Build the isolation into your protocol, and say so in the write-up.

Why Raw Agreement Is Not Enough

  • Two coders match on 74 of 100, which sounds respectable
  • Cohen’s 1960 example: two clinicians assigning blindly
    • At the observed rates, they still agree often
  • That reflects the number of categories, not judgment

Few categories, or one dominant category, and coincidences pile up fast.

Reliability First

Returning to Hayes and Krippendorff

“Conclusions from such data can be trusted only after demonstrating their reliability.”

Hayes and Krippendorff (2007, p. 77)

  • You have a number now. Does it clear the bar?
  • What would you report if it does not?
  • Whose trust is the standard protecting?

Informative, Not Impressive

“Do not try to spin a story where it looks like a study was highly informative when it was not.”

Lakens (2022, p. 11)

  • Say what your pilot could not settle
  • What does that cost the White Paper?

Cohen’s Kappa

  • Built for two coders and a nominal codebook (Cohen, 1960)
  • Observed agreement, less chance, over the room left
    • kappa = (observed - expected) / (1 - expected)
  • No better than chance: kappa is zero, whatever the raw rate
  • Perfect agreement: kappa is one

The codebook is credited only for what luck did not do.

Krippendorff’s Alpha

  • The more general instrument (Krippendorff, 2018)
    • Any number of coders, any level of measurement
    • Datasets with missing codes
  • Three coders, or an ordinal variable: alpha
  • For a two-coder nominal pilot they agree

v2v::reliability() computes either, selected by method.

Computing Alpha for Your Pilot

pilot <- coding_sample %>%
  slice_head(n = 100)

v2v::reliability(
  pilot$coder_a,
  pilot$coder_b,
  method = "alpha"
)

Take the first hundred messages, hand the two coders’ columns to reliability(), and ask for alpha rather than kappa.

A Worked Failure

  • The codebook said “high-energy” and never defined it
    • Coder A flagged messages carrying a Twitch emote
    • Coder B flagged messages containing all capitals
  • Both readings are reasonable, and both believed they coded one variable
Inter-coder reliability: Cohen's kappa

  Messages coded   100
  Cohen's kappa    0.236

What 0.236 Diagnoses

  • Observed 0.74, chance 0.66, kappa 0.24
  • Emote and all-caps messages overlap in Twitch chat
    • Two coders keying on different features often land together
    • They never coded the same concept
  • A loose codebook let two careful people measure two things

That is what kappa exposed and raw agreement concealed.

Reading Against a Standard

  • The Landis and Koch labels (1977)
    • 0.0 to 0.2 slight, 0.2 to 0.4 fair, 0.4 to 0.6 moderate
    • 0.6 to 0.8 substantial, above 0.8 almost perfect
  • Those labels are a vocabulary, not a law
  • Published work treats below 0.70 as too low, with 0.80 comfortable

Name your threshold in advance, not after seeing the number.

When Reliability Fails

  • 0.236 is an instruction, not a result
    • The coders finally meet and walk every disagreement
    • The disagreements are the data
  • Repair: state operationally what counts
  • Then pilot again, on fresh messages

Pilot, diagnose, revise, repeat, until it clears your threshold.

The Deliverable

Sampling Plan and Pilot · 75 points

  • The sampling frame and strata, and why each matters
  • The target n per stratum, the floor rule, and the seed
  • The pilot protocol: how many units, how coders were isolated
  • The coefficient, which one, and the threshold set beforehand
  • If it failed: the diagnosis and the revision it produced

Before Week 11

  • Submit the Sampling Plan and Pilot, with the computation
  • Read Chapter 11 with the graduate toggle on
    • Plus the assigned article: Wickham (2014), “Tidy data”
  • Write a journal entry, 450 to 500 words
  • Bring a laptop with R, VS Code and v2v working

Questions?