The Sample

S19 · Chapter 10 · MC 451 Research Methods in Mass Media

Dr. Alex Leith

Today’s Agenda

  • Why a census is off the table
  • What simple random sampling cannot fix
  • Strata, and choosing ones that matter
  • Drawing the sample reproducibly

A Hundred Hours of Coding

Do the arithmetic first

  • One message every ten seconds, by hand
    • The 35,267 in the working sample take a hundred hours
    • The full 2018 collection would outlast your degree
  • So you will code a sample

Drawing it well is the whole of today.

Population, Sample, Census

  • Population: the full set your conclusions are about
    • Here, every message in the corpus that week
    • In your study, whatever your prospectus said
  • Sample: the subset you actually examine
  • Census: every case, right when the population is small

The stream table is close to a census: roughly 32,000 snapshots, all kept.

Chat Is the Opposite Case

  • Far too much to code, so no census
    • The ordinary condition of content analysis
  • A sample makes the work finite
    • Without making the conclusions worthless
  • 1,500 messages can be sound or misleading

The difference is method, which means you decide it rather than get lucky.

Simple Random Sampling

  • Every message has an equal chance
    • Draw the number you need, and stop
  • Its virtue is that it is unbiased
    • No message is favored over any other
    • The sample carries no tilt built in by the researcher
  • A real guarantee, and for Twitch chat not enough

Why That Is Not Enough

The shape of the data

  • Chat volume is severely concentrated
    • Busiest channel: more than a million messages that week
    • Median channel: about 1,440
    • Quietest channel: one
  • That shape is right-skewed, a long thin tail

Draw at random and almost everything comes from the giants.

Your Turn

  • Draw 1,000 at random: how many from the one-message channel?
  • Is that sample biased? Is it representative?
  • Are those the same question?
  • Your study has a comparison at its center. Which side is rare?

The Second Dimension

  • Forty-three of the fifty channels are gaming channels
    • Bob Ross’s painting channel is the clearest exception
  • A random sample would be almost all gaming chat
    • Not from bias, but because that is where the messages are
  • For a study comparing the two, that is close to fatal

The Trap, Stated Plainly

  • Unbiased and still unrepresentative
    • It faithfully reproduces the imbalance at sample size
  • The study needs enough of every group to compare
  • Unbiased is a property of the procedure
    • It is not a guarantee about the result

Noticing this before you draw is the difference between a study and a mess.

Stratified Sampling

  • Do not draw from one undifferentiated pool
    • Divide the population into subgroups called strata
    • Draw from within each stratum separately
  • Every stratum gets a place in the sample
    • In the proportion you decide, not the one handed to you

Choosing Strata That Matter

  • Strata follow the question, not convenience
    • For the chat study that is the channel
    • A fixed number per channel admits small channels fully
  • Gaming status is a property of the channel
    • So stratifying by channel rescues the comparison
    • Every non-gaming channel is in by construction

The Corpus Is Itself Stratified

  • The fifty were not the fifty busiest
    • Eight were fixed anchors chosen for variety
    • Bob Ross among them, the deliberate non-gaming case
  • The other forty-two came from ten volume bands
    • Every band was sampled
  • The skew was not sampled away, it was sampled across

When a Stratum Is Too Small

  • Ask thirty of a channel that produced nine
    • The rule the tools follow: take everything it has
  • Nineteen of the fifty hold fewer than a thousand
    • A few hold only a handful
  • A stratum contributes its full weight, or its full self

Drawing the Sample

set.seed(409)

coding_sample <- v2v::sample_messages(chat, by = channel, n = 30)

by = channel makes each channel a stratum and n = 30 asks for thirty messages from each, which across fifty channels aims at roughly 1,500 messages.

That is large enough to support the study and small enough to code by hand.

Why set.seed() Is Not Optional

  • A draw is random: run it twice, get two samples
  • set.seed() fixes the generator’s starting point
    • The draw then comes out identically every time
  • Without it nobody can check your results
    • Including you, in six months

The number 409 is arbitrary. What matters is that one is set and recorded.

Looking Ahead to Thursday

  • You have a sample, and an untested codebook
  • Thursday is the pilot
    • Two coders, the same hundred messages, no conferring
    • We compute Cohen’s kappa and watch a codebook fail
  • Read Chapter 10’s second half before Thursday
  • Bring your codebook, because it is what goes on trial

Questions?