by = channel makes each channel a stratum and n = 30 asks for thirty from each, aiming at roughly 1,500 messages across fifty channels.
set.seed() Is Not Optional
A draw is random: run it twice, get two samples
set.seed() fixes the generator’s starting point
Without it your sample is unreproducible
Nobody can check your results, including you
The number is arbitrary, and 409 means nothing
What matters is that one is set and recorded, so the draw is frozen.
Where Power Meets the Data
Sampling is not only about capturing variation
Your prospectus specified a minimum n
This is the week that number meets your dataset
Need 128 and the data gives 10,000: stratification solves it
Need 128 and the data gives 80: the study is underpowered
Then you change the design, or justify the limitation explicitly.
Lakens on What Honesty Demands
“Do not try to spin a story where it looks like a study was highly informative when it was not. Instead, transparently evaluate how informative the study was given effect sizes that were of interest, and make sure that the conclusions follow from the data.”
Lakens (2022, p. 11)
Purposive sampling improves construct validity and does nothing for power. Conflating them gives a study that is coherent and indefensible.
The Pilot Test
A trial run before the real coding begins
A hundred messages is a reasonable size
Two people, the same codebook, then a comparison
The question it answers
Applied independently, do these rules give the same codes?
Reliability is the property that the codebook answers the same whoever holds it.
Independence Carries the Logic
No conferring, no comparing, no checking as they go
Not suspicion: the pilot measures the codebook
Discuss a hard message and you measure their memory
Not the clarity of the written rules
The rules must carry a classification alone
Build the isolation into your protocol, and say so in the write-up.
Why Raw Agreement Is Not Enough
Two coders match on 74 of 100, which sounds respectable
Cohen’s 1960 example: two clinicians assigning blindly
At the observed rates, they still agree often
That reflects the number of categories, not judgment
Few categories, or one dominant category, and coincidences pile up fast.
Reliability First
Returning to Hayes and Krippendorff
“Conclusions from such data can be trusted only after demonstrating their reliability.”
Hayes and Krippendorff (2007, p. 77)
You have a number now. Does it clear the bar?
What would you report if it does not?
Whose trust is the standard protecting?
Informative, Not Impressive
“Do not try to spin a story where it looks like a study was highly informative when it was not.”
Lakens (2022, p. 11)
Say what your pilot could not settle
What does that cost the White Paper?
Cohen’s Kappa
Built for two coders and a nominal codebook (Cohen, 1960)
Observed agreement, less chance, over the room left
kappa = (observed - expected) / (1 - expected)
No better than chance: kappa is zero, whatever the raw rate
Perfect agreement: kappa is one
The codebook is credited only for what luck did not do.
Krippendorff’s Alpha
The more general instrument (Krippendorff, 2018)
Any number of coders, any level of measurement
Datasets with missing codes
Three coders, or an ordinal variable: alpha
For a two-coder nominal pilot they agree
v2v::reliability() computes either, selected by method.
Computing Alpha for Your Pilot
pilot <- coding_sample %>%slice_head(n =100)v2v::reliability( pilot$coder_a, pilot$coder_b,method ="alpha")
Take the first hundred messages, hand the two coders’ columns to reliability(), and ask for alpha rather than kappa.
A Worked Failure
The codebook said “high-energy” and never defined it
Coder A flagged messages carrying a Twitch emote
Coder B flagged messages containing all capitals
Both readings are reasonable, and both believed they coded one variable