Week 5 · Chapters 5 and 6 · MC 501 Research Methods for Mass Communications
Dr. Alex Leith
Tonight
A theory tells you what to look at. Variables are how you look, and the question is what you are looking for.
First half: variables, hypotheses, and what it costs to decide what you were testing after you have already seen the data. Second half: writing the question, which is the hardest single thing in this course.
Everything here feeds the prospectus you draft next week.
Fair Warning, Second Half
Nobody writes a good research question on the first try. Nobody
Narrowing feels like giving up on the interesting version of your idea, and it is not
The version you arrive with tonight will almost certainly be too big, which is the normal starting point rather than a sign you misunderstood
We narrow them in the room, out loud, together. Bring the too-big one
Variables: The Building Blocks
Theories become testable through variables, concepts that take values
Independent variable: the proposed cause, or the predictor you measure
In this corpus: stream category, streamer tenure, viewer count.
Dependent variable: the outcome you think gets influenced
In this corpus: messages per minute, retention, subscription rate.
A Testable Hypothesis
In non-gaming streams, a higher proportion of chat messages are directed at the streamer than in gaming streams.
IV: category type, gaming or non-gaming
DV: whether a message addresses the streamer or the room
The whole procedure is visible in that sentence
Classify streams, code a sample, test the predicted relationship.
Swap the domain and the structure survives.
Mediators And Moderators
A mediator explains how or why, the mechanism in between
Category may affect directed messages through perceived pace.
Non-gaming streams feel conversational, which invites direct address.
A moderator changes strength or direction, under what conditions
The effect might hold on small streams and weaken on large ones.
Together they move you from “X relates to Y” to “X relates to Y, by this mechanism, under these conditions”.
HARKing, Precisely
Hypothesizing After Results are Known. You look at data, notice a pattern, then write as though you predicted it.
The test still reports a p-value. But the hypothesis came out of the data, so that p-value no longer carries the meaning it claims.
Nothing here requires dishonesty. It requires only ordinary forgetting, which is why the fix has to be structural rather than a matter of good character.
The Inferential Cost
“Failing to appreciate the difference can lead to overconfidence in post hoc explanations (postdictions) and inflate the likelihood of believing that there is evidence for a finding when there is not.”
Nosek et al. (2018, p. 2600)
A postdiction is an explanation produced after the fact. It can describe the same pattern a prediction would. The difference is that only the prediction was risking anything.
The Structural Fix
Pre-registration submits hypotheses, variables, and analysis plan publicly
Filed before collection, time-stamped, and permanent.
Drift becomes visible rather than deniable.
A V2V entry carries question, design, and criteria
Directional hypotheses; unit of analysis, platform, collection window.
Codebook or extracted variables; planned tests; sample size justification.
OSF’s standard template walks you through each section
The Garden Of Forking Paths
Gelman and Loken, 2014
The point is subtler than p-hacking
No conscious fishing is required for the problem to appear.
Many defensible choices, each contingent on the data
You would have chosen differently had the data come out differently.
Those unrealised paths still inflate the false-positive rate.
The hypothesis can be stated in advance and the problem survives
Forking Paths
“In this garden of forking paths, whatever route you take seems predetermined, but that’s because the choices are done implicitly.”
Gelman and Loken (2014, p. 464)
A study compares directed messages in gaming and non-gaming streams
Name one choice that comparison leaves implicit
What result would have made the author choose otherwise?
Multiple Comparisons
“it is possible to have multiple potential comparisons … without the researcher performing any conscious procedure of fishing through the data”
Gelman and Loken (2014, p. 460)
Intent is doing no work here
Where does this bite a content analysis?
Name one place a coding scheme could fork unnoticed
Preregistration’s Limits
“Preregistration may be practical in some fields and for some types of problems, but it cannot realistically be a general solution.”
Gelman and Loken (2014, p. 464)
This course teaches preregistration as a pillar
Where does their objection land hardest?
What survives it?
The Hardest Part
Running a test can be learned with practice. Asking well cannot be outsourced to anyone.
Question A: “How does livestreaming affect people?”
Question B: “Does the rate of chat messages per viewer differ between gaming and non-gaming streams?”
A is a career disguised as a question. B you can design, execute, and finish inside one semester.
Five Criteria For Questions
Specific and measurable: names platform, outcome, and population
Its concepts must survive operationalization.
The move from an abstract idea to something observable.
Answerable and open: finishable, and not already settled
Not decades of data or hundreds of interviews.
Your review tells you whether it is genuinely open.
It matters: contributes to theory, resolves a contradiction, or applies
The Question Template
Some combination of does there exist, plus a specific variable, plus a relationship, pattern, or difference, plus a bounded population, plus any confounds accounted for.
Is there a relationship between stream category and chat message rate among channels in the Twitch working corpus?
Every part of that sentence is doing work. Delete one and the study loses its shape.
The Narrowing Funnel
“I am interested in Twitch chat.” Too broad. Which aspect? Which streams?
“How viewers participate in chat.” Better, still vast.
“Whether participation differs across kinds of streams.” Which kinds? Measured how?
“Whether chat message volume differs between gaming and non-gaming streams.” Clearer.
A stratified sample from the 50-channel corpus, each message coded as directed or broadcast, tested across stream type, framed as social gratification.
Each step trades breadth for depth. Narrowing is discipline rather than weakness.
Narrowing Mistakes With Names
The Everything Study: “the effect of social media on society”
A research program, not a project.
The impossible comparison: Twitch viewers against non-viewers
Self-selection confounds it before you collect anything.
Circular and binary questions: loops, and forced either/or
“Popular because people watch them” defines itself in a loop.
“Streamer or game?” should ask how much each contributes.
Operationalizing A Parasocial Bond
The construct is not directly observable, so you choose among imperfect proxies and defend the one you chose.
Self-report scale: “I feel like this streamer understands me”
More reliable, and a behavior may not mean what you assume.
Language analysis: code chat for intimacy, disclosure, direct address
Naturalistic, labor-intensive, blind to viewers who say nothing.
Operationalizing Chat Sentiment
Automated scoring: fast and replicable
Blind to sarcasm, irony, and the coded vocabulary of emotes.
Human coding: catches context and nuance
Labor-intensive, and requires intercoder reliability testing.
Dimensional coding: valence and arousal scored separately
Better on emotional complexity, demands a more careful codebook.
There is no right answer here. You pick the proxy that fits the question and the resources you have, then defend its validity in writing.
Variables Hide In Descriptions
aura-lab.siue.edu/theories
Open Parasocial Interaction and read the TL;DR aloud. We pull the variables out of prose that never names them.
What the persona does: direct address, camera look, informal talk
Horton and Wohl called it an illusion of face-to-face conversation.
It varies by figure, by segment, and by platform.
What the audience does: loyalty, reliance, a sense of knowing
A scale exists for exactly this (Rubin, Perse & Powell, 1985).
The first list is candidate independent variables and the second is candidate dependent variables. The theory supplied both and labeled neither.
Naming The IV And DV
IV, persona behavior: how often the figure addresses viewers directly
Counted per minute, or coded present and absent per segment.
DV, audience response: intimacy and direct address in audience text
Second-person pronouns, first names, disclosure, repeat visits.
The split that matters: interaction during viewing, relationship between viewings
The two have been separated (Dibble, Hartmann & Rosaen, 2016).
Which one you measure decides what your DV can be.
The Pair Across Media
Livestream chat: persona talk and stream category as IV
DV is directed messages as a share of all messages.
Influencer comments: posting cadence and reply behavior as IV
DV is intimacy markers across the comment thread.
Social VR capture: proximity and apparent eye contact as IV
DV is turn-taking and disclosure in recorded speech.
Broadcast transcripts work too, and they cap what you can ask: the audience never appears in the data, so the DV has to come from somewhere else.
The pair travels. Pick the medium your question actually lives in.
Three Kinds Of Goal
Exploratory: “what is going on here?”
Generates hypotheses. Output is patterns and frameworks.
Descriptive: “what does the landscape look like?”
Output is frequencies, distributions, prevalence.
Explanatory: “why does this happen?”
Hypothesis-driven. Output is support or disconfirmation.
Research usually moves through them in that order. You do not need all three, and you do need to know which one you are attempting.
Question Or Hypothesis
A research question is interrogative and open
Use it when exploring, describing, or when the literature conflicts.
A hypothesis is declarative and predictive
Use it when theory predicts, or when prior work suggests an outcome.
Behind every hypothesis sits the null hypothesis
No relationship, no difference. The skeptical default.
It is what your test actually evaluates.
The question is never whether you can tell a story. It is whether this data could have come from a world with no effect in it.
Topic, Theory, Data
All three have to line up. Mismatch any one of them and the study collapses quietly.
Topic must be bounded, never a whole medium
“Chat participation in the 50-channel corpus”, not “livestreaming”.
Theory must predict patterns your data could show
Data must be accessible, sufficient, and appropriate
Political economy plus chat messages fails: the theory explains industry structure and the data captures individual expression. Swap to parasocial interaction and the three begin talking to each other.
Your Theory, Your Variables
aura-lab.siue.edu/theories
Open the theory you are building on. No theory settled yet? Take the closest one and treat tonight as a test of whether it holds.
The corpus we keep returning to is a shared example, not your data. Yours is whatever your question needs: a comment section, a subreddit, a news archive, a set of TikToks.
Two candidate IVs your theory says should matter
Name where each one lives in the content itself.
Two candidate DVs the same theory predicts
Say what a coder would look at to record it.
Then cut what you cannot code
If it needs a survey, it is out tonight.
What Content Analysis Reaches
Reachable: what is said, shown, posted, or counted
Word choice, frequency, address, timing, visible behavior.
Out of reach alone: what a person felt or intended
Loneliness, satisfaction, and attitude need a survey or interview.
The usual fix: move the construct to its visible trace
Not felt closeness, but expressions of closeness in the thread.
This is the constraint your prospectus inherits. A question you cannot code is a question you cannot answer with this method.
What The Prospectus Commits
Title and questions: descriptive, hinting at the key variables
One to three focused questions or hypotheses, not ten.
Framework and gap: two or three sentences each
Name the theory. Cite a few key sources on the gap.
Method and contribution: three or four sentences, then one or two
Data, coding, number of cases. Then what it adds.
Defending The Number
Your method section names a sample size, and next week it stops being an assertion
Statistical power: the chance of detecting a real effect
A study at 50% power is a coin flip even when the effect is there.
The conventional minimum is 80%, and it is a convention
Below it a significant result is less credible.
Underpowered studies inflate false positives and overestimate effects.
Lakens On The Convention
“the default recommendation to aim for 80% power lacks a solid justification.”
Lakens (2022, p. 6)
He treats justification by resource constraint as a last resort rather than a defense.
Come next week with a position on this: is it defensible to run a design you already know is underpowered?
Computing The Number
library(pwr)# Two-group t-test, medium effect (d = 0.5), 80% powerpwr.t.test(d =0.5, sig.level =0.05, power =0.80, type ="two.sample")
You should seen = 64 per group. That is what a sampling plan justification looks like when it is finished.
If you cannot collect that many, you revise the question, redesign for greater power, or state the limitation explicitly. Those are the only three honest options.
Before Week 6
Due tonight:Annotated Manuscripts (3) and CITI Certification
Read: Chapter 6 on the prospectus, Chapter 7 on Structured Listening
Keep the edition toggle on.
The assigned article: Lakens (2022), “Sample size justification”.
Write and bring: a journal entry of 450 to 500 words
Engage both the chapters and the reading.
Bring a draft research question. We narrow them in the room.