“I am interested in Twitch chat.” Too broad. What aspect? Which streams?
“How viewers participate in chat.” Better, still vast.
“Whether chat participation differs across kinds of streams.” Which kinds? Measured how?
The Narrowing Funnel (cont.)
“Whether chat message volume differs between gaming and non-gaming streams.”
“Draw a stratified sample from the 50-channel working corpus, code each message as directed at the streamer or broadcast to the room, and test whether the proportion of directed messages differs between gaming and non-gaming streams.”
Version 5 fits a semester. It answers something rather than everything.
Three Ways Narrowing Fails
Impossible comparison: viewers against non-viewers
The groups differ in many ways besides yours
Circular question: popular because people watch
It defines its terms in a loop
False binary: “is it the streamer or the game?”
Ask how much each contributes, and whether they interact
Choosing an Operationalization
Parasocial relationship (Horton & Wohl, 1956) is not directly observable.
More reliable, and a behavior may not reflect the bond
Language analysis: code what viewers write
Naturalistic, labor-intensive, blind to those who never post
There is no right answer. You defend the proxy you chose.
Three Research Goals
Exploratory: “what is going on here?”
Generates hypotheses, and its output is patterns
Descriptive: “what does the landscape look like?”
Output is frequencies and prevalence
Explanatory: “why does this happen?”
Hypothesis-driven, output is support or disconfirmation
Research on a topic tends to move through these in order. You need not do all three, and you must know which one you are attempting.
Question, Hypothesis, Null
Research question: interrogative and open
Use it when exploring, or when the literature conflicts
Hypothesis: declarative and predictive
Use it when theory or prior work points somewhere
The null sits behind every hypothesis
No relationship, no difference
It is what a test evaluates, and the sacred flaw again
Defending the Number
Statistical power: the chance of detecting a real effect
At 50% power it is a coin flip
The conventional minimum is 80%
Below it a significant result is less credible
Underpowered studies inflate positives and overestimate effects
In your prospectus the number is defended, not asserted
Lakens on the Convention
“the default recommendation to aim for 80% power lacks a solid justification.”
Lakens (2022, p. 6)
80% is a convention, not a derived optimum
Resource constraint is a last resort for Lakens
Which sets up tonight’s argument
Computing the Required n
Take an effect size from your Chapter 4 reading, set α = .05, and ask:
library(pwr)# Two-group t-test, medium effect (d = 0.5), 80% powerpwr.t.test(d =0.5, sig.level =0.05, power =0.80, type ="two.sample")
You should seen = 64 per group. That number, with the reasoning that produced it, is the sampling-plan justification your method section carries.
Justified, Not Conventional
Lakens (2022)
“the default recommendation to aim for 80% power lacks a solid justification.”
Lakens (2022, p. 6)
What would justify a target for your study?
He ranks resource constraints last. Fair?
One coder, one semester: where does that leave you?
Topic, Theory, and Data
Topic is bounded by population, time, or medium
“Livestreaming” is not a topic
“Chat participation in the 50-channel corpus” is
Theory must predict patterns your data could show
Data must be accessible, sufficient, and appropriate
Misaligned: a parasocial topic with political-economy theory. One explains industry structure, the other is individual expression.
Six Components
A descriptive title hinting at the key variables
One to three focused questions or hypotheses, not ten
A theoretical framework, two to three sentences, naming the theory
Six Components (cont.)
A gap statement, two to three sentences, citing a few key sources
A method overview, three to four sentences: data, coding, number of cases
An expected contribution, one to two sentences
Roughly 250 words. One page, with every sentence carrying weight.
Model Prospectus, Question and Frame
Research question. Does the proportion of chat messages directed at the streamer differ between streams in gaming categories and streams in non-gaming categories, among the 50 channels in the Twitch working corpus?
Theoretical framework. Uses and gratifications theory (Katz, Blumler, & Gurevitch, 1973) holds that audiences actively select media to satisfy specific needs, including social ones.
Model Prospectus, the Method
Sample: 1,500 messages, stratified
Balanced across gaming and non-gaming streams
Each message is the unit of analysis
Coding and reliability: three codes, one check
Directed, broadcast, or unclassifiable
Second coder on a 10 percent subsample, Krippendorff’s alpha
Contribution: does a survey finding leave a coded trace?
Prospectus as Pre-Registration
Pre-registration commits the plan before the data
Hypotheses, methods, and analysis plan, publicly
The Open Science Framework timestamps it
Without a fixed plan, the choices drift
Adjust until a result reaches significance (Simmons et al., 2011)
The source of researcher degrees of freedom
Your prospectus is an informal pre-registration
What a Coder Cannot See
Chapter 7
The vocabulary is invisible from outside
“LULW” reads as a typo rather than laughter
Thirty identical messages are copypasta, not noise
“@sodapoppin same” may be agreement or a stock reply
The scheme still sorts everything confidently
The buckets do not mean what you think
The alternative is thick description (Geertz, 1973)
Three Modes of Attention
Casual watching: exposure across the platform
No notes and no coding
You replace your assumptions with actual exposure
Focused watching: one dimension at a time
Chat behavior, viewer trajectories, host moves
Brief notes after each pass
Three Modes of Attention (cont.)
Analytical watching: pattern recognition across more streams
What recurs, and what clusters with what
What resists description entirely
Familiarize, then focus, then analyze. The sequence holds for any medium.
Field Notes, Four Kinds
Observational: what you saw
“Chat moved in bursts, a flood at every question”
Methodological: a measurement problem
“A coder reading the log cannot see the question”
Theoretical and comparative: framework, then difference
An observation connected to something you have read
“Gaming and non-gaming chat are not the same object”
One dated plain-text document, with a note every fifteen or twenty minutes.