Annotating the Literature and Theory as a Lens

Week 4 · Chapters 4 and 5 · MC 501 Research Methods for Mass Communications

Dr. Alex Leith

Tonight

Tonight is about what you are looking for, before you go looking for it.

First half: what makes an article worth keeping, and what to pull out of one when you read it, including a number you will need long before you expect to. Second half: theory, which decides what you look at and what you ignore.

You do not need a stack of articles yet. You need to know what you are reading for, because three annotated manuscripts are due the Monday after next.

This is a heavy evening

  • Two chapters, a statistics primer, and a catalogue of theories
  • The middle section involves effect sizes, and it will feel like a math class for about fifteen minutes
  • You are not being asked to derive anything. You are being asked to find a number in somebody else’s paper and write it down.
  • If the theory half feels abstract, hold on until the last two slides, where we read four real chat messages through two lenses

Decide this before the pile exists

Reading for annotation is a different act from reading for interest. You are extracting a fixed set of things from every manuscript, and extracting the same things every time is what makes the sources comparable to each other.

Which means the list comes first. Downloading twenty PDFs and deciding later what mattered gives you twenty documents you have to open twice.

What a weak review sounds like

Hamilton, Garretson, and Kerne (2014) studied Twitch as a participatory community. Sjöblom and Hamari (2017) examined viewer motivations. Hilvert-Bruce and colleagues (2018) studied social motivations for engagement.

  • Three disconnected summaries, and nothing about how they relate
  • Every sentence is driven by who wrote the study instead of by what is known
  • This is what annotation without synthesis produces
  • Most first drafts look exactly like it. It is a stage, not a verdict.

Where the citation goes

  • Default: make the claim the subject and put the citation at the end, in parentheses, as (Sjöblom & Hamari, 2017)
  • Lead with the name only when that study is the point: a landmark, a definition you are adopting, or a finding you are about to argue with
  • An author-led sentence announces a study. A claim-led sentence states a finding.
  • A page of “X (2014) found, Y (2017) found” is an annotated bibliography wearing the clothes of a review

What a strong review does

  • Groups sources by theme rather than by author or date
  • Names the convergence: what these studies together have established
  • Adds nuance: where they disagree, or where one qualifies another
  • Names a shared limitation, which is usually where your gap lives
  • Worked: survey research converges on livestream viewing being socially motivated, but all of it is self-report, so nobody has checked the behavioral trace

The same thing, written out

Livestream viewing is motivated less by what is on screen than by the social experience around it. Twitch streams function as virtual third places where informal communities form (Hamilton et al., 2014), and social and tension-release motivations are central to why viewers watch (Sjöblom & Hamari, 2017). Social interaction predicts not just watching but chatting and subscribing (Hilvert-Bruce et al., 2018). All three rely on self-report, leaving open whether the same patterns appear in behavioral chat-log data.

Why that version works

  • The claim leads, and the citations arrive behind it as support
  • Sources are grouped by what they argue, not by the year they appeared
  • The limitation is shared across the group, which makes it structural rather than a complaint about one paper
  • The gap is never asserted. It falls out of the limitation on its own.
  • Notice where it points: at a content analysis of a real chat log

What to mark in a manuscript

  • The research question, in the authors’ own words, and the theory behind it
  • The design: unit of analysis, sample, how the key construct was measured
  • The finding in one sentence, and the gap the authors claim to fill
  • The effect size and the sample size used to detect it
  • The limitation the authors admit, and the one they do not
  • One sentence you would be willing to quote in your own methods section

The number you are really there for

An effect size is a number saying how big a difference or relationship actually is, as opposed to whether it cleared a significance threshold.

You need one because before you can justify how much data to collect, you need a defensible estimate of the effect you expect to find. The most defensible source for that estimate is research that already measured it.

So you copy it out of their paper. That is the whole task.

What to write down

  • Cohen’s d, Cramer’s V, or a correlation r, whichever the authors report
  • The n they used, because an effect size without a sample size is not much use
  • Whether they ran a power analysis themselves, and what they assumed if so
  • Those numbers become the inputs to your own power analysis later in the term
  • A review that cannot supply an effect size estimate has not finished its job

Cohen on the hard part

“Researchers find specifying the ES the most difficult part of power analysis.”

Cohen (1992, p. 156)

ES is his abbreviation for effect size, and notice what he is conceding: the hardest part is not the arithmetic.

The literature review is how that difficulty becomes tractable. You are not inventing an estimate. You are borrowing one from people who already did the measuring.

What that looks like when you read

  • Start with one study close to the kind of thing you want to do
  • Find its primary effect size, and note whether the authors ran a power analysis themselves
  • Later in the term you hand those numbers to the pwr package, which turns them into the sample size you need
  • The harder case is coming: nobody has reported an effect for your question
  • We will build two strategies for justifying a sample size anyway

Discussion

On Cohen (1992), “A power primer”:

  • Cohen offers conventions for small, medium, and large effects. Conventions make a field legible. What do they cost it?
  • He says specifying the effect size is the hardest part. Why does that step, and not the arithmetic, stall people?
  • A paper this short has shaped decades of practice. What does its brevity buy, and what does it leave to the reader to get wrong?
  • Where would you expect to find an effect size for a question like yours, and what would you do if nobody has reported one?

Two minutes to think, then we take it as a group.

Four kinds of gap

  • Topical void: a phenomenon, population, or context nobody has studied
  • Methodological gap: studied, but with methods carrying an important limitation
  • Contradiction: two well-designed studies reaching opposite conclusions
  • Theoretical gap: explained through one lens where another would show more
  • One caution on the void: what looks novel is usually covered under other terminology

Articulating the gap

  1. Establish what is known. Show you understand the conversation.
  2. Identify the limitation. What is missing, contradictory, or unexplained.
  3. Argue for your study. How your design addresses exactly that.

Known, then limitation, then what this study does. The topic changes. The hinge does not.

The literature map

  • A markdown file in your project, organized by theme rather than alphabetically
  • One heading per theme, sources beneath it, plus a running section for gaps
  • It shows at a glance which themes are well covered and which are thin
  • It makes you decide how sources relate, which is where synthesis starts
  • Its structure usually becomes the structure of the written review
  • Start it this week with the first article you read. Three headings and one source under one of them is already a map.

If that feels out of reach

Start here. This is a floor, not a compromise:

  • Source A, one sentence: what did this study find?
  • Source B, one sentence: what did this study find?
  • Source C, one sentence: what did this study find?
  • The gap, one sentence: what do these three together leave unexplored?

Four sentences, three citations, one gap. Enough to anchor a prospectus, and the full synthesis paragraph is what it grows into rather than where you start.

Citation is ethical, not clerical

  • Cite specific empirical findings, theoretical concepts, and methodological approaches
  • Direct quotes always. Paraphrased ideas unless genuinely common knowledge.
  • No citation needed for “Twitch is a livestreaming platform”, for your own interpretations, or for your own data and analysis
  • Nobody is criticised for citing too much. Under-citation damages credibility.
  • In its serious forms it is plagiarism, which is a breach of trust before it is a rule

One finding, three explanations

Chapter 4 ended on a finding: livestream viewing looks more socially motivated than content motivated. Three research programs would explain that differently.

  • Uses and gratifications (Katz, Blumler, & Gurevitch, 1973): viewers are active agents meeting a social need a finished video cannot meet
  • Parasocial interaction (Horton & Wohl, 1956): viewers form one-sided bonds, and the social feeling is that bond being felt
  • Social identity (Tajfel & Turner, 1979): the community is a group, and chat is a space of belonging

Theory is not speculation

A theory is a formal, systematic explanation of relationships between concepts. Good ones are logically coherent, generalizable, falsifiable, and generative.

Social identity theory predicts that viewers rate their own community above rivals on subjective dimensions like authenticity, and that they respond defensively when the streamer or community is criticized. Change the domain and it predicts partisans rating aligned outlets as more credible.

The theory does not change. The domain does. That is what makes it testable.

The lens metaphor

  • A theory brings some things into focus and pushes others out of the frame
  • Appraisal theory points at individual cognitive evaluation, so you build surveys and experiments measuring individual reaction
  • Social identity theory points at group meaning, so you analyze how members talk about the stream and how boundaries get policed
  • Same phenomenon, different designs, neither one wrong
  • You cannot use every lens at once. That produces blur rather than richness.

Three paradigms

  • Social scientific: an objective reality, observable and measurable. Deductive. Asks “does X cause Y?”
  • Interpretive: reality socially constructed through interaction. Inductive. Asks “what does X mean to the people who experience it?”
  • Critical and cultural: power shapes what counts as knowledge. Asks “whose interests does X serve, and whose voices are silenced?”
  • Each treats theory differently: as a source of predictions, as an outcome of observation, or as a tool for exposing structure

Grand and middle-range

  • Grand theories explain human behavior across all contexts: Marxism, psychoanalysis, structuralism
  • Intellectually powerful, and too broad to design a study that could disprove them
  • Middle-range theories are specific enough to generate testable hypotheses and broad enough to travel beyond one case
  • Merton named the zone, and most communication research lives in it
  • Every theory in the catalogue ahead is middle-range

The catalogue, part one

  • Parasocial interaction (Horton & Wohl, 1956): one-sided bonds that feel intimate without reciprocity. Explains donating to a streamer you have never met.
  • Social identity (Tajfel & Turner, 1979): group membership shapes self-concept, producing in-group favoritism and boundary policing through emotes and slang
  • Agenda setting (McCombs & Shaw, 1972): media do not tell people what to think, but what to think about. The front page and the category rankings set an agenda.

The catalogue, part two

  • Framing (Goffman, 1974; Entman, 1993): the same facts, differently emphasized, produce different conclusions. Visible in how streamers title their own broadcasts.
  • Uses and gratifications (Katz, Blumler, & Gurevitch, 1973): audiences actively select media to meet needs. The lens behind most viewer-motivation research.
  • Cultivation (Gerbner & Gross, 1976): cumulative exposure shapes perceptions of social reality, most strongly where audiences lack direct experience

Two lenses, four messages

DragoBoi89: xQc actually playing well today wtf

emotelord: PogU PogU

realnamedfan: @xqcow do u remember when u said hi to me last week

streamfan22: chat is this real

Four messages from one stretch of a real broadcast. Read them once before we code them.

What each lens foregrounds

  • Uses and gratifications asks what need each viewer is meeting: evaluative commentary, collective emotional expression, direct acknowledgment, group sense-making
  • All four viewers are equally interesting, because each one shows a different gratification
  • Parasocial interaction asks what the bond looks like: realnamedfan’s @-address plus a memory claim is the textbook case
  • emotelord and streamfan22 never address the streamer, so this lens pushes them aside
  • Same data, two stories. Your prospectus commits to one of them.

Choosing what to read

  • Start from the keystone articles: the ones your other sources keep citing
  • Prefer studies close to your design, not only your topic. A content analysis of chat is worth more to you than a survey about streaming.
  • Take one that disagrees with the others if you can find it, because contradiction is a gap
  • Fit beats recency. A 2016 study measuring what you want to measure beats a 2024 one that does not.
  • Three is the floor, not the target

Before Week 5

  • Read three articles on your topic this week, with tonight’s list beside you. They become the Annotated Manuscripts, due Monday, September 28, with CITI.
  • Monday, September 21: the Librarian Visit Report. Tonight is the last class before it is due, and you arrange the meeting yourself: ours is in MC 500, or book a subject librarian at Lovejoy.
  • Read Chapter 5 (variables and hypotheses) and Chapter 6 (research questions), with the edition toggle on
  • Read the assigned article: Gelman and Loken (2014), “The statistical crisis in science”
  • Write your journal entry, 450 to 500 words, engaging both the chapters and the reading
  • Start CITI early. It takes longer than anyone expects.