Codebook Studio

S16 · Chapter 8 · MC 451 Research Methods in Mass Media

Dr. Alex Leith

Today’s Agenda

  • What a codebook is, and its five parts
  • Your unit of analysis, in one sentence
  • Three variables, defined twice each
  • Decision rules out of your edge-case log

What We Build Today

Studio session

  • A draft codebook, in a file, saved
    • Three variables, not notes about codebooks
    • It becomes the Codebook and Qual Memo
    • Later it becomes your White Paper method
  • Every number traces to a rule you write today
  • Open your field notes and edge-case log now

The Codebook Is the Instrument

  • The complete set of instructions for coding
    • Not paperwork assembled afterward
  • Think of it as source code
    • It specifies every operation and every conditional
    • Two coders running it on one input agree
  • The heart of a content analysis (Neuendorf, 2017)

Its clarity is the strongest predictor of whether coders agree.

Five Parts

  • Unit of analysis: what exactly one coded case is
  • Variables and categories: what you measure, and the values
  • Decision rules: what to do when a case will not sort
  • Examples: two or three prototypes per category
  • Special cases: recurring complications, thin at first

Two Rules for Every Variable

  • Exhaustive: every case can be coded
    • An “unclassifiable” catch-all guarantees it
  • Mutually exclusive: every case fits exactly one
    • Honest overlap destroys reliability
  • The fix: sharper definitions, or a precedence rule

More than one case in ten in the catch-all means the scheme is incomplete.

Step 1, Your Unit of Analysis

Do this first, right now, in one sentence.

  • “Code the chat” is not a unit of analysis
  • This is: one chat message, one row of chat_log
    • One sender, one message string, one timestamp
  • A vague unit makes every later count vague
  • Add the independence clause for each unit

Your Turn

  • We read a few aloud
  • Can the room point at exactly one thing in the data?
  • If your unit is a stream or an hour, how many cases?
  • Is that enough cases to compare two groups?

Step 2, Definitions Side by Side

For each variable, write both, in this order:

  • Conceptual: what it is meant to capture, in the abstract
  • Operational: exactly what a coder does to assign a value
  • Level: nominal, ordinal, interval, or ratio
  • Categories: each with a one-line description

Aim for three variables. Three is the floor for a workable study.

Model Codebook, Variable 1

Message target. Conceptual: the intended addressee of the message. Operational: after reading the message, the coder assigns one category. Categories: directed at streamer (at-mention of the channel name, or second-person address responding to the streamer); directed at another viewer (at-mention of a non-streamer account, or a reply to a specific prior message); broadcast to the room (no specific addressee); unclassifiable. Level: nominal.

Model Codebook, Variables 2 and 3

Message length. Conceptual: the verbal extent of the message. Operational: count of characters in the message string, including spaces and emote tokens. Derived, computed rather than judged. Level: ratio.

Contains emote. Conceptual: whether the message uses Twitch’s emote vocabulary. Operational: yes if it contains at least one token from the project’s emote list. Level: nominal.

Step 3, The Five Decision Rules

From the model codebook:

  • Precedence: explicit address wins over performance
  • Emote-only: broadcast, unless it carries an at-mention
  • Copypasta: code by content, like any other message
  • Non-English: code target if structure determines it
  • Bots: unclassifiable, and noted for possible exclusion

Write Your Own Rules Now

  • Open the edge-case log, and every entry becomes a rule
    • Format: the ambiguous case, then the single action
    • “Use your judgment” is not a rule
  • Start with the three you starred
    • The ones most likely to split two coders
  • Aim for four to six rules

Step 4, Examples and Special Cases

  • Two or three prototypical cases per category
    • Invented is fine, real messages are better
  • Examples do double duty
    • A reference for coding, and training for a second coder
  • Special cases stays thin today
    • Leave the heading in the file, empty, as a reminder

Deriving the Manifest Variables

Message length and emote presence are computed, not judged:

library(v2v)
library(dplyr)

chat <- twitch_chat()

chat %>%
  mutate(msg_length = nchar(message)) %>%
  summarise(mean_length = mean(msg_length),
            n_messages  = n())

Add a character count to every message, then report the average and the count. A derived variable needs an operational definition too, and this is it in executable form.

Common Problems Today

  • Categories that overlap
    • Hesitating between two means they are not exclusive yet
  • A latent variable with no observable criteria
    • You wrote a concept, not a recipe
  • Six variables: cut to three
    • Each is a reliability check and a results paragraph

Common Problems Today (cont.)

  • A catch-all doing all the work
    • Expecting a third unclassifiable means the scheme is wrong
  • Rules that restate the category
    • A rule resolves a conflict, it does not repeat a definition

Every one of these turns up today, which is why this is a studio and not a reading.

Checkpoint

You should now have a file containing:

  • One unambiguous unit of analysis
  • Three variables, each defined twice, with a level and categories
  • Categories that are exhaustive and mutually exclusive
  • Four to six decision rules, each traceable to an observation
  • Two or three examples per category, and an empty heading

Before Next Time

  • Finish and commit the codebook tonight
    • While today’s decisions are still in your head
  • Hand it to someone who was not here
    • No walk-through, and ask for ten coded messages
    • Note every hesitation, which is a missing rule
  • Codebook and Qual Memo is what this draft becomes

Questions?