Inferential: t-test, ANOVA, Regression

S25 · Chapter 13 · MC 451 Research Methods in Mass Media

Dr. Alex Leith

Today’s Agenda

  • What a hypothesis test actually asks
  • Choosing the test from your variables
  • Reading a p-value, and refusing to overread it
  • Effect size, and why big samples split the two

Where We Left Off

Inferencing Data, 100 points, due this week

  • Gaming channels averaged 28.49 characters
  • Non-gaming averaged 33.70, a gap of five characters
  • The histogram showed the gap, and could not say it was real

Today we settle it. No formulas to memorize, only questions to ask.

What a Test Actually Asks

Split a group of people in half at random and their average heights still differ a little. Samples wobble.

  • We measured 34,766 messages to learn about chat in general
    • The sample is the means, and the population is the point
  • The honest question is not whether there is a gap
    • It is whether chance is a convincing explanation for it

The Skeptic’s Move

  • A test starts from the opposite of what you suspect
    • The null hypothesis: no real difference, pure noise
    • Then: if that were true, how surprising is this gap?
  • Testing a coin works the same way
    • Assume it fair, flip a hundred times, ask how often it looks this lopsided
  • “Almost never” makes the fair-coin assumption hard to keep

The Variables Choose the Test

  • Continuous outcome, two groups: the t-test
  • Continuous outcome, three or more: ANOVA
  • Two categorical variables: chi-square, v2v::run_chi_square()
  • Ours: characters counted, gaming status has two values, so t-test

A test applied to the wrong kind of data produces a meaningless number.

Your Turn

  • State the null hypothesis for your question, in one sentence
  • What would make your finding nothing but sampling wobble?
  • What size of difference would change how you think about your topic?

Running the t-test

v2v::run_t_test(analysis, value = message_length, group = is_gaming)

One call: the data, the variable to compare, and the variable that splits it.

The output names a Welch t-test. The ordinary version assumes equal spread, and ours plainly differ, at sd 60.68 against 38.47. Welch drops that assumption.

Reading the Output

  Mean, gaming         28.49 characters
  Mean, non-gaming     33.70 characters
  Difference           -5.22 characters

  t                    -4.95
  df                   3768.7
  p-value              < .001
  Cohen's d            -0.13  (negligible)

t measures the gap in units of its own uncertainty. The sign is direction only.

What the p-value Means

  • If the two truly had the same mean length
    • A gap this large happens less than once in a thousand samples
    • The data are surprising if nothing is going on
  • Below .05 counts as statistically significant, by convention
    • A convention, not a law: .049 and .051 are the same evidence

What the p-value Does Not Mean

  • Not the probability the null hypothesis is true
  • Not the probability your result is a fluke
  • Not a measure of how large the difference is
  • Not a measure of whether the difference matters

A small p says the gap is probably not zero, and nothing about its size.

How Big Is the Difference

  • For size you need an effect size
    • Between two means, that is Cohen’s d
    • It states the gap in units of natural variation
  • Cohen’s benchmarks: 0.2 small, 0.5 medium, 0.8 large
  • Ours is -0.13, below small, in the range called negligible

Real and Trivial at Once

  • Five characters is fifteen percent of the non-gaming mean
    • Against standard deviations of 38 and 61, it is a ripple
  • So why was it significant? Sample size: 34,766
    • A large enough sample makes almost anything significant

Significance and size are different questions, and big samples pull them apart.

More Than Two Groups, ANOVA

four_games <- analysis %>%
  filter(game %in% c("Fortnite", "Hearthstone",
                     "Just Chatting", "League of Legends"))

summary(aov(message_length ~ game, data = four_games))

Four means, so a t-test will not do. Comparing them two at a time takes six tests, each with its own chance of a false positive.

Reading the F Table

              Df    Sum Sq  Mean Sq  F value  Pr(>F)
game           3    547281   182427    92.26  < 2e-16
Residuals  13429  26552407     1977

F compares variation between the group means against variation within them. If all four shared a mean, F would sit near one. Here it is 92.26.

The effect size is eta-squared, here 0.02: category explains about two percent of the variation. Real, and small.

From Difference to Model

summary(lm(message_length ~ is_gaming, data = analysis))
              Estimate  Std. Error  t value  Pr(>|t|)
(Intercept)     33.70       0.70      48.16   < .001
is_gamingTRUE   -5.22       0.74      -7.08   < .001

The intercept is the non-gaming mean. The coefficient, -5.22, is exactly the gap the t-test found, recast as a baseline plus an adjustment.

Checkpoint

You should now be able to:

  • Name the test your question calls for, and say why
  • Run v2v::run_t_test() and read the means, difference, p, and d
  • Say what a p-value does and does not claim
  • Report a significant result that is negligible, without inflating it
  • Know that aov() handles three or more groups, and lm() generalizes both

Before Next Time

  • Due this week: Inferencing Data [R], 100 points
  • Thursday is a lab, running your own test with me here
    • Bring your analysis table and last week’s summary
  • Re-read Chapter 13 on significance against importance
  • Thursday I assign the White Paper, 250 points

Questions?