Inferencing the Data [R]

Week 13 · Chapter 13 · MC 501 Research Methods for Mass Communications

Dr. Alex Leith

Today’s Agenda

  • The t-test that settles the central question
  • What a p-value does and does not say
  • ANOVA, regression, and the assumptions underneath
  • The White Paper and its poster, both assigned

What We Build Today

Working session

  • The t-test that settles the study’s question
    • Then an ANOVA, and a regression containing both
    • Then the checks that decide whether any of it holds
  • Two deliverables open this week, the paper and the poster
  • You leave with the sentence for your Results

The doing is one function call. The thinking around it takes the evening.

The Gap and the Doubt

  • Gaming averaged 28.49 characters, non-gaming 33.70
    • A difference of a little over five characters
  • The doubt is not idle
    • Split any group in two at random and the averages differ
  • The gap could be real, or ordinary wobble

The histogram cannot tell them apart, because both produce a gap.

What a Test Actually Asks

  • Nobody cares about those 34,766 messages as such
    • They were measured to learn about chat in general
  • The honest question is not whether a gap exists
    • It is whether chance is an unconvincing explanation
  • The test starts from the null hypothesis
    • That the two groups share one true mean length

The Coin Analogy

  • Assume the coin is fair
  • Work out how often a fair coin lands this lopsided
  • Judge the assumption by how surprising the data are

That is the whole logic, applied to two groups instead of one coin.

The Variables Choose the Test

  • Message length is ratio, an exact count of characters
  • Gaming status is nominal, with two values
  • A continuous outcome across two groups is the t-test
  • Two categorical variables instead: chi-square
  • The Hub keeps a test-decision tree, so consult it first

A test on the wrong kind of data produces a meaningless number.

Running the Test

v2v::run_t_test(
  analysis,
  value = message_length,
  group = is_gaming
)

One call takes the data, the variable to compare, and the variable that splits it into groups.

The Output

Welch two-sample t-test
message_length by is_gaming

  Mean, gaming         28.49 characters
  Mean, non-gaming     33.70 characters
  Difference           -5.22 characters

  t                    -4.95
  df                   3768.7
  p-value              < .001
  Cohen's d            -0.13  (negligible)

Why Welch, and Not Student

  • The ordinary t-test assumes equal spread
    • Week 12 showed sd of 60.68 against 38.47
    • Welch drops the assumption, so it is the safer default
  • t of -4.95 measures the gap in units of its own noise
    • The sign is direction only
  • df is not round because Welch computes it from both groups

Reading the p-value

  • If the two truly shared a mean length
    • A gap this large occurs less than once in a thousand
  • The data are surprising under that assumption
  • Below .05 counts as significant, and this clears it easily

Chance alone is an unconvincing explanation. That is the entire content.

What the p-value Does Not Say

  • Not the probability the null is true
    • The null is true or false; this is a statement about data
  • Not the probability the result is a fluke
  • Not a measure of how large the difference is
  • Not a measure of how much it matters

On the Threshold Itself

“Scientific conclusions and business or policy decisions should not be based only on whether a p-value passes a specific threshold.”

Wasserstein & Lazar (2016, p. 131)

  • .049 and .051 describe almost identical evidence
  • Ten uncorrected tests expect one false positive
  • Declare the tests and the correction in advance

Justify, Not Redefine

Lakens et al. (2018)

“we propose that researchers should transparently report and justify all choices they make when designing a study, including the alpha level.”

Lakens et al. (2018, p. 168)

  • Which choice in your study is least justified?
  • They answer a threshold with a demand. Enough?
  • Eighty-eight authors signed this. Does that matter?

Retiring the Phrase

“We conclude that the term ‘statistically significant’ should no longer be used.”

Lakens et al. (2018)

  • Write your result without the phrase
  • What did you lose, and what did you gain?

How Big Is the Difference

  • The p-value hands the question to an effect size
    • Between two means that is Cohen’s d, here -0.13
    • It states the gap in units of natural variation
  • About 0.2 small, 0.5 medium, 0.8 large (Cohen, 1988)
  • 0.13 is below small, and the word is negligible

Why It Reached Significance

  • The raw gap is fifteen percent of the non-gaming mean
    • Which does not sound negligible
  • But d measures it against sds of 38 and 61
  • n = 34,766

A large enough sample detects a difference far too small to care about.

Seeing It

means_ci <- analysis %>%
  filter(!is.na(is_gaming)) %>%
  group_by(is_gaming) %>%
  summarise(mean = mean(message_length, na.rm = TRUE),
            se   = sd(message_length, na.rm = TRUE) / sqrt(n()),
            .groups = "drop") %>%
  mutate(lower = mean - 1.96 * se,
         upper = mean + 1.96 * se)

Compute each group’s mean and standard error, then turn the error into a ninety-five percent interval you can plot.

Writing the Result Up

  • The intervals nearly vanish, and they do not overlap
    • That is the visual signature of significance
  • Both points still sit close together, far above zero
    • That is the negligible effect size made visible

Gaming channels produced shorter chat messages on average than non-gaming channels (28.49 vs. 33.70 characters), Welch’s t(3768.7) = -4.95, p < .001, Cohen’s d = -0.13.

Then Say It Plainly

The two kinds of channel do differ, but so slightly that they are better thought of as alike than as different.

A reader who cannot follow the notation still has to be able to follow the claim.

Four Groups, Four Means

game category       n     mean length
Fortnite          4000       33.23
Hearthstone       3004       36.82
Just Chatting     2429       39.03
League of Legends 4000       22.63
  • A t-test compares two means, and there are four
  • Six pairwise tests each carry a false-positive chance, and those pile up

ANOVA

four_games <- analysis %>%
  filter(game %in% c("Fortnite", "Hearthstone",
                     "Just Chatting", "League of Legends"))

summary(aov(message_length ~ game, data = four_games))
              Df    Sum Sq  Mean Sq  F value  Pr(>F)
game           3    547281   182427    92.26  < 2e-16 ***
Residuals  13429  26552407     1977

Reading F, Then Eta Squared

  • F compares variation between means against within groups
    • A shared mean would put F near one
    • Here F(3, 13429) = 92.26, with a minuscule p
  • League of Legends at 22.63 runs visibly shorter than 39.03
  • Eta squared is 0.02: category explains two percent

The Assumptions Underneath

  • ANOVA and ordinary regression assume homoscedasticity
    • Week 12 showed these groups plainly do not share variance
    • Which is why the two-group test used Welch
  • oneway.test() is the safer choice over pooled aov()
  • Run both and compare

If they disagree, the assumption was doing real work. Report the robust one.

Which Pair Actually Differs

  • A significant ANOVA says some pair differs, not which
  • Naming pairs takes post-hoc comparisons, and TukeyHSD() is standard
    • It corrects for the multiple looks you are taking
    • Six ad-hoc t-tests would lack that correction
  • Report the adjusted p-values, not the raw ones

With eta squared at 0.02, expect clean separations that matter little.

Regression Re-Derives the t-test

summary(lm(message_length ~ is_gaming, data = analysis))
              Estimate  Std. Error  t value  Pr(>|t|)
(Intercept)     33.70       0.70      48.16   < .001
is_gamingTRUE   -5.22       0.74      -7.08   < .001

The intercept is the non-gaming mean, and the coefficient is exactly the gap the t-test found, cast as a baseline plus an adjustment.

One Framework, Three Faces

  • A t-test is a regression with one two-level predictor
  • An ANOVA is one with a many-level predictor
  • All three are the general linear model
  • The regression’s t differs from Welch’s because it pools the spreads
    • The coefficient itself is identical
  • R² is about 0.001, so gaming status explains almost nothing

Residuals, and Earning a Predictor

  • Credibility rests on the residuals, and plot(model) draws four
    • Residuals against fitted: linearity and equal spread
    • A Q-Q plot: are the residuals normal?
    • Leverage: is one point carrying the result?
  • Whether a predictor earns its place is model comparison
    • Adjusted R², AIC, or a nested F-test, never raw R²

The White Paper and Its Poster

Both assigned this week · 200 points

  • The White Paper: IMRaD, APA 7th, effect sizes and power
    • Plus a pre-registration disclosure in the Methods
    • Its Discussion situates the finding in your literature
  • The poster: 36 by 24 inches, landscape, title to implications
    • A QR code to your pre-registration or portfolio
    • Your ninety-second pitch, built from one or two figures

Before Week 15

  • Submit Inferencing Data [R], 100 points
    • The test, the checks, the effect sizes, the interpretation
  • No class Week 14, for Thanksgiving break
  • Read Chapter 14 with the graduate toggle on
    • Plus the assigned article: Wasserstein and Lazar (2016)
  • Bring an outline and a poster sketch to Week 15

Questions?