Compute each group’s mean and standard error, then turn the error into a ninety-five percent interval you can plot.
Writing the Result Up
The intervals nearly vanish, and they do not overlap
That is the visual signature of significance
Both points still sit close together, far above zero
That is the negligible effect size made visible
Gaming channels produced shorter chat messages on average than non-gaming channels (28.49 vs. 33.70 characters), Welch’s t(3768.7) = -4.95, p < .001, Cohen’s d = -0.13.
Then Say It Plainly
The two kinds of channel do differ, but so slightly that they are better thought of as alike than as different.
A reader who cannot follow the notation still has to be able to follow the claim.
Four Groups, Four Means
game category n mean length
Fortnite 4000 33.23
Hearthstone 3004 36.82
Just Chatting 2429 39.03
League of Legends 4000 22.63
A t-test compares two means, and there are four
Six pairwise tests each carry a false-positive chance, and those pile up
ANOVA
four_games <- analysis %>%filter(game %in%c("Fortnite", "Hearthstone","Just Chatting", "League of Legends"))summary(aov(message_length ~ game, data = four_games))
Df Sum Sq Mean Sq F value Pr(>F)
game 3 547281 182427 92.26 < 2e-16 ***
Residuals 13429 26552407 1977
Reading F, Then Eta Squared
F compares variation between means against within groups
A shared mean would put F near one
Here F(3, 13429) = 92.26, with a minuscule p
League of Legends at 22.63 runs visibly shorter than 39.03
Eta squared is 0.02: category explains two percent
The Assumptions Underneath
ANOVA and ordinary regression assume homoscedasticity
Week 12 showed these groups plainly do not share variance
Which is why the two-group test used Welch
oneway.test() is the safer choice over pooled aov()
Run both and compare
If they disagree, the assumption was doing real work. Report the robust one.
Which Pair Actually Differs
A significant ANOVA says some pair differs, not which
Naming pairs takes post-hoc comparisons, and TukeyHSD() is standard
It corrects for the multiple looks you are taking
Six ad-hoc t-tests would lack that correction
Report the adjusted p-values, not the raw ones
With eta squared at 0.02, expect clean separations that matter little.
Regression Re-Derives the t-test
summary(lm(message_length ~ is_gaming, data = analysis))
Estimate Std. Error t value Pr(>|t|)
(Intercept) 33.70 0.70 48.16 < .001
is_gamingTRUE -5.22 0.74 -7.08 < .001
The intercept is the non-gaming mean, and the coefficient is exactly the gap the t-test found, cast as a baseline plus an adjustment.
One Framework, Three Faces
A t-test is a regression with one two-level predictor
An ANOVA is one with a many-level predictor
All three are the general linear model
The regression’s t differs from Welch’s because it pools the spreads
The coefficient itself is identical
R² is about 0.001, so gaming status explains almost nothing
Residuals, and Earning a Predictor
Credibility rests on the residuals, and plot(model) draws four
Residuals against fitted: linearity and equal spread
A Q-Q plot: are the residuals normal?
Leverage: is one point carrying the result?
Whether a predictor earns its place is model comparison
Adjusted R², AIC, or a nested F-test, never raw R²
The White Paper and Its Poster
Both assigned this week · 200 points
The White Paper: IMRaD, APA 7th, effect sizes and power
Plus a pre-registration disclosure in the Methods
Its Discussion situates the finding in your literature
The poster: 36 by 24 inches, landscape, title to implications
A QR code to your pre-registration or portfolio
Your ninety-second pitch, built from one or two figures
Before Week 15
Submit Inferencing Data [R], 100 points
The test, the checks, the effect sizes, the interpretation
No class Week 14, for Thanksgiving break
Read Chapter 14 with the graduate toggle on
Plus the assigned article: Wasserstein and Lazar (2016)