Bucket into six-hour windows, lump everything past the top five into “Other”, and sum viewers inside each bucket.
Drawing the Line
ggplot(viewers_over_time,aes(x = six_hour, y = total_viewers, color = category)) +geom_line(linewidth =0.9) +labs(title ="Concurrent viewers by game category",x ="Date (UTC, six-hour buckets)",y ="Total viewers in bucket",color ="Category") + v2v::scale_colour_v2v() + v2v::theme_v2v()
theme_v2v() applies publication-ready fonts, spacing and gridlines in one call.
What Dominates, and Why
The towering line is “Other”, the catch-all beyond the top five
Substantively: viewership really is spread across a long tail
As a caution: lumping fifty categories guarantees that result
There is a plain design cost too
The five named lines press into the bottom of the panel
A dominant series crowds out the rest. Noticing that is reading honestly.
Figure 2, a Bar for Counts
The question: when across the day are channels busiest?
Hour of day is 24 discrete categories, not a continuous sweep
What is measured is a simple count of messages
Counts across categories call for the bar chart
One bar each, height is the count, side by side
The Code
chat_by_hour <- analysis %>%mutate(hour =hour(timestamp)) %>%count(hour, name ="messages")ggplot(chat_by_hour, aes(x = hour, y = messages)) +geom_col(fill ="#2f7d8a") +labs(title ="Chat volume by hour of day",x ="Hour of day (UTC)", y ="Messages in sample") + v2v::theme_v2v()
geom_col() draws a bar whose height is a value already computed.
The Daily Pulse
Busiest through UTC midday, peaking at 13:00 with 2,201
Quietest in the small hours, bottoming at 03:00 with 833
A swing of well over two to one
Not mysterious: the 2018 audience sat in the Americas and Europe
Twenty-four numbers become a rhythm in one glance.
Effect Sizes, and Their Standing
“Effect sizes are the most important outcome of empirical studies.”
Lakens (2013, p. 1)
Every figure showing a comparison must report its effect size
A taller bar is a visual impression, not a claim
The effect size gives magnitude in a scale-invariant unit
The Most Important Outcome
Lakens (2013)
“Effect sizes are the most important outcome of empirical studies.”
Lakens (2013, p. 1)
Most important to whom?
Your field reports p first. Why?
What would change if journals inverted that?
Figure 3, a Histogram for Shape
The central question: how long is a chat message?
And does the answer depend on the kind of channel?
Not a total and not a trend, but a distribution
Where values cluster, and where they thin
The histogram slices the range into equal bins
Two overlaid, so the shapes compare directly
The Code
msglen <- analysis %>%filter(!is.na(is_gaming)) %>%mutate(length_shown =pmin(message_length, 120))ggplot(msglen, aes(x = length_shown, fill = is_gaming)) +geom_histogram(binwidth =5, position ="identity", alpha =0.55) +labs(x ="Message length (characters, capped at 120)",y ="Messages", fill ="Gaming channel") + v2v::scale_fill_v2v() + v2v::theme_v2v()
Drop the unlabeled messages, cap the displayed length, overlay at partial transparency.
Two Choices You Must Disclose
binwidth = 5 covers five characters per bar
Wider smooths the shape, narrower roughens it
The number is a judgment you report
pmin(message_length, 120) caps the display
Twitch’s real limit is 500, so the last bar is a pile-up
Capping keeps the bulk legible
An unannounced cap is a quiet distortion. The axis label says so.
What the Shape Shows
Both groups share one shape: heavily right-skewed
A tall stack of very short messages on the left
A long thin tail to the right
Twitch chat is mostly brief, a word or an emote
The most important thing the histogram reveals, and invisible in any summary number.