Describing the Data [R]

Week 12 · Chapter 12 · MC 501 Research Methods for Mass Communications

Dr. Alex Leith

Today’s Agenda

  • The grammar of graphics, in three parts
  • A line, a bar, and a histogram
  • Effect sizes, and the choices a figure discloses
  • Figures other people can actually read

What We Build Today

Working session

  • Three figures and one descriptives table
    • A line for change, a bar for counts, a histogram for shape
    • Each geometry chosen for a different kind of question
  • They go into your Results and onto your poster
  • Plus the rule that turns a picture into a claim

Expect every first version to look wrong. The fix is usually one line.

Rows You Cannot Read

  • The analysis table is complete and unreadable
    • Nobody scrolls 35,267 rows and learns anything
    • The numbers are present, the pattern is not visible
  • A figure turns a long column into a shape

Visualization is not decoration applied afterwards. It is part of the analysis.

The Grammar of Graphics

Three parts do most of the work (Wilkinson, 2005; Wickham, 2016):

ggplot(data, aes(x = ..., y = ...)) +
  geom_line()
  • Data: the table being shown
  • Aesthetic mappings: which column goes to which property
  • Geometry: the mark used to draw it

Swap geom_line() for geom_col() and the same data becomes bars.

Figure 1, a Line for Change

  • The question: how did viewership move across the week?
    • And how was it split across games?
  • Time on x, a rising and falling y, so a line chart
    • A line implies sequence, and real space between points
    • Right for time, wrong for unordered categories
  • The data needs shaping first

Shaping the Data

viewers_over_time <- streams %>%
  filter(!is.na(game)) %>%
  mutate(
    timestamp = as.POSIXct(date / 1000,
                           origin = "1970-01-01", tz = "UTC"),
    six_hour  = floor_date(timestamp, "6 hours"),
    category  = fct_lump_n(game, n = 5)
  ) %>%
  group_by(six_hour, category) %>%
  summarise(total_viewers = sum(viewers), .groups = "drop")

Bucket into six-hour windows, lump everything past the top five into “Other”, and sum viewers inside each bucket.

Drawing the Line

ggplot(viewers_over_time,
       aes(x = six_hour, y = total_viewers, color = category)) +
  geom_line(linewidth = 0.9) +
  labs(title = "Concurrent viewers by game category",
       x = "Date (UTC, six-hour buckets)",
       y = "Total viewers in bucket",
       color = "Category") +
  v2v::scale_colour_v2v() +
  v2v::theme_v2v()

theme_v2v() applies publication-ready fonts, spacing and gridlines in one call.

What Dominates, and Why

  • The towering line is “Other”, the catch-all beyond the top five
    • Substantively: viewership really is spread across a long tail
    • As a caution: lumping fifty categories guarantees that result
  • There is a plain design cost too
    • The five named lines press into the bottom of the panel

A dominant series crowds out the rest. Noticing that is reading honestly.

Figure 2, a Bar for Counts

  • The question: when across the day are channels busiest?
  • Hour of day is 24 discrete categories, not a continuous sweep
  • What is measured is a simple count of messages
  • Counts across categories call for the bar chart
    • One bar each, height is the count, side by side

The Code

chat_by_hour <- analysis %>%
  mutate(hour = hour(timestamp)) %>%
  count(hour, name = "messages")

ggplot(chat_by_hour, aes(x = hour, y = messages)) +
  geom_col(fill = "#2f7d8a") +
  labs(title = "Chat volume by hour of day",
       x = "Hour of day (UTC)", y = "Messages in sample") +
  v2v::theme_v2v()

geom_col() draws a bar whose height is a value already computed.

The Daily Pulse

  • Busiest through UTC midday, peaking at 13:00 with 2,201
  • Quietest in the small hours, bottoming at 03:00 with 833
  • A swing of well over two to one
  • Not mysterious: the 2018 audience sat in the Americas and Europe

Twenty-four numbers become a rhythm in one glance.

Effect Sizes, and Their Standing

“Effect sizes are the most important outcome of empirical studies.”

Lakens (2013, p. 1)

  • Every figure showing a comparison must report its effect size
  • A taller bar is a visual impression, not a claim
  • The effect size gives magnitude in a scale-invariant unit

The Most Important Outcome

Lakens (2013)

“Effect sizes are the most important outcome of empirical studies.”

Lakens (2013, p. 1)

  • Most important to whom?
  • Your field reports p first. Why?
  • What would change if journals inverted that?

Figure 3, a Histogram for Shape

  • The central question: how long is a chat message?
    • And does the answer depend on the kind of channel?
  • Not a total and not a trend, but a distribution
    • Where values cluster, and where they thin
  • The histogram slices the range into equal bins
  • Two overlaid, so the shapes compare directly

The Code

msglen <- analysis %>%
  filter(!is.na(is_gaming)) %>%
  mutate(length_shown = pmin(message_length, 120))

ggplot(msglen, aes(x = length_shown, fill = is_gaming)) +
  geom_histogram(binwidth = 5, position = "identity", alpha = 0.55) +
  labs(x = "Message length (characters, capped at 120)",
       y = "Messages", fill = "Gaming channel") +
  v2v::scale_fill_v2v() +
  v2v::theme_v2v()

Drop the unlabeled messages, cap the displayed length, overlay at partial transparency.

Two Choices You Must Disclose

  • binwidth = 5 covers five characters per bar
    • Wider smooths the shape, narrower roughens it
    • The number is a judgment you report
  • pmin(message_length, 120) caps the display
    • Twitch’s real limit is 500, so the last bar is a pile-up
    • Capping keeps the bulk legible

An unannounced cap is a quiet distortion. The axis label says so.

What the Shape Shows

  • Both groups share one shape: heavily right-skewed
    • A tall stack of very short messages on the left
    • A long thin tail to the right
  • Twitch chat is mostly brief, a word or an emote

The most important thing the histogram reveals, and invisible in any summary number.

The Descriptives

msglen %>%
  group_by(is_gaming) %>%
  summarise(
    n      = n(),
    mean   = mean(message_length, na.rm = TRUE),
    median = median(message_length, na.rm = TRUE),
    sd     = sd(message_length, na.rm = TRUE),
    .groups = "drop"
  ) %>%
  v2v::pretty_table()

Split by gaming status and compute the four numbers a Results section needs.

The Numbers

  • Non-gaming: n 3,457, mean 33.70, median 16, sd 60.68
  • Gaming: n 31,309, mean 28.49, median 17, sd 38.47

Read the means and non-gaming chat is longer. Read the medians, 16 against 17, and they are all but identical, pointing the other way.

When Mean and Median Disagree

  • The mean is sensitive to a long tail, the median is not
    • A few very long messages lift the average
    • The middle value is untouched
  • Non-gaming has the heavier tail, and its sd records it
  • The gap in means is the tail’s work

Lean on the mean alone and you misreport a typical message.

What to Report, and When

  • Two group means: Cohen’s d with a 95% confidence interval
  • A cross-tabulation: Cramer’s V
  • A correlation: r and its confidence interval
  • run_t_test() and run_chi_square() return caption-ready output

Put the effect size in the caption or a companion table, never a footnote.

Figures Other People Can Read

  • Roughly one in twelve men has a color-vision deficiency
    • Choose a colorblind-safe palette
    • Do not let color carry the message alone
    • Vary line type, or split into panels
  • Alt text is what a screen-reader user gets instead
    • With none, the figure is simply missing

State the chart type, the axes, and the takeaway, in fig-alt.

Checkpoint

You should now have:

  • Three rendered figures, each ending in v2v::theme_v2v()
  • A descriptives table with n, mean, median, and sd
  • fig-alt text written for all three
  • The binwidth and the cap named in your labels or captions
  • One sentence per figure, ready to paste into Results

Before Week 13

  • Submit Describing Data [R], 100 points
    • Figures, descriptives, alt text, and the disclosed choices
  • Read Chapter 13 with the graduate toggle on
    • The assigned article: Lakens et al. (2018), “Justify your alpha”
  • Write a journal entry, 450 to 500 words

The histogram left a precise question. Week 13 answers it.

Questions?