Week 11 · Chapter 11 · MC 501 Research Methods for Mass Communications
Working session
The least rewarding evening of the term, and where researchers spend most of their hours.
filter() keeps rows meeting a conditionselect() keeps or drops columnsmutate() adds or changes a computed columnarrange() reorders rows by a column’s valuesgroup_by() with summarize() collapses rows per groupEach takes a data frame and returns one, which is what lets the pipe chain them.
The first keeps only Bob Ross’s messages and drops every other row. The second keeps three columns and drops the rest.
as.POSIXct() turns a number into a date-time value, counting from the origin you give it.
date counts milliseconds since 1 January 1970
as.POSIXct() expects seconds, a thousand times coarserdate / 1000 converts first, and skipping it is how this fails.
# A tibble: 3 × 2
date timestamp
<dbl> <dttm>
1 1542578127023 2018-11-18 21:55:27
2 1542578135562 2018-11-18 21:55:35
3 1542578143610 2018-11-18 21:55:43
The type moved from <dbl> to <dttm>, and only from a date-time can you pull the hour, the weekday, or the calendar date.
str_length() counts the characters in a string, and mutate() writes the count into a new column.
# A tibble: 3 × 2
message message_length
<chr> <int>
1 "???????????????????????" 23
2 "TriEasy Clap TriEasy Clap TriEasy Cla…" 441
3 "cmonBruh" 8
message_length is ratio: true zero, equal intervals, averageableThat spread is the raw material, and Week 12 looks hard at its shape.
“Tidy data is a standard way of mapping the meaning of a dataset to its structure.”
Wickham (2014, p. 4)
Wickham (2014)
“Tidy data is a standard way of mapping the meaning of a dataset to its structure.”
Wickham (2014, p. 4)
Drop snapshots with no category, count how often each channel streamed each category, then keep each channel’s modal category, one row per channel.
Listing them states the rule, so the judgment is visible in the code.
left_join() keeps every chat row and attaches the matching is_gaming value, matching on the shared channel key.
##bobross in one source, bobross in anotherv2v fixture is normalized, and raw data is notWhen a join returns nothing, check the keys first.
# A tibble: 3 × 2
is_gaming n
<lgl> <int>
1 FALSE 3457
2 TRUE 31309
3 NA 501
Counting the new column is one line, and it is the line that catches the problem you did not anticipate.
NA Is the Correct OutcomeNANA records that the study cannot classify themThey sit the comparison out rather than being guessed into a side.
Four transformations, run in sequence from the raw fixture, rebuild the table every later result depends on.
Rows: 35,267
Columns: 8
$ id <int> 89, 551, 1033, 2094, 2803, …
$ channel <chr> "sodapoppin", "xqcow", "forsen", …
$ sender <chr> "madzee", "prometheanow", …
$ message <chr> "???????????????????????", …
$ date <dbl> 1542578127023, 1542578135562, …
$ timestamp <dttm> 2018-11-18 21:55:27, …
$ message_length <int> 23, 441, 8, 12, 206, …
$ is_gaming <lgl> TRUE, FALSE, TRUE, TRUE, FALSE, …
mutate() and filter() is a methodological decisionIdentify three decisions a reviewer could challenge and comment each.
data/raw/ is read-only after the first commit
data/processed/If it does not reproduce without intervention, you have a hidden dependency.
/ 1000could not find function "%>%": the tidyverse never loadedNA: the keys disagree, often a #object 'channel_type' not found: chunks run out of orderRun ?v2v::common_errors for the ones specific to this data.
You should now have:
chat object with 8 columns and 35,267 rowstimestamp of type <dttm> and a message_length of type <int>is_gaming with 3,457 FALSE, 31,309 TRUE, and 501 NA