
Lecture 12
Duke University
STA 199 - Fall 2026
September 30, 2026
Check your email (and your spam folder) for an email from TEAMMATES
Fill out your peer evaluation for your team members
If you fail to submit your feedback, you’re not eligible to receive the points they give you.
Go to your ae project in Positron.
If you haven’t yet done so, make sure all of your changes up to this point are committed and pushed, i.e., there’s nothing left in your source control pane.
If you haven’t yet done so, pull to get today’s application exercise file: ae-07-chronicle-scrape-I.qmd.
Work through the application exercise in class, and render, commit, and push your edits by the end of class.
Exam 1 is on Wednesday, starts promptly at 1:25 pm!
Covers all material up to and including today’s lecture
Sample exam + model solutions for select HW questions are posted
Exam will be in class, closed book, no calculators, no computers, just 1 notes sheet
Notes sheet: 8.5x11, both sides, hand written or typed, any content you want, must be prepared by you
We’ll direct you to your seat so get here early to allow time to get settled and read the instructions
blizzard_salary has 409 rows and 4 columns. Each row is a worker who completed the spreadsheet. Which statements are correct? Select all that apply.
percent_incr is numerical and discrete.salary_type is numerical.annual_salary is numerical. Correctperformance_rating is categorical and ordinal. CorrectA friend notices that the total number of observations in this plot seems lower than the total in blizzard_salary (409 workers). What might be going on here? Explain your reasoning.
Observations with missing values for both variables won’t be represented on the plot.
For Top performers, a team writes: “The relationship is positive, having a higher salary results in a higher percent increase. There is one clear outlier.” Which peer review notes are accurate and helpful? Select all that apply.

Which plot is produced by this code? Correct: D.
Based on this output, which statements must be true? Select all that apply.
# A tibble: 336,776 × 19
year month day arr_delay carrier dep_time sched_dep_time
<int> <int> <int> <dbl> <chr> <int> <int>
1 2013 1 1 11 UA 517 515
2 2013 1 1 20 UA 533 529
3 2013 1 1 33 AA 542 540
4 2013 1 1 -18 B6 544 545
5 2013 1 1 -25 DL 554 600
6 2013 1 1 12 UA 554 558
7 2013 1 1 19 B6 555 600
8 2013 1 1 -14 EV 557 600
9 2013 1 1 -8 B6 557 600
10 2013 1 1 8 AA 558 600
# ℹ 336,766 more rows
# ℹ 12 more variables: dep_delay <dbl>, arr_time <int>,
# sched_arr_time <int>, flight <int>, tailnum <chr>,
# origin <chr>, dest <chr>, air_time <dbl>, distance <dbl>,
# hour <dbl>, minute <dbl>, time_hour <dttm>
flights is a tibble. Correct.flights has 10 rows.flights has 8 columns.carrier is a character variable. Correct.flights.Which statements about the code and its result are true? Select all that apply.
# A tibble: 1 × 1
`sum(release_country == "United States")`
<int>
1 435
release_country equals "United States", resulting in logical values. Correct.release_country is not "United States" and counts the remaining rows.TRUE as 1 and each FALSE as 0. Correct.# A tibble: 6 × 6
rating Other Drama Action Comedy Horror
<fct> <int> <int> <int> <int> <int>
1 G 5 1 1 1 0
2 PG 38 13 10 18 0
3 PG-13 19 25 35 35 0
4 R 45 50 57 96 21
5 NC-17 1 2 0 1 0
6 Not Rated 4 11 4 6 1
Which statements are true? Select all that apply.
movies data frame to six.country_inflation has 38 rows and 34 columns: one row per country, a country column, and annual inflation columns for 1993–2025.
How many rows and columns does the result have? Show where each number comes from. What does each row represent, what did pivot_longer() do, and what does names_transform = as.numeric accomplish?
Rows: 1,254. Columns: 3. Each row represents one country in one year.
country column. Equivalently, 2025 − 1993 + 1 = 33 years.\[38 \times 33 = 38 \times 30 + 38 \times 3 = 1{,}140 + 114 = 1{,}254\]
The output columns are country + year + inflation_rate = 3:
country is retained and repeated for each year.year.inflation_rate.Missing inflation values still get rows; this code does not drop them.
country_inflation |>
filter(!is.na(`1993`), !is.na(`2025`)) |>
mutate(inf_ratio = `2025` / `1993`) |>
select(country, inf_ratio) |>
arrange(desc(inf_ratio)) |>
slice_head(n = 5)# A tibble: 5 × 2
country inf_ratio
<chr> <dbl>
1 New Zealand 2.20
2 Australia 1.64
3 Denmark 1.51
4 Ireland 1.50
5 United Kingdom 1.48
Part 1: What does the resulting data frame represent? What does New Zealand’s inf_ratio of 2.20 mean?
Part 2: What would change if the filter() step were removed? Why could that be misleading?
Part 1: These are the five countries with the largest 2025-to-1993 inflation ratios, in descending order.
arrange(desc(inf_ratio)) puts the largest ratios first. slice_head(n = 5) keeps the first five rows in that order.Part 2: Without filter(), six countries would have NA ratios because at least one of their two rates is missing.
arrange() puts these NAs last. For these data, the same five countries would still be returned.slice_head() do?It selects rows by their current positions. It does not sort or calculate a summary.
| Stage | Rows | What happened? |
|---|---|---|
| Start | 38 | One row per country |
filter(...) |
32 | Keep countries with both endpoint rates |
mutate(...) |
32 | Add one ratio to each row |
select(...) |
32 | Keep two columns |
arrange(desc(inf_ratio)) |
32 | Put largest ratios first |
slice_head(n = 5) |
5 | Keep rows 1–5 |
Removing only slice_head(n = 5) keeps all 32 eligible countries. It does not restore the six countries excluded by filter().