Lecture 1
Duke University
STA 199 - Fall 2026
August 24, 2026
Dr. Mine Çetinkaya-Rundel
Professor of the Practice
Director of First-Year Experience in Trinity College of Arts & Sciences
Associate Director of Undergraduate Studies, Statistical Science
Office: Old Chem 211C
Office hours:
Please share with at least two classmates…
04:00
Data science is an exciting discipline that allows you to turn raw data into understanding, insight, and knowledge.
We’re going to learn to do this in a modern and tidy way – more on that later!
This is a course on introduction to data science, with an emphasis on statistical thinking.
Today we’re going to explore some data together, following the data science cycle.
Our question: Are the stereotypes about country music actually true?


Scan the QR code or go to wooclap.com. Log in with your Duke NetID and enter the code OQPLQEK.
“Is it REALLY all about beer and trucks and girls and blue jeans? Are all the lyrics REALLY the same?”
— Grady Smith, country music critic
Grady Smith kept a spreadsheet of every country song that reached the Top 30 of Billboard’s Country Airplay chart.
Someone tidied it up and put it in a CSV (comma-separated values) file and released it as part of the TidyTuesday project.
More on what packages are on Thursday, but in a nutshell “get your tools out of the toolbox”:
Import the data called country-lyrics.csv in the folder called data using the read_csv() function:
# A tibble: 484 × 8
song artist featuring entered_top_30_in lyrics writers producer rough_order
<chr> <chr> <chr> <dbl> <chr> <chr> <chr> <dbl>
1 Wake … Craig… <NA> 2013 My fr… Josh O… Craig M… 1
2 Young… Kip M… <NA> 2013 Your … Kip Mo… Brett J… 2
3 Beat … Brett… <NA> 2013 Well … Brett … Brett E… 3
4 The H… Danie… <NA> 2013 She h… Brett … Brett J… 4
5 Every… Thomp… <NA> 2013 My mo… Keifer… RIch Re… 5
6 Goodn… Randy… <NA> 2013 Girl … Rob Ha… Derek G… 6
7 Cold … Josh … <NA> 2013 I hea… Brent … Mark Wr… 7
8 19 Yo… Dan +… <NA> 2013 It wa… Dan Sm… Dan Smy… 8
9 Hellu… Frank… <NA> 2013 Satur… Rodney… Marshal… 9
10 Chill… Cole … <NA> 2013 I got… Cole S… Jody St… 10
# ℹ 474 more rows
Every song in this dataset broke into the Top 30 in some year. One way to make sense of a variable like this is to visualize it.
country |>
count(entered_top_30_in) |>
mutate(prop = n / sum(n)) |>
ggplot(aes(x = entered_top_30_in, y = prop, group = 1)) +
geom_point() +
geom_line() +
scale_y_continuous(labels = percent_format(accuracy = 1)) +
scale_x_continuous(breaks = seq(2013, 2019, by = 1)) +
labs(
title = "When did these songs break into the Top 30?",
x = NULL,
y = NULL,
caption = "Country songs reaching the Top 30 of Billboard's Country Airplay chart."
) +
theme_minimal(base_size = 16)There are 484 songs in the data, but only 121 different artists. Let’s look at the 15 with the most Top 30 hits.
country |>
count(artist, sort = TRUE) |>
slice_head(n = 15) |>
ggplot(aes(y = fct_reorder(artist, n), x = n, fill = n)) +
geom_col(show.legend = FALSE) +
labs(
title = "Who had the most Top 30 country hits?",
y = NULL,
x = "Number of songs",
caption = "Country songs reaching the Top 30 of Billboard's Country Airplay chart."
) +
theme_minimal(base_size = 16)Songs are usually written by more than one person, so the writers variable holds a whole list in a single cell:
# A tibble: 484 × 2
song writers
<chr> <chr>
1 Wake Up Lovin You Josh Osborne, Matthew Ramsey, Trevo…
2 Young Love Kip Moore, Dan Couch, Westin Davis
3 Beat of the Music Brett Eldredge, Ross Copperman, Hea…
4 The Heart of Dixie Brett James, Caitlyn Smith, Troy Ve…
5 Everything I Shouldn't Be Thinking About Keifer Thompson, Shawna Thompson, D…
6 Goodnight Kiss Rob Hatch, Jason Sellers, Randy Hou…
7 Cold Beer With Your Name On It Brent Anderson, Clint Daniels
8 19 You + Me Dan Smyers, Shay Mooney, Danny Orton
9 Helluva Life Rodney Clawson, Chris Tompkins, Jos…
10 Chillin' It Cole Swindell, Shane Minor
# ℹ 474 more rows
Before we can visualize this variable, we need to tidy and transform it.
# A tibble: 517 × 2
writers n
<chr> <int>
1 Ashley Gorley 44
2 Shane McAnally 36
3 Ross Copperman 32
4 Josh Osborne 29
5 Rhett Akins 23
6 Hillary Lindsey 18
7 Jesse Frasure 18
8 Jon Nite 17
9 Zach Crowell 16
10 Luke Laird 14
# ℹ 507 more rows
country |>
separate_longer_delim(writers, delim = ", ") |>
count(writers, sort = TRUE) |>
slice_head(n = 15) |>
ggplot(aes(y = fct_reorder(writers, n), x = n, fill = n)) +
geom_col(show.legend = FALSE) +
scale_fill_taylor_c(album = "Speak Now") +
labs(
title = "Is country radio run by a small circle of hitmakers?",
y = NULL,
x = "Number of songs written",
caption = "Country songs reaching the Top 30 of Billboard's Country Airplay chart."
) +
theme_minimal(base_size = 16)The lyrics variable contains the entire text of each song:
“My friends call me up ’cause they know I’m down / Take me out to paint the town and help me get over it…”
We can use text mining techniques, like tokenizing to words, to explore it.
# A tibble: 15 × 2
word n
<chr> <int>
1 yeah 996
2 girl 781
3 baby 749
4 love 743
5 wanna 632
6 gonna 611
7 time 502
8 night 454
9 tonight 332
10 heart 293
11 home 272
12 kiss 263
13 world 261
14 gotta 256
15 eyes 240
We can also tokenize to bigrams (pairs of words):
# A tibble: 8 × 2
bigram n
<chr> <int>
1 in the 692
2 a little 513
3 on the 413
4 and i 381
5 i know 282
6 i don't 276
7 i can 254
8 in a 249
Hmm. Not very informative – these are all common words!
Let’s split each bigram in two and throw away the pairs built from common words:
country |>
select(lyrics) |>
unnest_tokens(bigram, lyrics, token = "ngrams", n = 2) |>
separate_wider_delim(bigram, delim = " ", names = c("word1", "word2")) |>
anti_join(stop_words, by = join_by(word1 == word)) |>
anti_join(stop_words, by = join_by(word2 == word)) |>
count(word1, word2, sort = TRUE) |>
slice_head(n = 20) |>
print(n = Inf)# A tibble: 20 × 3
word1 word2 n
<chr> <chr> <int>
1 yeah yeah 93
2 boo boo 81
3 ooh ooh 81
4 friday night 71
5 hey hey 45
6 tonight tonight 44
7 home home 36
8 yeah boy 33
9 saturday night 32
10 somebody's gotta 32
11 cold beer 31
12 dance floor 31
13 gonna love 31
14 country song 30
15 la la 28
16 wanna kiss 28
17 whoa whoa 28
18 da da 27
19 honky tonk 26
20 yeah baby 25
We could also ask a large language model to read every song and tell us what it’s about?
“What are common themes in these country song lyrics?” [Attached data file
country-lyrics.csv.] 24 August 2026. ChatGPT 5.6 Sol Light. chatgpt.com/share/e/6a8c7343-62e4-83ea-b114-715f9b847cea
Across the 484 songs in country-lyrics.csv, several recurring themes stand out:
Romantic love and attraction The dominant theme: falling in love, physical attraction, kissing, longing, and wanting to be together. Love-related language appears in roughly 94% of the songs.
Heartbreak and separation Many songs focus on breakups, regret, loneliness, missing someone, or being unable to move on. This appears in approximately 73% of the songs, often alongside love rather than separately from it.
Drinking and nightlife Beer, whiskey, bars, neon lights, and getting drunk frequently serve as settings or coping mechanisms—especially for heartbreak. Related language appears in about 60% of the songs.
Music and dancing Radios, songs, dancing, guitars, and jukeboxes are prominent. Music often brings couples together, evokes a memory, or provides the soundtrack for driving and partying. About 56% contain this imagery.
Roads, trucks, and escape Driving can represent romance, independence, returning home, or leaving a painful situation behind. Roads, trucks, and related imagery occur in approximately 55%.
Partying and carefree pleasure Friday nights, weekends, bonfires, dancing, and “good times” create an idealized space temporarily free from responsibility. This theme appears in about half the songs.
Faith, hope, and resilience References to God, prayer, heaven, belief, strength, and survival often accompany songs about hardship or gratitude. Related language occurs in approximately 43%.
Home, family, and belonging Home is both a physical place and an emotional anchor. Parents, family roots, and returning home appear in around 40% of the songs.
Rural and small-town identity Country roads, fields, porches, boots, farms, and small towns establish a distinct cultural setting and sense of authenticity. About 32% explicitly use this imagery.
Youth, memory, and nostalgia Songs frequently look back on summer, adolescence, first love, or an earlier and supposedly simpler time. Explicit nostalgic language appears in roughly 27%.
The broadest overarching theme is emotion tied to place: love, heartbreak, and personal identity are repeatedly expressed through small towns, roads, vehicles, bars, music, and home. The lyrics also contain a recurring tension between leaving and belonging—wanting to escape one’s circumstances while remaining emotionally attached to a person, hometown, or former life.
These percentages are approximate keyword-based measures, and the categories overlap: one song can simultaneously be about heartbreak, whiskey, driving, and home.
Ask anyone what country songs are about and you’ll likely hear: beer, trucks, dirt roads, blue jeans, tractors.
But we have the lyrics, so we can just check:
country |>
mutate(lyrics_lower = str_to_lower(lyrics)) |>
summarize(
truck = mean(str_detect(lyrics_lower, "truck")),
whiskey = mean(str_detect(lyrics_lower, "whiskey")),
beer = mean(str_detect(lyrics_lower, "beer")),
`dirt road` = mean(str_detect(lyrics_lower, "dirt road")),
`blue jeans` = mean(str_detect(lyrics_lower, "blue jeans")),
tractor = mean(str_detect(lyrics_lower, "tractor"))
) |>
pivot_longer(
everything(),
names_to = "stereotype",
values_to = "prop"
) |>
arrange(desc(prop))# A tibble: 6 × 2
stereotype prop
<chr> <dbl>
1 truck 0.120
2 whiskey 0.116
3 beer 0.0868
4 dirt road 0.0227
5 blue jeans 0.0227
6 tractor 0.0103
Only about 12% of these songs mention a truck, and 1% mention a tractor.
The words that actually dominate are girl, baby, love, and night.
The stereotype is memorable, not frequent – and that difference is only visible once you look at the data.
Tip
This is what we mean by statistical thinking: a claim that “feels true” is not a finding, but you can explore and verify the claim with data and the skills you’ll learn in this course.
All linked from the course website:
Both textbooks are free online and linked from the course website:
Computing:
R and Positron run in containers provided by Duke OIT – nothing to install
Bring a laptop (or Chromebook) to every class, fully charged – there aren’t enough outlets in the classroom for everyone
Daily in lecture, via Wooclap on your laptop or phone
Tracked for credit, but not based on correctness, only participation (and often they will be questions designed to make you think that might not have a single right answer!)
There are 25 lectures this semester: participate in at least 17 for full credit
Below that, the score is prorated, e.g. participating in 15/25 lectures gives (15/25) × 5% = 3%
Daily-ish in lecture
Not graded, but tracked for feedback on workflow
Hands-on practice with data analysis
A single exercise per lab, graded on turning in something reasonable + correctness
Completed in-person, in lab, in teams
Teams randomized each week until project teams assigned
Developed collaboratively, but turned in individually by the end of the lab session
Two lowest scores dropped
No late work accepted
Hands-on practice with data analysis
Some questions for practice with instant feedback by AI
1-2 randomly sampled questions to be graded by humans for correctness
Can consult with course team and peers, but completed and turned in individually by 11:59 pm on Wednesdays
Lowest score dropped
Up to 3 days late (-5% per day), no late work accepted after that
One-time late penalty waiver: Can be used on any homework assignment, no questions asked, must be requested from Dr. Knox before the deadline
Every assignment, same two steps:
Submit what you have even if you haven’t finished – you’ll get partial credit for your work.
Two exams + final, all in class:
All exams comprised of two parts:
Multiple choice questions
Short answer questions, on homework specifics (read: difficult/impossible to answer correctly without having done the homework)
Cover material from videos, readings, lectures, application exercises, homework, and labs
One sheet of notes for each exam, no larger than 8 1/2 x 11, both sides, must be prepared by you
No extensions and no make-up exams.
With documentation (Dean’s excuse or similar):
Improvement bonus: If you take all three exams, your final exam score replaces the lower of your two mid-semester exam scores – if the final score is higher.
Caution
Exam dates cannot be changed and no make-up exams will be given. If you can’t take the exams on Oct 7, Nov 18, and Dec 10, you should drop this class.
Dataset of your choice, method of your choice
Teamwork
Seven milestones, interim deadlines throughout semester
Final milestone: Presentation in lab and write-up, Thursday, Nov 5
Must be in lab, in-person to present
Peer feedback between teams on presentation content; peer evaluation within teams for contribution, alongside commit history
Some lab sessions allocated to project progress
Caution
The project presentation date (Nov 5) cannot be changed, and no late work is accepted for the project. You must complete the project to pass this class.
Randomized at first for weekly labs
Then selected by you or assigned by me if you don’t express a preference for project and remaining labs
Expectations and roles
For the project: Peer evaluation during teamwork and after completion
| Category | Percentage |
|---|---|
| Lectures (attendance + participation) | 5% |
| HW | 5% |
| Labs | 10% |
| Project | 15% |
| Exam 1 | 20% |
| Exam 2 | 20% |
| Final | 25% |
See course syllabus for how the final letter grade will be determined.
All linked from the course website at