HW 1 Rubric
Rubric
This is the rubric for AI feedback, not necessarily the rubric for grading.
Question 1
All code uses appropriate tidyverse functions and syntax.
Code produces a histogram with binwidth 5.
Code produces a histogram with binwidth 50.
Code produces a histogram with binwidth 500.
All plots have informative and human-readable titles and/or subtitles and axis labels.
Code style and readability: Line breaks after each +, proper indentation, spaces around = signs if they are present, and spaces after commas if they are present.
Narrative mentions the most appropriate binwidth among the choices.
- Note: Most appropriate binwidth among the choices: 50.
Narrative states a well-articulated and accurate reason for the chosen binwidth.
- Sample well-articulated and accurate reason: A much smaller binwidth results in too much detail in the histogram and a much larger binwidth hides important details binning too much of the data together.
Question 2
Code produces a boxplot with
weeks_on_charton either the x-axis or y-axis.Plot has an informative title and/or subtitle and a human-readable axis label on the axis where
weeks_on_chartis plotted.- Note: A label on the other axis is optional since it is not meaningful in a boxplot.
Narrative mentions shape: the distribution is right-skewed and unimodal.
Narrative mentions center: a roughly estimated value or range of values for the median, approximately 45 weeks.
Narrative mentions spread: a roughly estimated value or range of values for the IQR or for the bounds of the middle 50% of the distribution, approximately 15 to 124 weeks.
Code uses
filter()to identify all outliers.- Note: The cutoff should be around 285 but doesn’t have to be exactly that value.
Code uses
select()to keep onlyartist,song,weeks_on_chart, andrankin the outlier output.Code arranges the data frame in descending order of
weeks_on_chart.Narrative mentions something about expectedness/unexpectedness of outliers.
- Note: Any reasonable answer is acceptable, including “having had no expectations”.
Code style and readability: Line breaks after each
|>, line breaks after each+, proper indentation, spaces around=signs if they are present, and spaces after commas if they are present.
Question 3
Code creates
positions_below_peakwithmutate()usingrank - peak_rank.Code uses a single data wrangling pipeline to produce the requested output.
Code identifies the 10 songs farthest below their peak rank.
Code orders the songs so that the largest values of
positions_below_peakappear first.Code uses
select()to keepsong,artist,rank,peak_rank,positions_below_peak, andweeks_on_chartin the output.The output contains exactly 10 songs.
Narrative identifies a pattern supported by the output.
Narrative selects one song from the results and gives a plausible explanation for why it may be charting below its earlier peak.
The explanation is clearly presented as a hypothesis rather than a conclusion supported by these data alone.
Code style and readability: Line breaks after each
|>, proper indentation, spaces around=signs if they are present, and spaces after commas if they are present.
Question 4
Code produces a scatterplot with
geom_point().Code maps
weeks_on_chartto the x-axis andpositions_below_peakto the y-axis.Plot has an informative title and/or subtitle and human-readable axis labels that include appropriate units.
Narrative accurately describes the relationship between
weeks_on_chartandpositions_below_peak, such as a very weak positive relationship with substantial variability.Code uses a single pipeline to identify a song or songs that have been on the chart for a long time while remaining close to its peak rank.
Code uses
select()to displaysong,artist,rank,peak_rank,positions_below_peak, andweeks_on_chart.Narrative correctly identifies where the selected song appears on the scatterplot, such as far to the right and close to the bottom.
Code style and readability: Line breaks after each
|>and+, proper indentation, spaces around=signs if they are present, and spaces after commas if they are present.
Question 5
Code uses appropriate tidyverse functions and syntax.
Code identifies the genres with at least 10 songs on the chart and excludes genres with fewer than 10 songs from the plot.
Code produces side-by-side boxplots with
streamsplotted bygenre.Plot has an informative title and/or subtitle and human-readable axis labels for genre and weekly streams.
Narrative accurately compares typical weekly stream counts across the included genres, such as by comparing their medians.
Narrative accurately compares the variability of weekly stream counts across the included genres, such as by comparing their IQRs or overall spreads.
Narrative answers the initial question by explaining whether songs from some genres tend to have higher weekly stream counts than others.
Code uses a single pipeline to identify the single most-streamed song among the included genres.
- Note: There are multiple ways of identifying this song. A hard-coded list of eligible genres is also acceptable if it includes only genres with at least 10 songs.
Code identifies the correct song, artist, genre, and weekly stream count: “Earrings” by Malcolm Todd, a Pop song, with 25,092,481 streams.
Code style and readability: Line breaks after each
|>, line breaks after each+, proper indentation, spaces around=signs if they are present, and spaces after commas if they are present.
Question 6
Code creates
stream_levelwithmutate()andif_else(), classifying songs according to whetherstreamsis above the overall median.Code produces a segmented bar plot with one bar for each genre and fills the bars according to
stream_level.Code includes only genres with at least 10 songs on the chart.
The y-axis represents proportions from 0 to 1.
Plot has an informative title and/or subtitle and human-readable axis labels.
Narrative compares the proportions of above-median songs across the included genres using reasonable approximate values.
Narrative correctly identifies Pop as having the highest proportion and Electronic as having the lowest proportion of songs above the overall median.
Code style and readability: Line breaks after each
|>, line breaks after each+, proper indentation, spaces around=signs if they are present, and spaces after commas if they are present.