HW 2 Rubric
Rubric
This is the rubric for AI feedback, not necessarily the rubric for grading.
Question 1
Part a - Code calculates the number of rows with
nrow()and columns withncol()or usesdim()to calculate both at once.Part a - Output shows 100 rows and 22 columns.
Part b - Code counts the counties in each
county_typewithcount().Part b - Code displays the county type counts in descending order of the number of counties with the
sortargument incount()or in a subsequent step witharrange().Part b - Output shows 50 Rural - Non-Metro counties, 28 Rural - Metro counties, 16 Suburban counties, and 6 Urban counties.
Part c - Introductory narrative accurately explains that each row represents a North Carolina county and describes the dataset in a reasonable way.
Code style and readability: appropriate indentation, spaces around equals signs if they are present, spaces after commas, and readable code organization.
- Note: Two or four spaces / tabs are ok for indentation.
Question 2
Code makes either a histogram or density plot, or a boxplot as well as histogram or density plot of
pop_dens_2020.Code uses a single pipeline to calculate the median, first quartile, and third quartile of
pop_dens_2020.Output reports the median, first quartile, and third quartile of
pop_dens_2020as approximately 112.0, 60.7, and 214.0 people per square mile, respectively.Narrative identifies the distribution as unimodal.
Narrative identifies the distribution as right-skewed, extremely right-skewed, or strongly right-skewed.
Narrative correctly identifies the median as the typical value and the quartiles as the bounds of the middle 50%.
The reported median is approximately 112.0 people per square mile.
The reported first quartile is approximately 60.7 people per square mile.
The reported third quartile is approximately 214.0 people per square mile.
Code style and readability: appropriate pipeline formatting, indentation, spaces around equals signs if they are present, and spaces after commas.
- Note: Two or four spaces / tabs are ok for indentation.
Question 3
Part a - Code
filter()s for Durham County and selects the requested columns.- Note: No narrative needed for this part.
Part b - Code
filter()s for counties withpop_dens_2020 > 500, selects the requested columns, and arranges the output in descending order ofpop_dens_2020.Part b - Narrative identifies all nine counties with population density greater than 500 people per square mile and provides a reasonable geographic description of their locations and links to a map online.
Part c - Code
filter()s for the county with the highest population density usingmax()and selects the requested columns.- Note: No narrative needed for this part.
Code style and readability: line breaks after each |>, appropriate indentation, spaces around equals signs if they are present, and spaces after commas.
- Note: Two or four spaces / tabs are ok for indentation.
Question 4
Code uses a single data-wrangling pipeline.
Code
filter()s counties withp_family_sustaining_wage < 0.50.Code
count()s the qualifying counties within eachcounty_type.Code
filter()s for only county types with at least 10 qualifying counties.Code outputs only the county type and the number of qualifying counties.
Output reports 45 Rural - Non-Metro and 17 Rural - Metro counties that meet the requirement.
Narrative identifies the two qualifying county types (Rural Metro and Non-Metro) and correctly notes the other two that do not qualify (Urban and Suburban).
Code style and readability: line breaks after each |>, appropriate indentation, spaces around equals signs if they are present, and spaces after commas.
- Note: Two or four spaces / tabs are ok for indentation.
Question 5
Part a - Code creates
broadband_levelwithmutate()andif_else().Part a - Code assigns
"At least 75%"whenp_broadband >= 0.75and"Below 75%"otherwise.Part a - The new variable is stored in the
nc_countydata frame.Part b - Code produces one segmented bar for each
county_typewith bars filled according tobroadband_level.Part b - Plot has an informative title and/or subtitle and human-readable axis labels, especially y-axis label that reflects “proportion” or similar.
Part c - Code calculates conditional probabilities for below and at least 75% broadband access for each county type in a single pipeline.
Part d - Narrative compares the percentages across county types and correctly identifies Urban as highest, Suburban as next, and the two rural types as much lower, with values from Part c.
Code style and readability: line breaks after each |> and +, appropriate indentation, spaces around equals signs if they are present, and spaces after commas.
- Note: Two or four spaces / tabs are ok for indentation.
Question 6
Code creates
p_edu_hewithmutate().Code defines p_edu_he as p_edu_assoc + p_edu_ba + p_edu_mapl.
Code stores p_edu_he in the nc_county data frame.
Part a - Code uses a single pipeline to calculate the six requested overall statistics.
Part a - Output includes the minimum, first quartile, median, mean, third quartile, and maximum.
Part b - Code uses a single pipeline grouped by
county_typeto calculate the same six statistics as in Part a.Part b - Code arranges in ascending order of median.
Part b - Narrative states that Urban counties have the highest typical attainment and Rural - Non-Metro counties have the lowest typical attainment.
Code style and readability: line breaks after each |>, appropriate indentation, spaces around equals signs if they are present, spaces after commas, and an explicit .groups choice where appropriate.
- Note: Two or four spaces / tabs are ok for indentation.
Question 7
Code maps
p_edu_heto the x-axis andcounty_typeto the y-axis.Code produces boxplots with
geom_boxplot().Code overlays the observations with
geom_beeswarm().Boxplot outliers are displayed as open circles.
Boxplots use partially transparent fills or an equivalent transparency adjustment.
The x-axis is formatted as percentages.
Plot displays blue-ish colors for Rural counties, red-ish for Suburban, and brown-ish for Urban.
Plot has correct titles and labels.
Plot uses the correct theme and doesn’t have a legend.
Code style and readability: line breaks after each |> and +, appropriate indentation, spaces around equals signs if they are present, and spaces after commas.
- Note: Two or four spaces / tabs are ok for indentation.
Question 8
Code groups the data by
county_typebefore identifying outliers.Code uses a single pipeline to produce the outlier output.
Code calculates the IQR within each county type.
Code calculates lower and upper fences using 1.5 times the IQR subtracted from 25th percentile and added to 75th percentile.
Code filters observations below the lower fence or above the upper fence.
Output displays the columns
county,county_type, andp_edu_he, in that order.Output is arranged in descending order of
p_edu_he. Note: This requires that data is ungrouped beforearrange().Output correctly identifies Chatham, Anson, Watauga, Moore, and Orange as the outliers, with no Urban outliers.
Code style and readability: line breaks after each |>, appropriate indentation, spaces around equals signs if they are present, and spaces after commas.
- Note: Two or four spaces / tabs are ok for indentation.
Question 9
Narrative gives a reasonable answer for why
p_edu_heshould be on the x-axis andp_family_sustaining_wageon the y-axis.Code produces a scatterplot with
p_edu_heon the x-axis andp_family_sustaining_wageon the y-axis.The x and y-axes are formatted as percentages.
Narrative describes the association as positive in direction.
Narrative describes the form as approximately linear.
Narrative gives a reasonable description of strength.
Narrative gives defensible criteria for unusual observations.
Code used to filter for the outliers matches the described criteria.
Code style and readability: line breaks after each |> and +, appropriate indentation, spaces around equals signs if they are present, and spaces after commas.
- Note: Two or four spaces / tabs are ok for indentation.
Question 10
Part a - A third variable other than
county_typeis selected, correctly classified as numerical or categorical, and reason for selection is sufficiently articulated.Part b - A reasonable statement about how the third variable will be incorporated into the plot is provided.
Part b - The plot is modified with the third variable in the way described.
Part c - An appropriate description of the relationship is given.
Code style and readability: line breaks after each |> and +, appropriate indentation, spaces around equals signs if they are present, and spaces after commas.
- Note: Two or four spaces / tabs are ok for indentation.