HW 4 Rubric
Rubric
This is the rubric for AI feedback, not necessarily the rubric for grading.
Part 1 - Slam dunks
Question 1
Code reads the
2026 Draft Player Statssheet fromnba-draft-combine-2627.xlsxwithread_excel()or an equivalent readxl function.Code reads the
2026 Draft Pickssheet from the same workbook.Code excludes the three extraneous rows above the stats header and uses the correct column headings for both sheets.
- Note:
skip = 3is appropriate for the stats sheet. The picks sheet does not need rows skipped. Equivalent approaches are acceptable if they preserve the headers and player records.
- Note:
Missing values are represented as
NA, including blank cells and the-and-%placeholders.- Note: The
naargument inread_excel()is the simplest approach. Correct recoding after import is also acceptable.
- Note: The
All variable names in both data frames use snake_case, for example through
clean_names().- Note: Cleaning names in a separate pipeline is acceptable.
Code saves the two complete datasets as separate data frames without overwriting one with the other.
- Note: Names do not have to match the key.
The imported stats data frame has 78 rows and 16 columns, and the picks data frame has 50 rows and 7 columns.
- Note: These are checks on the saved data, not a requirement to write code reporting dimensions. Do not repeat feedback if an earlier import error explains a mismatch.
Output displays the first ten rows of the stats data frame.
Output displays the first ten rows of the picks data frame.
- Note for items 8–9: Direct tibble display or
slice_head(n = 10), or an equivalent approach is acceptable. Only columns that fit on the page need to be shown. Displaying a preview should not truncate the saved data frame.
- Note for items 8–9: Direct tibble display or
Code style and readability: line breaks after each
|>, appropriate indentation, spaces around=signs, and spaces after commas.- Note: Indentation with two spaces, four spaces, or tabs is acceptable.
Question 2
Part a. Code identifies players in the stats data but absent from the supplied picks data.
- Note: Use
anti_join()with stats on the left and picks on the right, or an equivalent approach.
- Note: Use
Part b. Code identifies players in the supplied picks data but absent from the stats data.
- Note: Use
anti_join()with picks on the left and stats on the right, or an equivalent approach.
- Note: Use
Both comparisons match records by
player.- Note:
by = join_by(player),by = "player", or an equivalent comparison is acceptable.
- Note:
Part a. Output is displayed and contains 33 unmatched stats records.
- Note: A tibble preview showing the total row count is acceptable; displaying all 33 rows is not required.
Part b. Output is displayed and contains the five unmatched picks: AJ Dybantsa, Mikel Brown Jr., Nate Ament, Sergio de Larrea, and Chris Cenac Jr.
Part c. Narrative reports both counts correctly: 33 stats-only records and 5 picks-only records.
- Note: If an earlier import or comparison error changes the counts, identify the underlying error and check whether the narrative matches the student’s output. Avoid repeating feedback for the same error.
Part c. The summary is a single, concise sentence that makes clear which count refers to each type of mismatch.
The response describes mismatches in the supplied data without concluding that all unmatched stats players went undrafted.
- Note: The picks data contain only the first 50 picks. Name differences can also cause mismatches, such as “AJ Dybantsa” versus “Anicet Dybantsa”. Identifying or fixing these name differences is not required for this question.
Code style and readability: line breaks after each
|>, appropriate indentation, spaces around=signs, and spaces after commas.- Note: Indentation with two spaces, four spaces, or tabs is acceptable.
Question 3
Part a. Code uses a single pipeline starting with the stats data, excludes players with missing positions, and groups by
pos.Part a. Code calculates
nas the number of players in each position, including players with missing weights.- Note: Do not filter out missing weights before counting. PG should have 19 players and SF should have 8; each includes one player with a missing weight.
Part a. Code calculates
mean_weightas the mean weight in pounds, excluding missing weights.- Note:
mean(weight_lbs, na.rm = TRUE)is appropriate. Without missing-value handling, PG and SF haveNAmeans.
- Note:
Part a. Code calculates
median_weightas the median weight in pounds, excluding missing weights.- Note:
median(weight_lbs, na.rm = TRUE)is appropriate. The median is required.
- Note:
Part a. Output is displayed with five rows and exactly four columns:
pos,n,mean_weight, andmedian_weight.Part a. Results are ordered by descending mean weight: C, PF, SF, SG, PG.
Part b. A one-sentence justification explains why the selected plot type is useful for comparing weight distributions across positions.
- Note: Boxplots, violin plots, jittered or dodged dot plots, and appropriately grouped density plots or histograms are acceptable.
Part b. Code creates the selected plot using only players with nonmissing positions and weights.
- Note: Filtering missing weights here is appropriate; filtering them before counting in Part a is not.
Part b. The plot has an informative title and readable axis labels, including weight units.
Part c. In at most three sentences, the narrative interprets the relationship between position and weight using both the summary statistics and visualization.
- Note: Centers tend to be heaviest and point guards lightest, with PF, SF, and SG in between. Accept other accurate comparisons supported by the results. If an earlier error affects the results, distinguish that error from the student’s interpretation rather than repeating the same feedback.
Code style and readability: line breaks after each
|>and+, appropriate indentation, spaces around=signs, and spaces after commas.- Note: Indentation with two spaces, four spaces, or tabs is acceptable. An explicit
.groups = "drop"insummarize()is also acceptable.
- Note: Indentation with two spaces, four spaces, or tabs is acceptable. An explicit
Question 4
Code joins the two datasets by
playerand retains only players present in both.- Note:
inner_join()withby = join_by(player),by = "player", or an equivalent approach is acceptable.
- Note:
The plot excludes players with missing
weight_lbs.The plot excludes players with missing
three_quarter_sprint_sec.- Note for items 2–3: Filtering before or after the join is acceptable. Deliberately handling missing values in the plotting code is also acceptable; unexplained warnings should receive feedback. With correct imports and matching, the plot contains 41 players.
Code produces a scatterplot with
weight_lbson the x-axis andthree_quarter_sprint_secon the y-axis.The plot has an informative, readable title.
Axis labels are human-readable and identify weight in pounds and sprint time in seconds.
Players whose
affiliationis"Duke"are colored Duke blue (#00539B).Players not from Duke are colored light gray (
gray70).All three Duke players are labeled with their names: Cameron Boozer, Isaiah Evans, and Maliq Brown.
- Note:
geom_label(),geom_text(), or an equivalent labeling approach is acceptable.
- Note:
Duke player labels do not overlap their own points.
- Note: Overlap with non-Duke points is allowed, and label positions need not exactly match the reference. Repelled labels are also acceptable.
The plot uses
theme_minimal().Code style and readability: line breaks after each
|>and+, appropriate indentation, spaces around=signs, and spaces after commas.- Note: Indentation with two spaces, four spaces, or tabs is acceptable. If an earlier error affects multiple checks, identify the underlying error rather than repeating the same feedback.
Part 2 - Shaken trusts
Question 5
In a 2–3 sentence response, the narrative identifies the main finding: 89% of U.S. adults in 2026 perceive government corruption as widespread, the highest percentage in Gallup’s 20-year series.
The narrative includes one accurate comparison across political groups or over time.
- Note: For example, 91% of Democrats versus 83% of Republicans, or an increase from 79% in 2025 to 89% in 2026. Other comparisons supported by the article are acceptable.
The narrative explains that the survey measures perceptions of corruption, not the actual extent of corruption.
Question 6
Part a. The downloaded data are from the chart “Perceived Government Corruption Spans Political Spectrum” and are saved in the
datafolder.- Note: The filename need not match the key. If the repository is unavailable, file location cannot be verified from the rendered answer alone.
Part b. Code reads the downloaded file with a function appropriate for its delimiter and correctly identifies the columns.
- Note: The source file used in the key is tab-delimited despite its
.csvextension, soread_tsv()is appropriate. Other download formats may require a different reader.
- Note: The source file used in the key is tab-delimited despite its
Part b. Code tidies the data into one observation per political group and year.
- Note: The resulting data should contain nine observations: three groups for each of 2024, 2025, and 2026. Equivalent tidy structures and variable names are acceptable.
Part b. Code converts the percentage values to numeric without losing observations or changing their meaning.
- Note: Removing
%withstr_remove()and then usingas.numeric(), or usingparse_number(), are acceptable. Either a 0–100 or 0–1 scale is fine if the plot labels are consistent.
- Note: Removing
Part c. Each group’s line connects its correct values in chronological order.
- Note: For 2024, 2025, and 2026, respectively: Republicans 87%, 81%, 83%; Democrats 57%, 76%, 91%; independents 78%, 85%, 90%.
Part c. Group colors match the reference: red for Republicans, dark blue for Democrats, and muted teal for independents.
- Note: The reference colors are
#F32735,#002169, and#6A8C89, respectively. Visually close colors are acceptable.
- Note: The reference colors are
Part c. Republican and Democratic lines are solid, and the independent line uses a dotted or short-dashed pattern matching the reference.
Part c. Points appear at the 2024 and 2026 endpoints for each group, with no point markers at 2025.
Part c. The x-axis displays 2024, 2025, and 2026 at equally spaced positions.
Part c. The y-axis spans 40% to 100%, with ticks every 10 percentage points.
- Note: Percent signs may appear on every tick or only the top tick, as in the reference. Watch for accidental labels such as 8,700% when the data already use a 0–100 scale.
Part c. The plot includes the title, survey question,
% Yeslabel, and note that party identification data are available only from 2024 onward.- Note: Equivalent wording, line breaks, fonts, and spacing are acceptable. Reproducing the Gallup logo is not required.
Part c. A readable legend above the plot identifies Republicans, Democrats, and independents in that order.
Part c. The 2026 endpoints are labeled 83, 91, and 90 for Republicans, Democrats, and independents, respectively.
- Note: Labels should be readable and associated with the correct lines; the 91 and 90 labels should not overlap.
Part c. The narrative explains which non-default features were matched and how the code produced them.
- Note: Examples include manual colors and line types, endpoint labels, axis scales, legend placement, a pale green background, and horizontal gridlines. Explanations should agree with the submitted code.
Part c. For any required feature that was not matched, the narrative describes what was tried and what did not work.
- Note: No such explanation is needed if all required features are matched. The plot does not need to be interactive.
Code style and readability: line breaks after each
|>and+, appropriate indentation, spaces around=signs, and spaces after commas.- Note: Indentation with two spaces, four spaces, or tabs is acceptable. If one import or transformation error affects several plot features, identify that underlying error rather than repeating the same feedback.
Part 3 - Stuck landings
Question 7
Part a. Code reads the two datasets and saves them as
gymnasticsandcontinents.Part a. Output displays the first ten rows of each data frame.
- Note for parts a–b: Direct tibble display or
slice_head(n = 10), or an equivalent approach is acceptable.
- Note for parts a–b: Direct tibble display or
Part b. Code uses
left_join()withgymnasticsas the left data frame andcontinentsas the right, withby = join_by(country == entity), or an equivalent join that preserves all gymnastics records.Part b. Code saves the result as a new, reasonably named data frame.
- Note: Any short but informative name is acceptable, such as
gymnastics_continentsorgym_cont.
- Note: Any short but informative name is acceptable, such as
Part b. Output displays the first ten rows of the joined data frame.
Part c. Code filters for vault routines (
filter(event == "Vault")).Part c. Code groups by
continentand summarizes the meanscorein a single pipeline.- Note: Performing the join within the same summary pipeline is also acceptable.
Part d. Narrative states the summary has 7 rows: six continents and an
NAgroup for unmatched countries, each summarizing vault routines.Part d. Narrative identifies South America as highest and Africa as lowest.
Part e. Code uses
anti_join()correctly to identify unmatched countries (Great Britain, AUT, Chinese Taipei).Part e. Narrative explains why each country failed to match and describes a reasonable fix for each.
- Note: Any reasonable fix, such as recoding names with
mutate()andcase_when()/if_else(), or explicitly assigning the correct continents to unmatched records, is acceptable. Implementing the fix is not required.
- Note: Any reasonable fix, such as recoding names with
Code style and readability: Line breaks after each |>, line breaks after each +, proper indentation, spaces around = signs if they are present, and spaces after commas if they are present.
Overall
These rubric items are not tied to a specific question but apply to the entire homework. Check for these if explicitly asked to provide overall feedback or if you notice them while reviewing the homework or as the student is finishing up working on the last question.
Rubric
Final .qmd and rendered PDF files exist in the remote GitHub repository.
At least three commits were made and pushed to the GitHub repository, and the submission is reasonably organized.
An AI-use disclosure states whether an LLM or generative AI tool was used (Yes or No).
If AI was used, the disclosure describes all the ways it was used.
If AI was used, the disclosure lists all tools used, including model names and versions if known.
Any non-AI sources used are cited, with the question(s) for which each source was used.