Explore a survey dataset.
Four tasks, one dataset: clean, summarize, visualize and interpret.
Entirely synthetic educational data • Beginner–Intermediate • About 25 minutes
Data and exercise files
Start with the raw data and attempt the tasks. The package also includes a data dictionary, clean reference data and Python/R solution scripts.
Each row is a person record. id: identifier; group: observational A/B group; minutes: daily learning time (minutes, blank/99 missing); transport: bus, walk or bike. The raw data contain intentional issues.
Clean the data
Find duplicate records, missing values and inconsistent category spellings. How many people and valid times remain?
Expected result and explanation
Remove one copy of the exact duplicate id=5 from 21 rows. Trim and lowercase transport. Blank and 99 are missing for minutes; do not replace them with zero. Result: 20 unique people, 18 valid times and 2 missing. The 99 rule belongs to this data dictionary, not every dataset.
Build a frequency table
Count transport categories in the clean data and calculate percentages. State your denominator.
Expected result and explanation
bus: 10 (50%), walk: 6 (30%), bike: 4 (20%). The denominator is 20 people; transport has no missing values. Missing times do not reduce this denominator to 18.
Plot and interpret
Create a bar chart of transport frequencies. Add a title, axis unit and data note.
Expected result and explanation
Bar heights should be 10, 6 and 4. The vertical axis counts people and starts at zero. Note: Synthetic educational data, n=20. Bus is the most frequent choice in these data; this does not establish population preferences.
Compare groups
Find the mean, median and sample standard deviation of time. Compare means for groups A and B.
Expected result and explanation
Valid n=18: mean 61.67, median 65, sample SD 31.11 minutes. A: n=9, mean 56.67; B: n=9, mean 66.67. B−A=10 minutes. Each group has one missing time. This descriptive difference is not evidence of causation or statistical significance.
Explore a real data dictionary
Learn to distinguish units, denominators and measures using the OpenAI Signals documentation.
Guide to reading a data dictionary ↗Continue with Python or R
Run the code on your computer. Python solutions require pandas and matplotlib; the R solution uses base R only.
Read descriptive statistics together
Work with the mini survey and compare your output with the expected result.
PythonCompare two group means
Work with the mini survey and compare your output with the expected result.
PythonCreate a bar chart from frequencies
Work with the mini survey and compare your output with the expected result.
RRead descriptive statistics together
Work with the mini survey and compare your output with the expected result.
RCompare two group means
Work with the mini survey and compare your output with the expected result.
RCreate a bar chart from frequencies
Work with the mini survey and compare your output with the expected result.