Mini exercises

Explore a survey dataset.

Four tasks, one dataset: clean, summarize, visualize and interpret.

Entirely synthetic educational data • Beginner–Intermediate • About 25 minutes

Data and exercise files

Start with the raw data and attempt the tasks. The package also includes a data dictionary, clean reference data and Python/R solution scripts.

Each row is a person record. id: identifier; group: observational A/B group; minutes: daily learning time (minutes, blank/99 missing); transport: bus, walk or bike. The raw data contain intentional issues.

1 / 4

Clean the data

Find duplicate records, missing values and inconsistent category spellings. How many people and valid times remain?

Expected result and explanation

Remove one copy of the exact duplicate id=5 from 21 rows. Trim and lowercase transport. Blank and 99 are missing for minutes; do not replace them with zero. Result: 20 unique people, 18 valid times and 2 missing. The 99 rule belongs to this data dictionary, not every dataset.

2 / 4

Build a frequency table

Count transport categories in the clean data and calculate percentages. State your denominator.

Expected result and explanation

bus: 10 (50%), walk: 6 (30%), bike: 4 (20%). The denominator is 20 people; transport has no missing values. Missing times do not reduce this denominator to 18.

3 / 4

Plot and interpret

Create a bar chart of transport frequencies. Add a title, axis unit and data note.

Expected result and explanation

Bar heights should be 10, 6 and 4. The vertical axis counts people and starts at zero. Note: Synthetic educational data, n=20. Bus is the most frequent choice in these data; this does not establish population preferences.

4 / 4

Compare groups

Find the mean, median and sample standard deviation of time. Compare means for groups A and B.

Expected result and explanation

Valid n=18: mean 61.67, median 65, sample SD 31.11 minutes. A: n=9, mean 56.67; B: n=9, mean 66.67. B−A=10 minutes. Each group has one missing time. This descriptive difference is not evidence of causation or statistical significance.

Explore a real data dictionary

Learn to distinguish units, denominators and measures using the OpenAI Signals documentation.

Guide to reading a data dictionary ↗

Continue with Python or R

Run the code on your computer. Python solutions require pandas and matplotlib; the R solution uses base R only.