Every number in this book comes from a real dataset or a clearly labeled
simulated (_sim) one — never from a made-up example. This appendix lists
every dataset the chapters actually use: what it is, how many rows it has,
which variables matter most, where it came from, and which chapter(s) put it
to work. Full documentation for each file — every variable defined, its
units, its missing-data pattern, and the exact download and checksum — lives
in that dataset’s codebook alongside the data itself.
11. Quick reference¶
| Dataset | Real or simulated | n | Kern-anchored | Used in |
|---|---|---|---|---|
kern_calenviroscreen | Real | 151 tracts | Yes | Ch. 1, 7, 9, 11 |
firstday_survey_sim | Simulated | 150 students | No | Ch. 1 |
kern_airquality | Real | 4,376 monitor-days | Yes | Ch. 2, 4, 5, 6, 12 |
kern_crops_sim | Simulated | 72 commodity-years | Yes | Ch. 3, 10, 11, 12, 13 |
nhanes_subset | Real | 10,000 people | No | Ch. 5, 6, 10 |
kern_energy_employment | Real | 10 years | Yes | Ch. 7, 13 |
fitness_tracker_sim | Simulated | 200 participants | No | Ch. 8, 10 |
Row counts above are exact counts of data rows (header row excluded) in each committed CSV, matching the count recorded in that file’s codebook. Details for each dataset — description, key variables, source, and license — follow below in the same order.
22. Real data¶
2.12.1 Kern County environmental burden — kern_calenviroscreen¶
One row per census tract in Kern County, scored by California’s
CalEnviroScreen 4.0 tool for pollution burden and population vulnerability.
The CalEnviroScreen score (ces_score) combines a Pollution Burden
component (air quality, traffic, pesticide use, and more) with a Population
Characteristics component (health outcomes and socioeconomic factors);
ces_percentile ranks that score among all California tracts, so a higher
percentile means a more pollution-burdened tract. Bakersfield-area tracts
rank among the state’s most burdened, which is why this file anchors the
book’s very first data-driven look at Kern County.
| File | data/processed/kern_calenviroscreen.csv |
| n | 151 census tracts (40 variables) |
| Key variables | tract, total_pop, ces_score, ces_percentile, pm25, poverty, asthma, hispanic_pct |
| Source | California OEHHA, CalEnviroScreen 4.0 Results (oehha |
| License | State of California / OEHHA open government data — public, attribute OEHHA |
| Used in | Ch. 1 (data & study design), Ch. 7 (confidence intervals), Ch. 9 (inference for proportions), Ch. 11 (chi-square) |
2.22.2 Kern County air quality — kern_airquality¶
Daily PM2.5 and ozone readings for 2023 from every EPA-reporting air-quality monitor in Kern County — 12 monitors, five of them in Bakersfield. Each row is one monitor’s reading for one pollutant on one day; the file mixes 1,554 PM2.5 rows and 2,822 ozone rows for 4,376 total. Winter temperature inversions trap PM2.5 in the southern San Joaquin Valley, producing the right-skewed pollution spikes this book uses to teach distribution shape, sampling distributions, and group comparisons.
| File | data/processed/kern_airquality.csv |
| n | 4,376 monitor-days (1,554 PM2.5 + 2,822 ozone; 15 variables) |
| Key variables | date, site_name, pollutant, daily_mean, daily_max, aqi |
| Source | U.S. EPA AirData / Air Quality System (epa |
| License | U.S. Government work — public domain; attribution requested |
| Used in | Ch. 2 (numerical summaries), Ch. 4 (probability), Ch. 5 (Normal model), Ch. 6 (sampling distributions & CLT), Ch. 12 (ANOVA) |
2.32.3 Kern County oil and jobs — kern_energy_employment¶
One row per calendar year (2015–2024), joining California crude-oil production (Kern produces roughly 70% of the state’s total, though EIA reports only a statewide figure) with Kern County’s labor force and the Bakersfield metro area’s oil-sector payroll employment. It is a small, real-data panel built specifically to give the “oilfield employment” hook this course’s grant narrative calls for.
| File | data/processed/kern_energy_employment.csv |
| n | 10 years, 2015–2024 (9 variables) |
| Key variables | year, ca_crude_oil_prod_thsd_bbl, kern_unemp_rate, ces_mining_logging_thsd, mining_logging_share_pct |
| Source | U.S. Energy Information Administration (eia.gov) + U.S. Bureau of Labor Statistics LAUS/CES (bls.gov) |
| License | U.S. Government work — public domain; citation requested |
| Used in | Ch. 7 (confidence intervals), Ch. 13 (correlation & regression) |
2.42.4 A national health snapshot — nhanes_subset¶
A 10,000-person teaching sample of demographics and body measurements from
the CDC’s National Health and Nutrition Examination Survey (NHANES),
2009–2012 cycles, distributed through the CRAN NHANES R package. This is
the book’s one non-Kern dataset — used where the teaching goal needs a
large general sample (Normal-model fitting, CLT demonstrations with a big
population, comparing means across sex or age).
| File | data/processed/nhanes_subset.csv |
| n | 10,000 people (14 variables) |
| Key variables | age, sex, race, height_cm, weight_kg, bmi, bp_sys, bp_dia |
| Source | CDC NHANES 2009–2012, via CRAN NHANES package (Pruim, 2015) |
| License | Underlying CDC data: public domain. NHANES package: GPL (≥ 2) |
| Used in | Ch. 5 (Normal model), Ch. 6 (sampling distributions & CLT), Ch. 10 (inference for means) |
33. Simulated classroom data (_sim)¶
These files are generated by a seeded, re-runnable script rather than downloaded. Each is built to have realistic structure — plausible distributions, a genuine treatment effect, honest missing data — while staying clearly and permanently labeled as simulated.
3.13.1 Opening-day class survey — firstday_survey_sim¶
A simulated “index card” survey of 150 students on the first day of a MATH 2200-like course at a commuter Hispanic-Serving Institution: major, class standing, commute time, work hours, sleep, study plans, and self-reported statistics anxiety. Mild, deliberately built-in relationships (more work hours associate with less sleep and less planned study time) make it useful for scatterplots, correlation, and missing-data practice — three of the optional fields carry realistic light missingness by design.
| File | data/processed/firstday_survey_sim.csv |
| n | 150 students (8 variables) |
| Key variables | major, year, commute_minutes, work_hours_week, stat_anxiety, sleep_hours, study_hours |
| Source | Simulated for this course; no external source |
| License | CC0-1.0 (synthetic, CSUB Department of Mathematics) |
| Used in | Ch. 1 (data & study design) |
3.23.2 Kern County crop panel — kern_crops_sim¶
A simulated Kern-flavored agricultural panel: 8 commodities (almonds,
pistachios, table and wine grapes, oranges, tangerines/mandarins, carrots,
cotton) across 9 crop years (2015–2023), one row per commodity-year. Real
USDA NASS Quick Stats data was the intended source, but the live NASS API
requires a free, human-registered key that the autonomous build could not
self-provision, so this file uses simulated values plausible for a top U.S.
agricultural county — magnitudes only, not a claim about actual Kern
production. Harvested acres, yield, production, and value are built to stay
internally consistent (production = harvested_acres × yield_per_acre),
which makes the file useful for teaching derived quantities as well as
ordinary summaries.
| File | data/processed/kern_crops_sim.csv |
| n | 72 commodity-years: 8 commodities × 9 years, 2015–2023 (12 variables) |
| Key variables | year, commodity, category, harvested_acres, yield_per_acre, production, value_usd |
| Source | Simulated (modeled on USDA NASS Quick Stats structure; live fetch blocked by a gated API key) |
| License | CC BY-SA 4.0 (synthetic data, CSUB Department of Mathematics); generator code MIT |
| Used in | Ch. 3 (categorical data & tables), Ch. 10 (inference for means), Ch. 11 (chi-square), Ch. 12 (ANOVA), Ch. 13 (correlation & regression) |
3.33.3 Wearable fitness study — fitness_tracker_sim¶
A simulated randomized study: 200 participants (100 treatment, 100 control) in an 8-week activity-coaching trial, with four wearable-style daily-average measures. A modest, built-in treatment effect (more steps and active minutes, lower resting heart rate, slightly more sleep in the treatment group) makes this the book’s dataset of choice for two-sample inference — Welch’s t-test, a difference-of-means confidence interval, and effect size.
| File | data/processed/fitness_tracker_sim.csv |
| n | 200 participants: 100 treatment / 100 control (6 variables) |
| Key variables | group, steps, active_minutes, resting_hr, sleep_hours |
| Source | Simulated for this course; no external source |
| License | CC0-1.0 (synthetic, CSUB Department of Mathematics) |
| Used in | Ch. 8 (hypothesis testing logic), Ch. 10 (inference for means) |