Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Every number in this book comes from a real dataset or a clearly labeled simulated (_sim) one — never from a made-up example. This appendix lists every dataset the chapters actually use: what it is, how many rows it has, which variables matter most, where it came from, and which chapter(s) put it to work. Full documentation for each file — every variable defined, its units, its missing-data pattern, and the exact download and checksum — lives in that dataset’s codebook alongside the data itself.


11. Quick reference

DatasetReal or simulatednKern-anchoredUsed in
kern_calenviroscreenReal151 tractsYesCh. 1, 7, 9, 11
firstday_survey_simSimulated150 studentsNoCh. 1
kern_airqualityReal4,376 monitor-daysYesCh. 2, 4, 5, 6, 12
kern_crops_simSimulated72 commodity-yearsYesCh. 3, 10, 11, 12, 13
nhanes_subsetReal10,000 peopleNoCh. 5, 6, 10
kern_energy_employmentReal10 yearsYesCh. 7, 13
fitness_tracker_simSimulated200 participantsNoCh. 8, 10

Row counts above are exact counts of data rows (header row excluded) in each committed CSV, matching the count recorded in that file’s codebook. Details for each dataset — description, key variables, source, and license — follow below in the same order.


22. Real data

2.12.1 Kern County environmental burden — kern_calenviroscreen

One row per census tract in Kern County, scored by California’s CalEnviroScreen 4.0 tool for pollution burden and population vulnerability. The CalEnviroScreen score (ces_score) combines a Pollution Burden component (air quality, traffic, pesticide use, and more) with a Population Characteristics component (health outcomes and socioeconomic factors); ces_percentile ranks that score among all California tracts, so a higher percentile means a more pollution-burdened tract. Bakersfield-area tracts rank among the state’s most burdened, which is why this file anchors the book’s very first data-driven look at Kern County.

Filedata/processed/kern_calenviroscreen.csv
n151 census tracts (40 variables)
Key variablestract, total_pop, ces_score, ces_percentile, pm25, poverty, asthma, hispanic_pct
SourceCalifornia OEHHA, CalEnviroScreen 4.0 Results (oehha.ca.gov/calenviroscreen)
LicenseState of California / OEHHA open government data — public, attribute OEHHA
Used inCh. 1 (data & study design), Ch. 7 (confidence intervals), Ch. 9 (inference for proportions), Ch. 11 (chi-square)

2.22.2 Kern County air quality — kern_airquality

Daily PM2.5 and ozone readings for 2023 from every EPA-reporting air-quality monitor in Kern County — 12 monitors, five of them in Bakersfield. Each row is one monitor’s reading for one pollutant on one day; the file mixes 1,554 PM2.5 rows and 2,822 ozone rows for 4,376 total. Winter temperature inversions trap PM2.5 in the southern San Joaquin Valley, producing the right-skewed pollution spikes this book uses to teach distribution shape, sampling distributions, and group comparisons.

Filedata/processed/kern_airquality.csv
n4,376 monitor-days (1,554 PM2.5 + 2,822 ozone; 15 variables)
Key variablesdate, site_name, pollutant, daily_mean, daily_max, aqi
SourceU.S. EPA AirData / Air Quality System (epa.gov/outdoor-air-quality-data)
LicenseU.S. Government work — public domain; attribution requested
Used inCh. 2 (numerical summaries), Ch. 4 (probability), Ch. 5 (Normal model), Ch. 6 (sampling distributions & CLT), Ch. 12 (ANOVA)

2.32.3 Kern County oil and jobs — kern_energy_employment

One row per calendar year (2015–2024), joining California crude-oil production (Kern produces roughly 70% of the state’s total, though EIA reports only a statewide figure) with Kern County’s labor force and the Bakersfield metro area’s oil-sector payroll employment. It is a small, real-data panel built specifically to give the “oilfield employment” hook this course’s grant narrative calls for.

Filedata/processed/kern_energy_employment.csv
n10 years, 2015–2024 (9 variables)
Key variablesyear, ca_crude_oil_prod_thsd_bbl, kern_unemp_rate, ces_mining_logging_thsd, mining_logging_share_pct
SourceU.S. Energy Information Administration (eia.gov) + U.S. Bureau of Labor Statistics LAUS/CES (bls.gov)
LicenseU.S. Government work — public domain; citation requested
Used inCh. 7 (confidence intervals), Ch. 13 (correlation & regression)

2.42.4 A national health snapshot — nhanes_subset

A 10,000-person teaching sample of demographics and body measurements from the CDC’s National Health and Nutrition Examination Survey (NHANES), 2009–2012 cycles, distributed through the CRAN NHANES R package. This is the book’s one non-Kern dataset — used where the teaching goal needs a large general sample (Normal-model fitting, CLT demonstrations with a big population, comparing means across sex or age).

Filedata/processed/nhanes_subset.csv
n10,000 people (14 variables)
Key variablesage, sex, race, height_cm, weight_kg, bmi, bp_sys, bp_dia
SourceCDC NHANES 2009–2012, via CRAN NHANES package (Pruim, 2015)
LicenseUnderlying CDC data: public domain. NHANES package: GPL (≥ 2)
Used inCh. 5 (Normal model), Ch. 6 (sampling distributions & CLT), Ch. 10 (inference for means)

33. Simulated classroom data (_sim)

These files are generated by a seeded, re-runnable script rather than downloaded. Each is built to have realistic structure — plausible distributions, a genuine treatment effect, honest missing data — while staying clearly and permanently labeled as simulated.

3.13.1 Opening-day class survey — firstday_survey_sim

A simulated “index card” survey of 150 students on the first day of a MATH 2200-like course at a commuter Hispanic-Serving Institution: major, class standing, commute time, work hours, sleep, study plans, and self-reported statistics anxiety. Mild, deliberately built-in relationships (more work hours associate with less sleep and less planned study time) make it useful for scatterplots, correlation, and missing-data practice — three of the optional fields carry realistic light missingness by design.

Filedata/processed/firstday_survey_sim.csv
n150 students (8 variables)
Key variablesmajor, year, commute_minutes, work_hours_week, stat_anxiety, sleep_hours, study_hours
SourceSimulated for this course; no external source
LicenseCC0-1.0 (synthetic, CSUB Department of Mathematics)
Used inCh. 1 (data & study design)

3.23.2 Kern County crop panel — kern_crops_sim

A simulated Kern-flavored agricultural panel: 8 commodities (almonds, pistachios, table and wine grapes, oranges, tangerines/mandarins, carrots, cotton) across 9 crop years (2015–2023), one row per commodity-year. Real USDA NASS Quick Stats data was the intended source, but the live NASS API requires a free, human-registered key that the autonomous build could not self-provision, so this file uses simulated values plausible for a top U.S. agricultural county — magnitudes only, not a claim about actual Kern production. Harvested acres, yield, production, and value are built to stay internally consistent (production = harvested_acres × yield_per_acre), which makes the file useful for teaching derived quantities as well as ordinary summaries.

Filedata/processed/kern_crops_sim.csv
n72 commodity-years: 8 commodities × 9 years, 2015–2023 (12 variables)
Key variablesyear, commodity, category, harvested_acres, yield_per_acre, production, value_usd
SourceSimulated (modeled on USDA NASS Quick Stats structure; live fetch blocked by a gated API key)
LicenseCC BY-SA 4.0 (synthetic data, CSUB Department of Mathematics); generator code MIT
Used inCh. 3 (categorical data & tables), Ch. 10 (inference for means), Ch. 11 (chi-square), Ch. 12 (ANOVA), Ch. 13 (correlation & regression)

3.33.3 Wearable fitness study — fitness_tracker_sim

A simulated randomized study: 200 participants (100 treatment, 100 control) in an 8-week activity-coaching trial, with four wearable-style daily-average measures. A modest, built-in treatment effect (more steps and active minutes, lower resting heart rate, slightly more sleep in the treatment group) makes this the book’s dataset of choice for two-sample inference — Welch’s t-test, a difference-of-means confidence interval, and effect size.

Filedata/processed/fitness_tracker_sim.csv
n200 participants: 100 treatment / 100 control (6 variables)
Key variablesgroup, steps, active_minutes, resting_hr, sleep_hours
SourceSimulated for this course; no external source
LicenseCC0-1.0 (synthetic, CSUB Department of Mathematics)
Used inCh. 8 (hypothesis testing logic), Ch. 10 (inference for means)