Statistical collaboration

Statistical collaboration and consulting

I work with researchers, departments and organizations whose question turns out to be a statistical one — data that are dependent, sparse, aggregated or machine-generated, and studies that have to be designed before they can be analyzed.

Areas

Where a collaboration usually starts

Forecasting and time series

Dependent, high-frequency and count-valued series: multi-frequency GARCH models, long-horizon volatility forecasting, and hybrid deep-learning forecasts of seasonal disease signals. Models with stated assumptions, and forecasts that are checked out of sample.

The statistics of AI systems

Statistical evaluation and benchmarking of systems whose outputs feed decisions — including large language models used as estimators, and model-generated labels used as the basis for inference. The question is how much uncertainty a number carries, and whether the decision built on it still holds.

Environmental and exposure data

Combining ground monitors, satellite retrievals and demographic records to measure how exposure differs between groups and places; low-cost sensor networks; and what can honestly be concluded once the data have been aggregated over areas that match no single variable.

Study design and measurement

Quantitative research design, assessment and outcome measurement, and an analysis plan agreed before the data are collected — including controlled studies where the measurement design decides whether an apparent effect is real.

Reproducible statistical computing

Analysis written as reproducible code in R and Python, browser-based computing through JupyterHub, and GPU and high-performance computing through the National Research Platform for work that outgrows a laptop.

Method

How the work runs

The question

What decision the number has to support, what the data can actually be asked, and what would count as an answer.

The method

The model, its assumptions and its limits, written down in plain language before the analysis rather than after it.

The analysis

Reproducible code, so every figure can be traced back to the data and the steps that produced it.

The result

A written account of what was found, what it does not establish, and what would change the conclusion — and, where the work is research, a manuscript with the collaborators named on it.

Record

Interdisciplinary work to date

Statistics is rarely the whole of a project. These are collaborations beyond my own research group that the work has already run through.

  • 2025 — 2026

    ELEVATE: Enhancing Learning Experiences Via AI Techniques

    Co-Principal Investigator on a California Education Learning Lab AI FAST Challenge project integrating generative AI and structured AI tutoring into mathematics, statistics and other university courses, with student research, faculty development and evaluation of AI-supported learning.

  • 2025 — present

    Engineering and computer science education

    Continuing work with Alberto C. Cruz, Maruti Mishra and J. J. Wang on structured AI tutoring, published through the Frontiers in Education Conference and AIxHEART, and accepted in the ASEE Computers in Education Journal.

  • 2026 — present

    Automated essay grading with Bakersfield College

    With Dr. Jonathan P. Brown, Professor of Mathematics at Bakersfield College: a resource-efficient local language model for an existing essay-grading application, developed on shared computing with Bakersfield College student researchers.

  • 2026

    Low-cost air-quality sensor networks

    Coauthor on an evaluation of fine-scale air-quality heterogeneity using a low-cost multipollutant sensor network in the Twin Cities, Minnesota, published in ACS ES&T Air with atmospheric and environmental health researchers.

  • 2026

    CSUB AWS Cloud & AI initiative

    Faculty collaborator on a campus cloud and AI project, contributing statistical evaluation and benchmarking, curriculum-integration planning and student-research applications.

  • 2026

    Data Science Exploration Challenge

    Co-organizer of a new high-school data science competition at the Lee Webb Math Field Day, including the dataset, the data dictionary, the participant materials and the judging framework.

Students are part of the work

Most of this I do with undergraduate researchers at CSU Bakersfield, and their names are on the resulting papers. Where a collaboration can carry a student, I try to make room for one.

Start with the data and the question

A short description of what you have measured and what you need to decide is enough for me to say whether I am the right person and what the work would involve. Colleagues, community organizations and student researchers are all welcome.