Research

Several areas, one working method.

Artificial intelligence in mathematics and statistics education — the largest strand of the work — alongside time series, stochastic modeling and financial econometrics; statistics for artificial intelligence; and environmental statistics and data science.

They share a working method more than a subject. Each one starts from data that are dependent, aggregated, or produced by a process I do not control, and asks what can honestly be estimated from them and how far the answer can be trusted. Across them all I work with students through the full research process, from data analysis and statistical computing to scientific writing and publication, and I collaborate with researchers in other disciplines on both methodological and applied problems.

Research area · principal focus

Artificial intelligence in mathematics and statistics education

The largest strand of current work — and where the external funding sits

Students now carry a capable answer machine in their pocket. The open question is whether it helps them learn — and answering it properly is a measurement problem before it is a technology problem.

I design AI tutors that are constrained to behave like tutors, asking before telling, staying inside the course material, and declining to hand over finished solutions, and then run controlled studies to find out what actually changes. Some of what first looks like a learning gain turns out to be an artifact of how the student work was produced, so the measurement design matters as much as the tool.

This strand carries ELEVATE ($150,000, California Education Learning Lab AI FAST Challenge, Co-Principal Investigator), the AI learning companion in the Title III CATALYST proposal, an accepted paper in the ASEE Computers in Education Journal, manuscripts under review at the Journal of Statistics and Data Science Education and Computers & Education: Artificial Intelligence, peer-reviewed conference papers and presentations at JSM, AERA, FIE and AIxHEART, and a sustained programme of faculty workshops. It also sits behind the ADSA AI Tutoring Study Working Group and the proposed general-education course Inside AI: The Mathematics of Large Language Models.

Published and accepted

Structured AI-Tutoring for Computer Architecture Courses

Accepted

Cruz, A. C., Yatawara, A., Mishra, M., & Wang, J. J.

ASEE Computers in Education Journal, 2026 FIE Special Issue. Accepted for publication, June 23, 2026.

WIP: Structured AI Tutoring in Engineering Education

Published

Cruz, A. C., Yatawara, A., Mishra, M., & Wang, J. J.

Proceedings of the 2025 IEEE Frontiers in Education Conference (FIE), 1–5.

Computer-Aided Instruction for K-12 Teachers: A Cognitive Apprenticeship Approach to LLM Integration

Published

Cruz, A. C., Yatawara, A., Mishra, M., & Wang, J. J.

2025 Artificial Intelligence x Humanities, Education, and Art (AIxHEART), 13–16.

Under review

From Answer Generators to Thinking Partners: Scaffolded LLM Tutors in Statistics Education

Yatawara, A., Wang, J., Cruz, A. C., & Mishra, M.

Journal of Statistics and Data Science Education.

From an Apparent Gain to a Composition Artefact: A Three-Condition Study of Generative AI in an Undergraduate Proof-Based Mathematics Course

Under editorial assessment

Yatawara, A., Wang, J., Cruz, A. C., & Mishra, M.

Computers & Education: Artificial Intelligence. Submitted.

Projects

  • 2025 – 2026

    ELEVATE: Enhancing Learning Experiences Via AI Techniques

    California Education Learning Lab, AI FAST Challenge. Interdisciplinary project integrating generative AI and structured AI tutoring into mathematics, statistics, and other university courses, including student research, faculty development, and evaluation of AI-supported learning.

  • 2026 –

    Local, task-specific language model for automated essay grading

    With Dr. Jonathan P. Brown, Professor of Mathematics, Bakersfield College. Developing a resource-efficient local language model for an existing automated essay-grading application, using shared computing resources and with Bakersfield College student researchers; investigating neural-network, small language model, and model-distillation approaches.

  • 2026 –

    ADSA AI Tutoring Study Working Group

    Alliance for Data Science and AI. A multi-institution working group examining AI tutoring and AI-supported learning in statistics and data science education.

Research area

Time series, stochastic modeling and financial econometrics

Financial and economic data arrive as long streams of numbers whose volatility, meaning how violently they move, changes over time and clusters in bursts. I build statistical models that describe that behavior and forecast it: how long a shock keeps mattering, whether good and bad news leave different marks, how information measured every few minutes relates to information measured monthly, and how to handle series that count events rather than measure quantities. The aim is models with stated assumptions, proven properties, and forecasts that hold up out of sample.

Published

Does trading volume improve long-term volatility forecasts? Evidence from the MF2-GARCH framework

Journal of Forecasting (2026).

The Multiplicative Factor Multi-Frequency Exponential GARCH ((MF)2-EGARCH)

With V. A. Samaranayake

Proceedings of the Joint Statistical Meetings, Business and Economic Statistics Section (2022).

The Asymmetric Hyperbolic Generalized Autoregressive Conditional Heteroscedastic (A-HYGARCH) Model

With V. A. Samaranayake

Proceedings of the Joint Statistical Meetings, Business and Economic Statistics Section (2021).

Under review

When Variance Arrives: Intraday Timing Information in Realized-Variance Forecasting

Under review

Journal of Econometrics.

The Timescale Structure of Volatility Memory

Under consideration

Journal of Financial Econometrics.

The Shape of Volatility Memory

Under consideration

Quantitative Finance.

A Multiplicative-Factor Multi-Frequency Conditional-Mean Model for Count Time Series

Journal of Time Series Analysis.

Quasi-Maximum Likelihood Estimation of EGARCH(p,q): Continuous Invertibility, Consistency, and Asymptotic Normality under Verifiable Conditions

Under editorial assessment

Econometric Theory.

Do Volatility Components Explain the Heavy Tails of GARCH Residuals?

Under consideration

International Review of Economics & Finance.

In progress

(MF)2-GARCH-A: Volatility Modeling with a Sign-Sensitive Long-Run Component

With V. A. Samaranayake

Multiple-Regime Hyperbolic GARCH (MR-HYGARCH)

With V. A. Samaranayake

Improving Influenza Forecasting: A GRU-CNN Hybrid Model with Yeo-Johnson Scaling

With Noah Gallego and Isuru Ratnayake

Related work was presented at the Joint Statistical Meetings, Nashville, 2025.

This line of work began with my doctoral dissertation, The Multiplicative Factor Multi-Frequency Exponential GARCH ((MF)2-EGARCH) (Missouri University of Science and Technology, 2023), written under Dr. V. A. Samaranayake.

Interactive

How long does a shock keep mattering?

One question runs underneath most of the time-series work: when volatility jumps, how does the effect fade? Competing models differ less in what they measure than in the shape they give that fading.

This simulator draws one flexible family of shapes — the stretched-exponential memory kernel w(τ) = exp(-(τ/T)^α) — against the pure exponential it generalizes at α = 1. Drag the two parameters. The half-life, the point at which half of a shock has gone, barely moves. The time it takes for the last one percent to disappear runs away.

That distance between the half-life and the tail is the territory of the manuscripts currently under consideration, The Shape of Volatility Memory and The Timescale Structure of Volatility Memory.

Research area

Statistics for artificial intelligence

Large language models now produce numbers, including forecasts, estimates, labels and ratings, that people use to make decisions. Those numbers are not measurements, and they arrive with no error bars. My work asks what statistical guarantees survive when a model, rather than an instrument or a survey, supplies the input: when a poor forecast can still support a good decision, how much hand-checked data is needed before model-generated labels are safe to do inference on, and how to report the remaining uncertainty honestly.

Published

Good portfolios from bad forecasts: The anatomy of LLM volatility estimates

Finance Research Letters, 109, 110572 (2026).

Under review

Empirical Asset Pricing via Large Language Models

The Journal of Finance and Data Science.

In progress

Prediction-Powered Inference with Model-Assisted Validation Labels

Projects

  • Summer 2025

    Locally hosted language models in higher education

    Chevron Summer Undergraduate Research Experience. Undergraduate research on retrieval-augmented generation, model fine-tuning, quantization, GPU-based deployment, and comparison with commercial cloud-based AI systems.

  • Fall 2026 –

    Cryptography and LLM text watermarking

    Undergraduate mathematics research with Chase Davis and Brian Martinez; currently at the initial research-development stage.

A cross-CSU NSF proposal in preparation, listed under Grants below, extends this area to the mathematical and statistical foundations of artificial intelligence, including reliability, robustness, uncertainty, interpretability, computational efficiency, and responsible AI.

Research area

Environmental statistics and data science

Air quality is not the same across a city, a county, or an income bracket. Measuring the difference is harder than it sounds: ground monitors are sparse and unevenly placed, satellite estimates cover everywhere but less precisely, and demographic data come averaged over areas that match neither. I develop statistical methods that combine these sources, quantify exposure gaps between groups over long periods, and state plainly what can and cannot be concluded once the data have been aggregated.

Published

Air pollution exposure inequality across U.S. income and racial groups, 2000-2023: A hybrid monitor-satellite analysis with two new shape diagnostics

Regpala, T., Palafox, E., & Yatawara, A.

Atmospheric Environment: X, 31, 100492 (2026).

Income-based exposure disparities across California counties, 2000 to 2023: A generalizable statistical framework

Ko, K., Rodriguez, C., & Yatawara, A.

Environmental Research: Health, 4(1), 011002 (2026).

Evaluating fine-scale air-quality heterogeneity using a low-cost multipollutant sensor network in Twin Cities, Minnesota

Abhayaratne, V., Hao, W., Ye, C., Yatawara, A., Hopke, P. K., Li, J., & Wang, Y.

ACS ES&T Air, 3(4), 1057–1068 (2026).

Tom Regpala and Eric Palafox are student co-first authors on the Atmospheric Environment: X paper. Kayla Ko is student first author and Christian Rodriguez a student coauthor on the Environmental Research: Health paper. They are all CSUB undergraduate researchers I supervise.

Under review

Lagged Causal Effects from Spatially Aggregated Data: A Change-of-Support Distributed-Lag Framework with Design Diagnostics

Under consideration

Spatial Statistics.

Projects

  • 2025 – 2026

    Air Pollution Disparities in California, 2000-2025: Examining Environmental Inequities Across Income, Poverty, and Minority Populations

    CSUB Environmental Studies CES Mini-Grant, as Principal Investigator. Supports undergraduate research in environmental statistics, air-pollution exposure inequality, statistical computing, and large-scale demographic and environmental data integration; student research from this project contributed to multiple peer-reviewed publications in environmental statistics and environmental health.

Behind the work

Research computing and infrastructure

Most of this work needs machines, and at a campus this size the computing has to be built as deliberately as the models. Since 2025 I have worked on that side alongside the research itself, so that undergraduates can run real analyses without first assembling an environment of their own.

  • CSUB Hub for Statistics, Applied AI and Data Science Innovation

    Founded 2025; I serve as Faculty Lead. The Hub supports statistics, applied AI, data science, student research, research computing, and cross-disciplinary collaboration. I established a physical home for it and developed computing infrastructure supporting undergraduate and faculty research.

  • National Research Platform and Nautilus

    Contributed to research-computing infrastructure supporting statistics, data science, and AI research at CSUB, including connection to the National Research Platform and Nautilus ecosystem for GPU and high-performance computing.

  • AWS Cloud and AI initiative

    Participated in CSUB's AWS Cloud and AI initiative, contributing statistical evaluation and benchmarking, curriculum-integration ideas, and student-training plans to a project supported by substantial AWS cloud-computing credits.

  • CSUB JupyterHub

    Browser-based statistical computing, built into the shared instructional system for MATH 2200 and MATH 1209 and used again in MATH 4230, so students work in R and Python without a local installation.

  • Shared computing with Bakersfield College

    Shared resources support the local essay-grading language model collaboration and the Bakersfield College student researchers working on it.

  • CSUB AI and Data Science Research Group

    Founded 2026. A cross-departmental faculty research group connecting statistics, computer science, artificial intelligence, institutional research, and related areas, with research meetings supporting collaborative projects, student involvement, and grant opportunities.

  • Statistics, Applied AI and Data Science Colloquium Series

    Founded 2025. Speakers so far have been Thomas A. DeFanti on CENIC AIR and the National Research Platform (November 2025), Bhash Abeysinghe of the American Institutes for Research on AI, NLP, and agentic systems (April 2026), and Emma Yates and Anna Winter of NASA Ames / BAER on Ozone Where We Live, community-based air-quality monitoring in California (September 2026).

Funding

Funded and pending grants

Funded

  • 2025 – 2026

    California Education Learning Lab, AI FAST Challenge

    Co-Principal Investigator·$150,000

    ELEVATE: Enhancing Learning Experiences Via AI Techniques. Interdisciplinary project integrating generative AI and structured AI tutoring into mathematics, statistics, and other university courses; included student research, faculty development, and evaluation of AI-supported learning.

  • 2025 – 2026

    CSUB Environmental Studies, CES Mini-Grant

    Principal Investigator

    Air Pollution Disparities in California, 2000-2025: Examining Environmental Inequities Across Income, Poverty, and Minority Populations. Funded environmental statistics project supporting undergraduate research, large-scale air-quality and demographic data analysis, and peer-reviewed scholarship.

  • 2026 – 2027

    CSUB Center for Entrepreneurship and Innovation, Faculty R&D Fellowship

    Faculty R&D Fellow / Project Lead·$7,000

    KernelStats: AI-Powered Statistical Analysis Platform. Competitive R&D fellowship supporting development, testing, intellectual-property planning, and commercialization of an AI-assisted statistical analysis platform. KernelStats is in private development.

  • 2026

    Amazon Web Services, Cloud and AI Research Credit Award

    Faculty Collaborator / Project Team Member·$200,000 in AWS cloud credits

    Credits awarded to the collaborative CSUB project. Contributed statistical evaluation and benchmarking expertise, curricular-integration planning, and student-research applications.

  • Summer 2025

    Chevron Summer Undergraduate Research Experience (SURE)

    Faculty Research Mentor·$7,000 faculty mentor award

    A Customizable AI Pipeline for Academic Excellence: Locally Hosted Language Models in Higher Education. Supervised undergraduate research on locally hosted large language models, retrieval-augmented generation, model fine-tuning, quantization, GPU-based deployment, and development of institution-specific AI applications.

Submitted, decision pending

  • 2026

    U.S. Department of Education, Title III Strengthening Institutions Program (SIP)

    AI Learning Companion Tool Developer (0.33 FTE)·Key Personnel

    CATALYST: Cultivating Adaptive Teaching And Learning with AI-Supported Technologies. Revised institutional proposal submitted June 2026; decision pending. Responsible for developing and maintaining the Adaptive AI Learning Companion, a scaffolded AI tutoring system for STEM gateway courses, building on prior mathematics and statistics AI-tutoring research and the ELEVATE project.

  • 2026

    U.S. Department of Education, Supporting Effective Educator Development (SEED)

    Proposal Team Contributor·Cruz, A. C. (PI), Yatawara, A., et al.

    FY 2026 SEED Competition, Assistance Listing 84.423A. Contributed quantitative research design, assessment, mathematics and statistics learning materials, analysis of implementation and participant outcomes, and proposal and budget development.

In preparation

  • 2026

    National Science Foundation, Mathematical Foundations of Artificial Intelligence (MFAI), NSF 24-569

    Co-Principal Investigator·PI: Jeremy Woods

    Cross-CSU interdisciplinary proposal investigating mathematical and statistical foundations of artificial intelligence, including reliability, robustness, uncertainty, interpretability, computational efficiency, and responsible AI. Planned submission: October 2026.

In private development

KernelStats

An all-in-one, AI-powered statistical analysis platform, currently in private development.

The work is supported by a Faculty R&D Fellowship from the CSUB Center for Entrepreneurship and Innovation (2026–2027), which funds development, testing, and intellectual-property and commercialization planning.

There is nothing more to share publicly yet.