Statistics · Time series · Applied AI · Environmental data

Anjana Yatawara

Assistant Professor of Statistics · California State University, Bakersfield

Founder & Faculty Lead, CSUB Hub for Statistics, Applied AI & Data Science Innovation

My research is organized around complementary areas, led by artificial intelligence in mathematics and statistics education; alongside time series, stochastic modeling and financial econometrics; statistics for artificial intelligence; and environmental statistics and data science. Across these areas I work with students throughout the research process, from data analysis and statistical computing to scientific writing and publication, and I collaborate with researchers across disciplines on both methodological and applied problems.

Research

Lines of work, one question

AI in education · Time series · Statistics of AI · Environmental exposure
  • AI in mathematics & statistics education

    This is the largest single strand of my current work. I run classroom studies on what generative AI actually does to mathematical and statistical reasoning — with control conditions, real assessments, and results that sometimes contradict the first impression. The applied half is scaffolded tutoring: systems that ask the next question instead of handing over the answer, built and tested in my own courses.

    It is the strand carrying the funding and the external work: ELEVATE, a $150,000 California Education Learning Lab AI FAST Challenge project on which I am Co-Principal Investigator; a Title III proposal in which I develop the adaptive AI learning companion; papers under review at the Journal of Statistics and Data Science Education and Computers & Education: AI; and an invited JSM 2026 presentation in Boston, From Answer Generators to Thinking Partners.

  • Time series, volatility & financial econometrics

    I build statistical models for dependent and dynamic data: multi-frequency and high-frequency volatility, count-valued time series, and long-memory and asymmetric GARCH structures. Alongside the models I work on the estimation theory that makes them usable — invertibility, consistency, and asymptotic normality under conditions that can actually be checked.

  • Statistics for artificial intelligence

    Large language models now produce numbers that people act on. I study what those numbers are worth: how reliable an LLM's estimate is, how much uncertainty it carries, and whether a decision built on it still holds.

  • Environmental statistics & data science

    Exposure to air pollution is distributed unevenly, and measuring that unevenness well is a statistical problem. With undergraduate researchers I combine ground monitors, satellite retrievals, and demographic records to quantify how exposure differs across U.S. income and racial groups and across California counties.

Publications

Selected work

Recent papers, with a plain-English note on each.

Good portfolios from bad forecasts: The anatomy of LLM volatility estimates

Published

Yatawara, A. (2026) · sole author

Finance Research Letters, 109, 110572.

Takes apart what a large language model's volatility estimates actually contain, and why portfolios built on them can come out well even when the forecasts themselves do not.

Does trading volume improve long-term volatility forecasts? Evidence from the MF2-GARCH framework

Published

Yatawara, A. (2026) · sole author

Journal of Forecasting.

Asks whether trading volume carries information a multi-frequency volatility model does not already have, and tests it at the long horizon where that claim usually goes untested.

Air pollution exposure inequality across U.S. income and racial groups, 2000–2023: A hybrid monitor–satellite analysis with two new shape diagnostics

Published

Regpala, T., Palafox, E., & Yatawara, A. (2026) · with undergraduate co-first authors Tom Regpala and Eric Palafox

Atmospheric Environment: X, 31, 100492.

Combines ground monitors with satellite estimates to track how pollution exposure has differed across U.S. income and racial groups over 24 years, and adds two diagnostics for the shape of the gap rather than only its size.

Income-based exposure disparities across California counties, 2000 to 2023: A generalizable statistical framework

Published

Ko, K., Rodriguez, C., & Yatawara, A. (2026) · with undergraduate first author Kayla Ko

Environmental Research: Health, 4(1), 011002.

Builds a reusable way to measure income-based differences in pollution exposure and runs it across California counties, so the method transfers to other states and pollutants.

WIP: Structured AI Tutoring in Engineering Education

Conference paper

Cruz, A. C., Yatawara, A., Mishra, M., & Wang, J. J. (2025)

Proceedings of the 2025 IEEE Frontiers in Education Conference (FIE), 1–5.

An early report from a line of work on AI tutors built to scaffold a student's reasoning step by step instead of handing back finished answers.

Teaching & mentoring

Statistics as work students do

In the classroom

I teach statistics as work students do, not a subject they watch me perform. Every course I run puts a computer in front of the student early: R and Python, run through CSUB JupyterHub so no one is held back by an install or a license. The data is real and usually messy, because the judgment calls that matter in statistics only appear once the cleaning is your own problem.

I ask for reproducible work. An analysis that cannot be rerun is an anecdote; I would rather read a notebook that runs than a number that happens to be right.

Undergraduate research

Undergraduates have worked with me in sustained individual research supervision since 2024, and several are authors on peer-reviewed journal articles, some as first or co-first author. The work runs from environmental statistics and exposure inequality to machine learning, forecasting, local language models, and cryptography — and it comes with the rest of it: research writing, poster and talk preparation, doctoral applications, and the long conversations about what comes next.

Contact

Email is the most reliable way to reach me.

ORCID 0009-0007-8506-7763 · Google Scholar · LinkedIn