Date: October 9, 2026
Speaker: Jamie Morgenstern, Assistant Professor, Paul G. Allen School of Computer Science & Engineering, University of Washington
Title: When Can We Trust a Prediction?

Abstract: When we ask whether a predictor can be trusted, we usually begin with a fixed notion of what it means to be right. In the first part of this talk, I will study this question through calibration: forecasts should agree with observed outcomes, not only overall but across many overlapping groups. Standard worst-case guarantees can deteriorate rapidly in high-dimensional settings and ignore structure that makes particular prediction problems easier. I will describe two recent results in online multicalibration. One gives an efficient algorithm that adapts to stochastic and piecewise-stationary data while retaining worst-case protection. The other shows that, for multiclass prediction, calibration guarantees can depend on the unknown intrinsic dimension of the outcome probabilities actually encountered rather than on the total number of classes.

I will then step back from the assumption that we already know how a prediction should be evaluated. Predictions and other candidate outputs are becoming extremely cheap to produce, but deciding whether they are any good can require expensive expertise, delayed outcomes, or measurements that only imperfectly capture what we care about. Those measurements may themselves depend on other models, datasets, benchmarks, and platforms. This is not merely a complication for calibration: it determines what being calibrated means. I will discuss data quality and AI entanglement as ways of understanding how evaluations are constructed, what dependencies they inherit, and what conclusions they can support. This talk draws on joint work with Zhiming Huang, Aaron Roth, and Claire Jie Zhang on calibration, and with Rachel Hong, Jevan Hutson, Sarah Huiyi Cen, and Tadayoshi Kohno on AI entanglement.