LLM Evaluation and Test-Time Scaling

Date:

Invited seminar on calibrated Bayesian evaluation of large language models and the formalization of test-time scaling as budgeted inference.