BEGIN:VCALENDAR
VERSION:2.0
PRODID:icalendar-ruby
CALSCALE:GREGORIAN
X-WR-CALNAME:Statistics and Data Science Seminar
X-WR-TIMEZONE:Eastern Time (US & Canada)
BEGIN:VEVENT
DTSTAMP:20260721T100924Z
UID:tag:localist.com\,2008:EventInstance_50587323478594
DTSTART:20251010T150000Z
DTEND:20251010T160000Z
DESCRIPTION:Speaker: Weijie Su (University of Pennsylvania)\n\nTitle: Do La
 rge Language Models (Really) Need Statistical Foundations?\n\nAbstract: In
  this talk\, we advocate for developing statistical foundations for large 
 language models (LLMs). We begin by examining two key characteristics that
  necessitate statistical perspectives for LLMs: (1) the probabilistic\, au
 toregressive nature of next-token prediction\, and (2) the inherent comple
 xity and black box nature of Transformer architectures. To demonstrate how
  statistical insights can advance LLM development and applications\, we pr
 esent two examples. First\, we demonstrate statistical inconsistencies and
  biases arising from the current approach to aligning LLMs with human pref
 erence. We propose a regularization term for aligning LLMs that is both ne
 cessary and sufficient to ensure consistent alignment. Second\, we introdu
 ce a novel statistical framework for analyzing the efficacy of watermarkin
 g schemes\, with a focus on a watermarking scheme developed by OpenAI for 
 which we derive optimal detection rules that outperform existing ones. Tim
 e permitting\, we will explore how statistical principles can inform rigor
 ous evaluation for LLMs. Collectively\, these findings demonstrate how sta
 tistical insights can effectively address several pressing challenges emer
 ging from LLMs.\n\nBiography: Weijie Su is an Associate Professor in the W
 harton Statistics and Data Science Department at the University of Pennsyl
 vania. He is a co-director of Penn Research in Machine Learning (PRiML) Ce
 nter. His research interests span statistical foundations of generative AI
 \, privacy-preserving machine learning\, high-dimensional statistics\, and
  optimization. He serves as an associate editor of Journal of the American
  Statistical Association\, Journal of Machine Learning Research\, Annals o
 f Applied Statistics\, Harvard Data Science Review\, Foundations and Trend
 s in Statistics\, Operations Research\, and Journal of the Operations Rese
 arch Society of China\, and he is currently guest editing a special issue 
 on Statistics for Large Language Models and Large Language Models for Stat
 istics in Stat. His work has been recognized with several awards\, such as
  the Stanford Anderson Dissertation Award\, NSF CAREER Award\, Sloan Resea
 rch Fellowship\, IMS Peter Hall Prize\, SIAM Early Career Prize in Data Sc
 ience\, ASA Noether Early Career Award\, ICBS Frontiers of Science Award i
 n Mathematics\, IMS Medallion Lectureship\, and Outstanding Young Talent A
 ward in the 2025 China Annual Review of Mathematics. He is a Fellow of the
  IMS.
GEO:42.362019;-71.087844
LOCATION:Building E18\, 304
SUMMARY:Statistics and Data Science Seminar
URL;VALUE=URI:https://calendar.mit.edu/event/stochastics-and-statistics-sem
 inar6
END:VEVENT
END:VCALENDAR
