MFAI: Mathematical Foundations of Alignment in Generative Artificial Intelligence

NSF Award Search · 01002526DB NSF RESEARCH & RELATED ACTIVIT · $1,000,000 · view on nsf.gov ↗

Abstract

Generative large language models (LLMs) and generative diffusion models (GDMs) have become known for generating data that can have an astounding resemblance to human-generated content. Yet, the content generated by these models can introduce serious risks in specific applications. These models are known to replicate biases of their training data, produce unsafe outputs, and generate content that is misleading, false, and reprehensible. This project tackles these challenges within the general framework of alignment. Large pretrained models for image and language generation are available in the public domain but they are generic. It is of interest to most users to retrain these models to adapt them to their specific goals and principles. Our success will make it possible to better incorporate, among others, fairness, safety, reliability, robustness, and truthfulness requirements. This is a necessary development for tools that will be deeply integrated into the social and economic fabrics of our country. Our technical approach builds on three properties of alignment problems in generative AI: (P1) Alignment problems are reinforcement (RL) problems in which the value function is known. This makes alignment easier than generic RL because most of the typical complications of general RL problems are related to the learning of the value function. (P2) Alignment problems in generative language models and generative diffusion processes share the same structure. The objective is to a

Key facts

NSF award ID
2502489
Awardee
University of Pennsylvania (PA)
SAM.gov UEI
GM1XX56LEP58
PI
Alejandro R Ribeiro
Primary program
01002526DB NSF RESEARCH & RELATED ACTIVIT
All programs
Artificial Intelligence (AI), Machine Learning Theory
Estimated total
$1,000,000
Funds obligated
$1,000,000
Transaction type
Standard Grant
Period
10/01/2025 → 09/30/2028