Peter Henderson
Peter Henderson is an assistant professor at Princeton University with appointments in Computer Science and the School of Public and International Affairs. He holds a J.D. and Ph.D. from Stanford. His research focuses on making AI systems safer and more accountable, drawing on methods from machine learning, reinforcement learning, and legal theory. His work has influenced policy at federal agencies including NIST and the DOJ, and has been cited in proceedings before the U.S. and Arizona Supreme Courts. His lab also partners with government agencies and public institutions to deploy AI tools that work toward the public good.
AI2050 Project
AI systems like ChatGPT are controlled by written rules — natural language documents that tell the model what it should and shouldn’t do. But these rules are often vague, contradictory, or produce unintended behavior. Henderson’s project builds tools to test these rules before they are deployed. Drawing on techniques from legal interpretation and machine learning, the project’s framework detects ambiguities in AI rule documents, predicts how rule changes will affect behavior, and identifies conflicts between rules and other parts of model training.
Assistant Professor, Princeton University
Hard ProblemAlignment