Why Measuring AI Is Hard
Led by: Arvind Narayanan
AI systems routinely achieve impressive scores on standardized tests and benchmarks, yet those results often tell us little about whether a tool will succeed in a real government setting.
This session explores the limitations of common AI benchmarks, the differences between predictive and generative AI systems, and why performance in laboratory tests often fails to predict success in real-world settings.
By the end of this workshop, participants will be able to:
- Explain why standard AI benchmarks and test scores may not predict how well an AI system will perform on real government tasks and workflows.
- Distinguish key evaluation challenges for predictive and generative AI systems and identify why each requires different approaches to measuring performance.
- Identify the factors agencies should consider when defining meaningful measures of AI effectiveness, reliability, and public value in real-world government settings.
This workshop is part of an InnovateUS Series called : Practical Approaches to Evaluating AI for Public Benefit
Click here to view all workshops from this seriesManaging AI Risk Across the Organization
December 16, 2026
Scaling Government Innovation
December 10, 2026
Evaluating AI Tools and Vendors
December 9, 2026
Building High-Performing Product Teams
December 4, 2026
When AI Creates New Risks
December 2, 2026
Communicating Across Language and Cultural Barriers: Equity, Access, and Multilingual Outreach
November 19, 2026