Try Before and After You Buy

Led by: Kathrin Frauscher & Patrick McLoughlin

How can agencies assess whether an AI tool is likely to work before committing significant resources? This session covers practical methods for testing AI systems prior to procurement or deployment, including baseline comparisons, pilot design, task-based evaluation, and the identification of acceptable and unacceptable risks. Also, performance can change as models evolve, workflows shift, staff adapt their practices, and systems encounter new contexts and use cases. This session explores how agencies can monitor performance over time, detect emerging problems, and understand when an initially successful implementation may require adjustment or reevaluation.

By the end of this workshop, participants will be able to:

  • Design practical pre-deployment evaluations that compare AI performance against current practices and establish meaningful baselines.
  • Identify acceptable and unacceptable risks through task-based testing and well-designed pilots before committing significant resources to AI adoption.
  • Develop approaches for monitoring AI performance over time and determining when changes in models, workflows, or use cases require adjustment, reevaluation, or retirement.
 
 
This workshop is part of an InnovateUS Series called : Practical Approaches to Evaluating AI for Public Benefit
Click here to view all workshops from this series
Kathrin Frauscher

Deputy Executive Director, Open Contracting Partnership

View bio
Patrick McLoughlin

Executive Director, Maryland Benefits, State of Maryland; Former MD State Chief Data Officer

View bio

Moderated By

Deborah Stine

Senior Fellow, Innovate US

View bio

Format: online

Date & Time: September 29, 2026, 2:00 PM ET

Duration: 60 mins

Register for free