Attention: You are using an outdated browser, device or you do not have the latest version of JavaScript downloaded and so this website may not work as expected. Please download the latest software or switch device to avoid further issues.
In 2013, Michigan deployed an automated system to detect unemployment-insurance fraud. Over the next two years it wrongly accused roughly 40,000 residents at an error rate as high as 93%, and families lost homes, wages, and tax refunds while the cases dragged through the courts for nearly a decade. The system had - ostensibly - been tested. It had never been evaluated.
As AI adoption spreads across education, healthcare, workforce services, and social benefits, AI-driven programs increasingly shape the public services people depend on. Yet most AI evaluations still measure model performance, like accuracy, speed, and benchmarks, rather than the real-world impact on the people these systems are meant to serve.
Join the Data Foundation and Center for Evidence Capacity Senior Fellow Lauren Damme, Ph.D., as she makes the case for why public-sector AI deployments need the rigor social science has built over decades of evaluating real programs, and unpacks five critical gaps in current AI evaluation practices:
Drawing on the latest empirical research across fields, Lauren will also share emerging efforts to close these gaps. Attendees will leave with a clear picture of where social science methods should be integrated into AI evaluation, and where the field most needs social scientists’ skills.
Attend this webinar if you are an evaluator, policy researcher, social scientist, and anyone responsible for testing and deploying AI in public and social service contexts.