How Companies Are Using Old-School Software to Grade AI Agents
Mirrored from The Information — AI for archival readability. Support the source by reading on the original site.
Measuring the quality of an AI model’s work can be tricky, especially when it comes to AI agents that can take a broad range of actions to autonomously solve a task. As businesses turn to agents to automate longer and more important tasks, they are getting savvier about how to evaluate the performance of those agents.
“Enterprises are getting more sophisticated,” said Brad Kenstler, director of general agents at Scale AI, where he focuses on developing evaluations for Scale’s enterprise customers. “They’re looking to deploy agents into much more business-critical operations.”
One surprisingly important way that enterprises are improving evaluations is by grading agents using software programs, rather than relying on another AI to serve as a judge. Scale is discussing developing these so-called “programmatic verifiers” with all its customers, including Mayo Clinic, real-estate and finance firm Howard Hughes and educational publishing company Cengage.
More from The Information — AI
-
Musk Responds to The Information’s Report About SpaceX Setting Up a Turbine Blade Factory
Aug 30
-
Exclusive: SpaceX Lays Groundwork For Turbine-Blade Factory to Solve Data Center Power Crunch
Aug 29
-
The Hugging Face Hack’s Chilling Postmortem
Aug 29
-
Why ‘Tax Alpha’ Is Silicon Valley’s New Obsession
Aug 29
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.