OpenAI aims to create AI benchmarks that better reflect real-world use cases

Apr 10, 2025

OpenAI has introduced a new initiative called the "Pioneers Program" aimed at developing AI benchmarks tailored to specific industries. The company says the goal is to create evaluation methods that better reflect real-world use cases in areas such as law, finance, and healthcare—domains where existing benchmarks fall short. According to OpenAI, current AI benchmarks are often flawed. They tend to measure tasks that are difficult to interpret or overly susceptible to manipulation—criticisms that have also been directed at OpenAI itself. As reported previously, the company has faced scrutiny over its involvement in funding and promoting a prominent math evaluation dataset. In the coming months, OpenAI plans to collaborate with multiple companies to build domain-specific evaluation tools. These benchmarks will eventually be released publicly. The first cohort includes select startups focused on practical AI applications. Participating companies will also have the opportunity to work with OpenAI on improving model performance via reinforcement fine-tuning, a method the company recently introduced for customizing expert-level language models.

AI News Without the Hype – Curated by Humans

As a THE DECODER subscriber, you get ad-free reading, our weekly AI newsletter, the exclusive "AI Radar" Frontier Report 6× per year, access to comments, and our complete archive.

AI news without the hype
Curated by humans.

Over 20 percent launch discount.
Read without distractions – no Google ads.
Access to comments and community discussions.
Weekly AI newsletter.
6 times a year: “AI Radar” – deep dives on key AI topics.
Up to 25 % off on KI Pro online events.
Access to our full ten-year archive.
Get the latest AI news from The Decoder.

Subscribe to The Decoder

OpenAI aims to create AI benchmarks that better reflect real-world use cases

AI News Without the Hype – Curated by Humans

AI news without the hypeCurated by humans.

AI news without the hype
Curated by humans.