Open Evaluation of AI's Ability to Publish iOS App
Researchers advocate for "open-world" evaluations of AI. An AI agent has successfully published an iOS app, highlighting new methods for measuring AI capabilities.

What happened?
A new study published on arXiv introduces the concept of "open-world" evaluations to measure frontier AI capabilities. Unlike traditional benchmark-based evaluation, "open-world" focuses on long-term, complex, and real-world tasks that require qualitative analysis. CRUX (Collaborative Research for Updating AI eXpectations) is a project aiming to regularly conduct such evaluations. In an initial test, an AI agent successfully developed and published a simple iOS app to Apple's App Store with only one avoidable manual intervention.
Key facts
”Benchmark-based evaluation remains important for tracking frontier AI progress. But it can both overstate and understate deployed capability because it privileges tasks that can be precisely specified, automatically graded, easy to optimize for, and run with low budgets and short”
”We advocate for a complementary class of evaluations, which we term open-world evaluations: long-horizon, messy, real-world tasks assessed through small-sample qualitative analysis rather than benchmark-scale automation.”
”As a first instance, we task an AI agent with developing and publishing a simple iOS application to the Apple App Store. The agent completed the task with only a single avoidable manual intervention, suggesting”
Why it matters
Traditional benchmarks risk over- or underestimating actual AI capability by favouring tasks that are easy to specify, automate, and optimise. This new method of "open-world" evaluations strives for a more complete picture of practical AI capacity by exposing it to real-world challenges. An AI agent navigating the complex steps of developing and publishing an app demonstrates progress in autonomy and problem-solving.
Who is affected?
Researchers in AI evaluation are affected as new methods are tested and established. AI developers and companies licensing AI models also gain a new reference point for assessing model performance. Finally, the public and regulatory authorities can get a clearer picture of what current frontier AI systems can actually perform in complex scenarios.
What else you should know
The study underscores the need for evaluation methods that reflect the use of AI systems in real environments to better understand their limitations and potential. The CRUX project indicates a commitment to systematising this approach for future evaluations.
Quick answers about this story
Vad har hänt?
När hände det?
Varför spelar det roll?
Vilka bolag berörs?
The link opens in a new window and leads to the publisher's own site.
Källan har spårats automatiskt från utgivaren via Aheadlines signalkedja.
AI-verktyg i artikeln
Topics
Get similar news straight to your inbox
The reader's room
Send in a question or an addition. The newsroom reads everything before it's published and replies when relevant. No AI-generated text – just people.
Sign in to submit a comment or question.
Read the article through your role
- Decide whether this affects strategy over 6–12 months or is just noise.
- Discuss with leadership: do we own the right question or does ownership need to move?
- Ask: what risk are we taking by NOT acting on this this quarter?
Generated angle — not editorial analysis of "Open Evaluation of AI's Ability to Publish iOS App"