Measuring AI Ability to Complete Long Software Tasks

Thomas Kwa, Ben West, Joel Becker, Amy Deng, Katharyn Garcia, M. Hasin, Sami Jawhar, Megan Kinniment, Nate Rush, Sydney von Arx, Ryan Bloom, Thomas Broadley, Haoxing Du, Brian Goodrich, Nikola Jurkovic, Luke Harold Miles, Seraphina Nix, T. Lin, Neev Parikh, David Rein, L. Sato, Hjalmar Wijk, Daniel M. Ziegler, Elizabeth Barnes, Lawrence Chan

arXiv:2503.14499 · 2026-07-27 공개 · arXiv · PDF

long-context tool-use task-completion ai-benchmarks autonomy model-reliability software-tasks re-bench

Abstract

Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of AI systems in terms of human capabilities, we propose a new metric: 50%-task-completion time horizon. This is the time humans typically take to complete tasks that AI models can complete with 50% success rate. We first timed humans with relevant domain expertise on a combination of RE-Bench, HCAST, and 66 novel shorter tasks. On these tasks, current frontier AI models such as Claude 3.7 Sonnet have a 50% time horizon of around 50 minutes. Furthermore, frontier AI time horizon has been doubling approximately every seven months since 2019, though the trend may have accelerated in 2024. The increase in AI models'time horizons seems to be primarily driven by greater reliability and ability to adapt to mistakes, combined with better logical reasoning and tool use capabilities. We discuss the limitations of our results -- including their degree of external validity -- and the implications of increased autonomy for dangerous capabilities. If these results generalize to real-world software tasks, extrapolation of this trend predicts that within 5 years, AI systems will be capable of automating many software tasks that currently take humans a month.

한국어 요약

한 줄 요약

AI가 소프트웨어 작업을 완료하는 시간을 인간 기준으로 측정하는 새로운 지표를 제안하고, 2019년 이후 AI의 시간 범위가 7개월마다 2배씩 증가하고 있음을 밝힘.

핵심 기여도

핵심 아이디어

기존 AI 벤치마크는 대부분 인공적인 작업에 초점이 맞춰져 있으며, 인간 수준의 능력을 정량적으로 비교하기 어렵다는 한계를 지닌다. 본 연구는 AI가 작업을 완료하는 시간을 인간 기준으로 측정하는 **50%-task-completion time horizon**이라는 새로운 지표를 제안한다. 이 지표는 AI가 50% 성공률로 작업을 완료하는 데 걸리는 시간을 인간의 평균 작업 시간과 비교하여 AI의 실질적 능력을 평가한다. 예를 들어, **Claude 3.7 Sonnet** 모델은 50% 성공률로 작업을 완료하는 데 약 50분이 걸리며, 이는 인간의 작업 시간 기준이다. 연구는 **RE-Bench, HCAST, SWAA** 데이터셋을 사용하여 AI의 시간 범위가 2019년 이후 **7개월마다 2배씩 증가**하고 있음을 밝혔다. 이는 AI가 **논리적 추론 능력, 도구 사용 능력, 실수 대응 능력**이 향상되었기 때문으로 분석된다.

기술적 접근법

주요 결과

의의 및 한계

AI의 실질적 능력을 인간 기준으로 정량적으로 평가할 수 있는 새로운 지표를 제안함으로써, AI의 진보 추세를 보다 명확히 파악할 수 있게 되었다. 특히, **50% time horizon**은 AI가 복잡한 작업을 자율적으로 수행하는 능력을 측정하는 데 유용하며, **SWE-bench Verified** 데이터셋을 활용한 외부 유효성 검증을 통해 추세의 일반화 가능성을 뒷받침하였다. 그러나 연구는 **구조화되지 않은 작업**(messy tasks)에서 AI 성능이 낮아지는 한계를 지닌다. 또한, AI의 자율성 증가가 **위험한 능력**(CBRN 개발 등)을 동반할 수 있다는 점도 언급된다. 따라서 AI의 능력 평가와 함께 안전 장치 개발이 필수적이다.

실용적 활용

본 연구는 소프트웨어 엔지니어링, 연구 개발, 보안 분야에서 AI의 자율 작업 능력을 평가하는 데 활용될 수 있다. 예를 들어, **SWE-bench Verified**와 같은 데이터셋을 사용하면 AI가 실제 개발 작업을 얼마나 빠르게 수행할 수 있는지를 평가할 수 있다. 또한, AI의 시간 범위 증가 추세를 예측함으로써, **5년 이내에 인간이 한 달이 걸리는 작업을 AI가 자동화할 수 있을 것**이라는 전망도 가능하다. 이는 AI 개발자와 정책 수립자에게 중요한 참고 자료가 될 수 있다.