Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?
The authors introduce BOTTLED, a benchmark for testing whether language-model agents can turn repeated tasks into cheaper reusable solutions.
TL;DR
- BOTTLED evaluates agents asked to solve an unlabeled workload under fixed time, compute and API budgets.
- Across ten models and three tasks, 48 of 60 runs scored below the lower bound of their model’s zero-shot 95% confidence interval.
- The paper reports large cost savings in one classification task, but results varied substantially across runs and tasks.
The paper calls the process of converting general model capability into a reusable, lower-cost task solution “bottling.” In its BOTTLED benchmark, agents receive an unlabeled workload and must complete it within fixed time, compute and language-model API budgets. [arXiv; Hugging Face Papers.] [1] [2]
Across ten models and three tasks, 48 of 60 bottling runs fell below the lower bound of the model’s zero-shot 95% confidence interval. The authors also report that one Opus 5 product-relevance task retained about 82% of zero-shot macro-F1 at roughly 657 times lower reported cost; that result is task-specific. [arXiv; Hugging Face Papers.] [1] [2]
Why it matters
The benchmark highlights a deployment question that raw task accuracy does not answer: whether an agent can reliably trade up-front work for lower cost over repeated tasks. Its varied outcomes caution against assuming that strong zero-shot performance predicts effective automation.
Editor's note
New benchmark preprint. Cost and quality findings are reported by the paper authors and vary by task. TODO-review: confirm Korean term for “bottling.”