Skip to content
AI DEEP 2 sources· 3 min· cluster 1· updated 22:01 UTC

Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?

The authors introduce BOTTLED, a benchmark for testing whether language-model agents can turn repeated tasks into cheaper reusable solutions.

TL;DR

  1. BOTTLED evaluates agents asked to solve an unlabeled workload under fixed time, compute and API budgets.
  2. Across ten models and three tasks, 48 of 60 runs scored below the lower bound of their model’s zero-shot 95% confidence interval.
  3. The paper reports large cost savings in one classification task, but results varied substantially across runs and tasks.

The paper calls the process of converting general model capability into a reusable, lower-cost task solution “bottling.” In its BOTTLED benchmark, agents receive an unlabeled workload and must complete it within fixed time, compute and language-model API budgets. [arXiv; Hugging Face Papers.] [1] [2]

Across ten models and three tasks, 48 of 60 bottling runs fell below the lower bound of the model’s zero-shot 95% confidence interval. The authors also report that one Opus 5 product-relevance task retained about 82% of zero-shot macro-F1 at roughly 657 times lower reported cost; that result is task-specific. [arXiv; Hugging Face Papers.] [1] [2]

Why it matters

The benchmark highlights a deployment question that raw task accuracy does not answer: whether an agent can reliably trade up-front work for lower cost over repeated tasks. Its varied outcomes caution against assuming that strong zero-shot performance predicts effective automation.

Editor's note

New benchmark preprint. Cost and quality findings are reported by the paper authors and vary by task. TODO-review: confirm Korean term for “bottling.”

Type to search

↑↓ navigate ↵ open esc close