Papers
arxiv:2608.31111

Aspire: Can Models Self-Evolve from Vague Goals?

Published on Aug 31
· Submitted by
wuyuhao
on Sep 3
#3 Paper of the day
Authors:
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,
,

Abstract

ASPIRE introduces a benchmark for self-evolving LLM agents from vague natural-language goals, revealing challenges in goal interpretation, data selection, and stable weight-level improvement.

Many important forms of human learning begin with a vague goal, such as "become a better physicist" or "improve at research." Learners must interpret the goal, identify capability gaps, decide how to learn, and determine whether they have actually improved. In contrast, existing work on LLM self-evolution typically begins with tasks and evaluation metrics specified by humans, reducing self-evolution to optimizing an explicit objective rather than deciding what and how to learn. We introduce ASPIRE, a benchmark for vague-goal-driven self-evolution. ASPIRE provides only a natural-language capability goal while downstream evaluation tasks remain hidden. The agent must operationalize the goal by choosing data and update methods, constructing training and validation signals, and deciding when to evaluate. ASPIRE supports both model-weight and agent-harness evolution in a unified interactive environment and evaluates the resulting systems on a hidden, expert-authored set of 520 items spanning six goals. Our experiments show that vague goals redirect search effort toward goal interpretation. Current agents routinely complete training and harness-editing loops, but weight-level gains remain sparse and unstable, and the strongest evolved harness remains below the engineered Qwen-Agent reference. Agents often train on mismatched data and trust narrow self-evaluations, so local gains fail to transfer to hidden evaluation and continued search and training can erase earlier improvements.

Community

Paper submitter

We introduce ASPIRE, a benchmark that asks whether models can self-evolve from only a vague capability goal, without access to downstream evaluation tasks. Across 520 expert-authored items spanning six goals, agents can complete training and harness-editing loops, but their gains remain sparse and unstable because their chosen data and self-evaluations often fail to match the hidden target. Project page: https://self-developing-agents.github.io/

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.31111
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2608.31111 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2608.31111 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2608.31111 in a Space README.md to link it from this page.

Collections including this paper 1