GPTProto

QwenGyre: An Elastic Reinforcement Learning Framework for Training xLong-Horizon Agents

HuggingFace Daily Papers(社区热门论文)·Sep 27, 2026, 8:00 AM·Qwen / Alibaba

View PDF HTML (experimental)

Abstract:Large language model (LLM) agents increasingly undertake extreme-long (xlong) horizon tasks, where a single execution can span hours, hundreds of model--environment interactions, and nearly 1M tokens per rollout. Applying online reinforcement learning (RL) to such executions poses two fundamental challenges: (1) severe execution variance and prolonged rollout delays cause massive GPU idling; and (2) complex non-linear branching generates massive trajectory redundancy, crippling training efficiency. To address these, we presents QwenGyre, an end-to-end framework for xlong-horizon online RL. QwenGyre elastically reallocates GPUs between rollout and training without interrupting live executions, while its trajectory processor reconstructs branching histories, scores partial progress, and deduplicates redundant paths to bound training costs. Scaled to our flagship model, Qwen~3.8 2.4T, with 700K tokens per rollout, QwenGyre yields a 6.0% absolute gain on NL2RepoBench (52.5% $\to$ 58.5%) in 48 steps. Across our evaluations on diverse domains of training datasets, QwenGyre delivers up to $1.85\times$ and $1.78\times$ speedups over Colocate and Async, respectively.
Subjects: Machine Learning (cs.LG); Distributed, Parallel, and Cluster Computing (cs.DC)
Cite as: arXiv:2609.33848 [cs.LG]
  (or arXiv:2609.33848v1 [cs.LG] for this version)
  https://doi.org/10.48550/arXiv.2609.33848

arXiv-issued DOI via DataCite (pending registration)

Submission history

From: Mouxiang Chen [view email]
[v1] Sun, 27 Sep 2026 19:02:03 UTC (535 KB)