‹ 返回 2026-06-16

LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling

LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling

▲ 3 北京航空航天大学、IQuest Research、Langboat、Renming University of China · 大语言模型、代码生成、推理优化、测试时计算 提交者 北京航空航天大学、IQuest Research、Langboat、Renming University of China 2026-06-16

Jian Yang, Shawn Guo, Wei Zhang, Tianyu Zheng, Yaxin Du, Haau-Sing Li, Jiajun Wu, Yue Song, Yan Xing, Qingsong Cai, Zelong Huang, Chuan Hao, Ran Tao, Xianglong Liu, Wayne Xin Zhao, Mingjie Tang, Weifeng Lv, Ming Zhou, Bryan Dai

摘要

LoopCoder-v2 是一款基于并行循环Transformer(PLT)架构的代码大语言模型。该模型通过引入跨循环位置偏移(CLP)和共享KV门控滑动窗口注意力(G-SWA)机制,实现了测试时计算的高效扩展,有效避免了标准循环Transformer的序列延迟和KV缓存内存膨胀问题。研究表明,在7B参数规模下,单次额外循环(R=2)能在代码生成、推理及智能体任务中带来显著的性能提升,而更多循环则因固定的位置偏移成本导致性能下降。

English Abstract

Looped Transformers scale latent computation by repeatedly applying shared blocks, but sequential looping increases latency and KV-cache memory with the loop count. Parallel loop Transformers (PLT) alleviate this cost through cross-loop position offsets (CLP) and shared-KV gated sliding-window attention, making loop count a practical design choice. We therefore study PLT loop-count selection through a gain--cost view: an extra loop may refine representations, but CLP also introduces a positional mismatch at each loop boundary. We instantiate this study by training LoopCoder-v2, a family of 7B PLT coders with different loop counts, from scratch on 18T tokens, followed by matched instruction tuning and evaluation. Empirically, the two-loop variant delivers broad gains over the non-looped baseline across code generation, code reasoning, agentic software engineering, and tool-use benchmarks, improving SWE-bench Verified from 43.0 to 64.4 points and Multi-SWE from 14.0 to 31.0 points. In contrast, variants with three or more loops regress, revealing a strongly non-monotonic loop-count effect. Our diagnostics show that loop 2 provides the main productive refinement, while later loops yield diminishing, oscillatory updates and reduced representational diversity. Because the CLP-induced mismatch remains roughly fixed as refinement gains shrink, the offset cost increasingly dominates. This gain--cost trade-off explains PLT's saturation at two loops and provides diagnostics for loop-count selection.