LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling
LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling
摘要
LoopCoder-v2 是一款基于并行循环Transformer(PLT)架构的代码大语言模型。该模型通过引入跨循环位置偏移(CLP)和共享KV门控滑动窗口注意力(G-SWA)机制,实现了测试时计算的高效扩展,有效避免了标准循环Transformer的序列延迟和KV缓存内存膨胀问题。研究表明,在7B参数规模下,单次额外循环(R=2)能在代码生成、推理及智能体任务中带来显著的性能提升,而更多循环则因固定的位置偏移成本导致性能下降。
English Abstract
Looped Transformers scale latent computation by repeatedly applying shared blocks, but sequential looping increases latency and KV-cache memory with the loop count. Parallel loop Transformers (PLT) alleviate this cost through cross-loop position offsets (CLP) and shared-KV gated sliding-window attention, making loop count a practical design choice. We therefore study PLT loop-count selection through a gain--cost view: an extra loop may refine representations, but CLP also introduces a positional mismatch at each loop boundary. We instantiate this study by training LoopCoder-v2, a family of 7B PLT coders with different loop counts, from scratch on 18T tokens, followed by matched instruction tuning and evaluation. Empirically, the two-loop variant delivers broad gains over the non-looped baseline across code generation, code reasoning, agentic software engineering, and tool-use benchmarks, improving SWE-bench Verified from 43.0 to 64.4 points and Multi-SWE from 14.0 to 31.0 points. In contrast, variants with three or more loops regress, revealing a strongly non-monotonic loop-count effect. Our diagnostics show that loop 2 provides the main productive refinement, while later loops yield diminishing, oscillatory updates and reduced representational diversity. Because the CLP-induced mismatch remains roughly fixed as refinement gains shrink, the offset cost increasingly dominates. This gain--cost trade-off explains PLT's saturation at two loops and provides diagnostics for loop-count selection.