AgentOdyssey:一种用于测试阶段持续学习智能体、具有开放性的长期目标文本生成游戏机制
AgentOdyssey: Open-Ended Long-Horizon Text Game Generation for Test-Time Continual Learning Agents
摘要
为了让智能体能够在测试时从与世界的互动中持续学习,它们必须能够有效探索、获取新的知识与技能、记住相关的经历,并能够进行长期规划。为了评估这些关键能力,我们提出了AgentOdyssey这一新的评估框架。该框架能够生成包含丰富实体、世界动态以及长期任务内容的开放式文本游戏。重要的是,AgentOdyssey打破了传统机器学习中的假设——即学习不会在测试时发生——它将智能体置于一个持续进行的长期环境中,使学习和推理过程贯穿整个运行过程。我们还提出了一种多方面的评估方法,不仅衡量游戏的进展,还评估智能体的知识获取能力、情景记忆、物体与行为的探索能力、行为多样性以及模型成本等。我们在生成的游戏中对多种智能体模式进行了评估。实验结果表明,智能体的关键能力存在一定的限制,同时也有一些因素会影响其有效的认知范围。虽然性能随着基础模型的提升而提高,但即使是最优秀的智能体也远远低于人类水平,因此仍有很大的改进空间。在各种智能体机制中,我们发现短期记忆对多种智能体模式都有益处,它是智能体在测试时训练的重要要素。
English Abstract
For agents to learn continuously from interaction with the world at test time, they must be able to explore effectively, acquire new world knowledge and skills, retain relevant episodic experiences, and plan over long horizons. To evaluate these key abilities of test-time continual learning agents, we introduce AgentOdyssey, a novel evaluation framework that procedurally generates open-ended text games with rich entities, world dynamics, and long-horizon tasks. Critically, AgentOdyssey goes beyond the conventional machine learning assumption that learning does not occur at test time by placing agents in a continuous, long-horizon setting that interleaves learning and inference throughout deployment. We further propose a multifaceted evaluation methodology that measures not only game progress but also offers diagnostic tests on world knowledge acquisition, episodic memory, object and action exploration, action diversity, and model cost. We evaluate diverse agent paradigms in the generated games. Our experimental results reveal critical limits in agents' key abilities, as well as factors that influence their meaningful horizon. Although performance scales with stronger base models, even the top agent remains far below human performance, leaving substantial headroom for improvement. Among agent mechanisms, we find that short-term memory benefits multiple agent paradigms and is an important component of agent test-time training.