WorldLines:长期展望的有状态实体化智能体的基准测试与建模
WorldLines: Benchmarking and Modeling Long-Horizon Stateful Embodied Agents
摘要
为了帮助人类在长期时间内在真实家庭中发挥作用,那些具有实体特性的智能体必须记住用户的习惯、环境状态以及过去的互动情况。现有的长期记忆评估标准主要侧重于以语言为中心的检索和问答任务,而那些针对实体智能体的评估标准则通常关注短期内的任务执行,并未对动态环境中的长期记忆使用进行考核。我们提出了WorldLines这一项目驱动的评估框架,用于评估长期范围内的家庭辅助功能。该框架能够构建包含对话、动作、执行反馈、物体和设备状态变化等信息的长期家庭记录,并将其转化为与记忆查询和实体任务规划相关的样本。此外,我们还提出了ObsMem这一基于观察者的记忆框架,它能够维护出具备视角的记忆信息以及与动作相关的状态记录,从而帮助做出正确的决策。实验表明,部分可观测性不足、环境状态被覆盖,以及将长期记忆转化为实体计划等问题仍然存在,而ObsMem则提供了更完善的参考架构来应对这些挑战。
English Abstract
To assist humans over extended periods in real homes, embodied agents must remember user routines, world states, and past interactions. Existing long-term memory benchmarks mainly evaluate language-centric retrieval and question answering, while embodied benchmarks often focus on short-horizon task execution without testing long-term memory use in dynamic environments. We introduce WorldLines, a project-driven benchmark for long-horizon embodied household assistance. It constructs temporally extended household traces with dialogues, actions, execution feedback, object and device state changes, and converts them into evidence-linked samples for Memory QA and Embodied Task Planning. We further propose ObsMem, an observer-grounded memory framework that maintains visibility-aware memories and action-native state trails for state-aware decisions. Experiments reveal persistent challenges in partial observability, overwritten world states, and translating long-term memory into embodied plans, while ObsMem offers a stronger reference architecture for this setting.