‹ 返回 2026-06-18

EgoCS-400K:用于世界模型的自中心性游戏数据集

EgoCS-400K: An Egocentric Gameplay Dataset for World Models

▲ 7 💬 1 2026-06-18

Rongjin Guo, Dong Liang, Yuhao Liu, Fang Liu, Tianyu Huang, Gerhard P. Hancke, Rynson W. H. Lau

摘要

从视频生成到交互式世界建模的转变,对数据提出了新的要求:除了带有字幕的视频之外,世界模型还需要包含时间上一致的视频-动作-语言轨迹信息,这些轨迹必须基于各种动作、摄像机运动、状态以及推动场景变化的事件来构建。然而,如此规模的数据获取却非常困难。网络视频数据集虽然提供了广泛的视觉信息,但缺乏可执行的动作数据和可靠的状态信息;而机器人数据集则提供了关于动作和状态的详细信息,不过其成本较高,且场景多样性有限。此外,现有的模拟器往往缺乏大规模的人机交互轨迹数据。在本文中,我们介绍了EgoCS-400K这一大规模的重放基础版以自我为中心的《反恐精英》数据集。该数据集基于公开的职业级《反恐精英》和《反恐精英2》比赛录像构建而成,能够保留人类的游戏行为轨迹,并支持数据的解析、重放、渲染以及时间对齐功能。我们提取了玩家状态、视角方向、移动路径、键盘/按钮输入、视角变化、武器使用方式、游戏事件以及每轮比赛的上下文信息,然后从这些轨迹中生成清晰的第一人称视频画面。EgoCS-400K包含超过400,000条第一人称视频画面,以及10,000小时的游戏数据,涵盖1,000场比赛中的40,000个回合,涉及13种地图和每轮比赛中的10种视角。该数据集支持多种交互式视觉建模任务,包括基于动作的未来预测、考虑状态和事件的场景展示、基于重放的数据标注,以及以自我为中心的行为理解等。通过将视觉观测数据与人类的动作、摄像机运动、游戏状态及事件相结合,EgoCS-400K成为了被动网络视频、可控的游戏模拟以及昂贵的现实世界数据之间的实用桥梁。

English Abstract

The shift from video generation to interactive world modeling places new demands on data: beyond captioned videos, world models require temporally aligned video-action-language trajectories grounded in the actions, camera motion, states, and events that drive future scene changes. However, such data is difficult to obtain at scale. Web video datasets offer broad visual coverage but lack executable actions and reliable states; robotic datasets provide action and state supervision but are costly and limited in scene diversity; and existing simulators often lack large-scale human-driven interaction trajectories. In this paper, we introduce EgoCS-400K, a large-scale replay-grounded egocentric Counter-Strike dataset for world models, built from public professional CS and CS2 match demos that preserve human gameplay trajectories and enable parsing, replaying, rendering, and temporal alignment. We extract player states, view directions, movements, keyboard/button inputs, view-angle changes, weapon usage, game events, and round-level context, and render clean first-person videos from the same trajectories. EgoCS-400K contains over 400,000 first-person videos and 10,000 hours of gameplay from more than 1,000 matches and 40,000 rounds, covering 13 maps and 10 player viewpoints per round. It supports a range of interactive visual modeling tasks, including action-conditioned future prediction, state- and event-aware scene rollout, replay-grounded captioning, and agent egocentric action understanding. By connecting visual observations with human actions, camera motion, game states, and events at scale, EgoCS-400K serves as a practical bridge between passive web videos, controllable game simulation, and costly real-world embodied data.