‹ 返回 2026-06-19

MaineCoon:构建实时音视频社交世界模型

MaineCoon: Pursuing A Real-Time Audio-Visual Social World Model

▲ 7 💬 1 2026-06-19

摘要

MaineCoon代表了首个用于社交场景的实时音视频自回归模型。通过创新的训练技术和推理框架,该模型能够实现高帧率以及长时段的生成效果。该模型由Qwen/Qwen2.5-Coder-32B-Instruct生成。随着越来越多的全球视频内容在社交平台上进行互动式消费,专为社交场景设计的视频生成模型非常重要,但之前的研究却忽视了这一领域。在这项研究中,我们首先定义了社交场景模型的定位,并构建了原型模型,以此作为实现这一目标的第一步。虽然之前的社交场景模型能够模拟物理环境或游戏世界,但它们仍然与以人类为中心的社交动态相脱离。为了填补这一空白,我们提出了MaineCoon——首个具有222亿个参数的实时音视频自回归模型,它能够实现实时流媒体生成以及亚秒级的互动效果,其帧率高达47.5 FPS,且可以在单个GPU上运行。据我们所知,MaineCoon也是首个专门为社交互动应用优化的实时音视频生成模型。为了实现高效稳定的训练过程,我们在MaineCoon中引入了多种创新技术,包括自重采样、跨模态表示对齐、领域感知偏好优化以及强化在线策略蒸馏等技术。我们还设计了首个代理式流媒体推理框架,该框架能够支持千秒级甚至更长时间的生成过程,同时通过代理式缓存管理和提示规划来减少性能波动。这些创新显著加快了训练过程,同时优化了实时推理性能。我们认为,这项工作不仅为高质量、低延迟、长时段的音视频自回归模型树立了新的最佳实践标准,也指出了下一代AI原生社交平台所需要的范式转变方向。

English Abstract

MaineCoon represents the first real-time audio-visual autoregressive model for social worlds, achieving high frame rates and long-horizon generation through novel training techniques and inference frameworks. Generated by Qwen/Qwen2.5-Coder-32B-Instruct As an increasing majority of global video content is consumed on social platforms for interactive social purposes, video generation models built for social worlds are important but largely overlooked by previous studies. In this work, we define the position of social world models and build a prototype model as the first step towards this goal. While previous world models successfully simulate physical environments or gaming world exploration, they remain fundamentally detached from human-centric social dynamics. To bridge this gap as the first step to social world models, we present MaineCoon, the first real-time audio-visual autoregressive model that has 22B parameters and is capable of real-time streaming generation and sub-second interaction, with a record-breaking frame rate of up to 47.5 FPS, on a single GPU. To the best of our knowledge, MaineCoon is also the first real-time audio-visual generation model specifically optimized for social-interactive applications. To enable efficient and stable training, we introduce several novel techniques into MaineCoon, including self-resampling , cross-modal representation alignment , domain-aware preference optimization , and reinforced online-policy distillation (ROPD). We also design the first agentic streaming inference framework that supports thousand-second-scale or even longer generation while mitigating drift with agentic cache management and prompt planing. These innovations significantly accelerate training while optimizing real-time inference performance. We believe this work not only sets a new state-of-the-art (SOTA) performance benchmark for high-quality, low-latency, and long-horizon audio-visual autoregressive model s, but also points out the paradigm shift desired for next-generation AI-native social platforms.