‹ 返回 2026-06-24

DataClaw0:从原始数据流中对多模态数据进行智能处理

DataClaw0: Agentic Tailoring Multimodal Data from Raw Streams

▲ 65 💬 1 2026-06-24

Cong Wan, Zeyu Guo, Zijian Cai, Jiangyang Li, SongLin Dong, Lin Peng, Xiangyang Luo, Zhiheng Ma, Yihong Gong

摘要

大量的非结构化多模态数据存在“数据熵”较高的问题,这阻碍了人类知识的有效获取以及高质量人工智能模型的训练过程。现有的被动注释方法依赖于启发式规则或通用虚拟语言模型,这些方法成本较高且效果单一,无法充分挖掘原始数据中蕴含的深层逻辑。我们则将数据处理提升为可学习的能力,提出了一种以代理式数据定制为核心的范式转变方式——这种方式能够主动对数据进行优化和结构化处理,从而满足不同用户需求及下游任务的需求。为了克服训练如此高级能力时面临的数据稀缺问题,我们设计了两阶段流程,将生成性语义合成与确定性事实锚点相结合,从而获得涵盖五个核心物理与数字领域的大规模数据集。在此基础上,DataClaw_0-9B模型将监督式精细调整与群体相对策略优化技术结合使用,实现了对复杂需求的高效处理。为了系统地量化这一能力,我们构建了DataClaw_0-val基准测试,这是第一个专门用于数据优化的基准测试。重要的是,我们以下游训练后的结果作为最终验证标准。在视频生成、现实世界VQA和GUI导航方面的测试表明,DataClaw_0能够提供信息密度高的定制数据,从而帮助模型在有限训练数据条件下高效适应新任务。项目页面:https://czjdsg.github.io/MakeAnyData

English Abstract

Massive unstructured multimodal streams suffer from high "data entropy," impeding both efficient human knowledge acquisition and high-quality AI post-training. Existing passive annotation paradigms, heavily reliant on heuristic rules or general VLMs, are costly, monotonous, and fail to unlock the deep procedural logic embedded in raw data. We elevate data processing to a learnable capability, proposing a paradigm shift towards Agentic Data Tailoring, which actively refining and structuring data to align with diverse user and downstream intents. To overcome the data scarcity bottleneck in training such high-order capabilities, we design a two-stage pipeline grounding generative semantic synthesis in deterministic Factual Anchors, yielding a large-scale dataset spanning five core physical and digital domains. Building upon this, DataClaw_0-9B model synergizes Supervised Fine-Tuning (SFT) with Group Relative Policy Optimization (GRPO), achieving robust alignment with complex refinement and tailoring intents. To systematically quantify this capability, we construct DataClaw_0-val, the first benchmark dedicated to data refinement. Crucially, we adopt downstream post-training as the ultimate validation touchstone. Evaluations on video generation, real-world VQA, and GUI navigation confirm that DataClaw_0 delivers high-information-density tailored data, facilitating efficient model adaptation to new tasks under limited training data regimes. Project page: https://czjdsg.github.io/MakeAnyData