剧本:让LLM智能体拥有音乐结构的感知
Libretto: Giving LLM Agents a Sense of Musical Structure
摘要
现在,生成式音乐系统能够根据文本提示生成令人印象深刻的音频作品。不过,这些音频作品很难被仔细分析、编辑或判断其音乐结构。我们提出了Libretto这一用于符号化音乐生成与修改的框架。Libretto采用了基于LLM的语法结构,包含明确的音高、声部以及小节级别的组织方式;随后,该框架会依据节奏、和声、旋律、音色、形式及变化等指标,对每一首音乐作品进行评估。这些相同的结构特征也支持检索、诊断、复制风险控制以及迭代式自我修改等操作。在填补空白、以参考信息为引导的完整作品生成、逐步变形以及教育性音乐创作等方面,Libretto能够将符号化音乐从纯粹的字符序列转化为可测量且可编辑的对象,从而让语言模型能够处理这种音乐作品。
English Abstract
Generative music systems can now produce impressive audio from text prompts, but audio outputs are difficult to inspect, edit, and diagnose as musical structure. We introduce Libretto, an agent-facing framework for symbolic music generation and revision. Libretto uses an LLM-native grammar with explicit onset slots, voices, and bar-level organization, then evaluates each piece in a corpus-calibrated statistical space over rhythm, harmony, melody, texture, form, and variation. The same structural axes support retrieval, diagnosis, copy-risk control, and iterative self-revision. Across gap filling, reference-guided full-piece generation, gradual morphing, and educational music generation, Libretto turns symbolic music from a raw token sequence into a measurable and editable object for language-model agents.