MIRA: A Musical Intent Refinement Agent for Aligning Text-to-Music Generation with User Intent

Zekai Liu, Zhilin Wang, Xuzheng He, Yu Cheng, Yang Yang

arXiv:2610.10355 · 2026-10-11 공개 · arXiv · PDF

music-generation prompt-refinement tree-search text-to-music rubric-evaluation intent-alignment mira-agent sunoo-mureka

Abstract

Text-to-music systems produce increasingly convincing audio, yet evaluation reveals little about whether the result matches user intent. A global text-audio relevance score can overlook the implicit intent in underspecified prompts and mask failures in specific requirements, such as instrumentation, structure, rhythm, or mood progression. To bridge this gap, we formulate text-to-music intent alignment as satisfying a per-request rubric of independently verifiable items covering both a request's explicit requirements and its implied musical intent. Scoring items individually makes evaluation diagnostic by intent source and musical dimension, rather than a single opaque score. We instantiate this as MuRA-Bench, a benchmark of real-world platform requests curated by music experts. We further propose MIRA (Musical Intent Refinement Agent), a test-time agent that first grounds a request's intent into rubrics, then searches over prompt revisions for a black-box generator under a bounded budget, iteratively generating music, verifying it against the rubrics, and using this feedback to guide a trajectory-aware tree search. Experiments across open-source and commercial backends show that MIRA improves intent alignment, enabling an open-source generator to achieve performance comparable to representative commercial systems (e.g. Suno and Mureka). Project page: https://mirareview.github.io/.

한국어 요약

한 줄 요약

MIRA는 사용자 의도와 음악 생성을 정렬하기 위해 MuRA-Bench 기반의 검증 가능한 루비릭 항목을 사용하는 테스트 시점 에이전트다.

핵심 기여도

핵심 아이디어

기존 텍스트-음악 생성 시스템은 전체적 관련성 점수만 제공하여 사용자 의도와의 정렬 여부를 명확히 알 수 없다. 이를 해결하기 위해, MIRA는 사용자 요청을 기반으로 **루비릭(rubric)** 항목을 생성하고, 생성된 음악이 이 루비릭을 충족하는지 검증하는 방식을 제안한다. 루비릭은 **명시된 제약 조건**과 **맥락에서 유추된 음악적 의도**를 모두 포함하며, 각 항목은 독립적으로 검증 가능하다. MIRA는 검증 결과를 바탕으로 **트리 구조의 피드백 메모리**를 사용해 향후 생성을 개선한다. 이는 기존의 생성 전 계획 또는 모델 내부 조절 방법과 달리, 생성 후 피드백을 기반으로 한 **검증자-가이드드 검색**을 구현한다.

기술적 접근법

주요 결과

의의 및 한계

MIRA는 텍스트-음악 생성의 **사용자 의도 정렬**을 정량적으로 평가하고 개선할 수 있는 새로운 프레임워크를 제시한다. MuRA-Bench는 실제 요청 기반의 루비릭을 통해 **명시적 제약과 암묵적 의도를 모두 평가**할 수 있다. MIRA는 생성기와 독립적으로 작동하므로, 다양한 블랙박스 시스템에 적용 가능하다는 장점이 있다. 그러나 MIRA는 **사전에 루비릭이 필요**하므로, 즉석 요청 처리에는 한계가 있을 수 있다. 또한, 검증기의 정확도가 최종 성능에 영향을 미칠 수 있다.

실용적 활용

MIRA는 음악 생성 플랫폼에서 사용자 요청을 더 정확히 반영하는 **개인 맞춤형 음악 생성**에 활용 가능하다. 또한, 음악 제작 도구나 AI 기반 음악 제작 서비스에서 **사용자 피드백을 기반으로 반복적으로 음악을 개선**하는 데 유용할 수 있다. 특히, MuRA-Bench는 음악 생성 모델의 평가 및 개선을 위한 **표준화된 벤치마크**로 활용될 수 있다.