Recent text-to-audio models generate high-quality audio, but often fail to follow realistic instructions involving multiple sound events and explicit temporal order. This gap stems from evaluation and training signals that emphasize global similarity or perceptual quality, with little supervision on event completeness and ordering. We propose an instruction-level evaluation and training framework that uses audio-aware large language models (ALLMs) as judges to explicitly verify target event presence and temporal relations in generated audio. We show modern ALLMs can reliably perform these judgments on multiple audio understanding benchmarks, and use their fine-grained feedback to construct preference pairs for direct preference optimization. To evaluate temporally structured generation, we introduce Sound Scene Story Benchmark (S3Bench), a narrative benchmark with multi-event temporal progression. Experiments consistently improve sound event completeness and temporal ordering across existing benchmarks and S3Bench.
Below, we present:
(1) Side-by-side comparisons of generated audio samples across multiple benchmarks,
highlighting how our method improves sound event completeness and temporal ordering compared to baseline systems;
(2) audio samples across multiple DPO iterations, demonstrating the progressive performance gains achieved through iterative preference optimization; and
(3) model-based evaluation using Meta AudioBox Aesthetics.
Our model maintains comparable production quality while achieving stronger instruction-following ability,
indicating no trade-off between instruction following and audio quality.
The audio samples and results shown here were independently reproduced and regenerated by the author. The work itself builds on publicly available models, datasets, and open-source components (see the full disclaimer ↓).
Note: The audio examples below are independently regenerated using public/open-source components and are not outputs from the original internship experiments. Read the full disclaimer ↓
For confidentiality reasons, audio samples generated during the internship project are not released on this page. The work itself builds on publicly available models, datasets, and open-source components. The audio examples provided here were independently regenerated using those public components and the implementation details described in the paper and the publicly released GitHub repository. No internal code, model checkpoints, generated samples, prompts, logs, or non-public training data from the internship project were used.
These examples are intended only to illustrate the qualitative evaluation format and the type of instruction-following behavior studied in the paper. They should not be interpreted as outputs from the original internship experiments or as an official company demo.