Audiobox is a unified AI audio generation research demo developed by Meta's Fundamental AI Research (FAIR) lab that explores the capability to generate multiple categories of audio—speech, natural sounds, and soundscapes—from text descriptions and optional voice prompts within a single model architecture. Users can provide a text description of desired audio content and, for speech generation, a reference voice recording to guide the output's vocal characteristics. The system aims to demonstrate that a single audio model can address speech synthesis, sound effect generation, and ambient audio production rather than requiring separate specialized models for each category.
Browse tools. Audiobox is primarily of interest to AI researchers exploring audio generation models, audio engineers and sound designers who want to experiment with text-to-audio generation for sound effects and environments, and creative technologists testing AI-generated voice and audio for multimedia projects. A typical experimental workflow involves writing a descriptive text prompt for the desired audio output, optionally providing a voice reference for speech synthesis, and generating the audio through the demo interface. Podcast creators, game audio designers, and film production teams exploring AI audio tools have experimented with Audiobox for ambient soundscape generation and sound effect prototyping
See related options.