Audio, VAD & Diarization
Collection
Voice Activity Detection, Diarization, and Audio Processing models • 28 items • Updated
Ungated mirrors of Meta's SAM-Audio model weights, converted to BF16 safetensors format and redistributed under the SAM License for easier access.
SAM-Audio (Segment Anything Model for Audio) is Meta AI's foundation model for isolating any sound in audio using text, visual, or temporal prompts. It can separate specific sounds from complex audio mixtures.
The -tv variants are optimized for target correctness and visual prompting.
| Model | Parameters | File Size | Original |
|---|---|---|---|
sam-audio-large-tv-bf16.safetensors |
3,715,221,638 | 6.92 GiB | facebook/sam-audio-large-tv |
sam-audio-base-tv-bf16.safetensors |
1,931,243,654 | 3.60 GiB | facebook/sam-audio-base-tv |
| File | Description |
|---|---|
sam-audio-large-tv-bf16.safetensors |
Large-TV model weights (BF16) |
sam-audio-base-tv-bf16.safetensors |
Base-TV model weights (BF16) |
config.json |
Model configuration |
LICENSE |
SAM License (required for redistribution) |
# With ComfyUI-FFMPEGA (automatic download)
# Set no_llm_mode = "audio_separate" and prompt = "vocals"
# Or standalone:
from sam_audio import SAMAudio
model = SAMAudio.from_pretrained("path/to/this/repo")
This model is distributed under the SAM License — see the LICENSE file. Key points:
Base model
facebook/sam-audio-large-tv