TUNENOODLE / Stem splitter / FX & Dialogue
BandIt v2 multilingual soundtrack separator
Separate Speech, Music and Sound effects with the BandIt v2 mode, informed by multilingual cinematic-audio research.
BandIt v2 multilingual soundtrack separator
BandIt v2 keeps the three soundtrack roles of Speech, Music and Sound effects while connecting the BandIt model family to the multilingual Divide and Remaster v3 dataset. The 2024 research expanded the dataset’s language coverage and revised how mixtures were constructed. MVSEP links this mode to the authors’ v2 implementation and lists a multilingual provider default. In Tunenoodle, choose the mode itself; there is no language-specific model selector. This makes it a relevant option to compare for multilingual interviews, subtitled scenes or material that changes language within one excerpt. Multilingual training is useful background, not a promise of equal results for every accent, language or recording condition. The output is separated audio and does not transcribe, translate or label the words. Evaluate each language transition as an audio event. Check quiet syllables, final consonants and speech over a musical entrance, and listen to the other stems for fragments of words. A loud clear sentence can survive while a short response is weakened. Keep the same original excerpt when comparing with BandIt Plus or DnR v3, and select the result that preserves the actual speech and scene details required by your edit.
How to separate your track
- 01
Choose BandIt v2 and separate the original multilingual passage.
- 02
Review each speaker or language transition in all three stems.
- 03
Keep useful speech excerpts and compare the same source with another cinematic mode where critical words remain incomplete.
What this mode gives you
| Parameter / option | What it changes |
|---|---|
| Start with the right source | Include a short response and a language or speaker change, not just a loud narrator. Submit the original mixed audio; selecting this mode does not require a language prompt. |
| Good for | Compare dialogue separation across a language switch in one scene. Hear short responses under a musical cue more clearly for manual review. Separate the speech layer of a multilingual interview while retaining score and location effects. |
| Listen for | Do quieter syllables survive as well as loud sentences? Does speech stay consistent through a language or speaker change? Are word fragments audible in Music or Sound effects after a musical entrance? |
A few useful answers
No language-specific model control is exposed in this mode. The app requests the provider’s mode without adding a language model option.
Your next step
All modesCrowd and dialogue separator
Separate a crowd-focused layer from a live recording, with the remaining performance in Other for comparison and editing.
Dialogue, music and effects separator
Create separate Speech, Music and Sound effects tracks from a mixed soundtrack with the Demucs4HT DnR mode.
BandIt Plus soundtrack separator
Separate Speech, Music and Sound effects with BandIt Plus, then compare the layers around the words and scene events that matter.
DnR v3 soundtrack separator
Separate Speech, Music and Sound effects with MVSEP’s DnR v3 mode for detailed comparison of mixed cinematic audio.
Braam effect separator
Isolate a braam-style cinematic hit to study its low-end impact, tonal body and decay apart from the surrounding arrangement.
Riser effect separator
Extract rising transition effects to study buildup shape, phrase length and the arrival of a drop or chorus.
FX separator
Separate an effects-focused layer from a mixed recording to inspect transitions, accents and textures alongside the remaining audio.
Uses credits
