TUNENOODLE / Stem splitter / FX & Dialogue

BandIt v2 multilingual soundtrack separator

Separate Speech, Music and Sound effects with the BandIt v2 mode, informed by multilingual cinematic-audio research.

BandIt v2 multilingual soundtrack separator

BandIt v2 keeps the three soundtrack roles of Speech, Music and Sound effects while connecting the BandIt model family to the multilingual Divide and Remaster v3 dataset. The 2024 research expanded the dataset’s language coverage and revised how mixtures were constructed. MVSEP links this mode to the authors’ v2 implementation and lists a multilingual provider default. In Tunenoodle, choose the mode itself; there is no language-specific model selector. This makes it a relevant option to compare for multilingual interviews, subtitled scenes or material that changes language within one excerpt. Multilingual training is useful background, not a promise of equal results for every accent, language or recording condition. The output is separated audio and does not transcribe, translate or label the words. Evaluate each language transition as an audio event. Check quiet syllables, final consonants and speech over a musical entrance, and listen to the other stems for fragments of words. A loud clear sentence can survive while a short response is weakened. Keep the same original excerpt when comparing with BandIt Plus or DnR v3, and select the result that preserves the actual speech and scene details required by your edit.

How to separate your track

  1. 01

    Choose BandIt v2 and separate the original multilingual passage.

  2. 02

    Review each speaker or language transition in all three stems.

  3. 03

    Keep useful speech excerpts and compare the same source with another cinematic mode where critical words remain incomplete.

What this mode gives you
What this mode gives you
Parameter / optionWhat it changes
Start with the right sourceInclude a short response and a language or speaker change, not just a loud narrator. Submit the original mixed audio; selecting this mode does not require a language prompt.
Good forCompare dialogue separation across a language switch in one scene. Hear short responses under a musical cue more clearly for manual review. Separate the speech layer of a multilingual interview while retaining score and location effects.
Listen forDo quieter syllables survive as well as loud sentences? Does speech stay consistent through a language or speaker change? Are word fragments audible in Music or Sound effects after a musical entrance?

A few useful answers

No language-specific model control is exposed in this mode. The app requests the provider’s mode without adding a language model option.

Your next step

All modes

Send feedback

Tell us what went wrong or what would help. Every message is read.

Type

Sending from /features/stem-splitter/bandit-v2