
The bottleneck is almost never the message. It is the fourth recording session.
A public health campaign gets written once, produced once, and then has to exist in six languages before it is any use to the people it was written for. Our team has run that process both ways, and the difference between the two routes is not artistic. It is a question of how many studio hours a station has, which in most places is the binding constraint on whether the message goes out at all.
Route one: re-record everything
The traditional method treats each language version as a new production.
You book the studio. You bring in a presenter who speaks the target language. You re-record the narration, then rebuild the music and effects underneath it, because the original bed was mixed into the original master and cannot be separated from the voice you no longer want.
For a well-funded broadcaster this is fine. For a community station running on a generator and a donated mixing desk, each additional language costs a session it does not have. In practice, this is where multilingual plans quietly become monolingual broadcasts. The Hausa version airs. The other five stay in a project document.
Route two: keep the bed, replace the voice
The alternative starts from a simple observation: the music, the ambience and the sound design are language-neutral. Only the narration needs to change.
If the original multitrack still exists, this is trivial — mute the voice track and record over it. The problem is that the multitrack usually does not exist. What survives is a finished MP3, handed over by a producer who has moved on, from a project that closed two funding cycles ago.
This is where separation changed the arithmetic. Feed that finished file to a vocal remover and it hands back two things: the narration alone, and the bed alone. The bed is what you want. Record the new language over it and the second version costs a fraction of the first.
A worked example
A station receives a three-minute maternal health spot as a single MP3. Original narration in English, music bed underneath, no session files. The producer runs it through a vocal remover, keeps the instrumental side, and hands the presenter a bed with the original timing intact. The presenter records in Kiswahili against that bed. Total studio time for version two: under an hour, most of it spent on the presenter’s takes rather than on rebuilding music that already existed.
Where route two fails
This route is not universally better, and a station that adopts a vocal remover without knowing the failure modes will produce something worse than route one.
Dense mixes leave traces. If the narration was recorded in a live room, or the bed sits very close to the voice in pitch, a vocal remover leaves artifacts. On a phone speaker nobody notices. On an FM transmission at volume, a listener does.
Timing is inherited, not negotiable. The new narration has to fit the pacing of the old. Languages are not equally compact. A sentence that runs four seconds in English can run seven in another language, and the bed will not wait. Usually the script gets adapted, or the version sounds rushed.
One escape exists, and it applies only to music a station owns outright. Stretch audio and it degrades; stretch notes and nothing breaks. Run the bed through mp3 to midi detection and the melody comes back as note data instead of a fixed recording — and note data will sit at a slower tempo, shift key, or return played on something else entirely. Any station holding the rights to its own theme can then rebuild it around however long the sentence turned out to be, rather than cutting the sentence down to fit an audio file it cannot change.
Two limits are worth stating. Detection follows a single melody line only, so simple station idents and sparse beds transcribe usefully while chords and layered arrangements do not. And ownership does not move: pushing somebody else’s recording through note detection leaves the rights in that song exactly where they already sat.
Rights do not transfer with the file. Running a production through a vocal remover does not grant permission to reuse the bed that comes out of it. If the music was licensed for one campaign, in one territory, for one period, that is what it is licensed for. This is worth settling before the session, not after the broadcast.
Quality degrades from the source down. Hand a vocal remover a file that three messaging apps have already re-compressed and it has less to work with. Common formats are accepted at file sizes well above a typical radio spot, so the limit is rarely the tool — it is whether anyone kept a clean copy.
The decision rule we use
Use route one when the material will carry a station’s reputation — flagship programming, anything archival, anything where a licensing question is unresolved.
Reach for route two when the constraint is exactly that: when the alternative is not a better version but no version, in a language a community actually speaks.
Most public information work sits in the second category. That is the honest reason the method matters. It is not that audio from a vocal remover sounds as good as a fresh production. It is that a slightly imperfect message in Kiswahili reaches people that a pristine message in English does not.
What it has bought us
- Five additional language versions from one production budget, on a campaign that had funded two.
- A usable archive: old finished files became reusable beds instead of dead weight.
- Presenter time spent on delivery rather than on waiting for a bed to be rebuilt.
- A cost per language low enough that adding a sixth stopped requiring a meeting.
None of that is a technical achievement. It is the removal of a specific, boring obstacle that was quietly deciding which communities got informed and which did not.

