Ask an AI about this

Optional, but it's what turns a good translation into a fully localized one. Audio Lab is where the game's voices get dubbed, line by line.

Audio Lab & voices

Optional, but it's what turns a good translation into a fully localized one. Audio Lab is where the game's voices get dubbed, line by line.

Dubbing is entirely optional. A text-only translation is already a complete, publishable thing, and most adaptations start (and many happily stay) that way. But hearing characters speak your language is a different level of immersion, and Audio Lab is the room where that happens: every voice clip in the game, lined up with its original audio and your translated script, waiting for a voice.

How dubbing works

The unit of work is the line: one spoken sentence or bark. Tranzio finds the voice clips inside the game's packed files, attaches each one to the script line it speaks, and lists them. You pick a line, record it, keep the take you like, and move on. Line by line, character by character, a dub grows.

VO_Ch02_Mira_014· 0:03.4
Wait. I have heard that voice before.
Dub
VO_Ch02_Mira_015· 0:02.1
Keep close to the wall.
Linked
CS_Ch02_Intro· 0:41.8
The harbour was quiet the night they came.(first of several)
6 subtitles
AMB_Harbour_loop.wem· 0:18.2
- no linked text -
Unlinked
The line list, with one line playing. Each row is numbered by its place in the list, then carries the clip name and length, the words it says, what state it's in, and its waveform. Once a line is translated the translation is what you read, with the original underneath it. The toolbar above counts how many clips the current filters left.
  1. Scan audio

    Sweeps the game's packed and loose files for voice clips, sound banks included, and indexes every one it finds. Run it once per game; re-run it after the game updates.

  2. Link texts

    Attaches each clip to the script line it speaks, so you can read the words while you record. Most clips link automatically; the rest you point at by hand.

  3. The labels rail

    Your characters, down the left. Labels are shared with the Translate page, so tagging a line as Mira in either place groups her voice clips here, with her own dubbing progress.

  4. The line you are dubbing

    Every screen that shows a voice line — the list, the inspector, the Studio player, a DAW clip — leads with the translation once the string has one, and keeps the original underneath as the reference. If it has no translation yet, the button under the line writes one without leaving the page.

  5. The audio editor

    Opening a line hands it to the multi-track editor: record as many takes as you like, trim them, shape the voice, and keep the best one.

There's no separate render step. The take you keep is stored as this line's override, and the Build page encodes it into the game's own audio format and packs it with the patch. Recording is shipping.

Link by names, at the top of the page, is the same step Translate a game runs last: it reads the map the game baked into its files (Wwise SoundBanksInfo, loose or inside the paks, then line ids, asset names, FMOD scene names and character names) and links clips to lines from it in seconds. Its dialog shows each file it reads and each route it tries as it goes. It is not guaranteed: a game that keeps that map in a shape Tranzio does not read links little or nothing, which is normal rather than a fault, and the dialog then points you at Link by listening for the rest.

What the badges mean

Every row wears one badge, and it's the fastest way to read a game's audio at a glance. Sorting by status turns the list into a to-do queue: unlinked clips first if you're still matching, undubbed ones first if you're recording.

  • Unlinked
    The clip was found in the game, but no script line is attached to it yet.
  • Linked
    The clip says exactly one game string, so you can see the words while you record.
  • 6 subtitles
    One long clip that speaks several strings, like a cutscene stem read in one take.
  • Dub
    Your recording replaces this clip in the patch. This is a finished line.
  • AI
    A generated voice stands in for now. Worth a listen before you ship it.
The five states, using the real badges from the line list.

Games pack music, ambience, and sound effects in the same place as speech, so a scan turns up plenty that nobody should ever dub. Flag those as no dubbing needed and they drop out of your coverage, which stops a finished dub from reading as 40% done.

Finding a clip

The bar above the list is how you find one clip among thousands. The search box on the left matches a file name, the words a clip says, or a line's name. On the right, the filter chooses which clips to show: Unlinked clips are not yet matched to the text they say, Linked ones are, Needs review is where the AI was not sure and wants your eye, and Overridden means you recorded your own voice for it. A game that ships voices in several languages also gets a voice language choice, and Sort changes the order of what is left.

The panel on the right

Clicking a clip opens its details on the right, top to bottom: facts about the file (where it lives, how long it is, how it was matched); what the AI heard, with the lines it thinks match, when a clip was transcribed; the line the clip says, with your translation under it, or the several lines a long clip covers; pictures and videos of the moment it plays; the labels you gave it; your own recording, with Record, Upload and Remove; and the scene and event that make the game play it.

Along the bottom, the player plays the clip you picked: play, previous and next on the left, the sound drawn as waves in the middle (click anywhere on it to jump), and on the right the voice language when there is more than one, the button that lets the AI hear the clip, Record, the sound editor, and a bookmark.

Working with the waveform

42% · click, drag, or use arrow keys to scrub

The waveform view used across Audio Lab. The filled part is played; click or drag anywhere to scrub. You'll live in this view while matching your takes to the original timing.

Hearing what a clip says

A clip with no subtitle attached, a link you do not trust, a take you want to check word for word: Link to text by AI, in the player at the bottom of the page, listens to the clip and shows what it says. It runs OpenAI's Whisper on your own computer, so the audio never leaves it, and the text can be copied straight into a line.

No model is part of Tranzio's installer. The first time you click Link to text by AI it shows the Whisper sizes, from a 45 MB tiny to a gigabyte large, and asks before downloading the one you pick. Small is the recommended middle: it hears most game dialogue well at a few seconds per clip. The English-only variants hear English better than the same size that knows every language, and the bigger sizes cope with accents, whispers and noisy scenes at several times the wait. You can keep several installed and switch, and remove any of them from the Tools page.

What Whisper hears is then matched against the game's text, ignoring punctuation, case and stage directions, and trying runs of neighbouring lines for a clip that says several in a row. A clear match links the clip on the spot. Anything less is shown as candidates with a match score, and you decide: link one, search for a different line, or say the clip has no line at all.

Speech to text

What the clip saysvo_keeper_gate_03.wem
The gate will not hold much longer. Get everyone to the docks before nightfall.
CloseCopy text
The first Link to text by AI on a machine: the question, the download, a few seconds of listening, and the line. Every later clip skips straight to the last two.

Whisper is told which language to hear from the clip's voice language, so switch the take to the right language in the player first. A clip it is asked to hear as English comes back as English, whatever was said.

Linking a whole game by listening

Link by listening, beside Link by names at the top of the page, does the same for every clip that has no line yet. Pick a model and start; it runs in the background, so you can close the dialog, move to another page or another game, and the pill in the corner keeps the count and reopens the dialog. Several games can run at once. Each clip is written to the game's database as it finishes, so stopping loses nothing already heard, and starting again skips what you have already answered.

Clear matches are linked as you watch. The rest land in the list marked Needs review: a clip with candidates to choose from, or one Whisper heard but nothing in the text came close to. Filter the list by Needs review, or sort with Needs review first, and the sidebar shows what was heard and the lines it came close to, each with its score. Link one, search for another, or say the clip has no line, which takes it out of the queue and out of the next run.

Heard by Whisper
AI heardWhisper small

It came from the kitchen, I think.

Linked to VO_AppleCore_03. Matched by: Whisper heard the words.
The sidebar for a clip waiting on you: what Whisper heard, the lines it came close to with their match scores, and the pick becoming a link. A run of several lines is one candidate, marked with how many it spans.

Start with the recommended model on a game's first pass and only reach for a bigger one on the clips it left for review. The big models are several times slower, and most of a game's dialogue does not need them.

Matching lines to game audio

Behind every line is an audio file inside the game. Link texts matches them up for you, and it starts from what the game itself recorded: a Wwise game's own soundbank metadata, an FMOD game's event paths from its strings bank, or a stock-Unreal game's DialogueWave assets, each of which says outright which clip speaks which line. Only when a game ships none of those records does the matching fall back to names, and a game that names its clips by number rather than by line leaves some to judgment. That's where the file picker comes in:

  • VX_M02S08_88.wem

    /game/audio/voice/

  • VX_M02S08_89.wem

    /game/audio/voice/

  • ambient_rain.wem

    /game/audio/sfx/

Picking the game audio file a line belongs to. Voice files usually live together in a voice folder, with names that group by scene.

Some clips aren't one line at all. A cutscene is often shipped as a single stem that speaks half a scene, and Tranzio marks those as a passage with the number of subtitles they carry, so a long clip never looks like a mislinked line. The row shows the first of those lines, but search reads all of them, so you can find a passage by anything said in it rather than only by its opening sentence.

A line that shows no audio at all is often spoken in a pre-rendered cutscene. Those are mixed once, voice and music together, and shipped as a single long recording, so none of their lines has a clip of its own to find. Tranzio matches the recording to the scene by name and marks it a passage carrying all of them. Names rarely agree letter for letter between the recording and the subtitles, so the comparison ignores punctuation and the scene-type tags the two sides spell differently.

The cutscene's video file is not where its sound is. The game plays the picture and the audio separately, and in most builds the video's own track is silent, so the recording Tranzio links is the one in the game's audio, named after the scene.

A passage like this is there to listen to and to match lines against, not to dub line by line: replacing it means re-recording the whole scene as one take. The subtitles translate and build normally.

After every run, Tranzio looks at the whole set of links once more for the shapes that almost always mean a rule matched too much: one line claimed by several clips in the same language, a passage that covers a whole chapter, a clip filed under one character whose line is spoken by another. Nothing is unlinked. Those clips get a Check this link row in the panel saying what looked off, and the link log says how many there were, so you check a handful rather than every line. A link you made yourself is never flagged.

One recording often says several subtitles in a row. In the link picker, use the + on each line, in the order the clip says them, and link them together as a passage; the sidebar then lists every line with its own Change and Remove, and an Add a line at the end for one the clip also says. A clip linked to one line grows into a passage the same way.

Dialogue, music, or a door creak

Most of a game's audio is not dialogue: for every spoken line there are dozens of footsteps, doors, gusts of wind and music stings. When Tranzio links audio to text it also works out what each clip is, from the game's own records: a clip linked to a subtitle is dialogue by definition, Wwise stamps every file with a language (or with "SFX", which means music and effects), FMOD projects sort their events into VO, Music and Ambience folders, and file names carry the rest. The result is the Sound type filter in the toolbar: pick Dialogue and the ten thousand door creaks leave the list. A clip's type and the reason for it are shown in the details panel, and a clip Tranzio cannot be sure about stays "Unclassified" rather than guessed.

The same pass keeps a Dialogue label on the Text page: every line that has audio linked to it gets the label automatically, so the text side can be filtered to spoken lines with one click too.

Games added under an older Tranzio can be sitting on clips and lines with no links at all, because the linker of that day did not understand the game. When you open such a game under a newer version, Tranzio offers once to try again: accept and it runs Link by names with its usual live log, decline and it will not ask again until the next version.

Where the game keeps its audio

Three arrangements, and you should not have to care which one your game uses. Many games keep each sound as its own file inside a pak, put there by audio middleware. Others use Unreal's own audio, where every sound lives inside an asset along with its sample rate and length. Tranzio lists both kinds together, plays both, and replaces both.

The third is a sound bank: one file holding hundreds or thousands of clips, each addressed by a number. Games built on FMOD work this way, Wwise games can embed their media in banks too, and Japanese titles often ship CRIWARE cue sheets (an .acb with the names and an .awb with the audio). A single game can ship well over ten thousand clips in a couple of hundred banks. Tranzio opens all of these for you, so what you see on the Audio page is the clips, listed by the names the bank or cue sheet gives them, with the voice language read from the file's own name or folder.

Clips in a Wwise bank can be replaced: the build rewrites the bank around your takes and ships it in place of the game's, so a dub against one of those clips lands like any other. Clips in an FMOD bank or a CRIWARE cue sheet can be listened to and matched to lines but not yet replaced, so a dub recorded against one of those will not ship; the clip panel says so before you record. Text in such a game translates and builds normally.

The difference only shows up at build time, and only in what gets written. A standalone sound file is replaced with your take. A sound inside an asset can't be: the asset is rebuilt around your take instead, with its length and channel count corrected to match what you actually recorded. Either way the original files in your game folder are never touched, and your take is encoded into the format that game expects before it goes in.

One limit worth knowing: a very long sound is sometimes streamed from a separate file next to the asset, and those cannot be replaced yet. Tranzio leaves them out of the list rather than offering a dub it cannot deliver.

A sound inside an asset also names the codec it was cooked with. Older games use Vorbis, which Tranzio encodes; newer Unreal games default to Bink Audio and, from 5.4, RAD Audio, which have no encoder outside the engine. A dub for one of those cannot be written, and putting the wrong codec in would play as noise, so Tranzio refuses it instead. The clip panel says so on the Dub row before you record, the scan log says it when the game is added, and the pre-build checks count every recorded line that will not ship, with the reason.

Recording a take

Opening a line hands it to the audio editor. Record, listen back, record again: takes stack up and cost nothing, so there's no reason to settle for the first one. When you have one you like, trim the silence off both ends and keep it.

Take 2 kept0:02.6

Trimmed to the words and saved as this line's dub.

One line, from the record button to the kept clip. Every raw take carries a breath at each end, and trimming to the words is what makes your delivery land where the original did.
In the editorWhat it's for
TakesEvery recording you make for this line, kept until you delete it. Only the one you select ships.
Trim and cutCut the silence off the ends, or drop a stumble out of the middle, without leaving the app.
VoiceShape a take by ear: deeper or higher, a bigger or smaller speaker, slower or faster, and heard through a phone, a radio, a megaphone, a robot, a monster. Clean it up, set its tone, put it in a room. Pick a voice with one press, save your own, and use it on every take a character speaks, so one voice sounds like one voice. What you hear is exactly what ships.
The original beside youThe game's own clip stays on screen, so you can hear the pacing you're matching and see whether your take runs long.

Length is the thing that catches people out. A dub that runs much longer than the clip it replaces can end up cut off or running over the next line, because the game's timing was built around the original. If your translation needs more syllables than the original, tighten the wording on the Translate page rather than rushing the delivery.

The sound editor's four steps

The sound editor lays its controls out in the order you use them, as four steps across the top. Capture is where you record: arm a track, press Record, wait for the countdown, say the line. Edit is where you cut away a bad start and move the take into place. Voice & FX is where you change how it sounds, a phone voice or a big hall, or leave it clean. Mix is where you set how loud it is next to the original before it goes into the game. Each step shows only the controls it needs, so the screen never fills up with everything at once.

CaptureEditVoice & FXMix

Capture. Press Record, wait for the countdown, say the line. Every try is kept.

Left to right, in the order you work. Each step shows only its own controls.

Left to right, in the order you work. You can jump to any step; the order is a suggestion, not a rule.

The tools

The tools down the left are the same as the video editor's, minus Link, since a game's line is already linked. Select (V) picks and moves takes. Hand (H) drags the timeline to another part. Split (S) cuts a take in two where you click, and deletes nothing by itself. Volume (G) shapes how loud a take or the dubbed-over original is over time, point by point: click the line to add a point, drag a point to change the level there, or drag the line between two points to move just that stretch up or down. The rest of the line stays where it is, and drops straight to the moved stretch at each end. Hold Shift while dragging to move all of it at once, which is the quick way to turn a take down. Record (R) starts and stops recording on the armed track. Undo takes back any of them.

Speaking a line instead of recording it

Not every line needs a performance, and not every translator has a microphone. Where you see AI voice, you can pick a voice, hear it, and have it read the line for you. It is in three places, and it puts the result where that place expects it:

WhereWhat using it does
Audio pageBecomes the clip's dub, the same as a recording would. The clip counts as dubbed and the build ships it.
Sound editorLands on the timeline as a take, so you trim it, shape its voice and draw its volume exactly as you would your own. A dub made on the Audio page shows up here too: opening the editor afterwards brings it in as a take called Current dub and mutes the earlier takes, so what you hear is what ships.
Studio editorLands in the take library, ready to place on the timeline and split.

The first choice is how good a voice you want, and it is three cards with a price each:

TierWhat it gives you
BasicClear and quick, in almost every language, and free. Enough for a first pass through a script, or to hear how a translation flows, however many lines that is.
StandardMore natural delivery with some character to it, in any language, from two credits a line.
ProReal intonation in over seventy languages, Georgian among them, from four credits a line. It also takes a direction: a short note on how to read the line (warm, tired, whispering) goes in beside the words. A language it does not read greys the card and says so.

The language is chosen for you: it is the language your adaptation translates into, and the panel says which voices it opened on. A tier that has no voice in it is greyed with the reason. If the app could not tell which language to use, the panel asks, once, and remembers.

Voices are listed with a male or female mark and a play button that says a short sample in your language, or in the voice's own where it has one. A long list opens folded to the first few, with a button for the rest; the voice you chose is always shown. The tier, the voice and the direction are remembered on this machine, so none of them is typed twice. Listening costs nothing, however many you try: the samples are made once and kept, so everybody hears the same file, and they are ready the moment the list appears. The voice you pick is remembered per tier and language, so it is a choice you make once rather than on every line.

The price is on the button. Speech is billed on the length of the text, which is known before anything runs, so Speak · 2 credits is what you are charged, not an estimate. Basic reads Speak · Free and charges nothing at all; the paid tiers cost at least their floor, and a line that fails costs nothing.

Once spoken, the line plays back and a single button does the one thing that place is for: make it the clip's dub on the Audio page, or add it as a take in the editors. The panel then says it was done, and on the Audio page the player switches to the dub at once. Anything an AI voice spoke keeps a small AI mark wherever takes are shown: the clip's badge on the Audio page, its take in the editor, its card in the Studio. Hover it to see which voice.

The words it speaks are the translated ones, and you can edit them in the box before pressing: a name that the voice mispronounces is often fixed by spelling it the way it sounds. That edit changes only what is spoken, never the translation itself.

Saving, and closing

The sound editor saves itself. It writes your session every thirty seconds and again the moment you leave, so a take is never waiting on you to remember. The word beside the game's name says which state it is in: autosaved when everything is written, and unsaved for the stretch between one save and the next.

Close while it reads unsaved and the editor asks first, because there are two sensible answers. Save and close keeps what you have just done. Close without saving goes back to the last save, which can be up to thirty seconds old, and is the way out when a take went wrong and you would rather it had not happened. Keep editing puts you back where you were. With nothing waiting to be written, closing simply closes.

Keeping the original under your voice

Some game files are one mix: the line, the music and the effects together. If you replace such a file with your take alone, the music goes with the old voice. Press Dub over on the Original row and the sound editor keeps the original underneath on a track called Dub over, the way a dubbed film keeps its score and its footsteps under the new voices. It is an ordinary track: cut it, mute it, set its volume. The monitor's With dub over button plays the bundle the way it will ship.

To make room for your voice, shape that track's volume line with the Volume tool (G), the same line every take has: click the line to add a point, drag a point down where your voice speaks and the original sinks there, drag it back up where the original should be heard in full. The slope between two points is how quickly it changes, so two points close together make a quick dip and two far apart a gentle one. Nothing is decided for you. Remove dub over goes back to replacing the original outright.

Going back to an earlier version

The sound editor keeps a version of your session after every change you make, and it keeps them for 30 days. Undo only reaches back as far as the editor has been open; this reaches back days. Open it with the History button at the bottom of the tool rail, under undo and redo.

The panel lists your versions by day, newest first, with what was done and at what time. The one at the top is the session as it stands now. Click any other and it appears on the timeline, marked at the top as a version you are only looking at: playback works, but nothing you do can change it. Restore this version makes it the session again, and Back to current leaves it alone. Escape does the same as Back to current.

Restoring is not destructive. The version you were on before is still in the list, so restoring the wrong one costs you a click, not a take. Your recordings are never deleted by a restore either: the WAV files stay on disk, so a version from last week still plays.

This history is private to this computer. It is never uploaded, never shared with the rest of the adaptation, and never shipped in a build.

Bookmarks, for coming back to a clip

A take you want to redo, a line whose delivery you are unsure about, a clip you need to check against the scene: bookmark it from the player at the bottom of the page and it joins the Marks tab beside the labels. Clicking it there loads the clip again.

Like the ones on the Text page, these are private to this computer. They are never uploaded and never shared with the rest of the adaptation.

The Marks tab on the Audio page. Bookmarks are listed newest first and show the clip's filename.

Partial dubs are normal

Dubbing every line of a big game is a massive job, often thousands of lines. Coverage is counted per character, so you can see exactly where the effort went and where it still has to go.

Voice coverage3,148 / 3,944 dubbed80%
  • Mira1,840 / 1,840
  • The Keeper902 / 1,204
  • Dockhand118 / 612
  • Narrator288 / 288
Dubbing coverage by character, growing over a few sessions. Most teams finish one voice at a time rather than spreading thin, because a fully dubbed lead reads as a real dub while four half-dubbed characters read as broken.

Many teams dub only the main story first and ship the adaptation as Text + partial audio. That's completely fine, players appreciate it, and the game page's voice coverage numbers set expectations honestly. You can keep recording and publish the rest as updates.

Voice coverage shows up on your game page automatically, so players always know what to expect before installing. And if what you want to dub is a video (a trailer, a cutscene, a walkthrough) rather than in-game lines, that's a different room: Studio is built exactly for that.