Ask an AI about this

Studio lets you dub over a video: bring in a clip, let Tranzio read its subtitles off the picture, and record your voice on top of them, line by line.

Studio: matching video & dubbing

Studio lets you dub over a video: bring in a clip, let Tranzio read its subtitles off the picture, and record your voice on top of them, line by line.

Studio answers a different question than the rest of the Workspace. Audio Lab dubs the voices inside a game; Studio dubs a video: a trailer you want to release alongside your translation, a cutscene, a tutorial, a walkthrough. You bring in a clip, and Studio lays out a video's subtitles and audio on a timeline so you can record your voice over it, line by line.

What Studio is, and who can use it

Game audio arrives as thousands of separate lines with no picture and no context. Studio plays a video of the game and lays its lines on a timeline, so you dub each one against the scene. A plain audio editor knows nothing about which line you are recording; a live translation overlay can only put text on the screen. Studio links every take to the game's own line, so the next build puts your voice where the original was.

Studio comes with the Pro and Studio plans, which start with a free trial you can cancel any time. On the Free plan the sound editor on the Audio page still records one line at a time.

Bringing a video in

  1. Pick a video

    Paste a YouTube link, or drop a file of your own. Either way Tranzio only asks the host for the name and the length. The subtitles come from the picture, so a link and a file behave the same from here on.

    youtube.com/watch?v=hollowtide-endingFetch
    From YouTube

    Hollowtide, the quiet ending

    Duration
    8:12492s
    Audio
    Stereo48kHz
    Captions
    EN146 lines

    Each caption becomes a timed line on the timeline, ready to translate.

    Pasting a link and fetching it. The tiles are what you get back: how long the clip runs, its audio, and where the subtitles will come from.

    Drop a video here

    or click to browse
    MP4, MOV, and WebM work best
    The drop area for your own clips. MP4, MOV, and WebM all work.
  2. Choose what comes with it

    Before the editor opens you decide whether to read the subtitles from the picture. Leave it on and Tranzio watches the video and turns every line it finds into a timed entry waiting to be translated and voiced. How long it takes depends on how much video there is, and the row says so before you agree to it: a few minutes for a scene, several hours for a whole playthrough read over a link. That is what the focus range below is for, and you can turn the reading off and do it later. You can also focus on part of the video by setting a start and end time, which is what you want when a two-hour walkthrough holds one ending you care about. Focusing hides nothing permanently; widen the range later and the rest comes back.

  3. Open in the editor

    The editor opens with everything lined up: the video on top, the lines that were read and your recordings below it on a shared timeline.

Tranzio never keeps the video. It plays from YouTube where it lives. Reading a link's subtitles does stream it through, but nothing is written to disk except small crops of the subtitle band, and each of those is deleted the moment it has been read. What is left at the end is the text. Import the same link twice and you get two separate sources, so you can focus one on the intro and the other on the ending.

Reading the subtitles off the picture

Games burn their subtitles into the video rather than shipping a caption track, so there is usually nothing to download. Tranzio watches the bottom of the picture instead and reads each line it finds, along with the moment it appeared and the moment it went away. That is where the timings come from, so they are measured rather than guessed, and it works the same on a link and on a clip you recorded yourself.

Subtitles read
Hollowtide, the quiet ending
492s
video read
214
frames read
38
lines found

38 lines found, 29 matched to the game's own text. The rest wait in the unmatched queue.

A read in progress. Watch the frame count climb far more slowly than the video position: a subtitle sits still for seconds at a time, and every one of those repeated frames is skipped.

The box says what is a subtitle; the reader looks at everything.Text inside the box that matches one of the game’s spoken lines becomes a line on the timeline, with the game’s own recording next to it. The rest of the picture is read too, on a slower rhythm: a menu, a chapter title, an objective or a loading screen that matches the game’s text is photographed for the Text page, so you can see where a string is used, and never becomes a line. A string keeps at most three pictures, never two of the same moment. The frame test in the region editor shows both: what the box read, and what the rest of the picture turned out to be.

It reads only the stretch you focused on. A playthrough recording runs for hours and almost never needs reading end to end, so setting a start and end on the import screen is the difference between a couple of minutes and most of an afternoon. A local file reads several times faster than a link, because a link is being fetched as it goes.

When it finishes, the lines are matched against the game's own text automatically. Anything that matched nothing is still placed on the timeline and waits in the unmatched queue in the editor, where you can attach it by hand.

Each line also gets the original delivery attached. Where a line matched a string the game has audio for, it gets the game's own recording, which is the line by itself. That covers every way a game links the two: a clip recorded for one line, a single clip covering a whole conversation, and a game whose clips were linked against a different language than the one being translated, which is commoner than it sounds. Everything else is cut out of the video using the times the read worked out, which carries whatever music and gunfire were playing over it but is still far better than nothing. Either way you can hear how a line was said before recording your own, and neither one asks you to attach a file by hand. A video with no audio track simply gets no clips, and those lines work as they always did.

Those clips stay on your machine. They are the video's audio rather than anything you made, so they are stripped out of an adaptation before it is published, along with every other local path.

Stopping is safe, and you can carry on later. The dialog keeps every line found up to that point and remembers where it got to, so the source shows a Continue button with the time it reached. Pressing it starts again from there and adds to what is already on the timeline rather than replacing it. That is what makes a long recording workable: read a scene, stop, come back to it.

Every imported source also carries a read button, which asks whether to carry on from where the last read stopped or to read the whole focused stretch again. The second replaces the lines already read from that video; the first only adds. Both say roughly how long they will take before you choose. Nothing needs deleting and re-importing to be read a second time.

While it runs the dialog shows the actual strip of the picture being read and the words got out of it. That is the quickest way to tell a working read from one aimed at scenery, and it is worth a glance in the first few seconds. The matched count beside it climbs as it goes: every line is looked up in the game's own text the moment it is read, rather than all at once at the end.

Each answer is remembered for that video, so a subtitle held on screen for four seconds is looked up once rather than twenty times, and a phrase that comes back an hour later is not looked up again at all. That record is kept beside the video's clips and is deleted with the source, which is also why carrying on a stopped read is cheap: it does not re-ask anything the first attempt already answered.

If a game puts its subtitles somewhere else, say so first. The reader looks at the bottom middle, which is right for nearly every game and useless for the rest. The read chooser on a source has a Set where the subtitles are option: it shows a frame with a box on it, you drag that box over the words (its middle moves it, its corners resize it), and Test this framereturns the exact strip the reader would see, the words it got, the game’s own wording of the line those words matched, and its recording to play back. Reading the two lines side by side is what catches a region one line too high, which can still match something. Use the scrubber to land on a frame that actually has a subtitle on it, since a silent shot proves nothing either way; it covers the stretch the source is focused on, because that is the only part that will be read.

The reader keeps a picture of each line as well.The moment a line matches one of the game’s strings it keeps one frame of the video where the line is said, so the Text page and the Audio page can show the moment beside the string: who says it, to whom, in what room. Beside the pictures, each video the line was found in plays right there from that moment, in the editor’s own player, for as long as the video is still in Studio. One picture per line per video, however long the line stays up and however many times the same video is imported; a different playthrough adds another. They are kept sealed with the rest of the vault and open in a gallery from the line’s row or its rail.

The same editor can blank outparts of the picture. Games put ammo counters, objective markers and keybind prompts in the same strip as their subtitles, and a recording often carries the channel’s watermark on top of that. All of it reads as text. Switch to Blank out, drag over anything that is not dialogue, and it is painted over before the picture is read, which is more certain than filtering its words out afterwards. The rectangles are saved with the region and listed so you can remove one.

The first read downloads the reading models, which takes a moment and happens once per language. After that everything runs on your own machine, so no frame of the video is ever sent anywhere.

The editor

The editor pairs the video with a multi-track timeline. Each caption becomes a line you translate and voice, and your recording sits on the lane underneath, split into chunks, one per line. Seeing them stacked is the whole point: you can tell at a glance whether your delivery lands where the original did.

0:000:040:080:120:16
Original
It never rains here.
Not since the harbour closed.
Come inside, you're soaked.
Your take
Take 1
Take 1
Take 2
Two lanes under one ruler: the original captions on top, your take below. The playhead sweeps both at once. The dashed amber chunk is one that hasn't been matched to a line yet.

A bar across the top walks you through the four steps in order, and the editor changes its advice as you move along it:

  1. Watch

    Play the clip end to end first. Get the rhythm, the pauses, the order the lines come in. Scrub the ruler to jump around.

  2. Record

    Record all your lines in one continuous take rather than one line at a time. The original audio ducks while you record so you can hear the timing you're matching, and gaps are fine because the next step splits on them.

  3. Split

    Your take gets cut into chunks at the silences, one chunk per line. Review the cuts. To add one, scrub to the moment and press ⌘ + B (or the Split at playhead button on the rail), or take the split tool and click a chunk: the blade snaps to the playhead and shows how long each piece will be. Merge two that shouldn't have been separated.

  4. Match

    Point each chunk at the line it replaces. This is what turns a long recording into a set of dubbed lines.

Matching chunks to lines

Matching is the step that makes the rest count, and it's mostly automatic. Auto-match walks every unmatched chunk and links it to the nearest caption in time, which is right almost every time when you recorded along with the video. What it gets wrong you fix by clicking a chunk and then clicking the line it belongs to, or by dragging the chunk onto it. The whole auto-match is one undo, so trying it costs nothing.

Match progress4/4
CtrlK
  • 0:04It never rains here.Take 1
  • 0:09Not since the harbour closed.Take 1
  • 0:14Come inside, you're soaked.Take 1
  • 0:19I'll put the kettle on.Take 1
Auto-match linking a take's chunks to their lines. The counter at the top is the number worth watching: a line with no chunk is a line nobody has dubbed yet.

Lines are yours to reshape, too. Split one that runs too long for a single breath, merge two that are really one sentence, or add a line by hand when the captions missed something. Short lines are the secret; nobody dubs a paragraph in one breath.

Attaching the original clip

A line can carry the game's own recording of it, which is what the A / B waveforms in the inspector compare: the original above, your take below. Press Attach original audio and a search opens over every clip the game scan found.

Search matches what is said first, then event, bank and file names, so a few words of the line usually finds it. When Tranzio already ties a clip to this line it is pinned to the top and marked Suggested. Arrow keys move through the results and Enter attaches, so a scene can be linked without leaving the keyboard, and the filters narrow to the suggestion or to clips nothing is linked to yet.

Nothing to search means the game's audio has not been scanned. Run a scan from the Audio page first, and the clips appear here.

The tools down the left

Six tools, one at a time. Select is the normal pointer: click a piece to pick it, drag it to move it. Hand drags the timeline itself, so you can look at another moment without moving anything. Split cuts a take in two where you click, so you can keep the good half; the cut deletes nothing. Volume shapes how loud a take is over time, point by point. Link ties a piece of video to the line of text it shows, and the number on it is how many pieces are still unlinked. Record records your voice from where the playhead is, after a short countdown. The letter on each tool is its key.

VHSGLR

Select V

Click a piece to pick it. Drag it to move it.

One tool is active at a time. The letter is its key.

Each tool lit in turn, with what it does. You never need more than one at a time, and the letter under it switches to it from the keyboard.

Above the tools sit Undo and Redo. They take back or put back the last cut, move, link or recording. If a tool did something you did not expect, undo first and read this page second.

Play, pause and move around

The bar along the bottom is how you move through the video. The big button plays and pauses (so does Space). The two arrows beside it skip ten seconds back or forward. Loop scene plays the current line's moment over and over, which is the easiest way to rehearse before you record. Speed slows the video down to catch fast speech or speeds it up to skim; recording always runs at normal speed. The clock shows where you are and how long the video is.

Takes: every attempt, kept

Each time you record a line, the recording is kept as a take. You can record the same line five times and pick the best one; the others cost nothing and sit in the Takes list until you delete them. Play each take, keep the one you like, and that is the one that goes into the game.

Line 214 · MELINA_GREET_01

ああ、褪せ人よ。黄金樹のふもとで、ずっと待っていた。

  • Take 1
    0:03
  • Take 2
    0:03
    Kept
  • Take 3
    0:04
One line, three takes. The kept one is marked; the others stay until you delete them.

Editor shortcuts

Recording goes much smoother when play, pause, and record are on keys. The full list lives in Keyboard shortcuts; these are the ones you'll use every minute:

ShortcutAction
SpacePlay / pause
RStart or stop recording
← / →Skip back / forward 5 seconds
Ctrl+KAuto-match every unmatched chunk
Ctrl+BSplit the chunk under the playhead
V / H / S / LTools: select, hand, split, link
Ctrl+ZUndo

Going back to an earlier version

Studio keeps a version of this video's document, the chunks, the lines they are matched to and the scene markers, after every change you make, and keeps them for 30 days. Undo only reaches back as far as the editor has been open; this reaches back days. Open it with the History button at the bottom of the tool rail, under undo and redo.

The panel lists your versions by day, newest first, with what was done and at what time. The one at the top is the video as it stands now. Click any other and it appears on the timeline, marked at the top as a version you are only looking at: playback works, but nothing you do can change it. Restore this version makes it the video again, and Back to current leaves it alone. Escape does the same as Back to current.

Restoring is not destructive. The version you were on before is still in the list, so restoring the wrong one costs you a click, not a take. Your recordings are never deleted by a restore either: the audio files stay on disk, so a version from last week still plays.

This history is private to this computer. It is never uploaded, never shared with the rest of the adaptation, and never shipped in a build.

Rule of thumb for which tool: Studio for a video, Audio Lab for the game. If the thing you're dubbing has a play button and a duration, it's Studio. If it's spoken lines the game triggers while playing, it's Audio Lab.

The two rooms do meet in one place. If a line in Studio is attached to the game's own audio file, the chunks you matched to it are mixed down and packed into the patch when you build, under the step called Compiling studio recordings. So a cutscene dubbed here can ship inside the game, not just alongside it.