One Long Video, Three Aspect Ratios: A Working Order for Cutting Social Variants

By Mega Deal Team

video repurposing captions social video content operations workflow

The problem is sequence, not software

One sixty-minute recording, three placements to feed. The tooling is cheap; what costs you the afternoon is running the operations in the wrong order and paying for the same correction three times.

Automated captions: what breaks

Recognition fails in predictable places. Knowing the categories lets you scan for them rather than read every line equally.

  • Proper nouns and brand names. The most damaging class: a mangled company name in burned-in text is what viewers screenshot. Anything invented or built from two words jammed together comes back wrong.
  • People's names. Speaker and guest names get normalized toward common spellings; every non-Anglophone name is a coin flip.
  • Spoken numbers. Currency, versions, dates, percentages, phone numbers. "Two point four" lands as "2.4" in one card and as words in the next.
  • Jargon and acronyms. Domain vocabulary sits outside the model's high-probability space; acronyms spelled letter by letter arrive as unrelated words.
  • Homophones. These survive review, being real words in grammatical positions: their/there, principal/principle, discrete/discreet.
  • Overlap and crosstalk. With two people talking at once, the transcript drops one, or interleaves fragments into a sentence nobody said.
  • Accented speech. Error rates climb unevenly across speakers, so a solo clip may be clean while the panel segment is a mess.
  • Filler words. Transcribed literally, "um" and false starts eat the card and make a competent speaker read as unsure.
  • Segmentation. The machine breaks on silence, not syntax, splitting a clause across two cards so the punchline lands with no setup.

A three-pass review method

Before transcription, build a term list — every product name, person name, acronym, and piece of domain vocabulary in the recording. Ten minutes with the deck gets most of it.

Pass one: find-and-replace against that list. Search each known-risk term and its plausible mishearings. Mechanical, fast, and it clears the worst class before you read anything in context.

Pass two: read through with the audio muted. The step people skip and the one that works. Listening while reading, your brain reconciles text against what you know was said: you hear the right word, see the wrong one, register no conflict. Muting removes that correction, and homophones, dropped negations, and broken segmentation surface immediately.

Pass three: timing. Audio back on, checking only that cards land with the speech and hold long enough to read. Fix breaks that split a clause. Delete filler rather than time it.

Reframing to vertical and square

Vertical is 9:16, square is 1:1, your source is almost certainly 16:9. Reframing is an editorial decision, and subject tracking makes it without knowing what the shot was for.

  • Two-shot interviews. A vertical crop holds one face. Where the segment's value is the exchange — reaction, interruption, disagreement — cropping to one person destroys it, and tracking that swings between speakers is worse.
  • Lower-thirds and graphics. Built for a horizontal frame, they crop at the edges or sit under interface elements. Names and titles must not be half-covered.
  • Screen-share and slides. Cropped to 9:16 the content is unreadable; scaled to fit, the text is too small for a phone.
  • Burned-in source text. Titles rendered into the master at horizontal proportions cannot be repositioned. Crop them or do not reframe.
  • Tracking swing. Tracking crops follow movement, so a speaker who gestures makes the frame lurch, reading as an operator losing control.
  • Safe-area collisions. Interface overlays occupy the bottom and one side of vertical placements, and those zones move.

Reframe, letterbox, or leave it horizontal

Reframe when one subject carries the segment, stays roughly still, and no graphic is load-bearing. Letterbox — source at full width in a vertical canvas, empty bands carrying a headline and captions — when the composition must stay intact: two-shots, legible slides, wide demos. It is treated as the lesser choice and it is not; a letterboxed clip you can read beats a cropped one with the point off-frame. Do not go vertical when the segment needs dense screen content, more than two people on camera, or graphics you cannot rebuild.

Where silence removal stops working

Automatic cut detection removes silent stretches. Having no model of meaning, it removes things that were doing work.

  • Meaningful breaths. The pause before a difficult answer is content. Strip it and the answer arrives with no weight.
  • Comic and rhetorical timing. A beat before a punchline is a beat, not dead air.
  • Clipped leading consonants. Cutting to the first detected sound of the next word shaves the plosive off the front, leaving a mushiness you hear as bad audio without being able to name it.
  • B-roll and music beds. Detection keyed to the voice track fights a continuous bed underneath, leaving audible seams.
  • Texture mismatch. Aggressive settings yield a dense jump-cut rhythm: native register on short-form vertical, unfinished on a customer-facing demo.

Set thresholds conservatively — longer minimum silence duration, generous padding either side of every cut. Let automation take the obvious two-second stalls and leave ambiguous ones alone: deleting a stall you left in takes seconds, while restoring timing the machine ate means returning to the master. A manual pass is non-negotiable on the first and last seconds of every clip, on B-roll transitions, and on any moment you flagged as a laugh or a deliberate pause.

Work in this order

  1. Select clips from the master. Mark in and out points, decide which placements each clip can serve.
  2. Lock the edit of each clip. Trims, internal cuts, pacing. Lock means locked.
  3. Transcribe and correct captions once, at the highest-quality source. One correction pass, one canonical caption file per clip.
  4. Reframe per placement. You now know which seconds ship, so you crop only surviving footage.
  5. Burn in captions, sized per placement. Same corrected text, different type size and safe-area position per aspect ratio.
  6. Export per placement.

Invert three and four and you transcribe three reframed variants instead of one master — finding the same mangled name three times, fixing it correctly in two. Invert five and four and you burn text at horizontal proportions into a frame you are about to crop. The rule underneath: anything expensive to redo happens once, as early as dependencies allow.

End to end

Log timecodes for candidates, marking which are single-subject, which carry screen content, which are exchanges. Cut and lock each clip. Then transcribe — this is where a browser-based editor with subtitling, such as VEED, fits, since trimming and generating a subtitle track in one tab avoids a round trip through a desktop suite for work this small. Run the term list, the muted read-through, then timing. Reframe per placement, burn captions sized per aspect ratio, export.

Name variants so lineage stays legible: project_topic_clipNN_ratio_vNN, for example q3webinar_onboarding_clip03_9x16_v02. The clip number ties back to your timecode log, the ratio makes wrong-placement uploads obvious, the version tells you which caption generation you have. Bump the version on every caption change and re-export each ratio together, so the three never drift.

Review checkpoint before publishing

Per variant, on a phone, sound off: cards inside the safe area and clear of interface overlays; no term from your list misspelled in burned-in text; no clause split across two cards; subject in frame with no tracking swing; first and last seconds intact; graphics legible at phone size. Anything that fails goes back to the step that owns it — caption errors to the corrected master file, never to the variant.

Back to Blog