Features

Everything Burnin does

Grouped the way a video happens: import it, transcribe it, edit it, export it. Everything on this page works with no account.

On the phone, with no signal

Transcription runs on the phone

Burnin does not send your video anywhere. Apple's on-device speech transcriber turns speech into text right on the iPhone, which is why captioning keeps working on a flight, underground, or with the data switched off.

Language packs, downloaded once

The first time you transcribe in a language, its model downloads. After that it stays on the phone, and every video in that language transcribes with no connection at all.

The speaker model, downloaded once

Speaker diarization uses an on-device FluidAudio model, around 100MB, that downloads once and then runs fully offline for every video afterward.

Automatic language detection

Burnin detects which language is being spoken on its own. A manual picker is there for when you would rather choose, or when a video mixes languages in a way detection cannot resolve on its own.

Nothing about the transcription, diarization, editing or export needs a connection. The one thing that does is the first-time download of a language pack or the speaker model, and each of those happens only once per model.

Transcription and captioning

Speech to readable caption chunks

A raw transcript is automatically segmented into caption-length lines, so what appears on screen reads at a glance instead of running on as one long sentence.

Gap detection that checks before it retries

Burnin looks for spans of the video that likely got dropped during transcription. It only re-transcribes a span when an independent speaker-diarization pass confirms real speech overlaps it, so it never invents captions over silence.

Speaker changes, labeled automatically

Diarization works out who is speaking when. Every caption where the speaker changes gets a "- " prefix, so a two-person conversation reads like one without any manual labeling.

Both models run detection and confirmation on-device

Segmentation and gap detection need no network signal at all, and the confirmation step that decides whether to retry a gap runs on the same on-device diarization pass used for speaker labeling.

How transcription and gap detection actually work

The caption editor

A full editor sheet for every caption

Tap a caption, or the add button floating over the list, to open a large editor sheet at full height from the first frame, with its own delete control and a "Done" to close it.

A new caption lands where you are

Add a caption and it is inserted at the video's current position, in chronological order alongside its neighbors, not always tacked onto the end of the list.

Drag-trim on a visual timeline

Each caption's start and end time is a handle on a timeline strip. Drag either edge to retime it against the video underneath, with a margin above and below every block so there is always room to start a pan.

A draggable playhead

Drag the playhead on the timeline to scrub the video directly, instead of only watching it play.

Neighbor-clamping

Dragging one caption's edge past its neighbor is not allowed. Captions cannot be made to overlap by accident.

Undo and redo, deep

Every text edit and every trim is on the undo stack, twenty to thirty steps back, so an experiment you don't like is one gesture from gone.

Auto-save, debounced

Edits are saved automatically roughly every 600ms as you make them. Closing the app mid-edit, or a call coming in, does not cost you the work.

Live preview with captions burned in

The video preview shows exactly what the exported file will look like, captions and all, updated as you edit.

Title and description, before you share

Written from your own transcript

Right after export, before the share sheet opens, Burnin offers a title and a short description with a couple of hashtags, generated from the captions you just edited, not re-transcribed from scratch.

Three tones to pick from

Neutral, balanced, or clickbait. Switching tone regenerates the title and description to match, and you can skip the screen entirely and export with just the video.

Apple Intelligence, on the device

Generation runs through Apple's on-device Foundation Models when available, so the transcript never leaves the phone to be summarized.

A fallback with no capability gate

On a device without Apple Intelligence, or if generation fails, an on-device heuristic takes over: it picks a named person, place, or organization from the transcript for the title, and the single most information-dense sentence for the description, never just the first line.

Crop and export

Four crop options

Original, 9:16 for Reels, TikTok and Shorts, 1:1 for square feed posts, or 16:9 widescreen. Pick the shape the destination actually wants.

Three export resolutions

720p, 1080p or 4K, with captions burned directly into the exported file rather than added as a separate track.

Captions burned in, not a subtitle file

The export is a single video file with the captions part of the picture, so it plays correctly anywhere without needing a subtitle track to be attached separately.

Projects

Every video becomes a project

Import a video and it becomes a project in your library, with its captions and edits kept alongside it.

Rename and delete

Rename a project to find it later, delete one you no longer need, or select several at once to clean up in bulk.

Recovers a missing source file

If the original video file goes missing locally, Burnin can recover it from your Photos library instead of leaving the project broken.

Stored on your phone, not a server

Projects are stored locally, on-device. There is no cloud sync and no account for them to belong to.

Requirements

PlatformiPhone and iPad (not iPad-optimized yet), iOS 18.0 or later
On-device transcriptioniPhone 12 and later. Older devices, and the Simulator, are not supported
Video formatsMP4, MOV
Max video length10 minutes
AccountNone. Not required, not offered
PriceNo ads, no subscription
One thing captions are not

Automatic transcription is not perfect on any device, on any app. Burnin's editor exists because every caption is meant to be checked, not accepted blindly, and gap detection is a heuristic that reduces missed speech rather than a guarantee that none is ever missed.