Audio and video
Upload a recording to a workflow or a chat and have it transcribed into text your agents can reason over
Overview
Audio and video uploads are transcribed into text automatically. From that point the transcript behaves exactly like text extracted from a PDF: it is stored, word-counted, and either inlined into the agent's context or retrieved from a knowledge base, following the same size rules as any other document.
This works for files attached in a chat, and for audio inputs on a START node.
Transcription happens at upload time, not at run time. By the time an execution starts, the transcript already exists, so a workflow never waits on speech-to-text.
Supported formats
| Kind | Extensions |
|---|---|
| Audio | mp3, wav, m4a, aac, ogg, flac, webm |
| Video | mp4, mov, mkv |
Video is accepted even though no transcription provider consumes video directly. The platform strips the video track and sends only the audio, which is what lets someone drop a screen recording straight onto a START node.
Setting it up
Transcription runs on an LLM config, and there are two requirements:
Tick Supports Audio on the config
In the LLM config, enable Supports Audio.
Give it a model that can actually transcribe
The model on the config is the transcription model. There is no global default and no separate transcription setting to fill in.
Which model you choose decides how the audio is sent:
| Model on the config | How it transcribes |
|---|---|
whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe | Dedicated transcription endpoint |
gpt-4o-audio-preview family | Chat completion with an audio part |
| Any Gemini model | Gemini inline audio |
A plain chat model such as gpt-4o | Refused |
A plain chat model on an audio-enabled config is a misconfiguration, and the platform treats it as one. It is rejected when you save the config, and again before any provider call, rather than failing partway through a workflow. If a config will not save with Supports Audio ticked, the model is the reason.
Endpoints are derived from the config's existing base URL, so there is no extra transcription endpoint to keep in sync.
How a recording is processed
Duration is measured first
The clip's length is probed before anything is sent to a provider, so a file over the cap is rejected without incurring cost.
Audio is extracted and compressed
Any video track is dropped and the audio is converted to mono 16kHz, roughly 4KB per second. This keeps long recordings cheap to move.
Long recordings are split
Anything longer than the chunk window is split into chunks and transcribed piece by piece.
Chunks are joined
The pieces are joined back into a single transcript, which then follows the normal file pipeline.
Limits
| Setting | Default | What it controls |
|---|---|---|
max_media_minutes | 120 | Hard cap on clip length, checked before any provider call |
stt_chunk_seconds | 600 | How long a chunk is before the recording is split |
stt_max_output_tokens | 32000 | Output cap when transcribing through a chat-audio model |
Long clips on a chat-audio model can hit the output cap. When a chat model produces the transcript, the text arrives as generated tokens, so a long recording can exceed stt_max_output_tokens. The platform fails the transcription in that case rather than returning a transcript that is quietly missing its tail. If you hit it, raise stt_max_output_tokens or lower stt_chunk_seconds. Do not work around it by ignoring the error, because the failure mode it prevents is silent data loss. The dedicated transcription endpoint does not have this problem.
Seeing it in the execution timeline
Each media input gets its own MEDIA row in the execution timeline, reporting the engine used, the clip length, how many chunks it was split into, the size of the transcript, and any token usage the provider measured.
The row exists because the work happened at upload time. Without it, a run would show a START node that mysteriously already had text, with nothing accounting for where it came from.
Token usage is only shown when the provider actually reported it. It is never estimated from the transcript's length or the clip's duration, so a missing figure means "not measured" rather than zero.
What happens to the transcript
Once the text exists, the standard input rules apply:
- Short transcripts are inlined directly into the agent's context.
- Long transcripts are indexed into a per-file knowledge base, and the agent retrieves the relevant passages instead of reading the whole thing.
The threshold and the retrieval behaviour are the same as for any large document. See Input limits for the caps that apply per execution.
Use cases
Meeting notes
Drop a recording on a START node. The transcript becomes the input to an agent that pulls out decisions, owners, and follow-up actions, then writes them somewhere.
Support call review
Transcribe a call, classify it with a Condition node, and route anything flagged as an escalation to a Human Task.
Screen recordings
A recorded walkthrough uploads as mp4, has its video track stripped, and the narration becomes searchable text.
Next steps
Memory and variable store
Persist and access data within and across workflow executions using MagOneAI's variable store
Execution artifacts and the Debugger
See exactly what an agent did, what was sent to the model, and what it recalled: durable per-run artifacts, the action ledger, blackboard recall, and captured prompts