Workflow Builder

Audio and video

Upload a recording to a workflow or a chat and have it transcribed into text your agents can reason over

Overview

Audio and video uploads are transcribed into text automatically. From that point the transcript behaves exactly like text extracted from a PDF: it is stored, word-counted, and either inlined into the agent's context or retrieved from a knowledge base, following the same size rules as any other document.

This works for files attached in a chat, and for audio inputs on a START node.

Transcription happens at upload time, not at run time. By the time an execution starts, the transcript already exists, so a workflow never waits on speech-to-text.

Supported formats

KindExtensions
Audiomp3, wav, m4a, aac, ogg, flac, webm
Videomp4, mov, mkv

Video is accepted even though no transcription provider consumes video directly. The platform strips the video track and sends only the audio, which is what lets someone drop a screen recording straight onto a START node.

Setting it up

Transcription runs on an LLM config, and there are two requirements:

Tick Supports Audio on the config

In the LLM config, enable Supports Audio.

Give it a model that can actually transcribe

The model on the config is the transcription model. There is no global default and no separate transcription setting to fill in.

Which model you choose decides how the audio is sent:

Model on the configHow it transcribes
whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribeDedicated transcription endpoint
gpt-4o-audio-preview familyChat completion with an audio part
Any Gemini modelGemini inline audio
A plain chat model such as gpt-4oRefused

A plain chat model on an audio-enabled config is a misconfiguration, and the platform treats it as one. It is rejected when you save the config, and again before any provider call, rather than failing partway through a workflow. If a config will not save with Supports Audio ticked, the model is the reason.

Endpoints are derived from the config's existing base URL, so there is no extra transcription endpoint to keep in sync.

How a recording is processed

Duration is measured first

The clip's length is probed before anything is sent to a provider, so a file over the cap is rejected without incurring cost.

Audio is extracted and compressed

Any video track is dropped and the audio is converted to mono 16kHz, roughly 4KB per second. This keeps long recordings cheap to move.

Long recordings are split

Anything longer than the chunk window is split into chunks and transcribed piece by piece.

Chunks are joined

The pieces are joined back into a single transcript, which then follows the normal file pipeline.

Limits

SettingDefaultWhat it controls
max_media_minutes120Hard cap on clip length, checked before any provider call
stt_chunk_seconds600How long a chunk is before the recording is split
stt_max_output_tokens32000Output cap when transcribing through a chat-audio model

Long clips on a chat-audio model can hit the output cap. When a chat model produces the transcript, the text arrives as generated tokens, so a long recording can exceed stt_max_output_tokens. The platform fails the transcription in that case rather than returning a transcript that is quietly missing its tail. If you hit it, raise stt_max_output_tokens or lower stt_chunk_seconds. Do not work around it by ignoring the error, because the failure mode it prevents is silent data loss. The dedicated transcription endpoint does not have this problem.

Seeing it in the execution timeline

Each media input gets its own MEDIA row in the execution timeline, reporting the engine used, the clip length, how many chunks it was split into, the size of the transcript, and any token usage the provider measured.

The row exists because the work happened at upload time. Without it, a run would show a START node that mysteriously already had text, with nothing accounting for where it came from.

Token usage is only shown when the provider actually reported it. It is never estimated from the transcript's length or the clip's duration, so a missing figure means "not measured" rather than zero.

What happens to the transcript

Once the text exists, the standard input rules apply:

  • Short transcripts are inlined directly into the agent's context.
  • Long transcripts are indexed into a per-file knowledge base, and the agent retrieves the relevant passages instead of reading the whole thing.

The threshold and the retrieval behaviour are the same as for any large document. See Input limits for the caps that apply per execution.

Use cases

Meeting notes

Drop a recording on a START node. The transcript becomes the input to an agent that pulls out decisions, owners, and follow-up actions, then writes them somewhere.

Support call review

Transcribe a call, classify it with a Condition node, and route anything flagged as an escalation to a Human Task.

Screen recordings

A recorded walkthrough uploads as mp4, has its video track stripped, and the narration becomes searchable text.

Next steps

MagOneAI© 2026 Magure, Inc.

On this page