Zeepit
Deep dive · 7 min read

What You Can Feed a Transcription App: Files, Links, Recordings and Text

A transcription app reads the audio track, so audio files (mp3, m4a, wav, ogg, opus, aac, flac) and video files (mp4, mov, mkv, avi) both work. Public links usually return a summary rather than a word-for-word transcript.

What audio and video files can be transcribed?

The format matters less than most people expect. A transcription app only reads the audio track, which is why audio files and video files are both valid sources.

Audio formats normally accepted:

  • mp3, m4a, aac — voice memos, exported meetings, most downloaded episodes
  • wav, caf, flac — uncompressed or lossless recordings
  • ogg, opus — WhatsApp and Telegram voice notes
  • amr, webm — older phone recorders and browser recordings

Video files work the same way: mp4, mov, mkv, avi. The image is discarded, so a screen recording of a webinar is no harder to handle than a short voice note. The practical details are in our guide on turning a video into a transcript.

In Zeepit the ceiling is 4 hours per recording, and there is nothing to convert beforehand: pick the file exactly as your phone saved it. If you have a working file in a rarer format, try it as it is — a file that plays on your phone will usually transcribe.

Can you paste a link instead of a file?

"Importing a link" covers two very different things, and only one of them generally works.

Public pages. Paste a YouTube link, a podcast link or a web page link into Zeepit and you get a summary of what it says, with nothing to download first. That is the fast way to judge whether a one-hour episode deserves your evening. For the word-for-word case, see how to get a YouTube video into text.

Links to your own file. A Dropbox, Drive or WeTransfer share URL is not a source you can paste. Phone apps do not sign in to someone else's storage and pull a file from an arbitrary address. Do this instead: open the link in the storage app, save or download the file to the phone, then pick it inside the transcription app or send it through the Share sheet.

Rule of thumb: a link gets you a summary, a file gets you the full transcript.

Step by step

  1. Pick the file, or share it

    Inside the app, choose a file from Files or Google Drive. Or start from the app that holds the audio — voice recorder, WhatsApp, Telegram, Drive — tap Share and choose Zeepit. Same gesture on iPhone and Android.

  2. Let the language be detected

    There is no language menu to set. Detection is automatic across dozens of languages, and you can ask for a translation afterwards if you need the text in another one.

  3. Choose transcript, summary or both

    The full word-for-word transcript answers "what exactly was said". The summary, in a short, medium or long version, answers "what do I need to remember".

  4. Keep or send the result

    Copy it, share it, or export a PDF, Word (DOCX) or plain-text file generated on the device. The audio itself is sent to the servers, transcribed and deleted right after: no archive, no reuse, no AI training on your recordings.

Voice notes, memos and meeting recordings you record yourself

Anything your phone can record is a valid source: a voice memo, a lecture, an interview at a café table, a meeting picked up from a laptop speaker, a voice note a colleague sent you.

One detail matters more than the rest: most phone transcription apps do not listen live. They do not sit inside a call and they do not capture a meeting as it happens. The pattern is always two steps — record with the recorder you already use, then send that file on through the Share button, which exists on both iPhone and Android.

The common cases are covered separately:

Recordings made by a laptop, a dictaphone or a camera count too, as long as you can get the file onto the phone first.

Is plain text a valid input as well?

Sometimes you already have the words and transcription is not the job. A pasted email thread, a chapter you copied out of a PDF, a transcript a colleague sent you: what you want is either a shorter version or a voice reading it.

Both are ordinary features rather than transcription. You can condense text into a short, medium or long summary — the trade-offs are in how to summarize any text — or have it read aloud on your phone while you walk or drive.

Translation sits between the two. An audio file recorded in one language can come back as text in another, and a piece of text can be translated without ever being spoken. Language detection is automatic in most apps, so you rarely need to declare which language you are feeding in.

What a phone transcription app cannot take

The honest boundaries, so you do not waste time looking for a setting that does not exist:

  • No live capture. Nothing is transcribed while it is happening. Record, then convert.
  • No desktop app, no web upload page. Everything happens on the phone; there is no computer step to prepare the file.
  • No sync between devices. Transcripts and summaries stay on the phone that made them. Move them yourself by copying, sharing, or exporting a PDF, Word (DOCX) or plain-text file.
  • No arbitrary download links. Link import is for public pages, not for private storage URLs.
  • No audio locked inside another app. If a streaming or protected track offers no share and no save option, no transcription tool can reach it.
  • Silence and music are not speech. A video with no spoken words produces nothing useful.
  • Very long sessions. Above the per-recording limit, split the file into parts with any recorder or trimming tool.

How to pick the best version of a source for an accurate transcript

When the same content exists in several versions, the closest one to the original wins. Accuracy comes from the recording, not from the file extension.

  • Prefer the recorder's own file over a phone re-recording of a loudspeaker.
  • Prefer the original export over a copy compressed for messaging: chat apps shrink attachments, and thin audio is harder to read.
  • Play the first and last few seconds before sending. Truncated or silent files are common and easy to spot.
  • For several speakers, one device placed near the group beats a phone at the far end of a long table.
  • Mono is fine. Voice memos are mono by default and transcribe as well as stereo.

If the audio is already recorded and you simply want it in text, the step-by-step version lives in audio file to text. And if a recording is genuinely bad — wind, crosstalk, a distant speaker — a better recording will always beat any post-processing trick.

Frequently asked questions

What is the best audio format for transcription?

mp3, m4a and wav all transcribe well, and accuracy depends on the recording rather than the container. Keep whatever your recorder produced instead of re-encoding it, since every conversion can shave off detail. Voice-note formats like ogg and opus work exactly as they are.

Can I transcribe a video without converting it to audio first?

Yes. Apps that accept mp4, mov, mkv and avi read only the audio track, so there is no conversion step and no separate audio export. The picture is ignored. File size mainly affects how long the transfer takes on a slow connection.

Can I paste a Google Drive or Dropbox link to transcribe a file?

Usually no. Link import is meant for public pages such as YouTube videos, podcast episodes and articles, and it returns a summary. For a file sitting in cloud storage, open it in the storage app, download it to the phone, then pick it or use the Share sheet.

How long can a single recording be?

In Zeepit, up to 4 hours per recording. Longer sessions need to be split into parts with any recorder or trimming tool. Your first conversions are free right after install, long recordings included, and the usage is shown in the app.

Can I transcribe files that are on my computer?

Not directly with a phone-only app: there is no desktop version and no web upload page. Send the file to your phone first, by cloud storage, AirDrop or a message to yourself, then pick it in the app. Exported PDF, Word or text files can travel back to the computer.

Read next

← Back to the blog