Study a hook and its delivery
Look at the opening sentence, the order of ideas and the closing call to action. Timestamps help distinguish the actual spoken hook from text that only appears on screen.
FROM SPOKEN VIDEO TO USEFUL TEXT
Need the words from a TikTok without replaying it sentence by sentence? Start with a public video link or a video or audio file. Review the spoken text with timestamps, correct the wording and choose a text or subtitle format for your next step.
A transcript captures speech. It is different from the post description, on-screen text or a translation.
KNOW WHAT TO EXPECT
You may need a single quote, a readable set of notes or subtitles for an edit. Start by checking the source and the output you actually need.
THE WORKFLOW
Keep the source, the spoken words and the subtitle timing connected. Review the result before publishing or quoting it.
Use the link to a public TikTok video. If the link cannot be read, use a video or audio file you already have. Choose the spoken language when you know it; otherwise leave language detection on automatic.
Read the spoken text in timed segments and check it against the original. Correct names, specialist terms and unclear phrases. Background music, overlapping speakers and fast speech can affect recognition.
Choose TXT for notes and writing, SRT for a conventional subtitle workflow, or VTT for web video. Editing a sentence changes its text; the original segment timing stays in place.
START WITH THE SOURCE YOU HAVE
Use the option that gets you to the spoken audio with the least extra work. You do not need to download a video simply to understand which output format you need.
Copy the link to the individual video, not the creator profile or a search-results page. A complete video URL or a supported TikTok short link is suitable for the link input. This route depends on the source video remaining accessible when it is processed.
A local copy is useful for your own recordings or when fetching a public link fails. Check the actual file format, duration and size first. Renaming a file extension does not convert its encoding. If only a section matters, trim the file in an editor before processing.
If the video already has readable captions and you only need a few words, checking that moment manually may be enough. Transcription becomes more useful when you need the full spoken text, searchable notes or a separate subtitle file. Existing captions can still contain errors, so confirm important wording with the audio.
UNDERSTAND THE OUTPUT
A TikTok transcript is a written record of the words spoken in a video. Speech recognition works from the audio track, so a video can have a transcript even when the creator did not add subtitles. It can help you find a phrase, take notes or prepare a subtitle file without repeatedly replaying the whole clip.
PUT THE WORDS TO WORK
A useful transcript makes the spoken information easier to inspect and organize. Keep the original source alongside anything you reuse.
Look at the opening sentence, the order of ideas and the closing call to action. Timestamps help distinguish the actual spoken hook from text that only appears on screen.
Extract steps, ingredients or explanations from a spoken tutorial. Search the text for a detail, then replay that moment to confirm quantities and instructions.
Use the timed transcript as a starting point for subtitles. Review wording and line length in your editor; speech timing alone does not guarantee comfortable reading speed.
Turn your own recorded explanation into an outline, a post or a brief. A transcript preserves the spoken wording; rewrite repetitions and unfinished sentences for the new format.
CHOOSE THE RIGHT FILE
Choose an export for what you will do next. Text files favor reading; subtitle files keep the time cues needed for playback.
| Format | What it contains | Best suited to |
|---|---|---|
| TXT | Plain text, without subtitle cues | Notes, outlines and documents |
| SRT | Numbered cues with start and end times | Video editors and subtitle workflows |
| VTT | WebVTT cues with start and end times | HTML video and web players |
CHECK BEFORE YOU REUSE
Clear speech is easier to recognize than a voice buried under music. Accents, switching languages, unfamiliar names and several people talking at once can produce errors. AI transcription is a first draft: listen to important passages before quoting them or publishing subtitles.
WHEN THE RESULT IS NOT WHAT YOU EXPECT
Identify whether the problem is the source, audio or using an export. If Generate is disabled and the panel says Preview, the processing service is not connected in that environment.
Open the link yourself and confirm it points to a video that is still public. Copy a fresh share link and check that it is not a profile URL. If the source remains restricted and you have your own accessible copy, use the file route instead.
The initial limits are five minutes and 50 MiB. Keep only the relevant section or export a smaller supported file from your editor. Avoid repeated heavy compression that makes voices harder to hear. If you split a recording, keep track of each clip’s original start time.
Listen for spoken words rather than music alone. Check that the recording includes its audio track and that the voice is audible. Text written in the video image is not speech and needs OCR; a speech tool cannot recover words that were never audible.
Replay those segments and correct them manually. Check the selected spoken language, especially for short or mixed-language clips. Background music and long pauses can confuse recognition. Removing repeated text is sensible only after checking that the speaker did not actually repeat it.
Speech segments are not always ideal reading units. Break long sentences into shorter cues in a subtitle editor, leave enough time to read them and preview the result with the video. Editing text inside this panel preserves the original cue times; it does not automatically retime subtitles.
SRT and VTT are separate text files, not videos with captions burned in. Import the subtitle file into a compatible player or editor and select its subtitle track. If the application cannot read the format, try its supported subtitle format rather than changing the filename extension.
TURN A DRAFT INTO SOMETHING YOU CAN USE
Review effort should match the consequence of an error. Personal notes may need a light check; a published quote, instruction or subtitle deserves closer listening.
Listen for negations, quantities, dates, names and units. Confusing “fifteen” with “fifty,” or missing “not,” can change a sentence’s meaning even when the text reads fluently. This is an illustrative error pattern, not a measured result from this tool.
Replay a little before and after the sentence you want. Keep the source link and timestamp with your notes so you can verify the wording again. If you shorten or rewrite a passage, distinguish your summary from the speaker’s original words.
For notes, check paragraph breaks and remove unwanted repetition. For subtitles, check timing, line length and whether cues obscure useful visual information. A valid subtitle file can still need editorial adjustments before it is ready to publish.
BEFORE YOU START
Paste a public video link into the TikTok Link tab, or select Upload File for a video or audio file. Choose a language, generate the transcript and review the timed text. Link processing depends on the source remaining publicly accessible. The tool panel displays whether transcription is available; the labelled format example is not a transcription of your video.
Yes, speech-to-text recognition uses the audio rather than requiring an existing subtitle track. It still needs audible speech. A music-only or silent clip may not produce a usable transcript, and on-screen writing requires a separate OCR tool.
The upload flow accepts MP4 video and MP3, M4A or WAV audio. Maximum duration is 5 minutes. MP4, MP3 and WAV support files up to 50 MiB; M4A audio is limited to 8 MiB. The Worker checks actual container bytes and duration, so a renamed or incompatible codec can be rejected. Files are sent only when you start transcription.
Timed transcript segments can be exported as SRT or VTT, while TXT contains the readable text without subtitle time codes. SRT is commonly used in editing workflows; VTT is suited to web players. Always preview subtitle timing and line length in the destination player or editor.
Transcription writes down speech in its original language. Translation changes those words into another language and is a separate operation. Selecting English tells recognition to expect English speech; it does not translate non-English audio.
The video may have been deleted, made private or restricted, or its media URL may have expired. Some posts are photo slideshows rather than spoken videos. If you have an accessible copy, uploading the file avoids the link-fetch step, but the audio still needs to contain recognizable speech.
No. A post caption is the written description and hashtags attached to the post. A transcript records words spoken in the audio. Subtitles add timing to those spoken words. A video can have all three, and their content may be very different.
Accuracy varies with recording quality, background sound, language and speech style. There is no single accuracy percentage that applies to every TikTok video. Treat the output as a draft, correct important wording against the source, and check subtitle timing before publishing.
No account or payment is required for the connected workspace. The limit is 3 attempts per device and 10 per IP per UTC day; failed attempts count. If the processing service is not connected, only the labelled format example is available.
File selection stays local until you start transcription. Processing sends media to private Cloudflare storage and Workers AI. The creating device cookie controls access; there is no public media link. Media and results expire after 24 hours, with background cleanup. Export files you want to keep.
Processing time depends on fetching the source, duration and current service load. The panel shows real task states. We have not yet established a representative speed benchmark, so the example does not prove a fixed turnaround time.
Each file or linked video is limited to five minutes. Longer recordings need to be trimmed into relevant clips first. Keep each clip’s original start time if you later combine transcripts. The current release does not provide batch processing or automatic splitting of videos longer than the limit.
Speech recognition reads the audio, not pixels in the video frame. Text displayed only on screen, such as a slide, product label or a silent text overlay, needs OCR. If a speaker reads those words aloud, they may appear in a speech transcript, but visual text is not extracted independently.
No. SRT is a separate subtitle file with timed text. Add it to a compatible player or editor, review its cue timing and export the video from that application if you want subtitles permanently rendered in the image. Editing text here does not burn captions into a video or automatically change cue times.
Read the words here; use the other tools to explore the account and the performance behind the video.