I Tried DupDub AI for Transcription, Here’s What Happened
I signed up for DupDub, ran a short video through its speech to text tool, and tried to export the result. This is a step by step account of what happened, with the exact screens I saw and a short note on what worked and what did not at each stage. I did not use any paid credits; everything here was done on the 10 free credits a new account gets.
Signing up
The sign up page offers Google, Facebook, or email. I went with email. You enter an address, set a password of at least six characters, then click Get Code and paste in the verification code that lands in your inbox. The whole thing took under two minutes and no card details were asked for.

Figure 1: DupDub sign up screen with Google, Facebook, and email options
What I liked
- No credit card required to get started.
- Social login options are there if you want to skip the OTP step.
- The left panel tells you upfront what the suite covers: voiceovers, TTS, speech to text, cloning, video editing.
What I did not like
- The email code step adds friction compared to tools that just send a magic link.
- The password minimum of six characters is on the weak side.
The onboarding questions
Before reaching the dashboard, DupDub asks three questions: who you are, what type of business you are in, and what content you plan to make. I picked Personal Use, Advertising, and Social Media Content. There is a progress bar and a back arrow, and each screen took a few seconds.

Figure 2: Question 1 of 3, user type

Figure 3: Question 2 of 3, business type

Figure 4: Question 3 of 3, content type
What I liked
- Clear progress indicator so you know it is only three screens.
- Single tap answers, nothing to type.
What I did not like
- There is no visible skip button, so you have to answer all three before you can do anything.
- I could not tell what actually changes based on my answers. The dashboard looked generic afterwards.
Opening the speech to text tool
The tool landing page is simple: a headline, a one line promise that it turns audio or video into editable text in seconds, and an upload box directly below.

Figure 5: Speech to text landing page
The upload box accepts a local file or a link. The link tabs cover YouTube, TikTok, Facebook, Instagram, X, Douyin, Kuaishou, Redbook, Bilibili, and a generic Other URL option. I uploaded a 10 second MP4 called Face_Shapes_vs_Frame_Styles.mp4, 2.48 MB. The language dropdown defaults to Auto detection, and a small note says it will also translate if the detected language differs from the one you select.

Figure 6: Upload panel with the test file loaded and language set to Auto detection
What I liked
- Paste a link support is broad, including Chinese platforms most Western tools ignore.
- Auto detection means you do not have to know the source language.
- The translate at the same time option is a nice bonus that surprised me. Most transcription tools make translation a separate paid step.
What I did not like
- No drag and drop hint or supported format list on the upload box itself.
- Auto detection is the default, which turned out to matter later.
Credit confirmation
Before anything runs, a confirmation box shows the file title, duration, size, the credit cost, and your remaining balance. My 10 second clip cost 0.2 credits, leaving 9.8 from the 10 free credits.

Figure 7: Transcription confirmation with credit consumption
What I liked
- Cost is shown before you commit. You will not be surprised by a drained balance.
- At 0.2 credits per 10 seconds, the free allowance works out to roughly eight minutes of audio, which is enough to test properly.
What I did not like
- The info icon next to Consumption does not explain how credits scale with longer files.
The transcript
Processing finished in a few seconds. The result screen shows the title, a language badge, duration, character count, and a timestamped transcript on the left with AI editing tools on the right: Rewrite, Summarize, Make longer, Make shorter, Ask AI to write, Convert text to speech, Download SRT, Copy text, Move to, and Rename.

Figure 8: Transcript result page
Here is the honest part. The transcript that came back was a single line reading [outro jingle], 14 characters in total. The language badge read Por, meaning DupDub decided my clip was Portuguese. The file name is English and the clip is 10 seconds long. The tool clearly picked up on the music bed and tagged it as a sound event, which is a fair thing to do, but the Portuguese label is wrong and there was no prompt to check or correct the detected language.
What I liked
- Non speech audio gets tagged rather than turned into gibberish words. That is the right behaviour.
- Auto segment and timestamp is on by default, so longer files would come back structured.
- The edit sidebar is genuinely useful. Summarize and Convert text to speech sitting next to the transcript saves a lot of copy pasting between tools.
What I did not like
- Language auto detection misfired on a short clip and the interface did not flag low confidence.
- For a 10 second video, 14 characters of output is thin. I would want to see at least a note that no speech was detected.
- There is no inline audio player to scrub through and check the transcript against the source.
Trying to download
I clicked Download SRT to grab the caption file. Instead of a download, a red banner appeared at the top of the screen.

Figure 9: Error shown when clicking Download SRT
The message said the file had expired and that I needed to retranscribe to download. This happened within minutes of the transcription finishing, in the same session. Retranscribing would cost another 0.2 credits. I did not expect a result to expire that quickly, and nothing on the result page warns you that downloads are time limited.
What I liked
Copy text still worked, so the content itself was not lost.
What I did not like
- This is the concrete flaw of the test. An export option that fails minutes after processing, with no warning and a re-run cost attached, is a real problem for anyone working with longer files.
- The error gives no time window, so you cannot plan around it.
My verdict
The setup and interface are among the smoother ones I have used. Cost transparency before every job, broad link support, and having summarise, rewrite, and text to speech right next to the transcript are all things I would happily keep using. The part that surprised me most was the built in translation offer on the upload screen, which competitors usually gate behind a higher plan.
The core job, though, had two problems in one short test: language detection got it wrong with no way to catch it, and the SRT download expired almost immediately. On a 10 second clip those are annoyances. On a 40 minute interview, paying credits twice because the first export timed out would be a genuine cost.
If you are evaluating DupDub, my advice is to set the language manually rather than trusting auto detection, and to download or copy your output the moment it finishes. I will run a longer clip with clear speech next and update this once I see how accuracy holds up on real dialogue.
