Skip to main content
Voice

Voice

Vresk can take your words in and read its own back. Dictation puts a microphone on the composer and turns speech into ordinary, editable text; read aloud puts a speaker on an answer and plays it to you. Both run entirely inside your browser — your voice is never uploaded, and no recording is ever stored.

Dictating a message

  1. Press Dictate

    The mic button sits in the composer toolbar, next to Attach (on a phone, tap the + button and choose Dictate). The first time, your browser asks permission to use the microphone.
  2. Speak

    A line under the input shows your words as they are recognized, and a ring around the button moves with your voice so you can see it is hearing you.
  3. Press Stop

    The finished transcript lands in the input box as plain text. Nothing is sent yet — read it, fix anything, then send it like any typed message. Editing it is the confirmation step; there is no separate approve button.

A single dictation can run up to ten minutes. When the last ten seconds remain, a countdown appears next to the live transcript so the stop never takes you by surprise.

The first dictation on a device downloads the speech model into your browser, and the strip under the input says so — with a running percentage — rather than pretending to listen while it downloads. Once the model is on your device, later dictations skip the download.

On Safari

Safari cannot show the live transcript while you speak. Recording works the same — the strip says “Listening…” while you talk, and your words appear all at once when you press Stop.

When a take doesn't land

A take that fails is not thrown away, and the notice under the input names what actually happened: the recording was too quiet, too short, no words were recognized, or the speech engine is still starting. Each has its own advice, because “too quiet” and “still loading” need different fixes.

When the recording itself was usable, the notice offers Try analysis again — it re-runs recognition on the same take, so a hiccup never costs you a re-recording — and Dismiss, which drops it. The kept take lives only in memory, and only for sixty seconds: after that it is discarded and the retry offer disappears with it.

Every notice also offers Copy diagnostics — a report you can paste into a bug report. It carries timings and device details, never the words you said.

Where your voice goes

Both directions of voice run inside the page you have open. Dictation captures audio on the page and hands the raw samples to a sandboxed frame that runs the recognition model locally. The audio is never uploaded to our servers and never written to disk — it exists in your browser’s memory while it is being worked on, and is freed when the transcript lands or when the sixty-second retry window above runs out.

Read aloud is the same shape in reverse: the answer’s text goes into a speech model running locally, audio comes out of your speakers, and nothing is kept. The one thing voice fetches over the network is the speech models themselves — once, then cached on your device.

Your device sets the pace

Because recognition runs on your own machine, a slower device updates the live transcript less often and takes longer after you press Stop. The words still arrive — and the strip reports what the engine is actually doing while you wait rather than pretending to listen.

Hearing an answer read aloud

Each answer’s controls — alongside the ones described in Answers — include a speaker button. Press it and a short Preparing spinner runs while the first audio is generated, bounded at about seven seconds; then the answer plays. Press the square to stop.

One message reads at a time — starting another stops the first. If audio generation falls behind playback mid-answer, the button shows a visible buffering pause and then resumes. You never get silent dead air with no sign of whether it finished.

It reads the prose. Code blocks and images are skipped rather than spelled out, and links are read as their text.

Two bounds worth knowing

Read aloud covers the first 8,000 characters of a message — a very long answer stops there. And a side-by-side comparison has no speaker button at all: reading two or three engines’ answers back to back in one voice would blur who said what.

Choosing a voice and speed

The small settings button next to the speaker opens the voice picker: six voices — US and UK accents, female and male — and four speeds, from 0.75× to 1.5×. On a phone the reply has no settings button: the same voices and speeds are in Settings, under General (tap Instructions in the account menu at the bottom of the chat list).

A change applies to your next play. An answer already being read keeps the voice and speed it started with rather than switching mid-sentence.

The choice is saved in this browser, so it holds across conversations on this device. It is not part of your account, and it does not follow you to another machine.