# Voice

Press a shortcut, say a task, press it again. Hover shows what it heard, the folder, the agent and its access, then starts a **new chat** after a short countdown (five seconds unless you change it).

> **Note:** Voice is off until you switch it on in **Settings → Voice**. Nothing is recorded or downloaded before you do.

## Setting it up

1. **Add your projects** (Settings → Projects): the folders voice may work in, and the other names you might say for them. See [Projects](https://tryhover.co/docs/projects.md).
2. **Pick how speech becomes text** (Settings → Voice), below.
3. **Optionally add cleanup.** Gemini, OpenAI or any OpenAI-compatible service, with your key and model, fixes punctuation and filler words. If it fails, the original text is used.
4. **Talk.** Press `Ctrl` + `Alt` + `Space` (you can change it; on a Mac it is Control-Option-Space and fixed) to start listening, speak, and press it again to finish. If you would rather hold the keys down, set **Settings → Voice → Voice Recording Mode** to "Hold to speak".

Voice uses your **default agent**, the one last picked in the new-task circle, with that agent's own model and settings.

## Local or cloud

| Mode | Where it runs | Languages | Needs |
| --- | --- | --- | --- |
| **Local** (Phonon) | On your computer. The recording never leaves it. | English only | Press Download on the Phonon card |
| **Cloud** (Groq) | The recording is sent to Groq | Detected for you | Your own key from console.groq.com |

Hover never switches between the two on its own. *Check key* tests a Groq key before you rely on it.

> **Note:** On a Mac voice uses Apple's own recognizer instead, on device where the language allows. There is no Phonon, Groq or cleanup there, and macOS asks for the Microphone and Speech Recognition the first time you use the shortcut.

## What the local model needs

| Platform | Download | On disk | Free space to set up |
| --- | --- | --- | --- |
| Windows x64 | 420 MB | 1.5 GB | 1.9 GB |
| Linux x64 | 523 MB | 1.8 GB | 2.3 GB |
| Linux arm64 | 410 MB | about 1.8 GB (not tested yet) | about 2.2 GB |

The model itself is 164 MB. The rest is the private Python and PyTorch runtime that Phonon's engine needs; Hover keeps it in its own folder and never touches a system Python. Windows also needs the Microsoft Visual C++ Redistributable (x64), and the card says so before anything is downloaded. *Remove* deletes the whole install.

## Starting, editing, cancelling

- Hover shows what it heard, the project it matched, the agent and its access, then counts down. **Settings → Voice → Start on its own** sets the wait: Off, 3, 5 (the default) or 10 seconds.
- **Edit the task** to stop the countdown, then press *Start* when you're ready.
- **Change the folder.** The folder pick on the card opens a list of your voice projects and the default workspace. Changing it stops the countdown, so press *Start* when you're ready.
- `Esc` cancels.
- **Settings → Voice → Try it** shows what would start, without starting anything.

## The aura

While Hover is listening or working on what you said, the card shows only an aura: a glowing ring of light, with no words. It swirls and swells with your voice while you speak, and swirls faster with a pulse while Hover works. `Esc` still cancels. Pick its colour in **Settings → Voice → Aura colour**, by name or with any hex colour.

## Attaching screenshots

Say "take a screenshot" (or "take a screenshot of this") while you speak, and Hover takes a picture of your screen and sends it with the task. You can say it as often as you like. Each one chimes, flashes the notch and says "Screenshot attached". The words themselves never reach the task.

- With **Cloud** speech the picture is taken while you speak. With **Local** speech it is taken when you finish.
- The preview shows the pictures, each with an × to drop it.

## Sending a task to Kiro Web

Say "use Kiro Web" or "use cloud agent" (also "run in the cloud" or "in Kiro Web") and the task goes to Kiro in Kiro Web, in Kiro's cloud, instead of running on this computer. English only.

- Those words are taken out of the task.
- The card shows Kiro with the cloud button on, so you can switch it off before *Start*.
- The card has a repository pick, with a search box: the folder's own repository, none, or a connected one. The folder pick and its default-workspace note are hidden.
- Every Kiro Web task has full access, so the card has no access pick for it.

See [Agents](https://tryhover.co/docs/agents.md) for what Kiro Web is.

## Dictating into a chat

With a chat open and its reply box open, use the shortcut with the pointer over the chat. What you say is written into the reply instead of starting a new task.

> **Heads up:** The recording is deleted once it has been turned into text. Cleanup only ever gets the text, never the audio.

---

Part of the [Hover docs](https://tryhover.co/docs.md) (Using Hover). HTML version: https://tryhover.co/docs/voice. Index for agents: https://tryhover.co/llms.txt