A press-to-dictate microphone button that fills a text field by voice, and stops claiming to listen the moment the browser has stopped. Reach for it beside any field somebody would rather speak than type: a search box, a comment or reply box, note and journal fields, a message composer, contact and support forms, meeting and consultation notes, inspection and field-service reports filled in on a phone with gloves or dirty hands, delivery and warehouse notes, clinical and veterinary notes, incident and maintenance logs, recipe and shopping lists, long-form description fields in a CMS or listing flow, translation and language-practice inputs, and anywhere dictation is the accessible alternative for someone who cannot comfortably type — motor impairment, RSI, a broken wrist, or simply a phone in one hand. Common asks it answers: "speech to text react", "voice input component", "dictation button", "react speech recognition component", "microphone button for input field", "web speech api react hook", "voice typing textarea", "speech to text search box", "react-speech-recognition alternative", "useSpeechRecognition hook", "voice dictation shadcn", "shadcn microphone input", "talk to type react", "record voice fill form", "webkitSpeechRecognition react". Official shadcn/ui has nothing for this and no combination of its parts reaches it: input and textarea are the fields themselves and know nothing about audio, button is a button, and there is no speech, microphone or recording primitive anywhere in the library. Distinct from every other -input component in this registry — otp-input, tag-input, masked-input, phone-input and the rest are the field, while this one sits beside a field you already have and writes into it, so it composes with any of them. The component turns on one fact that every hand-rolled version gets wrong: recognition ends by itself. The browser stops a session after a stretch of silence, and after a while regardless, and the only thing it tells your code is an end event. A button that tracks the boolean its own click set therefore keeps a pulsing red dot and the word Listening over a microphone that was handed back a minute ago, and the user keeps talking into nothing. Here every visible state is driven by the platform's own start, end and error events, the setting the user asked for is kept separate from whether audio is actually being captured (aria-pressed carries the first, data-active the second), and continuous mode starts a fresh session when the browser ends one — under a budget, so a machine with the microphone muted or a laptop that has gone offline cannot turn that into a hot loop, and so a genuine pause in the middle of a paragraph costs nothing. The restart also resets the result cursor, because a new session numbers its results from zero and a cursor carried over from the last one silently swallows the first words after every pause. Interim results are kept strictly out of the value: the service rewrites its guess as more audio arrives, so writing it into the field the user is editing changes their content under them and fills their undo history with words nobody typed — the guess is shown beside the button as a preview and only confirmed text is ever appended. Appending is its own small problem and is solved and exported as appendTranscript, because value + transcript welds every chunk onto the previous word and padding unconditionally with a space yields a stray gap before a dictated full stop. Feature detection reads the prefixed constructor as well as the standard one, since webkitSpeechRecognition is the spelling the browsers that actually ship this expose, and a missing API is a first-class state with its own copy rather than a dead button. The not-allowed error is untangled rather than taken at face value: it means four different things — an http origin, an iframe without allow="microphone", a stored block, and a prompt closed without an answer — and only the third is worth sending someone to their site settings for, so the other three say something true instead, and a permission changed in those settings is picked up live through the Permissions API change event rather than staying dead until a reload. Stop asks the service to deliver what it is still holding instead of aborting and losing the last thing that was said, and unmounting detaches the handlers and aborts, so navigating away inside a single-page app cannot leave the recording indicator lit with no control left that could turn it off. The recognition language is resolved from the document rather than left to the user agent, which is the one setting whose behaviour is not defined across browsers. Worth knowing before shipping it somewhere sensitive: the specification permits the audio to be sent to a remote service — the existence of a network error code is the platform admitting as much — so a dictated field may be data that has left the device. Ships as a hook (useSpeechInput) plus a button, controlled by pairing value with onValueChange or left to report through onTranscript, marked aria-disabled rather than disabled so a refusal stays reachable and can explain itself, with a permanently mounted polite live region for status and failures. Styled entirely with shadcn tokens (input, accent, ring, muted-foreground, destructive), so it follows light and dark mode, and its only dependency is lucide-react for the icons.
pnpm dlx shadcn@latest add "https://pulld.pages.dev/r/speech-input.json"