Skip to content
WriteVoice WriteVoice Start Creating

139 Voice Search Statistics October 2026

Updated 139 primary-source citations Refreshed monthly

On this page
  1. Key Takeaways
  2. How many people use voice assistants in 2026?
  3. What do people use voice search for?
  4. How are AI chatbots changing voice assistants?
  5. How accurate is speech recognition in 2026?
  6. Is voice typing replacing the keyboard?
  7. How big is the voice recognition market?
  8. Methodology and Sources

Voice search statistics for 2026 show people using voice on phones, PCs and in search, well beyond the smart speaker. When people use voice, they engage with Microsoft's Copilot on Windows twice as much as when they type, based on Microsoft's own data from March to August 2025 (Microsoft, Oct 2025). People in more than 200 countries and territories can now hold a voice and camera conversation with Google Search (Google, Mar 2026). Adoption is high; accuracy still varies by speaker and setting.

The figures below cover smart speaker and voice assistant adoption, what people ask by voice, how AI chatbots are rebuilding assistants, speech recognition accuracy, voice typing against the keyboard, and the size of the voice recognition market. Each one links to the organisation that collected it: survey firms, statistics offices, peer-reviewed papers, company filings and newsrooms, and market research firms named as such. Older foundational studies carry their year in the sentence. For keyboard speeds in more depth, see our typing speed statistics.

Key Takeaways

  • 39% of Americans aged 12 and older own a smart speaker in 2026, an estimated 112 million people, up from 7% in 2017 (Edison Research at SSRS, Mar 2026).
  • There are more than 600 million Alexa devices in use (Amazon, Jul 2026).
  • 58% of US consumers used voice search to find information on a local business in the previous 12 months, in a 2018 survey of 1,012 people (BrightLocal, Apr 2018).
  • 49% of US adults now use AI chatbots, up from 33% in 2024, when Pew asked the question differently (Pew Research Center, Jun 2026).
  • Gemini Live voice conversations are five times longer than text conversations on average (Google, May 2025).
  • 63% of Gemini app users now talk to it directly, including a growing number who use voice only (Google, Aug 2026).
  • Speaking into a phone was 2.93 times faster than typing on its keyboard in English, 153 against 52 words per minute, in a lab test (Ruan et al., IMWUT 2017, Dec 2017).
  • Phones carried 58.99% of worldwide web traffic in September 2026 (StatCounter, Sep 2026).
  • The most accurate model in the Open ASR Leaderboard's short-form English results, the open-source Cohere Labs Transcribe, averages a 5.42% word error rate, against 7.44% for OpenAI's Whisper Large v3 (Srivastav et al., Open ASR Leaderboard v4, Mar 2026).
  • In tests of five commercial speech recognition systems, 23% of audio snippets from Black speakers had a word error rate above 50%, the authors' threshold for an unusable transcript, against 1.6% of snippets from white speakers (Koenecke et al., PNAS 2020, Mar 2020).
  • A speech recognizer with a 3.4% word error rate for typical speakers had a 36.3% error rate for people with Parkinson's-related dysarthria (Hasegawa-Johnson et al., JSLHR 2024, Sep 2024).
  • 1.4% of Whisper API transcriptions in a 2023 test contained phrases or sentences that were never spoken, and 38% of those inventions carried explicit harms (Koenecke et al., ACM FAccT 2024, Jun 2024).
  • 14% of 1,003 UK GPs surveyed in August 2025 already used an ambient AI scribe, and 39% planned to adopt one soon (Blease et al., BMJ Health and Care Informatics 2026, Jul 2026).
  • Microsoft agreed to buy speech recognition company Nuance for $19.7 billion in cash in 2021, $56.00 a share and a 23% premium (Microsoft, Apr 2021).

How many people use voice assistants in 2026?

Voice assistants now ship on billions of devices, while the share of Americans with a smart speaker grew slowly.

Voice search statistics: share of Americans aged 12 and older who own a smart speaker, 2017 to 2026
Figure 1: Ownership climbed quickly after 2017, stalled for five years and rose again in 2026, a slow curve next to the spread of voice on phones, in cars and in AI apps. Source: Edison Research at SSRS, The Infinite Dial 2026, Mar 2026. Chart published by the source (slide 8, cropped from the PDF), captured 2 Oct 2026.

How many people own a smart speaker?

Pew's adult-only survey lands close to Edison's figure, and most owners have more than one device.

  • Smart speaker ownership held between 33% and 36% of Americans aged 12 and older every year from 2021 to 2025, with a dip to 34% in 2024 (Edison Research at SSRS, Mar 2026).
  • 35% of US adults say they have a smart speaker, such as an Amazon Echo or Apple HomePod, in a February 2026 survey (Pew Research Center, Jun 2026).
  • Other AI-enabled home devices are less common: 18% of US adults have a smart doorbell, 13% a robot vacuum and 11% a smart thermostat (Pew Research Center, Jun 2026).
  • In early 2020, 24% of US adults, about 60 million people, owned at least one smart speaker (NPR and Edison Research, Jan 2020).
  • The average smart speaker household owned 2.6 devices in early 2020, up from 2.3 a year earlier (NPR and Edison Research, Jan 2020).
  • 62% of US adults used a voice assistant on some device in 2022, counting smart speakers, smartphones and in-car systems (NPR and Edison Research, Jun 2022).
  • 57% of US voice command users gave voice commands every day in 2022 (NPR and Edison Research, Jun 2022).

How many voice assistant devices are in use worldwide?

Company-reported totals run from hundreds of millions to billions. They count hardware that can run an assistant, not people who use one every day.

  • Customers interacted with Alexa billions of times each week across hundreds of millions of Alexa-enabled devices, including TVs, PCs, hearables and wearables, at the end of 2020 (Amazon, Dec 2020).
  • Apple's installed base passed 2.5 billion active devices at the start of 2026, a count of devices rather than of Siri users (Apple, Jan 2026).
  • Google Assistant was available on more than 1 billion devices by January 2020 (Google, Jan 2020).
  • More than 500 million people used Google Assistant every month in January 2020, across smart speakers, smart displays, phones, TVs and cars in more than 90 countries (Google, Jan 2020).
  • Samsung's Bixby had more than 200 million users on nearly 400 million devices worldwide in 2022, after registered users more than doubled in a few years (Samsung, Nov 2022).

Who uses voice assistants most, by age and country?

In the US, people in their 30s and 40s lead; in Europe, the youngest adults do. Global adoption is uneven: the gap between EU countries is far wider than the gap between age groups.

  • 46% of US adults said they used digital voice assistants in 2017, most often on a smartphone (Pew Research Center, Dec 2017).
  • In 2017, 55% of US adults aged 18 to 49 used voice assistants, against 37% of those 50 and older (Pew Research Center, Dec 2017).
  • In 2026, 41% of US adults aged 30 to 49 have a smart speaker, against 33% of those 18 to 29, 36% of those 50 to 64 and 24% of those 65 and older (Pew Research Center, Jun 2026).
  • 16% of people aged 16 to 74 in the EU used a virtual assistant, through a smart speaker or an app, in the three months before they were surveyed in 2024 (Eurostat, Apr 2026).
  • Use ranged from 41% of people in Ireland, 29% in the Netherlands and 27% in Spain to 3% or less in Latvia and Bulgaria in 2024 (Eurostat, Apr 2026).
  • In the EU, 23% of people aged 16 to 24 used a virtual assistant in 2024, against 7% of those aged 65 to 74 (Eurostat, Apr 2026).
  • Men in the EU were slightly more likely than women to use one in 2024, at 17% against 15% (Eurostat, Apr 2026).

What do people use voice search for?

People talk to their devices when their hands are busy and the question is short: music, the weather, a timer, a nearby restaurant. Voice loses ground once a task turns private or complicated. The newest public task-level surveys are from 2018 to 2020 and nobody has published a broad replacement since, so read the dates.

Bar chart of the top ten requests US smart speaker owners made in a typical week in spring 2020, led by playing music, the weather, general questions and timers
Figure 2: Short, hands-free requests fill the weekly list, and owners whose routine changed and who worked from home during COVID-19 (purple bars) used their speakers in much the same way. Source: NPR and Edison Research, The Smart Audio Report, Spring 2020, Apr 2020. Chart published by the source (slide 48, cropped from the PDF), captured 2 Oct 2026.

What tasks do people hand to voice assistants?

Music comes first, then the weather, quick questions and timers. Phone users lean more on calls, texts and directions than speaker owners do.

  • In a spring 2020 survey, 85% of US smart speaker owners asked their speaker to play music in a typical week (NPR and Edison Research, Apr 2020).
  • The next most common weekly smart speaker requests in 2020 were the weather (74%), answering a general question (72%) and setting a timer or alarm (65%) (NPR and Edison Research, Apr 2020).
  • On smartphones the 2020 mix was different: 58% of phone voice assistant users checked the weather by voice in a typical week, 46% made a phone call and 39% read or replied to a text message (NPR and Edison Research, Apr 2020).
  • Smart speaker owners used their speaker for 10.8 different types of task in a typical week in 2020, up from 9.4 in 2019 (NPR and Edison Research, Apr 2020).
  • In Adobe's 2018 survey of more than 1,000 US consumers, 47% used voice for online search, 36% to make a call and 31% for smart home commands (Adobe, Sep 2018).
  • On phones, the top voice activity in Adobe's 2019 survey of 1,025 US consumers was searching online (50%), followed by getting directions (45%), the weather (39%) and composing texts or making calls (38%) (Adobe, Jun 2019).

How do people use voice to find local businesses and shop?

Mostly to look up a business: hours, directions, a phone number. Shopping by voice stayed a minority habit, used more for research and lists than for checkout.

  • Among 2018 local voice searchers, 46% looked up a local business by voice every day and a further 28% did so weekly (BrightLocal, Apr 2018).
  • Smart speaker users searched most often in 2018: 76% looked for local businesses at least weekly and 53% every day (BrightLocal, Apr 2018).
  • After a local voice search in 2018, 28% of consumers called the business, 27% visited its website and 19% went to the business in person (BrightLocal, Apr 2018).
  • In 2020, 38% of smartphone voice users and 31% of smart speaker owners asked their assistant to find restaurants or businesses nearby in a typical week (NPR and Edison Research, Apr 2020).
  • 20% of US smart speaker owners had shopped or ordered an item by voice in Adobe's 2019 survey, and 15% regularly reordered items that way (Adobe, Jun 2019).
  • In the same 2019 survey, voice was used more to prepare a purchase than to make one: 40% ran a product search by voice, 30% built shopping lists and 25% compared prices (Adobe, Jun 2019).
  • Amazon says Alexa+ can coordinate tens of thousands of services and devices, including table bookings through OpenTable and food orders from Grubhub and Uber Eats (Amazon, Jul 2026).

Why do people choose voice over typing, and what holds them back?

Hands-free use is the main draw, well ahead of the idea that speaking feels more natural. Privacy, plain lack of interest and assistants that mishear hold people back.

  • 55% of US voice assistant users said in 2017 that a major reason they use them is to interact with their devices without using their hands; 23% said assistants are fun and 22% that speaking feels more natural than typing (Pew Research Center, Dec 2017).
  • Among US adults who did not use voice assistants in 2017, 61% were not interested, 28% said none of their devices had one and 27% cited privacy concerns (Pew Research Center, Dec 2017).
  • Only 39% of US voice assistant users said in 2017 that their assistant responded accurately most of the time; 42% said some of the time and 16% not very often (Pew Research Center, Dec 2017).
  • In 2020, 33% of US smart speaker owners said their speaker misheard or misunderstood a request at least once a day, and 39% of phone assistant users said the same of their phone (NPR and Edison Research, Apr 2020).
  • Among Americans without a smart speaker in 2020, 66% were bothered that the devices are always listening, 65% worried hackers could use them to get into their home or personal information and 58% did not trust the makers to keep their information secure (NPR and Edison Research, Apr 2020).
  • 71% of US adults think more use of AI will make their personal information less secure; 3% think it will make it more secure (Pew Research Center, Jun 2026).
  • In May 2023 the FTC and the Justice Department charged Amazon with keeping children's Alexa voice recordings indefinitely, under a proposed order requiring Amazon to pay $25 million and delete the data (Federal Trade Commission, May 2023).

How are AI chatbots changing voice assistants?

Large language models rebuilt the voice assistant category in under three years. Usage data now centres on AI assistants, and early figures show people talk to them longer than they type.

Bar chart of the share of US adults who use AI chatbots such as ChatGPT, Gemini or Copilot, 2024 and 2026
Figure 3: About half of US adults now use AI chatbots; Pew notes the 2024 figure came from a differently worded question, asked only of people who had heard of chatbots. Source: Pew Research Center, Americans and AI 2026: Chatbots, Smart Devices and Views on Impact, Jun 2026. Chart published by the source, captured 2 Oct 2026.

How many people use AI assistants like ChatGPT, Gemini and Alexa+?

The largest assistants count users in the hundreds of millions to more than a billion. These measure chatbot use, not voice use, but they show how many people already talk to AI in some form.

  • ChatGPT leads US chatbot use at 44% of adults, ahead of Gemini at 24%, Copilot at 17%, Meta AI at 14% and Claude at 6% (Pew Research Center, Jun 2026).
  • About a quarter of US adults use chatbots every day (Pew Research Center, Jun 2026).
  • 52% of Americans aged 18 and older use at least one AI chatbot every week, many of them more than one (Edison Research at SSRS, Mar 2026).
  • ChatGPT has more than 900 million weekly active users and more than 50 million consumer subscribers (OpenAI, Feb 2026).
  • The Gemini app passed 1 billion monthly users, which Google calls the fastest-growing product in its history (Google, Aug 2026).
  • AI Mode in Google Search passed 1 billion monthly users one year after launch, with queries more than doubling every quarter (Google, May 2026).
  • Writing accounted for 40% of work-related ChatGPT messages in June 2025, and about two-thirds of writing requests asked it to edit or rework text the user already had rather than write from scratch (Chatterji et al., NBER 2025, Sep 2025).
  • Weekly use of AI chatbots for news rose from 7% to 10% across the markets surveyed, and the youngest age group uses them for news three times as often as the oldest, 17% against 5% (Reuters Institute, Jun 2026).

Do people talk to AI by voice, and for how long?

Yes. Platforms that publish voice data report longer sessions by voice than by text, and voice is spreading into camera and screen sharing.

  • One in five Gemini Live interactions goes beyond voice, adding a live camera feed or screen sharing (Google, Aug 2026).
  • OpenAI began rolling out two-way voice conversations in ChatGPT in September 2023, first to Plus and Enterprise users on iOS and Android, with five voices and its Whisper model transcribing what users say (OpenAI, Sep 2023).
  • OpenAI's gpt-realtime speech-to-speech model scored 82.8% on the Big Bench Audio reasoning test in OpenAI's own evaluation, up from 65.6% for its December 2024 model (OpenAI, Aug 2025).
  • When it made its Realtime API generally available, OpenAI cut gpt-realtime prices by 20%, to $32 per million audio input tokens and $64 per million audio output tokens (OpenAI, Aug 2025).

Are Alexa, Siri and Google Assistant being rebuilt around generative AI?

Yes. All three have been rebuilt on large language models or replaced by one since 2025, and Amazon is the first to report usage at scale.

  • Alexa+ costs $19.99 a month or comes free with Prime, is available to all customers in the US and Canada, and is in early access in eight more countries: Brazil, the UK, Mexico, Italy, Spain, Germany, Austria and France (Amazon, Jul 2026).
  • Hundreds of millions of customers now use the new Alexa experiences, and in the US, customers who have tried Alexa+ sign up for Prime at a nearly 25% higher rate (Amazon, Jul 2026).
  • Alexa for Shopping, which combines Alexa+ with Amazon's Rufus shopping assistant, nearly doubled its active users year over year in Q2 2026, with interactions up more than fivefold (Amazon, Jul 2026).
  • Apple began rolling out Siri AI, a rebuilt Siri running on Apple Foundation Models developed in collaboration with Google and its Gemini models, in beta in English on September 14, 2026, with five more languages due the following month (Apple, Sep 2026).
  • Apple Intelligence works in 16 languages and, in iOS 27, runs on iPhone 16 models or later plus iPhone 15 Pro and iPhone 15 Pro Max (Apple, Sep 2026).
  • Google began moving Google Assistant users on mobile to Gemini in March 2025, when the Gemini app was available in more than 40 languages and 200 countries, and said the classic Assistant would stop being accessible on most phones later that year (Google, Mar 2025).

How accurate is speech recognition in 2026?

Machines matched professional transcribers on a telephone-speech benchmark in 2017, and the best models have improved since. Error rates vary most by who is speaking: for Black speakers, people with speech disabilities, children and clinical conversations they run far higher than on clean benchmarks, as the studies below show.

Box plots of word error rate by race for Amazon's general and medical transcription services and Whisper on recordings of nurse visits with Black and white home-healthcare patients
Figure 4: On the same recordings of nurse visits, both Amazon services made more errors for Black patients than for white patients, so a good average accuracy figure says little about who gets misheard. Source: Zolnoori et al., JAMIA Open, Decoding disparities: evaluating automatic speech recognition system performance in transcribing Black and White patient verbal communication with nurses in home healthcare, Dec 2024. Chart published by the source (Fig. 2, CC BY 4.0), captured 2 Oct 2026.

What word error rates do the best speech recognition models reach?

On clean English test sets, the best systems now make about as many errors as careful professional transcribers. Accuracy drops on long recordings and on audio unlike the test sets, and every AI voice recognition tool built on these models inherits those limits. Word error rates compare only within the same test set, so each one below names its test.

  • Microsoft's system reached a 5.1% word error rate on the Switchboard telephone-conversation benchmark in August 2017, matching a multi-transcriber human process, a year after it reached the 5.9% rate Microsoft had measured for single professional transcribers (Microsoft Research, Aug 2017).
  • IBM's own study found the best professional transcriber, after quality checks, made 5.1% errors on Switchboard and 6.8% on the harder CallHome set, while IBM's system scored 5.5% and 10.3% (Saon et al., IBM Research 2017, Mar 2017).
  • OpenAI trained Whisper on 680,000 hours of transcribed web audio, 117,000 hours of it in 96 languages other than English (Radford et al., Whisper 2022, Dec 2022).
  • Benchmark scores overstate real-world accuracy: a Whisper model and a model trained only on LibriSpeech audiobooks scored within 0.1% of each other on LibriSpeech, yet Whisper made 55.2% fewer errors on average across 13 other test sets (Radford et al., Whisper 2022, Dec 2022).
  • On 25 recordings of broadcasts, phone calls and meetings, the best of five professional transcription services, a computer-assisted one, had an aggregate word error rate only 1.15 percentage points lower than Whisper's (Radford et al., Whisper 2022, Dec 2022).
  • The Open ASR Leaderboard, built by Hugging Face researchers with NVIDIA and University of Cambridge co-authors, compares 86 speech recognition systems from 26 organisations, 74 of them open source, on the same 12 datasets (Srivastav et al., Open ASR Leaderboard v4, Mar 2026).
  • On long recordings such as earnings calls, TED talks and interviews, closed commercial systems lead: ElevenLabs Scribe v2 averaged a 7.32% word error rate, against 9.73% for the best open model (Srivastav et al., Open ASR Leaderboard v4, Mar 2026).
  • Google pre-trained its Universal Speech Model on 12 million hours of unlabeled audio in more than 300 languages and, in its own tests, matched or beat Whisper with a labeled training set one seventh the size (Zhang et al., Google USM 2023, Mar 2023).

Do error rates differ by accent, race and language?

Yes, by a wide margin. Error rates track who is speaking and how much training audio exists for their language, and fine-tuning on the right speech closes much of the gap.

  • Across five commercial systems from Amazon, Apple, Google, IBM and Microsoft, the average word error rate was 0.35 for Black speakers and 0.19 for white speakers (Koenecke et al., PNAS 2020, Mar 2020).
  • The gap was widest for Black men, at 0.41 averaged across the five systems, against 0.17 for white women (Koenecke et al., PNAS 2020, Mar 2020).
  • The pattern held in a 2024 study of home-healthcare visit recordings: across 860 utterances from 10 patients, Amazon's general transcription service had a median word error rate of 50% for Black patients talking with nurses, against 33% for white patients (Zolnoori et al., JAMIA Open 2024, Dec 2024).
  • Fine-tuning on the speech of people with Parkinson's-related dysarthria cut the Speech Accessibility Project's baseline error rate on their speech to 23.7% (Hasegawa-Johnson et al., JSLHR 2024, Sep 2024).
  • Children are harder to transcribe: off the shelf, Whisper Large v3 had a 19.9% word error rate on the CSLU OGI Kids scripted-speech test set, which fell to 1.4% after fine-tuning on child speech (Fan et al., Interspeech 2024, Sep 2024).
  • Across languages on the FLEURS benchmark, Whisper's word error rate halves for every 16-fold increase in training audio, with a squared correlation of 0.83 between the two on a log scale (Radford et al., Whisper 2022, Dec 2022).
  • Meta's Massively Multilingual Speech project built one speech recognition model for 1,107 languages that more than halved Whisper's word error rate on 54 FLEURS languages, while training on a small fraction of Whisper's labeled data (Pratap et al., Meta MMS 2023, May 2023).
  • Mozilla's crowdsourced Common Voice dataset reached 42,593 hours of recorded speech in 295 languages in its September 2026 release, 29,295 hours of it validated by other volunteers (Mozilla Common Voice, Sep 2026).
  • By our count from Mozilla's release file, 108 of the 295 languages in Common Voice 27.0 have fewer than 10 validated hours, and 31 have less than one hour (Mozilla Common Voice, Sep 2026).

How accurate is speech to text in medical and noisy settings?

Less accurate than on benchmarks, and the output needs human review. Noise, distant microphones, several speakers and clinical vocabulary all push error rates up.

  • In 217 clinical notes from two US health systems, speech recognition output had 7.4 errors per 100 words, falling to 0.4 after medical transcriptionists edited it and 0.3 in the final notes physicians signed (Zhou et al., JAMA Network Open 2018, Jul 2018).
  • 96.3% of the raw speech recognition drafts in that study contained at least one error, and 5.7% of their errors were judged clinically significant (Zhou et al., JAMA Network Open 2018, Jul 2018).
  • A systematic review of 122 studies of speech recognition for clinical documentation from 1990 to 2018 found reported word error rates from 7.4% to 38.7%, and the share of documents with errors from 4.8% to 71% (Blackley et al., JAMIA 2019, Apr 2019).
  • A 2025 review of 29 studies of AI transcription in clinical settings found word error rates from 8.7% in controlled dictation to over 50% in conversations with several speakers (Ng et al., BMC Medical Informatics and Decision Making 2025, Jul 2025).
  • In nurses' home visits, the most accurate of four systems tested had a median word error rate of 39%, and for utterances under five words its average error rate reached 86% (Zolnoori et al., JAMIA Open 2024, Dec 2024).
  • In a 2025 test of 10 speech recognition setups on 200 dictated orthodontic records, word error rates ranged from 3.7% (GPT-4o transcription corrected by GPT-4o) to 33.9% (Dragon Professional Anywhere), background noise raised errors across systems, and clinically significant errors appeared with all of them (O'Kane et al., Journal of Dental Research 2025, Nov 2025).
  • Whisper's invented text hit speakers with aphasia harder: 1.7% of their audio segments produced hallucinations, against 1.2% for the control group, while Google's speech-to-text and Chirp models produced none on the same audio (Koenecke et al., ACM FAccT 2024, Jun 2024).
  • Whisper Large v2 had a 2.7% word error rate on LibriSpeech audiobooks but 25.5% on CHiME-6 dinner-party recordings, and 36.4% on meetings recorded by a single distant microphone against 16.9% for the same meetings on headset microphones (Radford et al., Whisper 2022, Dec 2022).

Is voice typing replacing the keyboard?

Speech beats thumbs on raw speed in lab tests, and phones now carry most web traffic. Medicine is where dictation use is best measured.

Line chart of words per minute for keyboard and speech input on a touchscreen phone, in English and Mandarin, from a lab study
Figure 5: Speech was far faster than the phone keyboard in both English and Mandarin, so the advantage holds across two very different keyboards. Source: Ruan et al., IMWUT 2017, Comparing Speech and Keyboard Text Entry for Short Messages in Two Languages on Touchscreen Phones, Dec 2017. Chart published by the source (Fig. 5, cropped from the PDF), captured 2 Oct 2026.

How much faster is speaking than typing?

Much faster in lab tests, in English and Mandarin alike, though speech leaves slightly more errors in the finished text. Ruan et al.'s lab study timed 48 university students transcribing short phrases, which the authors call upper-bound performance; composing a message while you think is slower. For keyboard speeds by age, device and training, see our typing speed statistics.

  • In Ruan et al.'s lab test on touchscreen phones, speech reached 123 words per minute in Mandarin Chinese against 43 on the Pinyin keyboard, 2.87 times faster (Ruan et al., IMWUT 2017, Dec 2017).
  • Speech also made fewer errors during entry: across both languages, the corrected error rate was 5.30% for speech against 11.22% for the keyboard (Ruan et al., IMWUT 2017, Dec 2017).
  • Speech left slightly more errors in the finished text, a 1.30% uncorrected error rate against 0.79% for the keyboard (Ruan et al., IMWUT 2017, Dec 2017).
  • Volunteers typing on their phones averaged 36.2 words per minute, with an uncorrected error rate of 2.3%, across 37,370 people who took an online typing test (Palin et al., MobileHCI 2019, Oct 2019).
  • On a physical desktop keyboard, 168,960 volunteers in an online typing test averaged 51.56 words per minute, and the fastest reached 120 or more (Dhakal et al., CHI 2018, Apr 2018).
  • People talk faster than either keyboard allows: in 2,438 recorded American telephone conversations, speakers averaged 164 words per minute across their own speaking turns, and in a larger set of recorded calls between strangers the topic alone moved the average between 152 and nearly 170 (Yuan et al., Interspeech 2006, Sep 2006).

How many people type and dictate on their phones?

Nearly every American adult has a smartphone, and phones carry most of the world's web traffic. No recent public primary source measures how many people dictate text rather than type it; the last US measure that split voice assistant use by device dates from 2017.

  • 91% of US adults own a smartphone, up from 35% in Pew's first measurement in 2011 (Pew Research Center, Nov 2025).
  • 16% of US adults are smartphone-only internet users, with a smartphone but no home broadband, rising to 27% of adults aged 18 to 29 and 34% of those in households earning under $30,000 (Pew Research Center, Nov 2025).
  • An estimated 262 million Americans aged 12 and older own a smartphone (Edison Research at SSRS, Mar 2026).
  • In 2017, 42% of US adults used a voice assistant on their smartphone, three times the 14% who used one on a computer or tablet (Pew Research Center, Dec 2017).
  • In the US the split is close to even: mobile devices took 49.1% of web traffic in September 2026 and desktops 48.47%, the first month since at least September 2025 with mobile ahead (StatCounter, Sep 2026).

Which professionals dictate most, and how much time does it save?

Doctors are the best-measured group. Ambient AI scribes, which listen to a visit and draft the note, save a few minutes of documentation time in trials, and the drafts still need checking.

  • At the time Microsoft agreed to buy it in 2021, Nuance's products were used by more than 55% of US physicians and 75% of US radiologists, and in 77% of US hospitals, by the companies' own count (Microsoft, Apr 2021).
  • The US had 42,000 medical transcriptionist jobs in 2025, and BLS projects employment to fall 4% by 2035; it names speech recognition software, which drafts a report for the transcriptionist to review, as the most common technology in the job (US Bureau of Labor Statistics, Aug 2026).
  • In a randomized trial with 238 UCLA outpatient physicians, one ambient AI scribe (Nabla) cut time spent in notes by 9.5% against usual care, while the other (Microsoft's DAX Copilot) made no significant difference (Lukac et al., NEJM AI 2025, Nov 2025).
  • Uptake was partial even inside that trial: DAX was used in 33.5% of 24,696 visits and Nabla in 29.5% of 23,653, and about 15% of physicians given a scribe never used it (Lukac et al., NEJM AI 2025, Nov 2025).
  • Before a randomized trial run in 2024 and 2025 across clinics in two states, 59% of the 66 clinicians enrolled already used Fluency Direct speech recognition software for their notes, and 4.5% used traditional dictation (Afshar et al., NEJM AI 2025, Nov 2025).
  • In the same trial, clinicians spent 0.36 fewer hours a day on notes with an ambient AI scribe, a secondary measure, and the authors report no loss in diagnosis, billing compliance or note quality (Afshar et al., NEJM AI 2025, Nov 2025).
  • In a randomized crossover trial of 160 outpatient clinicians comparing two ambient scribes, one saved 3.19 more minutes of note time per day than the other, and the two did not differ meaningfully in after-hours documentation time (Chowdhury et al., JAMIA 2026, Feb 2026).
  • Across 198,178 emergency department visits at four hospitals in 2025, ambient AI scribes were used in 4.3% and were associated with 1.6 fewer minutes of median physician documentation time per note, about half the 3.3-minute reduction associated with human scribes (Dutta et al., Annals of Emergency Medicine 2026, Jun 2026).
  • Among UK GPs who used an ambient scribe, 80% said it reduced time spent on documentation, but 32% reported errors often or always, and 14% reported errors with significant-to-critical implications (Blease et al., BMJ Health and Care Informatics 2026, Jul 2026).

How big is the voice recognition market?

Nobody agrees on the size of the voice recognition market, and the published estimates for 2025 differ by nearly a factor of two. Filings and deal prices are clearer: companies have paid billions for speech technology, and the listed voice AI companies are growing revenue at double-digit rates.

Bar chart of Mordor Intelligence's voice recognition market size estimate in USD billion for 2025 and 2026 and its forecast for 2031
Figure 6: Mordor Intelligence forecasts steep growth, but its 2025 starting point sits far above another firm's estimate for the same year, which is why every market number on this page names the firm that produced it. Source: Mordor Intelligence, Voice Recognition Market Size, Share and Growth, 2031, Sep 2026. Chart published by the source (CC BY 4.0), captured 2 Oct 2026.

What is the speech and voice recognition market worth?

It depends which firm you ask. The three estimates below define the market slightly differently and should never be averaged; they agree only that it is growing by roughly a fifth a year.

  • MarketsandMarkets valued the global speech and voice recognition market at $9.66 billion in 2025 and forecasts $23.11 billion by 2030, a 19.1% compound annual growth rate (MarketsandMarkets, Aug 2025).
  • MarketsandMarkets names voice search as the largest application in the market, driven by smartphones, smart speakers and AI assistants (MarketsandMarkets, Aug 2025).
  • Fortune Business Insights sizes the same market at $19.09 billion in 2025 and $23.70 billion in 2026, and forecasts $104.05 billion by 2034, a 20.30% compound annual growth rate (Fortune Business Insights, Sep 2026).
  • Fortune Business Insights projects the US speech and voice recognition market to reach $6.01 billion in 2026 (Fortune Business Insights, Sep 2026).
  • Speech recognition, as distinct from voice (speaker) recognition, will hold 66.40% of the market in 2026, according to Fortune Business Insights (Fortune Business Insights, Sep 2026).
  • Mordor Intelligence puts the voice recognition market at $18.39 billion in 2025 and $22.51 billion in 2026, and forecasts $61.78 billion by 2031, a 22.38% compound annual growth rate (Mordor Intelligence, Sep 2026).
  • Mordor Intelligence ranks Asia Pacific as the largest region, with 37.64% of voice recognition revenue in 2025 (Mordor Intelligence, Sep 2026).
  • Smartphones and tablets accounted for 39.17% of voice recognition revenue in 2025, the largest device category, according to Mordor Intelligence (Mordor Intelligence, Sep 2026).
  • Medical documentation is the fastest-growing voice recognition application in Mordor Intelligence's forecast, at a 23.39% compound annual growth rate to 2031 (Mordor Intelligence, Sep 2026).

How fast are voice AI companies growing?

Quickly. The listed voice companies report double-digit revenue growth, and ElevenLabs doubled its valuation in seven months.

  • SoundHound AI's revenue was $168.9 million in 2025, up 99% from 2024 (SoundHound AI, Feb 2026).
  • SoundHound AI's revenue for the second quarter of 2026 was $61.9 million, up 45% year over year; its CEO says that is 10 times its quarterly revenue when it listed in 2022 (SoundHound AI, Aug 2026).
  • Cerence, which supplies voice assistants to carmakers, reported fiscal 2025 revenue of $251.8 million and guided to $300 million to $320 million for fiscal 2026, a 23% rise at the midpoint that includes a patent licence payment from Samsung (Cerence, Nov 2025).
  • Cars built with Cerence technology made up 50% of worldwide auto production in the 12 months to June 2026 (Cerence, Aug 2026).
  • ElevenLabs ended 2025 with $350 million in annual recurring revenue and passed $500 million in the first four months of 2026 (ElevenLabs, May 2026).
  • A $300 million employee tender offer in September 2026 valued ElevenLabs at $22 billion, double its valuation at its February 2026 Series D (ElevenLabs, Sep 2026).
  • ElevenLabs' voice agents handle more than 15 million conversations a week, three times as many as in February 2026, and enterprise customers bring in 55% of its revenue (ElevenLabs, Sep 2026).

What have big tech companies paid for voice technology?

Microsoft's Nuance deal set the benchmark, and smaller voice AI companies are now buying their way into customer service.

  • Microsoft said the Nuance purchase would double its addressable market among healthcare providers, bringing its healthcare TAM to nearly $500 billion (Microsoft, Apr 2021).
  • When the deal closed on March 4, 2022, Microsoft recorded a total purchase price of $18.8 billion for Nuance, $16.3 billion of it goodwill (Microsoft, Jul 2022).
  • SoundHound AI valued its September 2025 purchase of Interactions, a customer-service AI company, at a preliminary $76.1 million (SoundHound AI, Mar 2026).
  • LivePerson, which SoundHound AI agreed to buy in April 2026, runs digital customer engagement that carries one billion customer messages a month (SoundHound AI, Apr 2026).
  • SoundHound AI completed the LivePerson acquisition on September 4, 2026, giving the combined company a customer base that includes 25 of the Fortune 100 and more than 750 patents (SoundHound AI, Sep 2026).

Glossary

  • ASR: automatic speech recognition, software that turns spoken audio into text.
  • Word error rate (WER): the share of words a system gets wrong (substituted, dropped or added) against a human reference transcript. Lower is better, and rates compare only within the same test set.
  • Switchboard, CallHome, LibriSpeech, CHiME-6, AMI, FLEURS: standard speech test sets: telephone conversations (Switchboard, CallHome), read audiobooks (LibriSpeech), dinner-party recordings (CHiME-6), meetings (AMI) and a multilingual benchmark (FLEURS).
  • Fine-tuning: further training of an existing model on a smaller, targeted set of audio, such as child speech.
  • Hallucination: text a speech recognition model outputs that was never spoken.
  • Smart speaker: a speaker with a built-in voice assistant, such as Amazon Echo, Google Nest or Apple HomePod.
  • Ambient AI scribe: software that listens to a clinical visit and drafts the clinician's note for review.
  • Dysarthria: a motor speech disorder in which weak or poorly controlled speech muscles make speech hard to understand.
  • Aphasia: a language disorder, often caused by a stroke, that affects speaking and understanding language.
  • GP: general practitioner, a family doctor in the UK.
  • API: application programming interface, the way developers connect their software to a model.
  • ARR: annual recurring revenue.
  • TAM: total addressable market.
  • Compound annual growth rate: the steady yearly growth rate that would take a market from its start value to its forecast value.

Methodology and Sources

Every number on this page comes from the organisation that produced it: survey firms (Edison Research, Pew Research Center, BrightLocal, Adobe), statistics offices and regulators (Eurostat, BLS, FTC), peer-reviewed and preprint research papers, company filings and newsrooms, and market research firms named as such. No figure was taken from another statistics roundup. Each link was fetched on 2 Oct 2026 and the number confirmed on the page or in the downloaded report. Market estimates disagree, so each one is attributed to its firm and shown separately, never averaged. Studies from 2017 to 2022 are included where nothing newer measures the same thing, and their year is stated in the sentence. The page is reviewed and refreshed every month.

Sources

Cite this page

Every figure here links to the study that produced it. Quote any of them with a link to this page or the original source.

WriteVoice. "139 Voice Search Statistics October 2026." WriteVoice, October 2026. https://www.writevoice.io/statistics/voice-search-statistics/
<a href="https://www.writevoice.io/statistics/voice-search-statistics/">Voice Search Statistics</a> (WriteVoice, October 2026)