
Almost every article written about voice search optimisation traces back to a single forecast: that half of all searches would be spoken by 2020. It was a prediction rather than a measurement, it did not come true on that timeline, and an enormous amount of advice has been built on top of it since. Starting from a number nobody verified is a bad way to plan technical work.
This is one of the questions agencies ask us most often. For the full picture, see the technical growth work behind this.
What is true is quieter and more useful. People genuinely do ask questions out loud — to phones, to speakers, and increasingly to AI assistants that answer in prose rather than reading out a result. None of that created a separate index you can optimise for. It changed which pages get selected, and how much of a page survives the trip to the answer.
There is no separate voice index
This is the single most important thing to understand, because it invalidates a lot of what gets sold as voice search optimisation. When someone asks a phone a question, the assistant queries the same index as everyone else and then extracts an answer from a result. There is no voice-specific crawler, no separate ranking system, and no markup that registers you as voice-ready.
So the work is not a new discipline bolted onto your site. It is ordinary technical SEO, judged by a harsher standard: instead of appearing in a list of ten links where a human decides what to click, your page either gets extracted as the answer or it does not exist for that query.
What genuinely differs about a spoken query
| Typed | Spoken or asked of an assistant | |
|---|---|---|
| Length | Two to four words | A full sentence, often eight to twelve words |
| Grammar | Keyword fragments | Natural phrasing, usually a question |
| Intent signals | Ambiguous, inferred from context | Stated explicitly — who, where, when, how much |
| Local weighting | Sometimes | Frequently, and often implicitly |
| Results returned | A page of options | One answer, occasionally two or three |
| Tolerance for a slow page | Some | Effectively none |
The last row matters more than it looks. An assistant answering out loud will not wait for a page that takes four seconds to become useful. Performance stops being a ranking factor among many and becomes a filter.
The work that actually helps
All of the following is ordinary engineering. None of it is voice-specific, which is exactly the point — it improves the site for every visitor and happens to be what assistants need.
- Answer the question in the first forty to sixty words of the section, before the context and the caveats. Extraction takes the opening, not the paragraph where you finally get to the point.
- Write headings the way people ask, not the way marketers write. "How much does website maintenance cost?" is retrievable. "Pricing" is not.
- Give every distinct question its own heading and its own self-contained answer, so a passage can be lifted without the surrounding page.
- Use FAQPage structured data for genuine question-and-answer content, so the pairing is machine-readable rather than inferred from layout. Google's introduction to structured data is the reference for what it will and will not act on.
- Fix Core Web Vitals at source rather than caching over them — measure against Google's own definitions rather than a plugin's score.
- Keep local details — address, hours, service area — in structured data and not only in an image or a footer.
One correction worth making, because a lot of published advice is now out of date: Google restricted FAQ rich results in 2023 to well-known government and health sites, so adding FAQPage markup to a commercial page will not produce those expandable results in the SERP any more. It is still worth adding. The markup makes your question-and-answer pairs explicit to any machine reading the page, and the systems doing the reading have multiplied since that change.
What is a waste of effort
- Writing a separate page for spoken queries. It is the same index and you have just created a duplicate.
- Stuffing conversational long-tail phrases into copy. Keyword stuffing is measurably counterproductive with the systems that extract answers.
- Chasing speakable schema unless you are a news publisher — support has always been narrow and market-limited.
- Rewriting a site's tone to sound like speech. Assistants read structure, not personality.
- Buying a tool that reports a voice search readiness score. There is nothing on the other side of the API to measure it against.
Where this actually went: assistants, not speakers
The interesting shift was not smart speakers. It was that asking a question out loud and asking an AI assistant became the same act. A spoken question to a phone now frequently reaches a system that reads several sources and composes an answer, rather than one that reads a single result aloud.
That changes the objective in a way worth being clear about. Ranking gets you considered. Being extractable gets you cited. A page that ranks third but states its answer plainly, in a self-contained passage, with its claims sourced, will be quoted more often than a page that ranks first and buries the answer under six paragraphs of preamble.
This is why the advice above is unglamorous and structural. Clear headings, direct answers, honest sourcing and a fast page are what both a human skim-reader and a machine extractor need. Optimising for one is optimising for the other.
A reasonable order to do it in
- Fix performance first. Nothing else matters if the page is not usable quickly.
- Audit your headings against real questions — pull them from your own support inbox and sales calls rather than a keyword tool.
- Restructure the pages that already rank on page two. They are closest to being selected and the rewrite is cheapest.
- Add structured data where the content genuinely is a question and answer, an organisation, a product or a service.
- Then measure. If nothing changes in three months, the problem is the content rather than the markup.
If you want this done rather than advised on, that is the distinction we work to — technical SEO implementation means shipped fixes rather than a report with recommendations. Performance in particular decays quietly as a site accumulates plugins and scripts, which is why it sits inside ongoing maintenance and support rather than being treated as a one-off project. Agencies selling this work to their own clients can partner with us and deliver it under their brand.


