Which AI vendors actually do voice, and how to choose

Text strength and voice strength are not the same thing. Grok publishes speech to speech at about 6p a minute. Anthropic sells no voice product at all.

John Kelleher
John Kelleher

Last checked 12 Sep 2026. Every price and date below came from the vendor's own pricing or deprecation page on that date, or from the Information Commissioner's Office. Prices are converted from US dollars at £1 = $1.35 on 11 Sep 2026; vendors bill in dollars and UK VAT applies on top.

Two prospects asked us the same thing this summer. One wanted to know whether we would augment a text-based sales agent with voice AI. The other wanted to know which channel to use per contact, phone, text or WhatsApp, because a chunk of their audience will not answer a call from a number they do not recognise.

Underneath both sits a harder question. If you have already chosen an AI vendor for writing, summarising and reasoning over your CRM data, does that choice carry over to voice?

It does not. Text strength and voice strength are different capabilities, priced in different units, shipped on different timelines, and in one case not shipped at all. Choosing a voice vendor by asking which model writes the best email is how businesses end up with a phone line that stumbles over interruptions.

What each vendor actually sells for voice

First-party voice capability published on each vendor's own pages, checked 12 Sep 2026. Third-party platforms that wrap these models are out of scope here.

VendorFirst-party voice productPublished price, convertedHow it is billed
OpenAI Realtime API, models gpt-realtime-2.1 and gpt-realtime-2.1-mini gpt-realtime-2.1: about £23.70 per million audio input tokens, about £47.40 per million audio output tokens, about 30p per million cached audio input tokens. The mini model: about £7.40 in, about £14.80 out. Per million tokens, audio and text billed separately. Text on gpt-realtime-2.1 is about £2.96 in and £17.78 out per million.
SpaceXAI (Grok) Speech to speech (grok-voice-think-fast-2.0), speech to text, text to speech Speech to speech about 5.9p per minute (about £3.56 an hour), plus a separate text input charge. Speech to text about 7p an hour over REST, about 15p streaming. Text to speech about £11.10 per million characters. Per minute or per hour of audio, the unit a contact centre already budgets in.
Google Live API, in Preview. The audio models named on Google's models page are Gemini 3.1 Flash Live, Gemini 2.5 Flash Live, Gemini 3.5 Live Translate, Gemini 3.1 Flash TTS and Gemini 3.5 Transcribe. Not verified in this session, so not quoted here. Gemini 3.8 Flash is not among the Live or native-audio models on that page. The Live API page states: "The Live API is in Preview."
Anthropic None. Not applicable. Anthropic's pricing page prices text, caching, batch, tool use and managed agent runtime across Fable 5.1, Mythos 5.1, Opus 5, Sonnet 5 and Haiku 4.5. No audio modality appears on it. A statement of scope, not a criticism.
Microsoft Copilot reaches staff and consumers as a voice-capable assistant. Not verified in this session. No first-party per-minute voice-agent price confirmed in this check, so treat Microsoft as a question for your account team rather than a costed option.
Apple Siri, a consumer assistant. Not applicable. Not a platform you build a business phone line on.

The dates that will move your build

OpenAI's deprecations page, checked 12 Sep 2026, carries two entries worth reading before signing a voice contract.

  • The Realtime API Beta shut down on 12 May 2026. The page notes that the beta and the released GA interfaces "have a few key differences", so a proof of concept written against the beta is not a proposal for what you would build today.
  • A whole generation of audio and realtime models shuts down on 20 Jan 2027: gpt-realtime, gpt-audio, gpt-4o-audio, gpt-4o-realtime, gpt-realtime-mini, gpt-audio-mini, gpt-4o-mini-realtime and gpt-4o-mini-audio. Separately, the transcription models whisper-1, gpt-4o-transcribe, gpt-4o-mini-transcribe and gpt-4o-transcribe-diarize shut down on 26 Feb 2027.

Commercially: a supplier quoting a voice build on gpt-realtime or whisper-1 today is quoting you a rebuild inside eighteen months. Ask which exact model ID they intend to run and check it against that list.

Why you cannot compare the headline prices directly

OpenAI bills realtime audio per million tokens. SpaceXAI bills speech to speech per minute. There is no published conversion between the two, because tokens per minute of speech depends on how fast people talk, how much the model says back and how much of the prompt is cached. So price the per-minute vendor from the rate card and measure the token-priced vendor in a pilot.

Worked example, 500 calls a month at four minutes each, so 2,000 minutes:

  • Speech to speech on grok-voice-think-fast-2.0: 2,000 minutes at about 5.9p is about £119 a month in model cost, plus the separate text input charge.
  • OpenAI Realtime: not derivable from the price list. Run 100 calls, read the token counts off the usage reporting, and multiply. Anyone who gives you a confident monthly figure before that has guessed.

Neither figure includes the part that usually costs more than the model: telephony. Numbers, SIP trunking, per-minute carriage, recording storage and the contact-centre platform all sit outside the AI bill, as does VAT. A pilot with £119 of model spend can still carry several hundred pounds of carriage, and carriage scales with minutes in a way caching will not rescue.

The decision method

Vendor choice is the last question, not the first. Five things decide it.

1. Channel preference per contact, before any model choice

Some people will not answer a phone. Hold a channel preference on the contact record and let the agent respect it, rather than defaulting everyone to voice because voice is the new capability. In our builds WhatsApp is usually a better default than SMS, domestically and internationally, on deliverability and because the thread persists where an SMS chain does not. That is a CRM data-model decision as much as an AI one, covered in our post on HubSpot WhatsApp options.

2. Latency and interruption handling

This is where voice vendors separate, and it does not show up on a price list. An agent that cannot be interrupted mid-sentence feels like an IVR and callers treat it like one. Test barge-in, two people speaking at once, and recovery from a caller who changes their mind halfway through a sentence, on your own accents and product vocabulary rather than the vendor's demo.

3. Telephony integration cost

Ask how the model reaches the phone network at all. Some vendors expose a session you bridge to a SIP trunk yourself. The cost of that bridge, and of who is on call when it drops, is a real line item.

4. Transcription and CRM logging

As far as the rest of the business is concerned, a call not logged against the record did not happen. Decide before the build whether you write a transcript, a summary or both, against which object, and what the next person sees when they open the record. Our post on what has to move when AI hands a customer over covers the handover payload.

5. Consent and call recording under UK law

Two separate things bite here, and businesses routinely conflate them.

The first is marketing calls. The ICO's guide to the Privacy and Electronic Communications Regulations, checked 12 Sep 2026, says the rules on automated calls are in regulation 19 "and are stricter. You must not make an automated marketing call, that is, a call made by an automated dialling system that plays a recorded message, unless the person has specifically consented to receive this type of call from you. General consent for marketing, or even consent for live calls, is not enough." Live marketing calls, under regulations 21, 21A and 21B, require screening against the Telephone Preference Service and Corporate TPS plus your own do-not-call list. PECR penalties run to £500,000, against the organisation or its directors.

Which of those two regimes a generative voice agent falls under is an open question, and one for your data protection officer rather than a vendor. It is not a recorded message in the ordinary sense, and that ICO guidance is currently marked as under review following the Data (Use and Access) Act. Get it answered in writing before an AI agent dials anyone for marketing.

The second is recording. Inbound service calls are not marketing, so the PECR marketing rules are not the issue, but UK GDPR transparency is: callers have to be told the call is recorded and why, and you need a lawful basis. If the transcript then feeds a model, that is a processing purpose to document and a retention period to set.

Procurement questions worth asking

  1. Which exact model ID will run this, and is it on the vendor's deprecation list?
  2. Is the price per minute or per token, and if per token, what did you measure in a pilot?
  3. What is the telephony cost, separately from the model cost?
  4. How does barge-in behave, demonstrated on our recordings rather than yours?
  5. What lands on the CRM record after the call, and who reads it next?
  6. What is the fallback when the model is unavailable mid-call?
  7. Which contacts should never be called, and where is that preference stored?

How we pick

SpotDev is an AI and digital transformation consultancy that builds real software. We are an OpenAI Select Partner and a Claude Registered Partner. Neither fact decides a voice build.

The method is to pick per workload. On a recent design we put one vendor on summarisation and a different one on client-facing email tone, because they were better at different jobs and nothing in the architecture required a single supplier. Voice is the clearest case of the same principle: the vendor with the strongest text model is not necessarily the right voice vendor, and one of the strongest text vendors sells no voice product at all.

To work that through against your own call volumes and contact data, start with the AI implementation page, or the AI for customer success page if the use case is support rather than sales. The vendor-specific detail sits on our OpenAI implementation and SpaceXAI implementation pages. There is a longer treatment of the Realtime API specifically in our post on OpenAI's voice API for UK business phone and support lines. When you are ready to price a build, Request a Quote.

Frequently asked questions

Does Anthropic offer a voice product?

No. As of 12 Sep 2026 Anthropic's pricing page lists no audio, speech-to-text, text-to-speech or realtime voice modality. It prices text, caching, batch, tool use and managed agent runtime only. If you want Claude in a voice workflow, the audio layer comes from elsewhere.

Which vendor publishes a per-minute voice price?

SpaceXAI does. Its developer pricing page lists speech to speech on grok-voice-think-fast-2.0 at about 5.9p per minute of audio, about £3.56 an hour, plus a text input charge. OpenAI prices realtime audio per million tokens instead, so a per-minute figure has to be measured rather than looked up.

Can we use Gemini 3.8 Flash for a voice agent?

Google's models page does not list Gemini 3.8 Flash among its Live API or native-audio models. The models named there for real-time dialogue are Gemini 3.1 Flash Live and Gemini 2.5 Flash Live, and Google's Live API page states that the Live API is in Preview.

Do we need consent before an AI agent phones someone?

For marketing calls, almost certainly yes in some form. The ICO's PECR guidance requires specific consent for automated marketing calls under regulation 19, and TPS and CTPS screening for live marketing calls under regulations 21, 21A and 21B. Whether a generative voice agent counts as an automated calling system playing a recorded message is not settled in that guidance, so get a written position from your data protection officer before dialling.

Will a voice build we commission today still run in 2027?

Only if it is on a current model. OpenAI's deprecations page lists gpt-realtime, gpt-audio, gpt-4o-audio, gpt-4o-realtime and their mini variants as shutting down on 20 Jan 2027, and whisper-1 plus the gpt-4o-transcribe family on 26 Feb 2027. Ask any supplier for the exact model ID before you sign.

Should we use phone, SMS or WhatsApp?

Per contact, not per campaign. Hold a channel preference on the record and let the agent respect it. We generally prefer WhatsApp over SMS for UK and international contacts, on deliverability and thread persistence, with voice reserved for contacts who have shown they will answer.

John Kelleher

John Kelleher

Author
John is the founder and the Chief Executive at SpotDev.

Stay Updated with Our Latest Insights

Get expert HubSpot tips and integration strategies delivered to your inbox.