What is agentic voice?
Every vendor says "agentic voice." Almost none of them mean a voice agent that another agent created an hour ago. This is about that version: what changed in September to make it real, what keeps a spawned agent inside the company's rules, and why it cannot be tied to the platform that built it.

Luke Miller
Co-founder

The agent that did not exist an hour ago
"Agentic voice" means two different things depending on who is selling it. Most vendors mean a voice agent that reasons and calls tools mid-call. Every production voice agent does that now. The label tells you nothing.
I mean the narrower version: a voice agent that is written, deployed and retired by another AI agent. A CRM workflow needs to phone a customer about an overdue invoice. The voice agent that makes the call did not exist an hour ago.
Why now
GPT-Live-1 shipped on 10 September and hands deeper reasoning to a backend you choose. When OpenAI's own voice model is designed to delegate the thinking, the voice model is a component.
A buyer now picks from three or four hosted voice models, five to ten reasoning backends in a given region, two or three frameworks people actually ship on, and then where it all runs: a hyperscale region, a national cloud, a VPC, on-prem. When an agent builds the voice agent, it makes every one of those picks with no person watching.
The rules have arrived
EU AI Act Article 50 has applied since 2 August 2026: anyone interacting with an AI system has to be told. In June the RBI published draft rules that would require Indian banks and NBFCs to disclose AI on customer-facing channels and hand over to a human on request. Saudi Arabia closed consultation on a Responsible AI Policy in May. The UAE has binding AI transparency rules inside DIFC and non-binding principles everywhere else.
Building one use case by hand still takes months of prompt work and vendor calls. Salesforce now creates voice agents from Agentforce Builder. Microsoft does it from Foundry Agent Service. Once a voice agent is something an API creates, it is something another agent creates.
UNMUTE stops a spawned agent going wrong
Any LLM can write a voice agent prompt in seconds. Left alone, it will also pick whatever model, region and tools it likes.
In UNMUTE the check is the manifest. The company writes it once: approved providers and models per role, languages, model and deployment regions, targets, tool kinds, tracing. You can write it in the editor, or run unmute skill install and let the /unmute-manifest skill interview you about what the company allows and save the contract. For an Indonesian collections agent it might look like this:
manifest: acme-corp
version: 1
models:
listen:
- provider: slng
allow: [slng/deepgram/nova:3-id]
languages:
allow: [id]
regions:
deployments:
- provider: slng
allow: [id]
targets:
allow: [slng, livekit]
tools:
kinds:
allow: [builtin, webhook]The coding agent starts from a package that already carries the contract:
shunmute manifest create acme-corp
unmute init collections-id --from-manifest
unmute validate collections-id
unmute compile collections-idPick a model or region the manifest does not list and validation fails before any output is written. The UNMUTE skill tells the assistant to fix its own choices and explain the conflict, not weaken the rule. The compile report records the contract name and revision, so an agent built by another agent leaves the same trail as one built by hand.
The manifest checks declared settings. It does not read local tool code, it cannot prove where a provider processes audio, and it is not signed yet. Rule packs for things like Indonesia's collections calling window (Monday to Saturday, 08:00 to 20:00 in the debtor's time zone, not on national holidays) are the next layer.
Design, selection, execution
I split voice agent operations into design time, selection time and execution time. The first two turn out to be one request. A workflow asks for a voice agent for a job. If one inside the manifest already fits, route to it. If not, compose one inside the same manifest. The workflow never needs to know which happened.
Execution is the call itself: per-turn routing between voice model and reasoning backend, caching, region choice, and a region boundary that requests cannot cross. SLNG runs that part today in 13 regions: us-east, us-west, br, eu-west, eu-north, gb, za, il, jp, sg, id, in, au.
Where it runs
An UNMUTE package compiles to four targets. LiveKit and Pipecat are code targets: you get a Python project to run on LiveKit Cloud, Pipecat Cloud or your own infrastructure. SLNG is a hosted target: compile writes a deployment body and SLNG runs the agent. Twilio compiles to a small app that ConversationRelay calls. A fresh package scaffolds to LiveKit with SLNG speech models bound by default, and the manifest's targets.allow decides which of the four a company permits.
An agent running on SLNG today can move to LiveKit Cloud or a company's own VPC without being rewritten, as long as it does not use SLNG-only tools such as the send_sms builtin.
That matters more when agents are doing the building. If a voice agent only runs on the platform that created it, whoever owns the platform owns the agent.
The manifest, the CLI and the four targets are all in the open-source UNMUTE repo. Save a contract with unmute manifest create, point your coding agent at it, and see what validation refuses.
Get all of our updates directly to your inbox. Sign up for our newsletter.