Quick answer
Every voice‑AI business faces the same decision: do you prioritize ultra‑low latency, higher accuracy, or a hybrid of both? Voxovo 1.0 gives you three native stacks—Speed, Smart, and Ultra—that let you tailor your experience to the conversation’s demands.
The Speed stack (pipeline_mode: 'voxovo1') is engineered for the fastest turn time. It uses fixed language and voice models and is ideal when you need to handle thousands of concurrent calls with minimal delay.
The Smart stack (pipeline_mode: 'voxovo1_smart') swaps the same fixed models for a higher‑accuracy engine. Latency is a bit higher, but the trade‑off is more reliable intent detection and fewer misunderstandings—perfect for complex sales or support scripts.
How Voxovo AI works on Voxovo
The Ultra stack (pipeline_mode: 'voxovo0_ultra') pushes the envelope with Gemini live speech‑to‑speech. It delivers instant knowledge snippets at call start and keeps the core audio processing on Speed or Smart, giving you the best of both worlds.
Across all three modes, Presence Buffer keeps callers engaged by pre‑caching filler audio while the AI thinks, eliminating the dreaded dead air that plagues many voice bots.
Live Sheet and Live Calendar are first‑party, zero‑setup tools that let agents pull data or book appointments on the fly. They integrate directly into the call UI, so the agent never has to switch tabs or context.
You bring your own telephony—Twilio, Telnyx, Plivo, or Vonage—so Voxovo never holds or resells numbers. This BYO model gives you full control over number ownership, compliance, and carrier costs.
Implementation checklist
Voxovo also supports agency and multi‑workspace tenancy, with an omni‑sk client portal that allows a single admin to manage several brands or teams from one dashboard.
In addition to the core pipelines, Voxovo offers Outbound campaigns with automatic call distribution (AMD), warm transfers, and multi‑agent squad handoff to ensure seamless customer journeys.
The platform tracks latency scorecards per turn—STT, LLM, TTS, and total turn milliseconds—so you can continuously optimize performance and meet SLA targets.
All voice traffic remains within the PSTN endpoints you control, helping you meet GDPR, CCPA, and other industry‑specific regulations with data residency that aligns to your carrier.
Languages are supported through the fixed model selection, so you can deploy the same pipeline for multiple locales without re‑training. This simplifies global rollouts compared to custom LLMs.