OneRange

Voice & Conversational AI

Learners understand and apply speech and voice AI in products and workflows. The course covers speech-to-text and text-to-speech, voice assistants, latency and UX, transcription, and use cases and limits. Participants finish able to evaluate and apply voice AI sensibly.

What this course covers

The course runs across 5 topics, each one a short adaptive session rather than a recorded lecture. The tutor explains the idea, works an example, checks understanding, and adjusts the next step to the answer given.

  • Speech-to-Text (STT) Pipelines and Transcription — Understand acoustic and language models. Hands-on practice configuring STT APIs, handling audio formats, and optimizing transcription accuracy using code.
  • Text-to-Speech (TTS) and Voice Synthesis — Explore speech synthesis techniques, SSML (Speech Synthesis Markup Language), and voice cloning. Implement a program to convert text into natural-sounding speech.
  • Building Voice Assistants and Dialogue Flows — Design and build a functional voice assistant. Connect STT and TTS engines with a natural language processing (NLP) backend to handle user intents.
  • Optimizing Latency and Real-Time Voice UX — Analyze latency bottlenecks in voice pipelines. Implement streaming audio input/output, chunking, and caching techniques to achieve sub-second response times.
  • Voice AI Evaluation, Limits, and Use Cases — Evaluate voice AI systems using Word Error Rate (WER) and latency metrics. Explore production limits, ethical considerations, and design patterns for high-value use cases.

Skills you build

Results are measured against named skills in the OneRange taxonomy of more than 10,000 skills, so a manager sees proficiency per skill rather than a completion tick. This course maps to Voice User Interface (VUI) Design, Voice & Audio Content Strategy, Understanding AI Limitations, Text-to-Speech (TTS), Speech Recognition.

  • Voice User Interface (VUI) Design
  • Voice & Audio Content Strategy
  • Understanding AI Limitations
  • Text-to-Speech (TTS)
  • Speech Recognition

Who it is for

Support, product, builders. The material is pitched at intermediate level, and takes roughly 120 minutes at a typical pace. Because every session adapts, someone who already knows a topic moves through it quickly instead of sitting through an explanation they do not need.

It sits in the Category: Multimodal track of the Vero AI Catalog, and can be assigned to one person, a team, or the whole company.

How it is delivered and assessed

Delivery is conversational and interactive, including code lab, guided lab exercises. There is no video to sit through and no slide deck to click past.

Understanding is checked with a lab assessment of 5 items, with a pass mark of 70%.

Administrators can copy this course into their own library and adapt it — edit the outline, change the duration, swap the assessment format, or ground it in internal documentation so answers cite the company's own source material.

Also in this track

  • Multimodal AI: Image, Audio & Video — This course shows how to use AI across images, documents, voice, and video responsibly. It covers image generation and editing, vision and document understanding, voice and audio, and video use cases. Learners finish able to work confidently beyond text.
  • Creative & Multimodal AI — Produce images, audio, and video with AI

Related

Frequently asked questions

How long does this course take?

It runs roughly 120 minutes at a typical pace. Because every session adapts, someone who already knows a topic moves through it quickly instead of sitting through an explanation they do not need.

How is it assigned to a team?

An administrator assigns it from the Vero dashboard to one person, a team, a department or the whole company, and sees progress and results per person and per skill.

Can we customise it with our own documents and terminology?

Yes. Administrators can copy this course into their own library and adapt it — edit the outline, change the duration, swap the assessment format, or ground it in internal documentation so answers cite the company's own source material.

Course code CAT-VOICE.