OneRange

Multimodal AI: Image, Audio & Video

This course shows how to use AI across images, documents, voice, and video responsibly. It covers image generation and editing, vision and document understanding, voice and audio, and video use cases. Learners finish able to work confidently beyond text.

What this course covers

The course runs across 10 topics, each one a short adaptive session rather than a recorded lecture. The tutor explains the idea, works an example, checks understanding, and adjusts the next step to the answer given.

  • Introduction to Multimodal AI Workflows — Understand the transition from text-only to multimodal AI. Learn the core capabilities of image, voice, and video models and set up the hands-on workspace environment.
  • Image Generation: Prompting and Style Controls — Master text-to-image generation. Practice writing descriptive prompts, adjusting aspect ratios, and applying specific artistic styles directly inside the AI tool.
  • Image Editing: Inpainting and Outpainting — Modify existing images. Learn how to use brush tools to add/remove objects (inpainting) and expand canvas boundaries (outpainting) to match an existing style.
  • Vision and Document Understanding — Upload complex documents, charts, and images. Practice extracting structured data, translating handwritten notes, and querying visual information using conversational AI.
  • Analyzing Complex Visual Layouts — Work with multi-page PDFs containing mixed media. Learn to prompt the AI to compare diagrams, analyze flowcharts, and compile data tables from visual inputs.
  • Voice and Audio: Speech-to-Text and Transcription — Convert spoken audio files into clean text. Practice setting up transcription parameters, identifying different speakers, and generating automated meeting summaries.
  • Voice Generation and Audio Synthesis — Generate high-quality voiceovers from text. Explore voice cloning, adjusting emotional tone, pacing, and choosing the right vocal profile for different audiences.
  • Video Generation from Text and Images — Create short video clips using text prompts and reference images. Learn how to control camera movement, panning, and zoom settings within the generative video tool.
  • Video Editing and Practical Use Cases — Apply AI video tools to practical workflows like social media content creation, prototyping, and presentations. Learn how to edit video elements and stitch clips together.
  • Responsible Multimodal AI and Ethical Guardrails — Navigate the ethical considerations of multimodal content, including deepfakes, copyright, and bias. Complete a scenario-based exercise on verifying synthetic media.

Skills you build

Results are measured against named skills in the OneRange taxonomy of more than 10,000 skills, so a manager sees proficiency per skill rather than a completion tick. This course maps to AI Ethics, Generative AI Concepts, Critical Thinking, Data Analysis, Content Strategy.

  • AI Ethics
  • Generative AI Concepts
  • Critical Thinking
  • Data Analysis
  • Content Strategy

Who it is for

Creative, marketing, all users. The material is pitched at intermediate level, and takes roughly 120 minutes at a typical pace. Because every session adapts, someone who already knows a topic moves through it quickly instead of sitting through an explanation they do not need.

It sits in the Category: Multimodal track of the Vero AI Catalog, and can be assigned to one person, a team, or the whole company.

How it is delivered and assessed

Delivery is conversational and interactive, including guided lab, branching scenario exercises. There is no video to sit through and no slide deck to click past.

Understanding is checked with a lab assessment of 5 items, with a pass mark of 70%.

Administrators can copy this course into their own library and adapt it — edit the outline, change the duration, swap the assessment format, or ground it in internal documentation so answers cite the company's own source material.

Also in this track

  • Voice & Conversational AI — Learners understand and apply speech and voice AI in products and workflows. The course covers speech-to-text and text-to-speech, voice assistants, latency and UX, transcription, and use cases and limits. Participants finish able to evaluate and apply voice AI sensibly.
  • Creative & Multimodal AI — Produce images, audio, and video with AI
  • AI for Marketing Teams — Create, scale, and measure marketing with AI

Related

Frequently asked questions

How long does this course take?

It runs roughly 120 minutes at a typical pace. Because every session adapts, someone who already knows a topic moves through it quickly instead of sitting through an explanation they do not need.

How is it assigned to a team?

An administrator assigns it from the Vero dashboard to one person, a team, a department or the whole company, and sees progress and results per person and per skill.

Can we customise it with our own documents and terminology?

Yes. Administrators can copy this course into their own library and adapt it — edit the outline, change the duration, swap the assessment format, or ground it in internal documentation so answers cite the company's own source material.

Course code CAT-MM.