Open-Source & Self-Hosted Models
This course helps technical teams evaluate, run, and govern open-source and self-hosted models. It covers open versus closed models, hosting and hardware, quantization, licensing, and data control and security. Learners finish able to decide when self-hosting makes sense and how to do it safely.
What this course covers
The course runs across 10 topics, each one a short adaptive session rather than a recorded lecture. The tutor explains the idea, works an example, checks understanding, and adjusts the next step to the answer given.
- Open-Source vs. Closed-Source Model Evaluation — Analyze trade-offs between proprietary APIs and open-weights models, focusing on performance, latency, and cost-benefit analysis for enterprise workloads.
- Licensing, Compliance, and Open-Source Governance — Navigate open-weights licenses (such as Llama 3, Apache 2.0, and OpenRAIL), understanding commercial use restrictions, derivative work obligations, and compliance risks.
- Hardware Selection and Infrastructure Sizing — Calculate VRAM requirements for model inference based on parameter size and precision. Compare cloud GPUs, on-premises clusters, and specialized hardware accelerators.
- Hands-On Model Quantization and Optimization — Apply quantization techniques (AWQ, GPTQ, and GGUF) to compress LLMs and run them efficiently on consumer-grade and enterprise hardware.
- Deploying Self-Hosted Inference Servers — Deploy high-throughput inference engines like vLLM or TGI locally and configure optimized serving parameters.
- Data Control, Privacy, and Network Security — Implement zero-trust network boundaries, secure local data pipelines, and prevent data exfiltration when deploying models in air-gapped or private VPC environments.
- Model Access Control and Enterprise API Integration — Set up secure API gateways, rate limiting, and role-based access control (RBAC) to safely expose self-hosted models to internal enterprise applications.
- Monitoring, Logging, and Observability in Self-Hosting — Set up monitoring pipelines to track token throughput, latency (TTFT), hardware utilization, and model drift in production environments.
- Architecting a Secure, Self-Hosted RAG System — Design and build an end-to-end Retrieval-Augmented Generation pipeline using a local vector database and a self-hosted LLM wrapper.
- Decision Framework: When to Self-Host — A synthesis exercise applying technical, financial, and regulatory constraints to formulate and defend a self-hosting roadmap for a mock enterprise scenario.
Skills you build
Results are measured against named skills in the OneRange taxonomy of more than 10,000 skills, so a manager sees proficiency per skill rather than a completion tick. This course maps to Large Language Models (LLMs), AI Risk Management, Model Quantization & Pruning, AI Governance, LLM Operations (LLMOps).
- Large Language Models (LLMs)
- AI Risk Management
- Model Quantization & Pruning
- AI Governance
- LLM Operations (LLMOps)
Who it is for
ML engineers, IT, security. The material is pitched at expert level, and takes roughly 120 minutes at a typical pace. Because every session adapts, someone who already knows a topic moves through it quickly instead of sitting through an explanation they do not need.
It sits in the Category: Open Source track of the Vero AI Catalog, and can be assigned to one person, a team, or the whole company.
How it is delivered and assessed
Delivery is conversational and interactive, including code lab, guided lab exercises. There is no video to sit through and no slide deck to click past.
Understanding is checked with a lab assessment of 5 items, with a pass mark of 70%.
Administrators can copy this course into their own library and adapt it — edit the outline, change the duration, swap the assessment format, or ground it in internal documentation so answers cite the company's own source material.
Also in this track
- Production AI Engineering — Take AI from prototype to reliable production
- AI Security Specialist — Defend AI systems and data
Related
- Browse the full Vero AI Catalog
- How the OneRange platform assesses and measures skill
- What OneRange Vero is
Frequently asked questions
How long does this course take?
It runs roughly 120 minutes at a typical pace. Because every session adapts, someone who already knows a topic moves through it quickly instead of sitting through an explanation they do not need.
How is it assigned to a team?
An administrator assigns it from the Vero dashboard to one person, a team, a department or the whole company, and sees progress and results per person and per skill.
Can we customise it with our own documents and terminology?
Yes. Administrators can copy this course into their own library and adapt it — edit the outline, change the duration, swap the assessment format, or ground it in internal documentation so answers cite the company's own source material.
Course code CAT-OSS.