- Link to Introduction to Local AI (101) https://www.youtube.com/watch?v=wRcByxXkJCQ - ODS https://github.com/osmantic/ods - Inference Engines blog series http://thelocalaibook.osmantic.com/ Find our speakers below: - https://x.com/TheAhmadOsman - https://x.com/MikeBradleyAI - https://x.com/alexocheema 00:00 Introduction to Local AI 201 01:38 Meet the Guests: Ahmed Osman & Michael from Osmantic 03:13 - Understanding Inference Engines for LLMs 04:41 Prefill vs. Decode: Two Phases of Inference Requests 05:20 Hardware Bottlenecks: Memory Bandwidth vs. Capacity 06:54 Multi-Device Hardware Arena Lab Overview 07:54 The Critical Role of Kernel and Engine Selection 08:52 Overview of Engine Families (llama.cpp, VLLM, SGLang) 10:57 Decision Guide: 9 Questions Before Picking Your Stack 12:15 Why Run Locally? Cost, Privacy, and Control 14:48 Live Demo: Single Instance Hardware Race 19:27 Live Demo: Burst Mode & Continuous Batching Performance 22:04 Software Optimization: M5 MacBook Pro Performance Analysis 24:19 Demo Race the Cloud: Local Qwen 3.6 vs. Remote GPT-5 26:41 Introduction to ODS (Open Deployment System) 31:32 Clustering Multiple GPUs with Exo 39:12 Alex (Founder of Exo) on Personal Computing V2 43:45 The Maturity of the Local AI Industry 45:51 Benchmarking with Local.ai: Intelligence vs. Speed Trade-offs 50:38 Testing Across 100+ Hardware Configurations 52:48 Outro and Future Local AI Series Announcements
- Link to Introduction to Local AI (101) https://www.youtube.com/watch?v=wRcByxXkJCQ
- ODS https://github.com/osmantic/ods
- Inference Engines blog series http://thelocalaibook.osmantic.com/
Find our speakers below:
- https://x.com/TheAhmadOsman
- https://x.com/MikeBradleyAI
- https://x.com/alexocheema
00:00 Introduction to Local AI 201
01:38 Meet the Guests: Ahmed Osman & Michael from Osmantic
03:13 - Understanding Inference Engines for LLMs
04:41 Prefill vs. Decode: Two Phases of Inference Requests
05:20 Hardware Bottlenecks: Memory Bandwidth vs. Capacity
06:54 Multi-Device Hardware Arena Lab Overview
07:54 The Critical Role of Kernel and Engine Selection
08:52 Overview of Engine Families (llama.cpp, VLLM, SGLang)
10:57 Decision Guide: 9 Questions Before Picking Your Stack
12:15 Why Run Locally? Cost, Privacy, and Control
14:48 Live Demo: Single Instance Hardware Race
19:27 Live Demo: Burst Mode & Continuous Batching Performance
22:04 Software Optimization: M5 MacBook Pro Performance Analysis
24:19 Demo Race the Cloud: Local Qwen 3.6 vs. Remote GPT-5
26:41 Introduction to ODS (Open Deployment System)
31:32 Clustering Multiple GPUs with Exo
39:12 Alex (Founder of Exo) on Personal Computing V2
43:45 The Maturity of the Local AI Industry
45:51 Benchmarking with Local.ai: Intelligence vs. Speed Trade-offs
50:38 Testing Across 100+ Hardware Configurations
52:48 Outro and Future Local AI Series Announcements