107Course · Foundation
Local AI and Local Models
Running intelligence on your own hardware.
Running intelligence on your own hardware. GPU economics, quantization formats (GGUF, AWQ, GPTQ), and the runtimes — Ollama, LM Studio, vLLM, llama.cpp — with a Lab that serves a real local model.
- 01
Run a modern model locally.
- 02
Understand GPU economics.
- 03
Choose a quantization defensibly.
- 04
Know when local beats API.
Prerequisites
- — How LLMs Work.
Pacing
~4 hours including the Lab.
- IModule
Why local
The cases.
- IIModule
Quantization
Fitting big models on smaller hardware.
- IIIModule
Inference engines
The runtimes.
A Note from the Faculty
Local models are the reason the field cannot be captured.