Local AI and Local Models107 · Module I · Lesson 03 of 11
Article · 12 min

Consumer vs. datacenter GPUs

The honest comparison.

Summary

The honest comparison. Consumer GPUs (RTX 4090, 5090) are cheap per FLOP but capped at 24-32 GB VRAM. Datacenter GPUs (H100, MI300X) cost 10-20x more but offer 80-192 GB VRAM and 3-5x bandwidth. Consumer is right for prototypes and small models; datacenter for serving large models to many users.

Objectives
  • 01State the VRAM ceiling of the top consumer GPU.
  • 02Compare $/token throughput consumer vs. datacenter.
  • 03Name the middle-ground options.
The Lesson

The numbers

RTX 5090: 32 GB VRAM, ~$2,000, ~1.8 TB/s bandwidth. H100 80GB: ~$25,000-$30,000, ~3.35 TB/s. On per-dollar throughput of small models, consumer wins. On per-dollar throughput of large models that don't fit on consumer hardware, datacenter is the only option.

Middle ground

Mac Studio M2 Ultra (192 GB unified memory, ~800 GB/s): unique product for local inference of large models on modest budgets. NVIDIA workstation cards (RTX 6000 Ada, 48 GB): expensive but consumer-adjacent. Rented cloud H100s: no capex, pay per hour.

Key Ideas
  • Consumer for < 30B params; datacenter for > 70B params.
  • Mac Studio is a genuinely unusual sweet spot.