Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series Mac
github.com
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series Mac
A custom Swift + Metal runtime for any Apple Silicon Mac, even the 8 GB ones, that runs the instruction-tuned Gemma 4 26B-A4B without loading the entire 14.3 GB model into memory.