Blog / Quantization
1 article

Quantization

All posts tagged with #Quantization

Your Next AI App Might Not Need the Cloud — POCKET 35B Proves It

Your Next AI App Might Not Need the Cloud — POCKET 35B Proves It

A 35-billion-parameter sparse MoE model that runs on iPhones and GPU-less PCs at 20 tokens per second. We break down why sparse Mixture-of-Experts finally makes on-device AI viable, what it means for privacy and cost, and which product features you should move local first.

Read Article