158
Tokens/sec aggregate
Gemma 3 270M, 16 concurrent users, measured.
NavyaAI Research - Measured Release - August 2026
Physical benchmarks for Gemma 3 small language models on a $499 Jetson Orin Nano 8GB: throughput, time-to-first-token, power, and the exact volume where the board beats a cloud API. Every number measured on hardware — including the configurations that failed.
By Ramachandra Vikas Chamarthi, Founder, NavyaAI. Benchmarks run June 2026, published August 7, 2026.
The short answer
For small models, yes — measured, not estimated. Gemma 3 270M sustains 158 tokens/sec aggregate across 16 concurrent users at 8W with sub-100ms first-token latency; Gemma 3 1B reaches 88 tokens/sec. At ~$14/month all-in, one board undercuts GPT-4o-mini once volume passes ~24M tokens/month. The limit is equally clear: 4B-class models fail under concurrent load on 8GB — we publish those failures with raw data instead of hiding them.
Slide 2 - The Numbers
158
Gemma 3 270M, 16 concurrent users, measured.
<100ms
42-91ms across all 270M load levels.
~24M
Where one $14/mo board beats GPT-4o-mini pricing.
Slide 3 - The Insight
Core Reality
270M-1B models serve real multi-user traffic on a $499 board under 10W. 4B models don't, on 8GB, under concurrency. The full report shows both sides with measured data.
Where the other 60-80% hides
Slide 4 - CTA
Instant access to the full measured report plus raw CSVs (cost model, crossover table, run statistics, and failure logs) — free, and we email you a permanent download link.
Gemma 3 270M, 1B, 4B and FunctionGemma via Ollama (Q4_K_M) on a Jetson Orin Nano 8GB: throughput, TTFT, latency percentiles, GPU utilization, RAM, power, and temperature at 1-16 concurrent users, in both serialized and true-parallel serving modes.
Configurations that failed are reported as failures with raw logs — zero completed requests for 4B-class models under concurrency. No estimated edge numbers appear anywhere in the report, and the CSVs ship unedited so you can check our work.
Teams weighing edge deployment for kiosks, factories, retail, robotics, or offline/private workloads — and anyone comparing a fleet of $499 boards against per-token API pricing or cloud GPU rentals.
Size your own fleet
Our free Edge LLM Sizing Agent uses this report's measured throughput, power, and cost data to tell you whether an edge fleet fits your token volume, latency target, and site count — and where the break-even sits against your current API bill.
Run the Edge LLM Sizing AgentRelated research
Edge is one answer to rising inference bills. Our AI Cost Report explains the other side: why token prices collapsed 99.7% while AI bills tripled, and where cloud spend actually hides.
Open the AI Cost Report