All Reports

NavyaAI Research - Measured Release - August 2026

Jetson Orin Nano LLM Benchmark: Edge Inference, Measured

Physical benchmarks for Gemma 3 small language models on a $499 Jetson Orin Nano 8GB: throughput, time-to-first-token, power, and the exact volume where the board beats a cloud API. Every number measured on hardware — including the configurations that failed.

By Ramachandra Vikas Chamarthi, Founder, NavyaAI. Benchmarks run June 2026, published August 7, 2026.

The short answer

Can a $499 board really serve production LLM traffic?

For small models, yes — measured, not estimated. Gemma 3 270M sustains 158 tokens/sec aggregate across 16 concurrent users at 8W with sub-100ms first-token latency; Gemma 3 1B reaches 88 tokens/sec. At ~$14/month all-in, one board undercuts GPT-4o-mini once volume passes ~24M tokens/month. The limit is equally clear: 4B-class models fail under concurrent load on 8GB — we publish those failures with raw data instead of hiding them.

Slide 2 - The Numbers

158

Tokens/sec aggregate

Gemma 3 270M, 16 concurrent users, measured.

<100ms

Time to first token

42-91ms across all 270M load levels.

~24M

Tokens/mo crossover

Where one $14/mo board beats GPT-4o-mini pricing.

Slide 3 - The Insight

Core Reality

Small models are the edge sweet spot — and we can prove where the edge ends.

270M-1B models serve real multi-user traffic on a $499 board under 10W. 4B models don't, on 8GB, under concurrency. The full report shows both sides with measured data.

Where the other 60-80% hides

  • Which model sizes actually serve concurrent users on 8GB
  • Full throughput/latency/power tables for 270M and 1B
  • The 4B and FunctionGemma failure data, unedited
  • Complete cost model: amortization, duty cycle, electricity
  • API vs edge crossover table from 0 to 100M tokens/month
  • Raw benchmark CSVs to rerun the analysis yourself

Slide 4 - CTA

Get the Benchmark Data Pack.

Instant access to the full measured report plus raw CSVs (cost model, crossover table, run statistics, and failure logs) — free, and we email you a permanent download link.

What was measured

Gemma 3 270M, 1B, 4B and FunctionGemma via Ollama (Q4_K_M) on a Jetson Orin Nano 8GB: throughput, TTFT, latency percentiles, GPU utilization, RAM, power, and temperature at 1-16 concurrent users, in both serialized and true-parallel serving modes.

What's honest about it

Configurations that failed are reported as failures with raw logs — zero completed requests for 4B-class models under concurrency. No estimated edge numbers appear anywhere in the report, and the CSVs ship unedited so you can check our work.

Who it's for

Teams weighing edge deployment for kiosks, factories, retail, robotics, or offline/private workloads — and anyone comparing a fleet of $499 boards against per-token API pricing or cloud GPU rentals.

Size your own fleet

Would these numbers hold for your workload?

Our free Edge LLM Sizing Agent uses this report's measured throughput, power, and cost data to tell you whether an edge fleet fits your token volume, latency target, and site count — and where the break-even sits against your current API bill.

Run the Edge LLM Sizing Agent

Related research

The cloud-side cost story.

Edge is one answer to rising inference bills. Our AI Cost Report explains the other side: why token prices collapsed 99.7% while AI bills tripled, and where cloud spend actually hides.

Open the AI Cost Report
Size your edge LLM fleet — free