Does NavyaAI provide the data center or colocation facility?
No — and that is deliberate. We deploy into your data center or the colocation facility you choose in your jurisdiction, so residency, physical access, and contractual control stay entirely yours. We advise on facility selection and handle everything from the rack up: hardware specification, deployment, serving stack, and operations.
How does on-prem AI deployment help with GDPR?
It removes the hardest problem — cross-border data transfer — by architecture. Prompts, documents, embeddings, and model outputs never leave your infrastructure in your jurisdiction, so there is no US-cloud transfer to assess, no third-party AI processor in the chain for that data path, and a much simpler DPIA. We design the deployment so your DPO can draw the data-flow diagram in one box.
What does a private LLM deployment cost to run?
From our measured data: an edge-class deployment runs about $14/month per board all-in; a 70B-class production stack served INT8 at high utilization measured ~$0.47 per million tokens plus a $1,500-$2,500/month operations floor; hardware CAPEX depends on tier. The consult produces a concrete TCO for your workload — the same math as our public on-prem cost calculator.
Which models can run on-premise?
Open-weight models — Llama, Gemma, Mistral, Qwen class — from 270M edge models to 70B+ production models, quantized to fit your hardware tier. Model choice is driven by your quality bar and measured throughput on your target hardware, which we benchmark before committing.
Can you operate the stack after deployment?
Yes — either full handover to your team (docs, runbooks, training) or a managed arrangement: monitoring, upgrades, capacity planning, and incident response, with all telemetry staying inside your network. Most regulated clients start managed and take over operations within a year.
We're not in the EU — is this still relevant?
Yes. The same architecture serves HIPAA-adjacent US healthcare workloads, financial-services data that cannot leave the firm, air-gapped and defense-adjacent environments, and any team whose unit economics beat API pricing at their volume — typically above ~24M tokens/month on small models per our published crossover data.