Models & Infrastructure
Intelligence you can upgrade. Hardware you own.
ZenithAI separates the experience your teams use from the models that power it. Four stable tiers run on open-weight models hosted on your GPUs, and can be exchanged as better models are released — without disrupting a single user.
Sec. 01 The tiers
The right engine for every task.
Routing between tiers is automatic — a quick lookup never occupies your deepest model, and a complex analysis is never short-changed by a light one. Users can also choose a tier explicitly per conversation.
Instant
Immediate responses for lookups, utilities and speech transcription.
Example class · compact models, 2–5BFast
The everyday workhorse: general questions, summaries and routine drafting.
Example class · efficient models, 4–9BSmart
Analysis, document reasoning with citations, coding, formal drafting.
Example class · capable models, 12–30BGenius
The deepest reasoning tier for complex, multi-step analytical work.
Example class · large open-weight models, sized to your infrastructureModels are examples, not commitments. Each deployment selects models from the current generation of leading open-weight families — such as Google’s Gemma line and comparable open models — matched to your hardware, languages and workload. When the state of the art moves, your tiers move with it.
Sec. 02 Model flexibility
Bring the model that fits. Change it when it doesn’t.
ZenithAI is built on open, standard model-serving technology rather than a proprietary inference stack. That single decision buys your organisation lasting freedom.
- 01No model lock-in — supported open-weight models can be evaluated and adopted per tier
- 02Bring your own model — organisation-specific or domain-tuned open models hosted in your deployment, subject to compatibility validation
- 03Residency without thrash — models are pinned to GPUs with health monitoring, so switching tiers doesn’t mean waiting for reloads
- 04Local embeddings too — document retrieval uses embedding models hosted on your own hardware, not an external embedding API
Sec. 03 Hardware
Sized to your organisation — not the other way around.
ZenithAI runs efficiently on commercially available GPU hardware. These are illustrative starting points; every deployment is sized during scoping against your user count, languages and workload.
Compact server
A single server with two workstation-class GPUs (e.g. 2 × 12–24 GB) runs the Instant, Fast and Smart tiers with document intelligence and OCR for a department or pilot group.
Dedicated GPU node
A dedicated node with data-centre GPUs (e.g. 2–4 × 48–80 GB) adds the Genius tier, larger context windows and headroom for concurrent departments.
Multi-node cluster
Multiple GPU nodes separate interactive tiers from heavy analytical workloads and OCR pipelines, supporting enterprise-wide rollout with capacity isolation.
Efficiency is an engineering discipline here. Tiering, GPU residency pinning and CPU-hosted embeddings extract full value from every card you buy — the reference deployment serves a working organisation’s daily load, including document intelligence and OCR, from a two-GPU footprint.
The next step
Get a sizing conversation, not a sales pitch.
Tell us your user count, languages and isolation requirements — we’ll propose tiers, models and hardware with a transparent cost picture.