Models & Infrastructure

Intelligence you can upgrade. Hardware you own.

ZenithAI separates the experience your teams use from the models that power it. Four stable tiers run on open-weight models hosted on your GPUs, and can be exchanged as better models are released — without disrupting a single user.

Sec. 01 The tiers

The right engine for every task.

Routing between tiers is automatic — a quick lookup never occupies your deepest model, and a complex analysis is never short-changed by a light one. Users can also choose a tier explicitly per conversation.

I

Instant

Immediate responses for lookups, utilities and speech transcription.

Example class · compact models, 2–5B
II

Fast

The everyday workhorse: general questions, summaries and routine drafting.

Example class · efficient models, 4–9B
III

Smart

Analysis, document reasoning with citations, coding, formal drafting.

Example class · capable models, 12–30B
IV

Genius

The deepest reasoning tier for complex, multi-step analytical work.

Example class · large open-weight models, sized to your infrastructure

Models are examples, not commitments. Each deployment selects models from the current generation of leading open-weight families — such as Google’s Gemma line and comparable open models — matched to your hardware, languages and workload. When the state of the art moves, your tiers move with it.

Sec. 02 Model flexibility

Bring the model that fits. Change it when it doesn’t.

ZenithAI is built on open, standard model-serving technology rather than a proprietary inference stack. That single decision buys your organisation lasting freedom.

  • 01No model lock-in — supported open-weight models can be evaluated and adopted per tier
  • 02Bring your own model — organisation-specific or domain-tuned open models hosted in your deployment, subject to compatibility validation
  • 03Residency without thrash — models are pinned to GPUs with health monitoring, so switching tiers doesn’t mean waiting for reloads
  • 04Local embeddings too — document retrieval uses embedding models hosted on your own hardware, not an external embedding API

Sec. 03 Hardware

Sized to your organisation — not the other way around.

ZenithAI runs efficiently on commercially available GPU hardware. These are illustrative starting points; every deployment is sized during scoping against your user count, languages and workload.

Example · Departmental

Compact server

A single server with two workstation-class GPUs (e.g. 2 × 12–24 GB) runs the Instant, Fast and Smart tiers with document intelligence and OCR for a department or pilot group.

Example · Organisation

Dedicated GPU node

A dedicated node with data-centre GPUs (e.g. 2–4 × 48–80 GB) adds the Genius tier, larger context windows and headroom for concurrent departments.

Example · Enterprise

Multi-node cluster

Multiple GPU nodes separate interactive tiers from heavy analytical workloads and OCR pipelines, supporting enterprise-wide rollout with capacity isolation.

Efficiency is an engineering discipline here. Tiering, GPU residency pinning and CPU-hosted embeddings extract full value from every card you buy — the reference deployment serves a working organisation’s daily load, including document intelligence and OCR, from a two-GPU footprint.

The next step

Get a sizing conversation, not a sales pitch.

Tell us your user count, languages and isolation requirements — we’ll propose tiers, models and hardware with a transparent cost picture.