Part of ZenithAI · Organisational knowledge

Turn your website into knowledge your organisation can talk to.

This is not a site search box. Your public pages become part of ZenithAI’s organisational knowledge — the same knowledge your teams already hold conversations with. Paste a URL and it is answered in plain questions, with a citation to the exact page, entirely on your own hardware. It is one of three ways knowledge gets into ZenithAI, alongside the documents you upload and the databases you connect. Nothing about your knowledge ever leaves your infrastructure.

Sec. 01 How it works

Paste a URL. It becomes knowledge. Then you just ask.

Three steps, one screen — and every stage runs inside your own environment. What you get at the end is not a search index to query, but knowledge your people talk to in plain language.

From a public URL to a cited answer, entirely on your infrastructure Step one: paste a public website URL and pick a depth. Step two: ZenithAI fetches, extracts to clean text and indexes each page on your own server, showing live progress. Step three: chat answers from those pages, with a citation linking back to the original page. ON YOUR INFRASTRUCTURE STEP 01 Paste a URL Pick a public page and a crawl depth. https://your-site.example/product STEP 02 Watch it index Fetch · extract to clean text · index — with live progress and honest counts. STEP 03 Ask, get citations Chat answers from those pages, each with a link back to the source. FETCH · EXTRACT · INDEX · CITE — NOTHING LEAVES YOUR WALLS

One screen, three steps. The fetch, the extraction and the retrieval all run on your side of the boundary.

Sec. 02 Choose the depth

One page, or a bounded crawl of the same site. Never off it.

You decide how much of a site becomes knowledge. Every option stays on the same website you chose — the crawler never wanders off-site, and it never follows a redirect that tries to leave.

Depth 01

This page

Imports exactly one page — the URL you pasted, and nothing else. The fastest way to add a single product page, policy or article.

Depth 02

Short crawl

Follows links within the same site to pull in a small cluster of related pages, up to a hard cap of 25 pages.

Depth 03

Standard crawl

A deeper same-site crawl for broader sections — documentation, a product line, a set of policies — still capped at 25 pages by design.

BoundLimitHow it is enforced
Pages per import25A deliberate ceiling — focused, reviewable knowledge, not an entire site scraped blind.
Size per page3 MBEnforced during download, streaming — a hostile or broken page is stopped mid-fetch.
Size per import5 MBA total budget across the whole import, counted live against a visible byte meter.
Typical page~secondsFetched, extracted and indexed end-to-end in seconds, so knowledge is ready almost immediately.

Bounds are a design choice, not a shortfall — they keep an import fast, focused and impossible to weaponise against your own server.

Sec. 03 Honest progress, cited answers

You see exactly what was imported — and what wasn’t.

The import reports discovered, fetched, extracted, indexed and failed as separate counts. An imported site never pretends to more coverage than it has. Illustrative frame below.

ZenithAI · knowledge.your-organisation.internalWorkspace: Product DeskImport: your-site.example
Import status
Discovered · 9
Fetching · 8
Extracted · 7
Usable knowledge · 7
Skipped · 1
Budget
412 KB / 5 MB

Status: fetching · Visited 8 · Fetched 8 · Extracted 7 · Usable knowledge 7 · Downloaded 412 KB / 5 MB

1 page skipped, reason shown:

/app — did not contain enough usable text (JavaScript-only shell). Its links were preserved as text; the page itself was not indexed.


Then, in chat:

What plans does the product offer?

Three plans are listed — Starter, Team and Enterprise — each with its own included seats and support tier.

Cited · Pricing — your-site.example/pricing

Fetched, extracted and answered on your infrastructure · illustrative data

Typical honest failures the import surfaces plainly: a page behind a rate limit, a JavaScript-only shell with no readable text, or a page a site’s own robots rules ask crawlers to leave alone.

Sec. 04 Keeping it fresh

When the page changes, one click updates it.

Each website document has its own details — last crawled, knowledge updated, page size, chunk count and a link to the original. Updates are yours to run, on your schedule, and a refresh keeps the document’s identity so existing citations stay valid.

If a fetch or extraction fails mid-update, the existing knowledge is kept untouched. The system never trades a working page for a broken fetch.

  • 01Check for updates — read-only. Fetches the live page and compares a content fingerprint, telling you “unchanged” or “updated version available”. It changes nothing.
  • 02Update knowledge — appears when the page has changed; re-imports only that one page. The document keeps its identity, so citations that already point to it stay valid.
  • 03Rebuild page — even when nothing changed, force a full re-extract and re-index of a single page.
  • 04Clean deletion — removing a website document takes its chunks, index and evidence with it. No residue is left behind.

Updates are always something you choose — check, update or rebuild — never a background process crawling your sources on its own.

Sec. 05 Security & privacy posture

The part competitors cannot honestly copy.

Every entry below is shipped behaviour, not intent. Importing from the web is exactly where a private platform earns its keep — so this is where the design is strictest.

WK-01100% on-premiseFetching, extraction, chunking, embedding, retrieval and answering all run on your own server. No cloud API, no third-party service, no telemetry.Enforced
WK-02Public web onlyPrivate, loopback, link-local, carrier-NAT and cloud-metadata addresses are refused — checked when you submit and re-checked at every connection, with the vetted address pinned. Only HTTP and HTTPS are accepted; every other scheme is rejected.Enforced
WK-03Same-site onlyThe crawler never follows links off the website you chose. A redirect that tries to leave the site is refused, not followed.Enforced
WK-04Bounded by designHard caps — 25 pages, 3 MB per page, 5 MB per import — enforced while streaming, so a hostile or broken page cannot exhaust your server.Enforced
WK-05Text is data, never instructionsContent from fetched pages is treated as untrusted document text. It cannot trigger tools, run searches or change how the assistant behaves — verified against adversarial injection tests.Enforced
WK-06Respects robots & identifies itselfThe importer honours a site’s robots rules and fetches under a clear, honest User-Agent — a well-behaved guest on the sites you point it at.Enforced
WK-07Links are knowledge, not targetsLinks found on a page, including external ones, are preserved as text so chat can quote them — but the system never follows them.Enforced
WK-08Private to the workspaceWebsite knowledge obeys the same tenancy rules as every other document — scoped to its workspace and owner, isolated from everyone else.Enforced

Extraction and retrieval use local, open-source components running on your hardware — there is no external model or service anywhere in the path. More on the wider posture → Security & governance.

Sec. 06 One knowledge layer

Documents, databases and websites — one knowledge your organisation talks to.

Importing a website is not a separate product. It is one of three ways knowledge enters ZenithAI, all feeding the same conversation, all carrying the same citations, all under the same privacy rules. Your people ask in plain language; they never have to learn where an answer lived.

  • 01Upload documents — PDF, DOCX, XLSX, CSV, TXT and Markdown, with passage-level retrieval and page citations. See document intelligence.
  • 02Connect databases — natural-language questions over your enterprise databases, schema-aware and read-bounded, scoped per workspace.
  • 03Import websites — public pages become citable knowledge documents, indexed into the same pipeline as every uploaded file.

Whichever way knowledge arrives, the assistant’s cite-or-abstain honesty applies to it unchanged — website pages simply become another kind of evidence your answers can point at.

The next step

See your own site answer back.

A demonstration takes under an hour. Bring a public URL of your own and watch it become private, cited knowledge on infrastructure you control.