Skip to main content
The one question your security team will ask is “where does our data go?” Shadow AI gives you three answers — the three Deployment modes. On the wire, a desktop agent differs only in the engine URL it is signed in to (engine_url) — but rolling a fleet out against your own engine takes more than that today; see Pointing desktop agents at your engine. “Self-hosted” means Hybrid or Air-gapped; it is not a fourth mode.
Cloud is the fastest way to start. Most security-conscious buyers run hybrid: scanning stays on their infrastructure, but they still get one clean cloud dashboard. Pick air-gapped only if you truly can’t allow outbound traffic.

Hybrid — scan on your servers, review in the cloud

Run the engine on your own infrastructure and let only the results flow to the cloud dashboard. On the engine, set:
Each scanned event is recorded locally first, then reliably forwarded to the cloud (ordered, retried automatically if the link drops — nothing is lost during an outage). Policies you author in the cloud are pulled down automatically. Each engine authenticates with its own token and can only write events for your organization.
By default (FORWARD_MODE=metadata) you keep per-request visibility and billing in the cloud dashboard while no prompt content leaves your network — only the verdict, threat type, counts and request details (such as the model and the end-user ID you pass) do. To see prompt previews in the cloud dashboard, opt in with FORWARD_MODE=content: that forwards a masked preview of each prompt and the detector’s explanation, which can quote the prompt.
Upgrading a Hybrid engine from v3.35.0 or earlier? The default changed from content to metadata — see the upgrade note.

Air-gapped — everything inside your network

Air-gapped is a shipped deployment mode, not a bespoke project. The engine runs with DEPLOYMENT_MODE=airgap (the third of the three engine modes: cloud, data_plane, airgap) entirely inside your network. Event forwarding and the license heartbeat are off in code, and the network layer blocks the rest. What you run:
  • The engine, via our Helm chart — including a default-deny egress NetworkPolicy overlay that blocks outbound traffic at the network layer, so “no data leaves” is enforced by Kubernetes, not by trust. GeoIP, release lookups and hosted ML inference are off in code, but the application can still attempt a few outbound calls when configured for them; the NetworkPolicy is what stops them, which is why your CNI must enforce it.
  • The dashboard — deployed by the same Helm chart (or our Docker Compose bundle), along with the database migrations and scheduled jobs.
  • An offline license — an Ed25519-signed, time-limited license file validated locally; no license server or phone-home required.
  • Detection — the deterministic detectors (regex + heuristic rules) and the prompt-injection classifier run in-process, the classifier on model weights you fetch yourself with your own Hugging Face account (PromptGuard does not ship them; see Third-Party Models). The chart’s optional judge service is bring your own GPU model server, not validated by PromptGuard; point the engine at a model server you run, or accept the reduced coverage — each scan lists the checks that could not run.
  • SSO against your own IdP — the dashboard authenticates directly via OIDC against your identity provider.
Everything you need to review activity — verdicts, threat types and decision counts — lives in your dashboard and its Postgres event log, so you never need to reach the cloud to operate. Masked content previews are the exception: an Air-gapped deployment stores none by default, and an Organization admin can turn them off independently of the operator. Set STORE_PREVIEWS=true if you want excerpts kept in your own database.

Get the deployment package

Contact sales@promptguard.co for the Helm chart, license, and rollout guidance for your environment.
When you do need to move data between an isolated site and another environment, you do it deliberately. The local event log is the system of record: it can be queried directly, and the import endpoint that ingests signed event bundles verifies both signature and tenant before accepting anything.
Tooling that packages the local event log into tamper-proof signed bundles for transfer across an air gap is available on request and on the near-term roadmap. If you need it for an isolated deployment, contact support@promptguard.co — we don’t ship a generic export script today, so don’t script against one.
In air-gapped mode your dashboard is the local instance — not promptguard.co. Combining data across sites is done with signed bundles, not a live link, and a tampered bundle is rejected.

Pointing desktop agents at your engine

For a fleet, write the engine_url key into managed configuration and every device on the machine uses your engine, locked, before it has talked to anyone. For one machine, sign in by URL:
What follows from pointing an agent at your own engine, in Hybrid and Air-gapped alike:
  • Your engine is shown no prompt content by default. A device pointed at an engine PromptGuard doesn’t run defaults to the metadata_only disclosure: no prompt text, attachment text, image or audio reaches it, only {category, count} device findings. Set disclosure to masked if you want your engine to see the (masked) prompt and the full detection that depends on it — see What the engine is shown.
  • You can take update traffic off PromptGuard. updater_enabled off stops every update check; update_url points the updater at your own mirror of our signed releases and is then the only source, with no GitHub fallback. Artifacts are still verified against the key compiled into the app, so mirroring is not re-signing. Block those hosts at egress either way. Details in Updates.
  • PROMPTGUARD_BASE_URL is not the same thing. As an environment variable it redirects only scans, and it never locks; the app’s update checks and enrollment follow the managed or signed-in URL, which without either is PromptGuard’s cloud. Push engine_url through your device management (Enrolling a Hybrid fleet) and nobody has to run anything: the value is locked, and enrollment goes to that engine too.
  • Hybrid enrollment needs a managed enrollment token. “Connect this device” opens a dashboard sign-in page, and a Hybrid engine serves no dashboard, so browser enrollment against it does not work. Enroll Hybrid devices with a token instead — see Enrolling a Hybrid fleet.
  • You can require that nothing goes unchecked. Set on_engine_unreachable to hold and a prompt your engine couldn’t be asked about is not forwarded to the AI tool at all — see If the engine cannot be reached.

Enrolling a Hybrid fleet

A Hybrid engine (DEPLOYMENT_MODE=data_plane) serves the enrollment endpoint, POST /api/v1/enroll, and none of the dashboard. So a Hybrid device enrolls from a token your device management delivers, with no browser step:
  1. Put two keys in the agent’s managed configuration — that section has the key schema and the per-OS locations:
  2. Install the desktop app. On first launch it sees the token, sends it to <engine_url>/enroll, and saves the device credential. It opens no browser and no dashboard page, and nobody has to click anything. Then it turns protection on the same way “Connect this device” does.
What you get is a managed device. Its fleet entry shows an attribution label — the OS account name that ran the app, or the hostname if there isn’t one — and not a verified identity. That’s the same as any device enrolled from an admin-minted token, because nobody signed in to prove who is at the machine. Managed configuration always wins: when an enrollment token is set, the app uses it even against the cloud engine, and it never offers browser sign-in. Keep these in mind:
  • The token has to exist on the engine that redeems it. A Hybrid engine checks tokens against its own database. A token minted in the promptguard.co dashboard lives in PromptGuard’s cloud database, so a Hybrid engine rejects it as Invalid or expired enrollment token (HTTP 403). A Hybrid engine doesn’t expose token minting itself (the fleet routes are part of the dashboard). Contact support@promptguard.co to provision tokens for a Hybrid rollout.
  • The dashboard can write the keys for you. Download MDM profile on the Shadow AI page mints an enrollment token and renders a macOS .mobileconfig with the keys above already filled in, plus on_engine_unreachable set from your organization’s fail-closed setting. For the Windows .reg or the Linux managed.json, call POST /dashboard/fleet/enforcement-artifact with {"platform": "windows"} or {"platform": "linux"}. Earlier releases wrote key names the agent did not read (base_url, enrollment_token) and, on Windows, a PromptGuard\Shadow subkey it never opened; if you pushed one of those, push a freshly downloaded artifact over it. Either way you can still write the keys yourself — a key the agent does not recognize is listed as an unknown key in Settings, under the settings your organization manages.
  • A refused token isn’t retried in a loop. The app tells the person to contact their IT administrator and waits for someone to press Try connecting again.

Which one is right for you

Cloud

Fastest to deploy, full real-time dashboard. Great for getting started.

Hybrid

Scanning on your infra, one cloud dashboard. The common choice for security-sensitive teams.

Air-gapped

Egress denied at the network layer — your own dashboard, signed bundles. For regulated or isolated environments.