Skip to main content
The detection engine is the same binary in every Deployment mode. There are three: Cloud (we run everything), Hybrid (the engine runs on your servers and reports to our dashboard) and Air-gapped (engine and dashboard both run on your servers). “Self-hosted” below means Hybrid or Air-gapped; where the two differ, the row names the mode. Self-hosting does not give you a cut-down firewall — the /guard engine, policy enforcement, audit trail, and RBAC run on your servers. What differs is the managed surface around it: metering, hosted inference, the control plane, and a small amount of proprietary tooling that is stripped from customer images. This page is the honest line-by-line.

Capability matrix

Legend: ✅ available · ⚙️ available but you operate/supply it · ⚠️ works, with the caveat stated · ❌ not available · — not applicable

Why the cloud-only items are cloud-only

The offensive red-team engine is proprietary attack tooling that is stripped from non-cloud images at build time (DEPLOYMENT_MODE ≠ cloud). It is internal IP we do not ship inside customer environments. The defensive detectors it exercises are fully present in self-host; only the built-in “attack yourself” harness is cloud-only. You can still run your own adversarial spot-checks against your live config by POSTing known-attack payloads to /guard in CI.
On Cloud we operate the classifier and LLM-judge inference for you. Self-hosted, ML_INFERENCE_MODE=local runs the prompt-injection and toxicity classifiers and the multi-turn embeddings inside the engine, from model weights you download yourself, with your own Hugging Face token and under each model’s licence, which you accept on your own account (Third-Party Models); PromptGuard does not ship or redistribute them. It never calls hosted inference. The LLM judges need a judge server inside your network: the Helm chart’s optional judge service (vLLM serving gpt-oss-safeguard-20b, which needs a GPU and weights you download; bring your own GPU model server, not validated by PromptGuard), or your own OpenAI-compatible server at LLM_GUARD_BASE_URL. Every detector that cannot run is listed on the scan result as unavailable, so reduced coverage is never read as a pass. ML_INFERENCE_MODE=off runs the (substantial) rule engine alone.
Cloud enforces per-plan monthly request quotas through Stripe-backed subscriptions. A self-hosted instance has no reason to phone a billing provider, so it runs on a signed, time-limited license (90 days by default) and is metered-but-uncapped: the counter increments for your own visibility, but requests are never hard-capped. No license behaves like the Free tier (20k requests/month); a lapsed/invalid license fails closed on billable requests (503 LICENSE_INACTIVE) while /health and /docs stay up.
Staged rollout, kill-switch, and auto-update for the Shadow AI desktop agents are control-plane features: they need the hosted release/telemetry backend to target cohorts and halt a bad release. An air-gapped fleet distributes agent updates through your own MDM instead.

Two things that behave differently self-hosted

Tenant isolation is enforced by Postgres as well as by the app. deploy/apply-migrations.sh gives a self-hosted Postgres the same schema and the same row-level security policies as Cloud (apart from the pgvector-backed attack_corpus table where pgvector is not installed). The API serves each request as promptguard_app, a database role those policies restrict to the request’s Organization, so one missing filter in the application cannot read another Organization’s rows. Org/project RBAC and IP allowlists are enforced by the application on top. Migrations and scheduled jobs keep a privileged role, so do not expose the database directly to untrusted clients. Run the script on install and on every upgrade; it stops with a non-zero exit at the first migration error, and the API must not be started until it has finished cleanly.
Pin a stable node identity. A license can carry a node fingerprint, which the instance compares with its own. Containerized deployments must set a stable PROMPTGUARD_NODE_ID (e.g. uuidgen once, then reuse) — the default hardware fingerprint changes on each container redeploy and would fail that comparison. PROMPTGUARD_NODE_ID is yours to set, so the fingerprint is a guard against accidental reuse, not a lock: the signature and the expiry date are what enforce the license.

Choosing a model

Hybrid and Air-gapped deployments are an Enterprise-tier capability. See the Compliance page for the air-gap egress audit, or contact sales to scope a deployment.

Next steps

Enterprise Setup

Organizations, SSO, RBAC, and audit logs

Compliance

Air-gap egress audit, build provenance, and certification status

Shadow AI Deployment Modes

Personal, fleet, and air-gapped agent deployment

Reliability

Fail-open vs fail-closed behavior