The detection engine is the same binary in every Deployment mode. There are three: Cloud (we run everything), Hybrid (the engine runs on your servers and reports to our dashboard) and Air-gapped (engine and dashboard both run on your servers). “Self-hosted” below means Hybrid or Air-gapped; where the two differ, the row names the mode. Self-hosting does not give you a cut-down firewall — the
/guard engine, policy enforcement, audit trail, and RBAC run on your servers. What differs is the managed surface around it: metering, hosted inference, the control plane, and a small amount of proprietary tooling that is stripped from customer images. This page is the honest line-by-line.Capability matrix
Legend: ✅ available · ⚙️ available but you operate/supply it · ⚠️ works, with the caveat stated · ❌ not available · — not applicable
Why the cloud-only items are cloud-only
Red-team attack engine & self-test UI
Red-team attack engine & self-test UI
The offensive red-team engine is proprietary attack tooling that is stripped from non-cloud images at build time (
DEPLOYMENT_MODE ≠ cloud). It is internal IP we do not ship inside customer environments. The defensive detectors it exercises are fully present in self-host; only the built-in “attack yourself” harness is cloud-only. You can still run your own adversarial spot-checks against your live config by POSTing known-attack payloads to /guard in CI.Managed ML and LLM-judge inference
Managed ML and LLM-judge inference
On Cloud we operate the classifier and LLM-judge inference for you. Self-hosted,
ML_INFERENCE_MODE=local runs the prompt-injection and toxicity classifiers and the multi-turn embeddings inside the engine, from model weights you download yourself, with your own Hugging Face token and under each model’s licence, which you accept on your own account (Third-Party Models); PromptGuard does not ship or redistribute them. It never calls hosted inference. The LLM judges need a judge server inside your network: the Helm chart’s optional judge service (vLLM serving gpt-oss-safeguard-20b, which needs a GPU and weights you download; bring your own GPU model server, not validated by PromptGuard), or your own OpenAI-compatible server at LLM_GUARD_BASE_URL. Every detector that cannot run is listed on the scan result as unavailable, so reduced coverage is never read as a pass. ML_INFERENCE_MODE=off runs the (substantial) rule engine alone.Stripe billing & quota metering
Stripe billing & quota metering
Cloud enforces per-plan monthly request quotas through Stripe-backed subscriptions. A self-hosted instance has no reason to phone a billing provider, so it runs on a signed, time-limited license (90 days by default) and is metered-but-uncapped: the counter increments for your own visibility, but requests are never hard-capped. No license behaves like the Free tier (20k requests/month); a lapsed/invalid license fails closed on billable requests (
503 LICENSE_INACTIVE) while /health and /docs stay up.Shadow AI agent auto-update & staged rollout
Shadow AI agent auto-update & staged rollout
Staged rollout, kill-switch, and auto-update for the Shadow AI desktop agents are control-plane features: they need the hosted release/telemetry backend to target cohorts and halt a bad release. An air-gapped fleet distributes agent updates through your own MDM instead.
Two things that behave differently self-hosted
Choosing a model
Hybrid and Air-gapped deployments are an Enterprise-tier capability. See the Compliance page for the air-gap egress audit, or contact sales to scope a deployment.
Next steps
Enterprise Setup
Organizations, SSO, RBAC, and audit logs
Compliance
Air-gap egress audit, build provenance, and certification status
Shadow AI Deployment Modes
Personal, fleet, and air-gapped agent deployment
Reliability
Fail-open vs fail-closed behavior