Built with Llama. PromptGuard’s prompt-injection detection uses Meta’s
Llama Prompt Guard 2.
How we use these models
The engine sends text to each model through an inference endpoint; no PromptGuard release ships model weights today. If a release starts to, this page will say how each licence’s text travels with the weights before that release ships.Detection ensemble
These run on scans as part of the toxicity and prompt-injection ensemble.Other models
Licence terms
Llama 4 Community License
Llama Prompt Guard 2 is made available under the Llama 4 Community License. Llama 4 is licensed under the Llama 4 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved.- The licence: https://www.llama.com/llama4/license/
- Your use of PromptGuard must also comply with Meta’s Llama 4 Acceptable Use Policy, which the licence incorporates and which we pass through to you: https://www.llama.com/llama4/use-policy/
OpenRAIL
KoalaAI/Text-Moderation is released under an Open Responsible AI License
(OpenRAIL); its model card declares the licence family without naming a specific
OpenRAIL variant. OpenRAIL licences permit commercial use but attach
use-based restrictions that bind everyone downstream of the model, including
you. You must not use PromptGuard, or the moderation results it produces, for
any use those restrictions prohibit. The OpenRAIL licences and their
restrictions are published by the RAIL initiative:
https://www.licenses.ai/ai-licenses
Creative Commons Attribution 4.0
cardiffnlp/twitter-roberta-base-hate-latest is by Cardiff NLP (Cardiff
University), published at
https://huggingface.co/cardiffnlp/twitter-roberta-base-hate-latest
under the Creative Commons Attribution 4.0 International licence:
https://creativecommons.org/licenses/by/4.0/.
PromptGuard uses the model unmodified.
Apache License 2.0
The models listed under the Apache License 2.0 are used under its terms: https://www.apache.org/licenses/LICENSE-2.0. None of their model repositories publishes a NOTICE file.Models we stopped using
facebook/roberta-hate-speech-dynabench-r4-target was removed from the
detection ensemble in September 2026 because it is published without a licence.
Its place is held by cardiffnlp/twitter-roberta-base-hate-latest.