AI Guardrails Index
AI Guardrails Categories
We broke AI safety down into 6 categories and curated datasets and models that demonstrate the state of AI guardrails using LLMs and other open source models.
Jailbreaking
Jailbreaking LLMs bypasses safety measures to generate harmful content, posing risks across industries. Learn how effective models resist attempts to bypass their safety controls and restrictions.
Best Model
Detect Jailbreak
Performance: 0.81
Top Models Comparison:
0.81
PII Detection
Exposing unredacted PII in AI applications risks compliance violations and privacy breaches. Learn how well models identify and mask PII to ensure compliance and privacy.
Best Model
Guardrails PII
Performance: 0.65
Top Models Comparison:
0.65
0.62
Content Moderation
Unchecked AI outputs can spread harmful content, posing reputational and compliance risks. Learn how well models filter toxic language and prevent the amplification of harmful content.
Best Model
Toxic Language
Performance: 0.72
Top Models Comparison:
0.72
0.60
Topic Restriction
LLMs can generate off-topic or unauthorized content, leading to misuse and compliance concerns. Learn how well identify deviation from topic boundaries and guidelines.
Best Model
Restrict to Topic (Hybrid)
Performance: 0.93
Top Models Comparison:
0.93
0.91
Competitors Check
The inadvertent creation or favoring of competitor mentions can impact brand equity and control. Learn how well models handle discussions of competing AI companies appropriately.
Best Model
Competitor Check
Performance: 0.67
Top Models Comparison:
0.67
0.64
Hallucination
AI hallucinations can result in inaccurate and misleading text that is nonetheless compelling and convincing. Learn how different models tend to generate false or unsupported information.
Best Model
Provenance LLM
Performance: 0.77
Top Models Comparison:
0.77
0.75
Model Leaderboard
A comprehensive visual comparison of how top-performing models stack up across key benchmarks like hallucinations, PII data exposure, and alignment with your AI strategy.
| Model | PII | Jailbreak | Content Moderation | On Topic | Competitor Check | Hallucinations |
|---|---|---|---|---|---|---|
| GLiNER for PII | 0.75 | 0.60 | 0.45 | 0s | 0.3 | F1 Score |
| Latency | 0.008s | |||||
| 0.016s | ||||||
| 0.024s | ||||||
| 0.032s | ||||||
| 0.04s | ||||||
| 0.048s | ||||||
| 0.056s | ||||||
| 0.064s | ||||||
| 0.072s |
Model: GLiNER for PII
Latency (CPU): 0.5031s
Latency (GPU): 0.0460s
Latency σ (CPU): 0.0707s
Latency σ (GPU): 0.0791s
Latency Best: 0.0678s
F1 Score: 0.6520
Model: Guardrails PII
Latency (CPU): 0.5002s
Latency (GPU): 0.0678s
Latency σ (CPU): 0.0791s
Latency σ (GPU): 0.0120s
Latency Best: 0.0678s
F1 Score: 0.6520
Model: Presidio
Latency (CPU): 0.0161s
Latency (GPU): 0.0150s
Latency σ (CPU): 0.0056s
Latency Best: 0.0150s
F1 Score: 0.3698
Deep dive into our findings
Learn more about our dataset curation process, our evaluation methodologies and our findings on the effectiveness of various guardrails.
Guardrails tested
6
Number of Datasets
32