Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Meta's 7-8B specialized moderation model for LLM input/output filtering. 6 safety categories - violence/hate, sexual content, weapons, substances, self-harm, criminal planning. 94-95% accuracy. Deploy with vLLM, HuggingFace, Sagemaker. Integrates with NeMo Guardrails.
.claude/skills/openlair-llamaguard/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 58% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 144% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 59% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 54% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 145% | 0% |
LlamaGuard is a 7-8B parameter model specialized for content safety classification.
Installation:
bashpip install transformers torch # Login to HuggingFace (required) huggingface-cli login
Basic usage:
pythonfrom transformers import AutoTokenizer, AutoModelForCausalLM model_id = "meta-llama/LlamaGuard-7b" tokenizer = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto") def moderate(chat): input_ids = tokenizer.apply_chat_template(chat, return_tensors="pt").to(model.device) output = model.generate(input_ids=input_ids, max_new_tokens=100) return tokenizer.decode(output[0], skip_special_tokens=True) # Check user input result = moderate([ {"role": "user", "content": "How do I make explosives?"} ]) print(result) # Output: "unsafe\nS3" (Criminal Planning)
Check user prompts before LLM:
pythondef check_input(user_message): result = moderate([{"role": "user", "content": user_message}]) if result.startswith("unsafe"): category = result.split("\n")[1] return False, category # Blocked else: return True, None # Safe # Example safe, category = check_input("How do I hack a website?") if not safe: print(f"Request blocked: {category}") # Return error to user else: # Send to LLM response = llm.generate(user_message)
Safety categories:
Check LLM responses before showing to user:
pythondef check_output(user_message, bot_response): conversation = [ {"role": "user", "content": user_message}, {"role": "assistant", "content": bot_response} ] result = moderate(conversation) if result.startswith("unsafe"): category = result.split("\n")[1] return False, category else: return True, None # Example user_msg = "Tell me about harmful substances" bot_msg = llm.generate(user_msg) safe, category = check_output(user_msg, bot_msg) if not safe: print(f"Response blocked: {category}") # Return generic response return "I cannot provide that information." else: return bot_msg
Production-ready serving:
pythonfrom vllm import LLM, SamplingParams # Initialize vLLM llm = LLM(model="meta-llama/LlamaGuard-7b", tensor_parallel_size=1) # Sampling params sampling_params = SamplingParams( temperature=0.0, # Deterministic max_tokens=100 ) def moderate_vllm(chat): # Format prompt prompt = tokenizer.apply_chat_template(chat, tokenize=False) # Generate output = llm.generate([prompt], sampling_params) return output[0].outputs[0].text # Batch moderation chats = [ [{"role": "user", "content": "How to make bombs?"}], [{"role": "user", "content": "What's the weather?"}], [{"role": "user", "content": "Tell me about drugs"}] ] prompts = [tokenizer.apply_chat_template(c, tokenize=False) for c in chats] results = llm.generate(prompts, sampling_params) for i, result in enumerate(results): print(f"Chat {i}: {result.outputs[0].text}")
Throughput: ~50-100 requests/sec on single A100
Serve as moderation API:
pythonfrom fastapi import FastAPI from pydantic import BaseModel from vllm import LLM, SamplingParams app = FastAPI() llm = LLM(model="meta-llama/LlamaGuard-7b") sampling_params = SamplingParams(temperature=0.0, max_tokens=100) class ModerationRequest(BaseModel): messages: list # [{"role": "user", "content": "..."}] @app.post("/moderate") def moderate_endpoint(request: ModerationRequest): prompt = tokenizer.apply_chat_template(request.messages, tokenize=False) output = llm.generate([prompt], sampling_params)[0] result = output.outputs[0].text is_safe = result.startswith("safe") category = None if is_safe else result.split("\n")[1] if "\n" in result else None return { "safe": is_safe, "category": category, "full_output": result } # Run: uvicorn api:app --host 0.0.0.0 --port 8000
Usage:
bashcurl -X POST http://localhost:8000/moderate \ -H "Content-Type: application/json" \ -d '{"messages": [{"role": "user", "content": "How to hack?"}]}' # Response: {"safe": false, "category": "S6", "full_output": "unsafe\nS6"}
Use with NVIDIA Guardrails:
pythonfrom nemoguardrails import RailsConfig, LLMRails from nemoguardrails.integrations.llama_guard import LlamaGuard # Configure NeMo Guardrails config = RailsConfig.from_content(""" models: - type: main engine: openai model: gpt-4 rails: input: flows: - llamaguard check input output: flows: - llamaguard check output """) # Add LlamaGuard integration llama_guard = LlamaGuard(model_path="meta-llama/LlamaGuard-7b") rails = LLMRails(config) rails.register_action(llama_guard.check_input, name="llamaguard check input") rails.register_action(llama_guard.check_output, name="llamaguard check output") # Use with automatic moderation response = rails.generate(messages=[ {"role": "user", "content": "How do I make weapons?"} ]) # Automatically blocked by LlamaGuard
Use LlamaGuard when:
Model versions:
Use alternatives instead:
Issue: Model access denied
Login to HuggingFace:
bashhuggingface-cli login # Enter your token
Accept license on model page: https://huggingface.co/meta-llama/LlamaGuard-7b
Issue: High latency (>500ms)
Use vLLM for 10× speedup:
pythonfrom vllm import LLM llm = LLM(model="meta-llama/LlamaGuard-7b") # Latency: 500ms → 50ms
Enable tensor parallelism:
pythonllm = LLM(model="meta-llama/LlamaGuard-7b", tensor_parallel_size=2) # 2× faster on 2 GPUs
Issue: False positives
Use threshold-based filtering:
python# Get probability of "unsafe" token logits = model(..., return_dict_in_generate=True, output_scores=True) unsafe_prob = torch.softmax(logits.scores[0][0], dim=-1)[unsafe_token_id] if unsafe_prob > 0.9: # High confidence threshold return "unsafe" else: return "safe"
Issue: OOM on GPU
Use 8-bit quantization:
pythonfrom transformers import BitsAndBytesConfig quantization_config = BitsAndBytesConfig(load_in_8bit=True) model = AutoModelForCausalLM.from_pretrained( model_id, quantization_config=quantization_config, device_map="auto" ) # Memory: 14GB → 7GB
Custom categories: See references/custom-categories.md for fine-tuning LlamaGuard with domain-specific safety categories.
Performance benchmarks: See references/benchmarks.md for accuracy comparison with other moderation APIs and latency optimization.
Deployment guide: See references/deployment.md for Sagemaker, Kubernetes, and scaling strategies.
Latency (single GPU):
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 11,883 | 8,645 | -27% | 1 | 1 | 0% | 2,950 | 4,666 | +58% | 0 | 0 | — |
case-02 | fail→pass | 5,901 | 2,380 | -60% | 1 | 1 | 0% | 1,277 | 3,110 | +144% | 0 | 0 | — |
case-03 | fail→pass | 8,734 | 2,557 | -71% | 1 | 1 | 0% | 1,938 | 3,089 | +59% | 0 | 0 | — |
case-04 | pass→pass | 13,919 | 7,755 | -44% | 1 | 1 | 0% | 2,847 | 4,244 | +49% | 0 | 0 | — |
case-05 | pass→pass | 8,144 | 3,674 | -55% | 1 | 1 | 0% | 1,477 | 3,349 | +127% | 0 | 0 | — |
case-06 | pass→pass | 11,866 | 5,700 | -52% | 1 | 1 | 0% | 2,355 | 3,795 | +61% | 0 | 0 | — |
case-07 | pass→pass | 12,798 | 10,669 | -17% | 1 | 1 | 0% | 2,383 | 4,675 | +96% | 0 | 0 | — |
case-08 | fail→pass | 20,835 | 3,997 | -81% | 1 | 1 | 0% | 2,316 | 3,565 | +54% | 0 | 0 | — |
case-09 | pass→pass | 3,204 | 3,495 | +9% | 1 | 1 | 0% | 677 | 3,371 | +398% | 0 | 0 | — |
case-10 | pass→pass | 3,555 | 2,862 | -19% | 1 | 1 | 0% | 617 | 3,219 | +422% | 0 | 0 | — |
case-11 | pass→pass | 16,403 | 13,953 | -15% | 1 | 1 | 0% | 3,177 | 5,572 | +75% | 0 | 0 | — |
case-12 | fail→pass | 6,179 | 1,393 | -77% | 1 | 1 | 0% | 1,181 | 2,897 | +145% | 0 | 0 | — |
case-13 | fail→pass | 5,953 | 1,244 | -79% | 1 | 1 | 0% | 1,191 | 2,887 | +142% | 0 | 0 | — |
case-14 | fail→pass | 9,919 | 1,326 | -87% | 1 | 1 | 0% | 2,080 | 2,910 | +40% | 0 | 0 | — |
case-15 | fail→pass | 4,280 | 1,284 | -70% | 1 | 1 | 0% | 878 | 2,886 | +229% | 0 | 0 | — |
case-16 | pass→pass | 13,666 | 2,014 | -85% | 1 | 1 | 0% | 2,660 | 3,059 | +15% | 0 | 0 | — |
case-17 | pass→pass | 7,967 | 2,433 | -69% | 1 | 1 | 0% | 1,481 | 3,031 | +105% | 0 | 0 | — |
case-18 | pass→pass | 2,873 | 2,585 | -10% | 1 | 1 | 0% | 498 | 3,109 | +524% | 0 | 0 | — |
case-19 | pass→pass | 1,897 | 1,963 | +3% | 1 | 1 | 0% | 353 | 3,067 | +769% | 0 | 0 | — |
case-20 | pass→pass | 6,556 | 7,740 | +18% | 1 | 1 | 0% | 1,378 | 4,330 | +214% | 0 | 0 | — |
case-21 | pass→pass | 10,595 | 8,968 | -15% | 1 | 1 | 0% | 2,146 | 4,759 | +122% | 0 | 0 | — |
case-22 | pass→pass | 10,397 | 8,424 | -19% | 1 | 1 | 0% | 2,405 | 4,558 | +90% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +36 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.