Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Fast structured generation and serving for LLMs with RadixAttention prefix caching. Use for JSON/regex outputs, constrained decoding, agentic workflows with tool calls, or when you need 5× faster inference than vLLM with prefix sharing. Powers 300,000+ GPUs at xAI, AMD, NVIDIA, and LinkedIn.
.claude/skills/openlair-sglang/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 428% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 193% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 333% | 0% |
| case-16 | ✗→✓ | ▲ Improved | 213% | 0% |
| case-02 | ✓→✓ | = Same ✓ | 102% | 0% |
High-performance serving framework for LLMs and VLMs with RadixAttention for automatic prefix caching.
Use SGLang when:
Use vLLM instead when:
Use TensorRT-LLM instead when:
bash# pip install (recommended) pip install "sglang[all]" # With FlashInfer (faster, CUDA 11.8/12.1) pip install sglang[all] flashinfer -i https://flashinfer.ai/whl/cu121/torch2.4/ # From source git clone https://github.com/sgl-project/sglang.git cd sglang pip install -e "python[all]"
bash# Basic server (Llama 3-8B) python -m sglang.launch_server \ --model-path meta-llama/Meta-Llama-3-8B-Instruct \ --port 30000 # With RadixAttention (automatic prefix caching) python -m sglang.launch_server \ --model-path meta-llama/Meta-Llama-3-8B-Instruct \ --port 30000 \ --enable-radix-cache # Default: enabled # Multi-GPU (tensor parallelism) python -m sglang.launch_server \ --model-path meta-llama/Meta-Llama-3-70B-Instruct \ --tp 4 \ --port 30000
pythonimport sglang as sgl # Set backend sgl.set_default_backend(sgl.OpenAI("http://localhost:30000/v1")) # Simple generation @sgl.function def simple_gen(s, question): s += "Q: " + question + "\n" s += "A:" + sgl.gen("answer", max_tokens=100) # Run state = simple_gen.run(question="What is the capital of France?") print(state["answer"]) # Output: "The capital of France is Paris."
pythonimport sglang as sgl @sgl.function def extract_person(s, text): s += f"Extract person information from: {text}\n" s += "Output JSON:\n" # Constrained JSON generation s += sgl.gen( "json_output", max_tokens=200, regex=r'\{"name": "[^"]+", "age": \d+, "occupation": "[^"]+"\}' ) # Run state = extract_person.run( text="John Smith is a 35-year-old software engineer." ) print(state["json_output"]) # Output: {"name": "John Smith", "age": 35, "occupation": "software engineer"}
What it does: Automatically caches and reuses common prefixes across requests.
Performance:
How it works:
Example (Agent with system prompt):
Request 1: [SYSTEM_PROMPT] + "What's the weather?"
→ Computes full prompt (1000 tokens)
Request 2: [SAME_SYSTEM_PROMPT] + "Book a flight"
→ Reuses system prompt KV cache (998 tokens)
→ Only computes 2 new tokens
→ 5× faster!python@sgl.function def structured_extraction(s, article): s += f"Article: {article}\n\n" s += "Extract key information as JSON:\n" # JSON schema constraint schema = { "type": "object", "properties": { "title": {"type": "string"}, "author": {"type": "string"}, "summary": {"type": "string"}, "sentiment": {"type": "string", "enum": ["positive", "negative", "neutral"]} }, "required": ["title", "author", "summary", "sentiment"] } s += sgl.gen("info", max_tokens=300, json_schema=schema) state = structured_extraction.run(article="...") print(state["info"]) # Output: Valid JSON matching schema
python@sgl.function def extract_email(s, text): s += f"Extract email from: {text}\n" s += "Email: " # Email regex pattern s += sgl.gen( "email", max_tokens=50, regex=r'[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}' ) state = extract_email.run(text="Contact john.doe@example.com for details") print(state["email"]) # Output: "john.doe@example.com"
python@sgl.function def generate_code(s, description): s += f"Generate Python code for: {description}\n" s += "```python\n" # EBNF grammar for Python python_grammar = """ ?start: function_def function_def: "def" NAME "(" [parameters] "):" suite parameters: parameter ("," parameter)* parameter: NAME suite: simple_stmt | NEWLINE INDENT stmt+ DEDENT """ s += sgl.gen("code", max_tokens=200, grammar=python_grammar) s += "\n```"
pythonimport sglang as sgl # Define tools tools = [ { "name": "get_weather", "description": "Get weather for a location", "parameters": { "type": "object", "properties": { "location": {"type": "string"} } } }, { "name": "book_flight", "description": "Book a flight", "parameters": { "type": "object", "properties": { "from": {"type": "string"}, "to": {"type": "string"}, "date": {"type": "string"} } } } ] @sgl.function def agent_workflow(s, user_query, tools): # System prompt (cached with RadixAttention) s += "You are a helpful assistant with access to tools.\n" s += f"Available tools: {tools}\n\n" # User query s += f"User: {user_query}\n" s += "Assistant: " # Generate with function calling s += sgl.gen( "response", max_tokens=200, tools=tools, # SGLang handles tool call format stop=["User:", "\n\n"] ) # Multiple queries reuse system prompt state1 = agent_workflow.run( user_query="What's the weather in NYC?", tools=tools ) # First call: Computes full system prompt state2 = agent_workflow.run( user_query="Book a flight to LA", tools=tools ) # Second call: Reuses system prompt (5× faster)
Few-shot prompting (10 examples in prompt):
Agent workflows (1000-token system prompt):
JSON decoding:
| Workload | vLLM | SGLang | Speedup | |----------|------|--------|---------| | Simple generation | 2500 tok/s | 2800 tok/s | 1.12× | | Few-shot (10 examples) | 500 tok/s | 5000 tok/s | 10× | | Agent (tool calls) | 800 tok/s | 4000 tok/s | 5× | | JSON output | 600 tok/s | 2400 tok/s | 4× |
python@sgl.function def multi_turn_chat(s, history, new_message): # System prompt (always cached) s += "You are a helpful AI assistant.\n\n" # Conversation history (cached as it grows) for msg in history: s += f"{msg['role']}: {msg['content']}\n" # New user message (only new part) s += f"User: {new_message}\n" s += "Assistant: " s += sgl.gen("response", max_tokens=200) # Turn 1 history = [] state = multi_turn_chat.run(history=history, new_message="Hi there!") history.append({"role": "User", "content": "Hi there!"}) history.append({"role": "Assistant", "content": state["response"]}) # Turn 2 (reuses Turn 1 KV cache) state = multi_turn_chat.run(history=history, new_message="What's 2+2?") # Only computes new message (much faster!) # Turn 3 (reuses Turn 1 + Turn 2 KV cache) state = multi_turn_chat.run(history=history, new_message="Tell me a joke") # Progressively faster as history grows
bash# Launch with draft model (2-3× faster) python -m sglang.launch_server \ --model-path meta-llama/Meta-Llama-3-70B-Instruct \ --speculative-model meta-llama/Meta-Llama-3-8B-Instruct \ --speculative-num-steps 5
python@sgl.function def describe_image(s, image_path): s += sgl.image(image_path) s += "Describe this image in detail: " s += sgl.gen("description", max_tokens=200) state = describe_image.run(image_path="photo.jpg") print(state["description"])
python# Automatic batching (continuous batching) states = sgl.run_batch( [ simple_gen.bind(question="What is AI?"), simple_gen.bind(question="What is ML?"), simple_gen.bind(question="What is DL?"), ] ) # All 3 processed in single batch (efficient)
bash# Start server with OpenAI API python -m sglang.launch_server \ --model-path meta-llama/Meta-Llama-3-8B-Instruct \ --port 30000 # Use with OpenAI client curl http://localhost:30000/v1/chat/completions \ -H "Content-Type: application/json" \ -d '{ "model": "default", "messages": [ {"role": "system", "content": "You are helpful"}, {"role": "user", "content": "Hello"} ], "temperature": 0.7, "max_tokens": 100 }' # Works with OpenAI Python SDK from openai import OpenAI client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY") response = client.chat.completions.create( model="default", messages=[{"role": "user", "content": "Hello"}] )
Text models:
Vision models:
100+ models from HuggingFace
NVIDIA: A100, H100, L4, T4 (CUDA 11.8+) AMD: MI300, MI250 (ROCm 6.0+) Intel: Xeon with GPU (coming soon) Apple: M1/M2/M3 via MPS (experimental)
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 4,249 | 4,196 | -1% | 1 | 1 | 0% | 803 | 4,239 | +428% | 0 | 0 | — |
case-02 | pass→pass | 12,827 | 9,162 | -29% | 1 | 1 | 0% | 2,575 | 5,200 | +102% | 0 | 0 | — |
case-03 | pass→pass | 9,452 | 3,854 | -59% | 1 | 1 | 0% | 1,660 | 4,247 | +156% | 0 | 0 | — |
case-04 | pass→pass | 12,165 | 4,945 | -59% | 1 | 1 | 0% | 2,095 | 4,426 | +111% | 0 | 0 | — |
case-05 | pass→pass | 9,746 | 4,735 | -51% | 1 | 1 | 0% | 1,608 | 4,355 | +171% | 0 | 0 | — |
case-06 | pass→pass | 2,677 | 1,517 | -43% | 1 | 1 | 0% | 482 | 3,803 | +689% | 0 | 0 | — |
case-07 | fail→pass | 6,977 | 2,132 | -69% | 1 | 1 | 0% | 1,335 | 3,915 | +193% | 0 | 0 | — |
case-08 | pass→pass | 5,282 | 1,713 | -68% | 1 | 1 | 0% | 1,168 | 3,825 | +227% | 0 | 0 | — |
case-09 | pass→pass | 3,360 | 1,790 | -47% | 1 | 1 | 0% | 595 | 3,865 | +550% | 0 | 0 | — |
case-10 | fail→pass | 4,522 | 2,987 | -34% | 1 | 1 | 0% | 919 | 3,975 | +333% | 0 | 0 | — |
case-11 | pass→pass | 4,236 | 3,102 | -27% | 1 | 1 | 0% | 788 | 4,125 | +423% | 0 | 0 | — |
case-12 | pass→pass | 11,495 | 6,082 | -47% | 1 | 1 | 0% | 2,403 | 5,126 | +113% | 0 | 0 | — |
case-13 | fail→fail | 9,951 | 6,559 | -34% | 1 | 1 | 0% | 2,140 | 4,907 | +129% | 0 | 0 | — |
case-14 | pass→pass | 5,647 | 2,473 | -56% | 1 | 1 | 0% | 1,148 | 4,014 | +250% | 0 | 0 | — |
case-15 | pass→pass | 4,858 | 4,936 | +2% | 1 | 1 | 0% | 1,029 | 4,008 | +290% | 0 | 0 | — |
case-16 | fail→pass | 6,436 | 2,016 | -69% | 1 | 1 | 0% | 1,246 | 3,900 | +213% | 0 | 0 | — |
case-17 | pass→pass | 2,889 | 2,078 | -28% | 1 | 1 | 0% | 648 | 3,922 | +505% | 0 | 0 | — |
case-18 | pass→pass | 5,032 | 1,740 | -65% | 1 | 1 | 0% | 977 | 3,831 | +292% | 0 | 0 | — |
case-19 | pass→pass | 11,148 | 2,386 | -79% | 1 | 1 | 0% | 2,171 | 3,995 | +84% | 0 | 0 | — |
case-20 | pass→pass | 11,166 | 3,322 | -70% | 1 | 1 | 0% | 1,998 | 4,116 | +106% | 0 | 0 | — |
case-21 | pass→pass | 12,846 | 12,605 | -2% | 1 | 1 | 0% | 2,509 | 6,089 | +143% | 0 | 0 | — |
case-22 | pass→pass | 9,624 | 2,315 | -76% | 1 | 1 | 0% | 1,882 | 3,875 | +106% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +18 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.