Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Open-source AI observability platform for LLM tracing, evaluation, and monitoring. Use when debugging LLM applications with detailed traces, running evaluations on datasets, or monitoring production AI systems with real-time insights.
.claude/skills/openlair-phoenix-observability/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 48% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 100% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 91% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 64% | 0% |
| case-13 | ✗→✓ | ▲ Improved | 350% | 0% |
Open-source AI observability and evaluation platform for LLM applications with tracing, evaluation, datasets, experiments, and real-time monitoring.
Use Phoenix when:
Key features:
Use alternatives instead:
bashpip install arize-phoenix # With specific backends pip install arize-phoenix[embeddings] # Embedding analysis pip install arize-phoenix-otel # OpenTelemetry config pip install arize-phoenix-evals # Evaluation framework pip install arize-phoenix-client # Lightweight REST client
pythonimport phoenix as px # Launch in notebook (ThreadServer mode) session = px.launch_app() # View UI session.view() # Embedded iframe print(session.url) # http://localhost:6006
bash# Start Phoenix server phoenix serve # With PostgreSQL export PHOENIX_SQL_DATABASE_URL="postgresql://user:pass@host/db" phoenix serve --port 6006
pythonfrom phoenix.otel import register from openinference.instrumentation.openai import OpenAIInstrumentor # Configure OpenTelemetry with Phoenix tracer_provider = register( project_name="my-llm-app", endpoint="http://localhost:6006/v1/traces" ) # Instrument OpenAI SDK OpenAIInstrumentor().instrument(tracer_provider=tracer_provider) # All OpenAI calls are now traced from openai import OpenAI client = OpenAI() response = client.chat.completions.create( model="gpt-4o", messages=[{"role": "user", "content": "Hello!"}] )
A trace represents a complete execution flow, while spans are individual operations within that trace.
pythonfrom phoenix.otel import register from opentelemetry import trace # Setup tracing tracer_provider = register(project_name="my-app") tracer = trace.get_tracer(__name__) # Create custom spans with tracer.start_as_current_span("process_query") as span: span.set_attribute("input.value", query) # Child spans are automatically nested with tracer.start_as_current_span("retrieve_context"): context = retriever.search(query) with tracer.start_as_current_span("generate_response"): response = llm.generate(query, context) span.set_attribute("output.value", response)
Projects organize related traces:
pythonimport os os.environ["PHOENIX_PROJECT_NAME"] = "production-chatbot" # Or per-trace from phoenix.otel import register tracer_provider = register(project_name="experiment-v2")
pythonfrom phoenix.otel import register from openinference.instrumentation.openai import OpenAIInstrumentor tracer_provider = register() OpenAIInstrumentor().instrument(tracer_provider=tracer_provider)
pythonfrom phoenix.otel import register from openinference.instrumentation.langchain import LangChainInstrumentor tracer_provider = register() LangChainInstrumentor().instrument(tracer_provider=tracer_provider) # All LangChain operations traced from langchain_openai import ChatOpenAI llm = ChatOpenAI(model="gpt-4o") response = llm.invoke("Hello!")
pythonfrom phoenix.otel import register from openinference.instrumentation.llama_index import LlamaIndexInstrumentor tracer_provider = register() LlamaIndexInstrumentor().instrument(tracer_provider=tracer_provider)
pythonfrom phoenix.otel import register from openinference.instrumentation.anthropic import AnthropicInstrumentor tracer_provider = register() AnthropicInstrumentor().instrument(tracer_provider=tracer_provider)
pythonfrom phoenix.evals import ( OpenAIModel, HallucinationEvaluator, RelevanceEvaluator, ToxicityEvaluator, llm_classify ) # Setup model for evaluation eval_model = OpenAIModel(model="gpt-4o") # Evaluate hallucination hallucination_eval = HallucinationEvaluator(eval_model) results = hallucination_eval.evaluate( input="What is the capital of France?", output="The capital of France is Paris.", reference="Paris is the capital of France." )
pythonfrom phoenix.evals import llm_classify # Define custom evaluation def evaluate_helpfulness(input_text, output_text): template = """ Evaluate if the response is helpful for the given question. Question: {input} Response: {output} Is this response helpful? Answer 'helpful' or 'not_helpful'. """ result = llm_classify( model=eval_model, template=template, input=input_text, output=output_text, rails=["helpful", "not_helpful"] ) return result
pythonfrom phoenix import Client from phoenix.evals import run_evals client = Client() # Get spans to evaluate spans_df = client.get_spans_dataframe( project_name="my-app", filter_condition="span_kind == 'LLM'" ) # Run evaluations eval_results = run_evals( dataframe=spans_df, evaluators=[ HallucinationEvaluator(eval_model), RelevanceEvaluator(eval_model) ], provide_explanation=True ) # Log results back to Phoenix client.log_evaluations(eval_results)
pythonfrom phoenix import Client client = Client() # Create dataset dataset = client.create_dataset( name="qa-test-set", description="QA evaluation dataset" ) # Add examples client.add_examples_to_dataset( dataset_name="qa-test-set", examples=[ { "input": {"question": "What is Python?"}, "output": {"answer": "A programming language"} }, { "input": {"question": "What is ML?"}, "output": {"answer": "Machine learning"} } ] )
pythonfrom phoenix import Client from phoenix.experiments import run_experiment client = Client() def my_model(input_data): """Your model function.""" question = input_data["question"] return {"answer": generate_answer(question)} def accuracy_evaluator(input_data, output, expected): """Custom evaluator.""" return { "score": 1.0 if expected["answer"].lower() in output["answer"].lower() else 0.0, "label": "correct" if expected["answer"].lower() in output["answer"].lower() else "incorrect" } # Run experiment results = run_experiment( dataset_name="qa-test-set", task=my_model, evaluators=[accuracy_evaluator], experiment_name="baseline-v1" ) print(f"Average accuracy: {results.aggregate_metrics['accuracy']}")
pythonfrom phoenix import Client client = Client(endpoint="http://localhost:6006") # Get spans as DataFrame spans_df = client.get_spans_dataframe( project_name="my-app", filter_condition="span_kind == 'LLM'", limit=1000 ) # Get specific span span = client.get_span(span_id="abc123") # Get trace trace = client.get_trace(trace_id="xyz789")
pythonfrom phoenix import Client client = Client() # Log user feedback client.log_annotation( span_id="abc123", name="user_rating", annotator_kind="HUMAN", score=0.8, label="helpful", metadata={"comment": "Good response"} )
python# Export to pandas df = client.get_spans_dataframe(project_name="my-app") # Export traces traces = client.list_traces(project_name="my-app")
bashdocker run -p 6006:6006 arizephoenix/phoenix:latest
bash# Set database URL export PHOENIX_SQL_DATABASE_URL="postgresql://user:pass@host:5432/phoenix" # Start server phoenix serve --host 0.0.0.0 --port 6006
| Variable | Description | Default | |----------|-------------|---------| | PHOENIX_PORT | HTTP server port | 6006 | | PHOENIX_HOST | Server bind address | 127.0.0.1 | | PHOENIX_GRPC_PORT | gRPC/OTLP port | 4317 | | PHOENIX_SQL_DATABASE_URL | Database connection | SQLite temp | | PHOENIX_WORKING_DIR | Data storage directory | OS temp | | PHOENIX_ENABLE_AUTH | Enable authentication | false | | PHOENIX_SECRET | JWT signing secret | Required if auth enabled |
bashexport PHOENIX_ENABLE_AUTH=true export PHOENIX_SECRET="your-secret-key-min-32-chars" export PHOENIX_ADMIN_SECRET="admin-bootstrap-token" phoenix serve
Traces not appearing:
pythonfrom phoenix.otel import register # Verify endpoint tracer_provider = register( project_name="my-app", endpoint="http://localhost:6006/v1/traces" # Correct endpoint ) # Force flush from opentelemetry import trace trace.get_tracer_provider().force_flush()
High memory in notebook:
python# Close session when done session = px.launch_app() # ... do work ... session.close() px.close_app()
Database connection issues:
bash# Verify PostgreSQL connection psql $PHOENIX_SQL_DATABASE_URL -c "SELECT 1" # Check Phoenix logs phoenix serve --log-level debug
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-22 | pass→pass | 7,009 | 1,994 | -72% | 1 | 1 | 0% | 1,250 | 3,294 | +164% | 0 | 0 | — |
case-01 | fail→pass | 17,379 | 10,074 | -42% | 1 | 1 | 0% | 3,471 | 5,134 | +48% | 0 | 0 | — |
case-02 | fail→fail | 13,446 | 11,264 | -16% | 1 | 1 | 0% | 2,705 | 5,164 | +91% | 0 | 0 | — |
case-03 | pass→pass | 14,690 | 11,084 | -25% | 1 | 1 | 0% | 2,941 | 5,304 | +80% | 0 | 0 | — |
case-04 | pass→pass | 9,877 | 5,821 | -41% | 1 | 1 | 0% | 1,822 | 4,010 | +120% | 0 | 0 | — |
case-05 | pass→pass | 11,065 | 6,629 | -40% | 1 | 1 | 0% | 1,926 | 4,309 | +124% | 0 | 0 | — |
case-06 | fail→pass | 8,923 | 2,939 | -67% | 1 | 1 | 0% | 1,717 | 3,437 | +100% | 0 | 0 | — |
case-07 | fail→pass | 10,956 | 3,653 | -67% | 1 | 1 | 0% | 1,929 | 3,676 | +91% | 0 | 0 | — |
case-08 | fail→pass | 12,078 | 5,119 | -58% | 1 | 1 | 0% | 2,464 | 4,041 | +64% | 0 | 0 | — |
case-09 | pass→pass | 4,629 | 2,209 | -52% | 1 | 1 | 0% | 939 | 3,471 | +270% | 0 | 0 | — |
case-10 | pass→pass | 3,117 | 2,538 | -19% | 1 | 1 | 0% | 602 | 3,513 | +484% | 0 | 0 | — |
case-11 | pass→pass | 4,332 | 2,916 | -33% | 1 | 1 | 0% | 810 | 3,565 | +340% | 0 | 0 | — |
case-12 | pass→pass | 8,814 | 7,192 | -18% | 1 | 1 | 0% | 1,718 | 4,577 | +166% | 0 | 0 | — |
case-13 | fail→pass | 4,251 | 2,071 | -51% | 1 | 1 | 0% | 751 | 3,377 | +350% | 0 | 0 | — |
case-14 | pass→pass | 17,622 | 3,048 | -83% | 1 | 1 | 0% | 3,170 | 3,640 | +15% | 0 | 0 | — |
case-15 | fail→pass | 10,201 | 4,367 | -57% | 1 | 1 | 0% | 2,100 | 3,855 | +84% | 0 | 0 | — |
case-16 | fail→pass | 7,332 | 2,231 | -70% | 1 | 1 | 0% | 1,396 | 3,397 | +143% | 0 | 0 | — |
case-17 | fail→pass | 3,970 | 2,008 | -49% | 1 | 1 | 0% | 733 | 3,372 | +360% | 0 | 0 | — |
case-18 | fail→pass | 11,697 | 3,853 | -67% | 1 | 1 | 0% | 2,102 | 3,780 | +80% | 0 | 0 | — |
case-19 | pass→pass | 4,724 | 3,012 | -36% | 1 | 1 | 0% | 869 | 3,572 | +311% | 0 | 0 | — |
case-20 | fail→pass | 6,449 | 2,868 | -56% | 1 | 1 | 0% | 1,306 | 3,622 | +177% | 0 | 0 | — |
case-21 | pass→pass | 8,036 | 3,854 | -52% | 1 | 1 | 0% | 1,326 | 3,640 | +175% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +45 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.