Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Data framework for building LLM applications with RAG. Specializes in document ingestion (300+ connectors), indexing, and querying. Features vector indices, query engines, agents, and multi-modal support. Use for document Q&A, chatbots, knowledge retrieval, or building RAG pipelines. Best for data-centric LLM applications.
.claude/skills/openlair-llamaindex/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 102% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 194% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 92% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 194% | 0% |
| case-05 | ✓→✓ | = Same ✓ | 185% | 0% |
The leading framework for connecting LLMs with your data.
Use LlamaIndex when:
Metrics:
Use alternatives instead:
bash# Starter package (recommended) pip install llama-index # Or minimal core + specific integrations pip install llama-index-core pip install llama-index-llms-openai pip install llama-index-embeddings-openai
pythonfrom llama_index.core import VectorStoreIndex, SimpleDirectoryReader # Load documents documents = SimpleDirectoryReader("data").load_data() # Create index index = VectorStoreIndex.from_documents(documents) # Query query_engine = index.as_query_engine() response = query_engine.query("What did the author do growing up?") print(response)
pythonfrom llama_index.core import SimpleDirectoryReader, Document from llama_index.readers.web import SimpleWebPageReader from llama_index.readers.github import GithubRepositoryReader # Directory of files documents = SimpleDirectoryReader("./data").load_data() # Web pages reader = SimpleWebPageReader() documents = reader.load_data(["https://example.com"]) # GitHub repository reader = GithubRepositoryReader(owner="user", repo="repo") documents = reader.load_data(branch="main") # Manual document creation doc = Document( text="This is the document content", metadata={"source": "manual", "date": "2025-01-01"} )
pythonfrom llama_index.core import VectorStoreIndex, ListIndex, TreeIndex # Vector index (most common - semantic search) vector_index = VectorStoreIndex.from_documents(documents) # List index (sequential scan) list_index = ListIndex.from_documents(documents) # Tree index (hierarchical summary) tree_index = TreeIndex.from_documents(documents) # Save index index.storage_context.persist(persist_dir="./storage") # Load index from llama_index.core import load_index_from_storage, StorageContext storage_context = StorageContext.from_defaults(persist_dir="./storage") index = load_index_from_storage(storage_context)
python# Basic query query_engine = index.as_query_engine() response = query_engine.query("What is the main topic?") print(response) # Streaming response query_engine = index.as_query_engine(streaming=True) response = query_engine.query("Explain quantum computing") for text in response.response_gen: print(text, end="", flush=True) # Custom configuration query_engine = index.as_query_engine( similarity_top_k=3, # Return top 3 chunks response_mode="compact", # Or "tree_summarize", "simple_summarize" verbose=True )
python# Vector retriever retriever = index.as_retriever(similarity_top_k=5) nodes = retriever.retrieve("machine learning") # With filtering retriever = index.as_retriever( similarity_top_k=3, filters={"metadata.category": "tutorial"} ) # Custom retriever from llama_index.core.retrievers import BaseRetriever class CustomRetriever(BaseRetriever): def _retrieve(self, query_bundle): # Your custom retrieval logic return nodes
pythonfrom llama_index.core.agent import FunctionAgent from llama_index.llms.openai import OpenAI # Define tools def multiply(a: int, b: int) -> int: """Multiply two numbers.""" return a * b def add(a: int, b: int) -> int: """Add two numbers.""" return a + b # Create agent llm = OpenAI(model="gpt-4o") agent = FunctionAgent.from_tools( tools=[multiply, add], llm=llm, verbose=True ) # Use agent response = agent.chat("What is 25 * 17 + 142?") print(response)
pythonfrom llama_index.core.tools import QueryEngineTool # Create index as before index = VectorStoreIndex.from_documents(documents) # Wrap query engine as tool query_tool = QueryEngineTool.from_defaults( query_engine=index.as_query_engine(), name="python_docs", description="Useful for answering questions about Python programming" ) # Agent with document search + calculator agent = FunctionAgent.from_tools( tools=[query_tool, multiply, add], llm=llm ) # Agent decides when to search docs vs calculate response = agent.chat("According to the docs, what is Python used for?")
pythonfrom llama_index.core.chat_engine import CondensePlusContextChatEngine # Chat with memory chat_engine = index.as_chat_engine( chat_mode="condense_plus_context", # Or "context", "react" verbose=True ) # Multi-turn conversation response1 = chat_engine.chat("What is Python?") response2 = chat_engine.chat("Can you give examples?") # Remembers context response3 = chat_engine.chat("What about web frameworks?")
pythonfrom llama_index.core.vector_stores import MetadataFilters, ExactMatchFilter # Filter by metadata filters = MetadataFilters( filters=[ ExactMatchFilter(key="category", value="tutorial"), ExactMatchFilter(key="difficulty", value="beginner") ] ) retriever = index.as_retriever( similarity_top_k=3, filters=filters ) query_engine = index.as_query_engine(filters=filters)
pythonfrom pydantic import BaseModel from llama_index.core.output_parsers import PydanticOutputParser class Summary(BaseModel): title: str main_points: list[str] conclusion: str # Get structured response output_parser = PydanticOutputParser(output_cls=Summary) query_engine = index.as_query_engine(output_parser=output_parser) response = query_engine.query("Summarize the document") summary = response # Pydantic model print(summary.title, summary.main_points)
python# Load all supported formats documents = SimpleDirectoryReader( "./data", recursive=True, required_exts=[".pdf", ".docx", ".txt", ".md"] ).load_data()
pythonfrom llama_index.readers.web import BeautifulSoupWebReader reader = BeautifulSoupWebReader() documents = reader.load_data(urls=[ "https://docs.python.org/3/tutorial/", "https://docs.python.org/3/library/" ])
pythonfrom llama_index.readers.database import DatabaseReader reader = DatabaseReader( sql_database_uri="postgresql://user:pass@localhost/db" ) documents = reader.load_data(query="SELECT * FROM articles")
pythonfrom llama_index.readers.json import JSONReader reader = JSONReader() documents = reader.load_data("https://api.example.com/data.json")
pythonfrom llama_index.vector_stores.chroma import ChromaVectorStore import chromadb # Initialize Chroma db = chromadb.PersistentClient(path="./chroma_db") collection = db.get_or_create_collection("my_collection") # Create vector store vector_store = ChromaVectorStore(chroma_collection=collection) # Use in index from llama_index.core import StorageContext storage_context = StorageContext.from_defaults(vector_store=vector_store) index = VectorStoreIndex.from_documents(documents, storage_context=storage_context)
pythonfrom llama_index.vector_stores.pinecone import PineconeVectorStore import pinecone # Initialize Pinecone pinecone.init(api_key="your-key", environment="us-west1-gcp") pinecone_index = pinecone.Index("my-index") # Create vector store vector_store = PineconeVectorStore(pinecone_index=pinecone_index) storage_context = StorageContext.from_defaults(vector_store=vector_store) index = VectorStoreIndex.from_documents(documents, storage_context=storage_context)
pythonfrom llama_index.vector_stores.faiss import FaissVectorStore import faiss # Create FAISS index d = 1536 # Dimension of embeddings faiss_index = faiss.IndexFlatL2(d) vector_store = FaissVectorStore(faiss_index=faiss_index) storage_context = StorageContext.from_defaults(vector_store=vector_store) index = VectorStoreIndex.from_documents(documents, storage_context=storage_context)
pythonfrom llama_index.llms.anthropic import Anthropic from llama_index.core import Settings # Set global LLM Settings.llm = Anthropic(model="claude-sonnet-4-5-20250929") # Now all queries use Anthropic query_engine = index.as_query_engine()
pythonfrom llama_index.embeddings.huggingface import HuggingFaceEmbedding # Use HuggingFace embeddings Settings.embed_model = HuggingFaceEmbedding( model_name="sentence-transformers/all-mpnet-base-v2" ) index = VectorStoreIndex.from_documents(documents)
pythonfrom llama_index.core import PromptTemplate qa_prompt = PromptTemplate( "Context: {context_str}\n" "Question: {query_str}\n" "Answer the question based only on the context. " "If the answer is not in the context, say 'I don't know'.\n" "Answer: " ) query_engine = index.as_query_engine(text_qa_template=qa_prompt)
pythonfrom llama_index.core import SimpleDirectoryReader from llama_index.multi_modal_llms.openai import OpenAIMultiModal # Load images and documents documents = SimpleDirectoryReader( "./data", required_exts=[".jpg", ".png", ".pdf"] ).load_data() # Multi-modal index index = VectorStoreIndex.from_documents(documents) # Query with multi-modal LLM multi_modal_llm = OpenAIMultiModal(model="gpt-4o") query_engine = index.as_query_engine(llm=multi_modal_llm) response = query_engine.query("What is in the diagram on page 3?")
pythonfrom llama_index.core.evaluation import RelevancyEvaluator, FaithfulnessEvaluator # Evaluate relevance relevancy = RelevancyEvaluator() result = relevancy.evaluate_response( query="What is Python?", response=response ) print(f"Relevancy: {result.passing}") # Evaluate faithfulness (no hallucination) faithfulness = FaithfulnessEvaluator() result = faithfulness.evaluate_response( query="What is Python?", response=response ) print(f"Faithfulness: {result.passing}")
python# Complete RAG pipeline documents = SimpleDirectoryReader("docs").load_data() index = VectorStoreIndex.from_documents(documents) index.storage_context.persist(persist_dir="./storage") # Query query_engine = index.as_query_engine( similarity_top_k=3, response_mode="compact", verbose=True ) response = query_engine.query("What is the main topic?") print(response) print(f"Sources: {[node.metadata['file_name'] for node in response.source_nodes]}")
python# Conversational interface chat_engine = index.as_chat_engine( chat_mode="condense_plus_context", verbose=True ) # Multi-turn chat while True: user_input = input("You: ") if user_input.lower() == "quit": break response = chat_engine.chat(user_input) print(f"Bot: {response}")
| Operation | Latency | Notes | |-----------|---------|-------| | Index 100 docs | ~10-30s | One-time, can persist | | Query (vector) | ~0.5-2s | Retrieval + LLM | | Streaming query | ~0.5s first token | Better UX | | Agent with tools | ~3-8s | Multiple tool calls |
| Feature | LlamaIndex | LangChain | |---------|------------|-----------| | Best for | RAG, document Q&A | Agents, general LLM apps | | Data connectors | 300+ (LlamaHub) | 100+ | | RAG focus | Core feature | One of many | | Learning curve | Easier for RAG | Steeper | | Customization | High | Very high | | Documentation | Excellent | Good |
Use LlamaIndex when:
Use LangChain when:
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 12,457 | 8,083 | -35% | 1 | 1 | 0% | 2,863 | 5,773 | +102% | 0 | 0 | — |
case-02 | fail→fail | 10,795 | 8,525 | -21% | 1 | 1 | 0% | 2,327 | 5,736 | +146% | 0 | 0 | — |
case-03 | pass→pass | 15,605 | 11,776 | -25% | 1 | 1 | 0% | 3,318 | 6,364 | +92% | 0 | 0 | — |
case-04 | pass→pass | 10,272 | 10,813 | +5% | 1 | 1 | 0% | 2,203 | 6,468 | +194% | 0 | 0 | — |
case-05 | pass→pass | 8,323 | 7,170 | -14% | 1 | 1 | 0% | 1,961 | 5,594 | +185% | 0 | 0 | — |
case-06 | pass→pass | 3,723 | 3,562 | -4% | 1 | 1 | 0% | 950 | 4,737 | +399% | 0 | 0 | — |
case-07 | pass→pass | 5,312 | 3,801 | -28% | 1 | 1 | 0% | 924 | 4,813 | +421% | 0 | 0 | — |
case-08 | pass→pass | 6,521 | 5,598 | -14% | 1 | 1 | 0% | 1,235 | 4,979 | +303% | 0 | 0 | — |
case-09 | pass→pass | 8,061 | 67,096 | +732% | 1 | 1 | 0% | 1,801 | 5,417 | +201% | 0 | 0 | — |
case-10 | pass→pass | 7,935 | 4,449 | -44% | 1 | 1 | 0% | 1,676 | 4,858 | +190% | 0 | 0 | — |
case-11 | pass→pass | 8,634 | 7,905 | -8% | 1 | 1 | 0% | 1,630 | 5,330 | +227% | 0 | 0 | — |
case-12 | pass→pass | 12,327 | 11,208 | -9% | 1 | 1 | 0% | 2,158 | 5,995 | +178% | 0 | 0 | — |
case-13 | pass→pass | 9,689 | 3,683 | -62% | 1 | 1 | 0% | 1,626 | 4,620 | +184% | 0 | 0 | — |
case-14 | pass→pass | 6,186 | 2,840 | -54% | 1 | 1 | 0% | 1,051 | 4,410 | +320% | 0 | 0 | — |
case-15 | fail→pass | 9,080 | 6,369 | -30% | 1 | 1 | 0% | 1,706 | 5,020 | +194% | 0 | 0 | — |
case-16 | pass→pass | 9,903 | 6,598 | -33% | 1 | 1 | 0% | 1,755 | 5,296 | +202% | 0 | 0 | — |
case-17 | pass→pass | 13,705 | 11,103 | -19% | 1 | 1 | 0% | 2,453 | 5,960 | +143% | 0 | 0 | — |
case-18 | pass→pass | 9,686 | 7,101 | -27% | 1 | 1 | 0% | 1,602 | 5,185 | +224% | 0 | 0 | — |
case-19 | pass→pass | 10,403 | 5,270 | -49% | 1 | 1 | 0% | 1,719 | 4,837 | +181% | 0 | 0 | — |
case-20 | pass→pass | 13,717 | 8,966 | -35% | 1 | 1 | 0% | 2,609 | 5,588 | +114% | 0 | 0 | — |
case-21 | pass→pass | 15,725 | 10,160 | -35% | 1 | 1 | 0% | 2,734 | 5,892 | +116% | 0 | 0 | — |
case-22 | pass→pass | 16,479 | 8,091 | -51% | 1 | 1 | 0% | 2,780 | 5,452 | +96% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.