▸case-01 I am building a RAG pipeline and need a Python script to handle my vector storage. Please write the code to create a new storage group, ingest a batch of text files with associated tags, and then perform a search that restricts the results based on those tags and specific text content. | fail→pass | 17,352 | 10,087 | -42% | 1 | 1 | 0% | 3,607 | 2,259 | -37% | 0 | 0 | — |
▸case-02 Can you provide a configuration guide with code examples showing how I should initialize my vector database differently when I am building the app locally versus when I deploy it live to my users? I need to know the recommended storage approaches for both environments. | fail→fail | 16,764 | 14,099 | -16% | 1 | 1 | 0% | 3,566 | 2,757 | -23% | 0 | 0 | — |
▸case-03 We are designing a system that serves multiple different clients and need to keep their data isolated in our vector store. Please provide an architectural overview and a code snippet showing how to initialize the storage containers for different clients, including how to configure the mathematical method used to calculate similarity between vectors. | fail→pass | 18,375 | 16,930 | -8% | 1 | 1 | 0% | 3,774 | 3,556 | -6% | 0 | 0 | — |
▸case-04 I am writing automated unit tests for my vector retrieval logic. I need the storage initialization to be as fast as possible and leave no artifacts on the filesystem. Provide the initialization code. | fail→fail | 9,670 | 5,550 | -43% | 1 | 1 | 0% | 1,784 | 1,279 | -28% | 0 | 0 | — |
▸case-05 When creating a new vector container for my embeddings, I want to use inner product instead of the default Euclidean calculation. Show me the exact configuration parameter to pass during initialization. | fail→fail | 7,007 | 3,356 | -52% | 1 | 1 | 0% | 1,416 | 760 | -46% | 0 | 0 | — |
▸case-06 I am setting up a vector storage group and need to calculate similarity using cosine distance. What is the specific string value I must provide in the configuration? | fail→pass | 5,016 | 2,433 | -51% | 1 | 1 | 0% | 917 | 579 | -37% | 0 | 0 | — |
▸case-07 I want to explicitly configure my vector storage to use Euclidean distance for similarity calculations. What is the exact string identifier required by the configuration? | fail→fail | 5,586 | 2,829 | -49% | 1 | 1 | 0% | 1,054 | 713 | -32% | 0 | 0 | — |
▸case-08 I have ingested documents into my Chroma vector database but realized some of the text needs to be corrected. I don't want to delete and recreate the collection. How can I modify the text of an existing document using its identifier? | fail→fail | 7,124 | 11,203 | +57% | 1 | 1 | 0% | 1,404 | 2,003 | +43% | 0 | 0 | — |
▸case-09 I am syncing a local folder of text files to my Chroma vector database. Some files are new, but others might already exist in the database and just have updated content. What single collection method can I use to add the new ones and update the existing ones simultaneously? | fail→fail | 4,350 | 4,649 | +7% | 1 | 1 | 0% | 753 | 1,128 | +50% | 0 | 0 | — |
▸case-10 I am setting up a new Python project for a RAG pipeline that requires vector storage. What are the two specific Python packages I need to add to my requirements.txt file? | fail→pass | 6,781 | 2,103 | -69% | 1 | 1 | 0% | 1,022 | 560 | -45% | 0 | 0 | — |
▸case-11 I have an existing vector collection. Provide the API methods required to insert new records, modify existing records, and remove obsolete records from the storage. | pass→pass | 11,859 | 6,913 | -42% | 1 | 1 | 0% | 2,460 | 1,668 | -32% | 0 | 0 | — |
▸case-12 I want to use a custom HuggingFace model to generate vectors when creating a new storage collection. How do I configure the collection so it automatically uses this specific model during document ingestion? | fail→fail | 13,379 | 8,941 | -33% | 1 | 1 | 0% | 2,571 | 2,214 | -14% | 0 | 0 | — |
▸case-13 I am configuring local file-based vector storage for my development environment. How do I specify the exact folder path where the database files should be saved on my machine? | fail→fail | 8,283 | 6,207 | -25% | 1 | 1 | 0% | 1,547 | 1,387 | -10% | 0 | 0 | — |
▸case-14 I need to attach custom configuration data to the vector storage container itself, not just to the individual documents inside it. How is this achieved during the initialization step? | fail→fail | 12,496 | 6,634 | -47% | 1 | 1 | 0% | 2,430 | 1,555 | -36% | 0 | 0 | — |
▸case-15 I am querying my vector database and need to restrict the search results exclusively to documents authored by 'Alice'. What specific filter parameter should I use in the query method? | fail→fail | 7,397 | 3,116 | -58% | 1 | 1 | 0% | 1,603 | 726 | -55% | 0 | 0 | — |
▸case-16 I want to search my vector store but only return results if the actual text of the document contains the word 'urgent'. What specific filter parameter allows this content-based restriction? | fail→fail | 9,628 | 4,248 | -56% | 1 | 1 | 0% | 1,745 | 1,082 | -38% | 0 | 0 | — |
▸case-17 I am deploying my vector database to a production environment. What specific setup mode is recommended, and what configuration is required to connect my application to it over the network? | fail→pass | 13,681 | 10,395 | -24% | 1 | 1 | 0% | 2,337 | 2,138 | -9% | 0 | 0 | — |
▸case-18 I am building a SaaS application with vector search. How should I structure the vector database to ensure data from different organizations remains strictly separated and cannot be queried together? | fail→pass | 16,125 | 15,342 | -5% | 1 | 1 | 0% | 2,701 | 3,093 | +15% | 0 | 0 | — |
▸case-19 I am moving my vector database setup from a temporary CI/CD pipeline to a local developer machine where I need data to survive restarts. What deployment mode change should I make? | fail→fail | 8,555 | 4,391 | -49% | 1 | 1 | 0% | 1,518 | 1,053 | -31% | 0 | 0 | — |
▸case-20 I am setting up a Pinecone vector database and want to use their new serverless architecture. How do I configure the index for serverless deployment using their Python SDK? | fail→fail | 7,780 | 7,212 | -7% | 1 | 1 | 0% | 1,642 | 1,609 | -2% | 0 | 0 | — |
▸case-21 I am using Weaviate and want to perform a hybrid search combining BM25 keyword search and vector similarity. Show me the GraphQL query syntax to achieve this. | fail→fail | 7,590 | 5,668 | -25% | 1 | 1 | 0% | 1,417 | 1,387 | -2% | 0 | 0 | — |
▸case-22 I need to create partitions within a Milvus collection to optimize my vector search performance. How do I define and load these partitions using PyMilvus? | fail→fail | 11,718 | 11,574 | -1% | 1 | 1 | 0% | 2,465 | 2,409 | -2% | 0 | 0 | — |