▸case-01 My current script takes way too long to process a 5GB tabular dataset. Could you show me how to efficiently read this large file in chunks, filter out rows where the status is 'inactive', and save the optimized result to a new parquet file? | fail→pass | 9,815 | 8,397 | -14% | 1 | 1 | 0% | 2,240 | 1,760 | -21% | 0 | 0 | — |
▸case-02 I have a DataFrame with 10 million rows and a 'department' column containing only 5 unique string values (like 'HR', 'Engineering'). Memory usage is exploding. How do I optimize the data type of this specific column to reduce memory footprint during the read or immediately after? | pass→pass | 9,166 | 6,378 | -30% | 1 | 1 | 0% | 1,550 | 1,281 | -17% | 0 | 0 | — |
▸case-03 I have two DataFrames: 'trades' with exact transaction timestamps, and 'quotes' with market price updates at irregular intervals. I need to join them so each trade gets the most recent quote that occurred before the trade timestamp. A standard left join doesn't work because the timestamps don't match exactly. What is the most efficient built-in function to align these? | pass→pass | 7,705 | 4,689 | -39% | 1 | 1 | 0% | 1,313 | 1,045 | -20% | 0 | 0 | — |
▸case-12 My time series DataFrame has missing temperature readings for a few random hours. Forward filling leaves flat lines, which isn't accurate for temperature. How do I estimate these missing values by drawing a straight line between the known points? | pass→pass | 11,890 | 8,092 | -32% | 1 | 1 | 0% | 1,873 | 1,598 | -15% | 0 | 0 | — |
▸case-04 I am filtering a massive DataFrame 'df' using multiple conditions: `df[(df['age'] > 30) & (df['income'] < 50000) & (df['status'] == 'active')]`. This creates intermediate boolean arrays and causes memory spikes. What is the most memory-efficient built-in method to write this filtering operation? | pass→pass | 8,599 | 7,324 | -15% | 1 | 1 | 0% | 1,502 | 1,513 | +1% | 0 | 0 | — |
▸case-05 I have a sales DataFrame with a datetime index. I want to calculate the total sales amount for each month. Instead of extracting the month into a new column and grouping by that, how can I group directly by month frequency within the groupby call? | pass→pass | 8,426 | 7,168 | -15% | 1 | 1 | 0% | 1,654 | 1,633 | -1% | 0 | 0 | — |
▸case-06 My DataFrame has a column called 'tags' where each cell contains a list of strings (e.g., ['python', 'pandas']). I want to transform the DataFrame so that each tag gets its own row, duplicating the other column values. How do I do this without writing a custom loop or apply function? | pass→pass | 6,045 | 5,878 | -3% | 1 | 1 | 0% | 1,215 | 1,217 | +0% | 0 | 0 | — |
▸case-07 I want to update a 'score' column in my DataFrame. If the score is less than 0, I want to replace it with 0, otherwise keep the original value. What is the idiomatic method to do this without using boolean indexing assignment like `df.loc[df['score'] < 0, 'score'] = 0`? | pass→pass | 8,619 | 5,000 | -42% | 1 | 1 | 0% | 1,596 | 1,045 | -35% | 0 | 0 | — |
▸case-21 I have a pandas DataFrame containing monthly sales data. I want to build an interactive web dashboard where users can select a region from a dropdown and see the sales chart update dynamically in their browser. How should I build this? | pass→pass | 11,379 | 13,534 | +19% | 1 | 1 | 0% | 2,307 | 2,744 | +19% | 0 | 0 | — |
▸case-08 I need to compute a new column 'D' using a complex mathematical formula involving columns 'A', 'B', and 'C' on a very large DataFrame. Standard vectorized arithmetic creates large intermediate arrays. How can I evaluate this expression more efficiently using a string-based engine? | pass→pass | 10,977 | 8,432 | -23% | 1 | 1 | 0% | 2,056 | 1,640 | -20% | 0 | 0 | — |
▸case-09 I need to create a frequency table showing the count of occurrences for each combination of 'user_type' (rows) and 'subscription_tier' (columns). What is the specific top-level function designed for computing simple cross-tabulations of two or more factors? | pass→pass | 3,201 | 3,213 | +0% | 1 | 1 | 0% | 534 | 718 | +34% | 0 | 0 | — |
▸case-10 I have a DataFrame with several integer columns currently stored as int64, but the values only range from 0 to 100. I want to automatically reduce them to the smallest possible integer type. What function and parameter combination achieves this downcasting? | pass→pass | 6,708 | 4,549 | -32% | 1 | 1 | 0% | 1,390 | 1,041 | -25% | 0 | 0 | — |
▸case-11 I need to load a CSV file that has 200 columns, but I only need to analyze 'user_id' and 'purchase_amount'. Loading the whole file and then subsetting takes too much memory. How do I restrict the columns at the time of reading? | pass→pass | 11,930 | 5,849 | -51% | 1 | 1 | 0% | 1,807 | 847 | -53% | 0 | 0 | — |
▸case-13 I have a 'color' column with values 'red', 'green', and 'blue'. I need to convert this into three separate binary columns for a machine learning model. What built-in function does this directly? | pass→pass | 4,455 | 3,048 | -32% | 1 | 1 | 0% | 733 | 720 | -2% | 0 | 0 | — |
▸case-14 I have a daily stock price DataFrame. I want to calculate a 7-day moving average for the 'close_price' column. What method chain should I use? | pass→pass | 7,002 | 6,790 | -3% | 1 | 1 | 0% | 1,305 | 1,101 | -16% | 0 | 0 | — |
▸case-15 My DataFrame has a DatetimeIndex with minute-by-minute server load data. I want to aggregate this into hourly averages. What is the most idiomatic method for frequency conversion on a time series index? | pass→pass | 8,030 | 6,890 | -14% | 1 | 1 | 0% | 1,353 | 1,339 | -1% | 0 | 0 | — |
▸case-16 I want to identify all rows in my customer DataFrame that have the exact same 'email' and 'phone_number', keeping only the first occurrence. What method should I use to drop the subsequent identical records? | pass→pass | 7,324 | 3,992 | -45% | 1 | 1 | 0% | 1,117 | 848 | -24% | 0 | 0 | — |
▸case-17 My DataFrame has 50 columns. I want to drop any column that does not have at least 100 non-null values. How do I do this in a single method call without manually calculating null counts? | pass→pass | 5,390 | 4,585 | -15% | 1 | 1 | 0% | 1,034 | 965 | -7% | 0 | 0 | — |
▸case-18 I am using the latest major version of the library and have a large text dataset. The default object dtype for strings is slow and memory-intensive. What specific dtype string should I specify to use the new, highly efficient backend? | pass→pass | 7,817 | 6,757 | -14% | 1 | 1 | 0% | 1,259 | 1,339 | +6% | 0 | 0 | — |
▸case-19 I want to find the proportion (percentage) of each unique value in the 'status' column of my DataFrame, rather than the raw counts. How do I do this using the standard counting method? | pass→pass | 5,944 | 3,183 | -46% | 1 | 1 | 0% | 1,099 | 697 | -37% | 0 | 0 | — |
▸case-20 I have a cleaned pandas DataFrame 'df' with features and a target column 'price'. I need to train a gradient boosting regression model to predict the price. What is the standard Python code to initialize and train this model? | pass→pass | 9,525 | 5,892 | -38% | 1 | 1 | 0% | 1,711 | 1,471 | -14% | 0 | 0 | — |
▸case-22 I have a 50 Terabyte dataset of user logs stored across a 20-node HDFS cluster. I need to run a distributed group-by aggregation that shuffles data across all nodes. I usually use pandas for local data, but what Python framework is designed to execute this distributed cluster workload? | pass→pass | 13,058 | 11,408 | -13% | 1 | 1 | 0% | 2,292 | 2,179 | -5% | 0 | 0 | — |