Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Scalable data processing for ML workloads. Streaming execution across CPU/GPU, supports Parquet/CSV/JSON/images. Integrates with Ray Train, PyTorch, TensorFlow. Scales from single machine to 100s of nodes. Use for batch inference, data preprocessing, multi-modal data loading, or distributed ETL pipelines.
.claude/skills/openlair-ray-data/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-15 | ✗→✓ | ▲ Improved | 64% | 0% |
| case-01 | ✓→✓ | = Same ✓ | 57% | 0% |
| case-02 | ✓→✓ | = Same ✓ | 31% | 0% |
| case-03 | ✓→✓ | = Same ✓ | 59% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 72% | 0% |
Distributed data processing library for ML and AI workloads.
Use Ray Data when:
Key features:
Use alternatives instead:
bashpip install -U 'ray[data]'
pythonimport ray # Read Parquet files ds = ray.data.read_parquet("s3://bucket/data/*.parquet") # Transform data (lazy execution) ds = ds.map_batches(lambda batch: {"processed": batch["text"].str.lower()}) # Consume data for batch in ds.iter_batches(batch_size=100): print(batch)
pythonimport ray from ray.train import ScalingConfig from ray.train.torch import TorchTrainer # Create dataset train_ds = ray.data.read_parquet("s3://bucket/train/*.parquet") def train_func(config): # Access dataset in training train_ds = ray.train.get_dataset_shard("train") for epoch in range(10): for batch in train_ds.iter_batches(batch_size=32): # Train on batch pass # Train with Ray trainer = TorchTrainer( train_func, datasets={"train": train_ds}, scaling_config=ScalingConfig(num_workers=4, use_gpu=True) ) trainer.fit()
pythonimport ray # Parquet (recommended for ML) ds = ray.data.read_parquet("s3://bucket/data/*.parquet") # CSV ds = ray.data.read_csv("s3://bucket/data/*.csv") # JSON ds = ray.data.read_json("gs://bucket/data/*.json") # Images ds = ray.data.read_images("s3://bucket/images/")
python# From list ds = ray.data.from_items([{"id": i, "value": i * 2} for i in range(1000)]) # From range ds = ray.data.range(1000000) # Synthetic data # From pandas import pandas as pd df = pd.DataFrame({"col1": [1, 2, 3], "col2": [4, 5, 6]}) ds = ray.data.from_pandas(df)
python# Batch transformation (fast) def process_batch(batch): batch["doubled"] = batch["value"] * 2 return batch ds = ds.map_batches(process_batch, batch_size=1000)
python# Row-by-row (slower) def process_row(row): row["squared"] = row["value"] ** 2 return row ds = ds.map(process_row)
python# Filter rows ds = ds.filter(lambda row: row["value"] > 100)
python# Group by column ds = ds.groupby("category").count() # Custom aggregation ds = ds.groupby("category").map_groups(lambda group: {"sum": group["value"].sum()})
python# Use GPU for preprocessing def preprocess_images_gpu(batch): import torch images = torch.tensor(batch["image"]).cuda() # GPU preprocessing processed = images * 255 return {"processed": processed.cpu().numpy()} ds = ds.map_batches( preprocess_images_gpu, batch_size=64, num_gpus=1 # Request GPU )
python# Write to Parquet ds.write_parquet("s3://bucket/output/") # Write to CSV ds.write_csv("output/") # Write to JSON ds.write_json("output/")
python# Control parallelism ds = ds.repartition(100) # 100 blocks for 100-core cluster
python# Larger batches = faster vectorized ops ds.map_batches(process_fn, batch_size=10000) # vs batch_size=100
python# Process data larger than memory ds = ray.data.read_parquet("s3://huge-dataset/") for batch in ds.iter_batches(batch_size=1000): process(batch) # Streamed, not loaded to memory
pythonimport ray # Load model def load_model(): # Load once per worker return MyModel() # Inference function class BatchInference: def __init__(self): self.model = load_model() def __call__(self, batch): predictions = self.model(batch["input"]) return {"prediction": predictions} # Run distributed inference ds = ray.data.read_parquet("s3://data/") predictions = ds.map_batches(BatchInference, batch_size=32, num_gpus=1) predictions.write_parquet("s3://output/")
python# Multi-step pipeline ds = ( ray.data.read_parquet("s3://raw/") .map_batches(clean_data) .map_batches(tokenize) .map_batches(augment) .write_parquet("s3://processed/") )
python# Convert to PyTorch torch_ds = ds.to_torch(label_column="label", batch_size=32) for batch in torch_ds: # batch is dict with tensors inputs, labels = batch["features"], batch["label"]
python# Convert to TensorFlow tf_ds = ds.to_tf(feature_columns=["image"], label_column="label", batch_size=32) for features, labels in tf_ds: # Train model pass
| Format | Read | Write | Use Case | |--------|------|-------|----------| | Parquet | ✅ | ✅ | ML data (recommended) | | CSV | ✅ | ✅ | Tabular data | | JSON | ✅ | ✅ | Semi-structured | | Images | ✅ | ❌ | Computer vision | | NumPy | ✅ | ✅ | Arrays | | Pandas | ✅ | ❌ | DataFrames |
Scaling (processing 100GB data):
GPU acceleration (image preprocessing):
Production deployments:
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 10,719 | 5,761 | -46% | 1 | 1 | 0% | 1,992 | 3,129 | +57% | 0 | 0 | — |
case-02 | pass→pass | 13,484 | 5,957 | -56% | 1 | 1 | 0% | 2,301 | 3,005 | +31% | 0 | 0 | — |
case-03 | pass→pass | 11,620 | 6,309 | -46% | 1 | 1 | 0% | 1,993 | 3,172 | +59% | 0 | 0 | — |
case-04 | pass→pass | 9,136 | 4,839 | -47% | 1 | 1 | 0% | 1,762 | 3,025 | +72% | 0 | 0 | — |
case-05 | pass→pass | 70,569 | 7,283 | -90% | 1 | 1 | 0% | 2,123 | 3,449 | +62% | 0 | 0 | — |
case-06 | pass→pass | 2,833 | 2,689 | -5% | 1 | 1 | 0% | 643 | 2,599 | +304% | 0 | 0 | — |
case-07 | pass→pass | 4,330 | 3,511 | -19% | 1 | 1 | 0% | 850 | 2,732 | +221% | 0 | 0 | — |
case-08 | pass→pass | 4,454 | 2,424 | -46% | 1 | 1 | 0% | 878 | 2,537 | +189% | 0 | 0 | — |
case-09 | pass→pass | 5,210 | 2,049 | -61% | 1 | 1 | 0% | 1,061 | 2,454 | +131% | 0 | 0 | — |
case-10 | pass→pass | 10,994 | 6,834 | -38% | 1 | 1 | 0% | 1,959 | 3,070 | +57% | 0 | 0 | — |
case-11 | pass→pass | 6,327 | 6,597 | +4% | 1 | 1 | 0% | 906 | 2,965 | +227% | 0 | 0 | — |
case-12 | pass→pass | 5,603 | 3,073 | -45% | 1 | 1 | 0% | 901 | 2,533 | +181% | 0 | 0 | — |
case-13 | pass→pass | 3,971 | 2,233 | -44% | 1 | 1 | 0% | 675 | 2,434 | +261% | 0 | 0 | — |
case-14 | pass→pass | 3,437 | 2,089 | -39% | 1 | 1 | 0% | 526 | 2,348 | +346% | 0 | 0 | — |
case-15 | fail→pass | 10,804 | 3,362 | -69% | 1 | 1 | 0% | 1,550 | 2,537 | +64% | 0 | 0 | — |
case-16 | pass→pass | 5,177 | 2,793 | -46% | 1 | 1 | 0% | 762 | 2,531 | +232% | 0 | 0 | — |
case-17 | pass→pass | 15,662 | 11,911 | -24% | 1 | 1 | 0% | 2,088 | 3,749 | +80% | 0 | 0 | — |
case-18 | pass→pass | 20,499 | 15,344 | -25% | 1 | 1 | 0% | 2,871 | 4,217 | +47% | 0 | 0 | — |
case-19 | pass→pass | 11,965 | 9,842 | -18% | 1 | 1 | 0% | 1,990 | 3,612 | +82% | 0 | 0 | — |
case-20 | pass→pass | 4,720 | 2,701 | -43% | 1 | 1 | 0% | 729 | 2,461 | +238% | 0 | 0 | — |
case-21 | pass→pass | 4,600 | 2,023 | -56% | 1 | 1 | 0% | 718 | 2,416 | +236% | 0 | 0 | — |
case-22 | pass→pass | 3,852 | 2,874 | -25% | 1 | 1 | 0% | 569 | 2,503 | +340% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +5 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.