Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Process large-scale data with Apache Spark. Use when a user asks to process big data, run distributed computations, build ETL pipelines, perform data analysis at scale, or use PySpark for data engineering.
.claude/skills/terminalskills-apache-spark/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 0% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 33% | 0% |
| case-21 | ✓→✓ | = Same ✓ | 13% | 0% |
| case-12 | ✓→✓ | = Same ✓ | 112% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 8% | 0% |
Apache Spark is the standard for distributed data processing. It handles batch processing, streaming, SQL, machine learning, and graph processing. PySpark provides a Python API. Runs on standalone clusters, YARN, Kubernetes, or managed services (Databricks, EMR, Dataproc).
bashpip install pyspark
python# etl/process.py — PySpark data processing from pyspark.sql import SparkSession from pyspark.sql import functions as F spark = SparkSession.builder \ .appName("DataPipeline") \ .config("spark.sql.adaptive.enabled", "true") \ .getOrCreate() # Read data df = spark.read.parquet("s3://bucket/raw/events/") # Transform processed = (df .filter(F.col("event_type").isin(["purchase", "signup"])) .withColumn("date", F.to_date("timestamp")) .withColumn("revenue", F.col("amount") * F.col("quantity")) .groupBy("date", "event_type") .agg( F.count("*").alias("event_count"), F.sum("revenue").alias("total_revenue"), F.countDistinct("user_id").alias("unique_users"), ) .orderBy("date") ) # Write results processed.write \ .mode("overwrite") \ .partitionBy("date") \ .parquet("s3://bucket/processed/daily_metrics/")
python# Register as SQL table df.createOrReplaceTempView("events") result = spark.sql(""" SELECT date_trunc('month', timestamp) as month, COUNT(DISTINCT user_id) as monthly_active_users, SUM(CASE WHEN event_type = 'purchase' THEN amount ELSE 0 END) as revenue FROM events WHERE timestamp >= '2025-01-01' GROUP BY 1 ORDER BY 1 """) result.show()
python# Real-time processing from Kafka stream = spark.readStream \ .format("kafka") \ .option("kafka.bootstrap.servers", "kafka:9092") \ .option("subscribe", "events") \ .load() parsed = stream.select( F.from_json(F.col("value").cast("string"), schema).alias("data") ).select("data.*") query = parsed \ .groupBy(F.window("timestamp", "5 minutes"), "event_type") \ .count() \ .writeStream \ .outputMode("update") \ .format("console") \ .start()
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-21 | pass→pass | 19,793 | 18,463 | -7% | 1 | 1 | 0% | 4,373 | 4,944 | +13% | 0 | 0 | — |
case-01 | fail→pass | 13,382 | 9,739 | -27% | 1 | 1 | 0% | 2,763 | 2,750 | -0% | 0 | 0 | — |
case-02 | fail→pass | 9,265 | 8,676 | -6% | 1 | 1 | 0% | 2,043 | 2,717 | +33% | 0 | 0 | — |
case-03 | fail→fail | 3,831 | 3,970 | +4% | 1 | 1 | 0% | 826 | 1,643 | +99% | 0 | 0 | — |
case-12 | pass→pass | 3,345 | 3,559 | +6% | 1 | 1 | 0% | 615 | 1,303 | +112% | 0 | 0 | — |
case-04 | pass→pass | 5,458 | 2,586 | -53% | 1 | 1 | 0% | 1,100 | 1,186 | +8% | 0 | 0 | — |
case-05 | fail→fail | 9,411 | 4,778 | -49% | 1 | 1 | 0% | 1,500 | 1,763 | +18% | 0 | 0 | — |
case-06 | pass→pass | 4,922 | 3,501 | -29% | 1 | 1 | 0% | 1,073 | 1,339 | +25% | 0 | 0 | — |
case-07 | pass→pass | 6,560 | 2,618 | -60% | 1 | 1 | 0% | 1,312 | 1,250 | -5% | 0 | 0 | — |
case-08 | pass→pass | 7,054 | 3,708 | -47% | 1 | 1 | 0% | 1,408 | 1,512 | +7% | 0 | 0 | — |
case-09 | pass→pass | 12,056 | 6,935 | -42% | 1 | 1 | 0% | 1,523 | 2,129 | +40% | 0 | 0 | — |
case-10 | pass→pass | 2,365 | 2,102 | -11% | 1 | 1 | 0% | 421 | 1,118 | +166% | 0 | 0 | — |
case-11 | pass→pass | 11,851 | 9,416 | -21% | 1 | 1 | 0% | 2,467 | 2,816 | +14% | 0 | 0 | — |
case-13 | pass→pass | 9,239 | 4,477 | -52% | 1 | 1 | 0% | 1,523 | 1,679 | +10% | 0 | 0 | — |
case-14 | pass→pass | 2,888 | 1,946 | -33% | 1 | 1 | 0% | 479 | 1,062 | +122% | 0 | 0 | — |
case-15 | pass→pass | 8,240 | 4,976 | -40% | 1 | 1 | 0% | 1,471 | 1,621 | +10% | 0 | 0 | — |
case-16 | pass→pass | 13,284 | 11,746 | -12% | 1 | 1 | 0% | 2,184 | 2,786 | +28% | 0 | 0 | — |
case-17 | pass→pass | 4,107 | 2,497 | -39% | 1 | 1 | 0% | 761 | 1,175 | +54% | 0 | 0 | — |
case-18 | pass→pass | 11,140 | 2,229 | -80% | 1 | 1 | 0% | 936 | 1,203 | +29% | 0 | 0 | — |
case-19 | pass→pass | 10,461 | 13,989 | +34% | 1 | 1 | 0% | 2,221 | 3,171 | +43% | 0 | 0 | — |
case-20 | pass→pass | 22,084 | 11,845 | -46% | 1 | 1 | 0% | 3,086 | 2,692 | -13% | 0 | 0 | — |
case-22 | pass→pass | 7,834 | 6,930 | -12% | 1 | 1 | 0% | 1,515 | 2,111 | +39% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.