Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Automotive Dfm Benchmarking expertise. Covers 1 topics: Dfm Benchmarking.
.claude/skills/pangzhenying2025-automotive-dfm-benchmarking/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 41% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 92% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 72% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 130% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 167% | 0% |
Benchmarking framework based on the Driver Foundation Model (DFM) concept for evaluating autonomous driving systems. DFM uses large-scale naturalistic driving data (NDD) to model human driver behavior distributions, providing a human-performance baseline for AD system evaluation. This skill supports scenario generation, performance benchmarking, and safety argument construction using NDD-derived metrics.
驾驶员基础模型 (Driver Foundation Model) 概念
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Core Idea:
Human drivers provide a safety baseline:
- Average driver: ~1 fatality per 10^8 km (developed countries)
- Good driver: ~10x safer than average
- AD must be at least as safe as good human driver
DFM Approach:
1. Collect large-scale NDD (7.5M+ aerial trajectories)
2. Model human driving behavior distributions
3. Extract scenario-specific performance baselines
4. Benchmark AD systems against human baselines
5. Quantify relative safety improvement
DFM as Foundation Model:
├── Pre-trained on massive NDD
├── Captures diverse driving styles and conditions
├── Fine-tunable for specific scenarios/regions
├── Provides probabilistic behavior predictions
└── Serves as benchmark generator and evaluator
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━python# NDD Processing Pipeline for DFM class NDDProcessor: """ Process Naturalistic Driving Data for DFM benchmarking. Supports aerial trajectory data (drone-based) and fleet data. """ def __init__(self, data_source: str): """ data_source options: - "aerial": Drone-based trajectory extraction (7.5M+ trajectories) - "fleet": Vehicle-mounted sensor data - "hybrid": Combined aerial + fleet data """ self.source = data_source def extract_driving_primitives(self, trajectories): """ Extract fundamental driving behaviors from trajectory data. Driving primitives: - Car-following (跟车) - Lane-changing (换道) - Merging (汇入) - Diverging (分流) - Crossing (交叉) - Free-driving (自由行驶) """ primitives = { "car_following": self.extract_car_following(trajectories), "lane_change": self.extract_lane_changes(trajectories), "merge": self.extract_merges(trajectories), "diverge": self.extract_diverges(trajectories), "crossing": self.extract_crossings(trajectories), "free_driving": self.extract_free_driving(trajectories), } return primitives def build_behavior_distributions(self, primitives): """ Build statistical distributions of driving behaviors. For car-following: - Time headway distribution: P(THW) - TTC distribution: P(TTC) - Speed distribution: P(v | context) - Acceleration distribution: P(a | context) - Lane offset distribution: P(offset | context) """ distributions = {} for primitive_type, data in primitives.items(): distributions[primitive_type] = { "thw": fit_distribution(data.thw_values), "ttc": fit_distribution(data.ttc_values), "speed": conditional_distribution(data.speeds, data.contexts), "acceleration": conditional_distribution(data.accels, data.contexts), "lateral_offset": fit_distribution(data.offsets), "jerk": fit_distribution(data.jerks), } return distributions def generate_benchmark_scenarios(self, distributions, n_scenarios=1000): """ Generate benchmark scenarios by sampling from behavior distributions. Importance sampling: over-sample from tail (critical) regions """ scenarios = [] for i in range(n_scenarios): # Sample scenario type based on exposure scenario_type = sample_weighted(distributions.keys(), weights=exposure_weights) # Sample parameters from distribution params = sample_from_distribution( distributions[scenario_type], sampling="importance", # over-sample tails criticality_weight=2.0 ) scenarios.append(BenchmarkScenario( type=scenario_type, parameters=params, human_baseline=distributions[scenario_type], )) return scenarios
DFM基准评测指标体系
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Safety Metrics (安全性指标):
├── Collision rate vs. human baseline
├── Near-miss rate (TTC < 1.5s events)
├── Safety-critical event rate
├── Minimum TTC distribution comparison
└── Emergency braking frequency
Comfort Metrics (舒适性指标):
├── Acceleration distribution vs. human
├── Jerk distribution vs. human
├── Lateral offset smoothness
├── Speed profile consistency
└── Ride quality index
Efficiency Metrics (效率指标):
├── Travel time vs. human baseline
├── Throughput at bottlenecks
├── Speed utilization (actual/limit ratio)
└── Lane utilization efficiency
Human-Likeness Metrics (类人性指标):
├── Trajectory similarity (Fréchet distance)
├── Decision timing similarity
├── Speed profile similarity (DTW distance)
├── Gap acceptance distribution similarity
└── Lane change timing similarity
Overall DFM Score:
DFM_score = w_s × Safety + w_c × Comfort + w_e × Efficiency + w_h × HumanLikeness
where: w_s > w_c > w_e > w_h (safety weighted highest)
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━python# DFM Benchmarking Protocol class DFMBenchmark: """ Benchmark AD system against human driver baseline using DFM. """ def __init__(self, dfm_model, ad_system): self.dfm = dfm_model # Trained DFM with human baselines self.ad = ad_system # AD system under test def run_benchmark(self, scenario_suite): """ Run complete benchmark suite. Returns: - Per-scenario comparison (AD vs. human) - Aggregate safety/comfort/efficiency scores - Failure mode analysis - Improvement recommendations """ results = [] for scenario in scenario_suite: # Get human baseline for this scenario human_baseline = self.dfm.predict_behavior(scenario) # Run AD system in same scenario ad_behavior = self.ad.simulate(scenario) # Compare comparison = self.compare_behaviors( human=human_baseline, ad=ad_behavior, scenario=scenario ) results.append(comparison) return self.aggregate_results(results) def compare_behaviors(self, human, ad, scenario): """Compare AD behavior with human baseline""" return { "scenario_id": scenario.id, "safety": { "ad_min_ttc": ad.min_ttc, "human_min_ttc_percentile": human.ttc_percentile(ad.min_ttc), "collision": ad.collision_occurred, "safety_score": self.compute_safety_score(ad, human), }, "comfort": { "ad_max_accel": ad.max_acceleration, "human_accel_percentile": human.accel_percentile(ad.max_acceleration), "ad_max_jerk": ad.max_jerk, "comfort_score": self.compute_comfort_score(ad, human), }, "human_likeness": { "trajectory_distance": frechet_distance(ad.trajectory, human.mean_trajectory), "speed_profile_dtw": dtw_distance(ad.speed_profile, human.mean_speed), "decision_timing_diff": abs(ad.decision_time - human.mean_decision_time), }, } def generate_report(self, results): """Generate benchmark report with visualizations""" report = { "overall_dfm_score": self.compute_overall_score(results), "safety_rating": self.rate_safety(results), "scenarios_worse_than_human": self.find_deficiencies(results), "scenarios_better_than_human": self.find_strengths(results), "improvement_priorities": self.prioritize_improvements(results), } return report
版本对比评测
├── Input: AD System v1.0, v2.0
├── Benchmark: Same DFM scenario suite
├── Output:
│ ├── Per-scenario performance delta
│ ├── Regression identification (v2 worse than v1)
│ ├── Improvement quantification
│ └── Overall DFM score trend
└── Use case: Release gate decision跨平台评测
├── Input: Multiple AD systems (OEM A vs. B vs. C)
├── Benchmark: Standardized DFM scenario suite
├── Output:
│ ├── Comparative safety ranking
│ ├── Comfort comparison
│ ├── Scenario-specific strengths/weaknesses
│ └── Industry positioning
└── Use case: C-NCAP, IIHS, consumer testingSOTIF证据生成
├── Input: AD system + DFM human baselines
├── Analysis: Per-scenario risk comparison
├── Output:
│ ├── Scenarios where AD safer than human → evidence
│ ├── Scenarios where AD less safe → risk
│ ├── Statistical safety argument
│ └── Residual risk quantification
└── Use case: ISO 21448 compliance, type approvalautomotive-scenario-driven-testing — Scenario-based V&V methodologyautomotive-sotif-hazard-scenario — SOTIF scenario constructionautomotive-e2e-safety-analysis — E2E AD safety analysisautomotive-china-l3-ads-compliance — L3 validation requirements| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 32,455 | 26,993 | -17% | 1 | 1 | 0% | 6,227 | 8,806 | +41% | 0 | 0 | — |
case-02 | fail→pass | 29,953 | 28,099 | -6% | 1 | 1 | 0% | 6,227 | 8,806 | +41% | 0 | 0 | — |
case-03 | fail→fail | 12,026 | 12,434 | +3% | 1 | 1 | 0% | 2,484 | 5,273 | +112% | 0 | 0 | — |
case-04 | pass→pass | 15,652 | 20,237 | +29% | 1 | 1 | 0% | 3,120 | 7,140 | +129% | 0 | 0 | — |
case-05 | pass→pass | 11,827 | 5,574 | -53% | 1 | 1 | 0% | 2,372 | 3,698 | +56% | 0 | 0 | — |
case-06 | fail→pass | 9,503 | 3,878 | -59% | 1 | 1 | 0% | 1,755 | 3,374 | +92% | 0 | 0 | — |
case-07 | pass→pass | 17,308 | 16,233 | -6% | 1 | 1 | 0% | 2,656 | 5,629 | +112% | 0 | 0 | — |
case-08 | fail→pass | 11,618 | 7,811 | -33% | 1 | 1 | 0% | 2,337 | 4,017 | +72% | 0 | 0 | — |
case-09 | fail→pass | 15,617 | 20,055 | +28% | 1 | 1 | 0% | 2,997 | 6,906 | +130% | 0 | 0 | — |
case-10 | fail→fail | 13,898 | 15,980 | +15% | 1 | 1 | 0% | 2,515 | 5,762 | +129% | 0 | 0 | — |
case-11 | fail→pass | 15,389 | 26,767 | +74% | 1 | 1 | 0% | 2,766 | 7,397 | +167% | 0 | 0 | — |
case-12 | fail→pass | 13,223 | 19,089 | +44% | 1 | 1 | 0% | 2,310 | 6,517 | +182% | 0 | 0 | — |
case-13 | pass→pass | 14,679 | 19,559 | +33% | 1 | 1 | 0% | 2,825 | 6,352 | +125% | 0 | 0 | — |
case-14 | fail→pass | 17,717 | 25,480 | +44% | 1 | 1 | 0% | 3,162 | 8,258 | +161% | 0 | 0 | — |
case-15 | fail→pass | 7,184 | 2,860 | -60% | 1 | 1 | 0% | 1,353 | 3,124 | +131% | 0 | 0 | — |
case-16 | fail→pass | 11,294 | 4,003 | -65% | 1 | 1 | 0% | 2,393 | 3,418 | +43% | 0 | 0 | — |
case-17 | fail→pass | 13,318 | 12,843 | -4% | 1 | 1 | 0% | 2,830 | 5,300 | +87% | 0 | 0 | — |
case-18 | pass→pass | 13,085 | 9,851 | -25% | 1 | 1 | 0% | 2,442 | 4,685 | +92% | 0 | 0 | — |
case-19 | fail→fail | 10,333 | 3,171 | -69% | 1 | 1 | 0% | 2,057 | 3,205 | +56% | 0 | 0 | — |
case-20 | pass→pass | 20,639 | 17,604 | -15% | 1 | 1 | 0% | 3,615 | 7,042 | +95% | 0 | 0 | — |
case-21 | pass→pass | 18,037 | 32,121 | +78% | 1 | 1 | 0% | 3,824 | 8,560 | +124% | 0 | 0 | — |
case-22 | pass→pass | 27,998 | 24,315 | -13% | 1 | 1 | 0% | 2,323 | 6,613 | +185% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +45 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.