Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Automotive E2E Safety Analysis expertise. Covers 1 topics: E2E Safety Analysis.
.claude/skills/pangzhenying2025-automotive-e2e-safety-analysis/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-06 | ✗→✓ | ▲ Improved | 131% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 77% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 94% | 0% |
| case-22 | ✗→✓ | ▲ Improved | 74% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 127% | 0% |
Safety analysis framework for end-to-end (E2E) autonomous driving systems that use neural networks for the complete perception-to-control pipeline. Addresses the unique safety challenges of E2E architectures including interpretability, verification, functional safety compliance, and SOTIF analysis for learned driving policies.
端到端自动驾驶安全分析框架
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Traditional Modular Stack:
Sensors → Perception → Prediction → Planning → Control
✓ Each module independently verifiable
✓ Clear failure mode attribution
✗ Information loss at interfaces
✗ Cumulative error propagation
End-to-End Architecture:
Sensors → [Neural Network] → Control
✓ No information loss (raw sensor to action)
✓ Potentially better performance (holistic optimization)
✗ Black-box: hard to verify/interpret
✗ No clear failure mode attribution
✗ ISO 26262 / SOTIF compliance challenges
Hybrid Architecture (现阶段主流):
Sensors → [E2E Backbone] → Structured Output → Safety Layer → Control
├── E2E handles perception + prediction + planning
├── Safety layer provides guardrails and override
├── Structured intermediate representations for interpretability
└── Fallback to rule-based system when confidence low
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━pythone2e_safety_challenges = { "interpretability": { "problem": "Cannot explain why a specific driving decision was made", "impact_on_safety": "Cannot perform systematic failure mode analysis", "mitigation_approaches": [ "Attention map visualization", "Intermediate representation extraction", "Concept-based explanations", "Counterfactual analysis", "Structured output heads (BEV, occupancy, trajectory)", ], }, "verification": { "problem": "Traditional V&V methods insufficient for DNN", "impact_on_safety": "Cannot guarantee behavior in unseen scenarios", "mitigation_approaches": [ "Massive scenario-based testing", "Formal verification of safety envelope", "Runtime monitoring and intervention", "Statistical safety arguments", "Neuron coverage and mutation testing", ], }, "functional_safety_compliance": { "problem": "ISO 26262 assumes decomposable system architecture", "impact_on_safety": "ASIL allocation and decomposition challenging", "mitigation_approaches": [ "Safety wrapper / safety cage architecture", "E2E as QM, safety layer as ASIL-rated", "Redundant conventional perception for monitoring", "ASIL decomposition at system level", "1oo2D architecture (E2E + rule-based)", ], }, "sotif_analysis": { "problem": "Triggering conditions for DNN are fundamentally different", "impact_on_safety": "Unknown-unsafe area potentially larger", "mitigation_approaches": [ "Out-of-distribution detection", "Uncertainty quantification (epistemic + aleatoric)", "Domain adaptation and generalization testing", "Adversarial robustness testing", "Continuous learning with safety constraints", ], }, "data_dependency": { "problem": "Model behavior determined by training data distribution", "impact_on_safety": "Bias, gaps, and distributional shift", "mitigation_approaches": [ "Training data coverage analysis", "Data augmentation for rare scenarios", "Sim-to-real transfer validation", "Geographic/cultural diversity in data", "Data quality monitoring pipeline", ], }, }
安全笼架构
┌────────────────────────────────────────────┐
│ Safety Cage │
│ ┌──────────────────────────────────────┐ │
│ │ E2E Neural Network │ │
│ │ Sensors → [Model] → Trajectory │ │
│ └──────────────┬───────────────────────┘ │
│ │ proposed trajectory │
│ ┌──────────────▼───────────────────────┐ │
│ │ Safety Monitor (ASIL-rated) │ │
│ │ ├── Collision check (TTC > threshold)│ │
│ │ ├── Kinematic feasibility │ │
│ │ ├── ODD boundary check │ │
│ │ ├── Traffic rule compliance │ │
│ │ └── Comfort envelope check │ │
│ └──────────────┬───────────────────────┘ │
│ ┌─────────┴─────────┐ │
│ │ Safe? │ │
│ Yes│ No │ │
│ ▼ ▼ │
│ [Execute E2E] [Execute Safe Fallback] │
│ ├── Maintain lane + brake │
│ ├── Emergency stop │
│ └── Handoff to driver │
└────────────────────────────────────────────┘双通道架构(1oo2D)
┌────────────────────────────────────────────┐
│ Path A: E2E Model (Performance Channel) │
│ Sensors → DNN → Trajectory A │
│ (High performance, QM or low ASIL) │
├────────────────────────────────────────────┤
│ Path B: Rule-Based (Safety Channel) │
│ Sensors → Classical Pipeline → Trajectory B │
│ (Conservative, ASIL-rated) │
├────────────────────────────────────────────┤
│ Arbitration Logic (ASIL-rated) │
│ ├── If A and B agree → Execute A (better) │
│ ├── If A and B disagree mildly → Execute B │
│ ├── If A proposes unsafe action → Override │
│ └── If both uncertain → MRM │
└────────────────────────────────────────────┘E2E系统SOTIF分析特殊考虑
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
1. DNN-Specific Triggering Conditions
├── Out-of-distribution inputs
│ ├── Novel objects (未见过的物体)
│ ├── Rare weather/lighting combinations
│ └── Geographic/cultural differences
├── Adversarial perturbations
│ ├── Physical adversarial patches
│ ├── Sensor spoofing attacks
│ └── Natural adversarial examples
├── Distribution shift
│ ├── Season/time-of-day shift
│ ├── Sensor aging/degradation
│ └── Map/infrastructure changes
└── Model uncertainty
├── Epistemic uncertainty (data gaps)
└── Aleatoric uncertainty (inherent noise)
2. E2E-Specific Hazardous Behaviors
├── Sudden trajectory change (mode collapse)
├── Freezing (model produces no output)
├── Imitation of human errors (from training data)
├── Overconfident wrong predictions
└── Inconsistent behavior across similar scenarios
3. Validation Approach
├── Scenario-based: >10M km equivalent simulation
├── Adversarial: Targeted attack scenarios
├── OOD detection: Calibrated uncertainty monitoring
├── Regression: Version-to-version comparison
└── Shadow mode: Real-world deployment monitoring
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━python# Functional Safety Strategy for E2E AD Systems fusa_strategy = { "system_level": { "asil": "ASIL_D", "approach": "ASIL decomposition at system level", "e2e_path": "QM or ASIL_A (performance, not safety-rated)", "safety_path": "ASIL_C/D (rule-based monitor + override)", "rationale": "E2E model cannot be developed per ASIL process, " "but system achieves ASIL through architectural decomposition", }, "safety_mechanisms": { "sm1_collision_monitor": { "asil": "ASIL_D", "function": "Check E2E trajectory for collision risk", "implementation": "Deterministic algorithm, MISRA C compliant", }, "sm2_kinematics_check": { "asil": "ASIL_C", "function": "Verify trajectory is physically feasible", "implementation": "Vehicle dynamics model with safety margins", }, "sm3_odd_monitor": { "asil": "ASIL_B", "function": "Monitor ODD compliance", "implementation": "Rule-based ODD boundary detection", }, "sm4_model_health": { "asil": "ASIL_B", "function": "Monitor E2E model inference health", "implementation": "Latency, output range, confidence monitoring", }, }, "safe_state": { "level1": "Maintain current lane + gradual braking", "level2": "Emergency braking + hazard lights", "level3": "Full stop in safe position", }, }
automotive-dfm-benchmarking — DFM framework for E2E evaluationautomotive-sotif-hazard-scenario — SOTIF scenario constructionautomotive-china-l3-ads-compliance — Chinese L3 regulatory requirementssensor-fusion-perception — Perception system fundamentals| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-04 | pass→pass | 14,020 | 18,575 | +32% | 1 | 1 | 0% | 2,585 | 5,866 | +127% | 0 | 0 | — |
case-05 | fail→fail | 31,618 | 23,900 | -24% | 1 | 1 | 0% | 1,560 | 7,574 | +386% | 0 | 0 | — |
case-01 | fail→fail | 25,901 | 29,823 | +15% | 1 | 1 | 0% | 5,089 | 8,554 | +68% | 0 | 0 | — |
case-02 | pass→pass | 14,196 | 19,600 | +38% | 1 | 1 | 0% | 3,220 | 6,299 | +96% | 0 | 0 | — |
case-03 | pass→pass | 12,275 | 16,742 | +36% | 1 | 1 | 0% | 2,291 | 5,303 | +131% | 0 | 0 | — |
case-06 | fail→pass | 18,795 | 21,922 | +17% | 1 | 1 | 0% | 2,676 | 6,187 | +131% | 0 | 0 | — |
case-07 | fail→fail | 14,260 | 18,514 | +30% | 1 | 1 | 0% | 2,394 | 5,785 | +142% | 0 | 0 | — |
case-08 | fail→pass | 19,860 | 23,197 | +17% | 1 | 1 | 0% | 3,820 | 6,754 | +77% | 0 | 0 | — |
case-09 | pass→pass | 14,334 | 14,896 | +4% | 1 | 1 | 0% | 2,330 | 5,026 | +116% | 0 | 0 | — |
case-10 | fail→fail | 13,946 | 19,161 | +37% | 1 | 1 | 0% | 2,544 | 6,036 | +137% | 0 | 0 | — |
case-11 | pass→pass | 14,132 | 16,224 | +15% | 1 | 1 | 0% | 2,661 | 5,583 | +110% | 0 | 0 | — |
case-12 | fail→pass | 16,921 | 25,578 | +51% | 1 | 1 | 0% | 4,151 | 8,035 | +94% | 0 | 0 | — |
case-13 | pass→pass | 16,324 | 22,191 | +36% | 1 | 1 | 0% | 3,173 | 6,562 | +107% | 0 | 0 | — |
case-14 | pass→pass | 14,118 | 17,942 | +27% | 1 | 1 | 0% | 2,761 | 5,819 | +111% | 0 | 0 | — |
case-15 | pass→pass | 16,505 | 17,297 | +5% | 1 | 1 | 0% | 3,451 | 6,032 | +75% | 0 | 0 | — |
case-16 | pass→pass | 9,592 | 14,335 | +49% | 1 | 1 | 0% | 1,869 | 5,222 | +179% | 0 | 0 | — |
case-22 | fail→pass | 14,208 | 12,737 | -10% | 1 | 1 | 0% | 2,813 | 4,889 | +74% | 0 | 0 | — |
case-17 | pass→pass | 7,692 | 10,570 | +37% | 1 | 1 | 0% | 1,538 | 3,944 | +156% | 0 | 0 | — |
case-18 | pass→pass | 9,829 | 39,156 | +298% | 1 | 1 | 0% | 1,881 | 5,732 | +205% | 0 | 0 | — |
case-19 | pass→pass | 19,211 | 28,732 | +50% | 1 | 1 | 0% | 3,875 | 8,523 | +120% | 0 | 0 | — |
case-20 | pass→pass | 16,955 | 19,121 | +13% | 1 | 1 | 0% | 3,205 | 6,279 | +96% | 0 | 0 | — |
case-21 | pass→pass | 24,611 | 33,420 | +36% | 1 | 1 | 0% | 4,193 | 7,481 | +78% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 21 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +18 percentage points is the difference between those two pass rates over the 21 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.