Install any skill in seconds. Free to start, no credit card required.
Get Started Free →China standard: Scenario Safety. # 场景安全评估框架 — ISO 34501/34502 + ISO 34503/34504/34505
.claude/skills/pangzhenying2025-china-scenario-safety/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-02 | ✗→✓ | ▲ Improved | 6% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 134% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 3% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-16 | ✗→✓ | ▲ Improved | -33% | 0% |
| 标准编号 | 名称 | 状态 | 推荐等级 | |---------|------|------|---------| | ISO 34501:2022 | 自动驾驶系统测试场景 术语 | 已发布 | P1 | | ISO 34502:2022 | 基于场景的安全评估框架 | 已发布 | P1 | | ISO 34502 GB征求意见稿 | 基于场景的安全评估框架(中国版) | 征求意见稿 | P3 | | ISO 34503/34504/34505 | 场景描述/分类/生成方法 | DIS阶段 | P3 |
场景安全评估三层模型
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Functional Scenario (功能场景)
└── 自然语言描述的抽象场景
└── 例:"高速公路上前车突然变道,暴露前方静止车辆"
Logical Scenario (逻辑场景)
└── 参数化描述,参数取值为范围/分布
└── 例:ego_speed ∈ [100,120] km/h,
target_speed = 0 km/h,
cut_out_ttc ∈ [2.0, 5.0] s
Concrete Scenario (具体场景)
└── 所有参数赋具体值的可执行场景
└── 例:ego_speed = 110 km/h,
target_speed = 0 km/h,
cut_out_ttc = 3.2 s
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ISO 34502 安全评估流程
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━
Phase 1: 场景识别
├── 从标准/法规提取
├── 从事故数据提取
├── 从自然驾驶数据提取
└── 专家知识补充
Phase 2: 场景描述与参数化
├── 功能场景定义
├── 逻辑场景参数化
└── 参数空间定义
Phase 3: 场景选择
├── 基于风险的优先级排序
├── 覆盖度分析
└── 测试资源分配
Phase 4: 场景执行
├── 仿真测试(批量执行)
├── 封闭场地测试(关键场景)
└── 开放道路测试(真实环境)
Phase 5: 安全论证
├── 通过率统计
├── 覆盖度论证
├── 残余风险评估
└── 安全案例构建
━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━━ISO 34502 ↔ DFM关联
├── ISO 34502定义了场景安全评估的工程框架
├── DFM(驾驶员基础模型)提供:
│ ├── 人类驾驶行为基线(作为安全参考)
│ ├── 场景暴露频率(基于大规模NDD)
│ ├── 场景参数分布(基于7.5M+轨迹数据)
│ └── 场景批评性量化指标
└── 组合使用:DFM为ISO 34502提供数据驱动的场景选择和安全论证ISO 3450x 场景标准族
├── ISO 34503: Specification of Operational Design Domain
│ └── ODD描述方法和分类框架
├── ISO 34504: Scenario Categorization
│ └── 场景分类方法(基于抽象层次)
└── ISO 34505: Scenario Generation and Selection
└── 场景生成和选择方法skills/china-standards/sotif/ — SOTIF标准集skills/china-standards/odd/ — ODD标准skills/automotive-scenario-driven-testing/ — 场景驱动测试方法skills/automotive-dfm-benchmarking/ — DFM基准评测| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 31,183 | 20,881 | -33% | 1 | 1 | 0% | 6,215 | 4,975 | -20% | 0 | 0 | — |
case-02 | fail→pass | 30,103 | 24,319 | -19% | 1 | 1 | 0% | 5,572 | 5,891 | +6% | 0 | 0 | — |
case-03 | fail→pass | 15,683 | 6,373 | -59% | 1 | 1 | 0% | 944 | 2,207 | +134% | 0 | 0 | — |
case-04 | fail→pass | 9,379 | 6,012 | -36% | 1 | 1 | 0% | 1,938 | 1,996 | +3% | 0 | 0 | — |
case-05 | pass→pass | 14,418 | 4,434 | -69% | 1 | 1 | 0% | 2,624 | 1,816 | -31% | 0 | 0 | — |
case-06 | fail→fail | 17,029 | 13,823 | -19% | 1 | 1 | 0% | 3,293 | 3,462 | +5% | 0 | 0 | — |
case-07 | pass→pass | 10,755 | 10,991 | +2% | 1 | 1 | 0% | 2,081 | 3,004 | +44% | 0 | 0 | — |
case-08 | pass→pass | 7,115 | 6,544 | -8% | 1 | 1 | 0% | 1,257 | 2,194 | +75% | 0 | 0 | — |
case-09 | pass→pass | 4,312 | 3,321 | -23% | 1 | 1 | 0% | 717 | 1,559 | +117% | 0 | 0 | — |
case-10 | pass→pass | 12,739 | 7,175 | -44% | 1 | 1 | 0% | 2,251 | 2,154 | -4% | 0 | 0 | — |
case-11 | pass→pass | 11,844 | 11,508 | -3% | 1 | 1 | 0% | 2,003 | 3,022 | +51% | 0 | 0 | — |
case-12 | pass→pass | 14,132 | 15,236 | +8% | 1 | 1 | 0% | 2,317 | 3,530 | +52% | 0 | 0 | — |
case-13 | pass→pass | 10,220 | 2,943 | -71% | 1 | 1 | 0% | 1,731 | 1,523 | -12% | 0 | 0 | — |
case-14 | pass→pass | 13,238 | 13,818 | +4% | 1 | 1 | 0% | 2,311 | 3,470 | +50% | 0 | 0 | — |
case-15 | fail→pass | 16,714 | 15,449 | -8% | 1 | 1 | 0% | 2,822 | 3,663 | +30% | 0 | 0 | — |
case-16 | fail→pass | 13,487 | 2,311 | -83% | 1 | 1 | 0% | 2,095 | 1,407 | -33% | 0 | 0 | — |
case-17 | pass→pass | 18,929 | 18,967 | +0% | 1 | 1 | 0% | 3,177 | 4,143 | +30% | 0 | 0 | — |
case-18 | fail→fail | 7,935 | 10,855 | +37% | 1 | 1 | 0% | 1,520 | 3,246 | +114% | 0 | 0 | — |
case-19 | pass→pass | 3,721 | 2,813 | -24% | 1 | 1 | 0% | 686 | 1,517 | +121% | 0 | 0 | — |
case-20 | pass→pass | 17,866 | 18,539 | +4% | 1 | 1 | 0% | 2,936 | 4,217 | +44% | 0 | 0 | — |
case-21 | pass→pass | 23,165 | 25,535 | +10% | 1 | 1 | 0% | 3,908 | 5,172 | +32% | 0 | 0 | — |
case-22 | pass→pass | 16,001 | 12,683 | -21% | 1 | 1 | 0% | 3,655 | 4,018 | +10% | 0 | 0 | — |
case-23 | pass→pass | 18,594 | 19,883 | +7% | 1 | 1 | 0% | 3,312 | 4,779 | +44% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 22 counted toward the lift figure. The other 1 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +22 percentage points is the difference between those two pass rates over the 22 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.