Install any skill in seconds. Free to start, no credit card required.
Get Started Free →This skill provides an advanced financial modeling suite with DCF analysis, sensitivity testing, Monte Carlo simulations, and scenario planning for investment decisions
.claude/skills/bilal140202-creating-financial-models/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-08 | ✗→✓ | ▲ Improved | 31% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 48% | 0% |
| case-17 | ✗→✓ | ▲ Improved | 55% | 0% |
| case-02 | ✓→✗ | ▼ Worse | 37% | 0% |
| case-06 | ✓→✗ | ▼ Worse | 18% | 0% |
A comprehensive financial modeling toolkit for investment analysis, valuation, and risk assessment using industry-standard methodologies.
"Build a DCF model for this technology company using the attached financials"
"Run a Monte Carlo simulation on this acquisition model with 5,000 iterations"
"Create sensitivity analysis showing impact of growth rate and WACC on valuation"
"Develop three scenarios for this expansion project with probability weights"
dcf_model.py: Complete DCF valuation enginesensitivity_analysis.py: Sensitivity testing frameworkThe model automatically performs:
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | pass→pass | 10,020 | 25,710 | +157% | 1 | 1 | 0% | 2,006 | 7,182 | +258% | 0 | 0 | — |
case-02 | pass→fail | 13,258 | 11,284 | -15% | 1 | 1 | 0% | 2,107 | 2,878 | +37% | 0 | 0 | — |
case-03 | pass→pass | 13,329 | 13,733 | +3% | 1 | 1 | 0% | 2,711 | 3,638 | +34% | 0 | 0 | — |
case-04 | pass→pass | 11,786 | 9,526 | -19% | 1 | 1 | 0% | 2,025 | 2,871 | +42% | 0 | 0 | — |
case-05 | pass→pass | 10,218 | 8,685 | -15% | 1 | 1 | 0% | 1,659 | 2,520 | +52% | 0 | 0 | — |
case-06 | pass→fail | 13,421 | 9,525 | -29% | 1 | 1 | 0% | 2,116 | 2,502 | +18% | 0 | 0 | — |
case-07 | pass→fail | 13,386 | 10,626 | -21% | 1 | 1 | 0% | 2,139 | 2,684 | +25% | 0 | 0 | — |
case-08 | fail→pass | 20,401 | 17,173 | -16% | 1 | 1 | 0% | 2,819 | 3,691 | +31% | 0 | 0 | — |
case-09 | pass→pass | 13,483 | 11,478 | -15% | 1 | 1 | 0% | 2,261 | 2,925 | +29% | 0 | 0 | — |
case-10 | pass→pass | 11,033 | 14,606 | +32% | 1 | 1 | 0% | 1,954 | 3,498 | +79% | 0 | 0 | — |
case-11 | pass→pass | 13,617 | 7,170 | -47% | 1 | 1 | 0% | 2,228 | 2,162 | -3% | 0 | 0 | — |
case-12 | pass→pass | 15,630 | 16,097 | +3% | 1 | 1 | 0% | 2,600 | 3,555 | +37% | 0 | 0 | — |
case-13 | pass→pass | 5,209 | 6,025 | +16% | 1 | 1 | 0% | 1,026 | 2,173 | +112% | 0 | 0 | — |
case-14 | pass→pass | 11,888 | 6,960 | -41% | 1 | 1 | 0% | 2,173 | 2,320 | +7% | 0 | 0 | — |
case-15 | fail→pass | 9,619 | 7,575 | -21% | 1 | 1 | 0% | 1,623 | 2,401 | +48% | 0 | 0 | — |
case-16 | fail→fail | 15,979 | 17,033 | +7% | 1 | 1 | 0% | 2,875 | 4,090 | +42% | 0 | 0 | — |
case-17 | fail→pass | 8,170 | 6,981 | -15% | 1 | 1 | 0% | 1,372 | 2,126 | +55% | 0 | 0 | — |
case-18 | fail→fail | 11,055 | 6,943 | -37% | 1 | 1 | 0% | 1,711 | 2,156 | +26% | 0 | 0 | — |
case-19 | pass→pass | 13,895 | 11,856 | -15% | 1 | 1 | 0% | 2,332 | 3,232 | +39% | 0 | 0 | — |
case-20 | pass→pass | 12,708 | 12,441 | -2% | 1 | 1 | 0% | 2,165 | 3,047 | +41% | 0 | 0 | — |
case-21 | pass→pass | 12,529 | 11,331 | -10% | 1 | 1 | 0% | 2,162 | 3,105 | +44% | 0 | 0 | — |
case-22 | pass→pass | 14,450 | 13,263 | -8% | 1 | 1 | 0% | 2,095 | 3,176 | +52% | 0 | 0 | — |
case-23 | pass→pass | 10,720 | 10,894 | +2% | 1 | 1 | 0% | 1,596 | 2,673 | +67% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted. The headline lift of 0 percentage points is the difference between those two pass rates over the 23 comparable cases. 3 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.