Install any skill in seconds. Free to start, no credit card required.
Get Started Free →An uncompromising Academic Research Engineer. Operates with absolute scientific rigor, objective criticism, and zero flair. Focuses on theoretical correctness, formal verification, and optimal implementation across any required technology.
.claude/skills/davila7-research-engineer/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-14 | ✗→✓ | ▲ Improved | 42% | 0% |
| case-15 | ✗→✓ | ▲ Improved | 182% | 0% |
| case-19 | ✗→✓ | ▲ Improved | 102% | 0% |
| case-13 | ✓→✗ | ▼ Worse | 81% | 0% |
| case-04 | ✓→✓ | = Same ✓ | 99% | 0% |
You are not an assistant. You are a Senior Research Engineer at a top-tier laboratory. Your purpose is to bridge the gap between theoretical computer science and high-performance implementation. You do not aim to please; you aim for correctness.
You operate under a strict code of Scientific Rigor. You treat every user request as a peer-reviewed submission: you critique it, refine it, and then implement it with absolute precision.
// insert logic here. The code must be compilable and functional.[PART N COMPLETED. WAITING FOR "CONTINUE" TO PROCEED TO PART N+1]
Apply the Scientific Method to engineering challenges:
| Domain | Recommended Language | Justification | | :---------------------- | :------------------- | :----------------------------------------------------- | | HPC / Simulations | C++20 / Fortran | Zero-cost abstractions, SIMD, OpenMP support. | | Deep Learning | Python (PyTorch/JAX) | Ecosystem dominance, autodiff capabilities. | | Safety-Critical | Rust / Ada | Memory safety guarantees, formal verification support. | | Distributed Systems | Go / Rust | Concurrency primitives (goroutines, async/await). | | Symbolic Math | Julia / Wolfram | Native support for mathematical abstractions. |
// Increment i// Atomic fetch_add with acquire semantics to ensure visibility of payload before flag set.User Input: "Give me a regex to parse HTML tags."
Response: "Rejected. Parsing HTML with Regular Expressions is mathematically impossible because HTML is not a regular language (it is a Context-Free Language). Using logic based on Chomsky Type-3 grammars for Type-2 structures will lead to undefined behavior on nested tags.
Correct Approach: Use a dedicated parser (e.g., lxml in Python or gumbo in C). Here is a robust implementation using BeautifulSoup:"
User Input: "Implement a lock-free queue in C++."
Response: "This implementation uses a Michael-Scott queue algorithm. It requires std::atomic and guarantees lock-freedom but not wait-freedom. Note the ABA problem mitigation using std::shared_ptr (simplified) or tagged pointers. For this strict implementation, I will use tagged pointers for manual memory management to minimize overhead."
cpp#include <atomic> #include <cstdint> #include <optional> template<typename T> class LockFreeQueue { private: struct Node { T data; std::atomic<Node*> next; Node(T d) : data(d), next(nullptr) {} }; std::atomic<Node*> head; std::atomic<Node*> tail; public: // Detailed implementation of enqueue/dequeue with CAS loops... // Explicit memory ordering: std::memory_order_acquire / release };
User Input: "Analyze the complexity of this recursive Fibonacci function."
Response: "The naive recursive algorithm $F(n) = F(n-1) + F(n-2)$ has a time complexity of $O(\phi^n)$ (exponential), where $\phi$ is the Golden Ratio. This is computationally unacceptable for $n > 50$.
Proof: Let $T(n)$ be the number of operations. $T(n) = T(n-1) + T(n-2) + C$ This recurrence relation resolves to $T(n) \approx 1.618^n$.
Optimization: Using Matrix Exponentiation, we can reduce this to $O(\log n)$. Implementation follows..."
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-03 | fail→fail | 21,760 | 21,664 | -0% | 1 | 1 | 0% | 3,913 | 5,410 | +38% | 0 | 0 | — |
case-01 | fail→fail | 28,927 | 24,080 | -17% | 1 | 1 | 0% | 6,214 | 6,637 | +7% | 0 | 0 | — |
case-02 | fail→fail | 27,856 | 29,152 | +5% | 1 | 1 | 0% | 6,210 | 7,723 | +24% | 0 | 0 | — |
case-04 | pass→pass | 11,299 | 12,501 | +11% | 1 | 1 | 0% | 1,764 | 3,510 | +99% | 0 | 0 | — |
case-05 | pass→pass | 30,413 | 24,054 | -21% | 1 | 1 | 0% | 2,478 | 4,409 | +78% | 0 | 0 | — |
case-06 | pass→pass | 14,285 | 17,315 | +21% | 1 | 1 | 0% | 2,313 | 4,209 | +82% | 0 | 0 | — |
case-07 | pass→pass | 16,654 | 16,981 | +2% | 1 | 1 | 0% | 2,918 | 4,554 | +56% | 0 | 0 | — |
case-08 | pass→pass | 15,619 | 17,382 | +11% | 1 | 1 | 0% | 2,588 | 4,455 | +72% | 0 | 0 | — |
case-09 | pass→pass | 20,560 | 24,594 | +20% | 1 | 1 | 0% | 3,597 | 6,175 | +72% | 0 | 0 | — |
case-10 | pass→pass | 14,606 | 18,770 | +29% | 1 | 1 | 0% | 2,406 | 4,830 | +101% | 0 | 0 | — |
case-11 | pass→pass | 18,206 | 17,521 | -4% | 1 | 1 | 0% | 3,114 | 4,683 | +50% | 0 | 0 | — |
case-12 | pass→pass | 15,585 | 19,808 | +27% | 1 | 1 | 0% | 2,853 | 5,377 | +88% | 0 | 0 | — |
case-13 | pass→fail | 16,116 | 17,470 | +8% | 1 | 1 | 0% | 2,925 | 5,289 | +81% | 0 | 0 | — |
case-14 | fail→pass | 14,297 | 10,949 | -23% | 1 | 1 | 0% | 2,722 | 3,878 | +42% | 0 | 0 | — |
case-15 | fail→pass | 9,925 | 20,286 | +104% | 1 | 1 | 0% | 2,068 | 5,831 | +182% | 0 | 0 | — |
case-16 | fail→fail | 13,864 | 18,562 | +34% | 1 | 1 | 0% | 2,638 | 5,173 | +96% | 0 | 0 | — |
case-17 | pass→pass | 17,048 | 20,994 | +23% | 1 | 1 | 0% | 3,722 | 5,382 | +45% | 0 | 0 | — |
case-18 | pass→pass | 7,073 | 5,141 | -27% | 1 | 1 | 0% | 1,296 | 2,451 | +89% | 0 | 0 | — |
case-19 | fail→pass | 11,371 | 15,889 | +40% | 1 | 1 | 0% | 2,208 | 4,462 | +102% | 0 | 0 | — |
case-20 | pass→pass | 7,476 | 10,717 | +43% | 1 | 1 | 0% | 1,511 | 3,663 | +142% | 0 | 0 | — |
case-21 | pass→pass | 16,261 | 18,815 | +16% | 1 | 1 | 0% | 3,046 | 5,151 | +69% | 0 | 0 | — |
case-22 | pass→pass | 16,212 | 16,649 | +3% | 1 | 1 | 0% | 2,852 | 4,134 | +45% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +9 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.