Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Create and edit file-based Warp software factory definitions, in a repository tree rooted at a factory.yaml. Use when authoring or changing that factory.yaml, Agent, Automation, Scorer, Runner, or Webhook files under that root, or its factory and agent skill trees, and when fixing Factory file diagnostics. Do not use for agent-definition Markdown that belongs to another tool, for a tree with no factory.yaml, or to operate a live factory or hand work to one through Factory MCP.
.claude/skills/warpdotdev-factory-files/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-04 | ✗→✓ | ▲ Improved | 233% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 141% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 30% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 143% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 82% | 0% |
A software factory can be defined by files in a repository. This skill covers authoring and editing those files, and validating them before you open a pull request.
warp-server owns the format. It publishes the schema for each version it supports and validates a tree with the same parser the apply path uses. This skill carries no copy of the format: a copy ships inside a Warp release, goes stale against the server, and then reports confident, wrong diagnostics. When the server cannot be reached, the answer is that the tree was not checked.
Use this skill for repository files. It is not the skill for operating a live factory: use factory-mcp to send work to a factory, inspect task status, or pull a task down locally. Playbooks under a factory's own skills/ directories tell that factory's agents how to do their job; editing one is a prompt change, not a schema change, so this skill's rules do not apply to their contents.
Every Factory tree is rooted at the directory containing factory.yaml. All paths below are relative to that root. A repository may register a subdirectory as the root, so find factory.yaml rather than assuming the repository root. Do not follow symlinks while looking: the server parses the repository tree, where a symlink is stored as its target path rather than its target's content.
If there is no factory.yaml, this is not a Factory tree and nothing here applies. agents/<name>/agent.md and similar paths are also used by other agent tooling; stop and say so rather than imposing this schema on them.
factory.yaml required, exactly one
agents/<name>/agent.md at least one; exactly one must be MAIN
agents/<name>/skills/** skills only that agent can use
automations/<name>/automation.md optional
runners/<name>.yaml optional
scorers/<name>/scorer.md optional; Markdown body is the rubric
benchmarks/<name>/suite.yaml optional; benchmark suite manifest
benchmarks/<name>/tasks/<name>.yaml optional; benchmark suite task
webhooks/<name>.yaml optional; custom webhook sources
skills/** skills every agent in the factory can useResource names come from the path, never from a field inside the file. Renaming an agent means moving its directory.
automations/<name>.md is a legacy flat form the parser still accepts. Create the directory form; when editing an existing flat file, leave it where it is unless the user asks you to normalize the tree.
factory.yaml, so you can seewhat is inherited and what is overridden.
after an Agent's or Automation's closing --- fence is its prompt; a Scorer's body is its rubric. Never fold either into frontmatter.
Read the tree's schemaVersion from factory.yaml; a tree that omits it is v1alpha1. Then fetch the schema for that version:
bashcurl -s https://app.warp.dev/api/v1/factory-files/schemas curl -s https://app.warp.dev/api/v1/factory-files/schemas/<schemaVersion>
The registry lists the versions the server supports. The version endpoint returns every document describing one version, keyed by file name: factory.schema.json for factory.yaml, agent.schema.json, automation.schema.json, runner.schema.json, scorer.schema.json and benchmark_suite.schema.json, benchmark_suite_task.schema.json, and webhook.schema.json for the corresponding resources, and common.schema.json for the definitions they share. Both endpoints are unauthenticated. They are exact for the version they describe: an unknown field is an error, and each enumerated value is one the server accepts today.
If the server does not publish the declared version, stop. Do not measure the tree against a version it does not claim to be, and never lower schemaVersion to make a check pass.
Read references/examples.md for worked examples of each resource, and references/scorers.md before writing or changing a Scorer. The field-by-field catalogue is not duplicated here any more; the fetched schema carries it, with a description on each field.
Run the bundled validator with Python 3.8 or newer, using the host's command (python3, python, or py -3). Quote both paths because an app-bundle path can contain spaces.
bashpython3 "{{skill_dir}}/scripts/validate_factory_files.py" "<factory-root>"
It selects the tree's resource files and submits them to the server, which runs the real parser. Add --json for machine-readable output and --server-root <url>, or WARP_SERVER_ROOT (or WARP_SERVER_ROOT_URL, the name an Oz sandbox exports), to point at a local, staging, or self-hosted server. No credential is required; WARP_API_KEY is forwarded when the environment already carries one, as an agent sandbox does, and dropped for a single retry if the server answers it with 401 or 403.
The exit code distinguishes three outcomes, and so must you:
0 the server checked the tree and found no problem.1 the server checked the tree and reported diagnostics. Fix every one andre-run until it is clean.
2 the tree was not checked. This is not a pass and not a failure; itsays nothing about the files at all.
On exit 2, say plainly that validation did not run and why. Do not describe the files as valid, correct, or ready, and do not substitute your own reading of the schema for a verdict. If you cannot reach a server and the change matters, say so and let the user decide.
On exit 0, repeat the sentence the validator prints rather than paraphrasing it into something stronger. A pass means the parser and the state-independent checks agreed; it does not mean the tree will apply.
Validation resolves no server state. Model IDs, environment IDs, secret names, runner names, Scorer model IDs, MCP server IDs, integration availability, and the values of Linear and Slack name aliases are all checked when the plan is applied. The response lists what it did not check, including any deferred name aliases; report that distinction rather than claiming a tree is fully verified.
If no Python 3 interpreter is available, do not install one or claim the tree was validated without the user's approval. Check the changed document against the fetched schema by hand and report that automated validation was unavailable.
When the Factory is already registered, a server plan remains the strongest available check. See references/validation.md for diagnostic codes and how to read them.
agentType: MAIN (or FOREMAN, its canonicalspelling). Zero or two is an error.
model and harness are mutually exclusive everywhere. model: <id> isshorthand for the Oz harness.
agentDefaults must declare one of them; agents and automations may declareneither and inherit.
secrets or mcpServers at agent or automation level replaces theinherited value; it does not merge.
providerand event.
schedule.cron_fired trigger needs either an inline schedule.cron or anon-empty filter.schedule_ids, and never both.
platform.linux.dockerImage. A runner with noplatform section defaults to Linux and will fail for that reason.
(provider, event) pair. Some fields havea friendlier authoring spelling that the server rewrites for you: GitHub baseBranches and prNumbers, Linear teams, projects, states and issues, and Slack channels, users and itemUsers. Each stands in for its canonical key, and declaring both is an error. The Linear and Slack ones name objects the server looks up at apply time, so they take a plain list of names rather than an in/not_in matcher.
secretName is required in everyauthMode and names an existing managed secret; the server does not generate a secret for a file-declared source, because nothing could then read it back to configure the sender.
signatureScheme is required when authMode: signature and rejectedotherwise.
authMode and signatureScheme cannot be changed on an existing source.Changing either means deleting the file and adding a new one under a different name.
filter.webhook_ids takes source UIDs the serverassigned, not the file names declared here.
It is tempting to bundle the schema, or to reimplement a few checks here so authoring works offline. Both have been tried and removed. A copy inside a Warp release is routinely older than the server it is used against, and a stale copy does not fail quietly: it reports a valid field as unknown, and an agent trying to get to a clean run deletes working configuration to satisfy it. That has already happened once, to Linear and Slack trigger aliases the server accepts.
Reporting that a tree was not checked costs a little. Reporting the wrong answer costs correct configuration. Fetch the format when you need it; say nothing when you cannot.
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→fail | 6,347 | 8,710 | +37% | 1 | 1 | 0% | 202 | 2,736 | +1254% | 0 | 0 | — |
case-02 | fail→fail | 4,504 | 5,525 | +23% | 1 | 1 | 0% | 207 | 2,457 | +1087% | 0 | 0 | — |
case-03 | fail→fail | 6,241 | 7,096 | +14% | 1 | 1 | 0% | 264 | 2,596 | +883% | 0 | 0 | — |
case-04 | fail→pass | 7,764 | 9,030 | +16% | 1 | 1 | 0% | 1,053 | 3,504 | +233% | 0 | 0 | — |
case-05 | pass→pass | 13,119 | 4,549 | -65% | 1 | 1 | 0% | 1,876 | 2,968 | +58% | 0 | 0 | — |
case-06 | fail→pass | 7,230 | 11,572 | +60% | 1 | 1 | 0% | 1,176 | 2,830 | +141% | 0 | 0 | — |
case-07 | pass→pass | 9,937 | 6,349 | -36% | 1 | 1 | 0% | 1,480 | 3,160 | +114% | 0 | 0 | — |
case-08 | fail→pass | 16,336 | 8,742 | -46% | 1 | 1 | 0% | 2,387 | 3,112 | +30% | 0 | 0 | — |
case-09 | fail→pass | 8,637 | 4,535 | -47% | 1 | 1 | 0% | 1,226 | 2,976 | +143% | 0 | 0 | — |
case-10 | pass→pass | 7,345 | 3,823 | -48% | 1 | 1 | 0% | 1,133 | 2,884 | +155% | 0 | 0 | — |
case-11 | fail→pass | 10,109 | 3,584 | -65% | 1 | 1 | 0% | 1,526 | 2,782 | +82% | 0 | 0 | — |
case-12 | fail→pass | 11,419 | 3,691 | -68% | 1 | 1 | 0% | 1,756 | 2,789 | +59% | 0 | 0 | — |
case-13 | fail→pass | 11,150 | 4,296 | -61% | 1 | 1 | 0% | 1,607 | 3,069 | +91% | 0 | 0 | — |
case-14 | fail→pass | 11,159 | 3,618 | -68% | 1 | 1 | 0% | 1,600 | 2,814 | +76% | 0 | 0 | — |
case-15 | pass→pass | 12,410 | 3,802 | -69% | 1 | 1 | 0% | 1,981 | 2,843 | +44% | 0 | 0 | — |
case-16 | fail→fail | 16,399 | 11,272 | -31% | 1 | 1 | 0% | 2,305 | 4,151 | +80% | 0 | 0 | — |
case-17 | fail→pass | 13,491 | 2,229 | -83% | 1 | 1 | 0% | 1,802 | 2,524 | +40% | 0 | 0 | — |
case-18 | fail→fail | 10,086 | 4,068 | -60% | 1 | 1 | 0% | 1,437 | 2,714 | +89% | 0 | 0 | — |
case-19 | pass→pass | 10,529 | 4,868 | -54% | 1 | 1 | 0% | 1,610 | 2,977 | +85% | 0 | 0 | — |
case-20 | fail→pass | 13,605 | 5,807 | -57% | 1 | 1 | 0% | 1,623 | 3,156 | +94% | 0 | 0 | — |
case-21 | pass→pass | 13,153 | 5,007 | -62% | 1 | 1 | 0% | 1,853 | 3,025 | +63% | 0 | 0 | — |
case-22 | fail→pass | 18,797 | 4,446 | -76% | 1 | 1 | 0% | 2,622 | 2,957 | +13% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted, and 19 counted toward the lift figure. The other 3 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +50 percentage points is the difference between those two pass rates over the 19 comparable cases.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.