Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Set up CIAgent regression testing for the AI agent in this repo — write a runner, record golden baselines, generate a test spec, and verify it. Use when the user asks to add tests, evals, or regression testing for their AI agent, or to set up CIAgent.
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-07 | ✗→✓ | ▲ Improved | -16% | 0% |
| case-08 | ✗→✓ | ▲ Improved | -17% | 0% |
| case-09 | ✗→✓ | ▲ Improved | -21% | 0% |
| case-10 | ✗→✓ | ▲ Improved | 26% | 0% |
| case-12 | ✗→✓ | ▲ Improved | -1% | 0% |
You are setting up CIAgent (pip install ciagent) so this repo's AI agent has recorded golden baselines and a runnable regression suite. The end state: the user can run ciagent test --runs 3 and see a stability report for their agent.
Work through the steps in order. Do not skip the cost gate in step 4.
openai, anthropic, langgraph,langchain) and for the function or endpoint that takes a user message and returns the agent's answer.
pip install "ciagent[openai]", [anthropic], [langgraph], or [all].
ciagent --version then ciagent doctor (it reports what ismissing; a missing spec is expected at this point).
Create agentci_runner.py at the repo root (or inside the package if the repo has one clear package):
pythondef run_for_agentci(query: str) -> str: """CIAgent entry point: one query in, final answer text out.""" # import the user's agent and invoke it ONCE, no chat history ... return final_answer_text
Rules:
capture, so LLM calls and tool calls are recorded automatically — do not build Trace objects unless the repo already produces them.
python -c "from agentci_runner import run_for_agentci; print(run_for_agentci('hello'))".
Write agentci_queries.txt, one query per line — 8 to 15 queries:
and existing tests for what it is supposed to handle).
limits, names) — those become deterministic checks in step 6.
Recording baselines runs the real agent once per query, on the user's API keys. State the query count and a cost ballpark, and ask the user to confirm before step 5. If there are no API keys or the user declines: write agentci_spec.yaml by hand instead (same queries, runner: set), validate with ciagent test --mock, and tell the user which step to resume later.
bashciagent bootstrap --runner agentci_runner:run_for_agentci \ --queries agentci_queries.txt --agent <agent-name> --yes
This runs every query, saves each trace as a golden baseline under ./baselines/<agent-name>/, and writes agentci_spec.yaml with path and cost budgets derived from the recorded traces. Read the printed answers as they stream by — if an answer is visibly wrong, that query should not be golden: fix the agent or the query, delete that baseline file, and rerun.
The generated spec has path and cost budgets but no correctness checks. Add a correctness: block per query, derived from the recorded baseline answers and the repo's docs/KB — never from what you wish the agent said:
yamlcorrectness: expected_in_answer: ["30 days"] # hard facts, AND any_expected_in_answer: ["$9.95", "9.95"] # phrasing variants, OR not_in_answer: ["I don't know"] # forbidden content
Check facts, not phrasing. If the repo has a knowledge-base directory, run ciagent generate-checks --kb <dir> --dry-run and review its candidates — every surviving candidate was already validated against the recorded goldens.
bashciagent test --mock # structure check, zero API calls ciagent test --yes --format json # live run (covered by step 4 approval) ciagent test --runs 3 --yes # stability report
Exit codes: 0 = pass (flaky-but-passing is 0), 1 = correctness failure in every run, 2 = infra/config error. In the stability report, flips labeled agent-variance mean the agent's answer changed (an agent problem); flips labeled judge-flake mean the eval itself is unstable (a check/judge problem).
If a check fails, fix the agent or fix a factually wrong check. Do not loosen a correct check to make the run green — report the failure to the user instead.
ciagent init scaffolds a GitHub Actions workflow (add --hook for apre-push hook if the user wants it).
agentci_queries.txt, agentci_spec.yaml, baselines/,and the workflow.
flaky (with its flip source), and that ciagent test --runs 3 is the command to watch after future agent changes.
Other measured skills in the registry, with their headline benchmark lift.