Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use when you explicitly want a second local or SSH remote Codex CLI session to prototype, debug, or review code, while your current session remains the primary owner of the final result.
.claude/skills/cnfjlhj-collaborating-with-codex/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 383% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 163% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 119% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 12% | 0% |
| case-11 | ✗→✓ | ▲ Improved | 43% | 0% |
bashpython scripts/codex_bridge.py --cd "/path/to/project" --PROMPT "Your task"
Remote Codex over SSH:
bashpython scripts/codex_bridge.py --ssh "server-alias" --cd "/remote/project" --PROMPT "Your task"
Output: JSON with success, SESSION_ID, agent_messages, and optional error.
usage: codex_bridge.py [-h] --PROMPT PROMPT --cd CD [--sandbox {read-only,workspace-write,danger-full-access}] [--SESSION_ID SESSION_ID] [--skip-git-repo-check]
[--return-all-messages] [--image IMAGE] [--model MODEL] [--yolo] [--profile PROFILE] [--ssh SSH] [--ssh-option SSH_OPTION]
[--codex-bin CODEX_BIN]
Codex Bridge
options:
-h, --help show this help message and exit
--PROMPT PROMPT Instruction for the task to send to codex.
--cd CD Set the workspace root for codex before executing the task.
--sandbox {read-only,workspace-write,danger-full-access}
Sandbox policy for model-generated commands. Defaults to `read-only`.
--SESSION_ID SESSION_ID
Resume the specified session of the codex. Defaults to `None`, start a new session.
--skip-git-repo-check
Allow codex running outside a Git repository (useful for one-off directories).
--return-all-messages
Return all messages (e.g. reasoning, tool calls, etc.) from the codex session. Set to `False` by default, only the agent's final reply message is
returned.
--image IMAGE Attach one or more image files to the initial prompt. Separate multiple paths with commas or repeat the flag.
--model MODEL The model to use for the codex session. This parameter is strictly prohibited unless explicitly specified by the user.
--yolo Run every command without approvals or sandboxing. Only use when `sandbox` couldn't be applied.
--profile PROFILE Configuration profile name to load from `~/.codex/config.toml`. This parameter is strictly prohibited unless explicitly specified by the user.
--ssh SSH Run Codex on a remote host via SSH. Value can be an SSH alias or user@host. When set, --cd and --image paths are remote paths.
--ssh-option SSH_OPTION
Extra ssh option, repeatable. Example: --ssh-option=-J --ssh-option=bastion
--codex-bin CODEX_BIN
Codex executable to run locally or on the remote host. Defaults to `codex`. `--remote-codex` is accepted as a backward-compatible alias.Always capture SESSION_ID from the first response for follow-up:
bash# Initial task python scripts/codex_bridge.py --cd "/project" --PROMPT "Analyze auth in login.py" # Continue with SESSION_ID python scripts/codex_bridge.py --cd "/project" --SESSION_ID "uuid-from-response" --PROMPT "Write unit tests for that"
Remote sessions use the same SESSION_ID, but continue them against the same SSH host:
bashpython scripts/codex_bridge.py --ssh "server-alias" --cd "/remote/project" --SESSION_ID "uuid-from-response" --PROMPT "Now implement it"
Use --ssh when the repository, runtime, or credentials needed for the task live on a remote server.
Requirements:
ssh <host> non-interactively.codex installed and available in the remote non-interactive shell PATH, or pass --codex-bin "/absolute/path/to/codex".--cd is a remote absolute path when --ssh is set.--image paths are also remote paths when --ssh is set. Upload local images first if Codex on the server needs them.Examples:
bash# Use an SSH config alias python scripts/codex_bridge.py --ssh "gpu-box" --cd "/srv/app" --PROMPT "Review this service for race conditions" # Use user@host and a remote codex path python scripts/codex_bridge.py --ssh "ubuntu@example.com" --codex-bin "/home/ubuntu/.local/bin/codex" --cd "/home/ubuntu/app" --PROMPT "Run the failing tests and diagnose" # Use a jump host python scripts/codex_bridge.py --ssh "private-box" --ssh-option=-J --ssh-option=bastion --cd "/workspace/repo" --PROMPT "Inspect the deployment scripts"
Do not use --ssh for local repositories. Do not pass --model or --profile unless the user explicitly requested them.
Prototyping (read-only, request diffs):
bashpython scripts/codex_bridge.py --cd "/project" --PROMPT "Generate unified diff to add logging"
Debug with full trace:
bashpython scripts/codex_bridge.py --cd "/project" --PROMPT "Debug this error" --return-all-messages
Remote debug with write access:
bashpython scripts/codex_bridge.py --ssh "server-alias" --cd "/remote/project" --sandbox workspace-write --PROMPT "Fix the failing integration test and summarize the patch"
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-01 | fail→pass | 44,237 | 7,904 | -82% | 1 | 1 | 0% | 361 | 1,745 | +383% | 0 | 0 | — |
case-02 | fail→fail | 6,607 | 73,708 | +1016% | 1 | 1 | 0% | 813 | 1,943 | +139% | 0 | 0 | — |
case-03 | fail→fail | 7,421 | 14,150 | +91% | 1 | 1 | 0% | 835 | 1,823 | +118% | 0 | 0 | — |
case-04 | fail→pass | 16,962 | 29,209 | +72% | 1 | 1 | 0% | 1,110 | 2,915 | +163% | 0 | 0 | — |
case-05 | fail→fail | 18,103 | 18,981 | +5% | 1 | 1 | 0% | 2,038 | 2,602 | +28% | 0 | 0 | — |
case-06 | fail→fail | 36,298 | 23,344 | -36% | 1 | 1 | 0% | 1,478 | 1,902 | +29% | 0 | 0 | — |
case-07 | fail→pass | 26,980 | 19,002 | -30% | 1 | 1 | 0% | 793 | 1,738 | +119% | 0 | 0 | — |
case-08 | fail→pass | 9,655 | 22,132 | +129% | 1 | 1 | 0% | 1,557 | 1,743 | +12% | 0 | 0 | — |
case-09 | pass→fail | 20,990 | 27,981 | +33% | 1 | 1 | 0% | 1,963 | 2,096 | +7% | 0 | 0 | — |
case-10 | pass→pass | 26,498 | 5,789 | -78% | 1 | 1 | 0% | 1,175 | 1,804 | +54% | 0 | 0 | — |
case-11 | fail→pass | 56,544 | 15,587 | -72% | 1 | 1 | 0% | 1,875 | 2,680 | +43% | 0 | 0 | — |
case-12 | fail→fail | 19,128 | 18,803 | -2% | 1 | 1 | 0% | 853 | 2,707 | +217% | 0 | 0 | — |
case-13 | fail→fail | 35,515 | 27,908 | -21% | 1 | 1 | 0% | 570 | 2,101 | +269% | 0 | 0 | — |
case-14 | fail→pass | 15,564 | 3,793 | -76% | 1 | 1 | 0% | 1,218 | 1,853 | +52% | 0 | 0 | — |
case-15 | fail→pass | 38,191 | 17,014 | -55% | 1 | 1 | 0% | 1,155 | 2,624 | +127% | 0 | 0 | — |
case-16 | fail→fail | 9,957 | 12,599 | +27% | 1 | 1 | 0% | 770 | 1,869 | +143% | 0 | 0 | — |
case-17 | pass→fail | 15,094 | 31,612 | +109% | 1 | 1 | 0% | 746 | 1,778 | +138% | 0 | 0 | — |
case-18 | fail→fail | 18,079 | 14,243 | -21% | 1 | 1 | 0% | 538 | 1,814 | +237% | 0 | 0 | — |
case-19 | fail→pass | 10,402 | 9,632 | -7% | 1 | 1 | 0% | 749 | 2,057 | +175% | 0 | 0 | — |
case-20 | fail→fail | 16,192 | 27,529 | +70% | 1 | 1 | 0% | 304 | 2,359 | +676% | 0 | 0 | — |
case-21 | pass→fail | 12,942 | 25,657 | +98% | 1 | 1 | 0% | 1,371 | 2,926 | +113% | 0 | 0 | — |
case-22 | pass→fail | 11,427 | 30,753 | +169% | 1 | 1 | 0% | 638 | 2,015 | +216% | 0 | 0 | — |
case-23 | fail→pass | 12,111 | 9,920 | -18% | 1 | 1 | 0% | 907 | 1,956 | +116% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 23 cases were attempted, and 10 counted toward the lift figure. The other 13 produced results that are not comparable between the two arms, so they are excluded from the headline rather than averaged into it. The headline lift of +22 percentage points is the difference between those two pass rates over the 10 comparable cases. 5 cases got worse with the skill loaded, and they are included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
| Model | Method | Date | Lift |
|---|---|---|---|
| gemini-3.6-flash | verified | 8/21/2026 | +74% |
Other measured skills in the registry, with their headline benchmark lift.