Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use this skill when the user is invoking doca_flow_dpa_perf on DPA-capable hardware (ConnectX-7 minimum supported, ConnectX-8 recommended, or BlueField-3) to measure rule update / disable rates on the DPA-offloaded DOCA Flow path — picking the active / passive device split, choosing workload-shape axes (burst, queue, completion threshold, workers, hash pipe algo, PSL tables), or reading Kops/sec iteration stats and the optional self-test. Trigger even when the user does not explicitly mention "d
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 115% | 0% |
| case-02 | ✗→✓ | ▲ Improved | 117% | 0% |
| case-03 | ✗→✓ | ▲ Improved | 17% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 36% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 71% | 0% |
doca_flow_dpa_perf)Where to start: This is a tool skill for invoking doca_flow_dpa_perf, the DPA-accelerated Flow performance tool. Open TASKS.md and start at ## configure to confirm DPA-capable hardware + VNF Flow mode + the active / passive device split, then ## run for the smoke-before-bulk flow with a small operation count before any sweep, then ## test for the eval-loop overlay that gates defensible Kops/sec numbers. Open CAPABILITIES.md when the question is what `doca_flow_dpa_perf` can measure, what the DPA preconditions are, which devices it runs on, or how to interpret update / disable / self-test output without fooling yourself. If DOCA is not installed yet, route to doca-setup first; if the device is not DPA-capable (no ConnectX-7+ or BlueField-3+) then this tool is the wrong surface and the right answer is doca-flow-perf.
The CLASSES of doca_flow_dpa_perf questions this skill is built to answer, each with one worked example. The class is the load-bearing piece; the worked example is one instance.
host / DPU-CPU Flow path for this question?" — worked example: "my workload programs path-selector entries via DOCA Flow; do I baseline with `doca_flow_dpa_perf` or with `doca_flow_perf`?". Answered by the DPA-vs-host boundary in CAPABILITIES.md ## Capabilities and modes and the device-preconditions table.
doesn't it change?" — worked example: "if I move my Flow rule update path to the DPA, what changes in the data plane for the packets themselves?". Answered by the DPA-Provider scope in CAPABILITIES.md ## Capabilities and modes.
worked example: "is my BlueField-2 DPA-capable?". Answered by the device-preconditions table in CAPABILITIES.md ## Capabilities and modes (BlueField-3 yes, BlueField-2 no; ConnectX-7 minimum supported, ConnectX-8 recommended, and later generations supported per the public guide and the shipped README on the user's install).
number of operations, iterations — to get a defensible Kops/sec number?" — worked example: "I want the median iteration time and standard deviation, not a single noisy first-iteration spike". Answered by the eval-loop overlay in TASKS.md ## test and the iteration-stats rule in CAPABILITIES.md ## Observability.
what does that mean?" — worked example: "the tool runs but the self-test step fails". Answered by the layered error taxonomy in CAPABILITIES.md ## Error taxonomy + the debug ladder in TASKS.md ## debug.
Flow-perf number for the same workload, in a way the next engineer can actually compare?" — worked example: "two Kops/sec numbers for what is supposedly the same workload". Answered by the four-tuple capture rule in CAPABILITIES.md ## Safety policy + the per-tool-name rule (the host tool and the DPA tool are different surfaces; their numbers are not interchangeable without naming which tool produced which).
This skill serves external operators, performance engineers, DOCA Flow application developers, and AI agents who need a defensible measurement of the DPA-offloaded Flow update path on DPA-capable hardware. Concretely:
workload onto the DPA versus keeping it on the host / DPU-CPU path, and wanting a number to compare.
operation, queue-size X, burst-size Y, N workers" baseline on a specific device + DOCA version so a downstream comparison is meaningful.
doca-dpa to land a DPA-offload of their Flow rule update path and wants to characterize what the device delivers.
the DPA-offloaded Flow path on device Y?" honestly — with a measured number, the command line that produced it, and the device + DOCA version + as-deployed environment that scopes it — instead of guessing from datasheet headlines.
It is not for users debugging the tool's source code, not a substitute for the live public DOCA Flow DPA Perf guide on docs.nvidia.com, not the place to learn the doca-flow or doca-dpa APIs (that audience belongs in doca-flow and doca-dpa), and not the right tool for the host / DPU-CPU Flow path (route to doca-flow-perf).
doca_flow_dpa_perf is shipped as a single CLI binary with DPA-side device code linked in. The skill uses the same kind: tool three-file shape as the rest of the bundle so the agent's task-verb contract is uniform across the bundle.
This skill governs invocation, output interpretation, and recommendation-of-routing for the doca_flow_dpa_perf CLI on DPA-capable hardware. The tool itself has both a host-side control (C-language ARGP + DOCA + DPDK code per the shipped flow_dpa_perf.c / flow_dpa_perf_core.c) and a DPA-side device component (DPA-side code on the shipped DPA device runtime). External users do not link any of this; what they configure is the JSON-config-or-CLI invocation surface. For the doca-dpa programming model behind the DPA-side execution engine, see doca-dpa; for the doca-flow API behind the pipeline the DPA path executes, see doca-flow.
Load this skill when the user is — or the agent needs to — invoke doca_flow_dpa_perf on a real host with DOCA installed and a DPA-capable device attached (or the public NGC DOCA container with the equivalent device passthrough) to measure update / disable rates on the DPA-offloaded Flow path. Concretely:
VNF Flow mode, recommended PF use, no SFs) before invoking the tool.
user's hardware (two-port BlueField-3 active + passive; one- port ConnectX-9 active only).
completion threshold, hash pipe algorithm, work policy, number of PSL tables, table size, number of workers).
shipped README's documented operations.
(median, max, standard deviation) captured.
the layered error taxonomy.
Do not load this skill for general DOCA orientation, Flow program API work, or installation. For those, use doca-public-knowledge-map, the matching libs/<library> skill, or doca-setup. Do not load it for the host / DPU-CPU Flow path — that audience belongs in doca-flow-perf.
This is a thin loader. Substantive material lives in two companion files:
CAPABILITIES.md — what doca_flow_dpa_perf measures(the DPA-Provider-on-DPA-device update / disable path specifically), the DPA-vs-host-path boundary, the device-preconditions table (ConnectX-7+ / BlueField-3+), the documented VNF-only Flow-mode rule, the PF-vs-VF-vs-SF rule (SFs not supported on DPA), the workload-shape axes (burst, queue, completion threshold, hash pipe algorithm, work policy, PSL tables, table size, workers), the operation axis (update vs disable-enable), the version overlay (this tool rides the doca-flow and doca-dpa versions it links against; the canonical rules live in doca-version), the layered error taxonomy (config-syntax / device-binding / dpa-precondition / workload-precondition / measurement-soundness / self-test / version / cross-cutting), the observability surface (iteration statistics, self-test path-selector verification, tcpdump-side traffic verification), and the safety posture (smoke-before-bulk, four-tuple capture, name the tool that produced the number).
TASKS.md — step-by-step workflows for the in-scope taskverbs: install (route to setup; the binary is shipped), configure (DPA-preconditions + active / passive device + workload-shape decision), build (route to install — the binary is shipped), modify (refuse — modify the invocation, not the binary), run (smoke before bulk), test (eval loop), debug (layered diagnosis), use (consume the captured number), plus a Deferred task verbs block routing out-of-scope questions and a Command appendix.
The skill assumes a host where DOCA is already installed (or the NGC DOCA container is running) on a DPA-capable device and the operator has the permissions to bind the device and allocate the DPA execution resources the tool needs.
This skill is agent guidance, not a samples or scripts bundle. To keep the boundary clean, it deliberately does not contain — and pull requests should not add:
the shipped README or installed --help documents. Read defaults from the README first, then fall back to the installed binary's --help. If neither defines a needed default, stop and request the operator's explicit value instead of guessing. The flag surface is install-specific within the documented surface; the documented invocations + --help on the installed version are the authoritative answer. Inventing a flag is the most common hallucination failure.
numbers. Output is device-, firmware-, DOCA-version-, workload-, and platform-specific; a pinned number for one platform misleads operators on a different platform / version. The shipped README's example numbers are illustrative, not a baseline the agent should quote as ground truth.
consume the tool's stdout / CSV. The output format is documented; if a user wants to script against it, the right answer is "read the live guide, write the parser against your installed version".
samples/ or reference/ subtree. This is a thinloader for a documented CLI; substantive material lives on the public page, in --help, and in the shipped README on the user's install.
SKILL.md first to confirm the user's questionis in scope (the user actually wants to invoke doca_flow_dpa_perf on DPA-capable hardware, not measure the host / DPU-CPU Flow path).
doca_flow_dpa_perf measures, the DPA-vs-hostboundary, the device-preconditions table, the workload- shape axes, the version overlay, the error taxonomy, the observability surface, and the safety posture, see CAPABILITIES.md.
workflow — install, configure, build, modify, run, test, debug, use — see TASKS.md.
doca-flow — the baselibrary whose pipeline this tool measures on the DPA path. The pipe / entry / rule surface this tool drives is created by doca-flow program code; the library's pipe attributes and capability surface are the upstream context.
doca-dpa — theprogramming model behind the DPA execution engine the tool runs on. When the user's question goes from "measure the DPA path" to "why is the DPA path doing this", that skill is the next stop.
doca-flow-perf — thehost / DPU-CPU Flow performance tool. The cross-tool comparison rule lives in CAPABILITIES.md ## Capabilities and modes: name which tool produced which number.
doca-flow-tune — the Flowtuning tool. A DPA-perf number is the kind of baseline doca-flow-tune then optimizes on top of, via a Flow-program modify-a-sample loop.
doca-public-knowledge-map —routing to the public DOCA Flow DPA Perf page on docs.nvidia.com and the rest of the public DOCA documentation set.
doca-version — canonicalDOCA version-handling rules. The ## Version compatibility section in this skill is a thin overlay on top.
doca-setup — env preparation,install verification, hugepages, NUMA awareness, and the I have no install yet path with the public NGC DOCA container.
doca-debug — the cross-cuttingdebug ladder. DPA-perf surfaces its own error taxonomy; when the cause turns out to be below DOCA, the taxonomy hands off to doca-debug.
doca-hardware-safety —the cross-cutting hardware-safety meta-policy this skill's ## Safety policy overlays.
Other measured skills in the registry, with their headline benchmark lift.