Install any skill in seconds. Free to start, no credit card required.
Get Started Free →Use this skill when the user is building, running, or interpreting the doca/tools/gpunetio_ib_write_bw client+server benchmark — a CUDA kernel on the client posts RDMA WRITE work requests through the doca-gpunetio device-side surface to measure sustained GPU-driven WRITE bandwidth on a GPU+IB-device pair. Trigger even when the user does not explicitly mention "doca-gpunetio-ib-write-bw" or "GPUNetIO" — typical implicit phrasings include "measure WRITE BW when the GPU posts the WRs", "BW swings b
| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-01 | ✗→✓ | ▲ Improved | 67% | 0% |
| case-04 | ✗→✓ | ▲ Improved | 111% | 0% |
| case-05 | ✗→✓ | ▲ Improved | 93% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 185% | 0% |
| case-12 | ✗→✓ | ▲ Improved | 304% | 0% |
Where to start: This is a tool skill for the GPUNetIO- flavored ib_write_bw benchmark shipped under doca/tools/gpunetio_ib_write_bw/ (a client + server pair, built from source against the installed DOCA via meson). It measures sustained RDMA WRITE bandwidth when the WRs are posted from a CUDA kernel through the doca-gpunetio device-side surface, with the GPU on the data path. Open TASKS.md and start at ## configure for the GPU-NIC pairing precondition and the build pattern; jump to ## run for the smoke-before-bulk flow. Open CAPABILITIES.md when the question is what this tool actually measures, how the result decomposes (GPU occupancy vs NIC issue rate vs link saturation), or how the result reads against the GPI sister tool and the upstream CPU-initiated `perftest` `ib_write_bw`. If DOCA is not installed yet, route to doca-setup first; if the user is still deciding between the GPI and GPUNetIO programming surfaces, the picture in ../../libs/doca-gpunetio/CAPABILITIES.md#capabilities-and-modes and ../../libs/doca-gpi/CAPABILITIES.md#capabilities-and-modes is the first stop.
The CLASSES of doca-gpunetio-ib-write-bw questions this skill is built to answer, each with one worked example. The class is the load-bearing piece; the worked example is one instance.
path deliver on this GPU-NIC pair?" — worked example: "measure sustained WRITE BW between two hosts with an H100 + ConnectX-7 on each side". Answered by the GPU-NIC pairing precondition in CAPABILITIES.md ## Capabilities and modes + the bring-up flow in TASKS.md ## configure + TASKS.md ## run. The same shape answers "measure GPUNetIO-driven WRITE BW between a host GPU and a BlueField DPU".
issue rate, or link saturation?" — worked example: "I see 120 Gbit/s on a 200 Gbit/s link; is the NIC saturated, am I CPU-bound on the client, or is the CUDA kernel not driving enough WRs in flight?". Answered by the throughput-decomposition rules in CAPABILITIES.md ## Observability + the eval-loop overlay in TASKS.md ## test.
initiated perftest ib_write_bw?" — worked example: "my team has a CPU-initiated WRITE BW number on this same NIC; should I expect the GPUNetIO number to match or be different?". Answered by the "GPU-initiated path adds (or removes) overhead vs the CPU-initiated path" rule in CAPABILITIES.md ## Capabilities and modes.
sustained-throughput workload class?" — worked example: "my application streams sensor data from GPU memory at line rate to a remote consumer". Answered by the "when GPUNetIO is the right surface vs GPI vs CPU- initiated" rule in CAPABILITIES.md ## Capabilities and modes + the use-side decision in TASKS.md ## use.
before quoting it?" — worked example: "three runs at the same flags gave 145, 187, and 160 Gbit/s; is the benchmark noisy or is my platform inconsistent?". Answered by the measurement-soundness rules in CAPABILITIES.md ## Error taxonomy layer 5 + the steady-state guidance in TASKS.md ## test.
binary to build and run?" — worked example: "my install has DOCA at one semver and CUDA at another; will the ToT-shipped `gpunetio_ib_write_bw` even link?". Answered by the version overlay in CAPABILITIES.md ## Version compatibility which cross-links the canonical detection chain in doca-version.
This skill serves external developers and performance engineers who need a reproducible measurement of sustained RDMA WRITE bandwidth when the WRs are posted from a CUDA kernel through doca-gpunetio, on the user's actual install and GPU-NIC pair. Concretely:
path or the host-initiated perftest-style path before committing an application design to one of them.
pinning, GPU PCIe placement, IB device choice, GID index, NIC firmware burn) by re-running this benchmark against the new state.
GPUNetIO-driven WRITE BW on this GPU-NIC pair today" artifact downstream consumers can cite.
for my sustained-throughput workload class" honestly — with a measured number, the build + invocation that produced it, and the GPU + NIC + DOCA version that scopes it — rather than guessing from datasheet headlines.
It is not for users debugging the doca-gpunetio library itself (route to ../../libs/doca-gpunetio/SKILL.md), and not a substitute for the perftest upstream ib_write_bw (which measures CPU-initiated WRITE BW).
The doca-gpunetio-ib-write-bw tool is shipped as C plus a CUDA .cu translation unit under doca/tools/gpunetio_ib_write_bw/, split into a client/ subtree and a server/ subtree. The verified surface (per client/{main.c,common.h,common.c,kernel.cu,perftest.c} and server/{main.c,common.h,common.c,perftest.c}): host-side build via meson against the installed DOCA pkg-config modules (doca-gpunetio, doca-rdma, doca-common); the device-side build via nvcc against the DOCA GPU NetIO device-side header set; the OOB descriptor exchange via a TCP socket between client and server. There is no Python / Rust / Go binding — the tool is a pair of CLI binaries. The skill's job is to keep the operator-side workflow language-neutral; the device-side CUDA surface is not wrappable in another language.
Load this skill when the user is — or the agent needs to — build and run the gpunetio_ib_write_bw client + server on real hosts with DOCA installed plus a CUDA Toolkit matched to the DOCA install, and a GPU + IB device pair on the host's PCIe topology. Concretely:
bandwidth between two hosts (or a host and a BlueField DPU) with the GPUNetIO surface.
surface for a class of workload vs the GPI programming surface (the doca-gpi library — doca/tools/ ships no GPI benchmark binary) or the classic CPU-initiated perftest path.
DOCA version + GPU + NIC + as-deployed environment + numbers) for later regression hunts.
the GPUNetIO + RDMA bring-up sequence under this tool's shipped scaffolding.
Do not load this skill for general DOCA orientation, library API work, or installation. For those, use doca-public-knowledge-map, ../../libs/doca-gpunetio/SKILL.md, or doca-setup. Do not load it for application-level end-to-end throughput either — this benchmark measures the WR-submission path through GPUNetIO, not the user's full pipeline.
This is a thin loader. Substantive material lives in two companion files:
CAPABILITIES.md — what the tool measures (thesustained-WRITE-BW primitive driven by a client-side CUDA kernel through doca-gpunetio), the runtime-surface selection rule (GPUNetIO vs GPI vs CPU-initiated), the GPU-NIC pairing precondition, the throughput-decomposition guide (GPU compute occupancy vs NIC issue rate vs link saturation), the version overlay (DOCA .pc PLUS CUDA Toolkit), the layered error taxonomy (config-syntax / build-time / GPU-NIC- pairing / GPUNetIO-lifecycle / RDMA-connection / measurement-soundness / version / cross-cutting), the observability surface (stdout report, DOCA log levels, OOB-socket exchange), and the safety overlay (the "GPU-side handle is a credential" rule from doca-gpunetio; the cross-cutting hardware-safety meta-policy).
TASKS.md — step-by-step workflows for the in-scopetask verbs: install (preconditions — DOCA install, CUDA Toolkit, GPU + NIC pair, OOB connectivity), configure (build-tree under doca/tools/gpunetio_ib_write_bw/ and the meson build wrapping the shipped DOCA), build (the meson setup + meson compile pattern from the public DOCA build documentation), modify (do not patch the shipped tool source; modify the invocation and the surrounding environment instead), run (smoke- before-bulk; client + server bring-up order; reading the per-iteration report), test (the eval loop — steady-state, NUMA placement, NIC saturation cross- check), debug (walk the error taxonomy layer by layer), use (how a BW result feeds a class-of- workload decision), plus a Deferred task verbs block routing out-of-scope questions.
The skill assumes a host where DOCA is already installed, a CUDA Toolkit matched to the install is present, and the operator has whatever privileges the public install profile expects for binding a doca_dev, a doca_gpu, and an OOB TCP socket.
This skill is agent guidance, not a samples or scripts bundle. To keep the boundary clean, it deliberately does not contain — and pull requests should not add:
beyond what the tool's shipped --help and main.c ARGP registration establish. The flag surface is small (device name, GPU PCIe address, GID index, server IP on the client side); the agent re-reads the binary's --help on the installed version before quoting flag strings. Throughput numbers are device-, firmware-, version-, and topology-specific.
that would compete with the shipped tool tree. The shipped client/{main.c,kernel.cu,perftest.c,common.{c,h}} and server/{main.c,perftest.c,common.{c,h}} files are the verified worked example; the agent's job is to route the user there and prescribe minimum-diff modification per the universal modify-a-sample workflow in doca-programming-guide.
consume the tool's stdout. The output format is small and documented in CAPABILITIES.md ## Observability; if the user wants to script against it, the right answer is "read the live source, write the parser against your installed binary".
samples/, bindings/, or reference/ subtree.This is a thin loader for a shipped tool tree; substantive material lives in the source tree and in the GPUNetIO library docs.
SKILL.md first to confirm the user'squestion is in scope (the user actually wants to measure sustained kernel-initiated WRITE BW through GPUNetIO, not learn GPUNetIO as a library or do a CPU-initiated measurement).
rule against the GPI sister tool and the CPU-initiated perftest, the throughput-decomposition guide, the version overlay, the error taxonomy, the observability surface, and the safety overlay, see CAPABILITIES.md.
install, configure,build, modify, run, test, debug, use — see TASKS.md.
../../libs/doca-gpunetio/SKILL.md —the library this tool wraps. The per-GPU doca_gpu context, the GPU-visible doca_gpu_eth_* and RDMA-side handles, the CUDA-side persistent-kernel pattern, the dual capability-discovery rule (DOCA cap-query AND cudaGetDeviceProperties), and the env preconditions (nvidia_peermem loaded, CUDA buffers registered with DOCA) live there.
../../libs/doca-rdma/SKILL.md —the underlying RDMA library. The RDMA queue this tool binds is created and connected via doca-rdma; the queue lifecycle, transport type (RC vs UC vs UD), permission matrix, and connection method are owned there.
../../libs/doca-verbs/SKILL.md —the raw-verbs escape hatch beneath doca-rdma / doca-gpunetio. This tool stays on the higher-level surfaces; doca-verbs is the right place only if the user needs a specific WR flag / QP attribute the GPUNetIO + RDMA surfaces do not expose.
../doca-gpunetio-ib-write-lat/SKILL.md —the latency analog of this tool. Same physical operation; same runtime framework; different metric class (BW vs latency). The two together carry the full GPUNetIO-side throughput / latency picture.
doca-gpi — the GPIprogramming surface (CUDA-kernel-initiated RDMA), the alternative runtime framework for the same physical operation. doca/tools/ ships no GPI ib_write_lat / ib_write_bw benchmark binary, so the GPI comparison is against the library surface, not a sibling tool. The selection rule in CAPABILITIES.md ## Capabilities and modes is the decision aid.
doca-version — thecanonical version-detection chain, four-way match rule, NGC container semantics, and headers-win-over-docs rule. The ## Version compatibility section in this skill is a thin overlay; the body lives there.
doca-setup — envpreparation, install verification, GPU + CUDA Toolkit pairing, nvidia_peermem load, hugepages, NUMA, and the I have no install yet path with the public NGC DOCA container.
doca-public-knowledge-map —routing to the public DOCA documentation set (DOCA GPU NetIO, DOCA RDMA pages on docs.nvidia.com) and the docs.nvidia.com/cuda/ pointer for the CUDA Toolkit.
doca-debug — thecross-cutting debug ladder. The tool surfaces its own error taxonomy; when the cause is below DOCA, the taxonomy hands off here.
doca-hardware-safety —the bundle-wide hardware-safety meta-policy. The ## Safety policy overlay cross-links it.
Other measured skills in the registry, with their headline benchmark lift.