Install any skill in seconds. Free to start, no credit card required.
Get Started Free →KubeSphere ServiceMesh extension management Skill (Istio + Kiali + Jaeger). Use this skill for the ServiceMesh extension Configuration (covers installation, uninstallation, status checks), troubleshooting (covers grayscale release, sidecar injection, topology/metrics, and tracing for Composed Apps Aka Custom Applications).
.claude/skills/kubesphere-kubesphere-servicemesh/SKILL.md| Test case | Without → With | Effect | Δ tokens | Δ turns |
|---|---|---|---|---|
| case-05 | ✗→✓ | ▲ Improved | 44% | 0% |
| case-06 | ✗→✓ | ▲ Improved | 33% | 0% |
| case-07 | ✗→✓ | ▲ Improved | 61% | 0% |
| case-08 | ✗→✓ | ▲ Improved | 70% | 0% |
| case-09 | ✗→✓ | ▲ Improved | 61% | 0% |
Integrates Istio, Kiali, and Jaeger to provide traffic governance for microservices. They operate on Composed App (backed by applications.app.k8s.io). A Service is governed when it has the annotation servicemesh.kubesphere.io/enabled: "true" and belongs to a Composed App.
Defines two core CRDs:
Istio handles traffic routing, Kiali provides topology visualization, and Jaeger enables tracing.
Check if ServiceMesh is already installed:
bashkubectl get installplans.kubesphere.io servicemesh --ignore-not-found
If found, upgrading is supported — just select a newer version in Step 1.
bashALL_VERSIONS=$(kubectl get extensionversions.kubesphere.io \ -l kubesphere.io/extension-ref=servicemesh \ -o jsonpath='{range .items[*]}{.spec.version}{"\n"}{end}' | sort -V) LATEST_STABLE=$(echo "$ALL_VERSIONS" | grep -v -E 'alpha|beta|rc' | tail -1) if [ -z "$LATEST_STABLE" ]; then LATEST_STABLE=$(echo "$ALL_VERSIONS" | tail -1) fi echo "Available versions:" echo "$ALL_VERSIONS" echo "" echo "Latest stable: $LATEST_STABLE"
This sets ALL_VERSIONS and LATEST_STABLE. Use SELECTED_VERSION for the version chosen.
Use the question tool:
$LATEST_STABLE (Recommended) — accept the auto-detected versionbashCLUSTER_DATA=$(kubectl get clusters.cluster.kubesphere.io \ -o jsonpath='{range .items[*]}{.metadata.name}{"\t"}{.status.conditions[?(@.type=="Ready")].status}{"\n"}{end}') READY_CLUSTERS=$(echo "$CLUSTER_DATA" | awk -F'\t' '$2 == "True" {print $1}') CLUSTER_COUNT=$(echo "$READY_CLUSTERS" | wc -l) HOST_CLUSTER=$(kubectl get clusters.cluster.kubesphere.io \ -l 'cluster-role.kubesphere.io/host' \ -o jsonpath='{.items[0].metadata.name}' || echo "") echo "Ready clusters:" echo "$READY_CLUSTERS" echo "" echo "Cluster count: $CLUSTER_COUNT" echo "Host cluster: $HOST_CLUSTER"
This sets READY_CLUSTERS, CLUSTER_COUNT, HOST_CLUSTER.
TARGET_CLUSTERS="$HOST_CLUSTER"question with multiple: true:TARGET_CLUSTERS="$READY_CLUSTERS"TARGET_CLUSTERS="$HOST_CLUSTER"$READY_CLUSTERSbashif kubectl get installplans.kubesphere.io gateway-api --ignore-not-found &>/dev/null; then GW_API_EXISTS=true echo "gateway-api InstallPlan exists → will set PILOT_ENABLE_GATEWAY_API=false" else GW_API_EXISTS=false echo "gateway-api not found → no compatibility config needed" fi echo "GW_API_EXISTS=$GW_API_EXISTS"
If GW_API_EXISTS=true, the InstallPlan config will include PILOT_ENABLE_GATEWAY_API: "false" to avoid conflicts.
bash./scripts/generate-installplan.sh "$SELECTED_VERSION" "$TARGET_CLUSTERS" "$GW_API_EXISTS"
This generates the YAML to /tmp/servicemesh-installplan.yaml, runs --dry-run=server, then prints the apply command.
> For configurable extension values (tracing sampling rate, storage backend, credentials, etc.), see references/extension-values.md.
Apply it:
bashkubectl apply -f /tmp/servicemesh-installplan.yaml
Tell the user "Installing". Then ask if they want to check status. If yes:
bash./scripts/check-status.sh poll
| Purpose | Command | |---|---| | Single snapshot | ./scripts/check-status.sh quick | | Wait until complete (5min timeout) | ./scripts/check-status.sh poll |
Logic:
Installed → ✓ successFailed → ✗ prints full status> ⚠ Always confirm with the user before proceeding.
bashif ! kubectl get installplans.kubesphere.io servicemesh --ignore-not-found &>/dev/null; then echo "ServiceMesh is not installed." exit 0 fi
Confirm with the user, then delete:
bashkubectl delete installplans.kubesphere.io servicemesh --ignore-not-found
Verify cleanup:
bash./scripts/verify-uninstall.sh
Success criteria:
extension-servicemesh> WARNING: Do NOT delete the InstallPlan. Only remove target clusters from the placement list.
Confirm which clusters to remove, compute remaining clusters, then patch:
bashkubectl patch installplans.kubesphere.io servicemesh --type='json' \ -p='[{"op": "replace", "path": "/spec/clusterScheduling/placement/clusters", "value": ["<REMAINING_CLUSTER_1>", "<REMAINING_CLUSTER_2>"]}]'
Success: patch returns OK + removed clusters no longer in .status.clusterSchedulingStatuses.
All scenarios below assume the namespace of the Composed App is known. Set it as NAMESPACE before proceeding.
bash# 1. Check Composed App status and governance annotation kubectl -n $NAMESPACE get applications.app.k8s.io -o yaml # Key things to verify: # - .status: health of composed components # - annotation servicemesh.kubesphere.io/enabled=true # 2. Verify pods with the injection annotation actually have istio-proxy sidecar kubectl -n $NAMESPACE get pods \ -l 'app.kubernetes.io/name,app.kubernetes.io/version,app,version' \ -o custom-columns='NAME:.metadata.name,APPLICATION:.metadata.labels.app\.kubernetes\.io/name,SHOULD_INJECT_ISTIO_PROXY:.metadata.annotations.sidecar\.istio\.io/inject,CONTAINERS:.spec.containers[*].name' # 3. If sidecar missing, check istiod injection logs kubectl logs -n extension-servicemesh -l app=istiod --tail=100
After completing the prerequisites, proceed to the specific scenario below.
When asking the user for <strategy-name>, refer to it as "grayscale release task name", not "strategy name".
Data flow: Strategy → VirtualService → Istio Proxy Sidecar → traffic routing.
After prerequisites:
bash# 1. Check Strategy reconciliation status and events kubectl -n $NAMESPACE describe strategies.servicemesh.kubesphere.io <strategy-name> # 2. Check the VirtualService controlled by this Strategy (linked via label) kubectl -n $NAMESPACE get virtualservice \ -l "servicemesh.kubesphere.io/controlled-by-strategy=<strategy-name>" -o yaml # The controller copies strategy.spec.template.spec to the VirtualService spec, # injecting only route.destination.host, route.destination.port.number, and match.port. # Compare key fields to verify sync: # spec.template.spec.hosts ↔ spec.hosts # spec.template.spec.http ↔ spec.http (port/host injected by controller) # spec.template.spec.tcp ↔ spec.tcp (port/host injected by controller) # When spec.governor is set, controller overrides all routes to 100% → governor version. # If the VirtualService spec does not reflect the template, sync failed. # 3. Verify actual workloads match the routing rules # (e.g., if routing to version: v2, check pods with that label exist) kubectl -n $NAMESPACE get pods -l "app=<app-name>,version=<target-version>" # 4. Check controller-manager sync logs for errors (last 100 lines) kubectl logs -n extension-servicemesh -l app=servicemesh-controller-manager \ --tail=100 | grep -iE "(error|reconcile|strategy|<strategy-name>)"
Before running commands, confirm with the user:
After prerequisites:
bash# Check servicemesh-apiserver logs for Kiali access errors kubectl logs -n extension-servicemesh -l app=servicemesh-apiserver --tail=100 # Check Prometheus / whizard-agent-proxy (only one will exist) kubectl get pods -n kubesphere-monitoring-system -l 'app.kubernetes.io/name in (prometheus, whizard-agent-proxy)'
Before running commands, confirm with the user:
After prerequisites:
bash# Check servicemesh-apiserver logs for Jaeger query errors kubectl logs -n extension-servicemesh -l app=servicemesh-apiserver --tail=100 # Check jaeger-query logs for storage backend connectivity kubectl logs -n extension-servicemesh -l app.kubernetes.io/component=query --tail=100 # Check jaeger-collector logs for storage backend connectivity kubectl logs -n extension-servicemesh -l app.kubernetes.io/component=collector --tail=100 # If backend.jaeger.storage.options.es.server-urls is the default, check opensearch kubectl get pods -n kubesphere-logging-system -l 'app.kubernetes.io/name=opensearch-data'
| Case | Status | Duration (ms) | Turns | Tokens | Tool calls | ||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Without | With | Δ | Without | With | Δ | Without | With | Δ | Without | With | Δ | ||
case-03 | fail→fail | 17,239 | 9,136 | -47% | 1 | 1 | 0% | 1,936 | 3,010 | +55% | 0 | 0 | — |
case-22 | pass→pass | 12,696 | 9,187 | -28% | 1 | 1 | 0% | 2,328 | 4,461 | +92% | 0 | 0 | — |
case-01 | fail→fail | 11,335 | 3,378 | -70% | 1 | 1 | 0% | 2,321 | 3,021 | +30% | 0 | 0 | — |
case-02 | fail→fail | 10,627 | 7,032 | -34% | 1 | 1 | 0% | 1,783 | 3,052 | +71% | 0 | 0 | — |
case-04 | fail→fail | 15,846 | 27,738 | +75% | 1 | 1 | 0% | 2,779 | 3,432 | +23% | 0 | 0 | — |
case-05 | fail→pass | 20,596 | 13,359 | -35% | 1 | 1 | 0% | 2,844 | 4,092 | +44% | 0 | 0 | — |
case-06 | fail→pass | 16,418 | 8,579 | -48% | 1 | 1 | 0% | 2,842 | 3,770 | +33% | 0 | 0 | — |
case-07 | fail→pass | 14,747 | 5,739 | -61% | 1 | 1 | 0% | 2,436 | 3,924 | +61% | 0 | 0 | — |
case-08 | fail→pass | 15,088 | 8,252 | -45% | 1 | 1 | 0% | 2,383 | 4,060 | +70% | 0 | 0 | — |
case-09 | fail→pass | 12,015 | 3,586 | -70% | 1 | 1 | 0% | 2,023 | 3,266 | +61% | 0 | 0 | — |
case-10 | fail→fail | 9,824 | 2,758 | -72% | 1 | 1 | 0% | 1,729 | 3,084 | +78% | 0 | 0 | — |
case-11 | fail→pass | 18,099 | 3,885 | -79% | 1 | 1 | 0% | 3,113 | 3,091 | -1% | 0 | 0 | — |
case-12 | fail→pass | 9,813 | 2,383 | -76% | 1 | 1 | 0% | 1,449 | 2,973 | +105% | 0 | 0 | — |
case-13 | fail→pass | 13,464 | 3,814 | -72% | 1 | 1 | 0% | 1,958 | 3,274 | +67% | 0 | 0 | — |
case-14 | fail→pass | 12,576 | 2,185 | -83% | 1 | 1 | 0% | 2,195 | 2,919 | +33% | 0 | 0 | — |
case-15 | pass→pass | 11,740 | 3,815 | -68% | 1 | 1 | 0% | 1,905 | 3,296 | +73% | 0 | 0 | — |
case-16 | fail→pass | 8,403 | 2,443 | -71% | 1 | 1 | 0% | 1,514 | 3,007 | +99% | 0 | 0 | — |
case-17 | pass→pass | 10,349 | 4,837 | -53% | 1 | 1 | 0% | 1,651 | 3,622 | +119% | 0 | 0 | — |
case-18 | fail→pass | 13,277 | 10,008 | -25% | 1 | 1 | 0% | 2,269 | 3,105 | +37% | 0 | 0 | — |
case-19 | fail→pass | 10,322 | 6,684 | -35% | 1 | 1 | 0% | 1,641 | 2,986 | +82% | 0 | 0 | — |
case-20 | pass→pass | 9,042 | 6,658 | -26% | 1 | 1 | 0% | 1,664 | 3,826 | +130% | 0 | 0 | — |
case-21 | pass→pass | 21,492 | 6,225 | -71% | 1 | 1 | 0% | 1,950 | 3,696 | +90% | 0 | 0 | — |
DecimalAI ran this skill against gemini-3.6-flash twice over the same eval suite — once with the skill loaded and once without — and compared the two runs case by case. 22 cases were attempted. The headline lift of +55 percentage points is the difference between those two pass rates over the 22 comparable cases. 1 case got worse with the skill loaded, and it is included in that figure.
Without the skill loaded, the model failed this case. With it loaded, the same prompt on the same model passed. This is one improved case from the latest verified run; every case, including any that regressed, is in the table above.
Other measured skills in the registry, with their headline benchmark lift.