---
name: lloydchang/debug
source: https://app.decimal.ai/s/lloydchang-debug@1/SKILL.md
source_sha256: 4113b0a8815a
---

# System Debugger Skill

## Overview

This skill provides comprehensive debugging capabilities for agents running in distributed Kubernetes environments with Temporal orchestration. It systematically diagnoses issues across the entire stack - from individual agent behavior to infrastructure health.

## Capabilities

### Component Debugging
- **Agents**: Analyze agent execution logs, performance metrics, and failure patterns
- **Temporal Workflows**: Inspect workflow execution history, timeouts, and activity failures  
- **Kubernetes Infrastructure**: Check pod health, resource utilization, and network connectivity
- **Integration Points**: Validate communication between agents, workflows, and external services

### Issue Types
- **Performance**: Slow execution, high latency, resource bottlenecks
- **Errors**: Agent failures, workflow crashes, API errors
- **Timeouts**: Workflow timeouts, agent inference delays
- **Connectivity**: Network issues, service discovery problems
- **Resource**: Memory/CPU exhaustion, storage issues
- **Behavior**: Unexpected agent responses, hallucination detection

## Usage Examples

### Basic Agent Debugging
```bash
python main.py debug \
  --target-component agents \
  --issue-type errors \
  --time-range 1h \
  --verbose
```

### Infrastructure Health Check
```bash
python main.py debug \
  --target-component infrastructure \
  --issue-type resource \
  --time-range 30m \
  --auto-fix
```

### Full System Analysis
```bash
python main.py debug \
  --target-component all \
  --issue-type performance \
  --time-range 2h \
  --namespace temporal \
  --verbose \
  --auto-fix
```

## Debugging Methodology

### 1. Information Gathering
- Collect metrics from monitoring endpoints
- Analyze Kubernetes pod status and logs
- Review Temporal workflow execution history
- Gather system resource utilization

### 2. Pattern Recognition
- Identify common failure patterns
- Detect anomalies in agent behavior
- Correlate issues across components
- Trend analysis over time ranges

### 3. Root Cause Analysis
- Trace issues through the call chain
- Identify bottlenecks and failure points
- Analyze dependencies and interactions
- Validate configuration and environment

### 4. Remediation
- Apply automatic fixes for common issues
- Generate detailed remediation plans
- Provide step-by-step resolution guides
- Create prevention strategies

## Integration Points

### Monitoring System
- Metrics API: `/monitoring/metrics`
- Alerts API: `/monitoring/alerts` 
- Health Checks: `/health`
- Audit Logs: `/audit/events`

### Temporal Integration
- Workflow History API
- Activity Execution Logs
- Task Queue Monitoring
- Worker Health Status

### Kubernetes Integration
- Pod Status and Logs
- Service Connectivity
- Resource Utilization
- Network Policies

## Output Format

The skill produces structured debugging reports with:
- **Findings**: Detailed analysis of discovered issues
- **Evidence**: Log snippets, metrics, and supporting data
- **Recommendations**: Actionable remediation steps
- **Metrics**: Summary statistics and health scores
- **Next Steps**: Follow-up actions and monitoring recommendations

## Auto-Fix Capabilities

When `auto_fix` is enabled, the skill can:
- Restart failing pods
- Clear stuck workflows
- Adjust resource limits
- Restart unhealthy agents
- Clear temporary cache issues

## Distributed System Considerations

This skill is designed for distributed environments and includes:
- Namespace isolation support
- Multi-cluster debugging capabilities
- Remote log aggregation
- Cross-component correlation
- Network connectivity validation

## Prevention and Regression

The skill helps prevent regressions by:
- Maintaining debug session history
- Tracking recurring issues
- Generating health baselines
- Creating monitoring alerts
- Providing documentation updates