Kubernetes Agent Troubleshooting
Comprehensive diagnostic checklist, failure recovery playbooks, RBAC remediation guides, and system health verification
This troubleshooting manual provides step-by-step procedures to diagnose and resolve common connectivity, workflow, authentication, and execution failures encountered when running k8s-autopilot.
1. Pre-Flight Diagnostic Checklist
Before troubleshooting individual operator errors, verify the baseline health of your environment:
# 1. Verify container runtime status
docker compose -f ~/.k8s-autopilot/docker-compose.yml ps
# 2. Check cluster reachability from your local shell
kubectl cluster-info
# 3. Stream real-time agent engine logs
docker compose -f ~/.k8s-autopilot/docker-compose.yml logs -f --tail=100
# 4. Verify system health via the install script
curl -LsSf https://raw.githubusercontent.com/talkops-ai/k8s-autopilot/main/scripts/install.sh | bash -s -- --status
2. Connectivity & Cluster Startup Issues
"Connection Refused (Kubernetes API)"
Error: dial tcp 127.0.0.1:6443: connect: connection refused
Root Cause: Your local kubeconfig points to 127.0.0.1 or localhost. Inside the Docker container, localhost refers to the container itself rather than your host machine.
Resolution:
- Docker Desktop (macOS / Windows): Open
~/.kube/configand replace127.0.0.1:6443withhost.docker.internal:6443. - Linux Hosts: Run the agent container with
--network hostmode so it shares the host's networking namespace directly. - Kind / Minikube: Run
kubectl config view --rawand ensure the server endpoint references the container network IP rather than the loopback interface.
"Kubeconfig File Not Found"
Error: Kubernetes configuration file not found at /root/.kube/config
Root Cause: The volume mount mapping your local kubeconfig into the container is missing or points to a non-existent path.
Resolution:
Ensure ~/.k8s-autopilot/docker-compose.yml mounts your configuration correctly:
volumes:
- ~/.kube/config:/root/.kube/config:ro
Verify the file exists on the host with ls -la ~/.kube/config.
"MCP Connection Refused / Subprocess Timeout"
Error: JIT MCP connection failed for helm-mcp-server or Connection timeout after 30s
Root Cause: The MCP tool subprocess failed to spawn, crashed during initialization, or exceeded its connection timeout.
Resolution:
- In Settings → Runtime Config, increase
MCP_TIMEOUT_CONNECTfrom30to60seconds. - If using Dockerized MCP servers over HTTP, verify both the agent and MCP containers share the same Docker network.
- Inspect standard error logs for the target subprocess by setting
LOG_LEVEL=DEBUG.
3. Workflow & Governance Issues
Agent Caught in "Validating / Fixing" Loops
Symptom: The agent enters an iterative loop of "Generating manifest..." ➔ "Lint failed..." ➔ "Fixing..." without making forward progress.
Root Cause: The underlying LLM is caught in a self-healing loop where its attempted fix introduces a secondary syntax error.
Resolution:
- Hard Limit Intervention: k8s-autopilot enforces a default limit of 2 self-healing retries per tool. The agent will stop and ask for your input.
- Manual Prompt Injection: Provide explicit constraints in your chat prompt:
Stop retrying. Use the Bitnami Redis chart with architecture=standalone and skip sentinel configuration.
"Approval Required" but No Buttons Appear in UI
Symptom: The agent conversation says it requires confirmation, but no interactive approval card is rendered.
Root Cause: The client frontend may have experienced a temporary WebSocket disconnect, or you are interacting via a headless CLI.
Resolution:
Type plain-text approval directly into the chat:
APPROVE
or
PROCEED
The agent's fallback intent parser accepts uppercase text confirmation as an alternative to UI clicks.
4. RBAC & Kubernetes Permission Failures
"Forbidden: User cannot create resource"
Error: deployments.apps is forbidden: User "system:serviceaccount:default:k8s-autopilot" cannot create resource "deployments" in API group "apps"
Root Cause: The Kubernetes ServiceAccount bound to k8s-autopilot lacks sufficient RBAC privileges.
Resolution:
Apply a ClusterRoleBinding granting required permissions:
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
name: k8s-autopilot-operator-role
rules:
- apiGroups: ["", "apps", "batch", "networking.k8s.io"]
resources: ["*"]
verbs: ["get", "list", "watch", "create", "update", "patch", "delete"]
- apiGroups: ["traefik.io", "argoproj.io", "monitoring.coreos.com"]
resources: ["*"]
verbs: ["*"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata:
name: k8s-autopilot-operator-binding
subjects:
- kind: ServiceAccount
name: k8s-autopilot
namespace: default
roleRef:
kind: ClusterRole
name: k8s-autopilot-operator-role
apiGroup: rbac.authorization.k8s.io
Helm "Release Not Found" Despite Release Being Deployed
Error: Error: release: "my-app" not found
Root Cause: Helm v3 stores release state in Kubernetes Secrets (or ConfigMaps) within the target release namespace. If the agent lacks RBAC permissions to read v1/Secrets in that namespace, Helm reports the release as non-existent.
Resolution:
Ensure the agent's ClusterRole includes secrets and configmaps in apiGroups: [""].
5. Database & Storage Triage
SQLite "Database is Locked" Error
Error: sqlite3.OperationalError: database is locked
Root Cause: High concurrency or multiple threads accessing SQLite simultaneously without Write-Ahead Logging (WAL).
Resolution:
- Ensure WAL mode is active: k8s-autopilot enables WAL by default.
- For multi-replica production environments, migrate to PostgreSQL:
K8S_AUTOPILOT_DATABASE_URL=postgresql://user:pass@postgres.monitoring:5432/autopilot
6. Enabling Debug Logging
To capture comprehensive diagnostic traces across all LangGraph nodes and MCP tool invocations:
- In
~/.k8s-autopilot/.env, set:LOG_LEVEL=DEBUG
LOG_MODE=json - Restart the agent daemon:
docker compose -f ~/.k8s-autopilot/docker-compose.yml up -d - Inspect structured logs for the exact JSON payloads transmitted between operators.
Next Steps
- Configuration Reference — Review all environment variables and timeout budgets.
- Approval Modes & Governance — Explore safety tiers and execution gating.
- TalkOps Discord Community — Connect with core maintainers and DevOps community engineers.