Get started Dashboard
AI in production · Updated
Part 1 of AI x Production: coding agents, MCP and autofix

How to Use Claude Code to Debug Production Issues

Explore with AI

Give Claude Code the production evidence it can't see by itself. Pipe in logs and stack traces, let it run read-only CLIs such as gh, aws, gcloud and sentry-cli, or connect an MCP server that queries your telemetry. Start in plan mode and ask for hypotheses with the evidence for each. Check recent commits, then have it write a failing test that reproduces the bug before it fixes the root cause. Keep production credentials read-only, and ship the fix only through a pull request you review.

On this page

Claude Code needs two things to debug a production issue well. The first is your code, which it reads from the repository. The second is the production evidence, which it can’t see until you give it a way in. You bring that evidence in by piping logs and stack traces into the session, by letting it run read-only command-line tools, or by connecting an MCP server that queries your telemetry.

After that, the workflow is the one a careful engineer follows. Read before editing. List hypotheses and check each one against the evidence and the recent commits. Reproduce the bug in a failing test, fix the root cause, and ship the fix as a pull request. This page walks through each step with real commands and prompts. It also covers the guardrails that stop a debugging session from changing production.

What debugging production with Claude Code means

Claude Code is Anthropic’s agentic coding environment. It runs in the terminal, and there are also versions for VS Code, JetBrains IDEs, the desktop and the web. It reads files, runs shell commands, edits code and works through a problem in a loop until it judges the work done. To debug production with it, you feed that loop the evidence of a live failure, and it traces the evidence back through your code to a cause.

The boundary matters. Claude Code needs a command, a pasted file, a URL you give it or an MCP server before it can see production data such as logs, metrics, traces or deploy history. When that evidence is missing, it has nothing to check its reasoning against, so ask it to show what each conclusion rests on. I built and led the Workers observability team at Cloudflare, and the lesson carries straight over: an agent can only reason over the telemetry you emit and hand to it.

Claude Code starts when you prompt it, schedule it or trigger it from an event: routines can fire on API calls or GitHub events, GitHub Actions can run it in CI, and channels can push alerts and webhooks into a session from an MCP server. For most teams, it is the tool you reach for once you already know something is wrong. Tools sold as AI SREs aim at a different job, and the related guide on what an AI SRE is covers that category.

Three ways to get production evidence into the session

Pipe logs and stack traces in directly

This is the fastest route. Filter the logs to the incident window first, then send them in:

grep "req_8f2c" /var/log/checkout/app.log | claude
git log --oneline -20 | claude -p "summarize these recent commits"

claude -p runs one query and exits, which suits quick questions and scripts. You can paste a stack trace straight into the prompt. You can also paste a screenshot of a dashboard or an error page with Ctrl+V (Alt+V on Windows and WSL), or give it an image path. Tell Claude whether the failure is intermittent or consistent, and give it the command that reproduces it if you have one.

Let it run the CLIs you already use

CLI tools are the most context-efficient way for Claude Code to reach external services. Tell it to use gh for pull requests and CI runs, and aws, gcloud or sentry-cli for your cloud and error data. Give those CLIs read-only credentials. Then write down the commands it can’t guess in CLAUDE.md, the file Claude reads at the start of every session:

# Production debugging
- Production logs: ./scripts/prod-logs.sh SERVICE MINUTES (read-only)
- Errors: use sentry-cli; the token in this shell is read-only
- Deploys are tagged with the commit SHA; map a deploy to its changes with git log
- Never run commands that write to production resources

Keep the file short. Claude follows a concise CLAUDE.md more reliably than a long one. Run /init to generate a starter file, and /context to confirm it loaded.

Connect your telemetry over MCP

If your observability tool or platform runs an MCP server, register it once and Claude can query it mid-conversation:

claude mcp add --transport http telemetry https://mcp.example.com/mcp

For a server the whole team shares, put it in .mcp.json at the repository root under the mcpServers key. A project-scoped server needs a one-time approval, so run /mcp to check its status and approve it. Once it is connected, you can pull MCP resources into a prompt with the @server:resource syntax, the same way you reference files with @.

Working an incident with Claude Code, step by step

  1. Start in plan mode. Run claude --permission-mode plan, or press Shift+Tab until the status bar shows plan mode on. In this mode Claude reads files and answers questions without making changes.
  2. Describe the symptom with its evidence. Say what broke, when it started, where the evidence is, and what fixed looks like. Ask for causes before any fix:
Checkout requests started failing after this morning's deploy.
The stack trace and recent log lines are in /tmp/incident/.
Failures are intermittent: some orders succeed.
Read the code paths in the trace and list the likely causes.
For each cause, say what evidence would confirm or rule it out.
Don't edit anything yet.
  1. Line the failure up with recent changes. Ask it to use git log to find commits since the last good deploy that touched the files in the trace, and to say which of them could produce this error.
  2. Test each hypothesis against evidence. Have it run the read-only queries or log searches that would confirm or rule out each cause, and quote the output it relied on. If it says a cause is confirmed without showing a log line, query result or commit, ask for the evidence.
  3. Reproduce the bug in a failing test. Ask it to write a test that fails for the reason it identified. Watching that test fail proves the diagnosis.
  4. Fix the root cause and verify. Leave plan mode by approving the plan or pressing Shift+Tab. Press Ctrl+G first if you want to edit the plan in your editor. Then ask it to fix the cause, run the test suite and fix any failures.
  5. Ship through review. Ask it to commit with a descriptive message and open a PR. Your normal review and CI decide the merge. You can find the session again later with claude --from-pr followed by the PR number.
~/app
$ claude
✳ 12 files · main · last session 2h ago
> checkout errors began after the last deploy. list hypotheses and the evidence for each
? for shortcuts
Ask in plan mode first, so Claude reads and reasons before it edits anything.

Keeping Claude Code read-only against production

Instructions in CLAUDE.md are guidance: they tell Claude how you work. Permissions, hooks and the sandbox are the parts that enforce limits whatever Claude decides. When production is involved, use the enforcing layer.

  • Give the shell read-only credentials. This is where I’d start: a read-only cloud role or API token limits what any command Claude runs can do against production.
  • Pick the permission mode on purpose. In Manual mode, Claude asks before file writes, Bash commands and MCP tool calls. From Claude Code v2.1.283, interactive terminal and VS Code sessions start in auto mode. There, a separate classifier model reviews most actions and blocks only what looks risky, such as scope escalation or unknown infrastructure. During an incident, press Shift+Tab to switch mode if you want to approve each command yourself.
  • Allowlist the safe commands. Use /permissions to pre-approve the read-only commands you trust, so you only get prompts for actions that matter. Run /permissions again to see the allow and deny rules actually in effect.
  • Sandbox the session. /sandbox turns on OS-level isolation that restricts filesystem and network access.
  • Know what deny rules miss. Bash rules match the literal command string. A Bash(rm *) deny rule doesn’t block /bin/rm or find -delete. For a hard guarantee, use a PreToolUse hook or the sandbox.

Holding the context window together during a long investigation

Claude’s context window holds every message, every file it reads and every command output. A single debugging session can use tens of thousands of tokens, and model performance drops as the window fills: Claude starts forgetting earlier instructions and making more mistakes. Incidents are where this bites, because logs are large.

  • Filter before you send. Grep for the request ID, the error class or the time window before you pipe anything in.
  • Delegate wide searches to a subagent. A prompt like “use a subagent to search the checkout logs for this request ID and report only the matching lines” runs the search in a separate context window, and only the summary comes back.
  • Watch the budget. /context breaks down what fills the window. Use /clear between unrelated incidents. If an investigation spans several sittings, claude --continue resumes the latest session in the directory and claude --resume lets you pick one from a list.
  • Move runbooks into skills. A procedure you need only during incidents belongs in a skill. Claude loads a skill when a request matches it, so it doesn’t sit in every session the way CLAUDE.md does.

Proving the fix before it ships

Claude stops when the work looks done. Unless it has a check it can run, you become the only verification, so give it something that returns pass or fail.

  • In the prompt: “write a failing test that reproduces the issue, then fix it, then run the tests.” Add “address the root cause, don’t suppress the error” when the likely quick fix is a guard clause.
  • Across the session: set the check as a /goal condition. A separate evaluator re-checks it after every turn.
  • As a hard gate: a Stop hook runs your check as a script and blocks the turn from ending until it passes.
  • With a second opinion: a verification subagent gets a fresh model to try to refute the fix. That way the agent that did the work isn’t the one grading it.

Ask for evidence, and don’t accept a claim of success on its own. The test output, the command it ran and what it returned are faster to review than rerunning everything yourself. After Claude’s checks pass, you can run /verify yourself to confirm the change against the running app.

Where production debugging with Claude Code goes wrong

  • A confident cause with no evidence. A stated cause with no log line, query or commit behind it is a guess. Fix: require every claim to name the log line, query output or commit behind it, and to mark anything unverified as unknown.
  • Raw logs flooding the context. Dumping an unfiltered log file pushes your instructions out of the model’s attention. Fix: filter to the incident window first, or hand the search to a subagent.
  • Patching the symptom. A null guard where the data should never be null hides the bug until it breaks somewhere else. Fix: ask why the value is missing, fix it at the source, add a regression test, and ask Claude to search the codebase for the same pattern.
  • Trusting CLAUDE.md to prevent writes. A rule in a markdown file is advice. Fix: read-only credentials, permission rules, a PreToolUse hook or the sandbox.
  • An MCP server that never shows up. The usual causes are .mcp.json placed inside .claude/ or servers listed under a servers key. Two others are a dismissed approval prompt and relative paths, which resolve against the directory you launched from. If /mcp shows the server connected with zero tools, choose Reconnect. If it still shows zero, run claude --debug=mcp and read the server’s stderr in ~/.claude/debug/<session-id>.txt.
  • Your own configuration getting in the way. If Claude behaves oddly mid-incident, start claude --safe-mode. It disables CLAUDE.md, skills, plugins, hooks and MCP servers, so you can tell whether a customisation is the cause.

Running production checks without a person at the keyboard

Claude Code can run on a schedule, which covers recurring checks such as looking for CI failures overnight. Routines run in the cloud, so they work when your computer is off, and they can also trigger on API calls or GitHub events. Desktop scheduled tasks run on your machine, GitHub Actions runs in your CI pipeline, and /loop polls within an open session. A scheduled run can’t ask you clarifying questions, so its prompt has to say what success looks like and what to do with the result. Keep these runs on read-only credentials too.

Giving Claude Code live production context with Polylane

Polylane connects your clouds, repositories and observability tools into one context graph. It gives that graph to Claude Code through a plugin that installs an MCP server, skills and an investigate command. With it, a question like what errored in production for the service this branch deploys is answered from your real telemetry and topology. Agent tools are read-only by default: a write needs an explicit header and a credential with the write scope. A safety model screens every write, and clients that support confirmation prompts ask you to approve each one before it runs.

Using Claude Code? See how to give Claude Code production context with Polylane.

Common questions.

Can Claude Code read my production logs directly?

Only through a path you give it. You can pipe a filtered log file in with a command like cat error.log | claude, let it run read-only CLIs such as aws, gcloud or sentry-cli, give it a URL to fetch, or connect an MCP server that queries your observability tool.

Is it safe to give Claude Code access to production?

It is much lower risk when every layer is read-only. Give its shell read-only credentials, allowlist only the read commands you trust with /permissions, and turn on /sandbox for OS-level isolation. Send every fix through a pull request that your review and CI gate.

How do I connect Claude Code to my observability tool?

If the tool has a CLI, install it and tell Claude to use it. That is the most context-efficient option. If it runs an MCP server, register it with claude mcp add --transport http NAME URL, or add it to .mcp.json at the repository root under mcpServers. Then run /mcp to approve it and check it is connected.

What is a good first prompt when an incident starts?

Start in plan mode with claude --permission-mode plan. Describe the symptom, when it started, where the stack trace and logs are, and whether it is intermittent. Ask for a list of likely causes, with the evidence that would confirm or rule out each one, and tell it not to edit anything yet.

Why does Claude Code sometimes give a confident but wrong root cause?

One reason is a full context window: as it fills, Claude forgets earlier instructions and makes more mistakes. Another is missing evidence: without logs, metrics or deploy history in the session, it has nothing to check its reasoning against. Require each claim to cite the command output behind it, filter logs before sending them, and use a subagent for wide searches.

How do I stop Claude Code from running a destructive command?

Use a PreToolUse hook or the sandbox. Bash deny rules match the literal command string, so a Bash(rm *) rule does not block /bin/rm or find -delete. Keeping its credentials read-only is a sensible backstop as well.

Can Claude Code watch production and start debugging on its own?

It starts when you prompt it, schedule it or trigger it from an event. Routines run in the cloud on a schedule or on API calls and GitHub events, GitHub Actions runs it in CI, channels push alerts and webhooks into a session from an MCP server, and /loop polls within an open session. For tools built around alert-driven investigation, see the related guide on what an AI SRE is.

How do I pick an incident investigation back up later?

Run claude --continue to resume the most recent session in the current directory, or claude --resume to pick from a list. If Claude opened a pull request during the session, claude --from-pr followed by the PR number opens the session picker filtered to that PR.

Sources

  1. Best practices for Claude Code
  2. Common workflows - Claude Code Docs
  3. Debug your configuration - Claude Code Docs
  4. Quickstart - Claude Code Docs
  5. Claude Code documentation index
  6. Polylane documentation

About the author

Boris Tane

Founder of Polylane

Boris Tane is the founder of Polylane. He previously founded Baselime, observability for the future of the cloud, which Cloudflare acquired. At Cloudflare he built and led the Workers observability team.

More in this series

AI x Production: coding agents, MCP and autofix

  1. 1 How to Use Claude Code to Debug Production Issues
  2. 2 How to keep AI coding agents from breaking production

Related

Nobody should be on-call. Polylane watches your infra, finds what broke, and writes the fix.

Get started for free