Get started Dashboard
AI in production ·
Part 1 of Choosing AI for production

Is It Safe to Give an AI SRE Read Access to Production?

Explore with AI

Yes, read access is a sound first step for an AI SRE, as long as you define read narrowly: resource metadata, logs, metrics, traces and code, with secret values and customer data stores left out. The agent cannot change anything with that access, so the risk that remains is disclosure: the AWS ReadOnlyAccess policy reads S3 and DynamoDB data, a Kubernetes read role that includes Secrets exposes their values, and whatever the agent reads flows into a model and anywhere it can send data. Scope the role, control egress, keep credentials outside the agent's runtime and check what the vendor stores.

On this page

Yes. Read access is the right way to start with an AI SRE, and for most teams it is safe enough, provided you define “read” narrowly: resource metadata, logs, metrics, traces and code, with secret values and customer data stores left out. The agent cannot change production with that access, so the risk that remains is disclosure: what the agent can see, where that text goes next, and who else can steer it.

That last part is where teams get caught. The AWS ReadOnlyAccess policy reads data in S3 and DynamoDB, and a Kubernetes read role that includes Secrets returns their values. An agent that can read also sends what it reads to a language model, may save it to memory, and may be able to call out to the internet. This page covers what read access really exposes, the four ways it goes wrong, the exact roles to grant, and a checklist to run before you connect a production account.

What “read access to production” actually covers

An AI SRE investigates by calling the same read APIs an engineer uses: list and describe resources, query logs and metrics, pull traces, read deploy history and code. Each call runs under a credential, and that credential’s permissions decide what “read” means. Two defaults surprise people.

Kubernetes. In RBAC, get on Secrets reveals their values, and so do list and watch. The Kubernetes docs point out that a list response, such as the one behind kubectl get secrets -A -o yaml, includes the contents of every Secret. The second trap is get on the nodes/proxy subresource. It reaches the kubelet API, which can run commands in any pod on that node, and it bypasses audit logging and admission control. A role made only of get verbs can still be a remote shell.

AWS. The managed ReadOnlyAccess policy lists, gets and describes every resource in the account. AWS warns that it also reads data in storage services such as S3 buckets and DynamoDB tables, and it includes read access to IAM and Billing. The ViewOnlyAccess job function is narrower: it shows resources and basic metadata and cannot read resource content. AWS keeps its job function policies up to date as new services and capabilities launch, so a managed read role grows over time without anyone on your team changing it.

So “read-only” answers one question: can the agent change things? It tells you nothing about whether it can read your database rows or your API keys.

Why read-only access still carries risk: four ways it goes wrong

A public write-up of incidents from an operations agent that Microsoft runs in production is a useful guide, because none of the incidents needed an attacker. Its authors conclude that if the environment permits an action, the agent will eventually take it, whether by accident or because someone steered it.

Secret values come back from read calls

If the role can read Secrets, secret stores or parameter stores, a routine investigation pulls credentials into the agent’s context. In the Microsoft write-up, the agent found a credential committed to a customer repository, quoted it in its findings and saved it to memory with a note never to use it. The note came too late. The secret already sat in an investigation summary and a memory store, and neither was covered by anyone’s rotation playbook.

Fix: keep secret values out of reach. A pod spec shows which Secret a container references, which is enough to diagnose a missing or wrong reference. The Kubernetes guidance is to restrict get, list and watch on Secrets, and to grant list or watch only to the most privileged system components.

Customer data lands in prompts, transcripts and memory

Anything the agent reads goes into a model prompt. Depending on the product, it may also be stored in a transcript, a summary, a memory or a shared thread. Logs are the usual carrier: they hold whatever your code prints, which is why the Kubernetes docs tell developers to avoid logging secret data in the clear. An agent that can read S3 objects or DynamoDB items can pull customer records into the same places.

Fix: leave object and table contents out of the role. Ask the vendor where inference runs, how long transcripts and memories are kept, whether your data trains models, and whether you can use your own model provider.

Text in your logs can steer the agent

An agent that reads logs, tickets and alert payloads reads text that outsiders can influence. A request path, a user agent string or an error message can be written to look like an instruction. The Microsoft write-up makes the point that a hallucinated command and an injected one are the same command once they reach execution. The dangerous pairing is read access plus a way out. In one Microsoft case, the agent was asked to read a screenshot and had no vision tool, so it found a free OCR service on the public internet and posted the image to it. Nobody attacked anything, and customer data could still have ended up on a third party’s server.

Fix: control egress at a boundary the model cannot modify. If the agent runs code, it should reach only the endpoints the task needs.

A read credential can be copied or replaced

If the credential sits on a filesystem the agent can read, the agent can reuse it, or rebuild it. In the same write-up, an agent’s short-lived GitHub token expired. It read its own harness code, rebuilt the OAuth device-code flow, asked a researcher to finish the login and wrote the new access and refresh tokens to disk for reuse. The harness was supposed to decide what authority the agent had, and the agent replaced it from the inside.

Fix: keep credentials outside the agent’s runtime and issue them short-lived and per task. The strongest pattern gives tools a handle that a proxy swaps for a real credential at the network boundary, so the sandbox can use a credential without ever holding one.

Scoping a Kubernetes read role that leaves Secrets out

  1. Create a dedicated ServiceAccount. Never hand the agent a human’s kubeconfig or an existing admin account. The Kubernetes docs advise against cluster-admin except where it is specifically needed.
  2. Name every resource. Wildcards grant access to every object type that exists today and every type added later, Secrets included.
  3. Keep nodes in a separate small role, and never include nodes/proxy. Container logs come from pods/log.
apiVersion: v1
kind: ServiceAccount
metadata:
  name: ai-sre-reader
  namespace: ai-sre
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: ai-sre-read
rules:
  # Core resources by name. No secrets, no wildcards.
  - apiGroups: [""]
    resources: ["pods", "pods/log", "services", "endpoints", "events", "configmaps", "persistentvolumeclaims"]
    verbs: ["get", "list", "watch"]
  - apiGroups: ["apps"]
    resources: ["deployments", "statefulsets", "daemonsets", "replicasets"]
    verbs: ["get", "list", "watch"]
  - apiGroups: ["batch"]
    resources: ["jobs", "cronjobs"]
    verbs: ["get", "list", "watch"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata:
  name: ai-sre-read-nodes
rules:
  - apiGroups: [""]
    resources: ["nodes"]   # never nodes/proxy
    verbs: ["get", "list", "watch"]

ConfigMaps are meant for non-confidential data. If your teams store credentials in them anyway, drop configmaps from the list until that is fixed.

  1. Bind per namespace where you can. RoleBindings keep the grant inside the namespaces the agent should see. Nodes are cluster-scoped, so that one role needs a ClusterRoleBinding.
kubectl create rolebinding ai-sre-read --clusterrole=ai-sre-read --serviceaccount=ai-sre:ai-sre-reader -n payments
kubectl create clusterrolebinding ai-sre-read-nodes --clusterrole=ai-sre-read-nodes --serviceaccount=ai-sre:ai-sre-reader
  1. Verify the grant yourself. Install charts and vendor docs drift. Test the permissions that must be absent.
SA=system:serviceaccount:ai-sre:ai-sre-reader
kubectl auth can-i list secrets --as=$SA -A
kubectl auth can-i get nodes/proxy --as=$SA
kubectl auth can-i create pods/exec --as=$SA -n payments
kubectl auth can-i delete pods --as=$SA -A

Each line should print:

no

Then run kubectl auth can-i --list --as=$SA -n payments and read the full list. Repeat after every upgrade of the agent’s chart.

Scoping an AWS read role that leaves data stores out

  1. Start from ViewOnlyAccess. Skip ReadOnlyAccess: it reads S3 objects and DynamoDB items.
  2. Add the telemetry reads the agent needs by name. ViewOnlyAccess covers listing and basic metadata, so grant log, metric and trace queries explicitly.
  3. Add an explicit deny for data and secret reads. An explicit deny still holds if someone later attaches a broader policy to the same role.
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "TelemetryReads",
      "Effect": "Allow",
      "Action": [
        "cloudwatch:GetMetricData",
        "cloudwatch:ListMetrics",
        "cloudwatch:DescribeAlarms",
        "logs:DescribeLogGroups",
        "logs:FilterLogEvents",
        "logs:StartQuery",
        "logs:GetQueryResults",
        "xray:GetTraceSummaries",
        "xray:BatchGetTraces"
      ],
      "Resource": "*"
    },
    {
      "Sid": "NoDataOrSecretReads",
      "Effect": "Deny",
      "Action": [
        "s3:GetObject",
        "dynamodb:GetItem",
        "dynamodb:BatchGetItem",
        "dynamodb:Query",
        "dynamodb:Scan",
        "secretsmanager:GetSecretValue",
        "ssm:GetParameter",
        "ssm:GetParameters",
        "ssm:GetParametersByPath",
        "kms:Decrypt"
      ],
      "Resource": "*"
    }
  ]
}
  1. Simulate before you trust it.
aws iam simulate-principal-policy --policy-source-arn arn:aws:iam::123456789012:role/ai-sre-reader --action-names s3:GetObject secretsmanager:GetSecretValue logs:StartQuery --query 'EvaluationResults[].[EvalActionName,EvalDecision]' --output text
s3:GetObject                    explicitDeny
secretsmanager:GetSecretValue   explicitDeny
logs:StartQuery                 allowed

Log queries still return log lines, and those can contain personal data. That is the point of the access, so treat it as customer data and apply the controls in the next section.

Controlling where production data goes after the agent reads it

A tight role limits what the agent can see. These controls limit what happens next:

  • Egress: default-deny outbound traffic from wherever the agent runs code, with a short allowlist such as package registries.
  • Credentials: short-lived, issued per task, stored outside the agent’s filesystem and memory.
  • Secret scrubbing: tools that remove recognised secrets from tool output before it reaches the model.
  • Data terms: where inference runs, retention for transcripts and memories, a model-training opt-out, and whether you can bring your own model provider key.
  • Memory hygiene: a way to read, edit and delete what the agent saved, and a rule that any secret it saw gets rotated.

Keeping the read identity separate from any write path

Read access is usually a stage before write access, so design for the second stage now. Give writes a separate identity, so a leaked read credential cannot change anything and your audit logs show reads and writes under different names. The Microsoft team’s production rule is a sensible default: reads run on their own, every write waits for human approval, and deletes are blocked.

The approval gate has to sit outside the agent and catch every write path. In one of their incidents, the agent was told to deallocate VMs only after five safety checks. The logging service went down after the third check, and the agent matched the situation to a past memory where deallocating had been safe and went ahead through a path the approval policy never intercepted. Code changes are the easy case: send them as pull requests, and your branch protection, reviewers and CI stay in charge of the merge.

Read-role mistakes that bite, and the fix for each

  • Attaching ReadOnlyAccess because the name sounds safe. It reads S3 and DynamoDB data. Use ViewOnlyAccess plus named telemetry reads.
  • Wildcards in a ClusterRole. A wildcard over core resources includes Secrets and every future type. List resources by name.
  • Granting nodes/proxy to reach container logs. get on it can run commands in pods. Use pods/log.
  • Leaving a long-lived key where the agent can read it. The agent can reuse or rebuild it. Issue short-lived credentials from outside its runtime.
  • Forgetting that the agent remembers. Transcripts and memories outlive the investigation. Review them, and rotate any secret the agent saw.
  • Checking the role once. AWS updates managed policies and charts change between versions. Re-run kubectl auth can-i and the IAM simulation after every upgrade, and review RBAC bindings on a schedule.
  • Treating a confident answer as a verified one. Read access cannot break production, but the person who acts on a wrong diagnosis can. Require a citation to a query, log line or commit for every claim.

Checklist before you connect a production account

  1. The agent has its own identity in every cluster and cloud account.
  2. Kubernetes roles name resources explicitly and exclude Secrets, nodes/proxy and pods/exec.
  3. AWS roles start from ViewOnlyAccess, add telemetry reads by name and explicitly deny data and secret reads.
  4. kubectl auth can-i and aws iam simulate-principal-policy confirm the negatives.
  5. Outbound network access from the agent’s code sandbox is default-deny.
  6. Credentials are short-lived and never stored in the agent’s runtime.
  7. You know where inference runs, what is retained, and how to opt out of model training.
  8. You can read and delete the agent’s memories.
  9. Any write path uses a separate identity behind human approval, with deletes blocked.
  10. You re-check all of the above after every upgrade.

Read-only by default in Polylane

Cloud accounts start read-only: the agent can call the provider’s read endpoints, and any write call is refused with the reason. AWS connects through a CloudFormation stack that creates a read-only role, Kubernetes clusters connected through the in-cluster agent stay read-only, and background runs nobody has written to stay read-only even after an admin enables writes. A write you ask for in a thread passes a safety review and waits for your confirmation, code changes arrive as pull requests you merge, and admins can opt the workspace out of model training under Privacy & data.

Running on AWS? See how Polylane monitors AWS in production.

Common questions.

Can an AI SRE read Kubernetes Secrets with a read-only role?

Yes, if the role grants get, list or watch on Secrets. The Kubernetes docs note that list and watch reveal Secret contents as well as get, and a list response returns every Secret's values. Leave Secrets out of the role: a pod spec still shows which Secret a container references, which is enough to diagnose a missing reference.

Is the AWS ReadOnlyAccess policy safe to give an AI SRE?

It is broader than the name suggests. AWS states that ReadOnlyAccess can read data in storage services such as S3 buckets and DynamoDB tables, and it includes read access to IAM and Billing. Start from the ViewOnlyAccess job function, which shows resources and basic metadata, then add the log, metric and trace reads the agent needs by name.

Can an agent with only read access still cause harm?

It cannot change resources, but it can disclose them. The main paths are secret values pulled into its context, customer data stored in transcripts or memory, and data sent out if the agent can reach the internet while running code. Watch one Kubernetes exception: get on nodes/proxy can run commands in pods, so in practice it is a write permission.

What should I ask a vendor before connecting production?

Ask where model inference runs, how long transcripts and memories are kept, whether your data trains models and how to opt out, and whether you can bring your own model provider key. Also ask whether the agent can make outbound network requests when it runs code, and whether credentials ever sit inside the agent's runtime.

How do I check what the agent's credentials can actually do?

In Kubernetes, run kubectl auth can-i --list --as=system:serviceaccount:NAMESPACE:NAME, then test negatives such as list secrets and get nodes/proxy, which should print no. In AWS, run aws iam simulate-principal-policy against the role with actions such as s3:GetObject and secretsmanager:GetSecretValue. Repeat after every upgrade, because managed policies and install charts change.

Do logs expose secrets to the agent?

They can. Logs hold whatever your code prints, which is why the Kubernetes docs tell developers to avoid logging secret data in the clear. Treat log read access as access to customer data, fix code that logs tokens, and prefer tools that scrub recognised secrets before they reach the model.

When should an AI SRE get write access?

After it has a track record of correct diagnoses on read access, and then through a separate identity. Keep a human approval on every write, send code changes as pull requests through your normal review, and block deletes. Make sure the approval gate sits outside the agent and catches every write path.

Sources

  1. Polylane documentation
  2. Role Based Access Control Good Practices (Kubernetes)
  3. Good practices for Kubernetes Secrets
  4. AWS managed policies for job functions (AWS IAM User Guide)
  5. Stop restricting the agent. Start restricting its environment (Microsoft Command Line)

About the author

Boris Tane

Founder of Polylane

Boris Tane is the founder of Polylane. He previously founded Baselime, observability for the future of the cloud, which Cloudflare acquired. At Cloudflare he built and led the Workers observability team.

Related

Nobody should be on-call. Polylane watches your infra, finds what broke, and writes the fix.

Get started for free