# Amazon Bedrock Managed Agents (OpenAI preview): setup and limits

> Amazon Bedrock Managed Agents, powered by OpenAI, entered preview on 29 Sep 2026. What it is, how to run a first session, and the preview limits to plan for.

By Boris Tane, Founder of Polylane · Published October 1, 2026 · 13 min read
Canonical: https://polylane.com/learn/ai-in-production/amazon-bedrock-managed-agents-openai-preview/

Amazon Bedrock Managed Agents (BMA), powered by OpenAI, went into public preview on 29 September 2026. AWS and OpenAI built it on a customised version of OpenAI's Agents API, and it runs stateful agents on OpenAI models inside your AWS account, each with its own IAM role and with supported API activity recorded in CloudTrail. You create a session on a regional bedrock-mantle endpoint, connect an execution environment that runs codex exec-server on your own host or on AgentCore Runtime, and send SigV4-signed messages; during preview it is API-only, runs in us-east-1, us-west-2 and us-east-2, and costs nothing extra beyond the inference and AWS resources it uses.

Amazon Bedrock Managed Agents (BMA), powered by OpenAI, is a managed service that runs stateful agents on OpenAI models inside your AWS account. AWS announced the public preview on 29 September 2026. Before that came a limited preview, announced on 28 April 2026 together with OpenAI models and Codex on Bedrock. The service holds the conversation, calls the model on Amazon Bedrock and picks the tools. The commands and tools themselves run on compute you supply: either a host you run or an Amazon Bedrock AgentCore Runtime.

BMA is built for multi-step work, such as investigating a codebase, processing documents or generating files. It suits teams that want each agent to have its own IAM role and leave a CloudTrail record. Because this is a preview, AWS says the features and APIs can change. Everything below comes from the announcement and the Amazon Bedrock user guide as published at launch.

## What Amazon Bedrock Managed Agents is

AWS's product page describes BMA as OpenAI models combined with the Codex harness and Amazon Bedrock AgentCore. You work with six pieces:

- **Session**: a stateful conversation with one agent. It sets the model, the instructions, the tools, an IAM role and an execution environment.
- **Turn**: the work done in response to one message. It can include reasoning, tool calls and generated output.
- **Execution environment**: compute you provide, where commands and local tools run. This is your own host or AgentCore Runtime.
- **Exec server**: the `codex exec-server` process that links the environment to BMA. It makes an outbound connection to the service.
- **Items and events**: items are the durable record of the conversation. The event stream reports progress while work happens.
- **Session role**: an IAM role that BMA assumes to run model inference for you and, when configured, to start an AgentCore Runtime.

The conversation that BMA manages and the files in your execution environment have separate lifecycles. Deleting a session leaves every file on your host or in your S3 buckets, so cleanup is two separate jobs.

## What the 29 September preview adds

According to the announcement, BMA manages how the model keeps state, chooses and uses tools, runs code, and coordinates work across many steps and decisions. Durable sessions keep messages, tool calls and intermediate results. You can come back later, add information and carry on from where the agent stopped. You can add reusable skills for specialised procedures and connect tools, including through Model Context Protocol (MCP) servers. Each agent runs with its own IAM role, supports human approval before consequential actions, and records supported API activity in AWS CloudTrail.

You reach the preview through APIs in three Regions: US East (N. Virginia), US West (Oregon) and US East (Ohio). During the preview, AWS charges nothing extra for BMA itself. You pay for model inference and for the AWS resources your agents use. AWS says pricing may change at general availability.

## How BMA differs from the agent setup you run today

**If you call OpenAI's hosted Agents API.** BMA signs requests with AWS credentials using Signature Version 4 and needs no OpenAI API key. Requests go to a regional `bedrock-mantle` endpoint. AWS warns that OpenAI-hosted Agents APIs and BMA may not offer the same features or change at the same time. It also warns that unknown or unsupported fields can cause a validation error. Port your request bodies one field at a time, checking each against the BMA reference.

**If you call models through Bedrock today.** The preview uses the `bedrock-mantle` endpoint. It does not use `bedrock-runtime`. The IAM actions sit under the same prefix, for example `bedrock-mantle:CreateAgentSession` and `bedrock-mantle:CreateInference`.

**If you wrote your own agent loop.** BMA takes over conversation state and the loop between model and tools. You still own the compute, the workspace, the operating-system user the tools run as, and the authorisation checks inside any tool that changes something outside the workspace.

## Choosing self-hosted compute or AgentCore Runtime

**Self-hosted compute** fits when you already have a development machine, container or compute environment. You provide the host, a workspace directory, network access and a running exec server. The host needs outbound HTTPS and WebSocket access to the BMA endpoint. The example CDK app creates only the IAM roles. It does not create a host.

**AgentCore Runtime** fits when you want managed runtime sessions and storage in your own account. BMA starts the Runtime when there is work, so you never attach an exec server by hand. The example stack creates:

- an ARM64 Runtime and its roles
- a VPC with private subnets, an S3 gateway endpoint and one NAT gateway
- versioned skills and outputs buckets with public access blocked
- S3 Files mounts, plus session storage at `/mnt/workspace`

It sets both the idle timeout and the maximum compute lifetime to 28,800 seconds, which is eight hours. These resources, the NAT gateway included, keep costing money when no turn is running.

Start with self-hosted compute while you learn the API. Move to AgentCore once you want compute that BMA can start by itself.

## What to install and configure first

Both examples need Node.js 20 or later with npm, AWS CLI version 2 (with `aws configure export-credentials`), Bash, `curl` with SigV4 support, `jq`, and Codex CLI 0.154.0 or later, which includes `codex exec-server`. AgentCore also needs Python 3, Docker with Linux ARM64 build support, and the Linux ARM64 Codex binary. You need that binary even when you deploy from a Mac.

Point every shell at the same account and Region:

```bash
export AWS_PROFILE=my-deployment-profile
export AWS_REGION=us-east-1
export AWS_DEFAULT_REGION="$AWS_REGION"
export BMA_REGION="$AWS_REGION"
export BMA_ENDPOINT="https://bedrock-mantle.${BMA_REGION}.api.aws"
aws sts get-caller-identity
```

Check that the account returned is the one you meant. An `AWS_REGION` or `AWS_DEFAULT_REGION` already set in your environment overrides the profile's default Region, so set both explicitly.

The scripts use the model `openai.gpt-5.6-luna` by default. Set `BMA_MODEL` to change it. The model list at `/v1/models` also includes models for other inference APIs, so a model appearing there doesn't prove it works with BMA. You need access to the model, and a session role that is allowed to call it.

## Running a first agent on your own host

Download the BMA example bundle from the user guide and extract it. It contains `self-hosted/`, `acr/` and an empty `bin/` directory. Save the Codex binary for your host as `bin/codex`, then follow these steps.

1. **Deploy the IAM roles.**

```bash
cd self-hosted
npm ci
npx cdk bootstrap
npm run deploy
```

The stack outputs `BmaAccessRoleArn`, the client role your scripts use, and `CustomerInferenceRoleArn`, the session role that BMA assumes. If another stack already owns the role named `BedrockManagedAgentsPreviewInferenceServiceRole`, deploy with `npm run deploy -- --parameters ManageCustomerInferenceRole=false`.

2. **Create a workspace and a session** using a profile that can use the client role.

```bash
export AWS_PROFILE=my-bma-client-profile
export BMA_WORKSPACE_DIRECTORY="$PWD/workspace"
mkdir -p "$BMA_WORKSPACE_DIRECTORY/skills"
./scripts/bma/0.create-session.sh
```

A successful response includes a session ID, an environment ID, the model and the session role ARN. The script saves these to `scripts/bma/.bma-session.env` with owner-only permissions. It stores no AWS access keys in that file.

3. **Attach the exec server.** Open a second terminal, use the same profile and the same directory:

```bash
export AWS_PROFILE=my-bma-client-profile
./scripts/bma/1.attach-exec-server.sh
```

Leave it running until its log confirms the connection. Registration and the WebSocket handshake are signed with SigV4, so knowing a session ID alone gives no access to the environment.

4. **Submit work** from the first terminal:

```bash
./scripts/bma/2.submit-turn.sh "Run: uname -a"
./scripts/bma/3.read-result.sh
```

Look for a command-execution item that shows the command ran on your host. Send a second message to the same session to continue the conversation. In this starter flow, run one turn at a time.

5. **Delete the session** with `./scripts/bma/4.delete-session.sh`, then stop the exec server. Files in the workspace stay on the host until you remove them.

## Reading what the agent did

A session is in one of three states: `idle`, `in_progress` or `failed`. An `idle` session isn't running a turn, but that doesn't prove the task worked. Read the items for the turn and check each command's exit code and output.

Items are the durable record, read with `GET /openai/v1/agents/sessions/SESSION_ID/items`. Use `order=asc` to read them in order, with a `limit` from 1 to 100. When `has_more` is true, pass `last_id` as the `after` cursor for the next page. A command item can include the command, working directory, output, exit code and duration. An MCP item includes the server label, tool name, arguments, result and status.

Events give you the live view. A `GET` on the session's `/events` path returns a stream of server-sent events (SSE). Open the stream before you submit a message, and parse SSE frames as they arrive. A dropped connection doesn't delete the session and may not cancel the work, so check the items after you reconnect.

When you submit a message, you get back an acceptance, possibly with an empty body. Completion comes later. If a network failure leaves you unsure whether the message arrived, read the session's activity before resubmitting, because the service may already have it. To stop a turn, post an `agent.session.input.cancel` event. Cancelling doesn't undo tool actions that have already finished, so check any outside system the agent wrote to.

## Adding skills and MCP tools

A skill is a directory that holds a `SKILL.md` file. It goes under a path listed in the session's `environment.capability_directories`. Add the skill before you create the session:

```bash
mkdir -p "$BMA_WORKSPACE_DIRECTORY/skills/hello-docs"
cat > "$BMA_WORKSPACE_DIRECTORY/skills/hello-docs/SKILL.md" <<'EOF'
---
name: hello-docs
description: Verify the documentation example in this workspace.
---
When asked to run hello-docs, write BMA_DOCS_VERIFIED to verified.txt,
read it back, and return that text.
EOF
```

Skills are plain files. They can ship in a container image, be mounted from storage or sit on a host. Any command a skill calls must exist wherever the skill runs.

Tools come from STDIO MCP servers that run inside the execution environment. Set an array like this one in `BMA_MCP_SERVERS_JSON` before you create the session:

```json
[
  {
    "server_label": "my_server",
    "transport": {
      "type": "stdio",
      "command": "/opt/tools/my-server",
      "args": [],
      "cwd": "/mnt/workspace",
      "env": {},
      "env_vars": []
    },
    "allowed_tools": ["lookup_record"]
  }
]
```

Each label must be unique. `env` passes explicit values, and `env_vars` names variables to inherit from the environment. Forward only what the tool needs, and keep credentials out of both: configuration can be stored with the session or printed in logs. Have the tool validate the arguments the model generates. On AgentCore, the MCP executable and its dependencies must be inside the container image, referenced by absolute container paths.

## Scoping the three identities

A session involves three identities, and AWS recommends a separate IAM role for each one.

1. **The caller** signs session and event requests. It needs the `bedrock-mantle` session actions and `iam:PassRole` on the exact session role, with the condition that `iam:PassedToService` equals `bedrock-mantle.amazonaws.com`. Setting `BMA_INFERENCE_ROLE_ARN` to a different role doesn't give the caller permission to pass it. A self-hosted exec server also needs `RegisterEnvironment` and `ConnectEnvironment`.
2. **The session role** is the one BMA assumes. It must be in the same account. Its trust policy allows the `bedrock-mantle.amazonaws.com` principal, with `aws:SourceAccount` and `aws:SourceArn` conditions, and its permissions allow `bedrock-mantle:CreateInference`. That action has no model-specific resource ARN, so narrow it with a `bedrock-mantle:Model` condition. For AgentCore, add `InvokeAgentRuntime` and `StopRuntimeSession`, scoped to your Runtime and its endpoint ARN.
3. **The execution environment** carries the real risk. Commands and MCP tools run with whatever access the environment has. Give it a dedicated workspace and a restricted operating-system user, plus only the files, credentials, network destinations and tools the task needs. Keep deployment credentials off it.

AWS says to treat instructions that arrive from documents, websites and tool responses as untrusted. For any action that changes something outside the workspace, enforce authorisation and any required human review inside your application or tool. The example session role covers every scenario in the bundle, so review its generated policy and narrow it before a real workload uses it.

I built and led the Workers observability team at Cloudflare, and on any new agent runtime I'd log the request ID, session ID and Region of every call from day one. AWS support asks for exactly those, along with the UTC time, HTTP status, model ID, exec-server version and environment type.

## Preview limits to plan around

- Only the `bedrock-mantle` endpoint works, and cross-Region inference profiles aren't supported.
- Text is the only documented input type.
- Subagents and programmatic tool calling (code mode) aren't supported. Leave both out of the session config.
- There are no dedicated APIs to list or fetch turns. Link work together using events and each item's `turn_id`.
- There is no built-in long-term memory. Memory shared across sessions needs a datastore you provision yourself.
- You can't use a customer-managed KMS key for the session data BMA stores.
- There is no console. You work through the APIs and the example stacks.

The example uses item pages of 1 to 100 (20 by default), at most 32 capability-directory entries, and polls for results every 2 seconds with a 300-second timeout. Account request limits, model capacity and token limits also apply. AWS adds that a field being accepted doesn't mean the feature behind it is supported, so test anything outside the documented API.

## What goes wrong in the first week

- **HTTP 400 when you create a session.** The request needs `agent.model`, `environment` and a `role_arn` from the same account. Scripts that name the role without passing its ARN fail here. So does any subagent config.
- **Access denied after a deploy that worked.** The identity that deploys and the identity that calls the API are different. Check `AWS_PROFILE` in every terminal, then confirm the caller can pass that exact role.
- **HTTP 404.** Check the Region, endpoint and path. Sessions live under `/openai/v1/agents/sessions` and models under `/v1/models`. A 404 can also mean your account doesn't have preview access.
- **The exec server never connects.** The binary must match the host's operating system and architecture and be version 0.154.0 or later. A running process doesn't prove a connection. Wait for the acknowledgement in its log.
- **The event stream stays silent.** It's a live SSE connection. Open it first, submit work from a second terminal, and read finished output from the items.
- **A skill or MCP server is missing.** Check the path inside the real host or container. On AgentCore, wait for S3 Files to sync. After changing the image, create a new session, because a running session keeps its old processes.
- **CloudFormation reports an existing IAM role.** Both examples use the same default role name. Deploy the second one with `ManageCustomerInferenceRole=false`.
- **A forgotten AgentCore stack keeps billing.** Download any outputs, delete the session, then run `npm run destroy` with your deployment profile. The stack's storage resources use destructive removal policies, so their objects are deleted with the stack.

## A checklist before a BMA agent touches real work

- You call one of the three preview Regions, and you sign for `bedrock-mantle` in that same Region.
- The caller, the session role and the execution identity are separate, and the inference permission is pinned to the models you use.
- The execution host runs as a restricted user in a dedicated workspace, with no deployment credentials.
- MCP servers expose only the tools in `allowed_tools`, and each tool validates its arguments and checks authorisation.
- Your client reads results from items, checks exit codes and checks session activity before retrying a submission.
- Example stacks you aren't using are destroyed, and you have a plan for when GA pricing arrives.

Running on AWS? See [how Polylane monitors AWS in production](https://polylane.com/for/aws/).

## Common questions

**When was Amazon Bedrock Managed Agents announced, and is it generally available?**

AWS announced the public preview on 29 September 2026. A limited preview was announced on 28 April 2026, together with OpenAI models and Codex on Bedrock. It is still in preview, and AWS says preview functionality and APIs can change.

**What does BMA cost during the preview?**

BMA itself has no extra charge during the preview. You pay for model inference and the AWS resources your agents use. The AgentCore example creates a NAT gateway and storage that keep billing while idle. AWS says pricing may change at general availability.

**Which Regions and endpoints does the preview support?**

US East (N. Virginia), US West (Oregon) and US East (Ohio), at https://bedrock-mantle.REGION.api.aws for us-east-1, us-west-2 and us-east-2. Sign requests with that endpoint's Region and the service name bedrock-mantle. Cross-Region inference profiles aren't supported in this preview.

**Which OpenAI models can a BMA agent use?**

The example scripts use openai.gpt-5.6-luna by default, and you can override it with BMA_MODEL. GET /v1/models lists the endpoint's catalogue, but AWS notes that a model appearing there doesn't prove it works with BMA. You also need access to the model and a session role allowed to call bedrock-mantle:CreateInference for it.

**Do I need an OpenAI API key to use BMA?**

No. Requests are signed with AWS Signature Version 4, using credentials from an AWS profile or the standard credential provider chain. If you use temporary credentials, the signature must include the session token, and the supplied shell and Python examples handle that for you.

**Can I reuse code written for OpenAI's hosted Agents API?**

Expect to adapt it. AWS says OpenAI-hosted Agents APIs and BMA may not offer the same features or change at the same time. Unknown or unsupported fields can cause a validation error. Session operations live under /openai/v1/agents/sessions on the bedrock-mantle endpoint.

**Does the preview support subagents, code mode or long-term memory?**

No. Subagents and programmatic tool calling (code mode) aren't supported, and AWS says not to enable them in session config. A session keeps context across its own turns, but memory shared across sessions needs a datastore you provision and authorise yourself.

**Is there a console for Bedrock Managed Agents?**

Not in the preview. You work through the REST API and the example bundle, which includes shell scripts and a Python client for creating sessions, submitting turns, reading items and deleting sessions.

## Sources

- [Amazon Bedrock Managed Agents, powered by OpenAI, is now available in preview (AWS What's New)](https://aws.amazon.com/about-aws/whats-new/2026/09/bedrock-managed-agents-preview/)
- [Amazon Bedrock now offers OpenAI models, Codex, and Managed Agents (Limited Preview)](https://aws.amazon.com/about-aws/whats-new/2026/04/bedrock-openai-models-codex-managed-agents/)
- [Amazon Bedrock Managed Agents product page](https://aws.amazon.com/bedrock/managed-agents-openai/)
- [Amazon Bedrock Managed Agents, powered by OpenAI (preview): User Guide](https://docs.aws.amazon.com/bedrock/latest/userguide/bedrock-managed-agents-openai.html)
- [Set up permissions and prerequisites](https://docs.aws.amazon.com/bedrock/latest/userguide/bedrock-managed-agents-openai-prerequisites.html)
- [Run your first agent with self-hosted compute](https://docs.aws.amazon.com/bedrock/latest/userguide/bedrock-managed-agents-openai-self-hosted.html)
- [Run your agent on Amazon Bedrock AgentCore Runtime](https://docs.aws.amazon.com/bedrock/latest/userguide/bedrock-managed-agents-openai-agentcore-runtime.html)
- [Work with sessions, events, and results](https://docs.aws.amazon.com/bedrock/latest/userguide/bedrock-managed-agents-openai-sessions.html)
- [Add skills and tools](https://docs.aws.amazon.com/bedrock/latest/userguide/bedrock-managed-agents-openai-skills-tools.html)
- [Security and IAM roles](https://docs.aws.amazon.com/bedrock/latest/userguide/bedrock-managed-agents-openai-security.html)
- [BMA preview REST API reference](https://docs.aws.amazon.com/bedrock/latest/userguide/bedrock-managed-agents-openai-api-reference.html)
- [Preview availability and limitations](https://docs.aws.amazon.com/bedrock/latest/userguide/bedrock-managed-agents-openai-quotas-limitations.html)
- [Troubleshooting](https://docs.aws.amazon.com/bedrock/latest/userguide/bedrock-managed-agents-openai-troubleshooting.html)

## About the author

Boris Tane is the founder of Polylane. He previously founded Baselime, observability for the future of the cloud, which Cloudflare acquired. At Cloudflare he built and led the Workers observability team.

## Related

- [How to keep AI coding agents from breaking production](https://polylane.com/learn/ai-in-production/how-to-keep-ai-coding-agents-from-breaking-production/): Stop AI coding agents breaking production: scoped credentials, branch protection they can't bypass, required checks that block, and production-aware review.
- [Is It Safe to Give an AI SRE Read Access to Production?](https://polylane.com/learn/ai-in-production/is-it-safe-to-give-an-ai-sre-read-access-to-production/): Read access for an AI SRE is a sound first step if you scope it. The risks are secrets, customer data and exfiltration. Here are the controls and a checklist.
- [How to Use Claude Code to Debug Production Issues](https://polylane.com/learn/ai-in-production/how-to-use-claude-code-to-debug-production-issues/): Debug production issues with Claude Code: get logs and traces into the session, test hypotheses against evidence, stay read-only and ship a verified fix.
- [Turn your app into a context graph](https://polylane.com/blog/turn-your-app-into-a-context-graph/): Agents that run software need a context graph of the app: every resource, what it connects to, the repository that deploys it and the team that owns it. How we built one on Durable Objects, keep it fresh across every provider without polling AWS, decide what counts as a change, and delete from it safely.
- [How to attach Vercel Sandbox to a Secure Compute network](https://polylane.com/learn/deployment-safety/vercel-sandbox-secure-compute-network-attachment/): Vercel Sandbox now supports Secure Compute. Attach a sandbox to your network for static egress IPs and VPC peering, with steps, costs and gotchas.

Get started with one command: `curl -fsSL https://polylane.com/setup | bash` installs the CLI, connects your coding agents, and creates the account.
