Get started Dashboard
Monitoring coverage ·

Cloudflare Basin: a guide to the serverless data platform

Explore with AI

Cloudflare Basin is a serverless analytics platform that Cloudflare made generally available on 1 October 2026, renaming its Data Platform. Basin Pipelines ingests events from Workers, HTTP or Logpush and transforms them with SQL, a managed Iceberg catalogue on R2 stores and maintains the tables, and Basin SQL queries them with no clusters to run. You pay only when it ingests, processes or queries data, and any Iceberg engine such as DuckDB, Snowflake or Spark can read the same tables without egress fees.

On this page

Cloudflare Basin is Cloudflare’s serverless analytics platform. It became generally available on 1 October 2026 and puts three products on one path. Basin Pipelines takes in events and transforms them with SQL. A managed Apache Iceberg catalogue stores them as tables in R2 and keeps those tables healthy. Basin SQL runs distributed queries over them. You don’t size clusters and there are no hourly charges: you pay when Basin ingests, processes or queries data.

If you used the Cloudflare Data Platform beta, nothing breaks. Cloudflare renamed the products, and the docs say existing resources and configurations keep working. This guide covers what Basin is for and how it differs from the tools people confuse it with. It then walks through a first pipeline you can ship from a Worker today, and the limits and gotchas that Cloudflare’s announcement and docs state.

What Cloudflare Basin is, and what changed on 1 October

Cloudflare announced the Data Platform during Birthday Week 2025. On 1 October 2026 it went GA under the Basin name, and each product got a new name:

  • Basin Pipelines, formerly Cloudflare Pipelines, receives events from Workers bindings, HTTP endpoints or Cloudflare Logpush. It transforms them with SQL and writes Iceberg tables, or JSON and Parquet files, to R2.
  • The Basin catalogue, formerly R2’s data catalogue product, is a managed Iceberg REST catalogue. It tracks table metadata and runs maintenance such as compaction and snapshot expiration.
  • Basin SQL, formerly R2 SQL, is a serverless, distributed, read-only SQL engine for Iceberg tables. It splits each query into smaller tasks and spreads them across Workers.

Two dates matter for budgets. Billing for all three products switched on for non-enterprise accounts on 3 August 2026, so any usage beyond the free allowances already appears on invoices. Then on 1 October, the per-stream ingest limit for Basin Pipelines rose from 5 MB/s to 1 GB/s.

1 GB/s
maximum ingest rate per Basin Pipelines stream
Up from 5 MB/s, per the 1 October 2026 Cloudflare changelog and the Basin Pipelines limits page.

The announcement post quotes up to 3GB/s per stream. The changelog and the limits page both list 1 GB/s, so plan around the limits page.

Who Basin suits

Basin fits teams whose events already pass through Cloudflare, and teams that want analytics data in an open format they control. The announcement and docs point to these uses:

  • Product and request analytics from Workers. A Worker sends a small JSON record per request through a binding, and a few minutes later you can answer questions about it with SQL.
  • Cloudflare logs you want to keep and query. Logpush can feed a pipeline that filters and reshapes HTTP logs with SQL, then stores them as compressed Parquet files or Iceberg tables. Dropping fields at ingestion cuts storage and keeps sensitive values out of the table.
  • Telemetry from the services your Worker calls. Cloudflare’s prompting guide shows records for D1 operations, Queue processing and R2 object activity. It warns not to assume a service has a direct Basin source, and recommends instrumenting the Worker that uses them.
  • Data teams that care about portability. Any Iceberg-compatible engine, including PyIceberg, DuckDB, Snowflake and Apache Spark, can read and write the same tables, and R2 charges no egress.

One customer quoted in the announcement replaced an AWS S3 and Athena setup with Pipelines, the catalogue and Basin SQL. That is the kind of workload Basin targets: event data in object storage with a query engine on top, and no infrastructure for you to run.

I built and led the Workers observability team at Cloudflare, and the first table I would build on Basin is plain request telemetry: the route, status and duration of each request, queryable with SQL. The setup below starts with a smaller version of exactly that.

How Basin differs from the tools you use today

Pipelines SQL shapes data on the way in, Basin SQL analyses it afterwards

Basin has two separate SQL layers, and mixing them up is the most common design mistake. Pipelines SQL runs on each event as it arrives. Every pipeline is a statement of the form INSERT INTO sink SELECT ... FROM stream, which you use to filter, reshape, cast and route records. These transforms are stateless today. Streaming aggregations and joins are on the roadmap. Basin SQL runs once the data has landed in an Iceberg table, and that is where GROUP BY, joins, window functions and reports belong.

An R2 sink and an Iceberg table are separate destinations

A pipeline can write plain JSON or Parquet files to an R2 bucket, or it can write Iceberg tables through the catalogue. Basin SQL can only query the Iceberg tables. If you plan to query with Basin SQL, choose the Iceberg sink when you create the pipeline.

Your tables stay readable by other engines

Basin stores data as Iceberg tables in your R2 bucket and exposes them through an Iceberg REST catalogue. Basin SQL is one engine that can query them. DuckDB, PyIceberg, Snowflake and Spark can query the same tables, and since July 2026 the catalogue has accepted read-only API tokens for connecting engines. R2 has no egress fees, so reading the data from another cloud or region carries no transfer cost.

Billing follows usage

Pipelines bills for SQL transforms and sink delivery. The catalogue bills for metadata operations and compaction. Basin SQL bills for the compressed data it scans. Standard R2 storage and operations charges apply on top of all three. Ingress into a stream is free.

Set up your first Basin pipeline from a Worker

These steps follow Cloudflare’s CLI guide. A Worker records the path of each request, Basin Pipelines writes it to an Iceberg table, and Basin SQL counts requests by path.

Before you start, you need:

  • A Cloudflare account on Workers Paid with R2 enabled. Basin SQL and the catalogue also need an active R2 subscription and a linked payment card.
  • Node.js and npm.
  • An R2 Account API token with Admin Read & Write permission. Create it under R2 object storage, Overview, Manage API tokens. The setup command asks for it, and Basin SQL uses the same token.
  1. Create a Hello World Worker and move into its directory:
npm create cloudflare@latest -- hello-world-pipeline --type=hello-world --lang=js --no-deploy --no-git --accept-defaults
cd hello-world-pipeline
npx wrangler login
  1. Run the interactive setup to create a stream, a sink and a pipeline:
npx wrangler basin pipelines setup --name request_events

In the prompts:

  • Turn off the HTTP endpoint, because the Worker sends events through a binding.
  • Build the schema with one required string field named path.
  • Choose the Iceberg data catalogue destination, then Advanced. Enter a unique R2 bucket name, such as request-events- plus a suffix. Use namespace default and table request_events.
  • Paste the R2 Account API token. Keep the default compression and the 100 MB file size, and set the roll interval to 60 seconds.
  • Confirm, then select Simple ingestion for the pipeline SQL.

Write down the bucket name and the stream ID from the output. If you lose the stream ID, list your streams and look for request_events_stream:

npx wrangler basin pipelines streams list
  1. Add the binding to the root object of wrangler.jsonc. The binding needs the stream ID, and the pipeline ID won’t work here. The older pipeline field is deprecated, so use stream:
{
  "pipelines": [
    {
      "binding": "EVENTS",
      "stream": "<STREAM_ID>"
    }
  ]
}
  1. Send one record per request from src/index.js. The Worker still returns its normal response:
export default {
	async fetch(request, env, ctx) {
		await env.EVENTS.send([{ path: new URL(request.url).pathname }]);
		return new Response("Hello World!");
	},
};

The binding handles authentication for ingestion, so keep the R2 token out of Worker code and wrangler.jsonc.

  1. Deploy, then open the URL Wrangler prints at / and at /example:
npx wrangler deploy
  1. Look up the warehouse name, set the SQL token and run a query:
npx wrangler basin catalog get <YOUR_BUCKET>
export WRANGLER_BASIN_SQL_AUTH_TOKEN="<YOUR_R2_ACCOUNT_API_TOKEN>"
npx wrangler basin sql query "<WAREHOUSE_NAME>" "SELECT path, COUNT(*) AS requests FROM default.request_events GROUP BY path ORDER BY requests DESC"

You should see rows for / and /example with their counts, and possibly /favicon.ico too. If the table is missing or the query returns nothing, wait a few minutes for the first file to be written and run it again.

If you use TypeScript, wrangler types generates types from the stream’s schema. That catches missing fields and type mismatches before you deploy. When you outgrow the interactive setup, Terraform resources cover the catalogue, the stream, the sink and the SQL that connects them.

Checking that events reach the table

A successful send() in the Worker only tells you the stream accepted the event. It doesn’t tell you a row reached the table. Events that don’t match the stream schema are accepted at ingestion and then dropped during processing, so check each stage:

  1. In the dashboard, go to Basin Pipelines, Pipelines, select the pipeline, and read the Metrics tab for records read and for records and files written.
  2. Open the Errors tab to see dropped events. Basin sorts them into four types: missing_field, type_mismatch, parse_failure and null_value.
  3. Run a small SELECT ... LIMIT 10 with Basin SQL.

The same data is available from the GraphQL Analytics API, which makes it easy to alert on dropped events. This query groups them by error type:

query GetPipelineUserErrors(
	$accountTag: String!
	$pipelineId: String!
	$datetimeStart: Time!
	$datetimeEnd: Time!
) {
	viewer {
		accounts(filter: { accountTag: $accountTag }) {
			pipelinesUserErrorsAdaptiveGroups(
				limit: 100
				filter: {
					pipelineId: $pipelineId
					datetime_geq: $datetimeStart
					datetime_leq: $datetimeEnd
				}
				orderBy: [count_DESC]
			) {
				count
				dimensions {
					date
					errorFamily
					errorType
				}
			}
		}
	}
}

The pipelinesUserErrorsAdaptive dataset adds detailed error descriptions for the last 24 hours. On a busy pipeline it can return a lot of data.

Keeping Iceberg tables fast with compaction and snapshot expiry

Every Iceberg write creates new files and a new snapshot. Without maintenance, queries have to open more and more small files, metadata grows, and query planning slows down. The catalogue automates two maintenance jobs.

Compaction merges small Parquet files into larger ones. The merged files appear with a compacted- prefix in the table’s /data/ directory. You can turn it on for a whole catalogue or for a single table:

npx wrangler basin catalog compaction enable my-bucket --target-size 128 --token $R2_CATALOG_TOKEN
npx wrangler basin catalog compaction enable my-bucket my-namespace my-table --target-size 256

The target size can be anywhere from 64 MB to 512 MB. The docs suggest 64 to 128 MB for latency-sensitive work, 128 to 256 MB for streaming ingest, and 256 to 512 MB for OLAP queries that scan a lot of data.

Snapshot expiration deletes old snapshots, along with any data files that only those snapshots referenced. A snapshot is removed only when it is older than --older-than-days (default 30) and also falls outside the most recent --retain-last snapshots (default 5):

npx wrangler basin catalog snapshot-expiration enable my-bucket --token $R2_CATALOG_TOKEN --older-than-days 7 --retain-last 10

Snapshot expiration is free. Compaction is billed per GB and per object processed. Each table has a Maintenance tab in the dashboard that shows the schedule, the next eligible time and a history of runs. Queue maintenance on that tab requests a compaction by hand. A queue request can fail with 40903 when an executor conflict blocks it, or with 42901 once the daily limit of accepted requests is reached.

What Basin SQL will and won’t run

Basin SQL covers most analytical SQL. It supports every standard join type, CTEs, subqueries, window functions with QUALIFY, GROUPING SETS, ROLLUP and CUBE, set operations, and more than 190 scalar and aggregate functions, including JSON functions and approximate aggregates. You can run it from Wrangler, from the API, or in the dashboard editor, which has autocomplete, a table browser, query plans and exportable results.

It also has hard limits:

  • It is read-only. INSERT, UPDATE, DELETE, CREATE, DROP and ALTER all fail with only read-only queries are allowed. Data gets into tables through Pipelines.
  • Some syntax is unsupported: OFFSET, a named WINDOW clause, UNNEST, PIVOT, UNPIVOT, LATERAL, nested parenthesised joins, and PERCENTILE_DISC (use PERCENTILE_CONT).
  • NOT IN on a nullable column is unsupported. Rewrite it as NOT EXISTS with a correlated subquery.
  • Only Parquet can be queried. CSV and JSON files can’t.
  • Some functions are budget-gated. MEDIAN, PERCENTILE_CONT, ARRAY_AGG, STRING_AGG, any aggregate with DISTINCT, and window functions get a memory estimate before the query runs. If too much data would be scanned, the query is rejected with a 400. Heavy joins and high-cardinality GROUP BY queries can also time out.
  • now() is quantised to 10 ms boundaries and always returns UTC.

The same fixes work almost every time:

  • Filter on a time range in WHERE.
  • Name the columns you need.
  • Add LIMIT.
  • On large datasets, use approx_distinct, approx_median and approx_percentile_cont.
  • Read the plan with EXPLAIN.

Error 40003 means invalid syntax. 40004 means an invalid query, such as an unknown column. 80001 is an edge connection failure and is worth retrying.

Mistakes that stall a first Basin rollout

  • The binding uses the pipeline ID. Workers bind to a stream. Fix: copy the stream ID from the setup output, or find it with wrangler basin pipelines streams list.
  • The R2 token ends up in the Worker. Fix: take it out. The binding authenticates ingestion. You only need the token for setup and for WRANGLER_BASIN_SQL_AUTH_TOKEN.
  • The table has no files yet. Fix: wait until the roll interval passes and the first file is written, then run the query again.
  • Rows go missing and the Worker shows no error. Events that don’t match the schema are dropped after ingestion. Fix: read the Errors tab or the GraphQL user error dataset, and use typed bindings.
  • The pipeline writes to a plain R2 sink and you expect to query it with SQL. Fix: create an Iceberg sink through the catalogue.
  • The schema is left until later. The announcement lists schema migrations and updatable configuration and Pipelines SQL as planned work. Fix: settle the schema before you create the stream.
  • Pagination uses OFFSET. Fix: page on a time range or a key in WHERE, with LIMIT.
  • You expect orphaned files to clean themselves up. Maintenance never removes files that no snapshot ever referenced. Fix: keep an eye on bucket size and remove leftover files yourself.
  • Snapshot retention is short on audit tables. Expiry deletes the history you would use for time travel. Fix: the docs suggest 30 to 90 days and 50 or more snapshots for compliance tables.

Account limits to plan around

Each account can have up to 20 streams, 20 sinks and 20 pipelines in Basin Pipelines. A single ingestion request can carry at most 5 MB, and each stream can ingest up to 1 GB/s. You can ask for higher limits through a form linked from the limits page. Compaction only handles Parquet data files.

What a month of Basin costs

The prices below come from Cloudflare’s pricing page for each product. Included usage resets every month.

Basin prices by product, from Cloudflare's pricing docs
ProductWhat you pay forIncluded each monthPrice beyond that
Basin Pipelines Stream ingress Unlimited Free
Basin Pipelines SQL transforms 50 GB $0.04 per GB
Basin Pipelines Sink to R2 as JSON 50 GB $0.03 per GB
Basin Pipelines Sink as Parquet or Iceberg 50 GB $0.06 per GB
Basin catalogue Metadata operations 1 million $9.00 per million
Basin catalogue Compaction data processed 10 GB $0.005 per GB
Basin catalogue Compaction objects processed 1 million $2.00 per million
Basin catalogue Snapshot expiration Free Free
Basin SQL Compressed data scanned 10 GB $0.0025 per GB, 10 MB minimum per query

The docs give two worked examples:

  • A pipeline that ingests 500 GB a month and filters it down to 300 GB written to an Iceberg table costs $33.00. That is $18.00 for transforms beyond the 50 GB allowance plus $15.00 for the Iceberg sink.
  • Storing 500 GB of Parquet data and scanning 50 GB of it with Basin SQL costs $7.45. That is $7.35 for R2 storage plus $0.10 for scanning.

The docs note that storage is usually the largest cost on big streaming workloads. Basin SQL bills at least 10 MB per query. Metadata commands such as EXPLAIN, SHOW and DESCRIBE scan no data and are free, and failed queries aren’t charged.

Building Basin pipelines with a coding agent

Cloudflare’s docs include a prompting guide for agents that build Basin workflows. It suggests connecting the Cloudflare documentation MCP server so the agent reads current commands, binding fields and limits. The Cloudflare Observability MCP server can help it inspect Worker logs and exceptions. The example prompts tell the agent to keep records small and to leave out request bodies, tokens, client IP addresses and SQL parameters.

~/app
$ claude
✳ 12 files · main · last session 2h ago
> add a Basin Pipelines stream for route, status and duration to this Worker
? for shortcuts
Point the agent at the Cloudflare docs MCP server before it creates any resources.

Check the agent’s work the same way you would check your own: confirm the event reached the stream, confirm the Iceberg sink wrote files, and confirm a LIMIT 10 query returns rows.

What Cloudflare says is coming next

The announcement lists this work as planned. None of it is available today:

  • Basin Pipelines: custom partitioning, schema migrations, updatable configuration and Pipelines SQL, Iceberg V3 with the Variant type, and stateful processing for streaming aggregations, joins and incrementally updated materialised views.
  • The catalogue: compaction that sorts and clusters data, finer auth controls for namespaces and tables, and jurisdiction support for data sovereignty.
  • Basin SQL: advanced statistics and adaptive scheduling, full DDL, and Iceberg V3 types including VARIANT and geospatial.

Running on Cloudflare? See how Polylane monitors Cloudflare in production.

Common questions.

What happened to Cloudflare Pipelines, R2's data catalogue and R2 SQL?

Cloudflare renamed them on 1 October 2026, when the Data Platform went GA as Basin. Cloudflare Pipelines is now Basin Pipelines, R2's data catalogue product is now the Basin catalogue, and R2 SQL is now Basin SQL. The docs say existing resources and configurations keep working.

Do I need to migrate my existing pipelines or tables?

No. The Basin overview page says existing resources and configurations continue to work after the rename. The one change to make is in Worker bindings: use the stream field in wrangler.jsonc, because the older pipeline field is deprecated.

What plan do I need to use Basin?

The CLI guide asks for a Cloudflare account on Workers Paid with R2 enabled. The pricing pages for Basin SQL and the catalogue say each also needs an active R2 subscription and a linked payment card. For setup and queries you need an R2 Account API token with Admin Read & Write permission.

How much data can a Basin Pipelines stream take?

Each stream can ingest up to 1 GB/s, up from 5 MB/s before 1 October 2026. Each ingestion request can carry at most 5 MB, and an account can have up to 20 streams, 20 sinks and 20 pipelines. You can request higher limits through Cloudflare's limit increase form.

Can I query Basin tables with DuckDB, Snowflake or Spark?

Yes. Basin stores Iceberg tables in R2 behind an Iceberg REST catalogue, and the announcement names PyIceberg, DuckDB, Snowflake and Apache Spark as compatible engines. Since July 2026 the catalogue has accepted read-only API tokens for connecting query engines, and R2 charges no egress.

Can Basin SQL insert data or create tables?

No. Basin SQL is read-only, so INSERT, UPDATE, DELETE, CREATE, DROP and ALTER fail with an error saying only read-only queries are allowed. Data gets into tables through Basin Pipelines, and full DDL support is on Cloudflare's roadmap.

Why is my Basin table empty after the Worker sends events?

The first Iceberg file isn't written until the roll interval passes, so wait a few minutes and query again. If rows are still missing, open the Errors tab for the pipeline. Events that fail the schema check (missing_field, type_mismatch, parse_failure or null_value) are accepted at ingestion and dropped during processing.

How is Basin SQL billed?

Basin SQL charges $0.0025 per GB of compressed data scanned, which is $2.50 per TB, after 10 GB included each month. Every query is billed for at least 10 MB. EXPLAIN, SHOW and DESCRIBE scan no data and are free, and failed queries aren't charged.

Sources

  1. Introducing Cloudflare Basin: an open, serverless data platform, now generally available (Cloudflare Blog)
  2. Basin overview (Cloudflare Docs)
  3. Basin get started with the CLI (Cloudflare Docs)
  4. Prompt an agent to build with Basin (Cloudflare Docs)
  5. Basin Pipelines ingest limit increased to 1 GB/s (Cloudflare Changelog)
  6. Basin Pipelines pricing (Cloudflare Docs)
  7. Basin Pipelines limits (Cloudflare Docs)
  8. Basin Pipelines metrics and analytics (Cloudflare Docs)
  9. Basin SQL pricing (Cloudflare Docs)
  10. Basin SQL limitations and best practices (Cloudflare Docs)
  11. Pricing for the Basin catalogue (Cloudflare Docs)
  12. Table maintenance for the Basin catalogue (Cloudflare Docs)

About the author

Boris Tane

Founder of Polylane

Boris Tane is the founder of Polylane. He previously founded Baselime, observability for the future of the cloud, which Cloudflare acquired. At Cloudflare he built and led the Workers observability team.

Related

Nobody should be on-call. Polylane watches your infra, finds what broke, and writes the fix.

Get started for free