Skip to Content

Cloudflare Pipelines + R2 Data Catalog Terraform 2026: Complete Apache Iceberg Pipeline Setup

Terraform setup for Cloudflare Pipelines, R2 Data Catalog, and Apache Iceberg tables
2026-05-05 21:27:47 Updated 2026-08-21 15:34:12.511586 — min read 233 views
Cloudflare Pipelines + R2 Data Catalog Terraform 2026: Complete Apache Iceberg Pipeline Setup
“Cloudflare Pipelines and R2 Data Catalog Terraform 2026 lets teams declare streams, SQL pipelines, sinks, and Apache Iceberg tables as infrastructure. The setup can route JSON events into Parquet-backed Iceberg tables, then query them with R2 SQL or other engines. Terraform improves repeatability, but schemas, credentials, costs, and failures still need operational ownership.

What You'll Learn

  • What Cloudflare Pipelines and R2 Data Catalog each contribute to a streaming data stack.
  • How Terraform resources connect buckets, streams, sinks, SQL pipelines, and scoped credentials.
  • Why Apache Iceberg output uses Parquet and how table features differ from raw R2 objects.
  • Which tests, limits, permissions, and cost checks belong before a production rollout.

What Terraform Support Adds in 2026

Cloudflare Pipelines and R2 Data Catalog Terraform 2026 support is an infrastructure-as-code path for a data pipeline that spans ingestion, SQL transformation, storage, and table management. Cloudflare’s May 4, 2026 changelog says Pipelines and R2 Data Catalog can be created and managed with Cloudflare Terraform provider version 5.19.0.

The practical change is not that Terraform turns a data platform into a single command. It gives a team a declared configuration for the resources that make the path work. A review can compare the proposed stream, sink, table, SQL, and permissions before an apply. A later run can reconcile the declared state with the account.

The official Cloudflare Terraform support announcement identifies four new resources for the Pipelines and catalog workflow. They cover the catalog, the input stream, the output sink, and the SQL pipeline that connects them. The full Terraform example also uses an R2 bucket and an account token resource.

That distinction matters when reviewing the old article’s “zero operational overhead” wording. Terraform reduces manual configuration and makes changes reviewable. It does not decide whether an event schema is correct, whether a token is over-permissioned, whether a small-file problem will affect queries, or whether an update can safely change a table.

Cloudflare’s current documentation is dated in 2026 and may change while the feature matures. Pin the provider range that your team has tested, read the resource documentation during upgrades, and keep the generated plan in the same review process as application code.

For adjacent context on Cloudflare’s storage and retrieval products, see Current Affair’s vector database comparison. A catalog table is a managed data asset, not merely another object path.

The Data Flow From Event to Iceberg Table

A useful mental model starts with the event and follows it to the query. A stream receives data through an HTTP endpoint or a Worker binding. A pipeline applies SQL to that stream. A sink writes the transformed result to R2 Data Catalog as an Apache Iceberg table. A query engine then reads table metadata and data files.

Each layer has a different responsibility. The stream defines how data enters and what shape the input is expected to have. The pipeline defines which fields are selected or transformed. The sink defines where the output goes and how it is written. The catalog provides table metadata and a namespace for Iceberg tables. The query engine reads the resulting table through its supported connection.

The Cloudflare Pipelines documentation describes Pipelines as a service that can ingest streaming data, transform it with SQL, and write to destinations such as R2 Data Catalog or R2. The destination choice should follow the query requirement. If the target is a managed Iceberg table, use the catalog sink. If the target is a raw file archive, an R2 sink may be more appropriate.

LayerPrimary responsibilityReview question
StreamReceive events and validate the input schemaCan producers send data in the declared shape?
PipelineApply SQL and connect input to outputDoes the transformation preserve the required fields?
SinkWrite processed data to R2 or R2 Data CatalogIs the format and destination correct?
CatalogManage the Iceberg table namespace and metadataCan approved engines discover and read the table?
Query engineRead and analyze the tableAre freshness, schema, and permissions visible to users?

The flow also defines failure boundaries. An event can fail schema validation before it reaches SQL. A SQL statement can select a field that no longer exists. A sink can reject credentials or a table configuration. A table can be present while recent data is still rolling into a file. Monitoring should identify which boundary failed instead of reporting only that “the pipeline is down.”

Current Affair’s operating-system layer analysis makes a similar point for agent systems. Clear boundaries make it easier to test and repair a system than a single label that hides several components.

Why Apache Iceberg Matters for R2 Data

R2 Data Catalog is useful here because the sink writes processed data as Apache Iceberg tables rather than treating every output as an unrelated file. Iceberg is a table format with metadata that tracks snapshots, schema, and the relationship between data files. Cloudflare’s sink documentation says Iceberg tables provide ACID transactions, schema evolution, and time travel capabilities for analytics workloads.

These features address problems that appear when files are written continuously. A query needs to know which files belong to a table version. A schema change needs a controlled interpretation. A correction may require a new snapshot rather than an opaque replacement of objects. Time travel can help a team inspect an earlier table state when the engine and workflow support it.

The feature belongs to the Iceberg table layer. It should not be read as a promise that every R2 object has ACID behavior, schema evolution, or time travel. A raw R2 sink that writes JSON or Parquet files is a different destination from an R2 Data Catalog sink that manages Iceberg output.

The R2 Data Catalog sink documentation states that the catalog sink supports Parquet only. JSON may be a valid input format for a stream or a valid output for an R2 file sink, but it is not the documented Iceberg output format for this sink.

Table format does not solve data quality. If an event uses the wrong unit, a missing identifier, or a changed meaning for a field, Iceberg can preserve the resulting records accurately while the analysis remains wrong. The input contract and validation tests still matter.

For a deeper look at retrieval infrastructure, Current Affair’s embeddings API comparison is a separate decision. Embeddings and Iceberg solve different problems and should not be mixed in the same architecture review.

The Terraform Resource Graph

The resource graph is easier to maintain when each object has one clear role. The R2 bucket stores the table data and metadata. The R2 Data Catalog resource enables catalog support on that bucket. The stream receives events. The sink describes the destination. The pipeline connects the stream and sink through SQL. An account token or another credential authorizes the sink.

The Cloudflare Pipelines Terraform guide lists the resource names and shows how they are connected. The provider requirement in that guide is v5.19.0+. The Terraform CLI prerequisite is >= 1.0. Those version markers should remain visible in a review because resource schemas can change between provider releases.

Terraform resourceRole in the stackDependency to verify
cloudflare_r2_bucketCreates the R2 bucket for pipeline outputAccount permission and unique bucket name
cloudflare_r2_data_catalogEnables the catalog on an R2 bucketCorrect bucket reference and catalog access
cloudflare_pipeline_streamDefines input format, schema, and endpoint or Worker accessProducer contract and authentication choice
cloudflare_pipeline_sinkDefines R2 or R2 Data Catalog outputParquet format for Iceberg and valid token
cloudflare_pipelineConnects stream to sink with SQLValid names, fields, and transformation logic

Terraform references should express dependencies rather than repeat names in several places. For example, the sink bucket can reference the R2 bucket resource and the table name can reference the catalog resource. That reduces the chance that an update changes one name while leaving another resource pointed at an old value.

The resource graph should also be split into reviewable modules or files when the team’s scale requires it. A small prototype can keep a single configuration. A shared platform may separate storage, permissions, ingestion, and pipeline definitions so a reviewer can see which change affects which boundary.

Do not treat the presence of a resource block as proof that the account has been deployed successfully. The evidence is the Terraform plan, the apply result, the resulting endpoint or identifier, and a test event that appears in the expected table.

Prerequisites and Least-Privilege Authentication

The full Cloudflare example assumes a Cloudflare account with R2 and Pipelines enabled. It also assumes Terraform and a provider version that contains the required resources. The guide’s general prerequisite is Terraform CLI >= 1.0. Its linked Wrangler guidance requires Node.js 16.17.0 or later for the documented command-line setup path.

Credentials need separate treatment for the Terraform provider and the pipeline sink. The provider token is used to manage infrastructure. The sink token is used by the pipeline to write to the catalog and storage. Giving one broad token to every component increases the effect of a leak or misconfiguration.

The Cloudflare Pipelines getting-started guide says the sink needs an R2 API token with Admin Read & Write permission in its setup flow. The Terraform guide shows a scoped account token resource and permission groups for the full infrastructure example. Follow the current documentation for the exact permission names in the account and provider version you use.

Secrets must remain outside the repository. Use sensitive Terraform variables or an approved secret manager. Protect Terraform state because it can contain resource details and, depending on the configuration, sensitive values. Avoid pasting a token into a plan artifact, issue tracker, or log.

Authentication settings also apply to the input stream. An HTTP endpoint that is open for a tutorial is not automatically suitable for production ingestion. Decide whether the producer can authenticate, how requests are rate-limited, how malformed input is rejected, and how the endpoint is monitored.

Make identity part of the data contract. If multiple producers send similar events, include a source identifier that can be audited. If the pipeline serves multiple tenants, partitioning and access rules must be tested at the table and query layers rather than assumed from a bucket name.

Creating the R2 Bucket and Data Catalog

The storage foundation is an R2 bucket with the Data Catalog enabled. The Terraform guide creates the bucket with cloudflare_r2_bucket, then passes its name into cloudflare_r2_data_catalog. A catalog should be treated as part of the table lifecycle. Destroying or replacing the wrong resource can affect the ability of engines to discover the data.

resource "cloudflare_r2_bucket" "pipeline_bucket" { account_id = var.cloudflare_account_id name = "pipeline-data" } resource "cloudflare_r2_data_catalog" "pipeline_catalog" { account_id = var.cloudflare_account_id bucket_name = cloudflare_r2_bucket.pipeline_bucket.name }

This is an illustrative fragment, not a deployment result. The account variable and provider credentials are example inputs. Run terraform plan in the intended account and inspect whether the bucket and catalog changes match the request.

Bucket naming, environment separation, and retention should be settled before data arrives. A development bucket can use sample events and short retention. A production bucket may need a formal ownership record, backup or export policy, and a process for handling table changes.

The catalog sink page says the specified namespace and table are created if they do not exist, and that sinks cannot be created for existing Iceberg tables. That makes naming and migration planning important. Do not point a new sink at a table name merely because the name looks correct. Confirm whether the table already exists and whether a new table or a supported migration path is intended.

Cloudflare’s public R2 Data Catalog documentation should be checked for catalog management details that sit outside the Pipelines resource graph. Keep table ownership and lifecycle decisions visible to both infrastructure and analytics teams.

Defining the Stream and Input Schema

A stream is the boundary where application events become pipeline data. Define the input format and field schema before writing the pipeline SQL. A small ecommerce-style schema might include an identifier, an event type, an optional product field, and an optional numeric amount. The names are examples. Real schemas should match the producer contract and the analytical question.

resource "cloudflare_pipeline_stream" "events" { account_id = var.cloudflare_account_id name = "events" format = { type = "json" } schema = { fields = [ { name = "entity_id", type = "string", required = true }, { name = "event_type", type = "string", required = true }, { name = "amount", type = "float64", required = false } ] } http = { enabled = true, authentication = true, cors = {} } worker_binding = { enabled = false } }

Make required fields truly required only when producers can supply them consistently. A field that is absent during a rollout can reject otherwise useful events. An optional field still needs a documented meaning for null or missing values.

The input schema should be tested with valid, missing, extra, and wrongly typed fields. Test a batch with one malformed event and observe whether the service rejects the batch, isolates the record, or reports an error. The handling behavior should be documented before a high-volume producer is connected.

Worker bindings and HTTP endpoints serve different integration patterns. A Worker binding can keep the producer inside Cloudflare’s application environment. An HTTP endpoint can serve external producers, but it needs authentication, origin controls where relevant, and monitoring for abuse. The Terraform resource can express the choice, but the surrounding application must enforce the policy.

Schema evolution is a process rather than a one-time field list. Record which producer version emits each change. Add compatibility tests before changing a required field. Query both old and new data in a staging table or test account before applying the change to a shared production table.

Choosing the Sink Format and Rolling Policy

The sink determines how processed data is written. For R2 Data Catalog, the documented output is Parquet because the target is an Apache Iceberg table. For a plain R2 sink, the documentation allows raw Parquet or JSON depending on the configuration. Do not copy an R2 file-sink example into an Iceberg catalog configuration without checking the sink type.

Compression and rolling policy affect storage layout and query behavior. The R2 Data Catalog sink page lists zstd as the default compression option and also lists snappy, gzip, lz4, and uncompressed. The right choice depends on the query engine, CPU budget, ingestion pattern, and storage objective.

Sink decisionDocumented optionTrade-off to test
Output for IcebergParquet onlyColumnar analytics compatibility and file size
Compressionzstd default, plus snappy, gzip, lz4, or uncompressedStorage ratio, write CPU, and scan behavior
Default roll interval300 seconds on the R2 Data Catalog sink pageLatency compared with file size and compaction work
Minimum roll interval60 seconds on the R2 Data Catalog sink pageSmall-file pressure and compaction interaction
Roll size example100 MB example on the sink pageMemory, file size, and query scan efficiency

The sink documentation explains the direction of the trade-off. Lower rolling values can produce more frequent, smaller writes and lower latency. Higher values can produce larger files and improve query performance. There is no single setting that is best for every workload.

Cloudflare’s getting-started interface may expose different advanced sample settings from the sink reference page. Treat those as setup choices that must be checked against the exact sink and current interface rather than as a universal platform limit. Record the value selected in the Terraform configuration and test it with representative event volume.

Iceberg compaction is part of the operational picture. Frequent small files can make maintenance and reads less efficient. A pipeline that appears healthy at the stream endpoint can still need attention if the table accumulates fragments faster than the catalog can compact them.

For another Cloudflare data-system perspective, see Current Affair’s embeddings API comparison. Storage format, indexing, and retrieval model are separate choices.

Wiring SQL and Applying Terraform

The pipeline resource connects the stream and sink with SQL. A simple ingestion statement selects fields from the stream and inserts them into the sink. More advanced pipelines can filter events, rename fields, derive values, or route different records to different destinations, but each transformation adds a schema and testing responsibility.

resource "cloudflare_pipeline" "events_to_iceberg" { account_id = var.cloudflare_account_id name = "events-to-iceberg" sql = "INSERT INTO ${cloudflare_pipeline_sink.iceberg.name} SELECT * FROM ${cloudflare_pipeline_stream.events.name}" }

Keep the SQL readable and test it against the declared schema. A wildcard select is convenient for a stable prototype, but explicit columns can make a production contract clearer when new fields appear. If the sink schema differs from the stream schema, write the projection explicitly and document the conversion.

The normal Terraform flow is terraform init, terraform plan, and terraform apply. Review the plan before apply. Confirm that resource names, account IDs, bucket references, namespace, table name, authentication mode, and SQL all point to the intended environment.

After apply, the documented guide outputs a stream endpoint. Send a small test batch, then inspect the R2 bucket for Iceberg metadata and data files. The getting-started guide says files may take a couple of minutes to appear. That observation window is a guide instruction, not a performance guarantee.

Do not call the configuration complete merely because Terraform returns success. Completion evidence should include a successful apply, an accepted test event, visible table metadata, a query result, and an error-path test. Save these results with the infrastructure change so a later reviewer can distinguish a declared resource from a working pipeline.

Current Affair’s coding-assistant comparison is a useful reminder that execution evidence matters. A configuration file can be syntactically valid while the real integration fails at credentials, data, or runtime behavior.

Querying With R2 SQL and Iceberg Engines

Once the sink has written table metadata and data files, query the table through R2 SQL or another engine that supports Apache Iceberg. Cloudflare’s getting-started guide shows a Wrangler R2 SQL command and also points to configuration examples for other Iceberg engines.

R2 SQL is one query option, not the only one. Spark and DuckDB are named in the Terraform support changelog as compatible query engines. The choice depends on workload size, interactive needs, existing data tools, governance, and the features the team needs from the engine.

npx wrangler r2 sql query "YOUR_WAREHOUSE_NAME" "SELECT entity_id, event_type, amount FROM default.events WHERE event_type = 'purchase' LIMIT 10"

The query is illustrative and uses a sample LIMIT 10. Replace the warehouse and table names with values from the deployed catalog. Do not claim a query result until the endpoint, table, credentials, and data have been tested in the intended account.

Query pathUseful forValidation point
R2 SQLQueries from Cloudflare’s command-line workflowWarehouse name, token, table, and SQL syntax
SparkExisting distributed analytics workflowsCatalog connection, permissions, and schema mapping
DuckDBLocal or lightweight analytical explorationIceberg connector and credential handling
Other Iceberg enginesTeams with an established table-format stackCompatibility with the catalog endpoint and table features

Freshness should be visible to the analyst. A query may not include an event that has been accepted by the stream if the sink has not rolled the file or the metadata is not yet visible. Expose ingestion time, table snapshot information, and producer event time where those fields are meaningful.

Query correctness also depends on schema evolution. Check whether a new field is nullable, whether a renamed field is projected correctly, and whether older snapshots remain readable. A successful query against today’s snapshot does not validate every historical table state.

Current Affair’s embeddings article and vector database article cover different data-access layers. Do not substitute a similarity index for a governed analytical table.

Testing, Failure Handling, and Cost Boundaries

A reliable pipeline needs tests at the resource, data, and query levels. Terraform validation checks that configuration is coherent. A test event checks the ingestion path. A table inspection checks the sink and catalog. A query checks whether an approved engine can read the result. A failure test checks whether the system reports and contains a bad input or credential.

Start with a small dataset that represents normal and edge cases. Include a valid record, a missing required field, an unexpected type, a late event, a duplicate event, and a schema change. Record whether each case is accepted, rejected, delayed, or transformed. Avoid using a synthetic success message as evidence that the whole path works.

Credentials deserve a dedicated failure test. Revoke or replace a test token in a non-production environment and confirm that the sink reports a clear error without exposing the secret. Verify that a read-only user cannot create or alter infrastructure. Confirm that a query token cannot write data when the policy says it should not.

Cost review should be based on the current Cloudflare pricing pages and the actual workload. Do not repeat a blanket “zero egress fees” promise as if it described every charge. Storage, requests, query execution, compute, retention, and external engine costs can still matter. A data movement policy should be reviewed with the account’s current commercial terms.

Operational ownership also includes Terraform state, provider upgrades, table compaction, endpoint abuse, dead-letter handling, and incident recovery. A managed sink can reduce the number of services the team runs, but it does not remove the need to decide who responds when the pipeline stops or data quality falls.

For systems that add agent-generated transformations, Current Affair’s agent execution guide is relevant. An automated proposal still needs schema validation, permission checks, and a review before it changes production data.

A Practical Terraform Rollout Checklist

Use the following checklist before treating a Cloudflare Pipelines and R2 Data Catalog configuration as ready for a wider rollout. First, confirm the provider version and Terraform prerequisite. Second, confirm the account, bucket, namespace, and table ownership. Third, review the stream schema with the teams that produce the data.

Next, inspect the sink type and output format. For an Iceberg table, confirm that the sink writes Parquet. Choose compression and rolling settings from measured workload needs, not from a copied tutorial value. Check the catalog token scope and make sure secrets are not stored in a public repository or an unprotected artifact.

Then review SQL against the declared fields. Test a small event batch, a malformed record, a duplicate, and a schema change. Confirm when files and table metadata become visible. Query the table with the intended engine and compare the result with the input event source.

Finally, define the day-two process. Decide who owns Terraform state, who approves provider upgrades, how table changes are migrated, how failed events are investigated, and how data is deleted or retained. Document the cost boundary and the conditions for pausing ingestion.

Cloudflare’s Terraform support provides a declarative way to connect these components. It does not turn data engineering into a copy-and-paste exercise. The value comes from a tested resource graph, narrow permissions, explicit schemas, observable transformations, and a table that analysts can query with understood freshness and history.

That is the accurate reading of this 2026 setup. Terraform can make the pipeline repeatable. Pipelines can move and transform events. R2 Data Catalog can manage Apache Iceberg tables. The team still owns the decisions that make the system safe, correct, and economical.

Frequently Asked Questions

Cloudflare Pipelines receives streaming events through an HTTP endpoint or Worker binding, applies SQL transformations, and sends the processed result to a configured sink. The sink can write to R2 Data Catalog as an Apache Iceberg table or to R2 as files, depending on the resource configuration.
Cloudflare’s Terraform documentation for this workflow requires the Cloudflare provider version 5.19.0 or later. The configuration should pin or constrain the provider version that the team has tested and should be reviewed when the provider is upgraded.
No. The R2 Data Catalog sink documentation states that Iceberg output uses Parquet only. JSON can be used as stream input or with an R2 file sink, but it is not the documented output format for an R2 Data Catalog Iceberg sink.
R2 Data Catalog manages Apache Iceberg table metadata and writes processed data as Iceberg tables. A raw R2 sink writes files such as JSON or Parquet without providing the same catalog table layer. Choose the destination based on how the data will be queried and managed.
Use a token with the permissions required by the current Cloudflare documentation and keep it outside source code. The Terraform example uses a scoped account token for sink authentication, while the getting-started guide uses an R2 API token with Admin Read and Write permission. Confirm the exact permission names for the provider and account configuration in use.
Rolling settings control how often or how large the sink files become. The R2 Data Catalog sink page documents a 300-second default interval and a 60-second minimum, while the getting-started interface shows separate advanced sample settings. Use the sink-specific current documentation and test the file-size and latency trade-off for the workload.
Review the Terraform plan, apply the configuration in the intended account, send a small valid test batch, inspect the R2 bucket for Iceberg metadata and data files, and query the table with R2 SQL or another compatible Iceberg engine. Test malformed input, credential failure, duplicates, and schema changes before production use.
SK Jabedul Haque
Written by

SK Jabedul Haque

Founder & Chief Editor

Building India's most trusted finance education platform — simplifying news, schemes and market trends so anyone can understand and invest confidently.

Read full bio

Never miss an update

Get our clearest explainers on schemes, markets and money — read what matters, without the noise.

Explore more articles
In this article