Cloudflare Pipelines + R2 Data Catalog Terraform 2026: Complete Apache Iceberg Pipeline Setup
What You'll Learn
- What Cloudflare Pipelines and R2 Data Catalog each contribute to a streaming data stack.
- How Terraform resources connect buckets, streams, sinks, SQL pipelines, and scoped credentials.
- Why Apache Iceberg output uses Parquet and how table features differ from raw R2 objects.
- Which tests, limits, permissions, and cost checks belong before a production rollout.
What Terraform Support Adds in 2026
Cloudflare Pipelines and R2 Data Catalog Terraform 2026 support is an infrastructure-as-code path for a data pipeline that spans ingestion, SQL transformation, storage, and table management. Cloudflare’s May 4, 2026 changelog says Pipelines and R2 Data Catalog can be created and managed with Cloudflare Terraform provider version 5.19.0.
The practical change is not that Terraform turns a data platform into a single command. It gives a team a declared configuration for the resources that make the path work. A review can compare the proposed stream, sink, table, SQL, and permissions before an apply. A later run can reconcile the declared state with the account.
The official Cloudflare Terraform support announcement identifies four new resources for the Pipelines and catalog workflow. They cover the catalog, the input stream, the output sink, and the SQL pipeline that connects them. The full Terraform example also uses an R2 bucket and an account token resource.
That distinction matters when reviewing the old article’s “zero operational overhead” wording. Terraform reduces manual configuration and makes changes reviewable. It does not decide whether an event schema is correct, whether a token is over-permissioned, whether a small-file problem will affect queries, or whether an update can safely change a table.
Cloudflare’s current documentation is dated in 2026 and may change while the feature matures. Pin the provider range that your team has tested, read the resource documentation during upgrades, and keep the generated plan in the same review process as application code.
For adjacent context on Cloudflare’s storage and retrieval products, see Current Affair’s vector database comparison. A catalog table is a managed data asset, not merely another object path.
The Data Flow From Event to Iceberg Table
A useful mental model starts with the event and follows it to the query. A stream receives data through an HTTP endpoint or a Worker binding. A pipeline applies SQL to that stream. A sink writes the transformed result to R2 Data Catalog as an Apache Iceberg table. A query engine then reads table metadata and data files.
Each layer has a different responsibility. The stream defines how data enters and what shape the input is expected to have. The pipeline defines which fields are selected or transformed. The sink defines where the output goes and how it is written. The catalog provides table metadata and a namespace for Iceberg tables. The query engine reads the resulting table through its supported connection.
The Cloudflare Pipelines documentation describes Pipelines as a service that can ingest streaming data, transform it with SQL, and write to destinations such as R2 Data Catalog or R2. The destination choice should follow the query requirement. If the target is a managed Iceberg table, use the catalog sink. If the target is a raw file archive, an R2 sink may be more appropriate.
| Layer | Primary responsibility | Review question |
|---|---|---|
| Stream | Receive events and validate the input schema | Can producers send data in the declared shape? |
| Pipeline | Apply SQL and connect input to output | Does the transformation preserve the required fields? |
| Sink | Write processed data to R2 or R2 Data Catalog | Is the format and destination correct? |
| Catalog | Manage the Iceberg table namespace and metadata | Can approved engines discover and read the table? |
| Query engine | Read and analyze the table | Are freshness, schema, and permissions visible to users? |
The flow also defines failure boundaries. An event can fail schema validation before it reaches SQL. A SQL statement can select a field that no longer exists. A sink can reject credentials or a table configuration. A table can be present while recent data is still rolling into a file. Monitoring should identify which boundary failed instead of reporting only that “the pipeline is down.”
Current Affair’s operating-system layer analysis makes a similar point for agent systems. Clear boundaries make it easier to test and repair a system than a single label that hides several components.
Why Apache Iceberg Matters for R2 Data
R2 Data Catalog is useful here because the sink writes processed data as Apache Iceberg tables rather than treating every output as an unrelated file. Iceberg is a table format with metadata that tracks snapshots, schema, and the relationship between data files. Cloudflare’s sink documentation says Iceberg tables provide ACID transactions, schema evolution, and time travel capabilities for analytics workloads.
These features address problems that appear when files are written continuously. A query needs to know which files belong to a table version. A schema change needs a controlled interpretation. A correction may require a new snapshot rather than an opaque replacement of objects. Time travel can help a team inspect an earlier table state when the engine and workflow support it.
The feature belongs to the Iceberg table layer. It should not be read as a promise that every R2 object has ACID behavior, schema evolution, or time travel. A raw R2 sink that writes JSON or Parquet files is a different destination from an R2 Data Catalog sink that manages Iceberg output.
The R2 Data Catalog sink documentation states that the catalog sink supports Parquet only. JSON may be a valid input format for a stream or a valid output for an R2 file sink, but it is not the documented Iceberg output format for this sink.
Table format does not solve data quality. If an event uses the wrong unit, a missing identifier, or a changed meaning for a field, Iceberg can preserve the resulting records accurately while the analysis remains wrong. The input contract and validation tests still matter.
For a deeper look at retrieval infrastructure, Current Affair’s embeddings API comparison is a separate decision. Embeddings and Iceberg solve different problems and should not be mixed in the same architecture review.
The Terraform Resource Graph
The resource graph is easier to maintain when each object has one clear role. The R2 bucket stores the table data and metadata. The R2 Data Catalog resource enables catalog support on that bucket. The stream receives events. The sink describes the destination. The pipeline connects the stream and sink through SQL. An account token or another credential authorizes the sink.
The Cloudflare Pipelines Terraform guide lists the resource names and shows how they are connected. The provider requirement in that guide is v5.19.0+. The Terraform CLI prerequisite is >= 1.0. Those version markers should remain visible in a review because resource schemas can change between provider releases.
| Terraform resource | Role in the stack | Dependency to verify |
|---|---|---|
cloudflare_r2_bucket | Creates the R2 bucket for pipeline output | Account permission and unique bucket name |
cloudflare_r2_data_catalog | Enables the catalog on an R2 bucket | Correct bucket reference and catalog access |
cloudflare_pipeline_stream | Defines input format, schema, and endpoint or Worker access | Producer contract and authentication choice |
cloudflare_pipeline_sink | Defines R2 or R2 Data Catalog output | Parquet format for Iceberg and valid token |
cloudflare_pipeline | Connects stream to sink with SQL | Valid names, fields, and transformation logic |
Terraform references should express dependencies rather than repeat names in several places. For example, the sink bucket can reference the R2 bucket resource and the table name can reference the catalog resource. That reduces the chance that an update changes one name while leaving another resource pointed at an old value.
The resource graph should also be split into reviewable modules or files when the team’s scale requires it. A small prototype can keep a single configuration. A shared platform may separate storage, permissions, ingestion, and pipeline definitions so a reviewer can see which change affects which boundary.
Do not treat the presence of a resource block as proof that the account has been deployed successfully. The evidence is the Terraform plan, the apply result, the resulting endpoint or identifier, and a test event that appears in the expected table.
Prerequisites and Least-Privilege Authentication
The full Cloudflare example assumes a Cloudflare account with R2 and Pipelines enabled. It also assumes Terraform and a provider version that contains the required resources. The guide’s general prerequisite is Terraform CLI >= 1.0. Its linked Wrangler guidance requires Node.js 16.17.0 or later for the documented command-line setup path.
Credentials need separate treatment for the Terraform provider and the pipeline sink. The provider token is used to manage infrastructure. The sink token is used by the pipeline to write to the catalog and storage. Giving one broad token to every component increases the effect of a leak or misconfiguration.
The Cloudflare Pipelines getting-started guide says the sink needs an R2 API token with Admin Read & Write permission in its setup flow. The Terraform guide shows a scoped account token resource and permission groups for the full infrastructure example. Follow the current documentation for the exact permission names in the account and provider version you use.
Secrets must remain outside the repository. Use sensitive Terraform variables or an approved secret manager. Protect Terraform state because it can contain resource details and, depending on the configuration, sensitive values. Avoid pasting a token into a plan artifact, issue tracker, or log.
Authentication settings also apply to the input stream. An HTTP endpoint that is open for a tutorial is not automatically suitable for production ingestion. Decide whether the producer can authenticate, how requests are rate-limited, how malformed input is rejected, and how the endpoint is monitored.
Make identity part of the data contract. If multiple producers send similar events, include a source identifier that can be audited. If the pipeline serves multiple tenants, partitioning and access rules must be tested at the table and query layers rather than assumed from a bucket name.
Creating the R2 Bucket and Data Catalog
The storage foundation is an R2 bucket with the Data Catalog enabled. The Terraform guide creates the bucket with cloudflare_r2_bucket, then passes its name into cloudflare_r2_data_catalog. A catalog should be treated as part of the table lifecycle. Destroying or replacing the wrong resource can affect the ability of engines to discover the data.
resource "cloudflare_r2_bucket" "pipeline_bucket" { account_id = var.cloudflare_account_id name = "pipeline-data" } resource "cloudflare_r2_data_catalog" "pipeline_catalog" { account_id = var.cloudflare_account_id bucket_name = cloudflare_r2_bucket.pipeline_bucket.name }This is an illustrative fragment, not a deployment result. The account variable and provider credentials are example inputs. Run terraform plan in the intended account and inspect whether the bucket and catalog changes match the request.
Bucket naming, environment separation, and retention should be settled before data arrives. A development bucket can use sample events and short retention. A production bucket may need a formal ownership record, backup or export policy, and a process for handling table changes.
The catalog sink page says the specified namespace and table are created if they do not exist, and that sinks cannot be created for existing Iceberg tables. That makes naming and migration planning important. Do not point a new sink at a table name merely because the name looks correct. Confirm whether the table already exists and whether a new table or a supported migration path is intended.
Cloudflare’s public R2 Data Catalog documentation should be checked for catalog management details that sit outside the Pipelines resource graph. Keep table ownership and lifecycle decisions visible to both infrastructure and analytics teams.
Defining the Stream and Input Schema
A stream is the boundary where application events become pipeline data. Define the input format and field schema before writing the pipeline SQL. A small ecommerce-style schema might include an identifier, an event type, an optional product field, and an optional numeric amount. The names are examples. Real schemas should match the producer contract and the analytical question.
resource "cloudflare_pipeline_stream" "events" { account_id = var.cloudflare_account_id name = "events" format = { type = "json" } schema = { fields = [ { name = "entity_id", type = "string", required = true }, { name = "event_type", type = "string", required = true }, { name = "amount", type = "float64", required = false } ] } http = { enabled = true, authentication = true, cors = {} } worker_binding = { enabled = false } }Make required fields truly required only when producers can supply them consistently. A field that is absent during a rollout can reject otherwise useful events. An optional field still needs a documented meaning for null or missing values.
The input schema should be tested with valid, missing, extra, and wrongly typed fields. Test a batch with one malformed event and observe whether the service rejects the batch, isolates the record, or reports an error. The handling behavior should be documented before a high-volume producer is connected.
Worker bindings and HTTP endpoints serve different integration patterns. A Worker binding can keep the producer inside Cloudflare’s application environment. An HTTP endpoint can serve external producers, but it needs authentication, origin controls where relevant, and monitoring for abuse. The Terraform resource can express the choice, but the surrounding application must enforce the policy.
Schema evolution is a process rather than a one-time field list. Record which producer version emits each change. Add compatibility tests before changing a required field. Query both old and new data in a staging table or test account before applying the change to a shared production table.
Choosing the Sink Format and Rolling Policy
The sink determines how processed data is written. For R2 Data Catalog, the documented output is Parquet because the target is an Apache Iceberg table. For a plain R2 sink, the documentation allows raw Parquet or JSON depending on the configuration. Do not copy an R2 file-sink example into an Iceberg catalog configuration without checking the sink type.
Compression and rolling policy affect storage layout and query behavior. The R2 Data Catalog sink page lists zstd as the default compression option and also lists snappy, gzip, lz4, and uncompressed. The right choice depends on the query engine, CPU budget, ingestion pattern, and storage objective.
| Sink decision | Documented option | Trade-off to test |
|---|---|---|
| Output for Iceberg | Parquet only | Columnar analytics compatibility and file size |
| Compression | zstd default, plus snappy, gzip, lz4, or uncompressed | Storage ratio, write CPU, and scan behavior |
| Default roll interval | 300 seconds on the R2 Data Catalog sink page | Latency compared with file size and compaction work |
| Minimum roll interval | 60 seconds on the R2 Data Catalog sink page | Small-file pressure and compaction interaction |
| Roll size example | 100 MB example on the sink page | Memory, file size, and query scan efficiency |
The sink documentation explains the direction of the trade-off. Lower rolling values can produce more frequent, smaller writes and lower latency. Higher values can produce larger files and improve query performance. There is no single setting that is best for every workload.
Cloudflare’s getting-started interface may expose different advanced sample settings from the sink reference page. Treat those as setup choices that must be checked against the exact sink and current interface rather than as a universal platform limit. Record the value selected in the Terraform configuration and test it with representative event volume.
Iceberg compaction is part of the operational picture. Frequent small files can make maintenance and reads less efficient. A pipeline that appears healthy at the stream endpoint can still need attention if the table accumulates fragments faster than the catalog can compact them.
For another Cloudflare data-system perspective, see Current Affair’s embeddings API comparison. Storage format, indexing, and retrieval model are separate choices.
Wiring SQL and Applying Terraform
The pipeline resource connects the stream and sink with SQL. A simple ingestion statement selects fields from the stream and inserts them into the sink. More advanced pipelines can filter events, rename fields, derive values, or route different records to different destinations, but each transformation adds a schema and testing responsibility.
resource "cloudflare_pipeline" "events_to_iceberg" { account_id = var.cloudflare_account_id name = "events-to-iceberg" sql = "INSERT INTO ${cloudflare_pipeline_sink.iceberg.name} SELECT * FROM ${cloudflare_pipeline_stream.events.name}" }Keep the SQL readable and test it against the declared schema. A wildcard select is convenient for a stable prototype, but explicit columns can make a production contract clearer when new fields appear. If the sink schema differs from the stream schema, write the projection explicitly and document the conversion.
The normal Terraform flow is terraform init, terraform plan, and terraform apply. Review the plan before apply. Confirm that resource names, account IDs, bucket references, namespace, table name, authentication mode, and SQL all point to the intended environment.
After apply, the documented guide outputs a stream endpoint. Send a small test batch, then inspect the R2 bucket for Iceberg metadata and data files. The getting-started guide says files may take a couple of minutes to appear. That observation window is a guide instruction, not a performance guarantee.
Do not call the configuration complete merely because Terraform returns success. Completion evidence should include a successful apply, an accepted test event, visible table metadata, a query result, and an error-path test. Save these results with the infrastructure change so a later reviewer can distinguish a declared resource from a working pipeline.
Current Affair’s coding-assistant comparison is a useful reminder that execution evidence matters. A configuration file can be syntactically valid while the real integration fails at credentials, data, or runtime behavior.
Querying With R2 SQL and Iceberg Engines
Once the sink has written table metadata and data files, query the table through R2 SQL or another engine that supports Apache Iceberg. Cloudflare’s getting-started guide shows a Wrangler R2 SQL command and also points to configuration examples for other Iceberg engines.
R2 SQL is one query option, not the only one. Spark and DuckDB are named in the Terraform support changelog as compatible query engines. The choice depends on workload size, interactive needs, existing data tools, governance, and the features the team needs from the engine.
npx wrangler r2 sql query "YOUR_WAREHOUSE_NAME" "SELECT entity_id, event_type, amount FROM default.events WHERE event_type = 'purchase' LIMIT 10"The query is illustrative and uses a sample LIMIT 10. Replace the warehouse and table names with values from the deployed catalog. Do not claim a query result until the endpoint, table, credentials, and data have been tested in the intended account.
| Query path | Useful for | Validation point |
|---|---|---|
| R2 SQL | Queries from Cloudflare’s command-line workflow | Warehouse name, token, table, and SQL syntax |
| Spark | Existing distributed analytics workflows | Catalog connection, permissions, and schema mapping |
| DuckDB | Local or lightweight analytical exploration | Iceberg connector and credential handling |
| Other Iceberg engines | Teams with an established table-format stack | Compatibility with the catalog endpoint and table features |
Freshness should be visible to the analyst. A query may not include an event that has been accepted by the stream if the sink has not rolled the file or the metadata is not yet visible. Expose ingestion time, table snapshot information, and producer event time where those fields are meaningful.
Query correctness also depends on schema evolution. Check whether a new field is nullable, whether a renamed field is projected correctly, and whether older snapshots remain readable. A successful query against today’s snapshot does not validate every historical table state.
Current Affair’s embeddings article and vector database article cover different data-access layers. Do not substitute a similarity index for a governed analytical table.
Testing, Failure Handling, and Cost Boundaries
A reliable pipeline needs tests at the resource, data, and query levels. Terraform validation checks that configuration is coherent. A test event checks the ingestion path. A table inspection checks the sink and catalog. A query checks whether an approved engine can read the result. A failure test checks whether the system reports and contains a bad input or credential.
Start with a small dataset that represents normal and edge cases. Include a valid record, a missing required field, an unexpected type, a late event, a duplicate event, and a schema change. Record whether each case is accepted, rejected, delayed, or transformed. Avoid using a synthetic success message as evidence that the whole path works.
Credentials deserve a dedicated failure test. Revoke or replace a test token in a non-production environment and confirm that the sink reports a clear error without exposing the secret. Verify that a read-only user cannot create or alter infrastructure. Confirm that a query token cannot write data when the policy says it should not.
Cost review should be based on the current Cloudflare pricing pages and the actual workload. Do not repeat a blanket “zero egress fees” promise as if it described every charge. Storage, requests, query execution, compute, retention, and external engine costs can still matter. A data movement policy should be reviewed with the account’s current commercial terms.
Operational ownership also includes Terraform state, provider upgrades, table compaction, endpoint abuse, dead-letter handling, and incident recovery. A managed sink can reduce the number of services the team runs, but it does not remove the need to decide who responds when the pipeline stops or data quality falls.
For systems that add agent-generated transformations, Current Affair’s agent execution guide is relevant. An automated proposal still needs schema validation, permission checks, and a review before it changes production data.
A Practical Terraform Rollout Checklist
Use the following checklist before treating a Cloudflare Pipelines and R2 Data Catalog configuration as ready for a wider rollout. First, confirm the provider version and Terraform prerequisite. Second, confirm the account, bucket, namespace, and table ownership. Third, review the stream schema with the teams that produce the data.
Next, inspect the sink type and output format. For an Iceberg table, confirm that the sink writes Parquet. Choose compression and rolling settings from measured workload needs, not from a copied tutorial value. Check the catalog token scope and make sure secrets are not stored in a public repository or an unprotected artifact.
Then review SQL against the declared fields. Test a small event batch, a malformed record, a duplicate, and a schema change. Confirm when files and table metadata become visible. Query the table with the intended engine and compare the result with the input event source.
Finally, define the day-two process. Decide who owns Terraform state, who approves provider upgrades, how table changes are migrated, how failed events are investigated, and how data is deleted or retained. Document the cost boundary and the conditions for pausing ingestion.
Cloudflare’s Terraform support provides a declarative way to connect these components. It does not turn data engineering into a copy-and-paste exercise. The value comes from a tested resource graph, narrow permissions, explicit schemas, observable transformations, and a table that analysts can query with understood freshness and history.
That is the accurate reading of this 2026 setup. Terraform can make the pipeline repeatable. Pipelines can move and transform events. R2 Data Catalog can manage Apache Iceberg tables. The team still owns the decisions that make the system safe, correct, and economical.
Frequently Asked Questions
SK Jabedul Haque
Building India's most trusted finance education platform — simplifying news, schemes and market trends so anyone can understand and invest confidently.
Read full bioNever miss an update
Get our clearest explainers on schemes, markets and money — read what matters, without the noise.
Explore more articles