Skip to main content
Matthias Broecheler
CEO of DataSQRL
View all authors

The Whole Pipeline in One File a Human Can Actually Read

· 6 min read
Matthias Broecheler
CEO of DataSQRL

Part 4 of 4: How a data engineering harness eliminates the AI coding agent errors that survive review.

Introduction​

"Why do I need that? Isn't Claude Code good enough?"

A human reviewing a complete data pipeline expressed in one readable SQL file >|

We get that question a lot. We are building an open-source data engineering harness, the tooling and guardrails that a coding agent uses to build data pipelines. Claude Code, Codex, and OpenCode already write plausible data pipeline code. So what is the harness for?

Parts 1, 2, and 3 were about generating the pipeline, validating it, and testing it. Part 4 is about the check that cannot be automated away: a person understanding what the pipeline actually means.

When an agent builds a conventional pipeline, the logic ends up spread across dozens or hundreds of files in two or three languages. Ingestion scripts, transformation jobs, table definitions, glue, API resolvers. Nobody holds that in their head. So the bugs where every line is correct and the meaning is wrong slip through, and the human review that was supposed to be the last line of defense collapses under fatigue.

Standard Integration Tests Obfuscate Common Errors

· 5 min read
Matthias Broecheler
CEO of DataSQRL

Part 3 of 4: How a data engineering harness eliminates the AI coding agent errors that survive standard testing.

Introduction​

"Why do I need that? Isn't Claude Code good enough?"

An event-time simulator replays data pipeline events at their original timestamps to test time-dependent behavior >|

We get that question a lot. We are building an open-source data engineering harness, the tooling and guardrails that a coding agent uses to build data pipelines. Claude Code, Codex, and OpenCode already write plausible data pipeline code. So what is the harness for?

Parts 1 and 2 were about generating the pipeline correctly and validating it. Part 3 is about the accomplice in every failure so far: the test that passed.

An agent writes a test, runs it against a fixed snapshot of data, sees green, and ships. But a pipeline's hardest bugs only exist in motion which standard tests don't catch, leaving you to troubleshoot in production.

The Bugs That Aren't in the Text: What Deep Relational Introspection Catches

· 7 min read
Matthias Broecheler
CEO of DataSQRL

Part 2 of 4: How a data engineering harness eliminates the AI coding agent errors that survive review.

Introduction​

"Why do I need that? Isn't Claude Code good enough?"

Deep introspection of the logic behind a data pipeline can reveal logical flaws

We get that question a lot. We are building an open-source data engineering harness, the tooling and guardrails that a coding agent uses to build data pipelines. Claude Code, Codex, and OpenCode already write plausible data pipeline code. So what is the harness for?

Part 1 was about the seams between systems, and how a transpiler generates them deterministically. Part 2 is about a harder problem.

Some bugs have no correctness condition in the query text at all. The condition lives in the relationship between a query, the data it reads, and how that data changes over time. An agent reads SQL/code as text, and at the text level these bugs are invisible until the production deployment fails.

The Pipeline Compiles, the Demo Works, and the Seams Are Quietly Broken

· 8 min read
Matthias Broecheler
CEO of DataSQRL

Part 1 of 4: How a data engineering harness eliminates the AI coding agent errors that survive review.

Introduction​

"Why do I need that? Isn't Claude Code good enough?"

We get that question a lot. We are building an open-source data engineering harness, the tooling and guardrails that a coding agent uses to build data pipelines. Claude Code, Codex, and OpenCode already write plausible data pipeline code. So what is the harness for?

A data pipeline spanning multiple systems with broken seams at the boundaries between them.

This series answers that. It looks at where general-purpose coding harnesses fail on data engineering work, and why those failures survive code review.

Part 1 is about the seams. A pipeline spans multiple data systems, and the same fact (a type, a name, an encoding) has to be restated in each one. Restating facts consistently across systems is strict rule-following. The thing doing the restating is a probabilistic model. It gets most of them right. The ones it gets wrong compile cleanly, pass the demo, and surface in production as overflowed numbers and fields that were never there.

A transpiler removes that entire class of error. It derives every boundary asset from one logical model, so the systems cannot disagree. The agent also writes less code, which means fewer tokens and fewer duplicated schema definitions cluttering its context.

AI in Data Engineering: Building Reliable Data Systems at Scale

· 13 min read
Matthias Broecheler
CEO of DataSQRL

AI is transforming data engineering. Coding agents can now generate SQL transformations, configure connectors, and define API schemas in minutes rather than days. But here's the catch: a query that works perfectly on test data may fail catastrophically when confronted with late-arriving events, schema evolution, or terabyte-scale volumes.

How do we ensure that AI-generated data systems meet the rigorous non-functional requirements that production data platforms demand? This article presents our framework for integrating AI coding agents into data engineering workflows while maintaining data quality, reliability, governance, and trust.

Agentic Data Engineering Harness

· 19 min read
Matthias Broecheler
CEO of DataSQRL

DataSQRL is an open-source data engineering harness that provides guardrails and feedback for AI coding agents to develop and operate data pipelines, data products, and data APIs autonomously. You can customize DataSQRL as the foundation of your agentic data platform. Our goal is to develop DataSQRL into a comprehensive data engineering harness for data platform automation.

DataSQRL harness architecture showing coding agent with framework, validator, and simulator feedback loops >

0.10 Release: Iceberg Mutations

· 3 min read
Ferenc Csaky
Apache Flink PMC
Matthias Broecheler
CEO of DataSQRL
SQRL 0.10 Release >

DataSQRL 0.10 has been released and the headline feature is supporting mutations for Iceberg tables. DataSQRL can now manage Apache Iceberg tables as sources and sinks.

Why is that a big deal? Up to this point, DataSQRL could read and write to Apache Iceberg tables, but you had to manage them explicitly. This new release makes it easy to share data through Apache Iceberg between DataSQRL pipelines.

Avoiding Duplicate Processing in Flink SQL Streaming Jobs

· 6 min read
Ferenc Csaky
Apache Flink PMC
Matthias Broecheler
CEO of DataSQRL

Flink SQL is a powerful abstraction layer that unifies batch and stream processing over semi-structured data. It extends the widely used SQL language with streaming constructs such as tumbling windows, session windows, and, more recently, process table functions. This enables non-experts in streaming technologies to express complex real-time data processing logic succinctly.

As a result, Flink SQL significantly lowers the barrier to entry for building real-time data systems.

However, developing streaming applications differs fundamentally from traditional SQL query processing. One key difference is that streaming jobs often have multiple sinks populated by a single pipeline, sharing large portions of common data processing logic.

While Flink SQL provides mechanisms to express this concisely—using views and statement sets—in practice, this often results in duplicate processing in the generated job graph.

DataSQRL 0.7 Release: The Data Delivery Interface

· 3 min read
Matthias Broecheler
CEO of DataSQRL
DataSQRL 0.7.0 Release >|

DataSQRL 0.7 marks a major milestone in our journey to automate data pipelines, thanks to significant improvements to the serving layer:

  • Support for the Model Context Protocol (MCP) for tooling and resource access
  • REST API support
  • JWT-based authentication and authorization

These features enable developers to build a wide range of production-ready data interfaces. This release also includes performance and configuration improvements to the serving layer of DataSQRL-generated pipelines.

You can find the full release notes and source code on our GitHub release page. To update your local installation of DataSQRL, simply pull the latest Docker image:

docker pull datasqrl/cmd:0.7.0

The Last Mile: Data Delivery​

Data delivery is the final and most visible stage of any data pipeline. It's how users, applications, and AI agents actually access and consume data. Most enterprise data interactions happen through APIs, making the delivery interface a critical component. At DataSQRL, we've invested heavily in automating the upstream parts of the pipeline: from Flink-powered data processing to Postgres-backed storage. With version 0.7, we turn our focus to the serving layer: introducing support for the Model Context Protocol (MCP) and REST APIs, as well as JWT-based authentication and authorization. These additions ensure seamless integration with most authentication providers and enable secure, token-based data access, with fine-grained authorization logic enforced directly in the SQRL script. This completes our vision of end-to-end pipeline automation, where consumption patterns inform data storage and processing—closing the loop between data production and usage.

Check out the interface documentation for more information.

Flink SQL Runner: Run Flink SQL Without JARs or Glue Code

· 3 min read
Matthias Broecheler
CEO of DataSQRL

Apache Flink has long been a powerhouse for streaming and batch data processing. And with the rise of Flink SQL, developers can now build sophisticated pipelines using a declarative language they already know. But getting Flink SQL applications into production still comes with friction: packaging JARs, managing connectors, injecting secrets, and wiring up deployment infrastructure.

FlinkSQL Runner >

Flink SQL Runner is here to change that. It's an open-source toolkit that simplifies development, deployment, and operation of Flink SQL applications—locally or in Kubernetes—without manual JAR assembly or scripting custom infrastructure pipelines.