Skip to main content

9 posts tagged with "DataSQRL"

Posts about DataSQRL

View All Tags

The Whole Pipeline in One File a Human Can Actually Read

· 6 min read
Matthias Broecheler
CEO of DataSQRL

Part 4 of 4: How a data engineering harness eliminates the AI coding agent errors that survive review.

Introduction​

"Why do I need that? Isn't Claude Code good enough?"

A human reviewing a complete data pipeline expressed in one readable SQL file >|

We get that question a lot. We are building an open-source data engineering harness, the tooling and guardrails that a coding agent uses to build data pipelines. Claude Code, Codex, and OpenCode already write plausible data pipeline code. So what is the harness for?

Parts 1, 2, and 3 were about generating the pipeline, validating it, and testing it. Part 4 is about the check that cannot be automated away: a person understanding what the pipeline actually means.

When an agent builds a conventional pipeline, the logic ends up spread across dozens or hundreds of files in two or three languages. Ingestion scripts, transformation jobs, table definitions, glue, API resolvers. Nobody holds that in their head. So the bugs where every line is correct and the meaning is wrong slip through, and the human review that was supposed to be the last line of defense collapses under fatigue.

Standard Integration Tests Obfuscate Common Errors

· 5 min read
Matthias Broecheler
CEO of DataSQRL

Part 3 of 4: How a data engineering harness eliminates the AI coding agent errors that survive standard testing.

Introduction​

"Why do I need that? Isn't Claude Code good enough?"

An event-time simulator replays data pipeline events at their original timestamps to test time-dependent behavior >|

We get that question a lot. We are building an open-source data engineering harness, the tooling and guardrails that a coding agent uses to build data pipelines. Claude Code, Codex, and OpenCode already write plausible data pipeline code. So what is the harness for?

Parts 1 and 2 were about generating the pipeline correctly and validating it. Part 3 is about the accomplice in every failure so far: the test that passed.

An agent writes a test, runs it against a fixed snapshot of data, sees green, and ships. But a pipeline's hardest bugs only exist in motion which standard tests don't catch, leaving you to troubleshoot in production.

The Bugs That Aren't in the Text: What Deep Relational Introspection Catches

· 7 min read
Matthias Broecheler
CEO of DataSQRL

Part 2 of 4: How a data engineering harness eliminates the AI coding agent errors that survive review.

Introduction​

"Why do I need that? Isn't Claude Code good enough?"

Deep introspection of the logic behind a data pipeline can reveal logical flaws

We get that question a lot. We are building an open-source data engineering harness, the tooling and guardrails that a coding agent uses to build data pipelines. Claude Code, Codex, and OpenCode already write plausible data pipeline code. So what is the harness for?

Part 1 was about the seams between systems, and how a transpiler generates them deterministically. Part 2 is about a harder problem.

Some bugs have no correctness condition in the query text at all. The condition lives in the relationship between a query, the data it reads, and how that data changes over time. An agent reads SQL/code as text, and at the text level these bugs are invisible until the production deployment fails.

The Pipeline Compiles, the Demo Works, and the Seams Are Quietly Broken

· 8 min read
Matthias Broecheler
CEO of DataSQRL

Part 1 of 4: How a data engineering harness eliminates the AI coding agent errors that survive review.

Introduction​

"Why do I need that? Isn't Claude Code good enough?"

We get that question a lot. We are building an open-source data engineering harness, the tooling and guardrails that a coding agent uses to build data pipelines. Claude Code, Codex, and OpenCode already write plausible data pipeline code. So what is the harness for?

A data pipeline spanning multiple systems with broken seams at the boundaries between them.

This series answers that. It looks at where general-purpose coding harnesses fail on data engineering work, and why those failures survive code review.

Part 1 is about the seams. A pipeline spans multiple data systems, and the same fact (a type, a name, an encoding) has to be restated in each one. Restating facts consistently across systems is strict rule-following. The thing doing the restating is a probabilistic model. It gets most of them right. The ones it gets wrong compile cleanly, pass the demo, and surface in production as overflowed numbers and fields that were never there.

A transpiler removes that entire class of error. It derives every boundary asset from one logical model, so the systems cannot disagree. The agent also writes less code, which means fewer tokens and fewer duplicated schema definitions cluttering its context.

0.10 Release: Iceberg Mutations

· 3 min read
Ferenc Csaky
Apache Flink PMC
Matthias Broecheler
CEO of DataSQRL
SQRL 0.10 Release >

DataSQRL 0.10 has been released and the headline feature is supporting mutations for Iceberg tables. DataSQRL can now manage Apache Iceberg tables as sources and sinks.

Why is that a big deal? Up to this point, DataSQRL could read and write to Apache Iceberg tables, but you had to manage them explicitly. This new release makes it easy to share data through Apache Iceberg between DataSQRL pipelines.

Flink SQL Runner: Run Flink SQL Without JARs or Glue Code

· 3 min read
Matthias Broecheler
CEO of DataSQRL

Apache Flink has long been a powerhouse for streaming and batch data processing. And with the rise of Flink SQL, developers can now build sophisticated pipelines using a declarative language they already know. But getting Flink SQL applications into production still comes with friction: packaging JARs, managing connectors, injecting secrets, and wiring up deployment infrastructure.

FlinkSQL Runner >

Flink SQL Runner is here to change that. It's an open-source toolkit that simplifies development, deployment, and operation of Flink SQL applications—locally or in Kubernetes—without manual JAR assembly or scripting custom infrastructure pipelines.

Defining Data Interfaces with FlinkSQL

· 4 min read
Matthias Broecheler
CEO of DataSQRL

FlinkSQL is an amazing innovation in data processing: it packages the power of realtime stream processing within the simplicity of SQL. That means you can start with the SQL you know and introduce stream processing constructs as you need them.

FlinkSQL API Extension >

FlinkSQL adds the ability to process data incrementally to the classic set-based semantics of SQL. In addition, FlinkSQL supports source and sink connectors making it easy to ingest data from and move data to other systems. That's a powerful combination which covers a lot of data processing use cases.

In fact, it only takes a few extensions to FlinkSQL to build entire data applications. Let's see how that works.

Building Data APIs with FlinkSQL​

/*+ engine(kafka) */
CREATE TABLE UserTokens (
userid BIGINT NOT NULL,
tokens BIGINT NOT NULL,
request_time TIMESTAMP_LTZ(3) NOT NULL METADATA FROM 'timestamp'
);

/*+query_by_all(userid) */
TotalUserTokens := SELECT userid, sum(tokens) as total_tokens,
count(tokens) as total_requests
FROM UserTokens GROUP BY userid;

UserTokensByTime(userid BIGINT NOT NULL, fromTime TIMESTAMP NOT NULL, toTime TIMESTAMP NOT NULL):=
SELECT * FROM UserTokens WHERE userid = :userid
AND request_time >= :fromTime AND request_time < :toTime ORDER BY request_time DESC;

UsageAlert := SUBSCRIBE SELECT * FROM UserTokens WHERE tokens > 100000;

This script defines a sequence of tables. We introduce := as syntactic sugar for the verbose CREATE TEMPORARY VIEW syntax.

The UserTokens table does not have a configured connector, which mean we treat it as an API mutation endpoint connected to Flink via a Kafka topic that captures the events. This makes it easy to build APIs that capture user activity, transactions, or other types of events.

Why Temporal Join is Stream Processing’s Superpower

· 8 min read
Matthias Broecheler
CEO of DataSQRL

Stream processing technologies like Apache Flink introduce a new type of data transformation that’s very powerful: the temporal join. Temporal joins add context to data streams while being efficient and fast to execute.

Temporal Join >

This article introduces the temporal join, compares it to the traditional inner join, explains when to use it, and why it is a secret superpower.

Table of Contents:

Let's Uplevel Our Database Game: Meet DataSQRL

· 5 min read
Matthias Broecheler
CEO of DataSQRL

We need to make it easier to build data-driven applications. Databases are great if all your application needs is storing and retrieving data. But if you want to build anything more interesting with data - like serving users recommendations based on the pages they are visiting, detecting fraudulent transactions on your site, or computing real-time features for your machine learning model - you end up building a ton of custom code and infrastructure around the database.

You need a queue like Kafka to hold your events, a stream processor like Flink to process data, a database like Postgres to store and query the result data, and an API layer to tie it all together.

DataSQRL Logo >

And that’s just the price of admission. To get a functioning data layer, you need to make sure that all these components talk to each other and that data flows smoothly between them. Schema synchronization, data model tuning, index selection, query batching … all that fun stuff.

The point is, you need to do a ton of data plumbing if you want to build a data-driven application. All that data plumbing code is time-consuming to develop, hard to maintain, and expensive to operate.

We need to make building with data easier. That’s why we are sending out this call to action to uplevel our database game. Join us in figuring out how to simplify the data layer.

We have an idea to get us started: Meet DataSQRL.