<?xml version="1.0" encoding="utf-8"?><?xml-stylesheet type="text/xsl" href="rss.xsl"?>
<rss version="2.0" xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/">
    <channel>
        <title>DataSQRL Blog</title>
        <link>https://docs.datasqrl.com/main/blog</link>
        <description>DataSQRL Blog</description>
        <lastBuildDate>Tue, 29 Sep 2026 00:00:00 GMT</lastBuildDate>
        <docs>https://validator.w3.org/feed/docs/rss2.html</docs>
        <generator>https://github.com/jpmonette/feed</generator>
        <language>en</language>
        <item>
            <title><![CDATA[The Whole Pipeline in One File a Human Can Actually Read]]></title>
            <link>https://docs.datasqrl.com/main/blog/p4-human-understanding-one-sql-file</link>
            <guid>https://docs.datasqrl.com/main/blog/p4-human-understanding-one-sql-file</guid>
            <pubDate>Tue, 29 Sep 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[Part 4 of 4: How a data engineering harness eliminates the AI coding agent errors that survive review.]]></description>
            <content:encoded><![CDATA[
<p><em>Part 4 of 4: How a data engineering harness eliminates the AI coding agent errors that survive review.</em></p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="introduction">Introduction<a href="https://docs.datasqrl.com/main/blog/p4-human-understanding-one-sql-file#introduction" class="hash-link" aria-label="Direct link to Introduction" title="Direct link to Introduction" translate="no">​</a></h2>
<p>"Why do I need that? Isn't Claude Code good enough?"</p>
<img src="https://docs.datasqrl.com/img/blog/harness_p4_human_validation.jpeg" alt="A human reviewing a complete data pipeline expressed in one readable SQL file >|" width="50%">
<p>We get that question a lot. We are building an <a href="https://docs.datasqrl.com/" target="_blank" rel="noopener noreferrer" class="">open-source data engineering harness</a>, the tooling and guardrails that a coding agent uses to build data pipelines. Claude Code, Codex, and OpenCode already write plausible data pipeline code. So what is the harness for?</p>
<p>Parts 1, 2, and 3 were about generating the pipeline, validating it, and testing it. Part 4 is about the check that cannot be automated away: a person understanding what the pipeline actually means.</p>
<p>When an agent builds a conventional pipeline, the logic ends up spread across dozens or hundreds of files in two or three languages. Ingestion scripts, transformation jobs, table definitions, glue, API resolvers. Nobody holds that in their head. So the bugs where every line is correct and the meaning is wrong slip through, and the human review that was supposed to be the last line of defense collapses under fatigue.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="1-assumptions-that-never-meet">1. Assumptions that never meet<a href="https://docs.datasqrl.com/main/blog/p4-human-understanding-one-sql-file#1-assumptions-that-never-meet" class="hash-link" aria-label="Direct link to 1. Assumptions that never meet" title="Direct link to 1. Assumptions that never meet" translate="no">​</a></h2>
<p>When ingestion, transformation, and serving live in separate files written at different times, each stage quietly assumes a different grain, unit, or shape. Nothing forces those assumptions to meet.</p>
<p><em>Example:</em> the transform emits one row per order line, expecting something downstream to aggregate. The serving query assumes one row per order and sums nothing. The API reports inflated totals, and the two assumptions never appear on the same screen.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="2-the-same-business-rule-implemented-twice-differently">2. The same business rule, implemented twice, differently<a href="https://docs.datasqrl.com/main/blog/p4-human-understanding-one-sql-file#2-the-same-business-rule-implemented-twice-differently" class="hash-link" aria-label="Direct link to 2. The same business rule, implemented twice, differently" title="Direct link to 2. The same business rule, implemented twice, differently" translate="no">​</a></h2>
<p>Scattered pipelines duplicate logic. A filter, a currency conversion, a status definition. The copies drift, and two parts of the system end up disagreeing about what a number means.</p>
<p><em>Example:</em> "active customer" means status equals active in the ingestion filter and status is not closed in an API query. Two endpoints return different customer counts, and no file shows both definitions together.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="3-units-currency-and-time-zone-drifting-between-stages">3. Units, currency, and time zone drifting between stages<a href="https://docs.datasqrl.com/main/blog/p4-human-understanding-one-sql-file#3-units-currency-and-time-zone-drifting-between-stages" class="hash-link" aria-label="Direct link to 3. Units, currency, and time zone drifting between stages" title="Direct link to 3. Units, currency, and time zone drifting between stages" translate="no">​</a></h2>
<p>The most expensive bugs are not broken code. They are correct code about the wrong meaning. Cents treated as dollars. UTC treated as local. One currency treated as another. The mismatch lives across a file boundary where the unit is never written down.</p>
<p><em>Example:</em> the stream layer stores an amount in cents and a serving view applies a tax rate as if it were dollars, producing charges off by a factor of 100. Each file is internally consistent. The error lives only in the gap between them.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="4-hidden-coupling-that-breaks-on-edit">4. Hidden coupling that breaks on edit<a href="https://docs.datasqrl.com/main/blog/p4-human-understanding-one-sql-file#4-hidden-coupling-that-breaks-on-edit" class="hash-link" aria-label="Direct link to 4. Hidden coupling that breaks on edit" title="Direct link to 4. Hidden coupling that breaks on edit" translate="no">​</a></h2>
<p>In a sprawling codebase, files depend on each other through implicit contracts. A column name, an ordering, a nullability. An agent changing one file cannot see what else relies on it.</p>
<p><em>Example:</em> the agent renames a column in a transform script to tidy things up. A resolver in a different directory that referenced the old name starts returning null. Nothing connects the two at edit time.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="5-a-filter-applied-on-one-path-and-forgotten-on-another">5. A filter applied on one path and forgotten on another<a href="https://docs.datasqrl.com/main/blog/p4-human-understanding-one-sql-file#5-a-filter-applied-on-one-path-and-forgotten-on-another" class="hash-link" aria-label="Direct link to 5. A filter applied on one path and forgotten on another" title="Direct link to 5. A filter applied on one path and forgotten on another" translate="no">​</a></h2>
<p>When the same source feeds several downstream paths through different files, a filter or a deduplication applied on one path and missing on another lets inconsistent data through. The omission is invisible because the paths are never read together.</p>
<p><em>Example:</em> an "exclude test accounts" filter is applied in the reporting path but missing from the path feeding a fraud model. The model quietly scores synthetic accounts, and no single file shows that one path is filtered and the other is not.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="6-reviewer-fatigue">6. Reviewer fatigue<a href="https://docs.datasqrl.com/main/blog/p4-human-understanding-one-sql-file#6-reviewer-fatigue" class="hash-link" aria-label="Direct link to 6. Reviewer fatigue" title="Direct link to 6. Reviewer fatigue" translate="no">​</a></h2>
<p>Human review is the real bottleneck once agents write pipelines quickly. Asking a person to work through thousands of lines across many files and languages produces fatigue and missed bugs, which is the opposite of validation.</p>
<p><em>Example:</em> a reviewer signs off on a 2,000-line, fifteen-file change after skimming, missing a one-line grain error in file nine. Nobody sustains attention across that surface.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="7-handoffs-that-lose-the-design">7. Handoffs that lose the design<a href="https://docs.datasqrl.com/main/blog/p4-human-understanding-one-sql-file#7-handoffs-that-lose-the-design" class="hash-link" aria-label="Direct link to 7. Handoffs that lose the design" title="Direct link to 7. Handoffs that lose the design" translate="no">​</a></h2>
<p>When the logic is distributed across many files and languages, no single person holds the whole picture. Handoffs lose context, and the next engineer or agent changes things blind to the original design.</p>
<p><em>Example:</em> the engineer who built the pipeline leaves. Their replacement cannot work out why a particular deduplication exists, removes it, and reintroduces the duplicate-records bug it was silently preventing.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="8-requirements-you-cannot-actually-verify">8. Requirements you cannot actually verify<a href="https://docs.datasqrl.com/main/blog/p4-human-understanding-one-sql-file#8-requirements-you-cannot-actually-verify" class="hash-link" aria-label="Direct link to 8. Requirements you cannot actually verify" title="Direct link to 8. Requirements you cannot actually verify" translate="no">​</a></h2>
<p>The point of review is confirming the pipeline does what was asked. You cannot confirm that against logic you cannot read as a whole, so requirements get marked done on faith.</p>
<p><em>Example:</em> a compliance rule requires closed accounts to be excluded from a report. Verifying it means tracing the rule across ingestion, transformation, and serving files. The reviewer assumes it is handled and ships a violation.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="summary">Summary<a href="https://docs.datasqrl.com/main/blog/p4-human-understanding-one-sql-file#summary" class="hash-link" aria-label="Direct link to Summary" title="Direct link to Summary" translate="no">​</a></h2>
<p>None of these is a bug a validator can flag, because every individual line is correct. The error is in the meaning: how stages align, what a number represents, an assumption two files quietly disagree about. The only thing that catches a misalignment of meaning is a person who understands the pipeline, and understanding is impossible when the logic is scattered across hundreds of files in several languages.</p>
<p>One readable file collapses that surface. The entire data flow, from ingest to transform to store to serve, is expressed as one declarative SQL logic in the language most data engineers already read. That makes the whole pipeline comprehensible in one sitting, which turns review from a fatiguing rubber stamp back into real scrutiny, and frees the engineer to spend that attention on the requirements and the meaning that agents and validators cannot judge for themselves.</p>
<p>As pipeline creation speeds up, human understanding becomes the bottleneck. Keeping the logic in one readable place is how you keep that bottleneck open.</p>
<p>It is open source, feel free to try it out yourself: <a href="https://github.com/DataSQRL/sqrl" target="_blank" rel="noopener noreferrer" class="">DataSQRL on GitHub</a>.</p>]]></content:encoded>
            <category>technical</category>
            <category>DataSQRL</category>
        </item>
        <item>
            <title><![CDATA[Standard Integration Tests Obfuscate Common Errors]]></title>
            <link>https://docs.datasqrl.com/main/blog/p3-testing-framework</link>
            <guid>https://docs.datasqrl.com/main/blog/p3-testing-framework</guid>
            <pubDate>Mon, 28 Sep 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[Part 3 of 4: How a data engineering harness eliminates the AI coding agent errors that survive standard testing.]]></description>
            <content:encoded><![CDATA[
<p><em>Part 3 of 4: How a data engineering harness eliminates the AI coding agent errors that survive standard testing.</em></p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="introduction">Introduction<a href="https://docs.datasqrl.com/main/blog/p3-testing-framework#introduction" class="hash-link" aria-label="Direct link to Introduction" title="Direct link to Introduction" translate="no">​</a></h2>
<p>"Why do I need that? Isn't Claude Code good enough?"</p>
<img src="https://docs.datasqrl.com/img/blog/harness_p3_test_simulator.jpeg" alt="An event-time simulator replays data pipeline events at their original timestamps to test time-dependent behavior >|" width="50%">
<p>We get that question a lot. We are building an <a href="https://docs.datasqrl.com/" target="_blank" rel="noopener noreferrer" class="">open-source data engineering harness</a>, the tooling and guardrails that a coding agent uses to build data pipelines. Claude Code, Codex, and OpenCode already write plausible data pipeline code. So what is the harness for?</p>
<p>Parts 1 and 2 were about generating the pipeline correctly and validating it. Part 3 is about the accomplice in every failure so far: the test that passed.</p>
<p>An agent writes a test, runs it against a fixed snapshot of data, sees green, and ships. But a pipeline's hardest bugs only exist in motion which standard tests don't catch, leaving you to troubleshoot in production.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="1-testing-on-a-static-snapshot-with-nothing-updating-mid-run">1. Testing on a static snapshot with nothing updating mid-run<a href="https://docs.datasqrl.com/main/blog/p3-testing-framework#1-testing-on-a-static-snapshot-with-nothing-updating-mid-run" class="hash-link" aria-label="Direct link to 1. Testing on a static snapshot with nothing updating mid-run" title="Direct link to 1. Testing on a static snapshot with nothing updating mid-run" translate="no">​</a></h2>
<p>The default agent test loads a fixed dataset, runs the pipeline once, and checks the output. Every bug that requires data to change during execution is invisible. This is the largest blind spot, because it makes time-dependent correctness untestable by construction.</p>
<p><em>Example:</em> the time-dependent join from Part 2, where a transaction gets enriched with a later version of an account, cannot reproduce on a static snapshot. Nothing updates while the test runs, so the naive join and the correct join produce identical output. The test proves the wrong code right.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="2-tests-that-depend-on-the-wall-clock">2. Tests that depend on the wall clock<a href="https://docs.datasqrl.com/main/blog/p3-testing-framework#2-tests-that-depend-on-the-wall-clock" class="hash-link" aria-label="Direct link to 2. Tests that depend on the wall clock" title="Direct link to 2. Tests that depend on the wall clock" translate="no">​</a></h2>
<p>Agents write tests against the system time, so the result depends on how fast the machine ran and when the test happened to execute. It passes locally, flakes in CI, and certifies nothing.</p>
<p><em>Example:</em> a windowed aggregation tested against the current timestamp produces different bucket boundaries on every run. The snapshot never matches twice, so the agent "stabilizes" it by loosening the assertion until it no longer tests anything.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="3-no-late-or-out-of-order-data-in-the-test-set">3. No late or out-of-order data in the test set<a href="https://docs.datasqrl.com/main/blog/p3-testing-framework#3-no-late-or-out-of-order-data-in-the-test-set" class="hash-link" aria-label="Direct link to 3. No late or out-of-order data in the test set" title="Direct link to 3. No late or out-of-order data in the test set" translate="no">​</a></h2>
<p>Real streams deliver records late and out of order. An agent's test data is conveniently sorted and on time, so lateness handling and retraction logic never get exercised.</p>
<p><em>Example:</em> an event that arrives after its window closed should either update the result or be dropped, depending on the policy. The test never includes one, so a pipeline that quietly discards late data passes while losing records in production.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="4-races-between-interleaved-streams">4. Races between interleaved streams<a href="https://docs.datasqrl.com/main/blog/p3-testing-framework#4-races-between-interleaved-streams" class="hash-link" aria-label="Direct link to 4. Races between interleaved streams" title="Direct link to 4. Races between interleaved streams" translate="no">​</a></h2>
<p>When two streams feed a join, correctness depends on how they interleave. An agent testing each stream in isolation never sees the race.</p>
<p><em>Example:</em> a transaction arrives in the same instant an account flips from "active" to "frozen," and which value the enrichment picks decides whether a fraud check fires. The agent's test loads every account before any transaction, so the race never happens and the bug ships.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="5-idle-sources-and-stalled-time">5. Idle sources and stalled time<a href="https://docs.datasqrl.com/main/blog/p3-testing-framework#5-idle-sources-and-stalled-time" class="hash-link" aria-label="Direct link to 5. Idle sources and stalled time" title="Direct link to 5. Idle sources and stalled time" translate="no">​</a></h2>
<p>A source that goes quiet stalls the progress of event time and can freeze a join indefinitely. An agent's test keeps every source busy, so the stall never appears.</p>
<p><em>Example:</em> in production, one low-volume source stops emitting overnight. Time stops advancing and the joined output silently halts. An always-busy test set cannot produce that condition.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="6-updates-and-deletes-never-tested">6. Updates and deletes never tested<a href="https://docs.datasqrl.com/main/blog/p3-testing-framework#6-updates-and-deletes-never-tested" class="hash-link" aria-label="Direct link to 6. Updates and deletes never tested" title="Direct link to 6. Updates and deletes never tested" translate="no">​</a></h2>
<p>Updates and deletes propagate through a streaming pipeline as retractions of earlier results. An agent that only tests inserts never checks that a later update or delete correctly replaces what came before.</p>
<p><em>Example:</em> a record is inserted, then deleted upstream. A pipeline tested only on inserts keeps emitting the deleted entity, and the API serves a row that no longer exists.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="7-production-bugs-that-cannot-be-turned-into-a-test">7. Production bugs that cannot be turned into a test<a href="https://docs.datasqrl.com/main/blog/p3-testing-framework#7-production-bugs-that-cannot-be-turned-into-a-test" class="hash-link" aria-label="Direct link to 7. Production bugs that cannot be turned into a test" title="Direct link to 7. Production bugs that cannot be turned into a test" translate="no">​</a></h2>
<p>When something breaks in production, the fix starts with reproducing it. Without timestamp-accurate replay, a time-dependent failure cannot be reconstructed, so the fix is a guess.</p>
<p><em>Example:</em> a customer reports numbers that drifted last Tuesday. The agent cannot recreate Tuesday's exact sequence of events, patches speculatively, and never confirms the fix worked.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="8-the-rare-scenarios-only-time-control-can-construct">8. The rare scenarios only time control can construct<a href="https://docs.datasqrl.com/main/blog/p3-testing-framework#8-the-rare-scenarios-only-time-control-can-construct" class="hash-link" aria-label="Direct link to 8. The rare scenarios only time control can construct" title="Direct link to 8. The rare scenarios only time control can construct" translate="no">​</a></h2>
<p>The scenarios most likely to cause an outage are a leap in event time, a burst, a duplicate replay, a backfill colliding with live data. None of them can be built out of static fixtures, so they go untested until they happen.</p>
<p><em>Example:</em> a backfill of historical data is replayed alongside the live stream, and the pipeline double-counts because it was never tested against overlapping time ranges.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="summary">Summary<a href="https://docs.datasqrl.com/main/blog/p3-testing-framework#summary" class="hash-link" aria-label="Direct link to Summary" title="Direct link to Summary" translate="no">​</a></h2>
<p>The thread through all eight is time. The bugs that matter in a data pipeline are time-dependent, and an agent testing on static, on-time, single-stream snapshots is testing the one regime where time cannot hurt it.</p>
<p>A simulator changes what the test runs against. It executes the real generated deployment assets in a container and replays events at their original timestamps, so time-consistent semantics hold and the same input always produces the same output. This allows testing those things standard integration tests cannot: mid-run updates, late and out-of-order data, interleavings and races, idle sources, retraction sequences, and faithful replay of a production incident. Tests run deterministically, on a laptop, in a tight loop with the agent before anything is proposed for deployment.</p>
<p>A race condition that would take weeks to surface in production becomes a test case you write on purpose.</p>
<p>The event-time simulator is part of the open source data engineering harness, so you can replay your own scenarios and break the pipeline yourself: <a href="https://github.com/DataSQRL/sqrl" target="_blank" rel="noopener noreferrer" class="">DataSQRL on GitHub</a>.</p>]]></content:encoded>
            <category>technical</category>
            <category>DataSQRL</category>
        </item>
        <item>
            <title><![CDATA[The Bugs That Aren't in the Text: What Deep Relational Introspection Catches]]></title>
            <link>https://docs.datasqrl.com/main/blog/p2-validator-relational-introspection</link>
            <guid>https://docs.datasqrl.com/main/blog/p2-validator-relational-introspection</guid>
            <pubDate>Mon, 14 Sep 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[Part 2 of 4: How a data engineering harness eliminates the AI coding agent errors that survive review.]]></description>
            <content:encoded><![CDATA[
<p><em>Part 2 of 4: How a data engineering harness eliminates the AI coding agent errors that survive review.</em></p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="introduction">Introduction<a href="https://docs.datasqrl.com/main/blog/p2-validator-relational-introspection#introduction" class="hash-link" aria-label="Direct link to Introduction" title="Direct link to Introduction" translate="no">​</a></h2>
<p>"Why do I need that? Isn't Claude Code good enough?"</p>
<img src="https://docs.datasqrl.com/img/blog/harness_p2_deep_introspection.jpeg" alt="Deep introspection of the logic behind a data pipeline can reveal logical flaws" width="50%">
<p>We get that question a lot. We are building an <a href="https://docs.datasqrl.com/" target="_blank" rel="noopener noreferrer" class="">open-source data engineering harness</a>, the tooling and guardrails that a coding agent uses to build data pipelines. Claude Code, Codex, and OpenCode already write plausible data pipeline code. So what is the harness for?</p>
<p>Part 1 was about the seams between systems, and how a transpiler generates them deterministically. Part 2 is about a harder problem.</p>
<p>Some bugs have no correctness condition in the query text at all. The condition lives in the relationship between a query, the data it reads, and how that data changes over time. An agent reads SQL/code as text, and at the text level these bugs are invisible until the production deployment fails.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="1-no-indexes-for-the-paths-the-api-actually-queries">1. No indexes for the paths the API actually queries<a href="https://docs.datasqrl.com/main/blog/p2-validator-relational-introspection#1-no-indexes-for-the-paths-the-api-actually-queries" class="hash-link" aria-label="Direct link to 1. No indexes for the paths the API actually queries" title="Direct link to 1. No indexes for the paths the API actually queries" translate="no">​</a></h2>
<p>An agent generating a database schema writes the table definitions and stops. It never reasons about which columns the API filters and sorts on, so the serving database has no index matching its access paths. The pipeline is fast on a thousand test rows and collapses into full table scans at production volume.</p>
<p><em>Example:</em> An endpoint that looks up a customer by email scans every customer row on every call. Nothing about the query is wrong. Tail latency goes from milliseconds to seconds as the table grows.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="2-the-wrong-primary-key-or-none-at-all">2. The wrong primary key, or none at all<a href="https://docs.datasqrl.com/main/blog/p2-validator-relational-introspection#2-the-wrong-primary-key-or-none-at-all" class="hash-link" aria-label="Direct link to 2. The wrong primary key, or none at all" title="Direct link to 2. The wrong primary key, or none at all" translate="no">​</a></h2>
<p>Moving a changing table into the database needs two things. A primary key that matches the upstream key, and a write mode that replaces rather than appends. Agents omit the key, pick the wrong columns, or append what is logically a changelog. You get duplicate rows that accumulate forever, or the last writer wins at the wrong grain.</p>
<p><em>Example:</em> An accounts table built from a feed of database changes is written with no primary key. Every account update inserts a new row instead of replacing the previous one, and the API starts returning three conflicting records for the same account.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="3-treating-a-changelog-as-an-event-stream-or-the-reverse">3. Treating a changelog as an event stream, or the reverse<a href="https://docs.datasqrl.com/main/blog/p2-validator-relational-introspection#3-treating-a-changelog-as-an-event-stream-or-the-reverse" class="hash-link" aria-label="Direct link to 3. Treating a changelog as an event stream, or the reverse" title="Direct link to 3. Treating a changelog as an event stream, or the reverse" translate="no">​</a></h2>
<p>The most consequential property of a source is what kind of table it is. An append-only stream of events, a changelog where later rows replace earlier ones, and a lookup table are three different things, and which operations are valid depends entirely on which one you have. An agent that cannot see the distinction ingests a stream of updates as if each were a new event, or collapses an event stream that was never meant to be deduplicated.</p>
<p><em>Example:</em> The agent reads an account update feed as a plain append stream. Every status change becomes another "account," and the count of active accounts inflates by the number of edits each account has ever received.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="4-joins-whose-answer-depends-on-when-they-run">4. Joins whose answer depends on when they run<a href="https://docs.datasqrl.com/main/blog/p2-validator-relational-introspection#4-joins-whose-answer-depends-on-when-they-run" class="hash-link" aria-label="Direct link to 4. Joins whose answer depends on when they run" title="Direct link to 4. Joins whose answer depends on when they run" translate="no">​</a></h2>
<p>Join a stream to a table that changes over time with a plain join, and the result depends on when the join executed rather than on the state of the world when the event happened. It also changes retroactively whenever the other table updates. The query reads correctly. The time semantics are wrong.</p>
<p><em>Example:</em> Transactions are joined to accounts with a regular join, so each transaction is enriched with whatever account classification exists at processing time. A Tuesday transaction gets labeled with Thursday's account type, and enrichments already written quietly change after the fact.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="5-broken-time-attributes-and-stalled-watermarks">5. Broken time attributes and stalled watermarks<a href="https://docs.datasqrl.com/main/blog/p2-validator-relational-introspection#5-broken-time-attributes-and-stalled-watermarks" class="hash-link" aria-label="Direct link to 5. Broken time attributes and stalled watermarks" title="Direct link to 5. Broken time attributes and stalled watermarks" translate="no">​</a></h2>
<p>Stream processing needs every source to carry an event timestamp and a watermark, the signal that tells the engine how far event time has advanced. That time attribute then has to survive joins and aggregations. Agents leave the watermark off a source, fall back to processing time, or break the attribute partway through. Temporal joins stall, windows never fire, or results stop being reproducible.</p>
<p><em>Example:</em> One of three joined sources is declared with an incorrect watermark. The join waits forever for a time that never advances and the pipeline emits nothing. There is no error, just an empty result that reads as "no matching data."</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="6-not-null-contracts-a-nullable-source-cannot-honor">6. <code>NOT NULL</code> contracts a nullable source cannot honor<a href="https://docs.datasqrl.com/main/blog/p2-validator-relational-introspection#6-not-null-contracts-a-nullable-source-cannot-honor" class="hash-link" aria-label="Direct link to 6-not-null-contracts-a-nullable-source-cannot-honor" title="Direct link to 6-not-null-contracts-a-nullable-source-cannot-honor" translate="no">​</a></h2>
<p><code>NOT NULL</code> is a contract that has to hold across every system. An agent declares a column non-null downstream of a source or a join that can legitimately produce null values. One unmatched row then fails an insert or aborts a batch.</p>
<p><em>Example:</em> An enriched column is declared NOT NULL in the database, but the outer join that produces it leaves the value null for unmatched transactions. The first unmatched row throws a constraint violation and halts the writer.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="7-ordering-that-does-not-actually-order-anything">7. Ordering that does not actually order anything<a href="https://docs.datasqrl.com/main/blog/p2-validator-relational-introspection#7-ordering-that-does-not-actually-order-anything" class="hash-link" aria-label="Direct link to 7. Ordering that does not actually order anything" title="Direct link to 7. Ordering that does not actually order anything" translate="no">​</a></h2>
<p>Deduplication and "latest version" logic depend on an ordering column that really does increase over time. Stable results depend on an order being defined at all. An agent that deduplicates on a column which can move in either direction, or that relies on an implicit order, produces results that change between runs and keeps the wrong version of a row.</p>
<p><em>Example:</em> The agent deduplicates products by ordering on price, meaning "the latest one." Price is not monotonic, so the row that survives is whichever one carries the highest price rather than the most recent update. The current product record is simply wrong.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="8-state-that-grows-forever">8. State that grows forever<a href="https://docs.datasqrl.com/main/blog/p2-validator-relational-introspection#8-state-that-grows-forever" class="hash-link" aria-label="Direct link to 8. State that grows forever" title="Direct link to 8. State that grows forever" translate="no">​</a></h2>
<p>Streaming aggregations and joins hold state. An agent that groups on an ever-growing key, omits a retention bound, or aggregates at the wrong grain produces a pipeline whose memory grows without limit until it falls over. That happens long after the demo passed.</p>
<p><em>Example:</em> The agent keys a running aggregation by transaction id. State grows by one entry per transaction forever, and the job degrades and then crashes weeks into production.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="summary">Summary<a href="https://docs.datasqrl.com/main/blog/p2-validator-relational-introspection#summary" class="hash-link" aria-label="Direct link to Summary" title="Direct link to Summary" translate="no">​</a></h2>
<p>None of these bugs is visible in the query text. They live in the semantics of the data flow: how data changes over time, what its key is, when its timestamp is valid, which system can run which operator, how much state a step accumulates. An agent reviewing that text approves every one of them. So does a human reviewing it.</p>
<p>A validator doesn't review text. It parses the pipeline into a relational plan and traverses it the way a query optimizer would, and every fix above falls out of that one traversal. It infers keys and pushes them into the write configuration. It tracks timestamps and watermarks through each join and aggregation. It classifies every table as a stream, a changelog, or a lookup, and checks each operator against that. It propagates nullability, verifies that ordering columns actually increase, and reads the query workload to choose indexes. It rejects any plan that puts an operation on a system that cannot run it, or serves data that was never materialized.</p>
<p>When something fails, the validator names the table, the column, and a suggested fix. That matters more than it sounds. In our testing, an agent handed structured feedback fixes the problem far more reliably than one reasoning backward from an opaque runtime error.</p>
<p>The validator is open source, and you can add your own rules: <a href="https://github.com/DataSQRL/sqrl" target="_blank" rel="noopener noreferrer" class="">DataSQRL on GitHub</a>.</p>]]></content:encoded>
            <category>technical</category>
            <category>DataSQRL</category>
        </item>
        <item>
            <title><![CDATA[The Pipeline Compiles, the Demo Works, and the Seams Are Quietly Broken]]></title>
            <link>https://docs.datasqrl.com/main/blog/p1-broken-at-seams</link>
            <guid>https://docs.datasqrl.com/main/blog/p1-broken-at-seams</guid>
            <pubDate>Mon, 31 Aug 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[Part 1 of 4: How a data engineering harness eliminates the AI coding agent errors that survive review.]]></description>
            <content:encoded><![CDATA[
<p><em>Part 1 of 4: How a data engineering harness eliminates the AI coding agent errors that survive review.</em></p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="introduction">Introduction<a href="https://docs.datasqrl.com/main/blog/p1-broken-at-seams#introduction" class="hash-link" aria-label="Direct link to Introduction" title="Direct link to Introduction" translate="no">​</a></h2>
<p>"Why do I need that? Isn't Claude Code good enough?"</p>
<p>We get that question a lot. We are building an <a href="https://docs.datasqrl.com/" target="_blank" rel="noopener noreferrer" class="">open-source data engineering harness</a>, the tooling and guardrails that a coding agent uses to build data pipelines. Claude Code, Codex, and OpenCode already write plausible data pipeline code. So what is the harness for?</p>
<img src="https://docs.datasqrl.com/img/blog/harness_p1_broken_seams.jpeg" alt="A data pipeline spanning multiple systems with broken seams at the boundaries between them." width="50%">
<p>This series answers that. It looks at where general-purpose coding harnesses fail on data engineering work, and why those failures survive code review.</p>
<p>Part 1 is about the seams. A pipeline spans multiple data systems, and the same fact (a type, a name, an encoding) has to be restated in each one. Restating facts consistently across systems is strict rule-following. The thing doing the restating is a probabilistic model. It gets most of them right. The ones it gets wrong compile cleanly, pass the demo, and surface in production as overflowed numbers and fields that were never there.</p>
<p>A transpiler removes that entire class of error. It derives every boundary asset from one logical model, so the systems cannot disagree. The agent also writes less code, which means fewer tokens and fewer duplicated schema definitions cluttering its context.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="1-the-same-column-typed-three-different-ways">1. The same column, typed three different ways<a href="https://docs.datasqrl.com/main/blog/p1-broken-at-seams#1-the-same-column-typed-three-different-ways" class="hash-link" aria-label="Direct link to 1. The same column, typed three different ways" title="Direct link to 1. The same column, typed three different ways" translate="no">​</a></h2>
<p>A column in a pipeline exists at least three times. Once in the stream processor, once in the database, once in the API schema. Every system has its own type system, and an agent hand-maps between them one file at a time. A 64-bit integer becomes a 32-bit one at the API. A timestamp with a time zone becomes a bare local timestamp. A fixed-precision decimal lands as a floating point number. Nothing errors. The numbers are just wrong at the edges.</p>
<p><em>Example:</em> an agent maps a transaction amount, stored as a 64-bit count of cents, to a 32-bit API type. Every transaction above roughly $21M wraps to a negative number. Every test built on small sample values passes.</p>
<p>A transpiler infers the type once from the logical model and generates the matching type for each system, so the three representations cannot drift apart.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="2-query-dialects-that-look-the-same-and-mean-different-things">2. Query dialects that look the same and mean different things<a href="https://docs.datasqrl.com/main/blog/p1-broken-at-seams#2-query-dialects-that-look-the-same-and-mean-different-things" class="hash-link" aria-label="Direct link to 2. Query dialects that look the same and mean different things" title="Direct link to 2. Query dialects that look the same and mean different things" translate="no">​</a></h2>
<p>Every engine speaks its own dialect. They share a vocabulary and disagree in the details: function names, null handling, division, casts, date arithmetic. An agent writes a transformation in one dialect and a serving query in another. Both run, neither complains, but they compute different answers.</p>
<p><em>Example:</em> an integer division floors in the database and returns a fraction in the stream engine. A "share of total" computed in the stream layer and recomputed in a database view disagree for the same row, and nothing in either system flags it.</p>
<p>A transpiler generates each engine's dialect from one logical plan. A single declared computation is translated consistently instead of agent-written twice.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="3-a-field-added-in-one-place-but-not-everywhere">3. A field added in one place but not everywhere<a href="https://docs.datasqrl.com/main/blog/p1-broken-at-seams#3-a-field-added-in-one-place-but-not-everywhere" class="hash-link" aria-label="Direct link to 3. A field added in one place but not everywhere" title="Direct link to 3. A field added in one place but not everywhere" translate="no">​</a></h2>
<p>Adding one output column is a multi-system edit. The transformation, the connector that moves the data, the database table definition, the API type, and every endpoint that exposes it all have to gain the field together. An agent updates two or three of them and forgets the rest. The column either never reaches the database or arrives and stays invisible to every consumer.</p>
<p><em>Example:</em> you ask for a merchant category on an enriched transaction stream. The agent updates the transformation and the database table but not the API type. The field does not exist for any consumer, and there is no error to point at.</p>
<p>With a transpiler you add the column once. Every downstream asset is regenerated from that one definition.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="4-connector-configuration-where-two-systems-meet">4. Connector configuration where two systems meet<a href="https://docs.datasqrl.com/main/blog/p1-broken-at-seams#4-connector-configuration-where-two-systems-meet" class="hash-link" aria-label="Direct link to 4. Connector configuration where two systems meet" title="Direct link to 4. Connector configuration where two systems meet" translate="no">​</a></h2>
<p>Wherever two systems meet, a connection has to be configured: names, formats, key fields, write mode. This is hand-written glue, and an agent gets it wrong in ways that never fail loudly. A key field that does not match the table's key. An append-only write for something that is logically an update. A missing time attribute.</p>
<p><em>Example:</em> the agent configures the write from the stream layer into the database but never aligns the update key. Concurrent updates race, and the stored row flickers between values.</p>
<p>A transpiler treats every cut between systems as a generated connection. Both sides of every boundary derive from the same source of truth.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="5-an-api-contract-the-data-model-cannot-honor">5. An API contract the data model cannot honor<a href="https://docs.datasqrl.com/main/blog/p1-broken-at-seams#5-an-api-contract-the-data-model-cannot-honor" class="hash-link" aria-label="Direct link to 5. An API contract the data model cannot honor" title="Direct link to 5. An API contract the data model cannot honor" translate="no">​</a></h2>
<p>When the API schema is hand-written separately from the queries behind it, the contract drifts away from the data. A field marked as always present that can return nothing. A list marked non-empty that can be empty. An argument whose type does not match the query parameter. The API now advertises guarantees the pipeline cannot keep, and consumers get errors where they were promised values.</p>
<p><em>Example:</em> the agent declares a customer's order summary as always non-empty. A customer with zero orders violates the contract and fails the whole query.</p>
<p>A transpiler builds the API schema from the data model. When you supply your own schema, it raises a compile error wherever schema and model disagree, rather than letting the mismatch reach consumers.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="6-internal-columns-leaking-into-the-public-interface">6. Internal columns leaking into the public interface<a href="https://docs.datasqrl.com/main/blog/p1-broken-at-seams#6-internal-columns-leaking-into-the-public-interface" class="hash-link" aria-label="Direct link to 6. Internal columns leaking into the public interface" title="Direct link to 6. Internal columns leaking into the public interface" translate="no">​</a></h2>
<p>Tables carry columns the system fills in itself: generated ids, ingest timestamps, computed values. An interface that accepts writes must ask clients for the columns they actually supply and nothing else. An agent hand-writing that interface routinely includes the internal ones, so valid writes are rejected, or omits a real one, so writes arrive incomplete.</p>
<p><em>Example:</em> the write interface for an event asks for the generated id and the ingest timestamp. Every well-formed client insert is rejected for sending fields the server computes on its own.</p>
<p>A transpiler derives the input contract from the table definition and excludes system-populated columns, so the interface matches what a client is meant to send.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="7-the-same-operation-exposed-three-ways-three-different-shapes">7. The same operation exposed three ways, three different shapes<a href="https://docs.datasqrl.com/main/blog/p1-broken-at-seams#7-the-same-operation-exposed-three-ways-three-different-shapes" class="hash-link" aria-label="Direct link to 7. The same operation exposed three ways, three different shapes" title="Direct link to 7. The same operation exposed three ways, three different shapes" translate="no">​</a></h2>
<p>Operations are usually published over more than one protocol. An agent hand-coding each one produces views that disagree. An operation present on one protocol and missing on another. A route using the wrong method for its payload. A path whose arguments do not match the query signature. Consumers on one protocol get a broken endpoint, or none at all.</p>
<p><em>Example:</em> a query that takes a filter is exposed as a request that carries the filter in URL parameters. The filter is silently truncated, and every call returns the unfiltered result set.</p>
<p>A transpiler generates every protocol from one authoritative API model, so names, paths, and signatures stay aligned across all of them.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="8-identifier-casing-quoting-and-reserved-words">8. Identifier casing, quoting, and reserved words<a href="https://docs.datasqrl.com/main/blog/p1-broken-at-seams#8-identifier-casing-quoting-and-reserved-words" class="hash-link" aria-label="Direct link to 8. Identifier casing, quoting, and reserved words" title="Direct link to 8. Identifier casing, quoting, and reserved words" translate="no">​</a></h2>
<p>Systems disagree about case sensitivity, quoting, and which words are reserved. An agent carrying a name across them by hand introduces mismatches that break the mapping from API field to stored column. A mixed-case name folded to lowercase by one system no longer matches the code looking for it. A name that is a reserved word parses on one engine and fails on the next.</p>
<p><em>Example:</em> a column named <code>order</code> compiles in the stream layer and produces invalid table definitions in the database. The agent "fixes" it by quoting the name in one place and leaving it unquoted in another, and the API field resolves to nothing.</p>
<p>A transpiler applies one identifier mapping across every generated asset, so a logical name resolves to the right form on every system.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="summary">Summary<a href="https://docs.datasqrl.com/main/blog/p1-broken-at-seams#summary" class="hash-link" aria-label="Direct link to Summary" title="Direct link to Summary" translate="no">​</a></h2>
<p>We could keep going down our list of seam-breaking errors we have seen in agentic implementations, but you get the idea: Don't generate probabilistically what can be derived deterministically from a shared ground truth.</p>
<p>That's what the transpiler in the data engineering harness does. It takes one verified logical model as the ground truth and generates every boundary asset from it: the stream transformations, the database schema and queries, the connector configuration at each cut, the API schema, and every protocol endpoint. A type is inferred once and projected into each system. A field added once propagates to every asset. A name maps the same way everywhere. Both ends of a connection come from the same definition.</p>
<p>Because a compiler does that mapping instead of a language model, the entire category of "the systems disagree at the seam" does not need review to catch. It cannot be expressed. The agent decides what the pipeline computes. The transpiler guarantees the boundaries agree.</p>
<p>This has the additional benefit that there is a lot less code for the agent to generate at the seams. That means faster implementation, fewer tokens, less context pollution, and less technical debt.</p>
<p>Want to learn more about what exactly gets generated? The transpiler is part of the open source data engineering harness, so you can read exactly what gets generated: <a href="https://github.com/DataSQRL/sqrl" target="_blank" rel="noopener noreferrer" class="">DataSQRL on GitHub</a>.</p>]]></content:encoded>
            <category>technical</category>
            <category>DataSQRL</category>
        </item>
        <item>
            <title><![CDATA[AI in Data Engineering: Building Reliable Data Systems at Scale]]></title>
            <link>https://docs.datasqrl.com/main/blog/ai-data-engineering</link>
            <guid>https://docs.datasqrl.com/main/blog/ai-data-engineering</guid>
            <pubDate>Fri, 05 Jun 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[AI is transforming data engineering. Coding agents can now generate SQL transformations, configure connectors, and define API schemas in minutes rather than days. But here's the catch: a query that works perfectly on test data may fail catastrophically when confronted with late-arriving events, schema evolution, or terabyte-scale volumes.]]></description>
            <content:encoded><![CDATA[
<p>AI is transforming data engineering. Coding agents can now generate SQL transformations, configure connectors, and define API schemas in minutes rather than days. But here's the catch: a query that works perfectly on test data may fail catastrophically when confronted with late-arriving events, schema evolution, or terabyte-scale volumes.</p>
<p>How do we ensure that AI-generated data systems meet the rigorous <strong>non-functional requirements</strong> that production data platforms demand? This article presents our framework for integrating AI coding agents into data engineering workflows while maintaining data quality, reliability, governance, and trust.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-scaling-challenge">The Scaling Challenge<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#the-scaling-challenge" class="hash-link" aria-label="Direct link to The Scaling Challenge" title="Direct link to The Scaling Challenge" translate="no">​</a></h2>
<p>Organizations are targeting 3-5x productivity improvements through AI-assisted development. The velocity is real, but it creates an unsustainable burden on traditional data engineering practices:</p>
<ul>
<li class=""><strong>Manual code review</strong> can't scale when agents generate dozens of pipeline changes daily</li>
<li class=""><strong>Ad-hoc data quality checks</strong> miss subtle issues that only manifest at scale or over time</li>
<li class=""><strong>Tribal knowledge</strong> about production constraints doesn't transfer to AI agents</li>
<li class=""><strong>Integration testing</strong> becomes a bottleneck when deployment velocity outpaces validation capacity</li>
<li class=""><strong>Operations and troubleshooting</strong> overwhelm teams when they have to manage dozens of pipelines in production</li>
</ul>
<p>The fundamental problem is coding agents optimize for functional correctness (e.g. does the query return the right results?) while production data systems require a much broader set of guarantees.</p>
<p>Without systematic guardrails, AI-generated data pipelines work in demos but fail in production and overwhelm the data engineering teams that have to fill the gaps.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="a-data-quality-governance-framework">A Data Quality Governance Framework<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#a-data-quality-governance-framework" class="hash-link" aria-label="Direct link to A Data Quality Governance Framework" title="Direct link to A Data Quality Governance Framework" translate="no">​</a></h2>
<p>Successfully integrating AI into data engineering requires a governance framework that addresses three dimensions: <strong>transparency</strong>, <strong>validation</strong>, and <strong>operations</strong>.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="transparency-exposing-data-lineage-and-reasoning">Transparency: Exposing Data Lineage and Reasoning<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#transparency-exposing-data-lineage-and-reasoning" class="hash-link" aria-label="Direct link to Transparency: Exposing Data Lineage and Reasoning" title="Direct link to Transparency: Exposing Data Lineage and Reasoning" translate="no">​</a></h3>
<p>AI agents need to operate within systems that expose their reasoning and the data flows they create. This serves two purposes: enabling human oversight and providing feedback for iterative refinement.</p>
<p>What does transparency look like for data engineering?</p>
<ul>
<li class=""><strong>Clear Transformations</strong>: Well-defined data transformations that make it obvious <em>what</em> was transformed and <em>why</em></li>
<li class=""><strong>Data lineage tracking</strong>: Trace every output field back through transformations to source systems</li>
<li class=""><strong>Computational DAGs</strong>: Visualize how data flows from ingestion through processing to serving</li>
<li class=""><strong>Schema inference</strong>: Make explicit the types, keys, and timestamps the agent has assumed</li>
<li class=""><strong>Execution plans</strong>: Show which engines execute which computations</li>
</ul>
<p>When an agent proposes a data pipeline, we should see not just the SQL code but the complete picture: where data originates, how it transforms, what guarantees apply at each stage, and how it ultimately reaches consumers.</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token plain">=== SpendingTransactions</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">ID:     default_catalog.default_database.SpendingTransactions</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">Type:   stream</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">Stage:  flink</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">Inputs: Transactions, Accounts, AccountHolders</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">Annotations:</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"> - temporal-join: uses event-time semantics</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"> - stream-root: Transactions</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">Primary Key: transactionId, tx_time</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">Timestamp  : tx_time</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">Schema:</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"> - transactionId: BIGINT NOT NULL</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"> - amount: DECIMAL(10,2) NOT NULL</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"> - tx_time: TIMESTAMP_LTZ(3) *ROWTIME* NOT NULL</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"> - creditor_name: VARCHAR NOT NULL</span><br></div></code></pre></div></div>
<p>This representation combines inferences from the logical layer (primary keys, timestamps) with physical mappings (execution stage, inputs) to give us complete visibility into agent-generated pipelines.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="validation-real-time-quality-assessment">Validation: Real-Time Quality Assessment<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#validation-real-time-quality-assessment" class="hash-link" aria-label="Direct link to Validation: Real-Time Quality Assessment" title="Direct link to Validation: Real-Time Quality Assessment" translate="no">​</a></h3>
<p>Every agent action must pass through validation layers that assess correctness before execution. Here's what makes data pipelines different from application code: bugs often produce <em>silently wrong results</em>: queries that execute successfully but return incorrect data.</p>
<p>Effective validation operates at multiple levels:</p>
<p><strong>Logical Validation</strong></p>
<ul>
<li class="">Syntax and schema verification</li>
<li class="">Data type inference and consistency checking</li>
<li class="">Primary key and timestamp propagation validation</li>
<li class="">Table type verification (stream vs. state semantics)</li>
</ul>
<p><strong>Physical Validation</strong></p>
<ul>
<li class="">Engine capability matching (can Postgres execute this temporal join?)</li>
<li class="">Data type mapping consistency across engines</li>
<li class="">Connector configuration verification</li>
<li class="">Topological constraint satisfaction (data must reach database before API can serve it)</li>
</ul>
<p><strong>Semantic Validation</strong></p>
<ul>
<li class="">Business rule assertions</li>
<li class="">Data quality constraints (nullability, referential integrity, value ranges)</li>
<li class="">SLA verification (latency bounds, freshness guarantees)</li>
</ul>
<p>The validation system needs to provide actionable feedback when checks fail. Rather than cryptic error messages, agents need comprehensive context and suggested fixes:</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token plain">ERROR: Temporal join requires timestamp column on probe side</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">Table 'Transactions' is used in a temporal join with 'Accounts'</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">but lacks a rowtime attribute.</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">Suggestion: Add a timestamp column with WATERMARK definition:</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">  tx_time TIMESTAMP_LTZ(3),</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">  WATERMARK FOR tx_time AS tx_time - INTERVAL '5' SECOND</span><br></div></code></pre></div></div>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="operations-continuous-monitoring-and-autonomous-troubleshooting">Operations: Continuous Monitoring and Autonomous Troubleshooting<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#operations-continuous-monitoring-and-autonomous-troubleshooting" class="hash-link" aria-label="Direct link to Operations: Continuous Monitoring and Autonomous Troubleshooting" title="Direct link to Operations: Continuous Monitoring and Autonomous Troubleshooting" translate="no">​</a></h3>
<p>Building pipelines is only half the challenge. When you're managing dozens of AI-generated data pipelines in production, operations becomes the bottleneck. Traditional monitoring approaches (dashboards, manual alerts, runbooks) can't scale when pipeline count grows faster than team headcount.</p>
<p>The framework needs to support autonomous operations:</p>
<p><strong>Continuous Monitoring</strong></p>
<ul>
<li class="">Real-time data quality assertions that validate business rules on every record</li>
<li class="">SLA tracking that measures end-to-end latency from source event to API availability</li>
<li class="">Schema drift detection that catches upstream changes before they break downstream consumers</li>
<li class="">Resource utilization monitoring that identifies capacity issues before they cause failures</li>
</ul>
<p><strong>Autonomous Troubleshooting</strong></p>
<ul>
<li class="">Automatic correlation of symptoms to root causes using lineage information</li>
<li class="">Self-healing for common failure modes (connector reconnection, checkpoint recovery, partition rebalancing)</li>
<li class="">Intelligent alerting that groups related issues and suppresses noise</li>
<li class="">Runbook automation that executes standard remediation steps without human intervention</li>
</ul>
<p><strong>Observability Integration</strong></p>
<ul>
<li class="">Structured logging that links every record to its source transformation</li>
<li class="">Distributed tracing across the complete pipeline (Kafka → Flink → Postgres → API)</li>
<li class="">Metrics export to existing observability platforms (Prometheus, Datadog, CloudWatch)</li>
</ul>
<p>The goal is a team of three data engineers being able to operate dozens pipelines in production. That's only possible when the harness handles routine operations autonomously and escalates only the issues that genuinely require human judgment.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="capturing-data-engineering-expertise">Capturing Data Engineering Expertise<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#capturing-data-engineering-expertise" class="hash-link" aria-label="Direct link to Capturing Data Engineering Expertise" title="Direct link to Capturing Data Engineering Expertise" translate="no">​</a></h2>
<p>The effectiveness of AI in data engineering depends on systematically capturing and encoding human expertise. This falls into three categories: <strong>domain knowledge</strong>, <strong>operational patterns</strong>, and <strong>failure modes</strong>.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="domain-knowledge-encoding">Domain Knowledge Encoding<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#domain-knowledge-encoding" class="hash-link" aria-label="Direct link to Domain Knowledge Encoding" title="Direct link to Domain Knowledge Encoding" translate="no">​</a></h3>
<p>Data engineers carry implicit knowledge about their data domains: which fields contain PII, how upstream systems behave during maintenance windows, what query patterns consumers actually use. This knowledge needs to be made explicit for agents to leverage.</p>
<p>What works:</p>
<ul>
<li class=""><strong>Schema annotations</strong>: Capture business semantics beyond technical types</li>
<li class=""><strong>Data contracts</strong>: Formalize expectations between producers and consumers</li>
<li class=""><strong>Quality rules</strong>: Encode domain-specific validity constraints</li>
<li class=""><strong>Access patterns</strong>: Document how data gets queried in practice</li>
</ul>
<div class="language-sql codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-sql codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token comment" style="color:rgb(98, 114, 164)">-- Domain knowledge encoded in table definition</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token comment" style="color:rgb(98, 114, 164)">/** Customer spending transactions enriched with merchant details.</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token comment" style="color:rgb(98, 114, 164)">    PII: contains customer_id (indirect identifier)</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token comment" style="color:rgb(98, 114, 164)">    Freshness SLA: &lt; 5 minutes from source event</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token comment" style="color:rgb(98, 114, 164)">    Primary consumer: Fraud detection system (latency-sensitive)</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token comment" style="color:rgb(98, 114, 164)">*/</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">SpendingTransactions :</span><span class="token operator">=</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SELECT</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><br></div></code></pre></div></div>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="operational-pattern-libraries">Operational Pattern Libraries<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#operational-pattern-libraries" class="hash-link" aria-label="Direct link to Operational Pattern Libraries" title="Direct link to Operational Pattern Libraries" translate="no">​</a></h3>
<p>Production data pipelines exhibit recurring patterns: CDC deduplication, temporal enrichment joins, windowed aggregations, slowly changing dimensions. We encode them as reusable patterns to reduce the occurrence of subtle bugs when agents try to recreate them.</p>
<div class="language-sql codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-sql codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token comment" style="color:rgb(98, 114, 164)">-- Pattern: CDC deduplication to current state</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">Accounts :</span><span class="token operator">=</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">DISTINCT</span><span class="token plain"> AccountsCDC </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ON</span><span class="token plain"> account_id </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ORDER</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BY</span><span class="token plain"> update_time </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">DESC</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token comment" style="color:rgb(98, 114, 164)">-- Pattern: Temporal enrichment join</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">EnrichedTransactions :</span><span class="token operator">=</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SELECT</span><span class="token plain"> t</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token operator">*</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> a</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">account_type</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">Transactions</span><span class="token plain"> t</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">JOIN</span><span class="token plain"> Accounts </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FOR</span><span class="token plain"> SYSTEM_TIME </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">OF</span><span class="token plain"> t</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">tx_time a</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ON</span><span class="token plain"> t</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">account_id </span><span class="token operator">=</span><span class="token plain"> a</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">account_id</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token comment" style="color:rgb(98, 114, 164)">-- Pattern: Tumbling window aggregation</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">HourlyMetrics :</span><span class="token operator">=</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SELECT</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    window_start</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> </span><span class="token function" style="color:rgb(80, 250, 123)">COUNT</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token operator">*</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">as</span><span class="token plain"> event_count</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">TABLE</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain">TUMBLE</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">TABLE</span><span class="token plain"> Events</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> DESCRIPTOR</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain">event_time</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">INTERVAL</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">'1'</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">HOUR</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">GROUP</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BY</span><span class="token plain"> window_start</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><br></div></code></pre></div></div>
<p>These patterns encode not just the SQL syntax but the semantic intent and operational characteristics. When an agent needs CDC deduplication, it applies the established pattern rather than improvising a potentially incorrect solution.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="failure-mode-documentation">Failure Mode Documentation<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#failure-mode-documentation" class="hash-link" aria-label="Direct link to Failure Mode Documentation" title="Direct link to Failure Mode Documentation" translate="no">​</a></h3>
<p>Every production incident represents encoded knowledge about what can go wrong. Systematically capturing failure modes and their resolutions creates a corpus that agents can learn from:</p>
<ul>
<li class=""><strong>Symptoms</strong>: How the failure manifested (data delays, incorrect aggregates, schema mismatches)</li>
<li class=""><strong>Root cause</strong>: The underlying issue (late data handling, join key mismatch, type coercion)</li>
<li class=""><strong>Resolution</strong>: How we fixed it</li>
<li class=""><strong>Prevention</strong>: What validation or pattern would have caught this earlier</li>
</ul>
<p>Over time, this corpus enables agents to anticipate failure modes and proactively avoid them.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-data-engineering-harness">The Data Engineering Harness<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#the-data-engineering-harness" class="hash-link" aria-label="Direct link to The Data Engineering Harness" title="Direct link to The Data Engineering Harness" translate="no">​</a></h2>
<p>Implementing governance, validation, and expertise capture requires purpose-built infrastructure. We call this a <strong>data engineering harness</strong>: a system that provides the guardrails and feedback loops coding agents need to produce production-grade data systems.</p>
<img src="https://docs.datasqrl.com/img/diagrams/agentic/harness_overview.png" alt="Data engineering harness architecture" width="80%">
<p>The harness has three integrated components:</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="conceptual-framework">Conceptual Framework<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#conceptual-framework" class="hash-link" aria-label="Direct link to Conceptual Framework" title="Direct link to Conceptual Framework" translate="no">​</a></h3>
<p>The framework provides a precise vocabulary for reasoning about data transformations:</p>
<p><strong>Logical Layer</strong>: Expresses <em>what</em> transformations are needed using SQL extended with stream processing semantics. The declarative nature enables deep introspection. We can analyze query structure, infer schemas, and validate semantics.</p>
<p><strong>Physical Layer</strong>: Represents <em>how</em> data gets processed through engine assignment and configuration. A cost-based optimizer maps logical operations to physical engines (Flink, Kafka, Postgres, Iceberg) while respecting capability constraints.</p>
<p>Why does this separation matter? Agents should reason about business logic (logical layer) while the harness handles infrastructure complexity (physical layer). This division produces higher quality results by keeping agent context focused on the problem domain.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="comprehensive-validation">Comprehensive Validation<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#comprehensive-validation" class="hash-link" aria-label="Direct link to Comprehensive Validation" title="Direct link to Comprehensive Validation" translate="no">​</a></h3>
<p>Validation operates continuously throughout the development lifecycle:</p>
<ul>
<li class=""><strong>Compile-time</strong>: Schema consistency, type safety, semantic correctness</li>
<li class=""><strong>Plan-time</strong>: Physical feasibility, capability matching, optimization validity</li>
<li class=""><strong>Test-time</strong>: Functional correctness against known inputs and expected outputs</li>
<li class=""><strong>Deploy-time</strong>: Configuration validity, resource availability, dependency satisfaction</li>
<li class=""><strong>Run-time</strong>: Data quality assertions, SLA monitoring, anomaly detection</li>
</ul>
<p>Each validation stage produces structured feedback that agents consume for iterative refinement. The harness transforms validation failures into actionable guidance.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="real-world-feedback">Real-World Feedback<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#real-world-feedback" class="hash-link" aria-label="Direct link to Real-World Feedback" title="Direct link to Real-World Feedback" translate="no">​</a></h3>
<p>Static validation catches many issues but can't substitute for execution feedback. The harness provides two mechanisms for real-world validation:</p>
<p><strong>Simulation</strong>: Execute pipelines locally with timestamp-accurate event replay. The simulator runs the complete stack (Flink, Kafka, Postgres) in Docker, enabling agents to test against realistic data volumes and timing scenarios. Crucially, simulation is deterministic: the same inputs always produce the same outputs, enabling reliable regression testing.</p>
<p><strong>Production Telemetry</strong>: Monitor deployed pipelines and correlate observations back to source code. When latency increases or data quality degrades, the harness links metrics to specific transformations, enabling autonomous troubleshooting.</p>
<img src="https://docs.datasqrl.com/img/diagrams/agentic/feedback_loops.png" alt="Feedback loops from harness components back to coding agent" width="100%">
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="implementation-the-datasqrl-approach">Implementation: The DataSQRL Approach<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#implementation-the-datasqrl-approach" class="hash-link" aria-label="Direct link to Implementation: The DataSQRL Approach" title="Direct link to Implementation: The DataSQRL Approach" translate="no">​</a></h2>
<p>DataSQRL implements this harness architecture as an open-source framework. Here's how the abstract governance principles translate to concrete tooling.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="sql-as-the-logical-layer">SQL as the Logical Layer<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#sql-as-the-logical-layer" class="hash-link" aria-label="Direct link to SQL as the Logical Layer" title="Direct link to SQL as the Logical Layer" translate="no">​</a></h3>
<p>DataSQRL uses SQRL as the logical representation, which is SQL extended with stream processing from Flink SQL and interface definitions. Why SQL?</p>
<ul>
<li class=""><strong>LLM familiarity</strong>: Most models are extensively trained on SQL</li>
<li class=""><strong>Human readability</strong>: Engineers can verify agent output without learning new syntax</li>
<li class=""><strong>Mathematical foundation</strong>: Relational algebra enables rigorous validation</li>
<li class=""><strong>Declarative introspection</strong>: We can analyze and transform queries programmatically</li>
</ul>
<div class="language-sql codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-sql codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token comment" style="color:rgb(98, 114, 164)">-- Agent-generated pipeline in SQRL</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">IMPORT</span><span class="token plain"> banking_data</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token operator">*</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token comment" style="color:rgb(98, 114, 164)">-- Deduplicate CDC stream to current state</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">Accounts :</span><span class="token operator">=</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">DISTINCT</span><span class="token plain"> AccountsCDC </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ON</span><span class="token plain"> account_id </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ORDER</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BY</span><span class="token plain"> update_time </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">DESC</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token comment" style="color:rgb(98, 114, 164)">-- Enrich transactions with temporal join</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">SpendingTransactions :</span><span class="token operator">=</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SELECT</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    t</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token operator">*</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    h</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">name </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> creditor_name</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">Transactions</span><span class="token plain"> t</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">JOIN</span><span class="token plain"> Accounts </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FOR</span><span class="token plain"> SYSTEM_TIME </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">OF</span><span class="token plain"> t</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">tx_time a</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ON</span><span class="token plain"> t</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">credit_account_id </span><span class="token operator">=</span><span class="token plain"> a</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">account_id</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">JOIN</span><span class="token plain"> AccountHolders </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FOR</span><span class="token plain"> SYSTEM_TIME </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">OF</span><span class="token plain"> t</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">tx_time h</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ON</span><span class="token plain"> a</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">holder_id </span><span class="token operator">=</span><span class="token plain"> h</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">holder_id</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token comment" style="color:rgb(98, 114, 164)">-- Define API endpoint</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">SpendingByAccount</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain">account_id STRING </span><span class="token operator">NOT</span><span class="token plain"> </span><span class="token boolean">NULL</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"> :</span><span class="token operator">=</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SELECT</span><span class="token plain"> </span><span class="token operator">*</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> SpendingTransactions</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">WHERE</span><span class="token plain"> debit_account_id </span><span class="token operator">=</span><span class="token plain"> :account_id</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ORDER</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BY</span><span class="token plain"> tx_time </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">DESC</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><br></div></code></pre></div></div>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="deterministic-transpilation">Deterministic Transpilation<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#deterministic-transpilation" class="hash-link" aria-label="Direct link to Deterministic Transpilation" title="Direct link to Deterministic Transpilation" translate="no">​</a></h3>
<p>The mapping from logical to physical layer happens through deterministic transpilation, not agent generation. This eliminates an entire class of subtle bugs:</p>
<ul>
<li class="">Schema mismatches between engines</li>
<li class="">Incorrect data type coercions</li>
<li class="">Missing index structures</li>
<li class="">Inconsistent serialization formats</li>
</ul>
<p>The transpiler generates deployment artifacts (Flink plans, Kafka topics, Postgres schemas, GraphQL models) that are guaranteed consistent with the logical definition. Agents focus on business logic while the harness handles infrastructure integration.</p>
<img src="https://docs.datasqrl.com/img/diagrams/agentic/complete_framework.png" alt="Complete framework showing transpilation from SQL to multiple engines" width="100%">
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="neuro-symbolic-optimization">Neuro-Symbolic Optimization<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#neuro-symbolic-optimization" class="hash-link" aria-label="Direct link to Neuro-Symbolic Optimization" title="Direct link to Neuro-Symbolic Optimization" translate="no">​</a></h3>
<p>Certain data engineering tasks are better handled by dedicated optimizers than LLM reasoning. DataSQRL implements a neuro-symbolic approach: agents handle high-level design while specialized solvers handle constraint satisfaction.</p>
<ul>
<li class=""><strong>Query Optimization</strong>: Apache Calcite's Volcano optimizer rewrites queries for performance</li>
<li class=""><strong>Physical Planning</strong>: Cost-based optimizer assigns operations to engines while respecting topological constraints</li>
<li class=""><strong>Index Selection</strong>: Lattice-based optimizer selects index structures that support query access patterns</li>
</ul>
<p>Agents can provide hints to guide optimization (e.g. forcing specific engine assignments or partition keys) but the optimizer ensures constraint satisfaction. This leverages LLM strengths (reasoning under uncertainty, creative problem-solving) while delegating deterministic optimization to purpose-built systems.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="continuous-evaluation">Continuous Evaluation<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#continuous-evaluation" class="hash-link" aria-label="Direct link to Continuous Evaluation" title="Direct link to Continuous Evaluation" translate="no">​</a></h3>
<p>The harness supports continuous evaluation through automated testing infrastructure:</p>
<div class="language-sql codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-sql codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token comment" style="color:rgb(98, 114, 164)">/*+test */</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">TransactionEnrichmentTest :</span><span class="token operator">=</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SELECT</span><span class="token plain"> creditor_name</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> </span><span class="token function" style="color:rgb(80, 250, 123)">COUNT</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token operator">*</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">as</span><span class="token plain"> tx_count</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> SpendingTransactions</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">GROUP</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BY</span><span class="token plain"> creditor_name</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ORDER</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BY</span><span class="token plain"> creditor_name</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><br></div></code></pre></div></div>
<p>Test definitions execute against known inputs with expected outputs captured as snapshots. The simulator replays events with precise timestamps, enabling tests for complex scenarios:</p>
<ul>
<li class="">Late-arriving events and watermark handling</li>
<li class="">Out-of-order data processing</li>
<li class="">Race conditions in temporal joins</li>
<li class="">Schema evolution compatibility</li>
</ul>
<p>Tests provide immediate feedback to agents and spot regressions in CI/CD infrastructure.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="organizational-implications">Organizational Implications<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#organizational-implications" class="hash-link" aria-label="Direct link to Organizational Implications" title="Direct link to Organizational Implications" translate="no">​</a></h2>
<p>Successfully integrating AI into data engineering requires organizational adaptation beyond tooling. What shifts as AI assumes greater responsibility for pipeline development?</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="from-implementation-to-oversight">From Implementation to Oversight<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#from-implementation-to-oversight" class="hash-link" aria-label="Direct link to From Implementation to Oversight" title="Direct link to From Implementation to Oversight" translate="no">​</a></h3>
<p>As agents handle routine implementation, data engineers shift focus toward:</p>
<ul>
<li class=""><strong>Architecture review</strong>: Evaluating agent-proposed designs against organizational patterns</li>
<li class=""><strong>Pipeline auditing</strong>: Reviewing agent-generated pipelines for correctness, efficiency, and compliance</li>
<li class=""><strong>Domain encoding</strong>: Capturing business knowledge in schemas, contracts, and quality rules</li>
<li class=""><strong>Failure analysis</strong>: Investigating production issues and encoding learnings</li>
<li class=""><strong>Governance evolution</strong>: Refining validation rules and operational criteria</li>
</ul>
<p>This shift parallels the evolution in other engineering disciplines where automation handles routine tasks while humans focus on judgment-intensive decisions.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="from-manual-testing-to-continuous-validation">From Manual Testing to Continuous Validation<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#from-manual-testing-to-continuous-validation" class="hash-link" aria-label="Direct link to From Manual Testing to Continuous Validation" title="Direct link to From Manual Testing to Continuous Validation" translate="no">​</a></h3>
<p>Traditional data pipeline testing, manual verification against sample datasets, can't scale with AI-accelerated development. Organizations need to invest in:</p>
<ul>
<li class=""><strong>Comprehensive test suites</strong> that encode expected behavior across edge cases</li>
<li class=""><strong>Automated regression detection</strong> that flags behavioral changes between versions</li>
<li class=""><strong>Production monitoring</strong> that validates data quality continuously</li>
<li class=""><strong>Anomaly detection</strong> that identifies novel failure modes</li>
</ul>
<p>The testing burden shifts from per-deployment verification to continuous infrastructure maintenance.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="from-tribal-knowledge-to-encoded-expertise">From Tribal Knowledge to Encoded Expertise<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#from-tribal-knowledge-to-encoded-expertise" class="hash-link" aria-label="Direct link to From Tribal Knowledge to Encoded Expertise" title="Direct link to From Tribal Knowledge to Encoded Expertise" translate="no">​</a></h3>
<p>AI agents can't access knowledge that exists only in engineers' heads. Organizations need to systematically externalize:</p>
<ul>
<li class="">Data domain semantics and business rules</li>
<li class="">Operational patterns and anti-patterns</li>
<li class="">Historical failure modes and resolutions</li>
<li class="">Consumer requirements and SLAs</li>
</ul>
<p>This documentation effort has value beyond AI enablement: it improves onboarding, reduces key-person dependencies, and creates institutional memory that persists through team changes.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-path-forward">The Path Forward<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#the-path-forward" class="hash-link" aria-label="Direct link to The Path Forward" title="Direct link to The Path Forward" translate="no">​</a></h2>
<p>AI-assisted data engineering is here. Organizations that successfully integrate AI into their data platforms will achieve significant productivity gains while maintaining the reliability that production systems demand.</p>
<p>AI integration requires infrastructure, not just tools. Coding agents operating without guardrails produce pipelines that work in demos but fail in production. Agents operating within a purpose-built harness with comprehensive validation, real-world feedback, and encoded expertise produce pipelines that meet production requirements.</p>
<p>The <strong>data engineering harness</strong> represents this infrastructure: a system that provides the governance, validation, and feedback loops necessary for AI-assisted data platform automation.</p>
<p><a href="https://github.com/DataSQRL/sqrl" target="_blank" rel="noopener noreferrer" class="">DataSQRL</a> implements this harness as an open-source framework. You can customize it to encode your domain knowledge, integrate your validation rules, and build an automated data platform tailored to your requirements.</p>
<p>The question is no longer whether AI will transform data engineering, but how we adapt our practices, tooling, and teams to harness its potential while maintaining the trust that data consumers depend on.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="getting-started">Getting Started<a href="https://docs.datasqrl.com/main/blog/ai-data-engineering#getting-started" class="hash-link" aria-label="Direct link to Getting Started" title="Direct link to Getting Started" translate="no">​</a></h2>
<p>To explore AI-assisted data engineering with DataSQRL:</p>
<ol>
<li class=""><a class="" href="https://docs.datasqrl.com/main/docs/intro/getting-started">Build a project from scratch</a> to understand harness components</li>
<li class=""><a class="" href="https://docs.datasqrl.com/main/docs/intro/examples">Explore example projects</a> demonstrating common patterns</li>
<li class=""><a class="" href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness">Read about the harness architecture</a> for detailed technical background</li>
<li class=""><a href="https://github.com/DataSQRL/sqrl" target="_blank" rel="noopener noreferrer" class="">Contribute to the open-source project</a> to shape the future of AI-assisted data engineering</li>
</ol>]]></content:encoded>
            <category>Agentic</category>
            <category>Data Engineering</category>
        </item>
        <item>
            <title><![CDATA[Agentic Data Engineering Harness]]></title>
            <link>https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness</link>
            <guid>https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness</guid>
            <pubDate>Wed, 22 Apr 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[DataSQRL is an open-source data engineering harness that provides guardrails and feedback for AI coding agents to develop and operate data pipelines, data products, and data APIs autonomously.]]></description>
            <content:encoded><![CDATA[
<p>DataSQRL is an open-source data engineering harness that provides guardrails and feedback for AI coding agents to develop and operate data pipelines, data products, and data APIs autonomously.
You can customize DataSQRL as the foundation of your agentic data platform. Our goal is to develop DataSQRL into a comprehensive data engineering harness for data platform automation.</p>
<img src="https://docs.datasqrl.com/img/diagrams/agentic/harness_overview.png" alt="DataSQRL harness architecture showing coding agent with framework, validator, and simulator feedback loops >" width="50%">
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-a-data-engineering-harness">Why a Data Engineering Harness?<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#why-a-data-engineering-harness" class="hash-link" aria-label="Direct link to Why a Data Engineering Harness?" title="Direct link to Why a Data Engineering Harness?" translate="no">​</a></h2>
<p>Coding agents are transforming software development. Tools like Claude Code, Copilot, and Codex can generate application code, write tests, and even refactor entire codebases. But data engineering presents unique challenges that general-purpose coding agents struggle to address.</p>
<p>The difference lies in <strong>non-functional requirements</strong>. When you build a data pipeline, functional correctness (does the query return the right results?) is just the starting point. Production data systems must also deliver:</p>
<ul>
<li class=""><strong>Data Quality</strong>: Consistent, accurate data with proper handling of late-arriving events, duplicates, and schema evolution</li>
<li class=""><strong>Scalability</strong>: Performance that holds up as data volumes grow from gigabytes to terabytes</li>
<li class=""><strong>Governance</strong>: Lineage tracking, access controls, and audit trails for regulatory compliance</li>
<li class=""><strong>Reliability</strong>: Exactly-once semantics, failure recovery, and graceful degradation under load</li>
<li class=""><strong>Cost Efficiency</strong>: Optimal resource utilization across compute, storage, and network</li>
</ul>
<p>A coding agent can generate a SQL query that produces correct results on a test dataset. But will that query perform at scale? Does it handle late data correctly? Will it maintain data quality guarantees when upstream schemas change? These are the questions that justify data engineering as its own discipline and that general-purpose coding agents are not equipped to answer consistently.</p>
<p>A data engineering harness provides the guardrails, feedback loops, and domain-specific constraints that coding agents need to produce production-grade data systems. Without a harness, you get code that works in demos but fails in production. With a harness, you get data pipelines that embody decades of hard-won data engineering and domain-specific knowledge.</p>
<p>DataSQRL is that harness. It encodes the conceptual framework of data systems, validates implementations against data engineering best practices, and provides real-world feedback through simulation and production telemetry. The goal is to constrain coding agents into building data systems you'd actually trust to run in production.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-datasqrl-harness">The DataSQRL Harness<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#the-datasqrl-harness" class="hash-link" aria-label="Direct link to The DataSQRL Harness" title="Direct link to The DataSQRL Harness" translate="no">​</a></h2>
<img src="https://docs.datasqrl.com/img/diagrams/agentic/feedback_loops.png" alt="Feedback loops from framework, compiler, testing, validators, and runtime back to the coding agent" width="100%">
<p>For the purposes of automating data platforms, a comprehensive harness captures data schemas, data processing, and data serving to consumers. Specifically, we are building a harness for <em>non-transactional</em> data processing and serving.</p>
<p>The harness provides the frame of reference for implementing safe, reliable data processing systems.
It captures the knowledge from <a href="http://infolab.stanford.edu/~ullman/dscb.html" target="_blank" rel="noopener noreferrer" class="">Database Systems: The Complete Book</a> combined with 25 years of data engineering experience.</p>
<p>DataSQRL breaks the harness into <em>logical</em> and <em>physical</em> layers.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="logical-layer">Logical Layer<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#logical-layer" class="hash-link" aria-label="Direct link to Logical Layer" title="Direct link to Logical Layer" translate="no">​</a></h3>
<p>The logical layer expresses what data transformations are needed to produce the desired results.</p>
<p>An obvious choice for the logical layer is <a href="https://en.wikipedia.org/wiki/Relational_model" target="_blank" rel="noopener noreferrer" class="">Codd's relational model</a> and its most popular implementation <a href="https://en.wikipedia.org/wiki/SQL" target="_blank" rel="noopener noreferrer" class="">SQL</a>.</p>
<p>The relational model is widely adopted, proven, and provides a solid mathematical foundation. Most LLMs are trained on lots of SQL code and related documentation. And it is easy for humans to read. Modern versions of SQL (e.g., the SQL:2023 standard) support semi-structured data (JSON), polymorphic table functions, and complex pattern matching to address the messy reality of data platforms.</p>
<p>While the relational model and SQL are a good starting point, we need two additions to achieve the expressibility that modern data platforms require.</p>
<h4 class="anchor anchorTargetStickyNavbar_Vzrq" id="1-dataflow">1. Dataflow<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#1-dataflow" class="hash-link" aria-label="Direct link to 1. Dataflow" title="Direct link to 1. Dataflow" translate="no">​</a></h4>
<p>The relational model uses set semantics. That is inconvenient for representing data flows which are important for data pipelines.</p>
<p>Jennifer Widom's <a href="http://infolab.stanford.edu/~arvind/papers/cql-vldbj.pdf" target="_blank" rel="noopener noreferrer" class="">Continuous Query Language</a> extends the relational model with data streams and relational operators for moving between streams and sets.</p>
<p><a href="https://nightlies.apache.org/flink/flink-docs-release-2.2/docs/dev/table/overview/" target="_blank" rel="noopener noreferrer" class="">Flink SQL</a>, based on <a href="https://calcite.apache.org/" target="_blank" rel="noopener noreferrer" class="">Apache Calcite</a>, is the most widely adopted implementation of this extended relational model. That's why we use Flink SQL as the basis of the logical layer in DataSQRL.</p>
<p>Using a declarative language for the harness has a number of advantages from concise representation to deep introspection, but a practical shortcoming is the fact that some data transformations are easier to express imperatively. Flink SQL overcomes this by supporting <a href="https://nightlies.apache.org/flink/flink-docs-release-2.2/docs/dev/table/functions/udfs/" target="_blank" rel="noopener noreferrer" class="">user defined functions</a> and <a href="https://nightlies.apache.org/flink/flink-docs-release-2.2/docs/dev/table/functions/ptfs/" target="_blank" rel="noopener noreferrer" class="">custom table operators</a> in programming languages like Java. This gives us a logical layer grounded in relational algebra with flexible extensibility to express complex data transformations imperatively.</p>
<p>DataSQRL builds on Flink SQL and adds 1) concise syntax for common transformations, 2) dbt-style templating, and 3) modular file management and importing. These features help with context management for LLMs by reducing the size of the active context that needs to be maintained during implementation and refinement.</p>
<div class="language-sql codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-sql codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token comment" style="color:rgb(98, 114, 164)">-- Ingest data from connected systems</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">IMPORT</span><span class="token plain"> banking_data</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">AccountHoldersCDC</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain"> </span><span class="token comment" style="color:rgb(98, 114, 164)">-- CDC stream from masterdata</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">IMPORT</span><span class="token plain"> banking_data</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">AccountsCDC</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain">       </span><span class="token comment" style="color:rgb(98, 114, 164)">-- CDC stream from database</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">IMPORT</span><span class="token plain"> banking_data</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">Transactions</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain">      </span><span class="token comment" style="color:rgb(98, 114, 164)">-- Kafka topic for transactions</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token comment" style="color:rgb(98, 114, 164)">-- Convert the CDC stream of updates to the most recent version</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">Accounts       :</span><span class="token operator">=</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">DISTINCT</span><span class="token plain"> AccountsCDC </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ON</span><span class="token plain"> account_id </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ORDER</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BY</span><span class="token plain"> update_time </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">DESC</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">AccountHolders :</span><span class="token operator">=</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">DISTINCT</span><span class="token plain"> AccountHoldersCDC </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ON</span><span class="token plain"> holder_id </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ORDER</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BY</span><span class="token plain"> update_time </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">DESC</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token comment" style="color:rgb(98, 114, 164)">-- Enrich debit transactions with creditor information using time-consistent join</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">SpendingTransactions :</span><span class="token operator">=</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SELECT</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    t</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token operator">*</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    h</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">name </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> creditor_name</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    h</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">type</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> creditor_type</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">Transactions</span><span class="token plain"> t</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">         </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">JOIN</span><span class="token plain"> Accounts </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FOR</span><span class="token plain"> SYSTEM_TIME </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">OF</span><span class="token plain"> t</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">tx_time a</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">           </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ON</span><span class="token plain"> t</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">credit_account_id </span><span class="token operator">=</span><span class="token plain"> a</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">account_id</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">         </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">JOIN</span><span class="token plain"> AccountHolders </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FOR</span><span class="token plain"> SYSTEM_TIME </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">OF</span><span class="token plain"> t</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">tx_time h</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">           </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ON</span><span class="token plain"> a</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">holder_id </span><span class="token operator">=</span><span class="token plain"> h</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">holder_id</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><br></div></code></pre></div></div>
<p>We call this SQL dialect <strong>SQRL</strong>. You can <a class="" href="https://docs.datasqrl.com/main/docs/sqrl-language">read the documentation</a> for a complete reference of the SQRL language.</p>
<h4 class="anchor anchorTargetStickyNavbar_Vzrq" id="2-serving">2. Serving<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#2-serving" class="hash-link" aria-label="Direct link to 2. Serving" title="Direct link to 2. Serving" translate="no">​</a></h4>
<p>In addition to data processing, a critical function of data platforms is serving data to consumers as data streams, datasets, or data APIs. Data APIs, in particular, are becoming more important with the rise of operational analytics and MCP (Model Context Protocol) for making data accessible to AI agents.</p>
<p>To support data serving, DataSQRL adds support for endpoint definitions via table functions and explicit relationships.</p>
<p>Table functions are part of the SQL:2016 standard and return entire tables as result sets computed dynamically based on provided parameters. In DataSQRL, table functions can be defined as API entry points.</p>
<div class="language-sql codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-sql codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token comment" style="color:rgb(98, 114, 164)">/** Retrieve spending transactions within the given time-range.</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token comment" style="color:rgb(98, 114, 164)">  from_time (inclusive) and to_time (exclusive) must be RFC-3339 compliant date time.</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token comment" style="color:rgb(98, 114, 164)">*/</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">SpendingTransactionsByTime</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">  account_id STRING </span><span class="token operator">NOT</span><span class="token plain"> </span><span class="token boolean">NULL</span><span class="token plain"> METADATA </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">'auth.accountId'</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">  from_time </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">TIMESTAMP</span><span class="token plain"> </span><span class="token operator">NOT</span><span class="token plain"> </span><span class="token boolean">NULL</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">  to_time </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">TIMESTAMP</span><span class="token plain"> </span><span class="token operator">NOT</span><span class="token plain"> </span><span class="token boolean">NULL</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"> :</span><span class="token operator">=</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SELECT</span><span class="token plain"> </span><span class="token operator">*</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> SpendingTransactions</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">WHERE</span><span class="token plain"> debit_account_id </span><span class="token operator">=</span><span class="token plain"> :account_id</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">  </span><span class="token operator">AND</span><span class="token plain"> :from_time </span><span class="token operator">&lt;=</span><span class="token plain"> tx_time</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">  </span><span class="token operator">AND</span><span class="token plain"> :to_time </span><span class="token operator">&gt;</span><span class="token plain"> tx_time</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ORDER</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BY</span><span class="token plain"> tx_time </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">DESC</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><br></div></code></pre></div></div>
<p>Furthermore, DataSQRL allows for explicit relationship definitions between tables which are important for API-based data access where results need to include related entities like <em>most recent orders</em> or <em>recommendations for movie category</em>. The relational model does not support traversing through an entity-relationship model, which is usually handled by an object-relational mapping layer when exposing an API. To avoid that extra complexity and impedance mismatch in our logical layer, DataSQRL provides first-class support for relationships.</p>
<div class="language-sql codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-sql codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token comment" style="color:rgb(98, 114, 164)">-- Create a relationship between holder and accounts filtered by status</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">AccountHolders</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">accounts</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">status</span><span class="token plain"> STRING</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"> :</span><span class="token operator">=</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SELECT</span><span class="token plain"> </span><span class="token operator">*</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> Accounts a</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">WHERE</span><span class="token plain"> a</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">holder_id </span><span class="token operator">=</span><span class="token plain"> this</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">holder_id</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">  </span><span class="token operator">AND</span><span class="token plain"> a</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">status</span><span class="token plain"> </span><span class="token operator">=</span><span class="token plain"> :</span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">status</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ORDER</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BY</span><span class="token plain"> a</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">account_type </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ASC</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><br></div></code></pre></div></div>
<p>With the addition of access functions and relationships, the logical layer maps directly to the entity-relationship model of GraphQL which DataSQRL uses as the logical representation for API-based data retrieval. This gives DataSQRL a highly expressive interface with a simple extension of the logical layer which retains conceptual simplicity of the harness.</p>
<p>The <a class="" href="https://docs.datasqrl.com/main/docs/interface">interface documentation</a> provides more details on the serving layer of DataSQRL.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="physical-layer">Physical Layer<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#physical-layer" class="hash-link" aria-label="Direct link to Physical Layer" title="Direct link to Physical Layer" translate="no">​</a></h3>
<p>The physical layer represents <em>how</em> the data gets processed and served. It's a translation of the logical layer into executable code that runs on actual data systems.</p>
<img src="https://docs.datasqrl.com/img/diagrams/agentic/complete_framework.png" alt="Complete DataSQRL framework showing data flow from sources through processing, storage, and serving layers" width="100%">
<h4 class="anchor anchorTargetStickyNavbar_Vzrq" id="pipeline-architecture">Pipeline Architecture<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#pipeline-architecture" class="hash-link" aria-label="Direct link to Pipeline Architecture" title="Direct link to Pipeline Architecture" translate="no">​</a></h4>
<p>With <a href="https://db-engines.com/" target="_blank" rel="noopener noreferrer" class="">hundreds of database systems</a> and many more data infrastructure choices, it is a daunting challenge to construct a simple and coherent physical layer that is flexible enough to cover the diverse needs of data platforms.</p>
<p>After analyzing a wide range of data platforms, we identified that the vast majority of implementations combine multiple data systems from these categories:</p>
<ul>
<li class=""><strong>Database</strong>: for storing and querying data, e.g., PostgreSQL, MySQL, SQLServer, Apache Cassandra, Clickhouse, etc.<!-- -->
<ul>
<li class=""><strong>Table Formats and Query Engines</strong>: For analytic data, separating compute from storage can save money and support multiple consumers. DataSQRL conceptualizes this as a "disintegrated database" with table formats for storage (e.g., Apache Iceberg, DeltaLake, Apache Hudi) and query engines for access (e.g., Apache Spark, Apache Flink, Snowflake).</li>
</ul>
</li>
<li class=""><strong>Data Processor</strong>: for batch or realtime transformation of data, e.g., Apache Spark, Apache Flink, etc.</li>
<li class=""><strong>Log/Queue</strong>: for reliably capturing data and moving it between data systems, e.g., Apache Kafka, RedPanda, Kinesis, etc.</li>
<li class=""><strong>Server</strong>: for capturing and exposing data through an API<!-- -->
<ul>
<li class=""><strong>Cache</strong>: sits between server and database to speed up frequent queries over less-frequently changing data.</li>
</ul>
</li>
</ul>
<p>We call each data system an <em>engine</em> and the above categories <em>engine types</em>. When looking at data platform implementations at the level of engine types, we see about 15 patterns emerge (the 10 most popular are <a href="https://docs.datasqrl.com/main/assets/files/architecture_pattern_overview-deb4664cba8278dbcf6bbe5611801042.png" target="_blank" class="">documented here</a>) that arrange those engines in a directed-acyclic graph (DAG) of data processing.</p>
<p>Hence, we use a computational DAG that models the flow of data from source to interface as the basis of our physical layer. Each node in the DAG represents a logical computation mapped to be executed by an engine. Thus, the physical layer provides an integrated view of the entire data flow.</p>
<h4 class="anchor anchorTargetStickyNavbar_Vzrq" id="transpiler">Transpiler<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#transpiler" class="hash-link" aria-label="Direct link to Transpiler" title="Direct link to Transpiler" translate="no">​</a></h4>
<p>While the physical layer gives the AI control over what engine executes which computation, the actual mapping of logical to physical plan is done by a deterministic transpiler built in Apache Calcite. This avoids subtle bugs in data mapping and execution. The results of the transpilation are deployment assets which are executed by each engine. For example, the transpiler generates the database schema and queries for Postgres.</p>
<p>In the transpiler component, we make the following simplifying assumptions:</p>
<ul>
<li class="">The database engines support a version of SQL (e.g., PostgreSQL, T-SQL) or a subset thereof (e.g., Cassandra Query Language)</li>
<li class="">The data processor supports a SQL-based dialect (e.g., Spark SQL, Flink SQL)</li>
<li class="">The log engine is Apache Kafka compatible (e.g., RedPanda, Azure EventHub)</li>
<li class="">The server has a GraphQL execution engine.</li>
</ul>
<p>This modular architecture allows new engines to be added by conforming to the engine type interface and implementing the transpiler rules in Calcite where needed. At the same time, it abstracts much of the physical plan mapping complexity from the AI, which produces higher quality results and preserves context for higher-level reasoning.</p>
<h4 class="anchor anchorTargetStickyNavbar_Vzrq" id="configuration">Configuration<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#configuration" class="hash-link" aria-label="Direct link to Configuration" title="Direct link to Configuration" translate="no">​</a></h4>
<p>DataSQRL uses a <code>package.json</code> file to configure the engines used to execute a data pipeline. The configuration file defines the overall pipeline topology, the individual engine configurations, and the compiler configuration. One file controls how the physical layer is derived and executed, making it simple for the AI to experiment with and fine-tune the physical layer.</p>
<h4 class="anchor anchorTargetStickyNavbar_Vzrq" id="interface">Interface<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#interface" class="hash-link" aria-label="Direct link to Interface" title="Direct link to Interface" translate="no">​</a></h4>
<p>For the data serving interface, we use GraphQL schema as the physical representation which bidirectionally maps to the access functions, table schema, and relationships defined in the logical plan by naming convention. GraphQL fields are mapped to SQL or Kafka queries based on their respective definitions in SQRL. This allows the AI to fine-tune the API within the GraphQL schema.</p>
<p>Furthermore, REST and MCP APIs can be explicitly or implicitly defined through GraphQL operations. Implicit definition traverses the GraphQL schema from root query and mutation fields. Explicitly defined operations are provided as separate GraphQL files.</p>
<p>Using GraphQL as the physical representation for the API combines simplicity with flexibility while benefiting from the prevalence of GraphQL in LLM training data.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="validator">Validator<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#validator" class="hash-link" aria-label="Direct link to Validator" title="Direct link to Validator" translate="no">​</a></h2>
<p>The harness gives AI coding agents a frame of reference to reason about data pipeline and data product implementations. DataSQRL provides validation to support that reasoning and give users tools to ensure the correctness and quality of the generated pipelines and APIs.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="verification--introspection">Verification &amp; Introspection<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#verification--introspection" class="hash-link" aria-label="Direct link to Verification &amp; Introspection" title="Direct link to Verification &amp; Introspection" translate="no">​</a></h3>
<p>Verification and introspection complement the harness by reinforcing the concepts, rules, and dependencies. DataSQRL provides validation at 3 levels: the logical layer, physical layer, and deployment assets (the code that gets executed by the engines).</p>
<h4 class="anchor anchorTargetStickyNavbar_Vzrq" id="logical">Logical<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#logical" class="hash-link" aria-label="Direct link to Logical" title="Direct link to Logical" translate="no">​</a></h4>
<p>At the logical level, the DataSQRL compiler verifies syntax, schemas, and data flow semantics. This ensures that the data pipeline is logically coherent and that data integration points (e.g., between the SQL definitions and GraphQL schema) are consistent.</p>
<p>One of the benefits of using relational algebra as the basis for our harness is the ability to run rules and deep traversals over the operators in the relational algebra tree. The DataSQRL compiler uses Apache Calcite's rule and RelNode traversal framework to validate timestamp propagation, infer primary keys and data types, validate table types, and more. This validation component can be extended with custom rules to validate domain-specific semantics and constraints.</p>
<p>The validation component was designed to provide comprehensive context and suggested fixes for validation errors. In our testing, this produces significantly better results compared to the AI coding agent having to look up and reason about encountered errors.</p>
<h4 class="anchor anchorTargetStickyNavbar_Vzrq" id="physical">Physical<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#physical" class="hash-link" aria-label="Direct link to Physical" title="Direct link to Physical" translate="no">​</a></h4>
<p>On compilation, DataSQRL produces the computational data flow DAG that represents the physical layer. DataSQRL generates a visual representation as shown above for human validation as well as a concise textual representation that is consumed by coding agents as feedback on their proposed solutions and to reinforce the conceptual data flow of the harness.</p>
<div class="language-text codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-text codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token plain">=== CustomerTransaction</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">ID:     default_catalog.default_database.CustomerTransaction</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">Type:   stream</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">Stage:  flink</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">Inputs: default_catalog.default_database._CardAssignment, default_catalog.default_database._Merchant, default_catalog.sources.Transaction</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">Annotations:</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"> - stream-root: Transaction</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">Primary Key: transactionId, time</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">Timestamp  : time</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">Schema:</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"> - transactionId: BIGINT NOT NULL</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"> - cardNo: VARCHAR(2147483647) CHARACTER SET "UTF-16LE" NOT NULL</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"> - time: TIMESTAMP_LTZ(3) *ROWTIME* NOT NULL</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"> - amount: DOUBLE NOT NULL</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"> - merchantName: VARCHAR(2147483647) CHARACTER SET "UTF-16LE" NOT NULL</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"> - category: VARCHAR(2147483647) CHARACTER SET "UTF-16LE" NOT NULL</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"> - customerId: BIGINT NOT NULL</span><br></div></code></pre></div></div>
<p>This representation of the physical layer combines the inferences from the logical layer with the mapping to execution engines to provide a source-to-interface definition of the data flow.</p>
<p>Validation at the physical level ensures that data type mappings are consistent and that the engine assignments are valid, i.e., that an assigned engine has the capabilities to execute a particular operator. DataSQRL uses a capabilities component that extracts all requirements from an operator (e.g., temporal join, or a particular function execution) and validates that the engine supports the corresponding capabilities.</p>
<h4 class="anchor anchorTargetStickyNavbar_Vzrq" id="deployment-assets">Deployment Assets<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#deployment-assets" class="hash-link" aria-label="Direct link to Deployment Assets" title="Direct link to Deployment Assets" translate="no">​</a></h4>
<p>The executable deployment assets are transpiled from the physical layer. Since the transpilation is deterministic, this yields better results than letting the coding agent generate them, and it keeps the harness concise. However, we generate all deployment assets in a text representation that the coding agent can easily consume as another source of feedback. This is particularly useful during troubleshooting where the deployment assets are the ultimate source of truth of what is being executed and allow the agent to reason "backwards" to the logical layer and how to fix it.</p>
<p>Specifically, we generate:</p>
<ul>
<li class=""><strong>Database</strong>: The database schema, index structures, and (parameterized) SQL queries for all views and API entrypoints.</li>
<li class=""><strong>Data Processor</strong>: The optimized physical plan and compiled execution plan.</li>
<li class=""><strong>Log</strong>: The topic definitions and filter predicates.</li>
<li class=""><strong>Server</strong>: The mapping from GraphQL fields to database or Kafka queries as well as operation definitions. Also, the GraphQL schema if it is not provided.</li>
</ul>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="optimization">Optimization<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#optimization" class="hash-link" aria-label="Direct link to Optimization" title="Direct link to Optimization" translate="no">​</a></h3>
<p>While LLMs' reasoning ability under uncertainty is outstanding, we have found LLMs to perform worse and less consistently on deterministic optimization and constraint satisfaction problems. This finding is supported by a rich body of research in <a href="https://en.wikipedia.org/wiki/Neuro-symbolic_AI" target="_blank" rel="noopener noreferrer" class="">neuro-symbolic AI</a> which researches the integration of neural networks (like LLMs) with symbolic computation (e.g., solvers, planners) and has documented how neural networks alone fall short for such tasks.</p>
<p>DataSQRL follows the neuro-symbolic approach and provides 3 types of planners for deterministic sub-tasks in the implementation and maintenance of data pipelines:</p>
<h4 class="anchor anchorTargetStickyNavbar_Vzrq" id="query-optimization">Query Optimization<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#query-optimization" class="hash-link" aria-label="Direct link to Query Optimization" title="Direct link to Query Optimization" translate="no">​</a></h4>
<p>Query rewriting and optimization is a well-established technique for producing high-performing physical plans from relational algebra. DataSQRL relies on Apache Calcite's Volcano optimizer and HEP rule engine for this purpose.</p>
<h4 class="anchor anchorTargetStickyNavbar_Vzrq" id="physical-planning">Physical Planning<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#physical-planning" class="hash-link" aria-label="Direct link to Physical Planning" title="Direct link to Physical Planning" translate="no">​</a></h4>
<p>The physical plan DAG is subject to a number of constraints forced by the real-world constraints of physical data movement. For example, when the API needs to serve data on request, that data must first be available in the database. It cannot be served directly from a data processor. These topological constraints combined with the capabilities of individual engines render many AI-proposed solutions invalid.</p>
<p>Hence, we implement a physical planner that uses a cost model with a greedy heuristic to assign logical operators to engines in a way that is consistent. The AI can provide hints to force the assignment of certain operators to specific engines which are added as constraints to the optimizer. This gives the AI control over allocations but shifts the burden of constraint satisfaction to a dedicated solver.</p>
<h4 class="anchor anchorTargetStickyNavbar_Vzrq" id="index-selection">Index Selection<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#index-selection" class="hash-link" aria-label="Direct link to Index Selection" title="Direct link to Index Selection" translate="no">​</a></h4>
<p>Efficiently querying data in the database or table format requires index structures (or partition + sort keys) that support the access paths to the data. Otherwise, we execute inefficient table scans.</p>
<p>Index structure selection is another optimization problem that is better handled by a dedicated optimizer. We use an adaptation of <a href="https://web.eecs.umich.edu/~jag/eecs584/papers/implementing_data_cube.pdf" target="_blank" rel="noopener noreferrer" class="">Ullman et al's lattice framework for data cube selection</a> since data cube selection and index selection are related problems with different optimization functions.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="real-world-feedback">Real World Feedback<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#real-world-feedback" class="hash-link" aria-label="Direct link to Real World Feedback" title="Direct link to Real World Feedback" translate="no">​</a></h2>
<p>A harness with complementary verification and introspection provides the foundation of a concise model for data pipelines with feedback on proposed solutions. However, that feedback is limited to the plan and does not account for the complexities of actual execution. Real-world feedback is critical for iterative refinement of production-grade implementations and troubleshooting issues that arise in operation.</p>
<p>DataSQRL provides two sources of real-world feedback: a simulator that's used at implementation time and telemetry collection from production deployments that captures the operational status of the pipeline.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="simulator">Simulator<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#simulator" class="hash-link" aria-label="Direct link to Simulator" title="Direct link to Simulator" translate="no">​</a></h3>
<p>The DataSQRL simulator executes the configured engines with the generated deployment assets within a Docker environment. The simulator can replay events and records at their original timestamp, allowing for deterministic reproducibility of real-world scenarios. This is important for creating realistic test cases as well as reproducing production issues for troubleshooting and regression testing.</p>
<p>By capturing and faithfully replaying records at their original timestamp, the simulator ensures time-consistent semantics of data flows and makes it simple to construct complex test cases for scenarios like race conditions.</p>
<p>Simulation is important in agentic coding workflows because it allows the agent to execute and refine the implementation in a feedback loop that is executed locally and can simulate scenarios that only occur rarely in production.</p>
<p>Read more about invoking the <a class="" href="https://docs.datasqrl.com/main/docs/compiler#test-command">simulator</a> and writing reproducible test cases.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="operations-and-telemetry">Operations and Telemetry<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#operations-and-telemetry" class="hash-link" aria-label="Direct link to Operations and Telemetry" title="Direct link to Operations and Telemetry" translate="no">​</a></h3>
<p>The most important source of real-world feedback is observing the deployed data pipeline in a production environment (or a closely approximated pre-prod environment). Observability is critical for assessing the health of the pipeline and troubleshooting any issues that may occur.</p>
<p>Logs and telemetry collection is a well-established practice for DevOps. What DataSQRL adds is the ability to link observed data back to the physical computation DAG so the agent can accurately reason about cause and effect. For data pipelines that execute across multiple engines, many complex errors arise at system boundaries and require reasoning across multiple systems. For example, an issue in the data processing layer may cause excessive writes to the database, degrading overall performance. To automate such troubleshooting, we need to correlate observations back to the physical data flow and logical layer.</p>
<p>DataSQRL currently assumes production operation in Kubernetes or Docker and provides hooks for extracting logs and telemetry data. That data is correlated back to the SQL code and configuration defining the pipeline via the deployment assets, allowing coding agents to reason about effective solutions for troubleshooting production issues autonomously.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="summary">Summary<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#summary" class="hash-link" aria-label="Direct link to Summary" title="Direct link to Summary" translate="no">​</a></h2>
<p>Data engineering is entering a new era of automation. Coding agents can now write SQL, configure pipelines, and deploy data systems. But without proper guardrails, they produce solutions that fail under the weight of production requirements.</p>
<p>DataSQRL is the data engineering harness that ensures AI coding agents produce high-quality pipelines. DataSQRL encodes decades of data engineering knowledge into a structured framework that coding agents can leverage to build production-grade data systems.</p>
<p>The harness provides three critical capabilities:</p>
<ol>
<li class="">
<p><strong>A Conceptual Framework</strong> grounded in relational algebra and stream processing that gives agents a precise vocabulary for reasoning about data transformations, with logical and physical layers that separate <em>what</em> from <em>how</em>.</p>
</li>
<li class="">
<p><strong>Comprehensive Validation</strong> at every level ensures that agent-generated code meets data engineering standards before it reaches production: from syntax and schema validation through physical plan verification to deployment asset generation.</p>
</li>
<li class="">
<p><strong>Real-World Feedback Loops</strong> through simulation and production telemetry that enable agents to iteratively refine implementations based on actual execution behavior, not just static analysis.</p>
</li>
</ol>
<p>For any organization pursuing data platform automation, a data engineering harness is foundational to avoid shipping data pipelines that fail in production. Without it, you're asking general-purpose coding agents to navigate the complex constraints of distributed data systems without a map. With DataSQRL, you're equipping them with the domain expertise and multiple sources of feedback to succeed.</p>
<p><a href="https://github.com/DataSQRL/sqrl" target="_blank" rel="noopener noreferrer" class="">DataSQRL is open-source</a> so you can customize it to build a self-driving data platform tailored to your organization.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="getting-started">Getting Started<a href="https://docs.datasqrl.com/main/blog/agentic-data-engineering-harness#getting-started" class="hash-link" aria-label="Direct link to Getting Started" title="Direct link to Getting Started" translate="no">​</a></h2>
<p>To try out DataSQRL:</p>
<ol>
<li class=""><a class="" href="https://docs.datasqrl.com/main/docs/intro/getting-started">Build a project from scratch with DataSQRL</a> to see how the components of DataSQRL work</li>
<li class="">Explore the <a href="https://github.com/datasqrl-colab/finance-demo" target="_blank" rel="noopener noreferrer" class="">AI generated data products</a> for a fictional bank based on <a href="https://github.com/datasqrl-colab/finance-data-catalog-demo" target="_blank" rel="noopener noreferrer" class="">the bank's catalog definition</a> as an example.</li>
<li class=""><a class="" href="https://docs.datasqrl.com/main/docs/intro">Read the documentation</a></li>
<li class=""><a href="https://github.com/DataSQRL/sqrl" target="_blank" rel="noopener noreferrer" class="">Check out the open-source project on GitHub</a></li>
</ol>]]></content:encoded>
            <category>Release</category>
        </item>
        <item>
            <title><![CDATA[0.10 Release: Iceberg Mutations]]></title>
            <link>https://docs.datasqrl.com/main/blog/release-datasqrl-0-10</link>
            <guid>https://docs.datasqrl.com/main/blog/release-datasqrl-0-10</guid>
            <pubDate>Fri, 10 Apr 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[" width="40%"/>]]></description>
            <content:encoded><![CDATA[
<img src="https://docs.datasqrl.com/img/blog/release_0.10.0.png" alt="SQRL 0.10 Release >" width="40%">
<p>DataSQRL 0.10 has been <a href="https://github.com/DataSQRL/sqrl/releases/tag/0.10.0" target="_blank" rel="noopener noreferrer" class="">released</a> and the headline feature is supporting mutations for Iceberg tables. DataSQRL can now manage Apache Iceberg tables as sources and sinks.</p>
<p>Why is that a big deal? Up to this point, DataSQRL could read and write to Apache Iceberg tables, but you had to manage them explicitly. This new release makes it easy to share data through Apache Iceberg between DataSQRL pipelines.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="data-fast-and-slow">Data Fast and Slow<a href="https://docs.datasqrl.com/main/blog/release-datasqrl-0-10#data-fast-and-slow" class="hash-link" aria-label="Direct link to Data Fast and Slow" title="Direct link to Data Fast and Slow" translate="no">​</a></h2>
<p>Let's back up a bit. Before 0.10 you could create tables in DataSQRL like this:</p>
<div class="language-sql codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-sql codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">CREATE</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">TABLE</span><span class="token plain"> Clickstream </span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    user_id STRING</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    event_time TIMESTAMP_LTZ</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token number">3</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"> </span><span class="token operator">NOT</span><span class="token plain"> </span><span class="token boolean">NULL</span><span class="token plain"> METADATA </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">'timestamp'</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    ad_id STRING</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    WATERMARK </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FOR</span><span class="token plain"> event_time </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> event_time </span><span class="token operator">-</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">INTERVAL</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">'5'</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SECOND</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><br></div></code></pre></div></div>
<p>If you don't specify a connector explicitly, DataSQRL will manage the table for you, create the Kafka topics, and wire everything up. This allows you to expose mutation endpoints in GraphQL that store mutation events in Kafka and process them in Flink, making it easy to build event-driven microservices.</p>
<p>Kafka is great if you need data fast – in milliseconds. But it requires a separate Kafka cluster, which is challenging to maintain and costly. For many data use cases, you don't need data that quickly. Minutes and hours are just fine.</p>
<p>That's where Apache Iceberg shines. It's a table format for sharing data that does not require a separate data system to operate. All you need is local or cloud storage.</p>
<p>With the 0.10 release, DataSQRL supports Apache Iceberg for managed tables. Use Kafka for the fast data and Iceberg for the slow data. This allows you to balance speed with operational simplicity and cost.</p>
<p>To make this possible, DataSQRL 0.10 introduces a breaking change: You have to explicitly annotate where you want the tables you create to be persisted: iceberg or kafka. Use the <code>engine</code> SQL hint:</p>
<div class="language-sql codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-sql codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token comment" style="color:rgb(98, 114, 164)">/*+ engine(iceberg) */</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">CREATE</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">TABLE</span><span class="token plain"> Clickstream </span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    user_id STRING</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    event_time TIMESTAMP_LTZ</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token number">3</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"> </span><span class="token operator">NOT</span><span class="token plain"> </span><span class="token boolean">NULL</span><span class="token plain"> METADATA </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">'timestamp'</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    ad_id STRING</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    WATERMARK </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FOR</span><span class="token plain"> event_time </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> event_time </span><span class="token operator">-</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">INTERVAL</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">'5'</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SECOND</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><br></div></code></pre></div></div>
<p>That's it. And make sure to upgrade your existing DataSQRL projects with <code>/*+ engine(kafka) */</code> on any managed tables as you migrate to 0.10.</p>
<p>There are a lot more goodies in the 0.10 release. Check out <a href="https://github.com/DataSQRL/sqrl/releases/tag/0.10.0" target="_blank" rel="noopener noreferrer" class="">the complete changelog</a> for details. Shout out to Ferenc for driving this release and to our newest team members Wellington and Mate for making their first contributions to DataSQRL. Thank you!</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="whats-next">What's Next?<a href="https://docs.datasqrl.com/main/blog/release-datasqrl-0-10#whats-next" class="hash-link" aria-label="Direct link to What's Next?" title="Direct link to What's Next?" translate="no">​</a></h2>
<p>We are working steadily toward the 1.0 release of DataSQRL. Most of the big features are in, and we are primarily focusing on hardening what is there: additional test coverage, handling edge cases, and extending support.</p>
<p>One area of work that's particularly interesting is partition strategy optimization. We are extending support for statistics to make that happen. More on this soon.</p>]]></content:encoded>
            <category>Release</category>
            <category>DataSQRL</category>
        </item>
        <item>
            <title><![CDATA[Avoiding Duplicate Processing in Flink SQL Streaming Jobs]]></title>
            <link>https://docs.datasqrl.com/main/blog/flink-sql-stream-ruleset</link>
            <guid>https://docs.datasqrl.com/main/blog/flink-sql-stream-ruleset</guid>
            <pubDate>Fri, 20 Feb 2026 00:00:00 GMT</pubDate>
            <description><![CDATA[Flink SQL is a powerful abstraction layer that unifies batch and stream processing over semi-structured data. It extends the widely used SQL language with streaming constructs such as tumbling windows, session windows, and, more recently, process table functions. This enables non-experts in streaming technologies to express complex real-time data processing logic succinctly.]]></description>
            <content:encoded><![CDATA[
<p>Flink SQL is a powerful abstraction layer that unifies batch and stream processing over semi-structured data. It extends the widely used SQL language with streaming constructs such as tumbling windows, session windows, and, more recently, process table functions. This enables non-experts in streaming technologies to express complex real-time data processing logic succinctly.</p>
<p>As a result, Flink SQL significantly lowers the barrier to entry for building real-time data systems.</p>
<p>However, developing streaming applications differs fundamentally from traditional SQL query processing. One key difference is that streaming jobs often have multiple sinks populated by a single pipeline, sharing large portions of common data processing logic.</p>
<p>While Flink SQL provides mechanisms to express this concisely—using views and statement sets—in practice, this often results in duplicate processing in the generated job graph.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-core-problem-why-duplication-happens">The Core Problem: Why Duplication Happens<a href="https://docs.datasqrl.com/main/blog/flink-sql-stream-ruleset#the-core-problem-why-duplication-happens" class="hash-link" aria-label="Direct link to The Core Problem: Why Duplication Happens" title="Direct link to The Core Problem: Why Duplication Happens" translate="no">​</a></h2>
<p>In Flink SQL, each sink maps to its own relational tree. These trees are:</p>
<ol>
<li class="">Optimized individually.</li>
<li class="">Combined afterward into a single job graph.</li>
</ol>
<p>By the time they are merged, the query optimization has introduced subtle difference in the shared data processing which render the subgraph merging ineffective. As a result, common processing logic, such as joins, gets duplicated.</p>
<p>Ironically, common SQL optimization techniques like predicate pushdown and projection pruning can make this worse in streaming contexts. While these optimizations are beneficial in traditional query processing (because they avoid computing unused data), they can fragment pipelines in streaming jobs and prevent subgraph sharing.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="example-clickstream-enrichment-with-two-aggregations">Example: Clickstream Enrichment with Two Aggregations<a href="https://docs.datasqrl.com/main/blog/flink-sql-stream-ruleset#example-clickstream-enrichment-with-two-aggregations" class="hash-link" aria-label="Direct link to Example: Clickstream Enrichment with Two Aggregations" title="Direct link to Example: Clickstream Enrichment with Two Aggregations" translate="no">​</a></h2>
<p>Consider a streaming application that:</p>
<ol>
<li class="">Ingests ad click events.</li>
<li class="">Enriches them with ad metadata.</li>
<li class="">Produces two separate aggregations.</li>
</ol>
<p>We perform a temporal join between the clickstream and the ad inventory to enrich each click event with ad metadata.</p>
<p>After enrichment, we compute two aggregations:</p>
<ul>
<li class="">Hourly tumbling window that counts number of clicks per ad category.</li>
<li class="">Daily tumbling window that counts clicks per advertiser.</li>
</ul>
<p>Both results are written to a PostgreSQL database for querying.</p>
<p>Below is a simplified Flink SQL script that implements our clickstream aggregation.</p>
<div class="language-sql codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-sql codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token comment" style="color:rgb(98, 114, 164)">-- Clickstream source</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">CREATE</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">TABLE</span><span class="token plain"> Clickstream </span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    user_id STRING</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    event_time TIMESTAMP_LTZ</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token number">3</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"> </span><span class="token operator">NOT</span><span class="token plain"> </span><span class="token boolean">NULL</span><span class="token plain"> METADATA </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">'timestamp'</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    ad_id STRING</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    WATERMARK </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FOR</span><span class="token plain"> event_time </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> event_time </span><span class="token operator">-</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">INTERVAL</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">'5'</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SECOND</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">WITH</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token string" style="color:rgb(255, 121, 198)">'connector'</span><span class="token plain"> </span><span class="token operator">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">'kafka'</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token comment" style="color:rgb(98, 114, 164)">-- Ad inventory source</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">CREATE</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">TABLE</span><span class="token plain"> AdInventory </span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">     ad_id STRING</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">     category STRING</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">     advertiser STRING</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">     launch_date TIMESTAMP_LTZ</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token number">3</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"> </span><span class="token operator">NOT</span><span class="token plain"> </span><span class="token boolean">NULL</span><span class="token plain"> METADATA </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">'timestamp'</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">     </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">PRIMARY</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">KEY</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain">ad_id</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"> </span><span class="token operator">NOT</span><span class="token plain"> ENFORCED</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">     WATERMARK </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FOR</span><span class="token plain"> launch_date </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> launch_date </span><span class="token operator">-</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">INTERVAL</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">'5'</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SECOND</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">WITH</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token string" style="color:rgb(255, 121, 198)">'connector'</span><span class="token plain"> </span><span class="token operator">=</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">'upsert-kafka'</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token comment" style="color:rgb(98, 114, 164)">-- Enrich clickstream with ads</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">CREATE</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">VIEW</span><span class="token plain"> EnrichedClicks </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SELECT</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    c</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">user_id</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    c</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">event_time</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    c</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">ad_id</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    a</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">category</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    a</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">advertiser</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> Clickstream </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> c</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">LEFT</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">JOIN</span><span class="token plain"> AdInventory </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FOR</span><span class="token plain"> SYSTEM_TIME </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">OF</span><span class="token plain"> c</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">event_time </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> a</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ON</span><span class="token plain"> c</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">ad_id </span><span class="token operator">=</span><span class="token plain"> a</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">ad_id</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BEGIN</span><span class="token plain"> STATEMENT </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SET</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token comment" style="color:rgb(98, 114, 164)">-- Hourly tumbling window: clicks per category</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">INSERT</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">INTO</span><span class="token plain"> hourly_category_clicks</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SELECT</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    category</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    window_start</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token function" style="color:rgb(80, 250, 123)">COUNT</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token operator">*</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> click_count</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">TABLE</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    TUMBLE</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">        </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">TABLE</span><span class="token plain"> EnrichedClicks</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">        DESCRIPTOR</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain">event_time</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">        </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">INTERVAL</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">'1'</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">HOUR</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">GROUP</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BY</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    category</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    window_start</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token comment" style="color:rgb(98, 114, 164)">-- Daily tumbling window: clicks per advertiser</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">INSERT</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">INTO</span><span class="token plain"> daily_advertiser_clicks</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SELECT</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    advertiser</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    window_start</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token function" style="color:rgb(80, 250, 123)">COUNT</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token operator">*</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> click_count</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">TABLE</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    TUMBLE</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">        </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">TABLE</span><span class="token plain"> EnrichedClicks</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">        DESCRIPTOR</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain">event_time</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">        </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">INTERVAL</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">'1'</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">DAY</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">GROUP</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BY</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    advertiser</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    window_start</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">END</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><br></div></code></pre></div></div>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="why-does-this-cause-duplicate-processing">Why Does this Cause Duplicate Processing?<a href="https://docs.datasqrl.com/main/blog/flink-sql-stream-ruleset#why-does-this-cause-duplicate-processing" class="hash-link" aria-label="Direct link to Why Does this Cause Duplicate Processing?" title="Direct link to Why Does this Cause Duplicate Processing?" translate="no">​</a></h3>
<p>When running this Flink SQL script, the generated job graph duplicates the temporal join.</p>
<img src="https://docs.datasqrl.com/img/blog/subgraph_elimination_before.png" alt="Duplicate Temporal Join" width="100%">
<p>Why? Because the two aggregations needs slightly different data: The hourly aggregation needs <code>category</code>. The daily aggregation needs <code>advertiser</code>. Thus, predicate pushdown and projection pruning cause the optimizer to generate slightly different pipelines for each sink. As a result, the temporal join is executed twice.</p>
<p>In traditional SQL systems, this behavior is desirable because it avoids processing unnecessary fields when you submit a query. In streaming systems, however, this approach is suboptimal because it considers each sink in isolation and not the overall data processing of the entire streaming job.</p>
<p>For our simple application, it is much more efficient to compute the temporal join once and enrich the clickstream with all the ad metadata we need downstream. That eliminates an operator and redundant data processing. More importantly, it cuts the number of state requests to RocksDB in half which is the primary bottleneck for this job.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="fixing-job-graph-duplication-in-flink-sql">Fixing Job-Graph Duplication in Flink SQL<a href="https://docs.datasqrl.com/main/blog/flink-sql-stream-ruleset#fixing-job-graph-duplication-in-flink-sql" class="hash-link" aria-label="Direct link to Fixing Job-Graph Duplication in Flink SQL" title="Direct link to Fixing Job-Graph Duplication in Flink SQL" translate="no">​</a></h2>
<p>Apache Flink already provides two important building blocks to eliminate this wasteful duplication:</p>
<ul>
<li class="">Subgraph elimination within the physical <code>RelNode</code> graph to remove duplicate processing.</li>
<li class="">Compile plans, which generate a static artifact representing the generated job graph from Flink SQL.</li>
</ul>
<p>The compile plan is especially useful because it allows validation of the generated job graph at compile time and assigns stable operator IDs. That's important for job evolution by preserving state mappings across job changes or Flink version upgrades.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="selective-rule-control-in-calcite">Selective Rule Control in Calcite<a href="https://docs.datasqrl.com/main/blog/flink-sql-stream-ruleset#selective-rule-control-in-calcite" class="hash-link" aria-label="Direct link to Selective Rule Control in Calcite" title="Direct link to Selective Rule Control in Calcite" translate="no">​</a></h3>
<p>The issue lies in how the Calcite optimizer applies certain optimization rules, particularly those related to projection pruning and filter pushdown. These are the most likely culprits for causing minor differences in the optimized versions of shared logical plans.</p>
<p>In many real-world streaming scenarios with multiple sinks, these rules prevent effective subgraph elimination because they cause subtle difference in the <code>RelNode</code> graph whereas the subgraph elimination requires strict equality.</p>
<p>The solution is to selectively disable specific Calcite rules so that intermediate views do not get optimized for each <code>RelNode</code> tree and remain identical. Identical <code>RelNode</code> trees are then removed during the subgraph elimination phase of the Flink SQL optimizer.</p>
<img src="https://docs.datasqrl.com/img/blog/subgraph_elimination_after.png" alt="Duplicate Temporal Join" width="100%">
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="how-do-i-use-this">How Do I use this?<a href="https://docs.datasqrl.com/main/blog/flink-sql-stream-ruleset#how-do-i-use-this" class="hash-link" aria-label="Direct link to How Do I use this?" title="Direct link to How Do I use this?" translate="no">​</a></h2>
<p>An easy way to remove job graph duplication is to use the <a href="https://docs.datasqrl.com/" target="_blank" rel="noopener noreferrer" class="">DataSQRL compiler</a> and configure the following in your project's <code>package.json</code>:</p>
<div class="language-json codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-json codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token plain">  </span><span class="token property">"compiler"</span><span class="token operator">:</span><span class="token plain"> </span><span class="token punctuation" style="color:rgb(248, 248, 242)">{</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token property">"predicate-pushdown-rules"</span><span class="token operator">:</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">"LIMITED_RULES"</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">  </span><span class="token punctuation" style="color:rgb(248, 248, 242)">}</span><br></div></code></pre></div></div>
<p>This ensures only a limited set of Calcite rules are applied during the compiled plan optimization.
<a href="https://docs.datasqrl.com/" target="_blank" rel="noopener noreferrer" class="">DataSQRL</a> is a data automation framework that compiles Flink SQL to data pipelines and it can compile your Flink SQL to a compiled plan with a single command you execute locally. For more information, check out the <a href="https://docs.datasqrl.com/docs/intro/getting-started" target="_blank" rel="noopener noreferrer" class="">getting started</a> tutorial.</p>
<p>As an alternative, you can selectively disable Calcite rules in your own instrumentation framework for Flink. Check out this <a href="https://github.com/DataSQRL/sqrl/blob/release-0.9/sqrl-planner/src/main/java/com/datasqrl/planner/FlinkPlannerConfigBuilder.java" target="_blank" rel="noopener noreferrer" class="">code snippet</a> to see what rules we are disabling. We highly recommend that you produce a compiled plan for your Flink SQL jobs for introspection and predictability. You can use the open-source <a href="https://github.com/DataSQRL/flink-sql-runner" target="_blank" rel="noopener noreferrer" class="">Flink SQL Runner</a> to execute compiled plans.</p>
<p>In effect, it allows you to define a streaming-specific optimization ruleset, rather than relying solely on optimizations designed primarily for traditional query workloads.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="whats-next">What's Next?<a href="https://docs.datasqrl.com/main/blog/flink-sql-stream-ruleset#whats-next" class="hash-link" aria-label="Direct link to What's Next?" title="Direct link to What's Next?" translate="no">​</a></h2>
<p>While the ruleset tweaks described above address most cases of subgraph duplication in streaming Flink SQL, we still have some more work to do for certain edge cases.</p>
<p>In particular, SQL constructs that introduce correlation variables (e.g. <code>UNNEST</code>) into the Calcite logical plan do not get deduplicated yet because correlation variable have a static counter that makes each variable unique.</p>
<p>We implemented a normalization algorithm for correlation variables and are looking for ways to contribute it directly to the Apache Flink project.</p>]]></content:encoded>
            <category>Flink</category>
        </item>
        <item>
            <title><![CDATA[DataSQRL 0.7 Release: The Data Delivery Interface]]></title>
            <link>https://docs.datasqrl.com/main/blog/datasqrl-0.7-release</link>
            <guid>https://docs.datasqrl.com/main/blog/datasqrl-0.7-release</guid>
            <pubDate>Sun, 27 Jul 2025 00:00:00 GMT</pubDate>
            <description><![CDATA[|" width="40%"/>]]></description>
            <content:encoded><![CDATA[
<img src="https://docs.datasqrl.com/img/blog/release_0.7.0.png" alt="DataSQRL 0.7.0 Release >|" width="40%">
<p>DataSQRL 0.7 marks a major milestone in our journey to automate data pipelines, thanks to significant improvements to the serving layer:</p>
<ul>
<li class="">Support for the Model Context Protocol (MCP) for tooling and resource access</li>
<li class="">REST API support</li>
<li class="">JWT-based authentication and authorization</li>
</ul>
<p>These features enable developers to build a wide range of production-ready data interfaces.
This release also includes performance and configuration improvements to the serving layer of DataSQRL-generated pipelines.</p>
<p>You can find the full release notes and source code on our <a href="https://github.com/DataSQRL/sqrl/releases/tag/0.7.0" target="_blank" rel="noopener noreferrer" class="">GitHub release page</a>.
To update your local installation of DataSQRL, simply pull the latest Docker image:</p>
<div class="language-bash codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-bash codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token plain">docker pull datasqrl/cmd:0.7.0</span><br></div></code></pre></div></div>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="the-last-mile-data-delivery">The Last Mile: Data Delivery<a href="https://docs.datasqrl.com/main/blog/datasqrl-0.7-release#the-last-mile-data-delivery" class="hash-link" aria-label="Direct link to The Last Mile: Data Delivery" title="Direct link to The Last Mile: Data Delivery" translate="no">​</a></h2>
<p>Data delivery is the final and most visible stage of any data pipeline. It's how users, applications, and AI agents actually access and consume data. Most enterprise data interactions happen through APIs, making the delivery interface a critical component. At DataSQRL, we've invested heavily in automating the upstream parts of the pipeline: from Flink-powered data processing to Postgres-backed storage. With version 0.7, we turn our focus to the serving layer: introducing support for the Model Context Protocol (MCP) and REST APIs, as well as JWT-based authentication and authorization. These additions ensure seamless integration with most authentication providers and enable secure, token-based data access, with fine-grained authorization logic enforced directly in the SQRL script. This completes our vision of end-to-end pipeline automation, where consumption patterns inform data storage and processing—closing the loop between data production and usage.</p>
<p>Check out the <a class="" href="https://docs.datasqrl.com/main/docs/interface">interface documentation</a> for more information.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="major-contributions">Major Contributions<a href="https://docs.datasqrl.com/main/blog/datasqrl-0.7-release#major-contributions" class="hash-link" aria-label="Direct link to Major Contributions" title="Direct link to Major Contributions" translate="no">​</a></h2>
<p>In addition to the flagship features—MCP, REST, and JWT support—which we’ll discuss in more detail in future blog posts, the 0.7.0 release contains a number of additional features and improvements.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="configuration--cli-improvements">Configuration &amp; CLI Improvements<a href="https://docs.datasqrl.com/main/blog/datasqrl-0.7-release#configuration--cli-improvements" class="hash-link" aria-label="Direct link to Configuration &amp; CLI Improvements" title="Direct link to Configuration &amp; CLI Improvements" translate="no">​</a></h3>
<ul>
<li class="">Refactored CLI and Flink configuration logic.</li>
<li class="">Improved error messages during package config validation.</li>
<li class="">Enforced predictable ordering and sorting of config keys.</li>
<li class="">Migrated deprecated config key naming (<code>-dir</code> → <code>-folder</code>), now with compile-time warnings.</li>
<li class="">Standardized configuration schema and structured logging to <code>/build/logs</code>.</li>
</ul>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="testing-infrastructure-enhancements">Testing Infrastructure Enhancements<a href="https://docs.datasqrl.com/main/blog/datasqrl-0.7-release#testing-infrastructure-enhancements" class="hash-link" aria-label="Direct link to Testing Infrastructure Enhancements" title="Direct link to Testing Infrastructure Enhancements" translate="no">​</a></h3>
<ul>
<li class="">Added new <code>sqrl-container-testing</code> module.</li>
<li class="">Converted tests to use AssertJ.</li>
<li class="">Increased test coverage for <code>SqrlConfig</code> and dependency mapping.</li>
<li class="">Fixed test runner error reporting and exit code handling.</li>
<li class="">Reworked dependent service startup to trigger post-compilation.</li>
</ul>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="flink-sql-runner-integration">Flink-SQL Runner Integration<a href="https://docs.datasqrl.com/main/blog/datasqrl-0.7-release#flink-sql-runner-integration" class="hash-link" aria-label="Direct link to Flink-SQL Runner Integration" title="Direct link to Flink-SQL Runner Integration" translate="no">​</a></h3>
<ul>
<li class="">Integrated <code>flink-sql-runner</code> into <code>DatasqrlRun</code>.</li>
<li class="">Temporarily merged <code>sqrl-test</code> module into <code>sqrl-run</code>.</li>
<li class="">Simplified Docker image setup (public <code>ghcr.io</code> images).</li>
<li class="">Updated submodule paths and version to <code>0.7.0</code>.</li>
</ul>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="authentication--api-enhancements">Authentication &amp; API Enhancements<a href="https://docs.datasqrl.com/main/blog/datasqrl-0.7-release#authentication--api-enhancements" class="hash-link" aria-label="Direct link to Authentication &amp; API Enhancements" title="Direct link to Authentication &amp; API Enhancements" translate="no">​</a></h3>
<ul>
<li class="">Added initial JWT-based authentication support.</li>
<li class="">Published documentation for JWT and Swagger-based OpenAPI specs for REST endpoints.</li>
<li class="">Added batch GraphQL mutation support with transactional semantics.</li>
<li class="">Replaced <code>GraphQLBigInteger</code> with native <code>Long</code> handling.</li>
</ul>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="kafka--runtime-improvements">Kafka &amp; Runtime Improvements<a href="https://docs.datasqrl.com/main/blog/datasqrl-0.7-release#kafka--runtime-improvements" class="hash-link" aria-label="Direct link to Kafka &amp; Runtime Improvements" title="Direct link to Kafka &amp; Runtime Improvements" translate="no">​</a></h3>
<ul>
<li class="">Kafka topic names now support templating.</li>
<li class="">Added an async OpenAI test use case and resolved snapshot issues.</li>
<li class="">Fixed intermittent WebSocket failures in <code>SubscriptionClient</code>.</li>
</ul>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="project-structure-and-ci-pipeline">Project Structure and CI Pipeline<a href="https://docs.datasqrl.com/main/blog/datasqrl-0.7-release#project-structure-and-ci-pipeline" class="hash-link" aria-label="Direct link to Project Structure and CI Pipeline" title="Direct link to Project Structure and CI Pipeline" translate="no">​</a></h3>
<ul>
<li class="">Simplified project structure and removed outdated dependency declarations.</li>
<li class="">Refactored CI pipeline and added automated GitHub package cleanup workflow.</li>
</ul>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="updated-dependencies">Updated Dependencies<a href="https://docs.datasqrl.com/main/blog/datasqrl-0.7-release#updated-dependencies" class="hash-link" aria-label="Direct link to Updated Dependencies" title="Direct link to Updated Dependencies" translate="no">​</a></h3>
<p>Upgraded versions for the following dependencies and patched critical vulnerabilities:</p>
<ul>
<li class="">Apache Flink and Flink connectors (e.g., Postgres CDC)</li>
<li class="">Vert.x and related plugins</li>
<li class="">Apache Iceberg, DuckDB, PostgreSQL JDBC</li>
<li class="">AWS SDK BOM</li>
<li class="">JSON Schema Validator, Netty, OpenCSV</li>
<li class="">Micrometer, Log4j, Reactor, Immutables, Testcontainers</li>
<li class="">Maven plugins: Enforcer, GPG, Build-helper</li>
</ul>]]></content:encoded>
            <category>Release</category>
        </item>
        <item>
            <title><![CDATA[Flink SQL Runner: Run Flink SQL Without JARs or Glue Code]]></title>
            <link>https://docs.datasqrl.com/main/blog/flinkrunner-announcement</link>
            <guid>https://docs.datasqrl.com/main/blog/flinkrunner-announcement</guid>
            <pubDate>Mon, 09 Jun 2025 00:00:00 GMT</pubDate>
            <description><![CDATA[Apache Flink has long been a powerhouse for streaming and batch data processing. And with the rise of Flink SQL, developers can now build sophisticated pipelines using a declarative language they already know. But getting Flink SQL applications into production still comes with friction: packaging JARs, managing connectors, injecting secrets, and wiring up deployment infrastructure.]]></description>
            <content:encoded><![CDATA[
<p>Apache Flink has long been a powerhouse for streaming and batch data processing. And with the rise of Flink SQL, developers can now build sophisticated pipelines using a declarative language they already know. But getting Flink SQL applications into production still comes with friction: packaging JARs, managing connectors, injecting secrets, and wiring up deployment infrastructure.</p>
<img src="https://docs.datasqrl.com/img/blog/flinksqlrunner_logo.png" alt="FlinkSQL Runner >" width="40%">
<p><a href="https://github.com/DataSQRL/flink-sql-runner/" target="_blank" rel="noopener noreferrer" class=""><strong>Flink SQL Runner</strong></a> is here to change that. It's an open-source toolkit that simplifies development, deployment, and operation of Flink SQL applications—locally or in Kubernetes—without manual JAR assembly or scripting custom infrastructure pipelines.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="from-sql-to-production-minus-the-plumbing">From SQL to Production, Minus the Plumbing<a href="https://docs.datasqrl.com/main/blog/flinkrunner-announcement#from-sql-to-production-minus-the-plumbing" class="hash-link" aria-label="Direct link to From SQL to Production, Minus the Plumbing" title="Direct link to From SQL to Production, Minus the Plumbing" translate="no">​</a></h2>
<p>Imagine you're writing a Flink SQL job that reads from Kafka, enriches the data, and sinks to Iceberg. In theory, it's just SQL. But in practice, production deployment requires:</p>
<ul>
<li class="">Assembling dependencies into a JAR</li>
<li class="">Writing YAML to configure connectors</li>
<li class="">Injecting secrets for different environments</li>
</ul>
<p>Flink SQL Runner eliminates those headaches. You get:</p>
<ul>
<li class=""><strong>Declarative execution</strong> with SQL scripts or compiled plans</li>
<li class=""><strong>Simple deployments</strong> on Kubernetes via Flink Operator</li>
<li class=""><strong>Environment isolation</strong> with variable substitution and UDF packaging</li>
</ul>
<p>All without leaving the SQL layer.</p>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="key-features">Key Features<a href="https://docs.datasqrl.com/main/blog/flinkrunner-announcement#key-features" class="hash-link" aria-label="Direct link to Key Features" title="Direct link to Key Features" translate="no">​</a></h3>
<ul>
<li class=""><strong>SQL and Plan Execution</strong>: Run raw SQL scripts or pre-compiled execution plans.</li>
<li class=""><strong>Kubernetes-Native</strong>: Built for the Flink Kubernetes Operator—deploy SQL jobs without writing infrastructure code.</li>
<li class=""><strong>Composable Toolkit</strong>: Use the pieces you need—Docker image, libraries, extensions—to suit your environment.</li>
<li class=""><strong>Environment Variable Substitution</strong>: Inject secrets and environment-specific config into SQL or plan files using <code>${ENV_VAR}</code> syntax.</li>
<li class=""><strong>UDF Infrastructure</strong>: Load custom JARs and register system functions easily.</li>
<li class=""><strong>Function Libraries</strong>: Drop-in UDFs for advanced math and OpenAI integration.</li>
</ul>
<h3 class="anchor anchorTargetStickyNavbar_Vzrq" id="flexible-and-extensible">Flexible and Extensible<a href="https://docs.datasqrl.com/main/blog/flinkrunner-announcement#flexible-and-extensible" class="hash-link" aria-label="Direct link to Flexible and Extensible" title="Direct link to Flexible and Extensible" translate="no">​</a></h3>
<p>Flink SQL Runner is not a monolith. You can:</p>
<ul>
<li class="">Run it standalone with Docker.</li>
<li class="">Deploy it with the Flink Kubernetes Operator.</li>
<li class="">Extend it via Maven or Gradle in your own Flink stack:</li>
</ul>
<div class="language-xml codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-xml codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token tag punctuation" style="color:rgb(248, 248, 242)">&lt;</span><span class="token tag" style="color:rgb(255, 121, 198)">dependency</span><span class="token tag punctuation" style="color:rgb(248, 248, 242)">&gt;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">  </span><span class="token tag punctuation" style="color:rgb(248, 248, 242)">&lt;</span><span class="token tag" style="color:rgb(255, 121, 198)">groupId</span><span class="token tag punctuation" style="color:rgb(248, 248, 242)">&gt;</span><span class="token plain">com.datasqrl.flinkrunner</span><span class="token tag punctuation" style="color:rgb(248, 248, 242)">&lt;/</span><span class="token tag" style="color:rgb(255, 121, 198)">groupId</span><span class="token tag punctuation" style="color:rgb(248, 248, 242)">&gt;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">  </span><span class="token tag punctuation" style="color:rgb(248, 248, 242)">&lt;</span><span class="token tag" style="color:rgb(255, 121, 198)">artifactId</span><span class="token tag punctuation" style="color:rgb(248, 248, 242)">&gt;</span><span class="token plain">flink-sql-runner</span><span class="token tag punctuation" style="color:rgb(248, 248, 242)">&lt;/</span><span class="token tag" style="color:rgb(255, 121, 198)">artifactId</span><span class="token tag punctuation" style="color:rgb(248, 248, 242)">&gt;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">  </span><span class="token tag punctuation" style="color:rgb(248, 248, 242)">&lt;</span><span class="token tag" style="color:rgb(255, 121, 198)">version</span><span class="token tag punctuation" style="color:rgb(248, 248, 242)">&gt;</span><span class="token plain">0.6.0</span><span class="token tag punctuation" style="color:rgb(248, 248, 242)">&lt;/</span><span class="token tag" style="color:rgb(255, 121, 198)">version</span><span class="token tag punctuation" style="color:rgb(248, 248, 242)">&gt;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token tag punctuation" style="color:rgb(248, 248, 242)">&lt;/</span><span class="token tag" style="color:rgb(255, 121, 198)">dependency</span><span class="token tag punctuation" style="color:rgb(248, 248, 242)">&gt;</span><br></div></code></pre></div></div>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="get-started">Get Started<a href="https://docs.datasqrl.com/main/blog/flinkrunner-announcement#get-started" class="hash-link" aria-label="Direct link to Get Started" title="Direct link to Get Started" translate="no">​</a></h2>
<p>The Flink SQL Runner project is open source and <a href="https://github.com/datasqrl/flink-sql-runner" target="_blank" rel="noopener noreferrer" class="">available on Github</a>.
Check out the README for more information on how to use and deploy the Flink SQL Runner.</p>
<p>Try it out, report issues, or contribute your own UDFs.</p>]]></content:encoded>
            <category>Flink</category>
            <category>DataSQRL</category>
        </item>
        <item>
            <title><![CDATA[Defining Data Interfaces with FlinkSQL]]></title>
            <link>https://docs.datasqrl.com/main/blog/flinksql-extensions</link>
            <guid>https://docs.datasqrl.com/main/blog/flinksql-extensions</guid>
            <pubDate>Fri, 09 May 2025 00:00:00 GMT</pubDate>
            <description><![CDATA[FlinkSQL is an amazing innovation in data processing: it packages the power of realtime stream processing within the simplicity of SQL.]]></description>
            <content:encoded><![CDATA[
<p><a href="https://nightlies.apache.org/flink/flink-docs-release-1.19/docs/dev/table/sql/overview/" target="_blank" rel="noopener noreferrer" class="">FlinkSQL</a> is an amazing innovation in data processing: it packages the power of realtime stream processing within the simplicity of SQL.
That means you can start with the SQL you know and introduce stream processing constructs as you need them.</p>
<img src="https://docs.datasqrl.com/img/blog/flinksql_extension_api.png" alt="FlinkSQL API Extension >" width="40%">
<p>FlinkSQL adds the ability to process data incrementally to the classic set-based semantics of SQL. In addition, FlinkSQL supports source and sink connectors making it easy to ingest data from and move data to other systems. That's a powerful combination which covers a lot of data processing use cases.</p>
<p>In fact, it only takes a few extensions to FlinkSQL to build entire data applications. Let's see how that works.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="building-data-apis-with-flinksql">Building Data APIs with FlinkSQL<a href="https://docs.datasqrl.com/main/blog/flinksql-extensions#building-data-apis-with-flinksql" class="hash-link" aria-label="Direct link to Building Data APIs with FlinkSQL" title="Direct link to Building Data APIs with FlinkSQL" translate="no">​</a></h2>
<div class="language-sql codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-sql codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token comment" style="color:rgb(98, 114, 164)">/*+ engine(kafka) */</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">CREATE</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">TABLE</span><span class="token plain"> UserTokens </span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">userid </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BIGINT</span><span class="token plain"> </span><span class="token operator">NOT</span><span class="token plain"> </span><span class="token boolean">NULL</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">tokens </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BIGINT</span><span class="token plain"> </span><span class="token operator">NOT</span><span class="token plain"> </span><span class="token boolean">NULL</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">request_time TIMESTAMP_LTZ</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token number">3</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"> </span><span class="token operator">NOT</span><span class="token plain"> </span><span class="token boolean">NULL</span><span class="token plain"> METADATA </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> </span><span class="token string" style="color:rgb(255, 121, 198)">'timestamp'</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token comment" style="color:rgb(98, 114, 164)">/*+query_by_all(userid) */</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">TotalUserTokens :</span><span class="token operator">=</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SELECT</span><span class="token plain"> userid</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> </span><span class="token function" style="color:rgb(80, 250, 123)">sum</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain">tokens</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">as</span><span class="token plain"> total_tokens</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token function" style="color:rgb(80, 250, 123)">count</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain">tokens</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">as</span><span class="token plain"> total_requests</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> UserTokens </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">GROUP</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BY</span><span class="token plain"> userid</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">UserTokensByTime</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain">userid </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BIGINT</span><span class="token plain"> </span><span class="token operator">NOT</span><span class="token plain"> </span><span class="token boolean">NULL</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> fromTime </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">TIMESTAMP</span><span class="token plain"> </span><span class="token operator">NOT</span><span class="token plain"> </span><span class="token boolean">NULL</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> toTime </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">TIMESTAMP</span><span class="token plain"> </span><span class="token operator">NOT</span><span class="token plain"> </span><span class="token boolean">NULL</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain">:</span><span class="token operator">=</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">                </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SELECT</span><span class="token plain"> </span><span class="token operator">*</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> UserTokens </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">WHERE</span><span class="token plain"> userid </span><span class="token operator">=</span><span class="token plain"> :userid</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">                </span><span class="token operator">AND</span><span class="token plain"> request_time </span><span class="token operator">&gt;=</span><span class="token plain"> :fromTime </span><span class="token operator">AND</span><span class="token plain"> request_time </span><span class="token operator">&lt;</span><span class="token plain"> :toTime </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ORDER</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BY</span><span class="token plain"> request_time </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">DESC</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">UsageAlert :</span><span class="token operator">=</span><span class="token plain"> SUBSCRIBE </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SELECT</span><span class="token plain"> </span><span class="token operator">*</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> UserTokens </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">WHERE</span><span class="token plain"> tokens </span><span class="token operator">&gt;</span><span class="token plain"> </span><span class="token number">100000</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><br></div></code></pre></div></div>
<p>This script defines a sequence of tables. We introduce <code>:=</code> as syntactic sugar for the verbose <code>CREATE TEMPORARY VIEW</code> syntax.</p>
<p>The <code>UserTokens</code> table does not have a configured connector, which mean we treat it as an API mutation endpoint connected to Flink via a Kafka topic that captures the events. This makes it easy to build APIs that capture user activity, transactions, or other types of events.</p>
<p>Next, we sum up the data collected through the API for each user. This is a standard FlinkSQL aggregation query and we expose the result in our API through the <code>query_by_all</code> hint which defines the arguments for the query endpoint of that table.</p>
<p>We can also explicitly define query endpoints with arguments through SQL table functions. FlinkSQL supports table functions natively. All we had to do is provide the syntax for defining the function signature.</p>
<p>And last, the <code>SUBSCRIBE</code> keyword in front of the query defines a subscription endpoint for requests exceeding a certain token count which get pushed to clients in real-time.</p>
<p>Voila, we just build ourselves a complete GraphQL API with mutation, query, and subscription endpoints.
Run the above script with DataSQRL to see the result. Save it as <code>usertokens.sqrl</code> next to a <code>usertokens-package.json</code> file that sets <code>script.main</code> to <code>usertokens.sqrl</code> (see the <a class="" href="https://docs.datasqrl.com/main/docs/configuration">configuration documentation</a>), then run:</p>
<div class="language-bash codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-bash codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token plain">docker run -it --rm -p 8888:8888 -v $PWD:/workspace datasqrl/cmd run usertokens-package.json</span><br></div></code></pre></div></div>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="relationships-for-complex-data-structures">Relationships for Complex Data Structures<a href="https://docs.datasqrl.com/main/blog/flinksql-extensions#relationships-for-complex-data-structures" class="hash-link" aria-label="Direct link to Relationships for Complex Data Structures" title="Direct link to Relationships for Complex Data Structures" translate="no">​</a></h2>
<p>And for extra credit, we can define relationships in FlinkSQL to represent the structure of our data explicitly and expose it in the API:</p>
<div class="language-sql codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-sql codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">User</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">totalTokens :</span><span class="token operator">=</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SELECT</span><span class="token plain"> </span><span class="token operator">*</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> TotalUserTokens t </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">WHERE</span><span class="token plain"> this</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">userid </span><span class="token operator">=</span><span class="token plain"> t</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">userid </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">LIMIT</span><span class="token plain"> </span><span class="token number">1</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><br></div></code></pre></div></div>
<p>The <code>User</code> table in this example is read from an upsert Kafka topic using a standard FlinkSQL <code>CREATE TABLE</code> statement.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="code-modularity-and-connector-management">Code Modularity and Connector Management<a href="https://docs.datasqrl.com/main/blog/flinksql-extensions#code-modularity-and-connector-management" class="hash-link" aria-label="Direct link to Code Modularity and Connector Management" title="Direct link to Code Modularity and Connector Management" translate="no">​</a></h2>
<p>Many FlinkSQL projects break the codebase into multiple files for better code readability, modularity, or to swap out sources and sinks. That requires extra infrastructure to manage FlinkSQL files and stitch them together.</p>
<p>How about we do that directly in FlinkSQL?</p>
<div class="language-sql codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-sql codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">IMPORT</span><span class="token plain"> source</span><span class="token operator">-</span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">data</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">User</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><br></div></code></pre></div></div>
<p>Here, we import the <code>User</code> table from a separate file within the <code>source-data</code> directory, allowing us to separate the data processing logic from the source configurations. It also enables us to use dependency management to swap out sources for local testing vs production.</p>
<p>And we can do the same for sinks:</p>
<div class="language-sql codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-sql codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token plain">EXPORT UsageAlert </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">TO</span><span class="token plain"> mysinks</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">UsageAlert</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><br></div></code></pre></div></div>
<p>In addition to breaking out the sink configuration from the main script, the <code>EXPORT</code> statement functions as an <code>INSERT INTO</code> statement and creates a <code>STATEMENT SET</code> implicitly. That makes the code easier to read.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="learn-more">Learn More<a href="https://docs.datasqrl.com/main/blog/flinksql-extensions#learn-more" class="hash-link" aria-label="Direct link to Learn More" title="Direct link to Learn More" translate="no">​</a></h2>
<p><a href="https://nightlies.apache.org/flink/flink-docs-release-1.19/docs/dev/table/sql/overview/" target="_blank" rel="noopener noreferrer" class="">FlinkSQL</a> is phenomenal extension of the SQL ecosystem to stream processing. With DataSQRL, we are trying to make it easier to build end-to-end data pipelines and complete data applications with FlinkSQL.</p>
<p>Check out the <a class="" href="https://docs.datasqrl.com/main/docs/intro/getting-started">complete example</a> which also covers testing, customization, and deployment. Or read the <a class="" href="https://docs.datasqrl.com/main/docs/sqrl-language">documentation</a> to learn more.</p>]]></content:encoded>
            <category>Join</category>
            <category>Flink</category>
            <category>DataSQRL</category>
        </item>
        <item>
            <title><![CDATA[DataSQRL 0.6 Release: The Streaming Data Framework]]></title>
            <link>https://docs.datasqrl.com/main/blog/datasqrl-0.6-release</link>
            <guid>https://docs.datasqrl.com/main/blog/datasqrl-0.6-release</guid>
            <pubDate>Wed, 07 May 2025 00:00:00 GMT</pubDate>
            <description><![CDATA[The DataSQRL community is proud to announce the release of DataSQRL 0.6. This release marks a major milestone in the evolution of our open-source project, bringing enhanced alignment with Flink SQL and powerful new capabilities to the real-time serving layer.]]></description>
            <content:encoded><![CDATA[
<p>The DataSQRL community is proud to announce the release of DataSQRL 0.6. This release marks a major milestone in the evolution of our open-source project, bringing enhanced alignment with Flink SQL and powerful new capabilities to the real-time serving layer.</p>
<img src="https://docs.datasqrl.com/img/blog/release_0.6.0.png" alt="DataSQRL 0.6.0 Release >" width="40%">
<p>You can find the full release notes and source code on our <a href="https://github.com/DataSQRL/sqrl/releases/tag/0.6.0" target="_blank" rel="noopener noreferrer" class="">GitHub release page</a>.
To get started with the latest compiler, simply pull the latest Docker image:</p>
<div class="language-bash codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-bash codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token plain">docker pull datasqrl/cmd:0.6.0</span><br></div></code></pre></div></div>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="a-new-chapter-flink-sql-integration">A New Chapter: Flink SQL Integration<a href="https://docs.datasqrl.com/main/blog/datasqrl-0.6-release#a-new-chapter-flink-sql-integration" class="hash-link" aria-label="Direct link to A New Chapter: Flink SQL Integration" title="Direct link to A New Chapter: Flink SQL Integration" translate="no">​</a></h2>
<p>With DataSQRL 0.6, we are embracing the Flink ecosystem more deeply than ever before. This release introduces a complete re-architecture of the DataSQRL compiler to build directly on top of Flink SQL's parser and planner. By aligning our internal model with Flink SQL semantics, we unlock a host of new capabilities and bring DataSQRL users closer to the vibrant Flink ecosystem.</p>
<p>This architectural shift allows DataSQRL to:</p>
<ul>
<li class=""><strong>Use Flink SQL syntax as the foundation</strong>, enabling more intuitive query definitions and easier onboarding for users familiar with Flink.</li>
<li class=""><strong>Extend Flink SQL with domain-specific features</strong>, such as declarative relationship definitions and functions to define the data interface.</li>
<li class=""><strong>Transpile FlinkSQL to database dialects</strong> for query execution.</li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="serving-layer-power-functions--relationships">Serving-Layer Power: Functions &amp; Relationships<a href="https://docs.datasqrl.com/main/blog/datasqrl-0.6-release#serving-layer-power-functions--relationships" class="hash-link" aria-label="Direct link to Serving-Layer Power: Functions &amp; Relationships" title="Direct link to Serving-Layer Power: Functions &amp; Relationships" translate="no">​</a></h2>
<p>DataSQRL 0.6 introduces first-class support for defining <strong>functions</strong> and <strong>relationships</strong> in your SQRL scripts. These constructs make it easier to model complex application logic in a modular, declarative fashion.</p>
<p>These features are purpose-built for powering LLM-ready APIs, event-driven architectures, and real-time user-facing applications.</p>
<p>Check out the <a class="" href="https://docs.datasqrl.com/main/docs/sqrl-language">language documentation</a> for details.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="developer-tooling">Developer Tooling<a href="https://docs.datasqrl.com/main/blog/datasqrl-0.6-release#developer-tooling" class="hash-link" aria-label="Direct link to Developer Tooling" title="Direct link to Developer Tooling" translate="no">​</a></h2>
<p>DataSQRL 0.6 provides a docker image for compiling, running, and testing SQRL projects. You can now quickly iterate and check the results. Or run automated tests in CI/CD.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="deployment-artifacts">Deployment Artifacts<a href="https://docs.datasqrl.com/main/blog/datasqrl-0.6-release#deployment-artifacts" class="hash-link" aria-label="Direct link to Deployment Artifacts" title="Direct link to Deployment Artifacts" translate="no">​</a></h2>
<p>DataSQRL 0.6 removes deployment profiles and instead generates all deployment artifacts in the <code>build/deploy/plan</code> folder. This makes it easier to integrate with Kubernetes deployment processes (e.g. via Helm) or cloud managed service deployments (e.g. via Terraform).</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="breaking-changes--migration-path">Breaking Changes &amp; Migration Path<a href="https://docs.datasqrl.com/main/blog/datasqrl-0.6-release#breaking-changes--migration-path" class="hash-link" aria-label="Direct link to Breaking Changes &amp; Migration Path" title="Direct link to Breaking Changes &amp; Migration Path" translate="no">​</a></h2>
<p>As this is a major release, <strong>DataSQRL 0.6 is not backwards compatible</strong> with version 0.5. The syntax and internal representation have been updated to align with Flink SQL and to support the new compiler architecture.</p>
<p>To help you transition, we’ve provided updated examples and migration guidance in the <a href="https://github.com/DataSQRL/datasqrl-examples" target="_blank" rel="noopener noreferrer" class="">DataSQRL examples repository</a>. We recommend starting with one of the updated use cases to get a feel for the new workflow.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="thanks-to-the-community">Thanks to the Community<a href="https://docs.datasqrl.com/main/blog/datasqrl-0.6-release#thanks-to-the-community" class="hash-link" aria-label="Direct link to Thanks to the Community" title="Direct link to Thanks to the Community" translate="no">​</a></h2>
<p>This release wouldn’t have been possible without the contributions, bug reports, and thoughtful feedback from our growing community. Whether you opened a pull request, filed an issue, or joined a discussion, thank you. Your support drives this project forward.</p>
<p>We’re excited to see what you build with DataSQRL 0.6. If you haven’t joined the <a class="" href="https://docs.datasqrl.com/main/community">community</a> yet, now’s a great time to get involved: star us on <a href="https://github.com/DataSQRL/sqrl" target="_blank" rel="noopener noreferrer" class="">GitHub</a>, try out the latest release, and share your thoughts.</p>
<p>Stay tuned for more updates, and happy building.</p>]]></content:encoded>
            <category>Release</category>
        </item>
        <item>
            <title><![CDATA[Why Temporal Join is Stream Processing’s Superpower]]></title>
            <link>https://docs.datasqrl.com/main/blog/temporal-join</link>
            <guid>https://docs.datasqrl.com/main/blog/temporal-join</guid>
            <pubDate>Mon, 10 Jul 2023 00:00:00 GMT</pubDate>
            <description><![CDATA[Stream processing technologies like Apache Flink introduce a new type of data transformation that’s very powerful: the temporal join. Temporal joins add context to data streams while being efficient and fast to execute.]]></description>
            <content:encoded><![CDATA[
<p>Stream processing technologies like Apache Flink introduce a new type of data transformation that’s very powerful: the temporal join. Temporal joins add context to data streams while being efficient and fast to execute.</p>
<img src="https://docs.datasqrl.com/main/img/blog/temporal_join.svg" alt="Temporal Join >" width="30%">
<p>This article introduces the temporal join, compares it to the traditional inner join, explains when to use it, and why it is a secret superpower.</p>
<p>Table of Contents:</p>
<ul>
<li class=""><a class="" href="https://docs.datasqrl.com/main/blog/temporal-join#review">The Join: A Quick Review</a></li>
<li class=""><a class="" href="https://docs.datasqrl.com/main/blog/temporal-join#tempjoin">The Temporal Join: Linking Stream and State</a></li>
<li class=""><a class="" href="https://docs.datasqrl.com/main/blog/temporal-join#tempinner">Temporal Join vs Inner Join</a></li>
<li class=""><a class="" href="https://docs.datasqrl.com/main/blog/temporal-join#efficient">Why Temporal Joins are Fast and Efficient</a></li>
<li class=""><a class="" href="https://docs.datasqrl.com/main/blog/temporal-join#summary">Summary</a></li>
</ul>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="review">The Join: A Quick Review<a href="https://docs.datasqrl.com/main/blog/temporal-join#review" class="hash-link" aria-label="Direct link to The Join: A Quick Review" title="Direct link to The Join: A Quick Review" translate="no">​</a></h2>
<p>Let's take a quick detour down memory lane and revisit the good ol' join operation. That trusty sidekick in your SQL utility belt helps you link data from two or more tables based on a related column between them.</p>
<p>Suppose we are operating a factory with a number of machines that roast and package coffee. We place sensors on each machine to monitor the temperature and detect overheating.</p>
<p>We keep track of the sensors and machines in two database tables.</p>
<p>The <code>Sensor</code> table contains the serial number and machine id that the sensor is placed on.</p>
<table><thead><tr><th>id</th><th>serialNo</th><th>machineid</th></tr></thead><tbody><tr><td>1</td><td>X57-774</td><td>501</td></tr><tr><td>2</td><td>X33-453</td><td>203</td></tr><tr><td>3</td><td>X54-554</td><td>501</td></tr></tbody></table>
<p>The <code>Machine</code> table contains the name of each machine.</p>
<table><thead><tr><th>id</th><th>name</th></tr></thead><tbody><tr><td>203</td><td>Iron Roaster</td></tr><tr><td>501</td><td>Gritty Grinder</td></tr></tbody></table>
<p>To identify all the sensors on the machine “Iron Roaster” we use the following SQL query which joins the <code>Sensor</code> and <code>Machine</code> tables:</p>
<div class="language-sql codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-sql codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SELECT</span><span class="token plain"> s</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">id</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> s</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">serialNo </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> Sensor s </span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">JOIN</span><span class="token plain"> Machine m </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ON</span><span class="token plain"> s</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">machineid </span><span class="token operator">=</span><span class="token plain"> m</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">id </span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">WHERE</span><span class="token plain"> m</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">name </span><span class="token operator">=</span><span class="token plain"> “Iron Roaster”</span><br></div></code></pre></div></div>
<p>Why are joins important? Without it, your data tables are like islands, isolated and lonely. Joins bring them together, creating meaningful relationships between data, and enriching data records with context to see the bigger picture.</p>
<p>By default, databases execute joins as <strong>inner</strong> joins which means only matching records are included in the join.</p>
<p>So, now that we've refreshed our memory about the classic join, let's dive into the exciting world of temporal joins in stream processing systems like Apache Flink.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="tempjoin">The Temporal Join: Linking Stream and State<a href="https://docs.datasqrl.com/main/blog/temporal-join#tempjoin" class="hash-link" aria-label="Direct link to The Temporal Join: Linking Stream and State" title="Direct link to The Temporal Join: Linking Stream and State" translate="no">​</a></h2>
<img src="https://docs.datasqrl.com/main/img/blog/delorean.jpeg" alt="Temporal Join DeLorean >" width="40%">
<p>Picture this: you're a time traveler. You have the power to access any point in time, past or future, at your will. Now, imagine that your data could do the same. Enter the Temporal Join, the DeLorean of data operations, capable of taking your data on a time-traveling adventure.</p>
<p>A Temporal Join is like a regular join but with a twist. It allows you to join a stream of data (the time traveler) with a versioned table (the timeline) based on the time attribute of the data stream. This means that for each record in the stream, the join will find the most recent record in the versioned table that is less than or equal to the stream record's time.</p>
<p>The versioned table is a normal state table where we keep track of data changes over time. That is, we keep older versions of each record around to allow the stream to match the correct version in time. Like time travel, temporal joins can make your head spin a bit. Let’s look at an example to break it down.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="tempinner">Temporal Join vs Inner Join<a href="https://docs.datasqrl.com/main/blog/temporal-join#tempinner" class="hash-link" aria-label="Direct link to Temporal Join vs Inner Join" title="Direct link to Temporal Join vs Inner Join" translate="no">​</a></h2>
<p>Back to our coffee roasting factory, we collect the temperature readings from each sensor in a data stream.</p>
<table><thead><tr><th>timestamp</th><th>sensorid</th><th>temperature</th></tr></thead><tbody><tr><td>2023-07-10T07:11:08</td><td>1</td><td>105.2</td></tr><tr><td>2023-07-10T07:11:08</td><td>2</td><td>83.1</td></tr><tr><td>...</td><td></td><td></td></tr><tr><td>2023-07-10T13:25:16</td><td>1</td><td>77.8</td></tr><tr><td>2023-07-10T13:25:16</td><td>2</td><td>83.5</td></tr></tbody></table>
<p>And we want to know the maximum temperature recorded for each machine.</p>
<p>Easy enough, let’s join the temperature data stream with the Sensors table and aggregate by machine id:</p>
<div class="language-sql codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-sql codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SELECT</span><span class="token plain"> s</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">machineid</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> </span><span class="token function" style="color:rgb(80, 250, 123)">MAX</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain">r</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">temperature</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> maxTemp </span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> SensorReading r </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">INNER</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">JOIN</span><span class="token plain"> Sensor s </span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ON</span><span class="token plain"> r</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">sensorid </span><span class="token operator">=</span><span class="token plain"> s</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">id </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">GROUP</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BY</span><span class="token plain"> s</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">machineid</span><br></div></code></pre></div></div>
<p>But here is a problem: What if we moved a sensor from one machine to another during the day? With an inner join, all of the sensor’s readings would be linked to the machine it was last placed on. So, if sensor 1 records a high temperature of 105 degrees in the morning and we move the sensor to the “Iron Roaster” machine in the afternoon, then we might see the 105 degrees falsely show up as the maximum temperature for the Iron Roaster. See how time played a trick on our join?</p>
<p>And this happens whenever we join a data stream with a state table that changes over time, like our sensors that get moved around the factory. What to do? Let’s call the temporal join to our rescue:</p>
<div class="language-sql codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-sql codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SELECT</span><span class="token plain"> s</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">machineid</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> </span><span class="token function" style="color:rgb(80, 250, 123)">MAX</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token plain">r</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">temperature</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> maxTemp </span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> SensorReading r </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">JOIN</span><span class="token plain"> Sensor </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FOR</span><span class="token plain"> SYSTEM_TIME </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">OF</span><span class="token plain"> r</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token identifier punctuation" style="color:rgb(248, 248, 242)">`</span><span class="token identifier">timestamp</span><span class="token identifier punctuation" style="color:rgb(248, 248, 242)">`</span><span class="token plain"> s</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">    </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ON</span><span class="token plain"> r</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">sensorid </span><span class="token operator">=</span><span class="token plain"> s</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">id </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">GROUP</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BY</span><span class="token plain"> s</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">machineid</span><br></div></code></pre></div></div>
<p>Pretty much the same query, just a different join type. Just a heads-up: the syntax for temporal joins in Flink SQL is more complex.</p>
<p>As a temporal join, we are joining each sensor reading with the version of the sensor record at the time of the data stream. In other words, the join not only matches the sensor reading with the sensor record based on the id but also based on the timestamp of the reading to ensure it matches the right version of the sensor record. Pretty neat, right?</p>
<p>Whenever you join a data stream with a state that changes over time, you want to use the temporal join to make sure your data is lined up correctly in time. Temporal joins are a powerful feature of stream processing engines that would be difficult to implement in a database.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="efficient">Why Temporal Joins are Fast and Efficient<a href="https://docs.datasqrl.com/main/blog/temporal-join#efficient" class="hash-link" aria-label="Direct link to Why Temporal Joins are Fast and Efficient" title="Direct link to Why Temporal Joins are Fast and Efficient" translate="no">​</a></h2>
<img src="https://docs.datasqrl.com/main/img/blog/flink_logo.svg" alt="Apache Flink >" width="30%">
<p>Not only do temporal joins solve the time-alignment problem when joining data streams with changing state, modern stream processors like Apache Flink are also incredibly efficient at executing temporal joins. A powerful feature with great performance? Sounds too good to be true. Let’s peek behind the stream processing curtain to find out why.</p>
<p>In stream processing, joins are maintained as the underlying data changes over time. That requires the stream engine to hold all the data it needs to update join records when either side of the join changes. This makes inner joins pretty expensive on data streams.</p>
<p>Consider our max-temperature query with the inner join: When we join a temperature reading with the corresponding sensor record, and that record changes, the engine has to update the result join record. To do so, it has to store all the sensor readings to determine which join results are affected by a change in a sensor record. This can lead to a lot of updates and hence a lot of downstream computation. It can also cause system failure when there are a lot of temperature readings in our data stream because the stream engine has to store all of them.</p>
<p>Temporal joins, on the other hand, can be executed much more efficiently. The stream engine only needs to store the versions of the sensor table that are within the time bounds of the sensor reading data stream. And it only has to briefly store (if at all) the sensor reading records to ensure they are joined with the most up-to-date sensor records. Moreover, temporal joins don’t require sending out a massive amount of updated join records when sensors change placement since the join is fixed in time.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="summary">Time to Wrap Up This Temporal Journey<a href="https://docs.datasqrl.com/main/blog/temporal-join#summary" class="hash-link" aria-label="Direct link to Time to Wrap Up This Temporal Journey" title="Direct link to Time to Wrap Up This Temporal Journey" translate="no">​</a></h2>
<p>We've reached the end of our time-traveling adventure through the universe of temporal joins. We've seen how they're like the DeLorean of data operations, zipping us back and forth through time to make sure our data matches up just right. We've also compared them to the good ol' inner join.</p>
<p>Temporal joins help us avoid the pitfalls of time-alignment problems when joining data streams with changing state. They're also super efficient, making them a great choice for high-volume, real-time data processing.</p>
<p>And that’s why the temporal join is stream processing's secret superpower.</p>
<p>DataSQRL makes using temporal joins a breeze. With its simplified syntax and smart defaults, it's like having a personal tour guide leading you through the sometimes bewildering landscape of stream processing. Take a look at our <a class="" href="https://docs.datasqrl.com/main/docs/intro/getting-started">Getting Started</a> to see a complete example of temporal joins in action or take a look at our <a class="" href="https://docs.datasqrl.com/main/docs/intro/examples">other tutorials</a> for a step-by-step guide to stream processing including temporal joins.</p>
<p>Happy data time-traveling, folks!</p>]]></content:encoded>
            <category>Join</category>
            <category>Flink</category>
            <category>DataSQRL</category>
        </item>
        <item>
            <title><![CDATA[Let's Uplevel Our Database Game: Meet DataSQRL]]></title>
            <link>https://docs.datasqrl.com/main/blog/lets-uplevel-database-datasqrl</link>
            <guid>https://docs.datasqrl.com/main/blog/lets-uplevel-database-datasqrl</guid>
            <pubDate>Mon, 15 May 2023 00:00:00 GMT</pubDate>
            <description><![CDATA[We need to make it easier to build data-driven applications. Databases are great if all your application needs is storing and retrieving data. But if you want to build anything more interesting with data - like serving users recommendations based on the pages they are visiting, detecting fraudulent transactions on your site, or computing real-time features for your machine learning model - you end up building a ton of custom code and infrastructure around the database.]]></description>
            <content:encoded><![CDATA[<p><strong>We need to make it easier to build data-driven applications.</strong> Databases are great if all your application needs is storing and retrieving data. But if you want to build anything more interesting with data - like serving users recommendations based on the pages they are visiting, detecting fraudulent transactions on your site, or computing real-time features for your machine learning model - you end up building a ton of custom code and infrastructure around the database.</p>
<p>You need a queue like Kafka to hold your events, a stream processor like Flink to process data, a database like Postgres to store and query the result data, and an API layer to tie it all together.</p>
<img src="https://docs.datasqrl.com/main/img/reference/full_logo.svg" alt="DataSQRL Logo >" width="30%">
<p>And that’s just the price of admission. To get a functioning data layer, you need to make sure that all these components talk to each other and that data flows smoothly between them. Schema synchronization, data model tuning, index selection, query batching … all that fun stuff.</p>
<p>The point is, you need to do a ton of data plumbing if you want to build a data-driven application. All that data plumbing code is time-consuming to develop, hard to maintain, and expensive to operate.</p>
<p>We need to make building with data easier. That’s why we are sending out this call to action to uplevel our database game. <strong>Join us in figuring out how to simplify the data layer.</strong></p>
<p>We have an idea to get us started: Meet DataSQRL.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="introducing-datasqrl">Introducing DataSQRL<a href="https://docs.datasqrl.com/main/blog/lets-uplevel-database-datasqrl#introducing-datasqrl" class="hash-link" aria-label="Direct link to Introducing DataSQRL" title="Direct link to Introducing DataSQRL" translate="no">​</a></h2>
<p>DataSQRL is a build tool that compiles your application’s data layer from a high-level data development language, dubbed SQRL.</p>
<p>Our goal is to create a new abstraction layer above the low-level languages often used in data layers, allowing a compiler to handle the tedious tasks of data plumbing, infrastructure assembly, and configuration management.</p>
<p>Much like how you use high-level languages such as Javascript, Python, or Java instead of Assembly for software development, we believe a similar approach should be used for data.</p>
<p>SQRL is designed to be a developer-friendly version of SQL, maintaining familiar syntax while adding features necessary for building data-driven applications, like support for nested data and data streams.</p>
<p>Check out this simple SQRL script to build a recommendation engine from clickstream data.</p>
<div class="language-sql codeBlockContainer_Ckt0 theme-code-block" style="--prism-color:#F8F8F2;--prism-background-color:#282A36"><div class="codeBlockContent_QJqH"><pre tabindex="0" class="prism-code language-sql codeBlock_bY9V thin-scrollbar" style="color:#F8F8F2;background-color:#282A36"><code class="codeBlockLines_e6Vv"><div class="token-line" style="color:#F8F8F2"><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">IMPORT</span><span class="token plain"> clickstream</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">Clickstream</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain"> </span><span class="token comment" style="color:rgb(98, 114, 164)">--Import clickstream data from Kafka</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">IMPORT</span><span class="token plain"> content</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">Content</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain">         </span><span class="token comment" style="color:rgb(98, 114, 164)">--Import content from CDC stream</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain" style="display:inline-block"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token comment" style="color:rgb(98, 114, 164)">/* Find next page visits within 10 minutes */</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">_CoVisits :</span><span class="token operator">=</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SELECT</span><span class="token plain"> b</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">url </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> beforeURL</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> a</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">url </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> afterURL</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">                    a</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">event_time </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> </span><span class="token identifier punctuation" style="color:rgb(248, 248, 242)">`</span><span class="token identifier">timestamp</span><span class="token identifier punctuation" style="color:rgb(248, 248, 242)">`</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">             </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> Clickstream b </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">INNER</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">JOIN</span><span class="token plain"> Clickstream a </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ON</span><span class="token plain"> b</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">userid</span><span class="token operator">=</span><span class="token plain">a</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">userid</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">                 </span><span class="token operator">AND</span><span class="token plain"> b</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">event_time </span><span class="token operator">&lt;</span><span class="token plain"> a</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">event_time </span><span class="token operator">AND</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">                     b</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">event_time </span><span class="token operator">&gt;=</span><span class="token plain"> a</span><span class="token punctuation" style="color:rgb(248, 248, 242)">.</span><span class="token plain">event_time </span><span class="token operator">-</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">INTERVAL</span><span class="token plain"> </span><span class="token number">10</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">MINUTE</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token comment" style="color:rgb(98, 114, 164)">/* Recommend pages that are visited shortly after */</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain"></span><span class="token comment" style="color:rgb(98, 114, 164)">/*+query_by_all(url) */</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">Recommendation :</span><span class="token operator">=</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">SELECT</span><span class="token plain"> beforeURL </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> url</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> afterURL </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> recommendation</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"></span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">                         </span><span class="token function" style="color:rgb(80, 250, 123)">count</span><span class="token punctuation" style="color:rgb(248, 248, 242)">(</span><span class="token number">1</span><span class="token punctuation" style="color:rgb(248, 248, 242)">)</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">AS</span><span class="token plain"> frequency </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">FROM</span><span class="token plain"> _CoVisits</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">                  </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">GROUP</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BY</span><span class="token plain"> beforeURL</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> afterURL</span><br></div><div class="token-line" style="color:#F8F8F2"><span class="token plain">                  </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ORDER</span><span class="token plain"> </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">BY</span><span class="token plain"> url </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">ASC</span><span class="token punctuation" style="color:rgb(248, 248, 242)">,</span><span class="token plain"> frequency </span><span class="token keyword" style="color:rgb(189, 147, 249);font-style:italic">DESC</span><span class="token punctuation" style="color:rgb(248, 248, 242)">;</span><br></div></code></pre></div></div>
<p>This little SQRL script imports clickstream data, identifies pairs of URLs visited within a 10-minute interval, and compiles these pairs into a set of recommendations, ordered by the frequency of co-visits.</p>
<img src="https://docs.datasqrl.com/main/img/diagrams/getting_started_diagram2.png" alt="Data pipeline >">
<p>DataSQRL then takes this script and compiles it into an integrated data pipeline, complete with all necessary data plumbing pre-installed. It configures access to the clickstream. It generates an executable for the stream processor that ingests, validates, joins, and aggregates the clickstream data. It creates the data model and writes the aggregated data to the database. It synchronizes timestamps and schemas between all the components. And it compiles a server executable that queries the database and exposes the computed recommendations through a GraphQL API.</p>
<p><strong>The bottom line: These 9 lines of SQRL code can replace hundreds of lines of complex data plumbing code and save hours of infrastructure setup.</strong></p>
<p>We believe that all this low-level data plumbing work should be done by a compiler since it is tedious, time-consuming, and error-prone. Let’s uplevel our data game, so we can focus on <strong>what</strong> we are trying to build with data and less on the <strong>how</strong>.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="join-us-on-this-journey">Join Us on this Journey<a href="https://docs.datasqrl.com/main/blog/lets-uplevel-database-datasqrl#join-us-on-this-journey" class="hash-link" aria-label="Direct link to Join Us on this Journey" title="Direct link to Join Us on this Journey" translate="no">​</a></h2>
<img src="https://docs.datasqrl.com/main/img/undraw/code.svg" alt="Join DataSQRL Community >" width="50%">
<p>We have the ambitious goal of designing a higher level of abstraction for data to enable millions of developers to build data-driven applications.</p>
<p>We <a href="https://github.com/DataSQRL/sqrl/releases/tag/v0.1.0" target="_blank" rel="noopener noreferrer" class="">just released</a> the first version of DataSQRL, and we recognize that we are at the beginning of a long, long road. We need your help. If you are a data nerd, like building with data, or wish it was easier, please <a href="https://github.com/DataSQRL/sqrl" target="_blank" rel="noopener noreferrer" class="">join us on this journey</a>. DataSQRL is an open-source project, and all development activity is transparent.</p>
<p>Here are some ideas for how you can contribute:</p>
<ul>
<li class="">Share your thoughts: Do you have ideas on how we can improve the SQRL language or the DataSQRL compiler? Jump into <a class="" href="https://docs.datasqrl.com/main/community">our community</a> and let us know!</li>
<li class="">Test the waters: Do you like playing with new technologies? Try out <a class="" href="https://docs.datasqrl.com/main/docs/intro/getting-started">DataSQRL</a> and let us know if you find any bugs or missing features.</li>
<li class="">Spread the word: Think DataSQRL has potential? Share this blog post and <a href="https://github.com/DataSQRL/sqrl" target="_blank" rel="noopener noreferrer" class="">star</a> DataSQRL on <a href="https://github.com/DataSQRL/sqrl" target="_blank" rel="noopener noreferrer" class="">Github</a>. Your support can help us reach more like-minded individuals.</li>
<li class="">Code with us: Do you enjoy contributing to open-source projects? Dive into <a href="https://github.com/DataSQRL/sqrl" target="_blank" rel="noopener noreferrer" class="">the code</a> with us and pick up a <a href="https://github.com/DataSQRL/sqrl/issues" target="_blank" rel="noopener noreferrer" class="">ticket</a>.</li>
</ul>
<p>Let’s uplevel our database game. With your help, we can make building with data fun and productive.</p>
<h2 class="anchor anchorTargetStickyNavbar_Vzrq" id="more-information">More Information<a href="https://docs.datasqrl.com/main/blog/lets-uplevel-database-datasqrl#more-information" class="hash-link" aria-label="Direct link to More Information" title="Direct link to More Information" translate="no">​</a></h2>
<p>You probably have a ton of questions now. How do I import my own data? How do I customize the API? How do I deploy SQRL scripts to production? How do I import functions from my favorite programming language?</p>
<p>Those are all great questions. Check out <a class="" href="https://docs.datasqrl.com/main/docs/intro">the documentation</a> for answers.</p>]]></content:encoded>
            <category>DataSQRL</category>
            <category>Community</category>
        </item>
    </channel>
</rss>