Agents write data code faster than humans can verify it
Writing code is no longer the bottleneck for data teams. Trusting and managing it is.
General-purpose coding agents solve data engineering tasks with dozens of files in multiple languages. The results look plausible but contain subtle bugs and misalignments that are hard to spot.
Automate data engineering without giving up control
Human Control
One SQL file your team can read.Not thousands of lines of Python, dbt, YAML, and glue. The whole pipeline, from ingest to API, is expressed in SQL you can understand, run, and approve.
Correctness
A compiler instead of guesswork.Deterministic generation removes errors at the seams between systems. Relational validation and event-time replay tests catch the bugs that code review misses.
Safety
Guardrails you can inspect and extend.Problems surface at compile time. Every compile writes out lineage, schemas, and the full data flow, so you can enforce your organization's policies automatically.
IMPORT banking_data.*;
-- Latest version of each account from the CDC stream
Accounts := DISTINCT AccountsCDC ON account_id ORDER BY update_time DESC;
-- Enrich transactions with the account as of transaction time
SpendingTransactions :=
SELECT t.*, h.name AS creditor_name, h.type AS creditor_type
FROM Transactions t
JOIN Accounts FOR SYSTEM_TIME AS OF t.tx_time a
ON t.credit_account_id = a.account_id
JOIN AccountHolders FOR SYSTEM_TIME AS OF t.tx_time h
ON a.holder_id = h.holder_id;
/** Spending transactions for the authenticated account
within [from_time, to_time). Exposed via GraphQL, REST, and MCP. */
SpendingTransactionsByTime(
account_id STRING NOT NULL METADATA FROM 'auth.accountId',
from_time TIMESTAMP NOT NULL,
to_time TIMESTAMP NOT NULL
) :=
SELECT * FROM SpendingTransactions
WHERE debit_account_id = :account_id
AND :from_time <= tx_time AND :to_time > tx_time
ORDER BY tx_time DESC;
Human Control: Review the Logic, Not the Plumbing
The agent's output is declarative SQL covering ingestion, transformation, storage, and a secure MCP/REST/GraphQL endpoint. Easy to understand, review, and argue about.
Aggregations, units, filters, and joins sit on one screen. One command runs the whole pipeline locally, API included, so you can inspect real results and experiment quickly.
Correctness: Deterministic Where It Matters
Language models are probabilistic. Mapping types, keys, schemas, and connectors across Flink, Kafka, Postgres, Iceberg, and the API layer requires strict rule-following that's better handled by a compiler.
Every deployment asset is generated from one logical model, so the systems cannot disagree.
-- Snapshot test, replayed at original event timestamps
/*+ test */
SpendingByHolderTest :=
SELECT creditor_name, COUNT(*) AS tx_count, SUM(amount) AS total
FROM SpendingTransactions
GROUP BY creditor_name ORDER BY creditor_name;
-- Assert that every transaction was enriched with its creditor
/*+ test(no_rows) */
NoUnenrichedTransactions :=
SELECT * FROM SpendingTransactions WHERE creditor_name IS NULL;
Correctness: Test the Pipeline in Motion
Static test fixtures miss the bugs that only appear over time. The simulator runs the real deployment artifacts and replays events at their original timestamps.
Late and out-of-order data, races between streams, idle sources, updates, and deletes become deterministic tests.
Why standard integration tests fall short →=== CustomerTransaction
Type: stream
Stage: flink
Inputs: _CardAssignment, _Merchant, sources.Transaction
Annotations:
- stream-root: Transaction
Primary Key: transactionId, time
Timestamp : time
Schema:
- transactionId: BIGINT NOT NULL
- cardNo: VARCHAR NOT NULL
- time: TIMESTAMP_LTZ(3) *ROWTIME* NOT NULL
- amount: DOUBLE NOT NULL
- merchantName: VARCHAR NOT NULL
- customerId: BIGINT NOT NULL
Safety: Every Pipeline Is Fully Inspectable
Each compile writes out the complete data flow: table types, inferred keys, timestamps, schemas, engine assignments, and every deployment asset.
Use these artifacts for lineage, impact analysis, and audit. Add custom rules that enforce your policies on PII, naming, retention, or residency for every pipeline the agent builds.

Safety: Guardrails by Construction
The compiler rejects invalid engine assignments, capability mismatches, and inconsistent data flows before anything is deployed.
API endpoints use parameterized queries, and authorization is bound to JWT claims in the SQL itself. Agent-built endpoints get these protections from the compiler, not from a prompt.
Your Harness, Your Organization
Every data organization has its own conventions, domain vocabulary, compliance requirements, and target infrastructure. DataSQRL is a harness you use to build an agent that honors those.
| Skills | How your team gathers requirements, plans, implements, tests, and deploys, plus domain and catalog knowledge |
| Validators & policies | Custom compiler rules for governance, security, and data quality |
| Functions & connectors | Your UDFs, sources, sinks, and formats |
| Engines & deployment | Flink, Kafka, Postgres, Iceberg on Docker, Kubernetes, or cloud |
| Coding agent | Claude Code, Codex, OpenCode, Pi, etc: your choice |
# The agent's inner loop, driven by the harness
# 1. Validate and explain the data flow
docker run --rm -v $PWD:/workspace \
datasqrl/cmd compile package.json
# 2. Event-time replay tests
docker run --rm -v $PWD:/workspace \
datasqrl/cmd test test-package.json
# 3. Deployment assets for K8s or cloud
ls build/deploy/plan
What a DataSQRL Data Engineering Agent Builds
Streaming & Batch Pipelines
CDC, temporal joins, windowed aggregations, and deduplication on Flink, Kafka, and Iceberg.
Data APIs
GraphQL, REST, and MCP endpoints generated from SQL, with authentication and authorization built in.
Data Products
Curated, documented, and tested datasets for analytics and downstream teams.
Operational Data
Low-latency data for applications and AI agents, including embeddings and LLM enrichment.

Runs on Infrastructure You Own
DataSQRL compiles to Flink, Kafka, Postgres, and Iceberg and deploys to Docker, Kubernetes, or managed cloud services. There is no proprietary runtime, and the harness is open source.
All you need is Docker and an API key from your LLM provider. Run the DataSQRL agent in your project folder:
- macOS
- Windows
- Linux
docker run -e ANTHROPIC_API_KEY -it --rm \
--detach-keys="ctrl-],ctrl-]" -e TERM -e COLORTERM \
-v "$PWD":/workspace -w /workspace datasqrl/core-agent
docker run -e ANTHROPIC_API_KEY -it --rm `
--detach-keys="ctrl-],ctrl-]" -e TERM -e COLORTERM `
-v "${PWD}:/workspace" -w /workspace datasqrl/core-agent
docker run -e ANTHROPIC_API_KEY -it --rm \
--detach-keys="ctrl-],ctrl-]" -e TERM -e COLORTERM \
-v "$PWD":/workspace -w /workspace datasqrl/core-agent
Using OpenAI, Bedrock, Azure, or Vertex AI? Swap in that provider's environment variables, as shown in the getting started guide.
Getting Started
The agent wraps the Pi coding agent with the DataSQRL framework and skills. Once it is running, tell it what you need in plain English:
Build a pipeline that ingests our order data from Kafka in real time and serves hourly revenue per product through an API.
You get SQL scripts with tests that you can read, run, and verify.
Want planning, iterative refinement, and deployment workflows? Install the advanced DataSQRL agent as a plugin for Claude Code, Codex, Cursor, or GitHub Copilot.
Getting Started GuidePlugin Docs