Skip to main content

DataSQRL Documentation

DataSQRL is an open-source data engineering harness for building data engineering agents designed around human control, correctness, and safety. It extends your coding agent (Claude Code, Codex, OpenCode, Pi, and others) with a SQL compiler, a validator, an event-time simulator, and your own skills and policies. The result is an agent you can trust with data pipelines, batch jobs, data APIs (REST, GraphQL, MCP), data products, and operational data.

DataSQRL harness: your skills and policies, a coding agent, and the DataSQRL framework packaged as one data engineering agent

How It Works​

  1. The agent writes SQL. The whole pipeline, from ingest to transform to store to serve, is expressed in SQRL: SQL extended with stream processing and API definitions. It stays readable enough for a human to review.
  2. The compiler validates and generates. DataSQRL checks the logical plan (schemas, keys, timestamps, table types) and the physical plan (engine capabilities, type mappings). It then generates every deployment asset from one model: Flink plans, Kafka topics, Postgres and Iceberg schemas, and GraphQL, REST, and MCP APIs.
  3. The simulator tests. Pipelines run locally with timestamp-accurate event replay, so time-dependent behavior becomes a deterministic test.
  4. You review and deploy. Compilation outputs such as the pipeline DAG and lineage support human review and automated policy checks. The artifacts run on open-source infrastructure you operate yourself.

For the full design, read the harness architecture.

Where to Go Next​

I want to…Go to
Try it on my own dataGetting Started: run the basic agent in Docker, or install the plugin for Claude Code, Codex, Cursor, or Copilot
See what it can buildExamples: data products for a retail bank, plus self-contained pipelines across many use cases
Compare DataSQRL to Flink, Spark, dbt, and other toolsFAQ: short answers to the questions data engineers ask most
Understand why a harness mattersThe four-part series: broken seams, relational introspection, event-time testing, and human understanding
Read and review the SQL an agent producesSQRL Language and Streaming Concepts
Connect my data sources and sinksConnectors
Shape the APIs and data productsInterface
Choose engines and deployConfiguration and Deployment: DataSQRL Cloud, managed cloud services, or Kubernetes
Add custom logicFunctions: the built-in library and your own UDFs
Compile, test, and run from the command lineCompiler
Inspect and validate what the compiler producesCompilation Output
Customize or extend the harness itselfHow DataSQRL Works

Documentation Map​

Start here

  • Getting Started: set up the DataSQRL agent and build your first pipeline
  • Examples: a gallery of what DataSQRL can build
  • FAQ: how DataSQRL compares to other tools, and common questions

Core concepts

Advanced

Community & Support​

DataSQRL is open source. Report bugs in GitHub Issues and ask questions or share feedback in GitHub Discussions. Contributions are welcome.