DataSQRL Documentation
DataSQRL is an open-source data engineering harness for building data engineering agents designed around human control, correctness, and safety. It extends your coding agent (Claude Code, Codex, OpenCode, Pi, and others) with a SQL compiler, a validator, an event-time simulator, and your own skills and policies. The result is an agent you can trust with data pipelines, batch jobs, data APIs (REST, GraphQL, MCP), data products, and operational data.
How It Worksβ
- The agent writes SQL. The whole pipeline, from ingest to transform to store to serve, is expressed in SQRL: SQL extended with stream processing and API definitions. It stays readable enough for a human to review.
- The compiler validates and generates. DataSQRL checks the logical plan (schemas, keys, timestamps, table types) and the physical plan (engine capabilities, type mappings). It then generates every deployment asset from one model: Flink plans, Kafka topics, Postgres and Iceberg schemas, and GraphQL, REST, and MCP APIs.
- The simulator tests. Pipelines run locally with timestamp-accurate event replay, so time-dependent behavior becomes a deterministic test.
- You review and deploy. Compilation outputs such as the pipeline DAG and lineage support human review and automated policy checks. The artifacts run on open-source infrastructure you operate yourself.
For the full design, read the harness architecture.
Where to Go Nextβ
| I want to⦠| Go to |
|---|---|
| Try it on my own data | Getting Started: run the basic agent in Docker, or install the plugin for Claude Code, Codex, Cursor, or Copilot |
| See what it can build | Examples: data products for a retail bank, plus self-contained pipelines across many use cases |
| Compare DataSQRL to Flink, Spark, dbt, and other tools | FAQ: short answers to the questions data engineers ask most |
| Understand why a harness matters | The four-part series: broken seams, relational introspection, event-time testing, and human understanding |
| Read and review the SQL an agent produces | SQRL Language and Streaming Concepts |
| Connect my data sources and sinks | Connectors |
| Shape the APIs and data products | Interface |
| Choose engines and deploy | Configuration and Deployment: DataSQRL Cloud, managed cloud services, or Kubernetes |
| Add custom logic | Functions: the built-in library and your own UDFs |
| Compile, test, and run from the command line | Compiler |
| Inspect and validate what the compiler produces | Compilation Output |
| Customize or extend the harness itself | How DataSQRL Works |
Documentation Mapβ
Start here
- Getting Started: set up the DataSQRL agent and build your first pipeline
- Examples: a gallery of what DataSQRL can build
- FAQ: how DataSQRL compares to other tools, and common questions
Core concepts
- SQRL Language: imports and exports, table functions and relationships, hints, subscriptions, and stream and state semantics
- Connectors: ingest from and export to Kafka, databases, data lakes, and files
- Interface: generated GraphQL, REST, and MCP APIs and data product tables, and how to customize them
- Configuration: engines, connectors, dependencies, and compiler options in
package.json, with a page for each engine (Flink, Kafka, Postgres, Iceberg, Iceberg query engines, Vert.x) and the default configuration - Functions: system and library functions, plus custom functions
- Compiler: the
compile,test, andruncommands, the compilation output they produce for review and validation, and how to deploy it - Streaming Concepts: time, watermarks, and other stream processing basics
Advanced
- How DataSQRL Works: internal architecture and advanced customization
- Compatibility: version compatibility and migration
Community & Supportβ
DataSQRL is open source. Report bugs in GitHub Issues and ask questions or share feedback in GitHub Discussions. Contributions are welcome.