Frequently Asked Questions
Why do I need DataSQRL in addition to my coding agent?โ
Claude Code, Codex, and similar agents write plausible data pipeline code, but data engineering has failure modes that general-purpose coding agents miss and that pass code review. DataSQRL adds a harness around your agent that catches them:
- The seams between systems. Types, schemas, and connector configurations are generated by a compiler instead of being handwritten in each system. Read more
- Bugs that aren't in the query text. Wrong keys, time-dependent joins, stalled watermarks, and unbounded state are caught by relational validation. Read more
- Bugs that only appear over time. Late data, races, and updates are tested with event-time replay. Read more
- Code nobody can review. The whole pipeline is expressed in SQL that a person can read in one sitting. Read more
Your agent keeps doing the reasoning. DataSQRL gives it guardrails and feedback, and it keeps you in control of what gets built.
How does DataSQRL compare to Apache Flink or Apache Spark?โ
DataSQRL doesn't replace them. It sits on top. Flink, Kafka, Postgres, Iceberg, DuckDB, Spark, Trino, and the API layer each need their own code and configuration. DataSQRL lets you define the whole pipeline once in SQL and compiles it into the assets each engine runs: Flink compiled plans, Kafka topics, database schemas and queries, and API definitions.
You keep the full power of each engine. The compiler writes out the artifacts at every level, so you can inspect exactly what runs where and tune it through configuration and hints. Flink is the processing engine today. The engine architecture is modular, so other data engines can be added. See Configuration and How DataSQRL Works.
How does DataSQRL compare to dbt or SQLMesh?โ
We share their belief that SQL is the right language for data work, but DataSQRL solves a different problem.
dbt helps people manage many SQL models across files, mostly through templating, which is string manipulation that does not understand the SQL. Coding agents already handle that kind of file organization well. DataSQRL focuses on what agents cannot guarantee on their own:
- Deep relational introspection. It parses the whole pipeline into a relational plan and validates keys, timestamps, table types, and engine capabilities.
- Event-time testing. It replays real-time and batch data deterministically.
- Local build and execution. Run the entire stack locally to inspect results and iterate quickly.
- End-to-end scope. Ingestion, streaming, storage, and API serving are covered in one model, not just warehouse transformations. Connectors are compiled deterministically and kept in sync.
SQLMesh parses SQL and adds features like column-level lineage, but it also targets batch transformations in a warehouse. Neither tool generates streaming pipelines, database integrations, or APIs.
How does DataSQRL compare to streaming databases like Materialize, RisingWave, or ksqlDB?โ
Streaming databases run incremental SQL inside a single system. DataSQRL is a compiler and harness that spans several systems. It splits a pipeline across a stream processor, a log, a database or table format, and an API server, depending on what each part needs. It also adds the pieces an agent needs to produce correct pipelines: validation, event-time replay tests, and inspectable compilation output. The engine architecture is extensible, so a streaming database could serve as an engine in a DataSQRL pipeline. But DataSQRL empowers you to combine the technologies you already trust to solve most of your data problems instead of having to adopt yet another database.
Does DataSQRL replace Airflow, Dagster, Fivetran, or Airbyte?โ
No, they are complementary. Orchestrators schedule and coordinate jobs, and ingestion tools move data into your platform. DataSQRL defines and compiles the integration, processing, and serving logic. It reads from the systems those tools write to, such as Kafka topics, databases, and Iceberg tables, and its compiled jobs can be deployed and scheduled with your existing tooling.
Is DataSQRL only for real-time streaming?โ
No. The same SQL runs as a streaming or a batch pipeline, depending on the Flink runtime mode. Batch pipelines can be split into sequential sub-batches. A common pattern is batch processing into Iceberg tables that are queried through DuckDB or Snowflake. You can also use DataSQRL to build data products and views in your existing data lake(house). See the examples for both styles.
Does it work with our existing infrastructure?โ
DataSQRL compiles to widely used open-source technologies: Apache Flink, Apache Kafka (and compatible systems such as Redpanda), PostgreSQL, and Apache Iceberg, with DuckDB, Spark, Trino, Snowflake, and others as query engines. It reads from and writes to your existing systems through connectors, and it deploys to Docker, Kubernetes, or managed cloud services. See Configuration for the supported engines.
Are we locked in? What if we stop using DataSQRL?โ
DataSQRL is open source under the Apache 2.0 license. Its output is standard deployment assets: Flink compiled plans, SQL schemas and queries, Kafka topic definitions, and GraphQL, REST, and MCP definitions. These run on standard open-source infrastructure with no proprietary runtime. You can also write and maintain SQRL by hand without an agent.
What if SQL can't express my logic?โ
Extend it. DataSQRL supports user-defined functions in Java. Simple functions can be single-file JBang scripts, and complex ones can be full Java projects that define new table operators. You can also build your own function libraries that capture your business semantics, so agents and people reuse the same building blocks.
In addition, DataSQRL supports SQL passthrough to the underlying engines for SQL features that the underlying engine supports but DataSQRL does not yet, so that the abstraction layer does not become a limiting factor.
How do we enforce our organization's standards and policies?โ
Customize the harness:
- Skills and conventions capture how your team gathers requirements, plans, tests, and deploys.
- Custom validation rules check governance, security, and data quality policies, such as PII handling or naming, on every compile.
- Compile artifacts, including the full data flow DAG, schemas, and lineage, feed audits and automated policy checks.
A semantic data catalog with sensitivity and regulatory tags gives the agent the context to respect those policies from the start.