Examples Gallery
A gallery of what you can build with the DataSQRL data engineering harness, from a data product suite for a retail bank to small, self-contained pipelines across a range of use cases. Every example is open source. You can read the SQL, run it locally, and run its tests.
Data Products for a Retail Bankโ
This example shows how DataSQRL works in an organizational setting. A fictional retail bank keeps a semantic data catalog of its source and enriched datasets. A coding agent with the DataSQRL harness uses that catalog to build data products: pipelines and APIs that serve specific business needs.
- Data Catalog: the bank's datasets, schemas, connectors, relationships, and data quality rules
- Data Products: two AI-generated data products built from that catalog
Explore the respective Github repositories to learn more.
Self-Contained Examplesโ
The DataSQRL Examples repository contains smaller, self-contained pipelines across a range of use cases. Each one is a standalone project with sample data and a README that explains how to run it. Pick the one closest to what you want to build:
| Example | What it builds | Look here for |
|---|---|---|
| Getting Started Examples | Seven minimal pipelines: Kafka to console, Kafka to Kafka, stream joins, files and Kafka to Iceberg (local or AWS Glue), and Avro with Schema Registry | Connector patterns and your first pipeline |
| Finance Credit Card Chatbot | Enriched transaction analytics for a GenAI chatbot, a credit card rewards program, and a batch variant that writes spending views to Iceberg for DuckDB and Snowflake | Enrichment, APIs for AI agents, batch to Iceberg |
| Clickstream AI Recommendation | Personalized content recommendations from clickstream data and LLM-generated vector embeddings | Vector embeddings and real-time recommendations |
| Healthcare Study | Three use cases over one shared catalog: a real-time API, analytics in Iceberg, and an enriched stream published to Kafka | One catalog feeding API, analytics, and streaming |
| Oil & Gas Agent Automation | A monitoring API for an AI agent plus an operations backend with an ingest mutation and a low-flow-rate alert subscription | Event-triggered agents and subscriptions |
| IoT Sensor Metrics | An event-driven microservice that ingests sensor readings and serves metrics and alerts | Ingest APIs and time-windowed metrics |
| Logistics Shipping | Real-time shipment tracking with locations, in about 30 lines of SQL | A compact streaming pipeline |
| Law Enforcement | An integrated view of drivers, vehicles, warrants, and BOLOs, with analytics and alerts for traffic stops | Combining databases and streams |
| Iceberg Data Deduplication | Compaction and deletion jobs for Iceberg tables | Maintaining data lake tables |
| User-Defined Functions | A custom function shipped with JBang or a Maven project | Extending SQL with your own logic |
The repository also includes a data generator for producing larger datasets for experiments and benchmarks.
Build Your Ownโ
Start from your own catalog or data sources and describe the data product you need. The Getting Started guide shows how to set up the DataSQRL agent.