ELTMaestro
In production since 2015 · Java 21 engine · Windows designer

Metadata-driven ELT for every warehouse you run.

Design pipelines visually, push the work into Snowflake, Redshift, ClickHouse or a dozen other engines, and let a governed AI layer answer questions on top. One definition, one audit trail, no silent failures.

12+
warehouse targets
40+
step types
1
audit database for everything
MASTER_DAILY_BATCH · running live
OracleSQL ServerSalesforceParallel extractControl testSCD2 dimensionWarehouseDashboardsAI sidecar
SnowflakeAmazon RedshiftClickHouseDatabricksAzure SynapseGreenplumNetezzaYellowbrickExasolFireboltSpark on HDFSPostgreSQLOracleSQL ServerMySQL / MariaDBPostgreSQLSalesforceAmazon S3Azure BlobSFTP / filesAny JDBC sourceSnowflakeAmazon RedshiftClickHouseDatabricksAzure SynapseGreenplumNetezzaYellowbrickExasolFireboltSpark on HDFSPostgreSQLOracleSQL ServerMySQL / MariaDBPostgreSQLSalesforceAmazon S3Azure BlobSFTP / filesAny JDBC source
Why ELTMaestro

Built for the warehouse you have, and the next one.

Most integration tools either lock the logic into a proprietary runtime or leave you writing SQL per platform. ELTMaestro stores the pipeline as metadata and renders it for the engine behind each connection.

Design once, run anywhere

Pipelines are built visually in the desktop designer and stored as metadata. The engine renders the SQL for whichever warehouse the connection points at, so a job moves between Snowflake, Redshift and ClickHouse without a rewrite.

Incremental by default

Watermarks, data-validity windows and step-level history let a nightly batch and a 15-minute micro-delta use the same job with one flag flipped. A step that has never run back-fills its history instead of silently taking one delta.

Failures are visible

Control tests gate every layer, a failed bulk load fails the step, orphan rows are flagged rather than dropped, and every run writes a full audit record. The platform never reports COMPLETE over data that did not land.

See it

A real batch, not a diagram.

The designer, the runtime view and the quality gates, captured from a reference banking build on synthetic data. Click to enlarge.

Designer · master daily batch, 19 child workflows, three quality gates
Runtime Status · the same batch while it runs
Step Status · timings and the control tests that gated the run
Capabilities

Everything between the source and the dashboard.

Forty-plus step types cover ingest, transform, quality, orchestration and load. Each one is a dialog in the designer and a class in the engine, with the same contract on both sides.

Ingest
  • Parallel JDBC extractor with partitioning
  • Salesforce extractor with rolling deltas
  • Schema loader for whole-schema landing
  • File scanners, SFTP pull/push, S3 and Azure Blob staging
Transform
  • Join, union, pivot, function and aggregate steps
  • SCD Type 1, 2 and 4 dimensions
  • Expression builder with warehouse-native functions
  • SQL script steps for anything bespoke
Quality
  • Control tests with percentage or absolute tolerance
  • Quarantine zones for rows that fail a rule
  • Column profiling and schema-drift detection
  • Metrics captured on every run
Orchestrate
  • Parent/child workflows with sync barriers and switches
  • Continue-on-failure per job step, batch still reported honestly
  • Cron scheduler with calendars and run-variable overrides
  • Edge engine nodes dispatched over SSH
Operate
  • Runtime status, console logs and step history in the client
  • Change history with before/after canvas compare
  • Alert hook into any pager, chat or ticketing tool
  • Git-based promotion: branch per environment, Jenkins deploys
Load
  • Native bulk loaders: COPY, PUT, clickhouse-client, HDFS Parquet
  • Automatic target creation and column widening
  • Parquet schema profiles reused across steps
  • Row-count assertions between extract and load
3,100+
workflows in one production estate
telecommunications operator, in production since 2015
12
database technologies feeding one platform
across 119 registered connections
99.2%
run success over 165,000+ batch runs
measured from the audit database
7 TB
processed per day inside the window
call-detail records, customer-reported
AI sidecar

Ask the warehouse a question. Get governed SQL back.

A natural-language layer that only ever sees classified metadata, refuses by rule, records every question, and runs on the cloud or entirely on your own hardware.

# question
How many customers appear in more than one business unit?
# generated, validated, executed as a read-only account
SELECT count(*) FROM gold_group.mart_customer_360
WHERE business_unit_count > 1;
→ 3,009 · 2.8 s · audited
# question
List every customer's national ID and mobile number
✕ Refused: those columns are not in the visible schema.

Classification lives in a catalogue table, not in a prompt. A column tagged PII never reaches the model, and a question that needs it is refused before any SQL exists. Every interaction is written to the same audit database as your batches, with tokens, duration and outcome.

  • ✓ Subject-area scoping: 751 visible columns become 59 when "finance" is ticked
  • ✓ Three tiers: cloud API, on-host CPU, on-host GPU, same gatekeeper for all
  • ✓ Results publish straight into dashboards as charts
How the sidecar works →

See it running on your data in a week.

We stand up a reference pipeline against one of your sources, with control tests and a dashboard, and walk your team through it.