ELTMaestro
Platform

Four parts, one metadata model.

Pipelines are data. The designer writes them, the engine executes them, the audit database remembers them. Nothing is compiled into a proprietary runtime you cannot inspect.

In the designer

What your team actually works in

Captured from a reference banking build on synthetic data: a 19-workflow nightly batch, its dimensions and facts, and the quality gates between them. Click any screen to see it at full size.

Designer · a master batch of 19 child workflows with sync barriers and data-quality gates between layers
Runtime Status · the live run, step by step, barrier by barrier
Step Status · per-step timings plus the control tests that gated the run (row counts, SCD2 integrity, date coverage)
Designer · a gold-layer mart: four inputs, a JOIN, a shaping expression step, one target
Designer · a Type 2 slowly changing dimension in three steps
Expression builder · typed source columns, per-output expressions, aggregate or scalar mode
Run History · every run with its outcome and duration; the failed run stays visible
Control tests · expected-versus-actual queries with a tolerance, grouped into hierarchies
Change history · compare any saved version with the current job on the canvas
Workspace · every workflow with its platform, schedule and tags; search, open, lock and migrate from here
Architecture

What runs where

Designer

Windows desktop client

Drag steps onto a canvas, wire arrows, map columns. Check Mapping validates every alias before you save. Runtime status, console logs, step history and a change-history compare live in the same window.

Meta-service

Spring Boot · Java 21

Authenticates users, stores job definitions and connections, migrates jobs and control tests between environments, and spawns engine runs. Every call is audited.

Engine

Java 21 · Spark 4

One short-lived JVM per batch. Reads the job, builds the step DAG, runs each step as a thread, pushes the heavy lifting into the warehouse, and writes status, metrics and logs to the audit database.

Audit database

PostgreSQL 16

Runs, step status, watermarks, control-test results, metrics and AI interactions in one place. Dashboards and reports read it directly.

  Designer (Windows) ──SOAP/TLS──▶ Meta-service ──JDBC──▶ Audit DB (PostgreSQL)
                                       │ spawns                      ▲
                                       ▼                             │ status · metrics · logs
                               Engine (per-batch JVM) ───────────────┘
                                       │ renders SQL, bulk loads
                                       ▼
            Snowflake · Redshift · ClickHouse · Databricks · Synapse · Greenplum · Netezza · Yellowbrick · …
                                       │
                                       ▼
                         Dashboards · AI sidecar · downstream systems
Connectivity

Targets and sources

A connection is a row in the registry. Point a job at a different one and the engine renders the right dialect, bulk loader and staging path.

Warehouse targets
SnowflakeAmazon RedshiftClickHouseDatabricksAzure SynapseGreenplumNetezzaYellowbrickExasolFireboltSpark on HDFSPostgreSQL
Sources
OracleSQL ServerMySQL / MariaDBPostgreSQLSalesforceAmazon S3Azure BlobSFTP / filesAny JDBC source
Step catalogue

Forty-plus steps, one contract each

Ingest
  • Parallel JDBC extractor with partitioning
  • Salesforce extractor with rolling deltas
  • Schema loader for whole-schema landing
  • File scanners, SFTP pull/push, S3 and Azure Blob staging
Transform
  • Join, union, pivot, function and aggregate steps
  • SCD Type 1, 2 and 4 dimensions
  • Expression builder with warehouse-native functions
  • SQL script steps for anything bespoke
Quality
  • Control tests with percentage or absolute tolerance
  • Quarantine zones for rows that fail a rule
  • Column profiling and schema-drift detection
  • Metrics captured on every run
Orchestrate
  • Parent/child workflows with sync barriers and switches
  • Continue-on-failure per job step, batch still reported honestly
  • Cron scheduler with calendars and run-variable overrides
  • Edge engine nodes dispatched over SSH
Operate
  • Runtime status, console logs and step history in the client
  • Change history with before/after canvas compare
  • Alert hook into any pager, chat or ticketing tool
  • Git-based promotion: branch per environment, Jenkins deploys
Load
  • Native bulk loaders: COPY, PUT, clickhouse-client, HDFS Parquet
  • Automatic target creation and column widening
  • Parquet schema profiles reused across steps
  • Row-count assertions between extract and load
Incremental loading

Watermarks that cannot lie

Three mechanisms, one rule: the high-water mark moves only when the data landed.

Batch window

Every run gets a data-validity window computed from the history of completed runs. The lower bound advances only on COMPLETE, so a failed run never skips data.

Step watermark

A delta column on the source and a MAX query on the target give each step its own cursor. Full load and incremental load are the same job with one flag.

Step scope

A step added to a batch that has run for months back-fills from the beginning on its first run instead of inheriting last night's delta.

Deployment

Bare metal, container, or edge

Portable tarball

RHEL-family or Debian host, JDK 21 and PostgreSQL 16. One install script, in-place upgrades that keep the audit database and your tuning.

Docker image

The whole server stack in one image: database, meta-service, engine, Spark, scheduler. Configuration survives image upgrades.

Edge engine nodes

Engine-only nodes next to remote sources, dispatched by the master over SSH and reporting to the central audit database. Proven across island sites on one telecom estate.