ETL ·· GRIDLENS

Data Pipeline & Lineage

A reproducible Extract → Transform → Load pipeline (etl/build_dataset.py) with enforced data-quality contracts feeds the warehouse. Deterministic seed = reproducible analytics.

last run: green
Pipeline Lineage
3 stages · 4 source tables → 1 warehouse
STAGE 01
Extract
Python · csv
Pull operational telemetry from 4 source systems into raw CSV.
→
STAGE 02
Transform
derive · validate
Compute capacity factor, CO₂, availability; enforce data-quality contracts.
→
STAGE 03
Load
JSON · SQLite
Emit warehouse snapshot; hydrate in-browser SQLite for self-service analytics.
Data Quality Contracts
Enforced at transform time — build fails if violated
✓capacity_factor ∈ [0, 1]
generation
✓mwh_generated ≥ 0
generation
✓commissioned_year ∈ (1900, 2026]
facilities
✓capacity_mw > 0
facilities
✓downtime_hours ≥ 0
incidents
✓region_id NOT NULL
regions
Warehouse Snapshot
loading…
Re-run the pipeline anytime with npm run etl. Output is deterministic (seeded), so dashboards stay reproducible across builds.