01Data that arrives.
Reliable by design.
I design reliable data pipelines and thoughtful architectures that move data efficiently, preserve its meaning, and make every destination ready for confident decisions.
Designing dependable AI data engineering systems that connect governed data, retrieval and intelligent automation.

I build the layer between data and decisions.
I am Okang Michael Ozeh, a data engineering and analytics specialist focused on designing reliable pipelines, well-structured data platforms and analytics-ready models.
My work spans batch processing, real-time streaming, change data capture, lakehouse architecture, dimensional modelling and business intelligence. The goal is consistent: move data efficiently from source to destination while preserving its quality, meaning and traceability.
I approach every project as an engineering problem. I define the data grain, design for failure and recovery, test critical boundaries, and document the decisions that make each system dependable and understandable.
My engineering loop
Understand the event. Protect its meaning. Separate fast operational paths from durable analytical ones. Test the boundaries. Make the result explainable. Then iterate.
Core toolkit
Systems, built end to end.
Two focused tracks covering reliable data platforms and decision-ready analytics products.
Upcoming Platform Architecture
Project Danube defines the governed CDC platform planned for the next implementation cycle.
01Data Engineering Platforms
Implemented streaming and batch systems that moved, recovered and served data reliably.
Project Alpine
Engineering notes+
Independent Python producers validated Kraken and Coinbase trade records and published each exchange to its own Kafka topic. Spark continuously normalized those streams into Bronze and a shared Silver Delta contract. Silver then served a low-latency alert path through Webhook/API, a dimensional model centred on fact_market_trades with asset, exchange and time dimensions, and dbt Gold marts for analytics users.
- Developed stateful PySpark Structured Streaming and continuous DataFrame transformations.
- Orchestrated dependencies, scheduling and execution with Databricks Workflows.
- Applied a cost-aware SDLC that separated local validation from cloud deployment.
2 independent exchange streams · 3 Silver consumption paths · 8 business-facing Gold marts
latest_market_pricesarbitrage_opportunitieshourly_ohlcexchange_spread_historymarket_volatilitystale_feed_monitordata_quality_summarymarket_activity
- Project Alpine unified live Kraken and Coinbase trades in a governed Delta Lakehouse, then served the trusted Silver stream to low-latency alerts, dimensional models and eight dbt Gold marts.
- One producer and one raw Kafka topic per exchange isolated schema or source failures and made new exchanges easier to add.
- The shared Silver contract scaled independently into alerting, dimensional modelling and Gold analytics without coupling their workloads.
Vienna Transit Pipeline
Engineering notes+
Public transit data is ingested with Python, scheduled and monitored through Airflow, stored in BigQuery and transformed with dbt into documented, analytics-ready models with visible lineage.
- Apache Airflow DAGs for coordinating dependent batch tasks.
- ELT design that decouples extraction from downstream transformation.
- Modular, version-controlled SQL modelling with dbt.
4 core pipeline stages · 1 orchestration control plane · 0 BI dependencies in the backend
- Vienna Transit Pipeline coordinated REST extraction, BigQuery loading and dbt modelling through Airflow, producing a repeatable backend flow with guarded dependencies and visible lineage.
- Layered transformations made failures easier to isolate and rerun.
- Scheduling, lineage and testing turned an analysis into a repeatable data product.
Cart Abandonment Engine
Engineering notes+
E-commerce click and cart events enter Kafka, load into Snowflake and pass through dbt models that identify abandonment behaviour. Airflow coordinates the warehouse workflow and exposes observable task boundaries.
- Continuous event-streaming architecture beyond scheduled batch jobs.
- Secure machine-to-machine authentication using RSA key pairs.
- Decoupled capture of explicit business events without polling or CDC.
4 decoupled platform layers · 0 source polling loops · 1 actionable abandonment product
- Cart Abandonment Engine replaced polling with Kafka and Snowflake-native ingestion, then modelled live cart events into actionable abandonment signals with dbt.
- Each layer exposed its own freshness, success and latency signals.
- RSA key-pair authentication protected machine-to-machine access between the streaming and warehouse layers.
Data Analytics & BI Projects
Modelled datasets and decision products that turned trusted data into measurable business insight.
Chocolate Sales Analytics
Engineering notes+
Python handles data profiling, cleaning and exploratory analysis across roughly 200K records. A Power BI star schema replaces flat-file reporting, while reusable DAX measures provide time intelligence and executive KPIs.
- Scalable star-schema design for analytical workloads.
- Data-quality rules across interconnected fact and dimension tables.
- Backend transformation architecture for complex business logic.
~200K transaction records modelled · +37% Wholesale average order value · 77% revenue from Australia and Brazil
- Chocolate Sales Analytics transformed roughly 200K transactions into a validated star schema that supported reusable DAX measures and revealed 37% higher Wholesale order value and 77% revenue concentration in Australia and Brazil.
- Discount percentage showed only 0.17 correlation with boxes shipped.
- The validated model kept product, customer, geography and time analysis consistent across the report.
Retail Sales Performance
Engineering notes+
Power Query standardizes three years of multi-location data before a Sales fact connects to Products, Locations, Customers, Salespeople and a dedicated Dates dimension. DAX measures layer revenue, cost, margin, YoY and running totals over the semantic model.
- Six-table star-schema design for multi-location sales analysis.
- Reusable DAX measures for revenue, cost, margin and time intelligence.
- Power Query preparation with a dedicated, filter-aware date dimension.
3 yrs sales history modelled · 6 tables in the star schema · 32.52% reported gross margin
- Retail Sales Performance modelled three years of multi-location activity in a six-table star schema, enabling reusable time intelligence across €25.66M in sales and a reported 32.52% margin.
- Los Angeles County led with €1.80M, more than twice the next county.
- Margins remained near 32 to 33% despite monthly revenue variation.
Cultural Heritage Tourism
Engineering notes+
T-SQL stages and validates the source, builds fact and dimension tables with foreign keys, and serves a Power BI semantic model with statistical DAX measures.
- Validating categorical assumptions before dimensional modelling.
- Using a junk dimension for changing site attributes.
- Building statistical DAX measures beyond simple aggregation.
15K validated source records · 50 heritage sites compared · 3 decision-focused dashboards
- Cultural Heritage Tourism validated 15K records across 50 sites and rebuilt a flawed categorical model into a trustworthy star schema that connected tourism demand, conservation risk and visitor experience.
- Overcrowding strongly reduced visitor satisfaction (r = -0.78).
- Maintenance spend tracked damage severity (r = 0.93), while environmental impact tracked visitor volume (r = 0.95).





