Next cohort coming. Enrolment is not open yet: join the waiting list, you will hear first and the launch price will be held for you. No payment now.
🗃 Intensive bootcamp - 14 weeks

Data Engineering Bootcamp

From raw data to production pipelines

Learn to design, build, automate and deploy modern data pipelines. Heavy on practice, light on theory: a hands-on lab in every module, on data close to real situations.

15
Modules
14
Weeks
10
Portfolio projects
1
Data platform
Who is it for?

For people who want to build

Some programming helps, but the path builds the skills step by step. Python or SQL is an advantage, not a requirement.

🌱
Motivated beginners
📊
Data Analysts
🤖
Data Scientists
💻
Developers
🔄
Career changers
🎓
Students
Programme

15 modules, 15 hands-on labs

Every module ends with a lab. You learn by building, not by listening.

1
Foundations
01

Data Engineering fundamentals

The Data Engineer's role, data life cycle, batch and streaming, ETL and ELT, Data Lake, Warehouse and Lakehouse.

Lab: Design a company's data architecture

02

SQL for Data Engineers

Joins, subqueries, CTEs, window functions, advanced aggregations, indexes, query tuning and relational modelling.

Lab: Analyse and transform a database of several million rows

03

Python for Data Engineering

Data structures, CSV, JSON and Parquet, Pandas, error handling, logging and processing large files.

Lab: Build a data processing pipeline in Python

2
Sources & storage
04

Databases

PostgreSQL, MySQL, MongoDB, schemas, keys, indexes, transactions, Python connectivity, security and access.

Lab: Build the database behind a business application

05

APIs & data collection

REST, authentication, pagination, rate limits, webhooks, automated ingestion and responsible web scraping.

Lab: Automatically collect data from several APIs

3
Pipelines & modelling
06

ETL & ELT

Extract, transform, load, data quality and validation, duplicates, missing values, incremental pipelines and idempotence.

Lab: A full ETL pipeline: API → Python → PostgreSQL

07

Data Warehouse & analytical modelling

OLTP and OLAP, dimensional modelling, fact and dimension tables, Star and Snowflake schemas, SCD, dbt.

Lab: Build a Data Warehouse for an e-commerce business

4
Scaling up
08

Big Data with Apache Spark

Spark architecture, PySpark, DataFrames, lazy evaluation, partitions, joins, window functions, Parquet and tuning.

Lab: Process several million rows with PySpark

09

Orchestration with Apache Airflow

DAGs, tasks, operators, scheduling, dependencies, XCom, retries, monitoring, backfill and good practice.

Lab: Orchestrate a data pipeline with Airflow

5
Cloud & Lakehouse
10

Data Engineering in the Cloud

AWS, Azure and Google Cloud, object storage, data lakes, cloud warehouses, IAM, serverless, BigQuery and cost control.

Lab: Build a data pipeline in the Cloud

11

Lakehouse & Databricks

Delta Lake, Delta tables, ACID transactions, schema enforcement and evolution, time travel, Medallion architecture.

Lab: Build a complete Lakehouse architecture

6
Industrialisation & real time
12

Docker, Git & CI/CD

Branches and pull requests, versioning, Dockerfile, images, containers, Docker Compose, secrets, tests and GitHub Actions.

Lab: Dockerise and automate a pipeline deployment

13

Data Quality, observability & monitoring

Data tests, schema validation, freshness, completeness, uniqueness, logging, alerting, lineage, SLAs and SLOs.

Lab: Automatically detect anomalies in a pipeline

14

Streaming & real-time data

Event-driven architecture, Apache Kafka, producers and consumers, topics, partitions, consumer groups, real-time processing.

Lab: Build a small real-time data pipeline

15

Data Engineering for AI

Feature pipelines, training pipelines, document ingestion, vector databases, RAG pipelines and embeddings.

Lab: Build the pipeline feeding an AI application

🔒

Rest of the programme locked

The detail of the remaining modules, the tools taught and each step's lab are available after enrolment.

Final project

An end-to-end data platform

Twelve steps, from collection to consumption, combining the technologies covered during the bootcamp.

✓ 1. Collection — APIs, files and databases
✓ 2. Ingestion — Python and automated pipelines
✓ 3. Data Lake — raw data storage
✓ 4. Transformation — Python, SQL and PySpark
✓ 5. Lakehouse — Bronze → Silver → Gold
✓ 6. Data Warehouse — analytical tables
✓ 7. Orchestration — Apache Airflow
✓ 8. Cloud — deploying the architecture
✓ 9. Data Quality — automated tests and validation
✓ 10. Monitoring — logs, alerts and supervision
✓ 11. BI — connection to a BI tool
✓ 12. Industrialisation — Git, Docker and CI/CD
🔒

The twelve steps of the final project

The full walkthrough, the practice datasets and the review are provided to enrolled participants.

Portfolio

Ten projects to show

By the end, every participant leaves with concrete work to present.

Python ETL pipeline
API ingestion pipeline
Data Warehouse
PySpark pipeline
Airflow-orchestrated pipeline
Lakehouse architecture
Cloud pipeline
Streaming pipeline
Data pipeline for an AI application
A complete end-to-end data platform
Technologies

The tools of the trade

The technologies actually used in companies, not teaching toys.

Languages

Python · SQL · Bash

Data

PostgreSQL · Pandas · PySpark · Parquet · dbt

Big Data

Apache Spark · Databricks · Delta Lake

Orchestration

Apache Airflow

Streaming

Apache Kafka

Cloud

AWS · Azure · Google Cloud

DevOps

Git · GitHub · Docker · CI/CD

Architecture

ETL · ELT · Data Lake · Warehouse · Lakehouse · Medallion

Outcome

What you will be able to do

Design a data architecture
Build ETL/ELT pipelines
Collect data from multiple sources
Work with SQL and Python
Process large volumes with Spark
Build a Data Warehouse and a Lakehouse
Orchestrate pipelines with Airflow
Use cloud technologies
Set up real-time pipelines
Monitor and secure pipelines
Industrialise with Git, Docker and CI/CD
Build a data platform from collection to consumption
Format

Your typical week

14 weeks, one module a week. Around 14 to 18 hours weekly, two thirds of it hands-on.

📚
Class & demos
2h · Live
🛠
Hands-on workshop
3h30 · Live
❓
Code review & Q&A
1h · Live · optional
💼
Project & practice
8 to 12h · Self-directed
Plans

Choose your plan

Four levels of support to match your pace and your goals.

Based in West or Central Africa?
Essential

Full access

One-off payment · Lifetime access
  • ✓ All 15 modules, delivered live
  • ✓ Real datasets
  • ✓ Exercises & labs
  • ✓ Final project reviewed
  • ✓ DataSAI Data Engineer certification
  • - Exercises and labs reviewed
  • - Your pipelines reviewed
  • - Working environment provided
  • - Discord community
Professional

Guided

One-off payment · Lifetime access
  • ✓ Everything in Essential
  • ✓ Exercises and labs reviewed
  • ✓ Your pipelines reviewed
  • ✓ Working environment provided
  • ✓ Discord community
  • - Group mentoring
  • - Pipeline portfolio
  • - End-to-end data platform
  • - Data Engineer interview preparation
Executive

Premium

One-off payment · Lifetime access
  • ✓ Everything in Career
  • ✓ One-on-one mentoring
  • ✓ Personalised architecture review
  • ✓ Audit of your data platform
  • ✓ Three months of career support
  • ✓ Partner company network
  • ✓ Priority access to future courses

Launch prices reserved for the waiting list. They will be guaranteed to pre-registered members when the cohort opens, with no payment before.

💬
Is the price a barrier? Message us on WhatsApp at +1 (904) 243-7876. Instalments, pricing adapted to your country, or a scholarship: we look at your situation and find a way. No motivated learner is turned away over money.

Ready to build pipelines?

Join the next cohort and build a complete data platform.