Skip to content
All paths

9-step roadmap

Data Engineer

You build the pipelines that move, clean, and store data, so analysts and scientists always have data they can trust. It's different from the Data Analyst path: they study the data, you make sure it arrives clean and on time.

Data Engineer roadmap

Follow the line, from first step to goal.

  1. SQL, and know it deeply

    Start

    SQL is the language for asking a database questions. For this job it's not optional — you'll live in it, so go past the basics and learn to write, read, and speed up serious queries.

  2. A language: Python

    Python is the glue of data work. You'll use it to move data around, clean it up, and connect the tools in your pipeline together.

  3. How data flows: batch vs streaming

    Learn the two main ways data moves. Batch means processing a big chunk on a schedule, like every night. Streaming means handling each event the moment it happens. Knowing when to use which shapes everything you build.

  4. Build ETL/ELT pipelines

    A pipeline takes data from one place, reshapes it, and drops it somewhere useful. ETL and ELT are just two orderings of those steps — extract, transform, load. This is the core of the job.

  5. Data warehouses & lakes

    A warehouse stores neat, organised data ready for analysis (like Snowflake or BigQuery). A lake stores raw data of every shape for later. Learn what each is for, and when to reach for which.

  6. Orchestration: scheduling jobs

    Real pipelines have many steps that must run in the right order, at the right time. A tool like Airflow schedules them, runs them, and tells you when a step fails, instead of you babysitting it by hand.

  7. Data quality & reliability

    Bad data quietly ruins every decision made from it. Learn to check data as it flows — catching missing values, duplicates, and sudden weird numbers — so problems get caught before anyone builds a report on them.

  8. Cloud data tools

    Most data work today runs on cloud platforms like AWS, Google Cloud, or Azure. Learn the data services one provider offers, since that's where you'll build and run your pipelines.

  9. Scaling to big data

    Goal

    When data gets too big for one machine, you spread the work across many. Learn the ideas behind tools like Spark, which crunch huge datasets by splitting the job up. This is the deep end — save it for last.