9-step roadmap
Data Engineer
You build the pipelines that move, clean, and store data, so analysts and scientists always have data they can trust. It's different from the Data Analyst path: they study the data, you make sure it arrives clean and on time.
Data Engineer roadmap
Follow the line, from first step to goal.
SQL, and know it deeply
StartSQL is the language for asking a database questions. For this job it's not optional — you'll live in it, so go past the basics and learn to write, read, and speed up serious queries.
A language: Python
Python is the glue of data work. You'll use it to move data around, clean it up, and connect the tools in your pipeline together.
How data flows: batch vs streaming
Learn the two main ways data moves. Batch means processing a big chunk on a schedule, like every night. Streaming means handling each event the moment it happens. Knowing when to use which shapes everything you build.
Build ETL/ELT pipelines
A pipeline takes data from one place, reshapes it, and drops it somewhere useful. ETL and ELT are just two orderings of those steps — extract, transform, load. This is the core of the job.
Data warehouses & lakes
A warehouse stores neat, organised data ready for analysis (like Snowflake or BigQuery). A lake stores raw data of every shape for later. Learn what each is for, and when to reach for which.
Orchestration: scheduling jobs
Real pipelines have many steps that must run in the right order, at the right time. A tool like Airflow schedules them, runs them, and tells you when a step fails, instead of you babysitting it by hand.
Data quality & reliability
Bad data quietly ruins every decision made from it. Learn to check data as it flows — catching missing values, duplicates, and sudden weird numbers — so problems get caught before anyone builds a report on them.
Cloud data tools
Most data work today runs on cloud platforms like AWS, Google Cloud, or Azure. Learn the data services one provider offers, since that's where you'll build and run your pipelines.
Scaling to big data
GoalWhen data gets too big for one machine, you spread the work across many. Learn the ideas behind tools like Spark, which crunch huge datasets by splitting the job up. This is the deep end — save it for last.