From framing questions and defining metrics, through SQL aggregation, joins, and window functions, pandas preprocessing, and matplotlib visualization, all the way to cohorts, funnels, RFM, and A/B testing. Across 30 lessons you'll go from looking at data to actually making decisions with it. Every result shown comes from real output run against a fixed dataset from a fictional e-commerce shop.
The 30 lessons are split into 6 chapters. We recommend working through them in order starting from Chapter 1, but feel free to skim just the parts you're curious about. * Lessons that use SQL, pandas, or matplotlib can't run in the browser. Try them in the Python 3.14 virtual environment you set up in lesson 1 (SQL uses the standard library's sqlite3). Small calculations that only need the standard library can still run right there.
Before writing a query, break your question down into a metric, a target, and a time period, define that metric in code, and build a map of which tables hold how many rows. We finish by covering pitfalls that can throw off your conclusions even with a correct query, like cancelled orders sneaking into totals and the trap of averaging averages.
Use aggregate functions and GROUP BY to see the whole picture and its breakdown, then shape the data you need with joins, subqueries, window functions, and date rounding. We finish by combining revenue, order count, AOV, and the top category into a single summary.
Load the data you pulled with SQL into pandas and shape it into an analyzable form by handling missing values, outliers, merges, and pivots. We finish by wrapping the whole process into a function so preprocessing gives the same result every time you run it.
Pick the right chart for your question and draw line, bar, and histogram charts with matplotlib, then combine two of them into a dashboard. We finish by looking at misleading chart tricks — like truncating the y-axis to exaggerate a difference — and how to avoid them.
Try out the analysis patterns you'll use again and again on the job — cohorts, funnels, RFM, time-series trends, and A/B testing — all on the same e-commerce data. We finish by using a statistical test to tell whether the difference between two options is real or just chance, so you don't adopt something just because the number went up.
Turn your analysis into a report structured as conclusion → evidence → next step, going beyond a mere observation to a recommendation that drives a decision. We finish by combining everything from extraction to write-up into a single pipeline that can run automatically on a schedule.
Once you finish all 30 lessons, move on to Intro to Python & Machine Learning, where you'll build predictive models on the same pandas foundation. All courses unlock with a membership.