df.understand()
Stage 4 · Day 22

Automate: profiling

Everything in Stages 1–3, done by hand, generated in one shot — the ideal first thing you run on an unfamiliar dataset.

# pip install ydata-profiling   (was pandas-profiling)
from ydata_profiling import ProfileReport

report = ProfileReport(df)
report.to_file('report.html')
The report — every tab maps to a concept you already learned
Overviewmaps to → Day 19

Rows, columns, % missing cells, duplicate rows, memory, and type counts — plus automatic warnings (high cardinality, high % missing, strong correlation).

Why the manual stages still matter

The report is only useful because you can read it — and you can read it because you learned Stages 1–3 by hand. Best practice: run it on a dataset, write down your observations, repeat across three or four.