Stage 4 · Day 22
Automate: profiling
Everything in Stages 1–3, done by hand, generated in one shot — the ideal first thing you run on an unfamiliar dataset.
# pip install ydata-profiling (was pandas-profiling)
from ydata_profiling import ProfileReport
report = ProfileReport(df)
report.to_file('report.html')The report — every tab maps to a concept you already learned
Overviewmaps to → Day 19
Rows, columns, % missing cells, duplicate rows, memory, and type counts — plus automatic warnings (high cardinality, high % missing, strong correlation).
Why the manual stages still matter
The report is only useful because you can read it — and you can read it because you learned Stages 1–3 by hand. Best practice: run it on a dataset, write down your observations, repeat across three or four.