Stage 2 · Day 20
Univariate
One column, in isolation. The only question is its type — and that alone picks the plot.
CATEGORICAL → how often each category appears
Categorical column — passenger class
Same data, two lenses: counts (raw frequency per category) vs a pie (each category's share of the whole).
NUMERICAL → the shape of the distribution
Histogram — age distribution
bins: 12
Drag the slider. Too few bins hide the shape; too many make it spiky and noisy. The bin count is the one knob you tune.
Skewness — what the distplot reveals
df.skew()≈ 0
Symmetric (normal)
Values pile in the middle; both tails are even. The classic bell.
Boxplot — the five-number summary & outliers
Drag the point's valuevalue 96 · OUTLIER
Anything past Q3 + 1.5·IQR is flagged an outlier (the same logic guards the low end with Q1 − 1.5·IQR). Push the point right and watch it turn coral.