df.understand()
Stage 2 · Day 20

Univariate

One column, in isolation. The only question is its type — and that alone picks the plot.

CATEGORICAL → how often each category appears
Categorical column — passenger class
Class 3491Class 1216Class 2184

Same data, two lenses: counts (raw frequency per category) vs a pie (each category's share of the whole).

NUMERICAL → the shape of the distribution
Histogram — age distribution
bins: 12
047930–7: 57–13: 3713–20: 5620–27: 7627–33: 9333–40: 7240–47: 7647–53: 4653–60: 2560–67: 1167–73: 373–80: 0age →

Drag the slider. Too few bins hide the shape; too many make it spiky and noisy. The bin count is the one knob you tune.

Skewness — what the distplot reveals
df.skew()≈ 0
Symmetric (normal)

Values pile in the middle; both tails are even. The classic bell.

Boxplot — the five-number summary & outliers
Q3 + 1.5·IQROutlier — beyond the whiskerQ1medianQ3minmax
Drag the point's valuevalue 96 · OUTLIER

Anything past Q3 + 1.5·IQR is flagged an outlier (the same logic guards the low end with Q1 − 1.5·IQR). Push the point right and watch it turn coral.