lochan-eda / User Guide / Preprocessing

Preprocessing Philosophy

The package adds a behaviour-analysis layer before preprocessing decisions are applied.

scikit-learn preprocessing

A scikit-learn preprocessing pipeline is a valid and useful approach. The problem is deciding whether the selected transformations are appropriate for the dataset.

from sklearn.compose import ColumnTransformer
from sklearn.impute import SimpleImputer
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import OneHotEncoder, StandardScaler

preprocessor = ColumnTransformer([
    (
        "numeric",
        Pipeline([
            ("imputer", SimpleImputer(strategy="median")),
            ("scaler", StandardScaler())
        ]),
        numeric_columns
    ),
    (
        "categorical",
        Pipeline([
            ("imputer", SimpleImputer(strategy="most_frequent")),
            ("encoder", OneHotEncoder(
                handle_unknown="ignore"
            ))
        ]),
        categorical_columns
    )
])

Behaviour before transformation

Dataset
   │
   ▼
Profile
   │
   ├──────────────┐
   ▼              ▼
Numerical    Categorical
behaviour      behaviour
   │              │
   └───────┬──────┘
           ▼
    Preprocessing
       workflow
           │
           ▼
     Model-ready data

Complementary design

lochan-eda complements preprocessing libraries rather than attempting to replace them.