03 / Developer Tool · Python Package

Lochan EDA
Preprocessing that starts with the data.

A reusable Python toolkit for understanding tabular data and turning its behaviour into practical preprocessing decisions before machine learning.

Package snapshot
VERSION
v2
LANGUAGE
Python
FOCUS
Tabular EDA + preprocessing
INTERFACE
Profiler / AutomatedEDA
Abstract

lochan-eda is designed around a simple idea: preprocessing should respond to the behaviour of the dataset instead of beginning with a fixed list of transformations.

01

The Problem

A typical tabular ML workflow quickly becomes a collection of repeated decisions: inspect missing values, separate numerical and categorical features, investigate outliers, understand distributions, choose transformations, and finally construct a preprocessing pipeline.

The code is repetitive. The difficult part is the reasoning behind the code.

lochan-eda is built around a simple idea: preprocessing should respond to the behaviour of the dataset, not begin with a fixed list of transformers.

02

Manual EDA

Every dataset becomes another notebook.

Without reusable tooling, the same investigation gets repeated for almost every dataset.

First inspect missing values. Then identify numerical and categorical columns. Then investigate distributions, outliers, cardinality, rare categories, and finally decide what transformations make sense.

The next dataset arrives and the process starts again.

REPETITIVE EDA
# inspect missing values
df.isnull().sum()

# inspect data types
df.dtypes

# numerical features
numeric_cols = df.select_dtypes(
    include="number"
).columns

# categorical features
categorical_cols = df.select_dtypes(
    exclude="number"
).columns

# inspect distributions
df[numeric_cols].describe()

# inspect categories
for col in categorical_cols:
    print(col, df[col].nunique())

# investigate outliers
# decide imputation
# decide scaling
# decide encoding
# build pipeline
# repeat for another dataset...
03

Blind Pipeline

A pipeline can be reusable and still be thoughtless.

GENERIC PREPROCESSING
preprocessor = ColumnTransformer([
    (
        "numeric",
        Pipeline([
            ("imputer", SimpleImputer(
                strategy="median"
            )),
            ("scaler", StandardScaler())
        ]),
        numeric_cols
    ),

    (
        "categorical",
        Pipeline([
            ("imputer", SimpleImputer(
                strategy="most_frequent"
            )),
            ("encoder", OneHotEncoder(
                handle_unknown="ignore"
            ))
        ]),
        categorical_cols
    )
])

This pipeline is perfectly valid. It is also easy to write without asking whether each transformation fits the data.

?Why median imputation?
?Why standard scaling?
?Why most-frequent categorical imputation?
?Why this encoding strategy?

Scikit-learn provides excellent preprocessing building blocks. The developer still has to determine which blocks make sense for the dataset.

04

Run the Workflow

Don't just read the workflow. Run it.

A complete notebook is included so the workflow can be inspected, executed, and modified with a real dataset.

01 · Load data02 · Profile03 · Inspect behaviour04 · Prepare05 · Transform06 · Model
Open notebook ↗
lochan_eda_example.ipynbJUPYTER NOTEBOOK
In [1]:
from lochan_eda import (
    Profiler,
    AutomatedEDA
)

profile = Profiler(df)
profile.report.save("report.pdf")
In [2]:
eda = AutomatedEDA()

X_train, X_test, y_train, y_test = eda.prepare(
    df,
    target="target",
    exclude=None,
    split=True,
    test_size=0.2,
    random_state=42,
    stratify=df["target"],
    in_return="ndarray"/"dataframe"/"tensor"
)
Output

Dataset inspected and transformed into model-ready training and testing data.

05

Data-First Preprocessing

Let the data influence the preprocessing.

Instead of starting with a fixed recipe, lochan-eda starts by understanding the dataset.

Numerical and categorical features are analysed according to their behaviour. Missingness, distributions, categories, outliers, and feature characteristics can then inform the preprocessing workflow.

The goal is not to hide preprocessing behind magic. It is to reduce repetitive investigation while keeping the reasoning visible and the workflow reusable.

01
DataFrame
Raw data
02
Behaviour
Understand
03
Decision
Choose
04
Pipeline
Transform
06

Workflow

From behaviour to model-ready data.

The package separates inspection from reusable transformation.

01 / INSTALL
pip install lochan-eda
02 / INSPECT
from lochan_eda import Profiler

profile = Profiler(df)

profile.overview()
03 / PREPARE
from lochan_eda import AutomatedEDA

eda = AutomatedEDA()

Xtr, Xte, ytr, yte = eda.prepare(
    df,
    target="target"
)
07

Feature Behaviour

Different features can require different treatment.

01

Numerical

Missing values, scale, distributions, and outliers can affect how numerical features should be prepared.

02

Categorical

Missing categories, cardinality, and rare values can influence how categorical features should be represented.

03

Dataset-level

Understanding the dataset before transformation gives the preprocessing workflow context instead of blindly applying the same recipe everywhere.

08

Fit and Transform

Learn on train. Reuse on test.

The important part of AutomatedEDA is not simply applying preprocessing. The workflow separates fit() from transform().

This allows preprocessing decisions to be learned from the training data and then reused when transforming another dataset.

FIT → TRANSFORM
eda.fit(X_train)

X_train = eda.transform(X_train)
X_test  = eda.transform(X_test)

prepare() wraps the common workflow when you want one entry point for target handling, train/test splitting, fitting, and transformation.

09

API Surface

Small public surface. Focused jobs.

01ORCHESTRATE

AutomatedEDA

Connects inspection, preprocessing decisions, train/test splitting, fitting, and transformation into one reusable workflow.

prepare()fit()transform()
02INSPECT

Profiler

A dataset-level entry point for understanding structure, distributions, missingness, and feature behaviour.

overview()numericalcategorical
03NUMERIC

Numerical

Works with numerical feature behaviour including imputation, outliers, scaling, summaries, and visual analysis.

imputer()outlier_manager()scaler()summary()plot()
04CATEGORICAL

Categorical

Handles categorical feature behaviour including missing values, rare categories, encoding, summaries, and plots.

imputer()rare_manager()encoder()summary()plot()
05QUALITY

Missing

Provides focused visualization for understanding missing-value patterns before preprocessing.

plot()
06OUTPUT

Report

Turns a profiling session into a shareable PDF report for documenting dataset behaviour.

save()
10

Why I Built It

From transformer-first to data-first preprocessing.

01

Inspect

Understand missing values, distributions, categories, outliers, and feature behaviour before choosing transformations.

02

Decide

Use what the dataset reveals to guide preprocessing rather than blindly applying the same recipe to every dataset.

03

Reuse

Turn those decisions into a consistent workflow that can be fitted on training data and reused on new data.

Package inventory
01AutomatedEDAConnects inspection, preprocessing decisions, train/test splitting, fitting, and transformation into one reusable workflow.
02ProfilerA dataset-level entry point for understanding structure, distributions, missingness, and feature behaviour.
03NumericalWorks with numerical feature behaviour including imputation, outliers, scaling, summaries, and visual analysis.
04CategoricalHandles categorical feature behaviour including missing values, rare categories, encoding, summaries, and plots.
05MissingProvides focused visualization for understanding missing-value patterns before preprocessing.
06ReportTurns a profiling session into a shareable PDF report for documenting dataset behaviour.