Lochan EDA
Preprocessing that starts with the data.
A reusable Python toolkit for understanding tabular data and turning its behaviour into practical preprocessing decisions before machine learning.
lochan-eda is designed around a simple idea: preprocessing should respond to the behaviour of the dataset instead of beginning with a fixed list of transformations.
The Problem
A typical tabular ML workflow quickly becomes a collection of repeated decisions: inspect missing values, separate numerical and categorical features, investigate outliers, understand distributions, choose transformations, and finally construct a preprocessing pipeline.
The code is repetitive. The difficult part is the reasoning behind the code.
lochan-eda is built around a simple idea: preprocessing should respond to the behaviour of the dataset, not begin with a fixed list of transformers.
Manual EDA
Every dataset becomes another notebook.
Without reusable tooling, the same investigation gets repeated for almost every dataset.
First inspect missing values. Then identify numerical and categorical columns. Then investigate distributions, outliers, cardinality, rare categories, and finally decide what transformations make sense.
The next dataset arrives and the process starts again.
# inspect missing values
df.isnull().sum()
# inspect data types
df.dtypes
# numerical features
numeric_cols = df.select_dtypes(
include="number"
).columns
# categorical features
categorical_cols = df.select_dtypes(
exclude="number"
).columns
# inspect distributions
df[numeric_cols].describe()
# inspect categories
for col in categorical_cols:
print(col, df[col].nunique())
# investigate outliers
# decide imputation
# decide scaling
# decide encoding
# build pipeline
# repeat for another dataset...Blind Pipeline
A pipeline can be reusable and still be thoughtless.
preprocessor = ColumnTransformer([
(
"numeric",
Pipeline([
("imputer", SimpleImputer(
strategy="median"
)),
("scaler", StandardScaler())
]),
numeric_cols
),
(
"categorical",
Pipeline([
("imputer", SimpleImputer(
strategy="most_frequent"
)),
("encoder", OneHotEncoder(
handle_unknown="ignore"
))
]),
categorical_cols
)
])This pipeline is perfectly valid. It is also easy to write without asking whether each transformation fits the data.
Scikit-learn provides excellent preprocessing building blocks. The developer still has to determine which blocks make sense for the dataset.
Run the Workflow
Don't just read the workflow. Run it.
A complete notebook is included so the workflow can be inspected, executed, and modified with a real dataset.
from lochan_eda import (
Profiler,
AutomatedEDA
)
profile = Profiler(df)
profile.report.save("report.pdf")eda = AutomatedEDA()
X_train, X_test, y_train, y_test = eda.prepare(
df,
target="target",
exclude=None,
split=True,
test_size=0.2,
random_state=42,
stratify=df["target"],
in_return="ndarray"/"dataframe"/"tensor"
)Dataset inspected and transformed into model-ready training and testing data.
Data-First Preprocessing
Let the data influence the preprocessing.
Instead of starting with a fixed recipe, lochan-eda starts by understanding the dataset.
Numerical and categorical features are analysed according to their behaviour. Missingness, distributions, categories, outliers, and feature characteristics can then inform the preprocessing workflow.
The goal is not to hide preprocessing behind magic. It is to reduce repetitive investigation while keeping the reasoning visible and the workflow reusable.
Workflow
From behaviour to model-ready data.
The package separates inspection from reusable transformation.
pip install lochan-edafrom lochan_eda import Profiler
profile = Profiler(df)
profile.overview()from lochan_eda import AutomatedEDA
eda = AutomatedEDA()
Xtr, Xte, ytr, yte = eda.prepare(
df,
target="target"
)Feature Behaviour
Different features can require different treatment.
Numerical
Missing values, scale, distributions, and outliers can affect how numerical features should be prepared.
Categorical
Missing categories, cardinality, and rare values can influence how categorical features should be represented.
Dataset-level
Understanding the dataset before transformation gives the preprocessing workflow context instead of blindly applying the same recipe everywhere.
Fit and Transform
Learn on train. Reuse on test.
The important part of AutomatedEDA is not simply applying preprocessing. The workflow separates fit() from transform().
This allows preprocessing decisions to be learned from the training data and then reused when transforming another dataset.
eda.fit(X_train)
X_train = eda.transform(X_train)
X_test = eda.transform(X_test)prepare() wraps the common workflow when you want one entry point for target handling, train/test splitting, fitting, and transformation.
API Surface
Small public surface. Focused jobs.
AutomatedEDA
Connects inspection, preprocessing decisions, train/test splitting, fitting, and transformation into one reusable workflow.
prepare()fit()transform()Profiler
A dataset-level entry point for understanding structure, distributions, missingness, and feature behaviour.
overview()numericalcategoricalNumerical
Works with numerical feature behaviour including imputation, outliers, scaling, summaries, and visual analysis.
imputer()outlier_manager()scaler()summary()plot()Categorical
Handles categorical feature behaviour including missing values, rare categories, encoding, summaries, and plots.
imputer()rare_manager()encoder()summary()plot()Missing
Provides focused visualization for understanding missing-value patterns before preprocessing.
plot()Report
Turns a profiling session into a shareable PDF report for documenting dataset behaviour.
save()Why I Built It
From transformer-first to data-first preprocessing.
Inspect
Understand missing values, distributions, categories, outliers, and feature behaviour before choosing transformations.
Decide
Use what the dataset reveals to guide preprocessing rather than blindly applying the same recipe to every dataset.
Reuse
Turn those decisions into a consistent workflow that can be fitted on training data and reused on new data.