This projects contains the architecture of a library that is scalable and functional for a predictive analytics task. First, we created the skeleton of the library, then, we created the first end-to-end prototype that loads the data, preprocess it, creates features, trains a naive model, performs hyperparameter tuning and evaluates the predictions with some metrics.
Since our pipeine is somewhat custom, this module provides a set of custom scikit-learn-compatible transformers designed to prepare vehicle data before it reaches a predictive model. Each transformer is a subclass of BaseEstimator and TransformerMixin.
-
MetadataAvgTransformer
- Purpose: Enriches the dataset with average metadata values (e.g., safety ratings, weights) based on the car's brand.
- Process:
- Uses a global
metadataDataFrame containing aggregated brand information. - Matches each entry’s
Make(brand) to the averaged columns (OVERALL_STARS,CURB_WEIGHT,MIN_GROSS_WEIGHT) and updates the DataFrame accordingly.
- Uses a global
-
CountryAssigner
- Purpose: Determines the country of origin for each car brand.
- Process:
- Utilizes
CountryMapperto mapMaketo a corresponding country. - Adds a new column
Countryto the dataset.
- Utilizes
-
AgeAssigner
- Purpose: Computes the vehicle’s age from its manufacturing year.
- Process:
- Calculates
Age = current_year - Year. - Drops the
Yearcolumn.
- Calculates
-
OutlierRemover
- Purpose: Filters out outliers from a specified numerical column to improve model robustness.
- Process:
- Calculates Interquartile Range (IQR) and determines upper and lower bounds.
- Retains only the rows whose values fall within these bounds.
from sklearn.pipeline import Pipeline
from sklearn.compose import ColumnTransformer
from sklearn.preprocessing import OneHotEncoder
from sklearn.tree import DecisionTreeRegressor
numeric_features = ["Mileage", "OVERALL_STARS", "CURB_WEIGHT", "MIN_GROSS_WEIGHT", "Age"]
categorical_features = ["Condition", "Country", "Make"]
preprocessing = ColumnTransformer([
("ohe", OneHotEncoder(handle_unknown='ignore'), categorical_features),
("num", "passthrough", numeric_features)
])
pipeline = Pipeline([
("metadata_avg", MetadataAvgTransformer(columns=["OVERALL_STARS", "CURB_WEIGHT", "MIN_GROSS_WEIGHT"])),
("country", CountryAssigner()),
("age", AgeAssigner()),
("outlier_remover", OutlierRemover(column="Mileage", threshold=3)),
("preprocessing", preprocessing),
("regressor", DecisionTreeRegressor(random_state=42))
])
pipeline.fit(X_train, y_train)
predictions = pipeline.predict(X_test)The DataPlotter class provides a set of visualization methods for exploratory data analysis of both categorical and numerical features. It allows you to easily plot bar charts, distributions of numerical columns, and distributions of categorical columns. It comes with default or custom configurations.
-
plot_barchart(data: pd.Series, title: str, xlabel: str, ylabel: str, figsize: tuple = None)- Description:
Creates a bar chart of a givenpd.Serieswith customizable title and axis labels. - Usage Example:
plotter = DataPlotter() counts = df['Make'].value_counts() plotter.plot_barchart(counts, title="Car Make Distribution", xlabel="Make", ylabel="Count")
- Description:
-
plot_numerical_distribution(data: pd.DataFrame, columns: list[str])- Description:
Plots the distribution (viaseaborn.kdeplot) of specified numerical columns. - Usage Example:
plotter = DataPlotter() numeric_cols = ['Mileage', 'Price'] plotter.plot_numerical_distribution(df, numeric_cols)
- Description:
-
plot_all_categorical(data: pd.DataFrame)- Description:
Detects and plots the distribution of all categorical columns in theDataFrame. Each categorical column is represented by a bar chart of its value counts. - Usage Example:
plotter = DataPlotter() plotter.plot_all_categorical(df)
- Description:
-
Default Configurations:
TheDataPlotterinitializes with a set of default configurations:figsize: (12, 8)color: "skyblue"title_fontsize: 16label_fontsize: 14tick_labelsize: 12
You can override these defaults by passing
figsizein the plotting methods or modifying the class to accept more parameters.
import pandas as pd
df = pd.DataFrame({
'Make': ['Ford', 'Toyota', 'Ford', 'Chevrolet', 'Toyota', 'Ford'],
'Condition': ['Good', 'Excellent', 'Fair', 'Excellent', 'Good', 'Good'],
'Mileage': [20000, 30000, 15000, 40000, 25000, 22000],
'Price': [18000, 22000, 16000, 25000, 21000, 19000]
})
plotter = DataPlotter()
plotter.plot_all_categorical(df)
plotter.plot_numerical_distribution(df, ['Mileage', 'Price'])
condition_counts = df['Condition'].value_counts()
plotter.plot_barchart(condition_counts, title="Condition Distribution", xlabel="Condition", ylabel="Count")The CountryMapper class provides a way to determine the country of origin for various car brands. Given a list of brands and a URL containing a reference table, it scrapes the webpage and constructs a mapping from brand names to countries.
-
Isolation:
Tests focus on one piece of functionality at a time. Each transformer is tested independently to confirm it performs its role correctly. -
Repeatability:
The tests can be run multiple times with the same results, ensuring deterministic outputs. Mocking ensures no external dependencies affect results. -
Readability and Maintainability: By using descriptive test function names and in-code comments, it’s easier for other developers to understand what’s being tested and why.
In test_metadata_avg_transformer, we:
- Monkeypatch the global
metadatavariable so the transformer uses thesample_metadatafixture. - Instantiate
MetadataAvgTransformerand fit-transform the sample data. - Check that new columns (
OVERALL_STARS,CURB_WEIGHT,MIN_GROSS_WEIGHT) appear. - Verify that the values for
Fordmatch those specified in the mock metadata.
Similarly, in test_country_assigner, we:
- Mock
CountryMappermethods to return predefined countries for each brand. - Fit-transform using
CountryAssigner. - Assert that the
Countrycolumn is created and matches expected values.
Lastly, test_age_assigner:
- Checks that
Ageis computed andYearis removed. - Ensures that the computed
Ageequalscurrent_year - Year.
Assuming you have pytest installed, you can run:
cd tests
pytest test_preprocessor.py -
Model Loading:
On startup, the application loads a pre-trained pipeline frombest_model.pkl. -
Data Validation with Pydantic:
TheCarFeaturesPydantic model defines the expected input schema for each request. -
Predict Endpoint (
/predict):
Sends a JSON object containing a single car's features. The API converts this to apandas.DataFrameand passes it through the pipeline to generate a price prediction. The response is a JSON object with a singlepredictionkey.Request Example:
{ "Make": "Ford", "Model": "F-150", "Year": 2020, "Mileage": 30000, "Condition": "Excellent" }