Skip to content
#

preprocessing-data

Here are 280 public repositories matching this topic...

Lightning-fast data preprocessing and feature engineering for Python, built on Polars. 108 sklearn-style transformers for imputation, encoding, scaling, outlier clipping, discretization and feature selection, with Pipeline support and ONNX export for low-latency inference.

  • Updated Sep 22, 2026
  • Python

Simple and automatic data cleaning in one line of code! It performs one-hot encoding, date & time casting to datetime dtype, detects binary columns, safely convert non-numeric columns to numeric dtypes, cleaning dirty/empty values, normalizing values and removing unwanted columns all in one line of code. Get your data ready for model training an…

  • Updated May 22, 2021
  • Python

A pure-Rust workspace for classical ML: scikit-learn-style preprocessing & models (datarust) plus one-call data profiling & quality reports (datarust-profile). Zero dependencies by default.

  • Updated Sep 14, 2026
  • Rust

This project develops an activity recognition model for a mobile fitness app using statistical analysis and machine learning. By processing smartphone sensor data, it extracts features to train models that accurately recognize user activities.

  • Updated Aug 6, 2024
  • Jupyter Notebook

Add this topic to your repo

To associate your repository with the preprocessing-data topic, visit your repo's landing page and select "manage topics."

Learn more