awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
scikit-learn-contrib avatar

scikit-learn-contrib/sklearn-pandas

0
View on GitHub↗
2,850 stars·420 forks·Python·17 views

Sklearn Pandas

Pandas integration with sklearn

Features

  • Data Manipulation - Pandas integration for scikit-learn pipelines.
  • Data Manipulation Libraries - Bridge between Pandas and Scikit-learn.
  • Data Processing Libraries - Bridge between scikit-learn transformers and pandas dataframes.

Star history

Star history chart for scikit-learn-contrib/sklearn-pandasStar history chart for scikit-learn-contrib/sklearn-pandas

How this analysis was created: This summary and feature list were written by an AI model that read the project's README and public documentation pages. Each feature links to the documentation it came from; stars, license and language come straight from the GitHub API. The model does not read the source code, and the analysis is refreshed when the project is re-analysed. Learn more on our About page.

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Start searching with AI

Frequently asked questions

What does scikit-learn-contrib/sklearn-pandas do?

Pandas integration with sklearn

What are the main features of scikit-learn-contrib/sklearn-pandas?

The main features of scikit-learn-contrib/sklearn-pandas are: Data Manipulation, Data Manipulation Libraries, Data Processing Libraries.

What are some open-source alternatives to scikit-learn-contrib/sklearn-pandas?

Open-source alternatives to scikit-learn-contrib/sklearn-pandas include: modin-project/modin — Modin is a distributed dataframe library and parallel data processing engine designed to handle large datasets that… pola-rs/polars — Polars is a high-performance columnar data processing library designed for efficient analytical workflows. It… vaexio/vaex — Vaex is a high-performance Apache Arrow DataFrame library and out-of-core data processing engine designed to handle… jmcarpenter2/swifter — A package which efficiently applies any function to a pandas dataframe or series in the fastest available manner. pydata/xarray — Xarray is a Python multidimensional array library and labeled dataset framework. It extends the NumPy data structure… iamseancheney/python_for_data_analysis_2nd_chinese_version — This project is an educational resource and a collection of instructional materials for performing data manipulation…

Open-source alternatives to Sklearn Pandas

Similar open-source projects, ranked by how many features they share with Sklearn Pandas.
  • vaexio/vaexvaexio avatar

    vaexio/vaex

    8,506View on GitHub↗

    Vaex is a high-performance Apache Arrow DataFrame library and out-of-core data processing engine designed to handle billion-row tabular datasets in Python. It functions as a lazy evaluation framework that defers computations and transformations until results are required, enabling the processing of datasets that exceed available system RAM by mapping files directly from disk. The project distinguishes itself as a tool for big data visualization and exploration, specifically integrated for use within interactive notebooks. It provides specialized capabilities for machine learning feature engin

    Python
    View on GitHub↗8,506
  • pola-rs/polarspola-rs avatar

    pola-rs/polars

    38,855View on GitHub↗

    Polars is a high-performance columnar data processing library designed for efficient analytical workflows. It functions as a structured data library that organizes information into typed columns, utilizing the Apache Arrow memory format to enable zero-copy data sharing and cache-friendly, vectorized operations. The engine is built to handle large-scale tabular datasets, providing both local and distributed analytical runtimes that scale from single-machine environments to multi-node clusters. The project distinguishes itself through a sophisticated lazy query engine that constructs abstract e

    Rustarrowdataframedataframe-library
    View on GitHub↗38,855
  • modin-project/modinmodin-project avatar

    modin-project/modin

    10,389View on GitHub↗

    Modin is a distributed dataframe library and parallel data processing engine designed to handle large datasets that exceed system memory. It functions as a distributed computing framework that parallelizes data manipulation tasks across multiple CPU cores or clusters to increase throughput and avoid memory errors. The project mirrors the Pandas API, allowing for the distribution of data workflows without changing core code logic. It utilizes a pluggable backend interface, which enables users to switch between different distributed execution engines to optimize performance based on available h

    Pythonanalyticsdata-sciencedataframe
    View on GitHub↗10,389
  • jmcarpenter2/swifterjmcarpenter2 avatar

    jmcarpenter2/swifter

    2,641View on GitHub↗

    A package which efficiently applies any function to a pandas dataframe or series in the fastest available manner

    Python
    View on GitHub↗2,641
See all 30 alternatives to Sklearn Pandas→