awesome-repositories.com
Blog
MCP
awesome-repositories.com

Discover the best open-source repositories with AI-powered search.

ExploreCurated searchesOpen-source alternativesSelf-hosted softwareBlogSitemap
ProjectMCP serverAboutHow we rankPress
LegalPrivacyTerms
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
Back to techascent/tech.ml.dataset

Open-source alternatives to Tech.ml.dataset

30 open-source projects similar to techascent/tech.ml.dataset, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best Tech.ml.dataset alternative.

  • alanmarazzi/pantheraalanmarazzi avatar

    alanmarazzi/panthera

    191View on GitHub↗

    Data-frames & arrays on Clojure

    Clojure
    View on GitHub↗191
  • aws/aws-sdk-pandasaws avatar

    aws/aws-sdk-pandas

    4,107View on GitHub↗

    aws-sdk-pandas is a Python library that integrates pandas dataframes with AWS services, acting as a cloud data ETL tool and data lake connector. It provides a unified interface to move and transform data between in-memory dataframes and cloud storage, databases, and data warehouses. The project distinguishes itself as a distributed compute orchestrator capable of submitting pandas-based workloads to EMR clusters and serverless processing environments. It further specializes in coordinating distributed data processing via Ray cluster initialization to handle datasets that exceed the memory of

    Pythonamazon-athenaamazon-sagemaker-notebookapache-arrow
    View on GitHub↗4,107
  • bididi-badidi/fyp-data-analysis-with-llmbididi-badidi avatar

    bididi-badidi/FYP-Data-Analysis-With-LLM

    10View on GitHub↗

    Human interpretation of data is inherently susceptible to cognitive biases. While Large Language Models (LLMs) act as automated data analysts, they often mirror user biases or training artifacts. This project introduces a "Bias-Contrastive" Agentic Framework that goes beyond simple text analysis.

    Python
    View on GitHub↗10
  • cdslaborg/paramontecdslaborg avatar

    cdslaborg/paramonte

    305View on GitHub↗

    ParaMonte: Parallel Monte Carlo and Machine Learning Library for Python, MATLAB, Fortran, C++, C.

    Fortran
    View on GitHub↗305

AI search

Explore more awesome repositories

Describe what you need in plain English — the AI ranks thousands of curated open-source projects by relevance.

Find more with AI search
  • data-centric-ai-community/fg-data-profilingData-Centric-AI-Community avatar

    Data-Centric-AI-Community/fg-data-profiling

    13,609View on GitHub↗

    This project is a data profiling and exploratory data analysis tool designed to generate automated quality reports for Pandas and Spark dataframes. It serves as a system for computing descriptive statistics, identifying correlations, and analyzing univariate and multivariate data patterns. The tool provides specialized capabilities for comparing different versions of datasets to identify changes in data quality and distributions. It includes a dedicated profiler for time-dependent data to extract statistical information such as seasonality and auto-correlation. The software covers a broad an

    Python
    View on GitHub↗13,609
  • desbordante/desbordante-coreDesbordante avatar

    Desbordante/desbordante-core

    484View on GitHub↗

    Desbordante is a high-performance data profiler that is capable of discovering many different patterns in data using various algorithms. It also allows to run data cleaning scenarios using these algorithms. Desbordante has a console version and an easy-to-use web application.

    C++anomaly-detectioncorrelationsdata-analytics
    View on GitHub↗484
  • generateme/fastmathgenerateme avatar

    generateme/fastmath

    280View on GitHub↗

    Fast primitive based math library

    Clojure
    View on GitHub↗280
  • go-gota/gotago-gota avatar

    go-gota/gota

    3,271View on GitHub↗

    Gota: DataFrames and data wrangling in Go (Golang)

    Go
    View on GitHub↗3,271
  • gonum/gonumgonum avatar

    gonum/gonum

    8,316View on GitHub↗

    Gonum is a numerical computing library for the Go programming language, providing a collection of packages for scientific computing, linear algebra, statistics, and optimization. It functions as a framework for performing complex numerical computations and solving systems of linear equations. The project includes a dedicated graph analysis framework for modeling network graphs and solving connectivity and pathfinding problems. It also provides a statistical analysis toolkit for computing descriptive and inferential statistics and estimating mixture entropy. The library's capability surface c

    Godata-analysisgogolang
    View on GitHub↗8,316
  • ibis-project/ibisibis-project avatar

    ibis-project/ibis

    6,574View on GitHub↗

    Ibis is a portable Python dataframe library and multi-backend query engine that provides a unified interface for executing data transformations across diverse compute engines. It functions as a Python SQL expression compiler and dialect transpiler, allowing users to define data logic once and execute it across cloud warehouses, embedded databases, and distributed clusters without rewriting code. The project distinguishes itself through a database backend abstraction that decouples transformation logic from the underlying execution engine. It enables polyglot data workflows by mixing raw SQL s

    Pythonbigqueryclickhousedatabase
    View on GitHub↗6,574
  • kalyanmurapaka45/article-web-scrapingKalyanMurapaka45 avatar

    KalyanMurapaka45/Article-Web-Scraping

    21View on GitHub↗

    This Python script is designed to scrape articles from The Guardian's technology section using their API. It fetches article data, extracts the titles and content, and then saves each article's content to separate text files. The text files are organized in a folder named with the current date…

    Jupyter Notebook
    View on GitHub↗21
  • kalyanmurapaka45/e-commerce-data-analysisK

    KalyanMurapaka45/E-Commerce-Data-Analysis

    0View on GitHub↗
    View on GitHub↗0
  • kalyanmurapaka45/end-to-end-image-scrapingKalyanMurapaka45 avatar

    KalyanMurapaka45/End-to-End-Image-Scraping

    14View on GitHub↗

    The "Image Scraper" is a Flask web application that allows users to search for images on Google and download them directly to their local machines. The project leverages web scraping techniques to fetch the image URLs from Google search results and then download the images to a specified directory.

    Jupyter Notebook
    View on GitHub↗14
  • kalyanmurapaka45/indian-restaurants-data-analysisKalyanMurapaka45 avatar

    KalyanMurapaka45/Indian-Restaurants-Data-Analysis

    8View on GitHub↗

    This repository contains a Power BI data analysis project on Indian restaurants, enabling you to delve into restaurant data, customer preferences, and regional trends.

    View on GitHub↗8
  • kalyanmurapaka45/virat-kohli-score-analyticsKalyanMurapaka45 avatar

    KalyanMurapaka45/Virat-Kohli-Score-Analytics

    12View on GitHub↗

    This project leverages Power BI to create a dynamic and visually appealing analytics dashboard focused on the cricket performances of the legendary Virat Kohli. Gain insights into his batting trends, run-scoring patterns, and statistical analysis over time.

    View on GitHub↗12
  • khanhnamle1994/spotify-artists-analysisK

    khanhnamle1994/spotify-artists-analysis

    0View on GitHub↗
    View on GitHub↗0
  • khanhnamle1994/world-cup-2018K

    khanhnamle1994/world-cup-2018

    0View on GitHub↗
    View on GitHub↗0
  • mastodonc/kixi.statsMastodonC avatar

    MastodonC/kixi.stats

    368View on GitHub↗

    A library of statistical distribution sampling and transducing functions

    Clojure
    View on GitHub↗368
  • modin-project/modinmodin-project avatar

    modin-project/modin

    10,389View on GitHub↗

    Modin is a distributed dataframe library and parallel data processing engine designed to handle large datasets that exceed system memory. It functions as a distributed computing framework that parallelizes data manipulation tasks across multiple CPU cores or clusters to increase throughput and avoid memory errors. The project mirrors the Pandas API, allowing for the distribution of data workflows without changing core code logic. It utilizes a pluggable backend interface, which enables users to switch between different distributed execution engines to optimize performance based on available h

    Pythonanalyticsdata-sciencedataframe
    View on GitHub↗10,389
  • netflix/pigpenNetflix avatar

    Netflix/PigPen

    565View on GitHub↗

    Map-Reduce for Clojure

    Clojure
    View on GitHub↗565
  • pandas-dev/pandaspandas-dev avatar

    pandas-dev/pandas

    49,039View on GitHub↗

    Pandas is a high-performance data analysis library that provides a comprehensive framework for manipulating, cleaning, and transforming structured datasets. It centers on labeled one-dimensional and two-dimensional data structures, allowing users to construct, filter, and reshape tabular information while performing complex arithmetic and logical operations. The library distinguishes itself through a sophisticated indexing engine that enables automatic data alignment during calculations and relational merges. By utilizing a block-based memory layout, it optimizes cache locality for vectorized

    Pythonalignmentdata-analysisdata-science
    View on GitHub↗49,039
  • pathwaycom/pathwaypathwaycom avatar

    pathwaycom/pathway

    62,959View on GitHub↗

    Pathway is a high-performance data processing framework designed for building unified batch and streaming pipelines. It functions as an orchestrator for complex data transformations, utilizing a differential dataflow engine to process updates incrementally. By treating static datasets and continuous event streams with identical logic, the platform ensures exactly-once processing semantics and consistent results across diverse data sources. The framework distinguishes itself through its specialized support for real-time artificial intelligence and retrieval-augmented generation. It features in

    Pythonbatch-processingdata-analyticsdata-pipelines
    View on GitHub↗62,959
  • pola-rs/polarspola-rs avatar

    pola-rs/polars

    38,855View on GitHub↗

    Polars is a high-performance columnar data processing library designed for efficient analytical workflows. It functions as a structured data library that organizes information into typed columns, utilizing the Apache Arrow memory format to enable zero-copy data sharing and cache-friendly, vectorized operations. The engine is built to handle large-scale tabular datasets, providing both local and distributed analytical runtimes that scale from single-machine environments to multi-node clusters. The project distinguishes itself through a sophisticated lazy query engine that constructs abstract e

    Rustarrowdataframedataframe-library
    View on GitHub↗38,855
  • project-ryoma/ryomaproject-ryoma avatar

    project-ryoma/ryoma

    406View on GitHub↗

    AI Powered Data Agent framework, a comprehensive solution for data analysis, engineering, and visualization.

    Python
    View on GitHub↗406
  • rocketlaunchr/dataframe-gorocketlaunchr avatar

    rocketlaunchr/dataframe-go

    1,287View on GitHub↗

    DataFrames for Go: For statistics, machine-learning, and data manipulation/exploration

    Godata-sciencedataframedataframes
    View on GitHub↗1,287
  • scicloj/tableclothscicloj avatar

    scicloj/tablecloth

    365View on GitHub↗

    Dataset manipulation library built on the top of tech.ml.dataset

    Clojure
    View on GitHub↗365
  • simonw/datasettesimonw avatar

    simonw/datasette

    11,198View on GitHub↗

    Datasette is a tool for publishing and sharing SQLite databases as public websites. It functions as a data publishing system that provides searchable interfaces and JSON APIs to expose the contents of SQLite files. The project enables both server-side and client-side execution. It can operate as an API server or as a database browser that runs entirely within a web browser using WebAssembly, allowing for serverless database access. The system supports a variety of deployment strategies, including containerized images for cloud hosting and a local development server for testing. It includes c

    Pythonasgiautomatic-apicsv
    View on GitHub↗11,198
  • starpig1129/ai-data-analysis-multiagentstarpig1129 avatar

    starpig1129/AI-Data-Analysis-MultiAgent

    1,762View on GitHub↗

    DATAGEN is a powerful brand name that represents our vision of leveraging artificial intelligence technology for data generation and analysis. The name combines "DATA" and "GEN"(generation), perfectly embodying the core functionality of this project - automated data analysis and research through…

    Python
    View on GitHub↗1,762
  • stripe/veneurstripe avatar

    stripe/veneur

    1,746View on GitHub↗

    A distributed, fault-tolerant pipeline for observability data

    Go
    View on GitHub↗1,746
  • usefathom/fathomusefathom avatar

    usefathom/fathom

    8,005View on GitHub↗

    Fathom is a privacy-focused website analytics server written in Go. It monitors website traffic and page views without collecting personal data or using intrusive cookies, providing a self-hosted alternative for traffic monitoring. The system utilizes a Preact-based dashboard interface for visualizing traffic patterns and reports. Data is persisted in a SQL database analytics store, with support for MySQL, PostgreSQL, and SQLite. The project covers the collection of visitor data via lightweight tracking snippets and the management of that data through a pluggable storage layer. It includes m

    Goanalyticsfathomfathom-analytics
    View on GitHub↗8,005