30 open-source projects similar to rfordatascience/tidytuesday, ranked by shared indexed features. Tags may describe platforms or build tools rather than the same primary purpose. Check each project’s use case, license, and deployment requirements before treating it as a replacement.
This project is a public health dataset providing historical and real-time COVID-19 case and death counts across the United States. It consists of a collection of CSV files containing time-series pandemic data organized by date, state, and county. The dataset includes specialized records for institutional outbreaks, tracking infection and death rates within correctional facilities, colleges, and universities. It also provides statistics on excess mortality to estimate total pandemic impact and survey-based data on mask usage prevalence across different counties. To facilitate geographic anal
splearn: package for signal processing and machine learning with Python. Contains tutorials on understanding and applying signal processing.
This project is a curated collection of technical reference materials and study guides designed for machine learning interview preparation. It provides comprehensive resources for candidates pursuing engineering roles, focusing on deep learning, production infrastructure, and large-scale system design. The repository distinguishes itself through an architecture that combines theoretical research with industrial case studies. It utilizes a pattern-based approach to system design, breaking down complex deployments—such as recommendation engines, search ranking, and ad click prediction—into reus
Slides, scripts and materials for the Machine Learning in Finance Course at NYU Tandon, 2022
source code from the book Genetic Algorithms with Python by Clinton Sheppard
🐍 Quick reference guide to common patterns & functions in PySpark.
Ways of doing Data Science Engineering and Machine Learning in R and Python
This project is a static site forum archive and housing market dataset. It serves as a read-only repository of preserved community discussions focused on residential property pricing and real estate trends, stored as a collection of Markdown files to ensure long-term data preservation and portability. The system functions as a web content preservation tool that converts historical forum threads into pre-rendered HTML files for serverless access. This approach organizes forum content archiving and housing market analysis into a permanent digital format to prevent data loss. The archive is bui
Corpora is a library of curated, domain-specific text corpora and structured datasets. It provides collections of categorized nouns, adjectives, and verbs in JSON format, intended for use as static training or testing data for conversational agents and automated systems. The project organizes linguistic data across diverse fields, including science, art, and geography. These datasets are distributed as standalone files using a language-neutral schema and a static JSON data model, allowing for direct import into applications without an API layer. The library supports chatbot content generatio
Grav is a flat-file content management system that eliminates the need for a traditional database by storing site content and configuration in human-readable Markdown and YAML files. Built as a modular PHP web framework, it uses a hierarchical page routing system where the physical directory structure directly determines the site's URL paths. The platform is distinguished by its event-driven plugin architecture and a command-line interface that prioritizes system administration, deployment, and maintenance tasks. It utilizes a blueprint-driven system to generate administrative forms from stru
This project is a public health data repository and open government data archive. It provides structured historical records, daily updates on infection trends, and official epidemiological statistics for pandemic monitoring. The resource includes geospatial health data, such as standardized mapping files and administrative boundary data used to visualize regional health restrictions and emergency measures. It also serves as a monitor for public health spending by publishing records of government supply contracts related to emergency health responses. The system utilizes a decoupled data dist
This project is a public health dataset and epidemiological repository providing historical pandemic statistics. It functions as a structured archive of medical records tracking global cases and vaccination rates for public health analysis and longitudinal research. The repository serves as a health data archive where records are stored in comma-separated values to ensure portable data export and visualization. This allows for the retrieval of historical pandemic data to study infectious disease trends and healthcare outcomes. The archive utilizes flat-file storage and a versioned file hiera
This project is a comprehensive collection of Python programming education materials, including tutorials, exercises, and curated code samples. It serves as a learning curriculum and software engineering toolkit, utilizing Jupyter Notebooks to combine executable code with descriptive educational text. The repository provides practical implementation guides for building large language model applications, such as retrieval-augmented generation systems, stateful AI agents, and machine learning workflows. It distinguishes itself by offering a structured approach to agentic coding workflows, cover
This project is a comprehensive geographic location dataset and reference library providing standardized data for countries, states, and cities. It serves as a source of truth for regional hierarchies, ISO codes, coordinates, and timezone information, available as both a relational SQL database and a document-based JSON library. The project includes a custom dataset export tool that functions as a filtering engine. This allows for the generation of tailored geographic files in JSON, CSV, and GeoJSON formats by selecting only the specific regions or fields required. The dataset covers global
A list of colleges and universities offering degrees in data science.
Towhee is a framework that is dedicated to making neural data processing pipelines simple and fast.
:truck: Agile Data Preparation Workflows made easy with Pandas, Dask, cuDF, Dask-cuDF, Vaex and PySpark
Train and Deploy an ML REST API to predict crypto prices, in 10 steps
A general purpose recommender metrics library for fair evaluation.
A PyTorch and TorchDrug based deep learning library for drug pair scoring. (KDD 2022)
Featuretools is an automated feature engineering library and data transformation framework written in Python. It automatically generates machine learning feature vectors from multi-table datasets by applying synthesis patterns to relational and timestamped data. The system functions as a distributed feature synthesis engine, allowing the process of creating feature vectors to scale across multiple cores or clusters to handle large-scale datasets. The library supports the synthesis of multi-table datasets, time series feature generation, and the creation of custom machine learning primitives
Cleanlab is a data-centric AI library and toolkit designed to improve machine learning model performance by detecting label errors and increasing overall dataset quality. It implements a confident learning framework that iteratively refines label noise estimates by comparing model predictions with estimated label probabilities to identify mislabeled examples. The project provides specialized utilities for active learning optimization, allowing for the selection of the most impactful examples for labeling or re-labeling. It also includes an outlier detection tool to identify atypical data poin
Move fast from data science prototype to pipeline. Capture, analyze, and transform messy notebooks into data pipelines with just two lines of code.
Karate Club: An API Oriented Open-source Python Framework for Unsupervised Learning on Graphs (CIKM 2020)
Little Ball of Fur - A graph sampling extension library for NetworKit and NetworkX (CIKM 2020)
Albumentations is a computer vision image augmentation library designed to increase training data diversity for deep learning models. It provides a toolset for applying geometric and color transformations to images and annotations, including a specialized collection of 3D operations for volumetric data used in medical and scientific imaging. The library functions as an image mask and bounding box transformer, automatically updating masks, bounding boxes, and keypoints when images undergo geometric changes. This ensures that spatial alterations remain synchronized across images and their assoc
This project is a data science curriculum and instructional syllabus designed to teach the fundamental principles and tools of the field. It provides a structured set of learning materials, including R programming courseware and guides for statistical learning. The materials focus on the practical application of data science, covering data cleaning, visualization, and exploratory data analysis. It includes resources for mastering specific techniques such as linear regression, classification, and unsupervised learning. The curriculum is organized into a modular sequence of educational modules
ProgrammingVTuberLogos is a collection of high-quality graphic assets and icons representing various programming languages, developer tooling, and software ecosystems. It serves as a standardized library of PNG images designed for use in digital media. The repository provides a set of logos specifically tailored for digital broadcast asset libraries, such as livestream overlays. These assets are also used for software branding and technical content creation, including tutorials and educational documentation. The collection supports the retrieval of stylized imagery for programming presentati
This project is a curated directory and registry of functional open source web and mobile applications. It serves as a categorized catalog of high-quality software projects and scripts available for discovery, learning, and contribution. The directory organizes these resources into a structured index, allowing users to find tools for automation, data processing, and general utility tasks. It provides a reference for exploring modern application architecture through human-verified public code repositories. The collection covers several capability areas, including automation and scheduling, da