awesome-repositories.comCategoriiBlog
MCP
awesome-repositories.com

Descoperă cele mai bune repository-uri open source cu căutare AI.

ExploreazăCăutări recomandateAlternative open-sourceSoftware self-hostedBlogHartă site
ProiectServer MCPDespreCum realizăm clasamentulPresă
LegalConfidențialitateTermeni
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
Back to netflix/pigpen

Open-source alternatives to PigPen

30 open-source projects similar to netflix/pigpen, ranked by how many features they have in common. Compare stars, activity and what each one does to find the best PigPen alternative.

  • twitter/summingbirdAvatar twitter

    twitter/summingbird

    2,125Vezi pe GitHub↗

    Streaming MapReduce with Scalding and Storm

    Scala
    Vezi pe GitHub↗2,125
  • addthis/hydraAvatar addthis

    addthis/hydra

    436Vezi pe GitHub↗
    Java
    Vezi pe GitHub↗436
  • alanmarazzi/pantheraAvatar alanmarazzi

    alanmarazzi/panthera

    191Vezi pe GitHub↗

    Data-frames & arrays on Clojure

    Clojure
    Vezi pe GitHub↗191
  • apache/beamAvatar apache

    apache/beam

    8,612Vezi pe GitHub↗

    Apache Beam is a distributed data pipeline framework and unified data processing model designed to handle both bounded batch data and unbounded real-time streams. It provides a system for building scalable, data-parallel workflows that operate across compute clusters using a single programming model. The framework utilizes a cross-runner pipeline abstraction that decouples the data processing logic from the underlying execution backend, allowing the same pipeline to run on different distributed compute engines. It supports multi-language pipeline development by translating high-level code fro

    Java
    Vezi pe GitHub↗8,612
  • aws/aws-sdk-pandasAvatar aws

    aws/aws-sdk-pandas

    4,107Vezi pe GitHub↗

    aws-sdk-pandas is a Python library that integrates pandas dataframes with AWS services, acting as a cloud data ETL tool and data lake connector. It provides a unified interface to move and transform data between in-memory dataframes and cloud storage, databases, and data warehouses. The project distinguishes itself as a distributed compute orchestrator capable of submitting pandas-based workloads to EMR clusters and serverless processing environments. It further specializes in coordinating distributed data processing via Ray cluster initialization to handle datasets that exceed the memory of

    Pythonamazon-athenaamazon-sagemaker-notebookapache-arrow
    Vezi pe GitHub↗4,107

Căutare AI

Explorează mai multe repository-uri excelente

Descrie ce ai nevoie în limbaj simplu — AI-ul sortează mii de proiecte open source selectate în funcție de relevanță.

Find more with AI search
  • bididi-badidi/fyp-data-analysis-with-llmAvatar bididi-badidi

    bididi-badidi/FYP-Data-Analysis-With-LLM

    10Vezi pe GitHub↗

    Human interpretation of data is inherently susceptible to cognitive biases. While Large Language Models (LLMs) act as automated data analysts, they often mirror user biases or training artifacts. This project introduces a "Bias-Contrastive" Agentic Framework that goes beyond simple text analysis.

    Python
    Vezi pe GitHub↗10
  • bkirwi/coastAvatar bkirwi

    bkirwi/coast

    60Vezi pe GitHub↗

    Experiments in Streaming

    Scala
    Vezi pe GitHub↗60
  • cdslaborg/paramonteAvatar cdslaborg

    cdslaborg/paramonte

    305Vezi pe GitHub↗

    ParaMonte: Parallel Monte Carlo and Machine Learning Library for Python, MATLAB, Fortran, C++, C.

    Fortran
    Vezi pe GitHub↗305
  • damballa/parkourAvatar damballa

    damballa/parkour

    255Vezi pe GitHub↗

    Hadoop MapReduce in idiomatic Clojure.

    Clojure
    Vezi pe GitHub↗255
  • data-centric-ai-community/fg-data-profilingAvatar Data-Centric-AI-Community

    Data-Centric-AI-Community/fg-data-profiling

    13,609Vezi pe GitHub↗

    This project is a data profiling and exploratory data analysis tool designed to generate automated quality reports for Pandas and Spark dataframes. It serves as a system for computing descriptive statistics, identifying correlations, and analyzing univariate and multivariate data patterns. The tool provides specialized capabilities for comparing different versions of datasets to identify changes in data quality and distributions. It includes a dedicated profiler for time-dependent data to extract statistical information such as seasonality and auto-correlation. The software covers a broad an

    Python
    Vezi pe GitHub↗13,609
  • datasalt/pangoolAvatar datasalt

    datasalt/pangool

    57Vezi pe GitHub↗

    Tuple MapReduce for Hadoop: Hadoop API made easy

    Java
    Vezi pe GitHub↗57
  • desbordante/desbordante-coreAvatar Desbordante

    Desbordante/desbordante-core

    484Vezi pe GitHub↗

    Desbordante is a high-performance data profiler that is capable of discovering many different patterns in data using various algorithms. It also allows to run data cleaning scenarios using these algorithms. Desbordante has a console version and an easy-to-use web application.

    C++anomaly-detectioncorrelationsdata-analytics
    Vezi pe GitHub↗484
  • espertechinc/esperAvatar espertechinc

    espertechinc/esper

    875Vezi pe GitHub↗

    Esper Complex Event Processing, Streaming SQL and Event Series Analysis

    Java
    Vezi pe GitHub↗875
  • generateme/fastmathAvatar generateme

    generateme/fastmath

    280Vezi pe GitHub↗

    Fast primitive based math library

    Clojure
    Vezi pe GitHub↗280
  • go-gota/gotaAvatar go-gota

    go-gota/gota

    3,271Vezi pe GitHub↗

    Gota: DataFrames and data wrangling in Go (Golang)

    Go
    Vezi pe GitHub↗3,271
  • gonum/gonumAvatar gonum

    gonum/gonum

    8,316Vezi pe GitHub↗

    Gonum is a numerical computing library for the Go programming language, providing a collection of packages for scientific computing, linear algebra, statistics, and optimization. It functions as a framework for performing complex numerical computations and solving systems of linear equations. The project includes a dedicated graph analysis framework for modeling network graphs and solving connectivity and pathfinding problems. It also provides a statistical analysis toolkit for computing descriptive and inferential statistics and estimating mixture entropy. The library's capability surface c

    Godata-analysisgogolang
    Vezi pe GitHub↗8,316
  • ibis-project/ibisAvatar ibis-project

    ibis-project/ibis

    6,574Vezi pe GitHub↗

    Ibis is a portable Python dataframe library and multi-backend query engine that provides a unified interface for executing data transformations across diverse compute engines. It functions as a Python SQL expression compiler and dialect transpiler, allowing users to define data logic once and execute it across cloud warehouses, embedded databases, and distributed clusters without rewriting code. The project distinguishes itself through a database backend abstraction that decouples transformation logic from the underlying execution engine. It enables polyglot data workflows by mixing raw SQL s

    Pythonbigqueryclickhousedatabase
    Vezi pe GitHub↗6,574
  • ibmstreams/streamsx.topologyAvatar IBMStreams

    IBMStreams/streamsx.topology

    29Vezi pe GitHub↗

    Develop streaming applications for IBM Streams in Python, Java & Scala.

    Java
    Vezi pe GitHub↗29
  • kalyanmurapaka45/article-web-scrapingAvatar KalyanMurapaka45

    KalyanMurapaka45/Article-Web-Scraping

    21Vezi pe GitHub↗

    This Python script is designed to scrape articles from The Guardian's technology section using their API. It fetches article data, extracts the titles and content, and then saves each article's content to separate text files. The text files are organized in a folder named with the current date…

    Jupyter Notebook
    Vezi pe GitHub↗21
  • kalyanmurapaka45/e-commerce-data-analysisK

    KalyanMurapaka45/E-Commerce-Data-Analysis

    0Vezi pe GitHub↗
    Vezi pe GitHub↗0
  • kalyanmurapaka45/end-to-end-image-scrapingAvatar KalyanMurapaka45

    KalyanMurapaka45/End-to-End-Image-Scraping

    14Vezi pe GitHub↗

    The "Image Scraper" is a Flask web application that allows users to search for images on Google and download them directly to their local machines. The project leverages web scraping techniques to fetch the image URLs from Google search results and then download the images to a specified directory.

    Jupyter Notebook
    Vezi pe GitHub↗14
  • kalyanmurapaka45/indian-restaurants-data-analysisAvatar KalyanMurapaka45

    KalyanMurapaka45/Indian-Restaurants-Data-Analysis

    8Vezi pe GitHub↗

    This repository contains a Power BI data analysis project on Indian restaurants, enabling you to delve into restaurant data, customer preferences, and regional trends.

    Vezi pe GitHub↗8
  • kalyanmurapaka45/virat-kohli-score-analyticsAvatar KalyanMurapaka45

    KalyanMurapaka45/Virat-Kohli-Score-Analytics

    12Vezi pe GitHub↗

    This project leverages Power BI to create a dynamic and visually appealing analytics dashboard focused on the cricket performances of the legendary Virat Kohli. Gain insights into his batting trends, run-scoring patterns, and statistical analysis over time.

    Vezi pe GitHub↗12
  • khanhnamle1994/spotify-artists-analysisK

    khanhnamle1994/spotify-artists-analysis

    0Vezi pe GitHub↗
    Vezi pe GitHub↗0
  • khanhnamle1994/world-cup-2018K

    khanhnamle1994/world-cup-2018

    0Vezi pe GitHub↗
    Vezi pe GitHub↗0
  • marcelotto/jsonld-exAvatar marcelotto

    marcelotto/jsonld-ex

    84Vezi pe GitHub↗

    An implementation of JSON-LD for Elixir

    Elixir
    Vezi pe GitHub↗84
  • marcelotto/rdf-exAvatar marcelotto

    marcelotto/rdf-ex

    125Vezi pe GitHub↗

    An implementation of RDF for Elixir

    Elixir
    Vezi pe GitHub↗125
  • marcelotto/sparql-exAvatar marcelotto

    marcelotto/sparql-ex

    44Vezi pe GitHub↗

    An implementation of SPARQL for Elixir

    Elixir
    Vezi pe GitHub↗44
  • mastodonc/kixi.statsAvatar MastodonC

    MastodonC/kixi.stats

    368Vezi pe GitHub↗

    A library of statistical distribution sampling and transducing functions

    Clojure
    Vezi pe GitHub↗368
  • modin-project/modinAvatar modin-project

    modin-project/modin

    10,389Vezi pe GitHub↗

    Modin is a distributed dataframe library and parallel data processing engine designed to handle large datasets that exceed system memory. It functions as a distributed computing framework that parallelizes data manipulation tasks across multiple CPU cores or clusters to increase throughput and avoid memory errors. The project mirrors the Pandas API, allowing for the distribution of data workflows without changing core code logic. It utilizes a pluggable backend interface, which enables users to switch between different distributed execution engines to optimize performance based on available h

    Pythonanalyticsdata-sciencedataframe
    Vezi pe GitHub↗10,389