9 مستودعات
Systems for executing large-scale data processing tasks with support for optimization and distributed architectures.
Explore 9 awesome GitHub repositories matching data & databases · Distributed Data Processing Engines. Refine with filters or upvote what's useful.
Apache Spark is a unified distributed data processing engine designed for large-scale data analysis and computation graphs. It functions as a distributed machine learning framework, a graph processing system, a real-time stream processor, and a SQL analytics engine. The system enables the execution of distributed SQL querying, large-scale graph analysis, and real-time stream analytics across clusters of machines. It also provides a scalable environment for implementing machine learning algorithms and predictive model development on massive datasets. The engine incorporates relational query e
Functions as a unified engine for executing large-scale data analysis and computation graphs across clusters.
This project is a collection of foundational machine learning algorithms and data science tools implemented in Python. It focuses on building the logic of these tools using basic programming primitives rather than relying on specialized libraries. The implementation covers several core domains, including a linear algebra library for matrix and vector operations, a statistical analysis toolkit for probability and hypothesis testing, and a framework for map-reduce distributed processing. It also includes implementations for natural language processing, graph theory for network analysis, and var
Implements distributed data processing systems using map-reduce techniques to handle large datasets.
Pentaho Kettle هو منصة مؤسسية لدمج البيانات (ETL) مصممة لاستخراج وتحويل وتحميل البيانات بين المصادر المتباينة وقواعد البيانات المستهدفة. يعمل كمنظم قائم على البيانات الوصفية يستخدم مصمماً مرئياً لسير العمل لإنشاء وإدارة تسلسلات معقدة من مهام البيانات وخطوط أنابيب التحويل. يتميز النظام بمحرك معالجة بيانات موزع، يقوم بتنفيذ أعباء العمل عبر مجموعات من عقد الخادم لزيادة الإنتاجية. يستخدم بنية قائمة على الإضافات، مما يسمح بتوسيع المنصة عبر ملفات JAR خارجية لتوفير الاتصال بقواعد بيانات وخدمات سحابية متنوعة. تغطي المنصة مجموعة واسعة من قدرات دمج البيانات، بما في ذلك التحميل بالجملة، وإدارة الملفات عن بُعد، وتحويل هيكل البيانات. توفر أدوات للتحقق من جودة البيانات، وأتمتة خطوط الأنابيب، وإدارة دورة حياة الوظائف، إلى جانب أدوات مراقبة لتتبع صحة الخادم وحالة التنفيذ في الوقت الفعلي.
Implements a processing engine that executes large-scale data transformation tasks with support for distributed architectures.
Storm is a distributed stream processing framework designed to execute unbounded computations across a cluster to process real-time data streams. It functions as a data pipeline orchestrator that allows users to define and deploy declarative data flow graphs connecting streaming sources to processing components. The system operates as a multi-tenant distributed compute engine that isolates workloads and limits resource usage across shared clusters using dedicated pools and access control. It is also a secure distributed processing engine that employs encrypted node communication and SSL-secur
Implements a processing engine with encrypted node communication and SSL-secured management interfaces.
Data-Juicer is an open-source framework for cleaning, filtering, deduplicating, and transforming multimodal datasets to prepare them for training large language and vision models. It functions as a distributed data pipeline engine that runs processing jobs across Ray clusters, handling billions of samples with automatic operator fusion and adaptive parallelism. The framework provides a library of operators that leverage large language models for semantic extraction, filtering, and data synthesis within processing pipelines. The project distinguishes itself through a YAML-based data recipe sys
Runs data processing jobs across Ray clusters, handling billions of samples with automatic operator fusion and adaptive parallelism.
Cube Studio هو منصة MLOps سحابية ومنظم ذكاء اصطناعي يعتمد على Kubernetes ومصمم لدورة حياة تعلم الآلة بالكامل. يوفر إطار عمل للتدريب الموزع لضبط النماذج على نطاق واسع، ومدير موارد GPU لافتراضية الأجهزة، ومنظم لخطوط أنابيب تعلم الآلة يستخدم رسوم بيانية موجهة غير دورية (DAGs) لإدارة سير العمل من البداية إلى النهاية. تتميز المنصة بخادم استنتاج LLM متخصص يدعم التوليد المعزز بالاسترجاع (RAG) وبناء قواعد المعرفة الخاصة. كما يتميز بنظام مخصص للضبط الخاضع للإشراف والتعلم التعزيزي لنماذج اللغات الكبيرة، مدعوماً بأدوات مرئية للبحث عن المعاملات الفائقة (Hyperparameters). يغطي النظام نطاقاً واسعاً من القدرات التشغيلية، بما في ذلك تصنيف البيانات متعددة الوسائط، وخطوط أنابيب البيانات الموزعة، وجدولة أحمال العمل عبر مجموعات متعددة. كما يوفر بيئات تطوير تفاعلية تعتمد على المتصفح، وإدارة صور الحاويات، وسجل نماذج لإصدار ونشر واجهات برمجة تطبيقات استنتاج قابلة للتوسع مع تقسيم حركة المرور. تتضمن البنية التحتية مراقبة صحة المجموعات (Cluster Health) والتحكم في الوصول القائم على الأدوار مع تكامل تسجيل الدخول الموحد (SSO).
Executes distributed jobs to import heterogeneous data and extract features using big data processing engines.
This project is a dataset management framework and cross-framework data loader that provides a unified interface for reading data formats compatible with TensorFlow, JAX, and PyTorch. It serves as a library of curated public datasets provided as data streams and includes tools for building, versioning, and documenting large-scale datasets. The system differentiates itself through a distributed data processing engine capable of managing massive datasets across clusters using parallelized pipelines. It utilizes builder-based construction to standardize how data is downloaded and prepared, while
Integrates with parallel processing engines like Apache Beam to process large-scale data across distributed clusters.
This project is an interactive, web-based notebook environment designed for distributed data science and large-scale computing. It serves as a development tool for executing code and performing data analysis specifically within the Apache Spark framework, providing a browser-based interface that combines code execution with reactive data visualization. The platform distinguishes itself through its deep integration with distributed infrastructure, allowing users to manage cluster resources, configure runtime dependencies, and isolate execution processes for individual notebooks. It supports co
Provides an interactive environment for running code, queries, and data processing jobs using a pre-configured distributed computing engine.
Data engineering practice repository providing tutorials, distributed processing engines, and Python data pipeline automation scripts. The system encompasses automated data validation, distributed compute aggregation, embedded columnar querying, lazy evaluation planning, partitioned storage export, and cloud storage retrieval. The capability surface covers cloud integration and storage, data engineering and pipelines, data processing and analytics, data quality and testing, database and storage, file management, and monitoring and observability.
Runs lazy evaluations, transformations, and numerical aggregations across large-scale tabular datasets.