7 个仓库
CLI interfaces that orchestrate multi-stage workflows including configuration, training, evaluation, and visualization.
Distinct from CLI Execution: Distinct from CLI Execution: focuses on orchestrating a complete multi-stage pipeline, not just running a single command.
Explore 7 awesome GitHub repositories matching development tools & productivity · Pipeline Execution Interfaces. Refine with filters or upvote what's useful.
PaddleX is a PaddlePaddle-based framework for building, deploying, and fine-tuning AI model pipelines, with pre-built support for computer vision, OCR, document analysis, and time series tasks. It offers a toolkit of ready-to-use pipelines for image classification, object detection, segmentation, and pose estimation, alongside an end-to-end OCR document analysis pipeline that extracts text, tables, formulas, and layout information. The platform also includes a dedicated time series forecasting pipeline for analyzing historical data to detect anomalies, classify patterns, and predict future val
Executes pre-built processing pipelines by specifying name, input file, and target device in a single terminal command.
Anomalib is a PyTorch-based library for visual anomaly detection, offering a modular framework, a comprehensive model zoo, and a benchmarking suite designed for industrial defect detection. It provides a wide range of algorithms—including generative, discriminative, teacher-student, and vision-language approaches—that support unsupervised, few-shot, and zero-shot settings. The library enables deployment through model export to ONNX and OpenVINO for edge devices, and includes a no-code web application for training and inference. It also features a command-line interface for orchestrating multi
Orchestrates the complete anomaly detection workflow through a command-line interface.
ZenML is an extensible machine learning orchestration framework designed to manage the end-to-end lifecycle of data pipelines and AI agent workflows. It functions as a durable orchestrator that executes machine learning tasks as directed acyclic graphs, ensuring that every step is containerized for consistent performance across local, cloud, and hybrid infrastructure. By decoupling pipeline code from underlying compute and storage backends, the platform allows developers to define infrastructure-agnostic stacks that remain portable across diverse environments. The project distinguishes itself
Triggers and manages pipeline runs from a dashboard interface by deploying ad-hoc runners into configured compute environments.
ZenML is an orchestration platform designed for building, deploying, and monitoring reproducible machine learning pipelines and agentic workflows. It provides a unified framework that manages the entire lifecycle of machine learning assets, from data processing and model training to the deployment of persistent inference services. By decoupling pipeline logic from underlying compute and storage, the platform enables teams to transition workflows seamlessly from local development environments to production-grade cloud infrastructure. The platform distinguishes itself through a service-oriented
Triggers machine learning pipelines directly from a web dashboard by spawning ephemeral jobs.
这是一个中文自然语言处理工具包,提供了一套用于分词、词性标注和命名实体识别的工具。它包括一个用于分析词间句法和语义关系的神经依存句法分析器,以及一个用于使用标注数据集创建自定义语言模型的机器学习训练套件。 该工具包的独特之处在于其部署灵活性,提供了一个 Docker 化服务器和一个通过 API 暴露处理能力的 Web 服务接口。它支持使用预训练模型,并允许集成外部词库和词典扩展以提高分析准确性。 该项目广泛涵盖了完整的语言任务流水线,包括句子分割、句法依存映射和语义角色标注。这些功能可通过命令行界面、独立模块或集成分析流水线使用。 核心逻辑采用 C++ 实现,并提供 Python 和 Java 的官方语言绑定。
Provides a command-line interface to execute a full pipeline of text processing.
Arroyo is a high-performance stream processing platform built in Rust. It executes continuous SQL queries on streaming data with event-time semantics, enabling accurate windowed aggregations, joins, and stateful computations on unbounded event streams. The platform uses native Rust execution for high throughput and low latency, with periodic checkpointing for exactly-once fault tolerance and horizontal scaling across distributed workers. The system integrates deeply with Kafka for reading and writing topics with exactly-once delivery and supports change data capture (CDC) from MySQL and Postg
Starts a stream processing pipeline directly from the command line, accepting SQL from standard input or as an argument.
Git stats is a command-line utility and reporting tool that analyzes project activity and generates statistical insights from source code version control data. It functions as a Git history analysis tool and repository analytics generator, processing historical commit logs and file modification patterns to track how codebases grow and change over time. The application operates through a command-line interface execution pipeline that parses raw repository logs and commit streams directly into structured data records. It includes an incremental activity aggregator that rolls up individual commi
Orchestrates multi-stage analysis pipelines from repository scanning to report generation via CLI arguments.