awesome-repositories.com
Blog
MCP
awesome-repositories.com

Descubre los mejores repositorios open-source con nuestra búsqueda potenciada por IA.

ExplorarBúsquedas curadasAlternativas open-sourceSoftware autohospedableBlogMapa del sitio
ProyectoAcerca deCómo clasificamosPrensaServidor MCP
Aviso legalPrivacidadTérminos
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
h2oai avatar

h2oai/h2ogpt

0
View on GitHub↗
12,016 estrellas·1,313 forks·Python·apache-2.0·9 vistash2o.ai↗

H2ogpt

h2oGPT is a self-hosted platform designed for running large language models and executing retrieval-augmented generation workflows locally. It provides a comprehensive web interface that allows users to index private document collections into searchable databases, enabling context-aware question answering and summarization without exposing sensitive data to external services.

The platform distinguishes itself by offering a modular architecture that supports both local model execution and connections to external inference servers. It facilitates the development of autonomous agents capable of performing multi-step tasks by delegating actions to various tools and models. Beyond simple chat, the system includes capabilities for fine-tuning models on local hardware and managing the full lifecycle of predictive assets, from data ingestion and feature engineering to model deployment and performance monitoring.

The software covers a broad range of enterprise-grade requirements, including document intelligence for extracting structured data from unstructured files, multi-GPU training support, and robust access control mechanisms. It provides tools for model explainability, compliance tracking, and collaborative experiment management to ensure transparency and reproducibility in machine learning workflows.

The project is designed for containerized deployment, utilizing standard configuration files to ensure consistent execution across local and cloud environments.

Features

  • Generative AI Dashboards - Provides a comprehensive dashboard for interacting with models, managing documents, and executing RAG workflows.
  • Retrieval Augmented Generation - Connects private document collections to language models for context-aware question answering and summarization.
  • Local Model Execution - Executes large language models directly on local hardware to maintain data privacy and control.
  • Self-Hosted AI Platforms - Offers a self-hosted interface for running large language models locally to perform document analysis.
  • Local Language Model Hosting - Hosts generative AI models on local hardware to ensure data privacy and full control over sensitive information.
  • Document Indexing - Indexes local files into searchable databases to enable context-aware retrieval-augmented generation.
  • RAG Frameworks - Indexes private documents into searchable databases to enable context-aware responses without external data exposure.
  • Agentic Workflow Orchestration - Coordinates multi-step tasks by delegating actions to external tools and models within a unified execution environment.
  • Autonomous Agent Orchestration - Orchestrates autonomous agents that combine reasoning with external tools to execute complex multi-step workflows.
  • Generative Content Tools - Generates situation-specific responses and summaries based on private document collections.
  • Intelligent Document Processing - Extracts structured data and insights from unstructured files using automated processing and intelligent character recognition.
  • Remote Inference Offloaders - Offload model execution to third-party engines while maintaining centralized control over application logic and request routing across your infrastructure.
  • Model Fine-Tuning - Enables customization and training of open-source models on local infrastructure for domain-specific tasks.
  • Multi-GPU Training Utilities - Distribute intensive deep learning workloads across multiple graphics processing units to accelerate model training times for large datasets.
  • Private Cloud Hosting - Hosts end-to-end private AI platforms on-premises or in private clouds to ensure data sovereignty.
  • Agentic Workflow Automation - Connect automated AI capabilities to enterprise platforms and document repositories to perform complex searches and execute tasks across existing business processes.
  • Autonomous Task Execution - Perform multi-step tasks by delegating actions to external models and tools within an experimental environment to test complex logic.
  • Inference Backends - Supports interchangeable model execution engines to allow flexible switching between local hardware and external API providers.
  • Model Deployment Pipelines - Deploys models to production via real-time or batch endpoints with support for testing and updates.
  • No-Code Training Interfaces - Build and tune models for text, image, video, and audio data using a no-code interface and expert-curated training techniques.
  • General Purpose Models - Framework for building and deploying open-source large language models.
  • Language Model Development - Open-source generative AI for private model ownership.
  • Document and Unstructured Extraction - Extracts structured data from unstructured documents using optical character recognition and machine learning.
  • Model Deployment Management - Manages the full lifecycle of document processing models including deployment, monitoring, and logging.
  • LLM Performance Monitoring - Tracks accuracy, bias, and data drift in production to ensure models remain effective.
  • AI Monitoring - Evaluates model outputs through automated testing and risk monitoring to ensure safety and compliance.
  • Automated Machine Learning Tools - Perform machine learning tasks and create deep learning models using specialized engines that remove the need for manual coding or complex configuration.
  • Automated Lineage Capturers - Captures metadata and event logs throughout the machine learning lifecycle to ensure traceability and reproducibility.
  • Model Lineage Trackers - Records comprehensive event logs and versioning data throughout the machine learning lifecycle to ensure reproducibility and compliance.
  • Data Preparation - Processes diverse unstructured data types like text, images, and audio to prepare them for predictive modeling.
  • Dataset Integration - Pull data from various external storage systems and cloud repositories to prepare information for model training and analysis.
  • External Model Connectors - Centralize machine learning models from various frameworks to simplify the deployment, tracking, and monitoring of your predictive assets.
  • Feature Stores - Integrate with pipelines to reduce redundant data ingestion while providing metadata-driven recommendations for feature discovery and generation.
  • Information Extraction - Identifies and isolates key entities like names and invoice numbers from documents using natural language processing.
  • Model Deployment Pipelines - Exports and deploys finished models to external environments for production serving.
  • Model Interpretability Tools - Provides visual and analytical tools to interpret and explain machine learning model decisions.
  • Model Validation Tools - Verifies model performance through backtesting and drift monitoring to maintain reliability.
  • Managed Cloud Deployments - Provides managed infrastructure for hosting and scaling machine learning environments in the cloud.
  • Compliance and Governance - Automates data versioning and lineage tracking to ensure adherence to privacy and regulatory standards.
  • Identity and Access Management - Controls user and group access to projects and model artifacts to ensure secure collaboration.
  • Human-in-the-Loop Systems - Integrates human-in-the-loop evaluation to refine model performance based on real-world outcomes.
  • Inference Latency Optimizers - Minimizes data transfer delays by placing models near core application logic.
  • Predictive Machine Learning Analytics - Select, parameterize, and train optimal machine learning algorithms based on user-defined target variables and specific business requirements.
  • Model Training Optimizers - Automates parameter adjustment and result analysis to improve model precision through repeated testing.
  • Team Collaboration Management - Offers collaborative tools for data science teams to version, compare, and register machine learning models.
  • Feature Engineering Tools - Automates the identification and calculation of data features to enhance predictive model performance.
  • Tabular Predictive Models - Enables automated tabular predictions without requiring manual model training or complex infrastructure.
  • RESTful APIs - Provide access to document processing, model training, and scoring capabilities through a standard interface to connect services with existing applications.
  • AI Application Deployment Platforms - Share and deploy artificial intelligence applications across an organization using a centralized marketplace to accelerate internal innovation.
  • Cloud Infrastructure Deployment - Supports containerized deployment across diverse cloud and on-premises environments to ensure portability.
  • Container Deployment Configurations - Simplifies platform deployment and dependency management using standard Docker Compose configurations.

Historial de estrellas

Gráfico del historial de estrellas de h2oai/h2ogptGráfico del historial de estrellas de h2oai/h2ogpt

Búsqueda con IA

Explora más repositorios increíbles

Describe lo que necesitas en lenguaje sencillo: la IA clasifica miles de proyectos open-source curados por relevancia.

Start searching with AI

Alternativas open-source a H2ogpt

Proyectos open-source similares, clasificados según cuántas características comparten con H2ogpt.
  • kilo-org/kilocodeAvatar de Kilo-Org

    Kilo-Org/kilocode

    15,616Ver en GitHub↗

    Kilocode is an autonomous engineering platform designed to orchestrate AI agents for complex software development tasks. It functions as a comprehensive system for automating coding, testing, and repository management by integrating directly with your codebase and terminal. The platform provides a unified gateway for model orchestration, allowing for the management of agentic workflows, event-driven automation, and persistent session state across distributed development environments. The platform distinguishes itself through its federated task management and policy-based access control, which

    TypeScriptaiai-ageai-coding
    Ver en GitHub↗15,616
  • letta-ai/lettaAvatar de letta-ai

    letta-ai/letta

    21,168Ver en GitHub↗

    Letta is a framework for building, deploying, and managing autonomous AI agents that maintain persistent state across long-term interactions. It provides a comprehensive suite of primitives for defining agents with configurable personas, modular memory blocks, and tool-use capabilities, enabling them to retain user preferences and conversation history over extended sessions. The platform distinguishes itself through its advanced memory management and orchestration capabilities. It allows agents to autonomously update their own memory, perform retrieval-augmented generation, and coordinate com

    Pythonaiai-agentsllm
    Ver en GitHub↗21,168
  • mastra-ai/mastraAvatar de mastra-ai

    mastra-ai/mastra

    21,221Ver en GitHub↗

    Mastra is an orchestration framework designed for building, deploying, and managing autonomous AI agents and multi-agent systems. It provides a comprehensive suite of primitives for creating resilient AI applications, including durable workflow orchestration, event-driven agent loops, and semantic memory management. By integrating these core components, the platform enables developers to build complex, multi-step processes that can reason about goals and execute tasks without manual intervention. The framework distinguishes itself through its focus on observability and secure, isolated execut

    TypeScriptagentsaichatbots
    Ver en GitHub↗21,221
  • datawhalechina/llm-cookbookAvatar de datawhalechina

    datawhalechina/llm-cookbook

    24,263Ver en GitHub↗

    This repository is a comprehensive set of tutorials and examples for building software powered by large language models. It serves as an application development guide and a prompt engineering framework, providing instructional content for integrating model logic with user interfaces and external data sources. The project provides technical walkthroughs for specialized workflows, including the implementation of retrieval augmented generation using vector databases and semantic search. It includes guidance on adapting pre-trained model weights through fine-tuning with private datasets and the o

    Jupyter Notebookcookbookllm
    Ver en GitHub↗24,263
Ver las 30 alternativas a H2ogpt→

Preguntas frecuentes

¿Qué hace h2oai/h2ogpt?

h2oGPT is a self-hosted platform designed for running large language models and executing retrieval-augmented generation workflows locally. It provides a comprehensive web interface that allows users to index private document collections into searchable databases, enabling context-aware question answering and summarization without exposing sensitive data to external services.

¿Cuáles son las características principales de h2oai/h2ogpt?

Las características principales de h2oai/h2ogpt son: Generative AI Dashboards, Retrieval Augmented Generation, Local Model Execution, Self-Hosted AI Platforms, Local Language Model Hosting, Document Indexing, RAG Frameworks, Agentic Workflow Orchestration.

¿Qué alternativas de código abierto existen para h2oai/h2ogpt?

Las alternativas de código abierto para h2oai/h2ogpt incluyen: kilo-org/kilocode — Kilocode is an autonomous engineering platform designed to orchestrate AI agents for complex software development… letta-ai/letta — Letta is a framework for building, deploying, and managing autonomous AI agents that maintain persistent state across… mastra-ai/mastra — Mastra is an orchestration framework designed for building, deploying, and managing autonomous AI agents and… datawhalechina/llm-cookbook — This repository is a comprehensive set of tutorials and examples for building software powered by large language… asyncfuncai/deepwiki-open — This platform is an automated documentation and codebase analysis system designed to generate structured wikis,… xiaolincoder/cs-base — CS-Base is a comprehensive educational platform and technical repository designed to support software engineers in…