awesome-repositories.com
博客
awesome-repositories.com

通过 AI 驱动的搜索,发现最优秀的开源仓库。

探索精选搜索开源替代品自托管软件博客网站地图
项目关于排名机制媒体报道MCP 服务器
法律隐私政策服务条款
© 2026 Bringes Technology SRL·VAT RO45896025·hello@awesome-repositories.com
·
langfuse avatar

langfuse/langfuse

0
View on GitHub↗
29,190 星标·3,027 分支·TypeScript·12 次浏览langfuse.com/docs↗

Langfuse

Langfuse is an open-source observability and evaluation platform designed for language model applications. It provides a centralized system for tracking execution traces, monitoring performance metrics, and managing prompt templates. By capturing hierarchical units of work and telemetry data, the platform enables developers to debug complex application lifecycles and analyze token usage, latency, and model interactions in production environments.

The platform distinguishes itself through an integrated evaluation framework that allows for systematic benchmarking and automated scoring of model outputs. Users can perform comparative experimentation by running multiple prompt or model versions side-by-side, and convert production traces into versioned test datasets to validate performance against ground truth. A dedicated prompt management system further decouples logic from application code, offering a playground for refinement and dynamic fetching of versioned templates.

Beyond core observability, the project supports a comprehensive suite of administrative and operational tools, including organizational access controls, identity provider integration, and automated workflow triggers. It is built for flexible deployment, supporting containerized orchestration in private, cloud, or Kubernetes-based environments to ensure data control and high-availability scaling.

The platform is designed for self-hosting and provides infrastructure-as-code templates to facilitate consistent environment setup. It integrates with standard observability ecosystems through open telemetry support and offers programmatic interfaces for headless management and automated deployment workflows.

Features

  • LLM Observability - Monitors and debugs language model applications by tracking prompts, completions, latency, and token usage.
  • AI Observability and Evaluation - Provides systematic experiments and automated scoring against datasets to validate model performance and output quality.
  • Prompt Registries - Decouples prompt logic from application code by serving versioned templates through a managed interface for dynamic retrieval.
  • Automated Trace Evaluation - Executes automated scoring and custom logic against captured traces to validate output quality against defined performance benchmarks.
  • Observability Instrumentation - Captures hierarchical units of work and execution flows from model operations to visualize performance and cost across distributed components.

AI 搜索

探索更多 awesome 仓库

用简单的语言描述您的需求 —— AI 将根据相关性为您从数千个精选开源项目中进行排序。

Start searching with AI
  • Prompt Management Workflows - Centrally manages, versions, and tests prompt templates to ensure consistent model behavior across environments.
  • Prompt Playgrounds - Provides an interactive environment to refine prompts and model parameters before deploying them to production applications.
  • Distributed Tracing Instrumentation - Captures nested execution flows by linking parent-child relationships across distributed components to visualize complex application lifecycles.
  • Trace-to-Dataset Converters - Converts production traces into test items to build evaluation sets based on real-world application performance.
  • Experiment Tracking - Enables the execution of versioned experiments to evaluate historical data snapshots and track performance changes over time.
  • AI Evaluation Frameworks - Provides a system for running systematic experiments, benchmarking model outputs, and automating quality scoring.
  • Prompt Templates - Retrieves production-labeled prompt templates from a central store to separate prompt logic from application code.
  • Self-Hosted Infrastructure - Enables self-hosting and scaling of observability platforms within private or cloud environments for full data control.
  • Monitoring and Observability - Streams internal performance metrics and trace data to external observability tools while optimizing transmission through batching and sampling.
  • Human-in-the-Loop Systems - Runs automated and human-in-the-loop assessments using custom judges to validate model performance.
  • AI and Machine Learning - Platform for LLM engineering, tracing, and evaluation.
  • AI Observability and Evaluation - Platform for LLM observability, evals, and prompt management.
  • Application Development - Platform for tracing, evaluation, and prompt management.
  • Development Frameworks - Observability platform for debugging and monitoring LLM applications.
  • Evaluation and Observability - Platform for tracing, evals, and prompt management.
  • Evaluation Frameworks - Tracks LLM metrics, observability, and manages prompts.
  • Generative AI - Listed in the “Generative AI” section of the Free For Dev awesome list.
  • LLM Development Frameworks - Engineering platform for LLM observability, metrics, and prompt management.
  • LLM Evaluation Frameworks - Observability, metrics, and evaluation management for LLM applications.
  • LLM Evaluation Tools - Observability platform for tracing, prompt management, and human annotation.
  • LLM Frameworks and Libraries - Engineering platform for LLM observability, metrics, and prompt management.
  • Model Evaluation and Benchmarking - Observability and analytics solution for LLM-based applications.
  • Observability And Analytics - Open-source analytics platform for monitoring LLM application performance.
  • Observability and Evaluation - Integrated platform for tracing, evaluation, and prompt management.
  • LLM Development Frameworks - Platform for debugging, analyzing, and iterating on LLM applications.
  • Monitoring and Observability - Observability platform for debugging and analyzing LLM applications.
  • Testing and Observability - Platform for LLM observability, metrics, and evaluation.
  • 安全与隐私 - Listed in the “Security And Privacy” section of the Llm Course awesome list.
  • Private Infrastructure Hosting - Supports hosting within isolated network environments to maintain data control and security.
  • API Request Authentication - Validates client identity using keys to secure access to platform resources.
  • Identity Provider Integrations - Connects to external identity providers to manage user access and enforce single sign-on protocols across the organization.
  • Experimentation Frameworks - Executes multiple prompt or model versions side-by-side to measure the impact of configuration changes.
  • Organization Management - Allows for the creation and administration of organizational structures and user access permissions through external interface calls.
  • Kubernetes Orchestration - Orchestrates application containers and dependencies within cluster environments using standardized packaging.
  • Data Encryption - Secures stored information using encryption at rest to protect configuration and application data from unauthorized access.
  • Data Schema Validation - Checks dataset items against defined structures to ensure input and output consistency across team contributions.
  • Execution Metadata - Enriches execution logs with custom attributes and user identifiers to enable granular filtering and performance analysis across sessions.
  • Test Data Management - Organizes collections of inputs and expected outputs into versioned groups to facilitate consistent application testing and evaluation.
  • Workflow Automation Triggers - Triggers actions based on performance events and provides interfaces for external agents to query data and execute tasks.
  • Cloud Infrastructure Deployment - Automates the provisioning of compute and storage resources on major cloud platforms.
  • Workload Scheduling and Scaling - Supports high-availability deployments using container orchestration to manage large-scale observability data and traffic.
  • Identity Provisioning - Synchronizes user accounts and permissions from external identity systems to ensure consistent access management.
  • Trace Metadata - Attaches custom metadata, user identifiers, and session tags to execution traces to provide context for debugging and analysis.
  • Distributed Tracing - Ingests execution traces and telemetry data to monitor distributed AI application pipelines and agent behavior.
  • OpenTelemetry Exporters - Integrates with standard observability ecosystems to collect and export telemetry data from any language or runtime.
  • Telemetry Ingestion - Buffers high-volume telemetry data in background queues to maintain host application responsiveness during intensive monitoring tasks.
  • Star 历史

    langfuse/langfuse 的 Star 历史图表langfuse/langfuse 的 Star 历史图表

    常见问题解答

    langfuse/langfuse 是做什么的?

    Langfuse is an open-source observability and evaluation platform designed for language model applications. It provides a centralized system for tracking execution traces, monitoring performance metrics, and managing prompt templates. By capturing hierarchical units of work and telemetry data, the platform enables developers to debug complex application lifecycles and analyze token usage, latency, and model interactions in production environments.

    langfuse/langfuse 的主要功能有哪些?

    langfuse/langfuse 的主要功能包括:LLM Observability, AI Observability and Evaluation, Prompt Registries, Automated Trace Evaluation, Observability Instrumentation, Prompt Management Workflows, Prompt Playgrounds, Distributed Tracing Instrumentation。

    langfuse/langfuse 有哪些开源替代品?

    langfuse/langfuse 的开源替代品包括: comet-ml/opik — Opik is an observability and evaluation platform designed for generative AI applications and agentic workflows. It… helicone/helicone — Helicone is an AI gateway and observability platform designed to intercept, manage, and monitor interactions with… agenta-ai/agenta — Agenta is a Prompt Ops lifecycle manager and prompt management platform that decouples prompt engineering from… arize-ai/phoenix — Arize Phoenix is an LLM observability platform and evaluation framework designed to capture execution traces and… mastra-ai/mastra — Mastra is an orchestration framework designed for building, deploying, and managing autonomous AI agents and… latitude-dev/latitude-llm — This project is a self-hosted AI monitoring stack that functions as an LLM observability platform, AI evaluation…

    Langfuse 的开源替代方案

    相似的开源项目,按与 Langfuse 的功能重合度排序。
    • comet-ml/opikcomet-ml 的头像

      comet-ml/opik

      17,787在 GitHub 上查看↗

      Opik is an observability and evaluation platform designed for generative AI applications and agentic workflows. It provides a centralized environment for tracing execution flows, managing prompt templates, and monitoring production performance, allowing teams to gain visibility into complex model interactions and tool usage without requiring manual application code changes. The platform distinguishes itself through its integrated approach to the AI development lifecycle, combining distributed trace instrumentation with automated evaluation frameworks. It supports model-as-a-judge scoring, syn

      Pythonevaluationhacktoberfesthacktoberfest2025
      在 GitHub 上查看↗17,787
    • helicone/heliconeHelicone 的头像

      Helicone/helicone

      5,830在 GitHub 上查看↗

      Helicone is an AI gateway and observability platform designed to intercept, manage, and monitor interactions with large language models. By acting as a reverse-proxy, it provides a centralized layer for routing requests across multiple AI providers, allowing developers to maintain consistent application logic while gaining deep visibility into model performance, usage, and costs. The platform distinguishes itself through a robust suite of traffic management and prompt engineering tools. It enables policy-driven control, including automatic failover between providers, rate limiting, and edge-b

      TypeScript
      在 GitHub 上查看↗5,830
    • agenta-ai/agentaAgenta-AI 的头像

      Agenta-AI/agenta

      3,860在 GitHub 上查看↗

      Agenta is a Prompt Ops lifecycle manager and prompt management platform that decouples prompt engineering from application code. It serves as a centralized system for developing, versioning, and deploying prompt templates and model configurations across different environments. The platform functions as an AI agent orchestrator with a visual interface for building agent workflows and connecting models to external tools. It further acts as an evaluation framework and observability tool, utilizing OpenTelemetry to capture execution traces, monitor latency, and track token costs. The system cove

      TypeScriptagentsevaluationllm-as-a-judge
      在 GitHub 上查看↗3,860
    • arize-ai/phoenixArize-ai 的头像

      Arize-ai/phoenix

      8,605在 GitHub 上查看↗

      Arize Phoenix is an LLM observability platform and evaluation framework designed to capture execution traces and monitor large language model applications. It serves as a prompt management system for versioning and testing templates, and as a self-hosted AI operations infrastructure for managing telemetry and experiments. The platform differentiates itself through a specialized embedding visualization tool used to detect data drift and optimize vector search. It provides a comprehensive evaluation suite that utilizes judge-based evaluators and ground-truth datasets to score model outputs, and

      Jupyter Notebookagentsai-monitoringai-observability
      在 GitHub 上查看↗8,605
    查看 Langfuse 的所有 30 个替代方案→