For an open source observability platform for llm applications, the strongest matches are langfuse/langfuse (Langfuse is the exact open-source platform requested, offering comprehensive), confident-ai/deepeval (Deepeval is an evaluation and testing framework for LLM) and wandb/client (Weights & Biases provides robust experiment tracking, LLM observability). comet-ml/comet-llm and arize-ai/phoenix round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.
We curate open-source GitHub repositories matching “open source alternatives to langfuse”. Results are ranked by relevance to your query — pick filters below to narrow, or refine with AI.
Langfuse is an open-source observability and evaluation platform designed for language model applications. It provides a centralized system for tracking execution traces, monitoring performance metrics, and managing prompt templates. By capturing hierarchical units of work and telemetry data, the platform enables developers to debug complex application lifecycles and analyze token usage, latency, and model interactions in production environments. The platform distinguishes itself through an integrated evaluation framework that allows for systematic benchmarking and automated scoring of model
Langfuse is the exact open-source platform requested, offering comprehensive LLM tracing, prompt management, cost and token analytics, evaluation frameworks, and self-hosted deployment options.
Deepeval is a framework for testing and evaluating large language model applications. It provides a suite of tools for executing automated regression tests, validating model output quality against defined standards, and tracing the execution of complex agent workflows. By integrating these capabilities into development pipelines, the platform ensures consistent performance and reliability throughout the software lifecycle. The platform distinguishes itself through its focus on programmatic validation and observability. It utilizes secondary language models to score output quality and employs
Deepeval is an evaluation and testing framework for LLM applications that includes execution tracing and observability, though it focuses more on regression testing and validation than being an all-in-one production analytics and prompt management platform like Langfuse.
This project is a collection of utilities designed for machine learning experiment tracking, data versioning, and the observability of large language model applications. It provides a client for recording hyperparameters and metrics during training to visualize performance trends and compare different model versions. The tool includes a model evaluation framework that uses custom scorers and automated judges to assess the quality of generated text outputs. It also provides observability tools to monitor and debug the execution flow and runtime behavior of language model applications. The sys
Weights & Biases provides robust experiment tracking, LLM observability, prompt management, and evaluation capabilities, making it a well-established alternative for monitoring language model applications.
Comet LLM is an observability platform and evaluation framework designed for large language model applications and agentic workflows. It functions as a system for tracing, monitoring, and debugging execution flows while providing tools for prompt optimization and the enforcement of AI safety guardrails. The platform distinguishes itself through a combination of model-based scoring and heuristic metrics to quantify output quality and detect hallucinations. It includes a dedicated prompt and agent optimizer with an interactive playground for refining templates and tool configurations. For retri
Comet LLM is an observability and evaluation platform for LLM applications that supports tracing, debugging, and prompt management, serving as a direct alternative for monitoring AI workflows.
Arize Phoenix is an LLM observability platform and evaluation framework designed to capture execution traces and monitor large language model applications. It serves as a prompt management system for versioning and testing templates, and as a self-hosted AI operations infrastructure for managing telemetry and experiments. The platform differentiates itself through a specialized embedding visualization tool used to detect data drift and optimize vector search. It provides a comprehensive evaluation suite that utilizes judge-based evaluators and ground-truth datasets to score model outputs, and
Arize Phoenix is an LLM observability platform that delivers execution tracing, prompt management, cost analytics, and evaluation tools in a self-hostable package directly aligned with your search.
Opik is an observability and evaluation platform designed for generative AI applications and agentic workflows. It provides a centralized environment for tracing execution flows, managing prompt templates, and monitoring production performance, allowing teams to gain visibility into complex model interactions and tool usage without requiring manual application code changes. The platform distinguishes itself through its integrated approach to the AI development lifecycle, combining distributed trace instrumentation with automated evaluation frameworks. It supports model-as-a-judge scoring, syn
Opik is an open-source observability and evaluation platform for LLM applications that supports tracing, prompt management, and analytics, making it a fitting alternative for monitoring generative AI workflows.
MLflow is an open-source MLOps and LLMops platform supporting model evaluation, prompt engineering, and trace monitoring, though it is broader than a dedicated Langfuse alternative.
Agenta is a Prompt Ops lifecycle manager and prompt management platform that decouples prompt engineering from application code. It serves as a centralized system for developing, versioning, and deploying prompt templates and model configurations across different environments. The platform functions as an AI agent orchestrator with a visual interface for building agent workflows and connecting models to external tools. It further acts as an evaluation framework and observability tool, utilizing OpenTelemetry to capture execution traces, monitor latency, and track token costs. The system cove
Agenta provides prompt management, LLM tracing, cost monitoring, and evaluation capabilities, making it a relevant alternative for LLM observability though with a strong emphasis on prompt operations and agent workflows.
OpenLLMetry is an OpenTelemetry-based observability framework and instrumentation library for generative AI applications. It provides toolsets for tracing and monitoring large language model workflows, capturing telemetry from model providers, agent frameworks, and vector databases using standardized semantic conventions. The project distinguishes itself by providing a specialized evaluation and experimentation suite that associates user feedback and prompt version hashes with specific execution traces. It includes a system for tracking model reasoning paths and enforcing security guardrails
OpenLLMetry is an OpenTelemetry-based observability library and tracing framework for LLM applications that covers telemetry, token tracking, and evaluation, though it functions primarily as an instrumentation library rather than a full standalone analytics platform.
HyperDX is an OpenTelemetry observability platform that provides centralized log management, distributed tracing, and a self-hosted monitoring stack. It functions as a unified system for collecting, indexing, and visualizing logs, metrics, and traces from cloud and container environments. The platform distinguishes itself with specialized tooling for large language model monitoring and session replay, allowing user interactions in the browser to be linked to backend telemetry. It employs schema-less JSON parsing to index structured logs dynamically and uses source maps to resolve minified sta
HyperDX is a self-hostable OpenTelemetry observability platform that includes specialized tooling for LLM performance monitoring and tracing, serving as a viable alternative for application monitoring though lacking dedicated prompt management features.
Evidently is an AI observability platform and evaluation framework designed to quantify the performance of machine learning models and large language models. It functions as a monitoring tool for detecting data drift and quality degradation in tabular datasets, while providing a specialized analyzer for the faithfulness and correctness of retrieval augmented generation systems. The project distinguishes itself through an evaluation framework that utilizes judge models and custom rubrics to score language model outputs. It includes tools for iterative prompt optimization and the generation of
Evidently is an open-source ML and LLM observability and evaluation framework that supports monitoring and testing, though it focuses more heavily on model quality and evaluation than on being an all-in-one Langfuse alternative for production tracing and prompt management.
This project is a self-hosted AI monitoring stack that functions as an LLM observability platform, AI evaluation framework, and OpenTelemetry trace analyzer. It is designed to capture and analyze LLM traces, sessions, and telemetry to monitor AI agent performance. The platform distinguishes itself as a Model Context Protocol server, exposing workspace functions as tools for AI coding agents. It enables the conversion of failing production traces into test datasets for regression testing and utilizes semantic-based session clustering to discover emerging user behavior patterns. The system cov
This project is a self-hostable LLM observability platform and evaluation framework that supports tracing, cost analysis, and OpenTelemetry integration, though it focuses heavily on agent workflows and Model Context Protocol features compared to standard platforms.
Helicone is an AI gateway and observability platform designed to intercept, manage, and monitor interactions with large language models. By acting as a reverse-proxy, it provides a centralized layer for routing requests across multiple AI providers, allowing developers to maintain consistent application logic while gaining deep visibility into model performance, usage, and costs. The platform distinguishes itself through a robust suite of traffic management and prompt engineering tools. It enables policy-driven control, including automatic failover between providers, rate limiting, and edge-b
Helicone is an AI gateway and observability platform that provides robust tracing, token analytics, and prompt management features, making it a strong alternative for monitoring LLM applications even though it operates primarily as a reverse-proxy.
Wandb is a centralized platform for machine learning experiment tracking, model registry management, and workflow orchestration. It provides a comprehensive suite of tools for logging, visualizing, and versioning training metrics, model artifacts, and hyperparameter sweeps to ensure reproducibility across development cycles. The platform also functions as an observability tool for large language model applications, enabling the tracing of execution steps, token usage, and reasoning processes. The project distinguishes itself through its event-driven automation capabilities, which allow users
Wandb is a comprehensive machine learning experiment tracking and MLOps platform that includes LLM tracing and analytics features, serving as an alternative though it focuses more broadly on traditional model training than dedicated prompt management.
The platform for LLM evaluations and AI agent testing
Langwatch is a self-hostable LLM operations platform that covers evaluations, testing, and observability, serving as a solid alternative for monitoring and analyzing AI application workflows.
Coze-loop is an optimization platform and orchestration management suite for large language model agents. It functions as a comprehensive environment for the development, debugging, evaluation, and monitoring of AI agent performance. The project provides a dedicated prompt engineering playground for real-time iteration and validation of model responses. It includes an evaluation framework that runs automated assessments against datasets to generate performance metrics and verify output accuracy. The system covers observability through real-time execution tracing and historical analysis of ag
Coze-loop provides LLM tracing, prompt playgrounds, and agent evaluation features tailored for AI applications, fitting well as an observability and optimization platform even though its primary focus is on agent orchestration.
| Repository | Stars | Language | License | Last push |
|---|---|---|---|---|
| langfuse/langfuse | 29.2K | TypeScript | NOASSERTION | |
| confident-ai/deepeval | 13.7K | Python | apache-2.0 | |
| wandb/client | 11.1K | Python | MIT | |
| comet-ml/comet-llm | 19.7K | Python | Apache-2.0 | |
| arize-ai/phoenix | 8.6K | Jupyter Notebook | other | |
| comet-ml/opik | 17.8K | Python | apache-2.0 | |
| mlflow/mlflow | 26.6K | Python | Apache-2.0 | |
| agenta-ai/agenta | 3.9K | TypeScript | other | |
| traceloop/openllmetry | 7.2K | Python | Apache-2.0 | |
| hyperdxio/hyperdx | 9.3K | TypeScript | mit |