For an open source platform for debugging LLM applications, the strongest matches are arize-ai/phoenix (Arize Phoenix is an open-source LLM observability and evaluation), traceloop/openllmetry (OpenLLMetry is an OpenTelemetry-based instrumentation and tracing framework tailored) and comet-ml/opik (Opik is a self-hostable LLM observability and evaluation platform). lmnr-ai/lmnr and langfuse/langfuse round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.
We curate open-source GitHub repositories matching “open source alternatives to langsmith”. Results are ranked by relevance to your query — pick filters below to narrow, or refine with AI.
Arize Phoenix is an LLM observability platform and evaluation framework designed to capture execution traces and monitor large language model applications. It serves as a prompt management system for versioning and testing templates, and as a self-hosted AI operations infrastructure for managing telemetry and experiments. The platform differentiates itself through a specialized embedding visualization tool used to detect data drift and optimize vector search. It provides a comprehensive evaluation suite that utilizes judge-based evaluators and ground-truth datasets to score model outputs, and
Arize Phoenix is an open-source LLM observability and evaluation platform that delivers tracing, prompt debugging, dataset evaluation, and self-hostable monitoring for AI workflows.
OpenLLMetry is an OpenTelemetry-based observability framework and instrumentation library for generative AI applications. It provides toolsets for tracing and monitoring large language model workflows, capturing telemetry from model providers, agent frameworks, and vector databases using standardized semantic conventions. The project distinguishes itself by providing a specialized evaluation and experimentation suite that associates user feedback and prompt version hashes with specific execution traces. It includes a system for tracking model reasoning paths and enforcing security guardrails
OpenLLMetry is an OpenTelemetry-based instrumentation and tracing framework tailored for generative AI applications, though as a library rather than a standalone platform it requires integration into your own application setup.
Opik is an observability and evaluation platform designed for generative AI applications and agentic workflows. It provides a centralized environment for tracing execution flows, managing prompt templates, and monitoring production performance, allowing teams to gain visibility into complex model interactions and tool usage without requiring manual application code changes. The platform distinguishes itself through its integrated approach to the AI development lifecycle, combining distributed trace instrumentation with automated evaluation frameworks. It supports model-as-a-judge scoring, syn
Opik is a self-hostable LLM observability and evaluation platform that provides tracing, prompt debugging, dataset evaluation, and performance monitoring for AI applications and agent workflows.
Lmnr is an LLM observability platform and evaluation framework designed for tracing, logging, and monitoring language model executions. It provides the tools necessary to debug agent behavior, analyze performance, and identify failure patterns in AI agents. The platform differentiates itself through a trace-to-dataset pipeline that converts production logs into labeled test sets for regression testing. It includes a prompt-variant replay engine to compare different prompts or models side-by-side and a state-cached debugging system to replay agent loops without restarting the process. The sys
Lmnr is an open-source LLM observability platform that provides tracing, prompt debugging, dataset evaluation, and self-hostable deployment for monitoring agent workflows.
Langfuse is an open-source observability and evaluation platform designed for language model applications. It provides a centralized system for tracking execution traces, monitoring performance metrics, and managing prompt templates. By capturing hierarchical units of work and telemetry data, the platform enables developers to debug complex application lifecycles and analyze token usage, latency, and model interactions in production environments. The platform distinguishes itself through an integrated evaluation framework that allows for systematic benchmarking and automated scoring of model
Langfuse is an open-source LLM observability and evaluation platform that provides tracing, prompt debugging, cost and latency monitoring, self-hosting options, and SDK integration for monitoring agent workflows.
Comet LLM is an observability platform and evaluation framework designed for large language model applications and agentic workflows. It functions as a system for tracing, monitoring, and debugging execution flows while providing tools for prompt optimization and the enforcement of AI safety guardrails. The platform distinguishes itself through a combination of model-based scoring and heuristic metrics to quantify output quality and detect hallucinations. It includes a dedicated prompt and agent optimizer with an interactive playground for refining templates and tool configurations. For retri
Comet LLM is an observability and evaluation platform tailored for debugging execution flows, monitoring performance, and optimizing prompts in language model applications, fitting the category well despite omitting explicit OpenTelemetry mention in the provided details.
MLflow is an open-source machine learning lifecycle platform that includes robust LLM tracing, evaluation tools, and prompt debugging capabilities, though its roots as a general MLOps suite make it slightly broader than a dedicated agent observability platform.
Agenta is a Prompt Ops lifecycle manager and prompt management platform that decouples prompt engineering from application code. It serves as a centralized system for developing, versioning, and deploying prompt templates and model configurations across different environments. The platform functions as an AI agent orchestrator with a visual interface for building agent workflows and connecting models to external tools. It further acts as an evaluation framework and observability tool, utilizing OpenTelemetry to capture execution traces, monitor latency, and track token costs. The system cove
Agenta is an open-source prompt management and LLM observability platform that supports tracing, evaluation, and monitoring, making it a strong fit despite leaning heavily toward prompt engineering and workflow orchestration.
HyperDX is an OpenTelemetry observability platform that provides centralized log management, distributed tracing, and a self-hosted monitoring stack. It functions as a unified system for collecting, indexing, and visualizing logs, metrics, and traces from cloud and container environments. The platform distinguishes itself with specialized tooling for large language model monitoring and session replay, allowing user interactions in the browser to be linked to backend telemetry. It employs schema-less JSON parsing to index structured logs dynamically and uses source maps to resolve minified sta
HyperDX is an OpenTelemetry-based observability platform that includes specialized tooling for LLM performance monitoring and session tracing, though it functions primarily as a broad application monitoring stack rather than a dedicated LLM evaluation platform.
Evidently is an AI observability platform and evaluation framework designed to quantify the performance of machine learning models and large language models. It functions as a monitoring tool for detecting data drift and quality degradation in tabular datasets, while providing a specialized analyzer for the faithfulness and correctness of retrieval augmented generation systems. The project distinguishes itself through an evaluation framework that utilizes judge models and custom rubrics to score language model outputs. It includes tools for iterative prompt optimization and the generation of
Evidently is an AI observability and evaluation platform that supports LLM testing, prompt debugging, and model monitoring, though it is more heavily focused on evaluation and data drift than a full end-to-end tracing suite.
Aim is an open-source platform for logging, visualizing, and comparing machine learning training runs and LLM traces. It provides a remote tracking server and a comparison UI, functioning as an ML experiment tracker, AI workflow logger, and LLM trace recorder that captures prompts, generations, and tool calls from AI applications. The platform distinguishes itself through a run-based data model with local SQLite storage, real-time metric streaming, and a plugin-based explorer system that supports specialized visual analysis of metrics, images, audio, and text. It offers a Python SDK with cont
Aim is an open-source machine learning experiment tracker and LLM tracing platform that captures prompts and generations, though it focuses more on traditional ML experiment management and lacks comprehensive evaluation and OpenTelemetry features.
Giskard is an evaluation framework, testing library, and quality monitoring system for large language models and AI agents. It serves as a toolkit for quantifying model performance and reliability, providing specialized capabilities for validating retrieval-augmented generation pipelines. The project distinguishes itself through an automated red teaming tool and security scanner designed to identify vulnerabilities, prompt injections, and safety risks. It utilizes adversarial probing and synthetic edge case generation to quantify model robustness and detect information disclosure. The platfo
Giskard provides model evaluation, testing, and quality monitoring for AI agents and language models, fitting the testing and evaluation aspects of the category while leaning less on real-time production tracing.
The platform for LLM evaluations and AI agent testing
LangWatch is a self-hostable open-source platform providing comprehensive LLM tracing, prompt debugging, and dataset evaluation tools with multi-language SDK support.
Helicone is an AI gateway and observability platform designed to intercept, manage, and monitor interactions with large language models. By acting as a reverse-proxy, it provides a centralized layer for routing requests across multiple AI providers, allowing developers to maintain consistent application logic while gaining deep visibility into model performance, usage, and costs. The platform distinguishes itself through a robust suite of traffic management and prompt engineering tools. It enables policy-driven control, including automatic failover between providers, rate limiting, and edge-b
Helicone is an open-source AI gateway and observability platform that handles LLM monitoring, cost tracking, and request routing, though it is primarily designed around a reverse-proxy gateway architecture rather than a dedicated evaluation and tracing suite.
| Repository | Stars | Language | License | Last push |
|---|---|---|---|---|
| arize-ai/phoenix | 8.6K | Jupyter Notebook | other | |
| traceloop/openllmetry | 7.2K | Python | Apache-2.0 | |
| comet-ml/opik | 17.8K | Python | apache-2.0 | |
| lmnr-ai/lmnr | 2.6K | TypeScript | apache-2.0 | |
| langfuse/langfuse | 29.2K | TypeScript | NOASSERTION | |
| comet-ml/comet-llm | 19.7K | Python | Apache-2.0 | |
| mlflow/mlflow | 26.6K | Python | Apache-2.0 | |
| agenta-ai/agenta | 3.9K | TypeScript | other | |
| hyperdxio/hyperdx | 9.3K | TypeScript | mit | |
| evidentlyai/evidently | 7.1K | Jupyter Notebook | apache-2.0 |