For a tracing tool for debugging AI agents, the strongest matches are different-ai/openwork (Openwork is an LLM agent orchestration platform with built-in), confident-ai/deepeval (Deepeval is an LLM evaluation framework that includes tracing) and langchain-ai/deepagents (Deepagents provides a full observability suite for capturing execution). comet-ml/opik and coze-dev/coze-loop round out the shortlist. Each is ranked by relevance to your query, popularity and recent activity.
Tools for monitoring and inspecting the execution flow of large language model agent processes.
Openwork is an LLM agent orchestration platform and cross-platform desktop application designed for building and running automated workflows. It serves as a local AI agent host and session manager, allowing users to connect local project folders to various large language models and remote cloud workers. The project distinguishes itself through a local-first execution model that enables agents to process files directly on a host machine. It implements human-in-the-loop permissioning to intercept agent resource requests, requiring explicit user approval before accessing specific local system fi
Openwork is an LLM agent orchestration platform with built-in execution timeline visualization, action auditing, and debug logging, enabling step-by-step inspection of agent workflows and error tracing, directly matching the need for an agent execution debugging tool.
Deepeval is a framework for testing and evaluating large language model applications. It provides a suite of tools for executing automated regression tests, validating model output quality against defined standards, and tracing the execution of complex agent workflows. By integrating these capabilities into development pipelines, the platform ensures consistent performance and reliability throughout the software lifecycle. The platform distinguishes itself through its focus on programmatic validation and observability. It utilizes secondary language models to score output quality and employs
Deepeval is an LLM evaluation framework that includes tracing of agent workflow executions, allowing inspection of inputs, outputs, and errors, which fits your need for an agent execution tracing and debugging tool, though it emphasizes evaluation over a dedicated step-by-step debugger.
Deepagents is an LLM agent orchestration platform and stateful application server designed for deploying and managing AI agents built with computational graphs. It provides a containerized runtime environment that handles agent execution, state persistence, and the versioning of AI assistants. The platform distinguishes itself through deep integration with the Model Context Protocol, allowing agents to function as servers that expose tools and capabilities to external clients. It features a sophisticated observability suite for capturing execution traces, performing LLM-based evaluations agai
Deepagents provides a full observability suite for capturing execution traces of LLM agents, enabling step-by-step inspection of actions, inputs, outputs, and errors—exactly what you need for debugging agent workflows.
Opik is an observability and evaluation platform designed for generative AI applications and agentic workflows. It provides a centralized environment for tracing execution flows, managing prompt templates, and monitoring production performance, allowing teams to gain visibility into complex model interactions and tool usage without requiring manual application code changes. The platform distinguishes itself through its integrated approach to the AI development lifecycle, combining distributed trace instrumentation with automated evaluation frameworks. It supports model-as-a-judge scoring, syn
Opik is an observability and evaluation platform purpose-built for tracing and debugging agentic workflows, offering step-by-step execution views, input/output inspection, error tracing, and execution span hierarchies that directly match the need to inspect each action and debug agent behavior.
Coze-loop is an optimization platform and orchestration management suite for large language model agents. It functions as a comprehensive environment for the development, debugging, evaluation, and monitoring of AI agent performance. The project provides a dedicated prompt engineering playground for real-time iteration and validation of model responses. It includes an evaluation framework that runs automated assessments against datasets to generate performance metrics and verify output accuracy. The system covers observability through real-time execution tracing and historical analysis of ag
Coze-loop is an optimization and orchestration platform for LLM agents that includes real-time execution tracing, historical analysis, and debug monitoring, directly providing the step-by-step execution inspection, I/O per step, and error tracing this search requires.
Langfuse is an open-source observability and evaluation platform designed for language model applications. It provides a centralized system for tracking execution traces, monitoring performance metrics, and managing prompt templates. By capturing hierarchical units of work and telemetry data, the platform enables developers to debug complex application lifecycles and analyze token usage, latency, and model interactions in production environments. The platform distinguishes itself through an integrated evaluation framework that allows for systematic benchmarking and automated scoring of model
Langfuse is an open-source observability platform that captures hierarchical execution traces of LLM applications, including agent steps, and provides detailed per-step input/output and error inspection, which directly matches the need for an agent execution tracing and debugging tool.
Arize Phoenix is an LLM observability platform and evaluation framework designed to capture execution traces and monitor large language model applications. It serves as a prompt management system for versioning and testing templates, and as a self-hosted AI operations infrastructure for managing telemetry and experiments. The platform differentiates itself through a specialized embedding visualization tool used to detect data drift and optimize vector search. It provides a comprehensive evaluation suite that utilizes judge-based evaluators and ground-truth datasets to score model outputs, and
Arize Phoenix is an LLM observability platform that captures execution traces of LLM applications, including agent step sequences, with inspection of inputs, outputs, and errors per step, timeline visualization, and filtering capabilities.