Capability
Evaluate and trace LLM apps
Evaluate and trace LLM apps: 6 AI tools in this graph can do it, including Langfuse, promptfoo, Phoenix, Ragas and DeepEval.
Trace what an LLM app does and score its answers against test sets.
Takes in Text · Data → Gives out Data
Markdown version of https://godooo.ai/en/capability/llm-evaluation. Every page on this site has one: add .md to its address.
Tools that do it · 6
| Tool | How well | Pricing | Open source | How you run it |
|---|---|---|---|---|
| LangfuseProduct | main job * | Open source, free | Yes | Web · Self-host · API |
| promptfooCLI | main job * | Open source, free | Yes | Terminal · Runs offline · Self-host |
| PhoenixFramework | main job * | Yes | Self-host · Runs offline · API · Terminal | |
| RagasFramework | main job * | Yes | Runs offline · Self-host | |
| DeepEvalFramework | main job * | Yes | Self-host · Terminal · API | |
| Pydantic AIFramework | supported * | Yes | Terminal · API · Self-host · Runs offline |
* proposed by a machine, awaiting calibration
Questions
What is Evaluate and trace LLM apps?
Trace what an LLM app does and score its answers against test sets. Evaluate and trace LLM apps: 6 AI tools in this graph can do it, including Langfuse, promptfoo, Phoenix, Ragas and DeepEval.
Which tools are built for “Evaluate and trace LLM apps”?
Langfuse, promptfoo, Phoenix, Ragas and DeepEval.
Are there free tools for “Evaluate and trace LLM apps”?
Langfuse, promptfoo, Phoenix, Ragas, DeepEval and Pydantic AI.
Which open-source tools can do “Evaluate and trace LLM apps”?
Langfuse, promptfoo, Phoenix, Ragas, DeepEval and Pydantic AI.
Which tools for “Evaluate and trace LLM apps” run offline?
promptfoo, Phoenix, Ragas and Pydantic AI.