GOD.EXEA god's-eye view of the AI universe
2026Drag to orbit · right-drag to pan · scroll to zoom
GOD.EXE is booting…

Capability

Evaluate and trace LLM apps

Evaluate and trace LLM apps: 6 AI tools in this graph can do it, including Langfuse, promptfoo, Phoenix, Ragas and DeepEval.

Trace what an LLM app does and score its answers against test sets.

Takes in Text · Data → Gives out Data

Markdown version of https://godooo.ai/en/capability/llm-evaluation. Every page on this site has one: add .md to its address.

Tools that do it · 6

ToolHow wellPricingOpen sourceHow you run it
LangfuseProductmain job *Open source, freeYesWeb · Self-host · API
promptfooCLImain job *Open source, freeYesTerminal · Runs offline · Self-host
PhoenixFrameworkmain job *YesSelf-host · Runs offline · API · Terminal
RagasFrameworkmain job *YesRuns offline · Self-host
DeepEvalFrameworkmain job *YesSelf-host · Terminal · API
Pydantic AIFrameworksupported *YesTerminal · API · Self-host · Runs offline

* proposed by a machine, awaiting calibration

Questions

What is Evaluate and trace LLM apps?

Trace what an LLM app does and score its answers against test sets. Evaluate and trace LLM apps: 6 AI tools in this graph can do it, including Langfuse, promptfoo, Phoenix, Ragas and DeepEval.

Which tools are built for “Evaluate and trace LLM apps”?

Langfuse, promptfoo, Phoenix, Ragas and DeepEval.

Are there free tools for “Evaluate and trace LLM apps”?

Langfuse, promptfoo, Phoenix, Ragas, DeepEval and Pydantic AI.

Which open-source tools can do “Evaluate and trace LLM apps”?

Langfuse, promptfoo, Phoenix, Ragas, DeepEval and Pydantic AI.

Which tools for “Evaluate and trace LLM apps” run offline?

promptfoo, Phoenix, Ragas and Pydantic AI.