
Evaluations can be aggregated across executions to be used as KPIs

Retrieval Evaluations can be run directly on application traces

Inferences that contain generative records can be fed into evals to produce evaluations for analysis

Adding evaluations on traces can highlight problematic areas that require further analysis

End-to-end evaluation flow

In the above screenshot you can see how poor retrieval directly correlates with hallucinations