AI Agent

Evaluations

Sets of questions with expected answers; a run scores the real engine against them.

Two people in a yoga studio
AI Agent
  1. AI Agent → Evaluations → New set. Add cases: a question and the answer you expect (or import a CSV).
  2. Run. Each question goes through the real engine with your knowledge, rules and question sets; the model scores each answer 1 to 5 against yours, and the run keeps every answer and its cost.
  3. Change something, run again, compare.
TipThe Playground is for trying one conversation; evaluations are for catching regressions after you edit sources or rules.