Skip to content

Tests Analysis and Meta Evaluation

zeroneverload edited this page Apr 21, 2026 · 1 revision

Tests, Analysis, and Meta Evaluation

Deutsch | English

Tests

In Tests, you can run:

  • single tests (using each system's selected primary model)
  • profile runs (multiple prompts from one profile)
  • batch matrix runs (multiple systems x multiple profiles)

Analysis

In Analysis, you can filter by:

  • system
  • model
  • prompt

Use detail views to inspect results and optional metadata.

Runs

In Runs, you can:

  • filter and inspect run-level data
  • delete selected runs
  • export selected/system/all run data

Live Dashboard

The live dashboard shows task progress, current mode, stream events, and a visual speed gauge.

Important:

  • the visual gauge is partly for UX/fun and may differ slightly from final stored metrics
  • reliable evaluation is based on saved run results in Runs/Analysis

Meta Evaluation (mini tutorial)

  1. set filters and create dataset
  2. choose local judge system + model
  3. run local meta evaluation
  4. optionally open cloud exposure and use shared endpoints

Local judge

Local only, but result quality strongly depends on judge model quality. Very small models usually produce weaker evaluator output.

Cloud exposure

Supports token-protected or public mode. Use responsibly and close exposure when done.

Meta evaluation is experimental.

Clone this wiki locally