-
Notifications
You must be signed in to change notification settings - Fork 0
Tests Analysis and Meta Evaluation
zeroneverload edited this page Apr 21, 2026
·
1 revision
In Tests, you can run:
- single tests (using each system's selected primary model)
- profile runs (multiple prompts from one profile)
- batch matrix runs (multiple systems x multiple profiles)
In Analysis, you can filter by:
- system
- model
- prompt
Use detail views to inspect results and optional metadata.
In Runs, you can:
- filter and inspect run-level data
- delete selected runs
- export selected/system/all run data
The live dashboard shows task progress, current mode, stream events, and a visual speed gauge.
Important:
- the visual gauge is partly for UX/fun and may differ slightly from final stored metrics
- reliable evaluation is based on saved run results in Runs/Analysis
- set filters and create dataset
- choose local judge system + model
- run local meta evaluation
- optionally open cloud exposure and use shared endpoints
Local only, but result quality strongly depends on judge model quality. Very small models usually produce weaker evaluator output.
Supports token-protected or public mode. Use responsibly and close exposure when done.
Meta evaluation is experimental.