Chat completions
Evaluating Models
Planned: compare models on your own gold set, with cost and latency.
Planned: not available yet. Tracked as TI-77.
What this will do
A guide to comparing models and agents on your own gold set: run the same prompts through several models, score the answers, and read cost and latency next to quality. No evaluation guide exists yet.