Same brief, same brand layer, measured cost and latency, deterministic brand gates, and a human 10-dimension rubric you can review and correct live.
Primary: best run per model (phase-1 skill preferred when unreviewed). Secondary: best run per skill. Quality is human-corrected; cost/time never silently buy rank.
Each cell is the canonical run for that pair. Click through to review.