Cognaptus Model Evidence Monitor

Model Comparisons Monitor

A source-preserving dashboard for comparing model-specific benchmark signals with explicit input-price data, provenance, scope, and confidence.

Updated 2026-09-03T11:44:12+0800 Sources and scope are retained per metric

Metric Views

7

Source Boundaries

Preserved

Price Basis

Declared per metric

Missingness

Explicit

How To Read This

This page is meant to answer a practical question:

How does model input price relate to a selected, explicitly sourced evaluation metric?

The current legacy view uses two sources:

  1. Arena leaderboard performance and vote data as an external benchmark signal
  2. Cognaptus pricing data from OpenRouter as a practical cost surface

New datasets retain their own metric definitions, scope, timestamps, price basis, sample sizes, and missingness. Arena preference, Artificial Analysis metrics, and task-level benchmarks are displayed separately rather than normalized into a universal score.

Model Metric Versus Input Price

Select one source-specific metric at a time. The chart does not combine Arena, Artificial Analysis, or task-level benchmarks into a common score.

Matched record Pareto frontier Low-confidence match

Missing or unplottable data

0

Records with explicit missingness, or a non-positive cost that cannot appear on a logarithmic axis.

    Low-confidence matches

    0

    Visible on the chart with a dashed ring; inspect provenance before relying on the join.

      Secondary Signal: Arena Chat Preference

      Arena reflects comparative user preference in open-ended chat. It is retained as a separate, narrower signal and is not included in the metric selector or interpreted as general model capability.

      Snapshot: 2026-08-29 20:39:59 +0800. Live Arena leaderboard refresh succeeded.

      View the Arena text leaderboard