RANK.AI / INFERENCE

Inference

Throughput, time-to-first-token, goodput, and cost at useful concurrency.

Tables mapped
2
Published
Collecting
Weekly

Tokens per second

Observed throughput by model, accelerator, framework, and concurrency.

Weekly

Cost per million tokens

Serving cost at target output speeds and concurrency.

SOURCE-GATED

No unverified leaderboard is published here.

Rank.ai is resolving entity identity, licensing, historical revisions, and primary evidence before this category is published. The table definitions above are the collection map—not illustrative data.