EXPERT INSIGHTS
LLM leaderboard
Use model benchmarks and specifications to build a shortlist. Then test the candidates on the documents, questions, and tool calls your application will handle.
A public score is a starting point. It does not establish how a model will perform in your workflow.
MMLU Benchmarks
MMLU evaluates knowledge and problem-solving across 57 academic and professional tasks. It is not a benchmark limited to grade-school mathematics.
HumanEval+
Extended version of HumanEval with more complex programming challenges across multiple languages to test code quality.
GPQA Evaluation
Graduate-level expert knowledge evaluation designed to test advanced reasoning in specialized domains.
MT-Bench Analysis
Multi-turn benchmarking that evaluates conversation abilities, reasoning, and instruction following across complex dialogues.
SWE Benchmarks
Software engineering tests including code generation, debugging, and algorithm design to measure programming capabilities.
GSM8K Reasoning
Grade school math word problems requiring multi-step reasoning to evaluate logical thinking and problem-solving capabilities.
Compare the measurements that matter
Choose the models and measurements relevant to your task. Before relying on a score or specification, check the original source, model version, and measurement conditions.
Treat an unavailable value as missing information, not as a zero. Do not infer that a model supports a capability merely because another model from the same provider does.
Provider
Model Name
MMLU Score
Parameters
Context Size
Price
AI21
Jamba Large 1.7
N/A
N/A
256,000
3.5
AI21
Jamba Mini 2
N/A
N/A
256,000
0.25
Amazon
Nova 2 Lite
80.9%
N/A
1,000,000
0.85
Amazon
Nova 2 Omni
N/A
N/A
1,000,000
0.85
Amazon
Nova 2 Pro
N/A
N/A
1,000,000
3.438
Amazon
Nova 2 Sonic
N/A
N/A
1,000,000
0.935
Amazon
Nova Premier
N/A
N/A
1,000,000
5
Amazon
Nova Pro
69.1%
N/A
300,000
1.4
Anthropic
Claude 3.5 Haiku
63.4%
N/A
200,000
1.6
Anthropic
Claude 3.7 Sonnet
80.3%
N/A
200,000
7
Test your shortlist on real work
Choose representative examples and define an acceptable result before running the comparison. Keep the inputs and evaluation rules consistent across models.
Track incorrect answers, missing evidence, invalid tool calls, elapsed time, and the cost of obtaining an accepted result. Include cases where the system should ask for clarification or stop.
Need to evaluate a complete workflow?
Bring a sample input, the expected output, and the systems involved. We can discuss how to test the model and the surrounding workflow together.
MMLU Scores
Compare results for the benchmark version and evaluation setting recorded with each entry.
Reported Token Throughput
Compare measurements only after checking the provider, test conditions, and measurement date.
Published Model Pricing
Check input and output prices separately, including the billing unit and any conditions attached to the rate.
Published Context Limits
Check the exact model and endpoint before using a context limit in your application design.
This animation illustrates example speeds. It is not a live measurement of a model or provider.
1200
t/s
The quick brown fox jumps over the lazy dog. Meanwhile, a clever rabbit watches from nearby bushes, intrigued by the scene unfolding before its eyes. The fox continues its playful pursuit, demonstrating remarkable agility and grace in motion. As the sun sets on the horizon, the forest comes alive with the sounds of nature, creating a symphony of rustling leaves and gentle breezes. The fox pauses, alert to these changes, its ears perked up to catch every subtle noise in the surroundings.
200
t/s
The quick brown fox jumps over the lazy dog. Meanwhile, a clever rabbit watches from nearby bushes, intrigued by the scene unfolding before its eyes. The fox continues its playful pursuit, demonstrating remarkable agility and grace in motion. As the sun sets on the horizon, the forest comes alive with the sounds of nature, creating a symphony of rustling leaves and gentle breezes. The fox pauses, alert to these changes, its ears perked up to catch every subtle noise in the surroundings.
40
t/s
The quick brown fox jumps over the lazy dog. Meanwhile, a clever rabbit watches from nearby bushes, intrigued by the scene unfolding before its eyes. The fox continues its playful pursuit, demonstrating remarkable agility and grace in motion. As the sun sets on the horizon, the forest comes alive with the sounds of nature, creating a symphony of rustling leaves and gentle breezes. The fox pauses, alert to these changes, its ears perked up to catch every subtle noise in the surroundings.
Values reset every 5 seconds to demonstrate different speeds
Compare any two LLM models side by side across different metrics, including MMLU, GPQA, HumanEval, DROP, Context Size, Parameters, Input Price, Output Price, Inference Speed, Throughput, and Latency.
Metric
Provider
MMLU Score
GPQA Score
HumanEval Score
Context Size
Parameters
Input Price
Throughput
Latency
Claude 3.5 Haiku
Anthropic
63.4%
40.8%
75.6%
200,000
N/A
0.8
49.093
0.689
Claude 3.7 Sonnet
Anthropic
80.3%
65.6%
92.1%
200,000
N/A
3
N/A
N/A
