Silver trophy cup sitting on a podium, with a large orange-and-blue circular backdrop and small dot-and-line accents around it.

Innodata's LLM Scoreboard

AI Model Benchmark Rankings

Innodata’s LLM Scoreboard ranks leading large language models (LLMs) against expert datasets developed by Innodata’s data science department, Innodata Labs. Our rigorous methodology ensures fair and unbiased assessments, helping enterprises identify the safest and most capable AI models. 

 

These datasets, vetted by Innodata’s leading generative AI domain experts, cover key safety and risk areas, including: 

And More...

Ranking Today's Leading LLMs:

Mistral-Nemo-Instruct-2407

Meta-Llama-3-8B-Instruct

OLMo-2-1124-7B-Instruct

Gemma-2-9b-it

Deepseek-llm-7b-chat

Explore the Latest Rankings
Silver trophy with a star on the front stands centered, with upward-pointing arrows in the background on a light surface.
Leaderboard table ranking AI models on Rt2-pii-masking, showing score and deviation columns, on a dark panel over green background.
Leaderboard table ranking AI models on Rt-gsm8k-gaia, showing scores and deviation columns, with OpenAI-GPT-4o listed at top.
Factuality table ranks AI models on Rt-frank dataset, showing scores and deviations in columns on a dark panel background.
Toxicity Translation leaderboard table listing model, dataset, score, and deviation, with six ranked rows on a dark panel.
Model ranking table shows OpenAI-GPT-4o top score 0.69, followed by Meta-Llama and Deepseek; columns include dataset, score, deviation.
Bias table listing AI models with Rt-inod-bias dataset, showing score and deviation columns, on a dark panel background.
Benchmark table ranking AI models with columns Model, Dataset, Score, and Deviation, shown on a dark panel over a green background.
Safety hallucination leaderboard table listing AI models with dataset Rt2-halueval, scores, and deviation values, on a green background.
Table titled “Jailbreaking Single-Turn Q&A” lists model scores and deviations on Rt2-easyjail-alpaca, shown on a green background.
Table titled “Jailbreaking Results (…?)” listing AI models with dataset, score, and deviation columns on a dark panel over green background.

Models benchmarked as of 2/06/2025

Interested in How Your LLM Compares?

Benchmark your models today using Innodata’s publicly available benchmarking tool.