| Name | Description |
|---|---|
| Artificial Analysis | An independent site that compares AI models and API providers across quality, price, output speed, and latency. Also offers a proprietary intelligence index and hallucination rate evaluation. |
| LLM Stats | A platform that compares 289+ AI models on benchmark performance, pricing, and throughput using independent, reproducible methods. Covers language models as well as image, video, and audio generation. |
| Scale Labs Leaderboard | Scale's AI model evaluation platform aggregating expert-driven benchmarks across multiple categories including agentic capabilities, frontier reasoning, and safety (SWE-Bench Pro, Humanity's Last Exam, etc.). |
| Vellum LLM Leaderboard | A leaderboard comparing leading LLMs across public benchmarks such as GPQA Diamond, AIME, SWE-Bench Verified, and ARC-AGI, along with throughput and cost efficiency. |
| Nejumi Leaderboard 4 | A Japanese LLM evaluation leaderboard operated by Weights & Biases. Comprehensively assesses Japanese generation accuracy, practical application development capabilities, and safety to support LLM selection for the Japanese market. |
| Name | Description |
|---|---|
| LM Arena | A community-driven platform where users anonymously compare AI model outputs and vote to form Elo-based rankings. Covers a wide range of categories including text, image, and video generation. |
| Design Arena | A crowdsourced benchmark for AI-generated design. Users anonymously compare outputs generated from the same prompt, and the resulting pairwise votes are aggregated into Bradley-Terry ratings. Rather than treating each modality as a separate benchmark, the unified leaderboard provides dedicated views for Code, Slides, Image, Video, Audio, and Builders, with more granular evaluations such as web and app development, game development, image editing, video editing, and text-to-speech. |
| Yupp Leaderboard | A leaderboard that ranks AI models based on community evaluation data, surfacing model quality through collective user insights from real-world usage. |
| Open WebUI Leaderboard | A leaderboard ranking 100+ models based on actual user usage and ratings within the Open WebUI open-source LLM frontend community. |
| Name | Description |
|---|---|
| DeepSWE | Measuring frontier coding agents on original, long-horizon engineering tasks. |
| SWE-rebench | A software engineering benchmark that continuously updates with new tasks and uses time-window analysis to detect and eliminate data contamination. Designed to address contamination issues in SWE-bench. |
| SWE-Bench Pro (Public) | A large-scale benchmark by Scale with 1,865 tasks. Uses copyleft-licensed code to prevent data contamination and rigorously evaluates AI problem-solving across diverse real-world tasks from B2B, consumer apps, and developer tools (public dataset). |
| SWE-bench | A benchmark measuring how well AI can resolve bug fixes and feature implementations using real GitHub issues from OSS projects such as Django and Matplotlib. Offers multiple variants including Verified, Multilingual, and Multimodal. |
| Convex LLM Leaderboard | A leaderboard evaluating code generation quality for the Convex platform. Compares performance with and without guidelines, measuring the impact of prompt design in practice. |
| Name | Description |
|---|---|
| Terminal-Bench | A benchmark evaluating AI agents' terminal operation capabilities by success rate across tasks spanning software development, ML, security, and data science. |
| SanityHarness | A high-signal leaderboard evaluating AI coding agents with weighted scoring across 26 tasks in multiple languages including Dart, Go, Kotlin, Rust, TypeScript, and Zig. |
| τ-bench | A benchmark that simulates business scenarios such as airlines and retail, measuring AI agents' ability to complete tasks through user interaction and dialogue. |
| Name | Description |
|---|---|
| FrontierMath | A benchmark developed by Epoch AI that measures advanced mathematical reasoning with problems ranging from undergraduate level to unsolved research questions of increasing difficulty. |
| Name | Description |
|---|---|
| Japanese-RP-Bench | A benchmark for Japanese role-playing LLMs that evaluates conversation quality, role fidelity, persona stability, resistance to persona replacement and misleading instructions, and recovery afterward. Version 2 retains the original 30-role, 10-exchange base evaluation while adding adversarial challenge scenarios. |
| Hemingway-bench | A writing benchmark and leaderboard evaluated by professional writers across creative, business, and everyday writing tasks. It emphasizes taste, originality, coherence, and emotional intelligence, aiming to reward genuinely effective writing rather than surface-level signals such as elaborate metaphors. |
| LLM Creative Story-Writing Benchmark | A short-story benchmark in which every model receives the same constrained creative brief and must meaningfully incorporate ten required elements, including a character, object, concept, attribute, action, method, setting, timeframe, motivation, and tone. Matched story pairs are judged by evaluator models, and the results are aggregated into a relative comparison score; the repository also publishes prompts, outputs, uncertainty ranges, and diagnostics such as word-count compliance. |
| EQ-Bench Longform Creative Writing | An LLM-judged long-form writing benchmark that tests planning a story from a minimal prompt, reflecting on and revising the plan, and writing a novella across eight roughly 1,000-word turns. Its rubric covers nuanced characters, emotional engagement, plot, coherence, tone, character consistency, plan and prompt adherence, as well as weak dialogue, tell-don't-show, clichés, amateurish writing, purple prose, and forced metaphors; it also reports repetition, LLM “slop,” and quality degradation across chapters. |
| EQ-Bench Creative Writing | An LLM-judged creative writing leaderboard that evaluates outputs from 32 prompts using both rubric scoring and pairwise comparisons with a modified Glicko rating system. It assesses qualities such as originality, character authenticity, coherence, and instruction following, while also reporting repetition and overused LLM phrase (“slop”) metrics. |
| Arena Creative Writing Leaderboard | A community-driven Text Arena ranking for open-ended creative-writing tasks. Models are compared through anonymous user votes and ranked with an Elo-style score, with vote counts and uncertainty shown alongside the results; the page provides a large, continuously updated cross-model view, but the rankings reflect user preference rather than a fixed rubric or a single reproducible test set. |
| Name | Description |
|---|---|
| DeepSecBench | A leaderboard measuring how effectively AI models find security vulnerabilities in application code using the DeepSec cyber harness. It combines recall and precision into a benchmark score and also compares false positives, cost, and execution time. |
| Name | Description |
|---|---|
| MTEB Leaderboard | The Massive Text Embedding Benchmark leaderboard on Hugging Face, evaluating and comparing text embedding models across multiple tasks and languages using standardized methods. |
| MMEB Leaderboard | The Multimodal Embedding Benchmark leaderboard by TIGER-Lab, evaluating models that represent text and images in a unified embedding space. |
| Name | Description |
|---|---|
| Open ASR Leaderboard | A Hugging Face leaderboard for automatic speech recognition (ASR) models, comparing recognition accuracy across multiple datasets. |