Automated auditing pipeline for LLM and agent benchmarks — surfaces task ambiguity, environment conflicts, and evaluation bugs.
-
Updated
May 27, 2026 - HTML
Automated auditing pipeline for LLM and agent benchmarks — surfaces task ambiguity, environment conflicts, and evaluation bugs.
This project aims to address this gap by conducting a systematic, controlled study of human versus LLM-generated text detectability using paired question–answer datasets. Rather than proposing a novel detection architecture, the focus is on analyzing detection robustness, failure modes, and the impact of adversarial humanization strategies.
7 Claude Code skills for software architecture review (Python, web, cloud, microservices). Includes A/B benchmarks against unskilled baseline, assertion-graded eval suite, and interactive dashboards.
Live, open-source benchmark for comparing AI coding agents on real GitHub issues
Find & compare 300+ LLMs by benchmark, speed, latency, price & capability — plus a free JSON API. Aggregates OpenRouter, Artificial Analysis & Design Arena.
Benchmark methodology, task sets, and evaluation results for RA²R
A/B evaluate any LLM task with and without Ejentum cognitive injection. n8n workflow + TypeScript module.
OpenCode TUI plugin for model benchmark scores, token costs, and token efficiency.
A local LLM benchmarking framework designed to evaluate model performance across multiple backends (LM Studio, Ollama, etc.), including metrics for speed, quality, and instruction adherence. Supports structured test runs, result analysis, and reproducible evaluations.
Interaction Net Equivalence Testing for LLMs
A living, evidence-based catalog of benchmarks for personalized LLMs and AI agents—covering preference alignment, long-term memory, tool use, safety, privacy, and multimodal adaptation.
Deep20Bench uses the classic Twenty Questions game to test how well an AI can identify a hidden subject. Success requires strategic questioning, reasoning, and broad world knowledge.
Experimental data and technical audit of LLM content detectors (GPT-5, Claude 4) vs. humanization algorithms.
Open benchmarks for local LLMs — same prompts, same suites, run on whatever hardware you've got, results compose into one ladder. Crowdsourced data + agent runbook for friends to clone-and-run.
Qwen 3.8 is LIVE NOW! Can It Survive 3 Brutal Tests? (Qwen 3.8 Max Benchmarks) - Technical guide, 2.4T parameter specifications, token pricing, and Canvas execution test prompts.
Every LLM leaderboard on one scale — 53 benchmarks, 6 arenas and live pricing, weighted by what you actually need.
Stock Bench is an LLM benchmarking system where LLMs compete in a prediction market, making bets on how well they’ll perform on tasks. Thus making it possible to measure each model's performance, as well as how accurate and self-aware each model is about their own performance.
Coding evals for agentic Drupal & Elm work — mechanical grading, hidden holdouts, and the first legacy-stack (Drupal 7) agent benchmark
Compare AI coding models with coding-focused benchmarks weighted your way — cross-verified, contradiction-flagged, in Turkish and English
Benchmarks, setup guides, and workarounds for running LLMs on Apple Silicon. Tested on M1 Max 64 GB — covers Ollama, MLX, quantization, memory management, and model recommendations.
Add a description, image, and links to the llm-benchmarks topic page so that developers can more easily learn about it.
To associate your repository with the llm-benchmarks topic, visit your repo's landing page and select "manage topics."