Real Benchmarks

Planned

My personal feedback benchmarks for AI models. Scores from how models feel in real work, plus optional community notes. Not a sponsored or company-biased leaderboard.

Status: planned. Structure below is a product outline only. No live scores yet.

Feedback over marketing

Rankings will reflect real prompts and real preferences across capabilities you actually care about: reasoning, coding, agents, cost, latency, trust, and more.

Models

Seven task groups by output job so understanding, generation, extraction, and speech stay comparable on their own terms.

LLM ModelsVLM ModelsImage ModelsVideo ModelsOCR ModelsSpeech-to-TextText-to-Speech

Leaderboard filters

LLM-focused capability filters, adaptable to other modalities. Outline only; filters are not live.

AllReasoningCodingAgentic CodingMathematicsData AnalysisLanguageInstruction FollowingTool UseMultilingualLong ContextCreative WritingMulti-Turn

Insights

Views for comparing tradeoffs: overall quality, per-capability depth, cost, latency, and trust.

CostOverallReasoningCodingAgentic CodingMathematicsData AnalysisLanguageIFLatencyTool UseMultilingualLong ContextVoiceTrust
Free call