Digital Intelligence Institute Benchmark Report
A High-Cognitive Benchmark for Turkish-Language Performance of Large Language Models
HCBfT-TR1000
High-Cognitive Benchmark for Turkish  ·  An LLM test measuring Turkish verbal reasoning, cultural context, colloquial language, everyday practical knowledge, domain expertise, and visual culture
Accuracy Ranking
Accuracy Ranking — Percentage success rate (%)
Excellent ≥65%
Good 50–65%
Average 35–50%
Low 20–35%
Weak <20%
20%: random guessing
Efficiency — response time vs. accuracy
Free (:free)
Paid
↳ Bubble size is proportional to total output tokens. Top-left = fast and accurate.
Response Integrity format compliance, extraction, and errors
Response Breakdown — correct / wrong / unparsed (questions)
Correct
Wrong
Unparsed / errored
↳ "Unparsed": the model gave no valid answer, or the request failed permanently.
Instruction Following — how the answer was matched (%)
Stability and Breakdowns where each model falls short
Section Scores — multiple choice vs. short answer (%)
Multiple choice (A–E)
Short answer (open-ended)
↳ Multiple choice has a 20% random floor; short answer has a floor of zero. A wide gap between the two columns shows the model is eliminating among options rather than genuinely recalling the answer.
Accuracy by Topic (%)
Cells show accuracy percentage; parentheses show correct/total
Option Bias models' letter preference vs. the answer key
Prediction Distribution — how often each model picked A–E
↳ The grey bars show the answer key's actual distribution. If a model's bars deviate noticeably from the key, that points to position bias.
Detailed Model Scorecard
CI: Wilson 95% confidence interval. Coverage: share of questions that received a valid response from the provider; if low, the headline score is not a reliable measurement. Among answered: accuracy computed only over questions that received a response. Strict format: share of questions where the model wrote only the requested answer with no explanation. Cost: actual USD amount reported by OpenRouter (0 for free models).
Run Details method and limitations