A High-Cognitive Benchmark for Turkish-Language Performance of Large Language Models
HCBfT-TR1000
High-Cognitive Benchmark for Turkish · An LLM test measuring Turkish verbal reasoning, cultural context, colloquial language, everyday practical knowledge, domain expertise, and visual culture
Section 01Accuracy Ranking
Accuracy Ranking — Percentage success rate (%)
Excellent ≥65%
Good 50–65%
Average 35–50%
Low 20–35%
Weak <20%
20%: random guessing
Efficiency — response time vs. accuracy
Free (:free)
Paid
↳ Bubble size is proportional to total output tokens. Top-left = fast and accurate.
Section 02Response Integrity format compliance, extraction, and errors
↳ "Unparsed": the model gave no valid answer, or the request failed permanently.
Instruction Following — how the answer was matched (%)
Section 03Stability and Breakdowns where each model falls short
Stability Range — always correct → majority vote → at least one (%)
Always correct (lower bound)
Majority vote (reported score)
Correct on at least one attempt (upper bound)
↳ A wide gap between the three values means the model is unstable: it answers the same question differently across attempts.
Consistency Across Repeats (%)
Same answer across all attempts
Average hit share (all attempts)
↳ High consistency is expected at temperature 0; low values point to provider-side nondeterminism or the model's own instability.
Vision Section — text vs. visual accuracy (%)
Text section
Vision section
↳ Vision questions are only asked of multimodal models; for models that cannot accept image input, this section is skipped and never counts toward the score. Cross-model ranking should be based on the text section.
Vision Subtypes — which task type is hardest (%)
Section Scores — multiple choice vs. short answer (%)
Multiple choice (A–E)
Short answer (open-ended)
↳ Multiple choice has a 20% random floor; short answer has a floor of zero. A wide gap between the two columns shows the model is eliminating among options rather than genuinely recalling the answer.
Accuracy by Topic (%)
Cells show accuracy percentage; parentheses show correct/total
Section 04Option Bias models' letter preference vs. the answer key
Prediction Distribution — how often each model picked A–E
↳ The grey bars show the answer key's actual distribution. If a model's bars deviate noticeably from the key, that points to position bias.
Section 05Detailed Model ScorecardClick column headers to sort
↳ CI: Wilson 95% confidence interval.
Coverage: share of questions that received a valid response from the provider; if low, the headline score is not a reliable measurement.
Among answered: accuracy computed only over questions that received a response. Strict format: share of questions where the model wrote only the requested answer with no explanation. Cost: actual USD amount reported by OpenRouter (0 for free models).