Digital Intelligence Institute Oyun Karnesi
Büyük Dil Modelleri İçin Türkçe Çok Ajanlı Oyun Değerlendirmesi · 2026-08-01 16:21 UTC
HCBfT-Games
High-Cognitive Benchmark for Turkish  ·  Diplomat, Tüccar, Vezir ve Âşık oyunlarında stratejik akıl yürütme, müzakere, aldatma direnci ve yaratıcı Türkçe üretimi ölçen çok ajanlı değerlendirme
Koşum demo-hcbft-games  ·  tohum 20260801  ·  tarafsız özetleyici google/gemini-3.6-flash-lite  ·  jüri anthropic/claude-haiku-4.5, openai/gpt-5.6-mini
Yarışmacı Model
10
4 oyun · 240 maç
Lider
109,6
claude-opus-5
Kompozit Ortalama
100,0
aralık 91,0–109,6 · ölçek 100 ± 15
Toplam Maç
240
toplam 23,5 sa oyun süresi
Geçersiz Eylem
%2,96
en yüksek: deepseek-v4-flash %5,16
Toplam Maliyet
$120,98
20.034 istek · 62,3M token
Oyun Karneleri her oyunda genel puan ve sıralama
Diplomat
Diplomacy ilhamlı
Stratejik iletişim · ittifak ve ihanet
110,0
En yüksek HCB-100 claude-opus-5
01 claude-opus-5 110,0
02 gpt-5.6-sol 106,6
03 kimi-k3 104,9
04 gpt-5.6-terra 101,5
05 claude-sonnet-5 99,8
06 grok-4.5 99,8
07 glm-5.2 96,4
08 deepseek-v4-flash 94,8
09 gemini-3.6-flash 94,8
10 qwen3.7-max 91,4
50 maç 10 model hücre n 20/20 aralık 91,4–110,0
Tüccar
Catan ilhamlı
Fayda analizi · pazarlık ve takas
109,8
En yüksek HCB-100 claude-opus-5
01 claude-opus-5 109,8
02 gpt-5.6-sol 106,7
03 kimi-k3 103,7
04 gpt-5.6-terra 102,1
05 grok-4.5 100,6
06 claude-sonnet-5 99,1
07 glm-5.2 96,0
08 deepseek-v4-flash 96,0
09 gemini-3.6-flash 96,0
10 qwen3.7-max 89,9
50 maç 10 model hücre n 20/20 aralık 89,9–109,8
Vezir
Avalon İlhamlı
Aldatma · tutarlılık ve çıkarım
110,3
En yüksek HCB-100 claude-opus-5
01 claude-opus-5 110,3
02 gpt-5.6-sol 107,5
03 kimi-k3 104,7
04 gpt-5.6-terra 101,7
05 grok-4.5 100,6
06 claude-sonnet-5 99,3
07 glm-5.2 96,4
08 gemini-3.6-flash 95,1
09 deepseek-v4-flash 94,9
10 qwen3.7-max 89,5
40 maç 10 model hücre n 20/20 aralık 89,5–110,3
Âşık
Ozan Atışması
Yaratıcı Türkçe · ölçü ve uyak
108,4
En yüksek HCB-100 claude-opus-5
01 claude-opus-5 108,4
02 gpt-5.6-sol 104,6
03 kimi-k3 102,7
04 grok-4.5 100,8
05 gpt-5.6-terra 100,8
06 claude-sonnet-5 98,9
07 deepseek-v4-flash 96,9
08 gemini-3.6-flash 96,9
09 glm-5.2 96,9
10 qwen3.7-max 93,1
100 maç 10 model hücre n 20/20 aralık 93,1–108,4
⌐ Büyük sayı, o oyundaki en yüksek HCB-100 puanıdır. Çubuklar modelleri koşumun en düşük ve en yüksek puanı arasında konumlandırır; 100 çizgisi havuz ortalamasıdır. Vezir'de standardizasyon rol sınıfı içinde yapılır (spec §8.2).
Genel Sıralama kompozit HCB-100
01
claude-opus-5 anthropic/claude-opus-5
109,6
GA 106,4–112,7
Diplomat 110,0Tüccar 109,8Vezir 110,3Âşık 108,4
02
gpt-5.6-sol openai/gpt-5.6-sol
106,3
GA 103,5–109,3
Diplomat 106,6Tüccar 106,7Vezir 107,5Âşık 104,6
03
kimi-k3 moonshotai/kimi-k3
104,0
GA 100,9–106,8
Diplomat 104,9Tüccar 103,7Vezir 104,7Âşık 102,7
04
gpt-5.6-terra openai/gpt-5.6-terra
101,5
GA 98,6–104,5
Diplomat 101,5Tüccar 102,1Vezir 101,7Âşık 100,8
05
grok-4.5 x-ai/grok-4.5
100,5
GA 97,7–103,0
Diplomat 99,8Tüccar 100,6Vezir 100,6Âşık 100,8
06
claude-sonnet-5 anthropic/claude-sonnet-5
99,3
GA 96,0–102,5
Diplomat 99,8Tüccar 99,1Vezir 99,3Âşık 98,9
07
glm-5.2 z-ai/glm-5.2
96,5
GA 93,6–99,2
Diplomat 96,4Tüccar 96,0Vezir 96,4Âşık 96,9
08
gemini-3.6-flash google/gemini-3.6-flash
95,7
GA 92,9–98,8
Diplomat 94,8Tüccar 96,0Vezir 95,1Âşık 96,9
09
deepseek-v4-flash deepseek/deepseek-v4-flash
95,6
GA 92,6–99,0
Diplomat 94,8Tüccar 96,0Vezir 94,9Âşık 96,9
10
qwen3.7-max qwen/qwen3.7-max
91,0
GA 88,3–93,8
Diplomat 91,4Tüccar 89,9Vezir 89,5Âşık 93,1
Elo — sağlamlık kontrolü
OyunModelElonKazanma oranı
Diplomatanthropic/claude-opus-51.623,120%55,0
Diplomatmoonshotai/kimi-k31.581,420%35,0
Diplomatopenai/gpt-5.6-sol1.569,720%30,0
Diplomatanthropic/claude-sonnet-51.527,220%30,0
Diplomatx-ai/grok-4.51.498,220%15,0
Diplomatopenai/gpt-5.6-terra1.469,820%15,0
Diplomatdeepseek/deepseek-v4-flash1.463,520%20,0
Diplomatgoogle/gemini-3.6-flash1.437,820%15,0
Diplomatqwen/qwen3.7-max1.426,320%10,0
Diplomatz-ai/glm-5.21.403,020%25,0
Tüccaropenai/gpt-5.6-sol1.632,620%40,0
Tüccaranthropic/claude-opus-51.621,320%35,0
Tüccarmoonshotai/kimi-k31.561,220%30,0
Tüccarx-ai/grok-4.51.535,120%20,0
Tüccardeepseek/deepseek-v4-flash1.472,920%25,0
Tüccaropenai/gpt-5.6-terra1.460,220%25,0
Tüccarz-ai/glm-5.21.448,920%25,0
Tüccaranthropic/claude-sonnet-51.446,020%25,0
Tüccargoogle/gemini-3.6-flash1.432,920%15,0
Tüccarqwen/qwen3.7-max1.388,820%10,0
Veziranthropic/claude-opus-51.649,120%70,0
Veziropenai/gpt-5.6-sol1.619,020%55,0
Vezirmoonshotai/kimi-k31.543,120%40,0
Vezirx-ai/grok-4.51.541,520%45,0
Veziropenai/gpt-5.6-terra1.515,420%45,0
Veziranthropic/claude-sonnet-51.474,720%35,0
Vezirdeepseek/deepseek-v4-flash1.456,520%40,0
Vezirz-ai/glm-5.21.451,620%20,0
Vezirgoogle/gemini-3.6-flash1.436,820%35,0
Vezirqwen/qwen3.7-max1.312,320%15,0
Âşıkmoonshotai/kimi-k31.639,720%70,0
Âşıkopenai/gpt-5.6-sol1.605,220%65,0
Âşıkx-ai/grok-4.51.565,620%60,0
Âşıkanthropic/claude-opus-51.557,520%55,0
Âşıkopenai/gpt-5.6-terra1.533,620%55,0
Âşıkz-ai/glm-5.21.497,320%50,0
Âşıkdeepseek/deepseek-v4-flash1.470,820%45,0
Âşıkqwen/qwen3.7-max1.408,620%40,0
Âşıkgoogle/gemini-3.6-flash1.367,420%30,0
Âşıkanthropic/claude-sonnet-51.354,420%30,0
z-puanı ve Elo sıralamaları uyuşmuyor: diplomat, tuccar, vezir, asik. Bu uyuşmazlık kendi başına raporlanmaya değerdir (spec §8.2).
Diplomat Stratejik iletişim · ittifak ve ihanet
Ham puan ve %95 bootstrap güven aralığı
ModelnHam ortalamaGA altGA üstStd. hatazHCB-100
anthropic/claude-opus-5200,3700,3290,4110,0210,665110,0
openai/gpt-5.6-sol200,3500,3100,3890,0200,440106,6
moonshotai/kimi-k3200,3400,3070,3730,0170,327104,9
openai/gpt-5.6-terra200,3200,2830,3570,0190,101101,5
anthropic/claude-sonnet-5200,3100,2670,3510,022-0,01199,8
x-ai/grok-4.5200,3100,2800,3420,016-0,01199,8
z-ai/glm-5.2200,2900,2540,3250,018-0,23796,4
deepseek/deepseek-v4-flash200,2800,2470,3140,017-0,35094,8
google/gemini-3.6-flash200,2800,2450,3160,018-0,35094,8
qwen/qwen3.7-max200,2600,2310,2890,015-0,57591,4
Koltuk bazında ortalama — denge denetimi
ModelKoltuknHam ortalama
anthropic/claude-opus-5050,421
anthropic/claude-opus-5150,299
anthropic/claude-opus-5250,384
anthropic/claude-opus-5350,377
anthropic/claude-sonnet-5050,363
anthropic/claude-sonnet-5150,356
anthropic/claude-sonnet-5250,300
anthropic/claude-sonnet-5350,220
deepseek/deepseek-v4-flash050,280
deepseek/deepseek-v4-flash150,320
deepseek/deepseek-v4-flash250,322
deepseek/deepseek-v4-flash350,198
google/gemini-3.6-flash050,324
google/gemini-3.6-flash150,252
google/gemini-3.6-flash250,270
google/gemini-3.6-flash350,273
moonshotai/kimi-k3050,318
moonshotai/kimi-k3150,335
moonshotai/kimi-k3250,343
moonshotai/kimi-k3350,364
openai/gpt-5.6-sol050,358
openai/gpt-5.6-sol150,359
openai/gpt-5.6-sol250,360
openai/gpt-5.6-sol350,324
openai/gpt-5.6-terra050,315
openai/gpt-5.6-terra150,369
openai/gpt-5.6-terra250,304
openai/gpt-5.6-terra350,293
qwen/qwen3.7-max050,286
qwen/qwen3.7-max150,257
qwen/qwen3.7-max250,217
qwen/qwen3.7-max350,280
x-ai/grok-4.5050,343
x-ai/grok-4.5150,322
x-ai/grok-4.5250,286
x-ai/grok-4.5350,289
z-ai/glm-5.2050,329
z-ai/glm-5.2150,321
z-ai/glm-5.2250,243
z-ai/glm-5.2350,267
İkincil metrikler — model ortalaması
ModelMerkezBirimMesaj sayısıDuyuru sayısıHamle sayısıDestek sayısıHamle başarı oranıElendiSöz verdiSöz tutma oranıİhanet oranı
anthropic/claude-opus-55,5505,30082125%72,80,0004,600%82,3%18,9
anthropic/claude-sonnet-55,1004,75083134%50,70,0504,950%55,7%44,5
deepseek/deepseek-v4-flash5,0504,60083124%42,50,0004,650%39,1%60,5
google/gemini-3.6-flash5,0504,85082135%43,40,0004,350%41,6%59,3
moonshotai/kimi-k35,4504,90082124%65,10,0004,500%70,6%29,9
openai/gpt-5.6-sol5,2005,20082124%66,80,0004,700%77,2%21,9
openai/gpt-5.6-terra5,3505,10083134%58,90,0004,550%60,3%40,3
qwen/qwen3.7-max4,9504,50083134%34,60,0004,400%29,6%71,0
x-ai/grok-4.55,1505,00082125%53,90,0004,150%52,8%47,4
z-ai/glm-5.25,0004,70093124%43,90,0004,950%42,4%57,8
İkili anlamlılık matrisi — permütasyon testi, Holm–Bonferroni
Model AModel Bn (A)n (B)Farkpp (Holm)Yeterli nAnlamlılık
anthropic/claude-opus-5anthropic/claude-sonnet-520200,0600,0581,000
anthropic/claude-opus-5deepseek/deepseek-v4-flash20200,0900,0030,118
anthropic/claude-opus-5google/gemini-3.6-flash20200,0900,0030,123
anthropic/claude-opus-5moonshotai/kimi-k320200,0300,2901,000
anthropic/claude-opus-5openai/gpt-5.6-sol20200,0200,5071,000
anthropic/claude-opus-5openai/gpt-5.6-terra20200,0500,0901,000
anthropic/claude-opus-5qwen/qwen3.7-max20200,1100,0010,027*
anthropic/claude-opus-5x-ai/grok-4.520200,0600,0351,000
anthropic/claude-opus-5z-ai/glm-5.220200,0800,0060,224
anthropic/claude-sonnet-5deepseek/deepseek-v4-flash20200,0300,2981,000
anthropic/claude-sonnet-5google/gemini-3.6-flash20200,0300,3091,000
anthropic/claude-sonnet-5moonshotai/kimi-k32020-0,0300,2971,000
anthropic/claude-sonnet-5openai/gpt-5.6-sol2020-0,0400,1981,000
anthropic/claude-sonnet-5openai/gpt-5.6-terra2020-0,0100,7331,000
anthropic/claude-sonnet-5qwen/qwen3.7-max20200,0500,0781,000
anthropic/claude-sonnet-5x-ai/grok-4.520200,0001,0001,000
anthropic/claude-sonnet-5z-ai/glm-5.220200,0200,4981,000
deepseek/deepseek-v4-flashgoogle/gemini-3.6-flash20200,0001,0001,000
deepseek/deepseek-v4-flashmoonshotai/kimi-k32020-0,0600,0220,829
deepseek/deepseek-v4-flashopenai/gpt-5.6-sol2020-0,0700,0150,577
deepseek/deepseek-v4-flashopenai/gpt-5.6-terra2020-0,0400,1341,000
deepseek/deepseek-v4-flashqwen/qwen3.7-max20200,0200,3951,000
deepseek/deepseek-v4-flashx-ai/grok-4.52020-0,0300,2081,000
deepseek/deepseek-v4-flashz-ai/glm-5.22020-0,0100,6891,000
google/gemini-3.6-flashmoonshotai/kimi-k32020-0,0600,0260,896
google/gemini-3.6-flashopenai/gpt-5.6-sol2020-0,0700,0170,638
google/gemini-3.6-flashopenai/gpt-5.6-terra2020-0,0400,1411,000
google/gemini-3.6-flashqwen/qwen3.7-max20200,0200,4171,000
google/gemini-3.6-flashx-ai/grok-4.52020-0,0300,2291,000
google/gemini-3.6-flashz-ai/glm-5.22020-0,0100,6991,000
moonshotai/kimi-k3openai/gpt-5.6-sol2020-0,0100,7151,000
moonshotai/kimi-k3openai/gpt-5.6-terra20200,0200,4451,000
moonshotai/kimi-k3qwen/qwen3.7-max20200,0800,0020,069
moonshotai/kimi-k3x-ai/grok-4.520200,0300,2171,000
moonshotai/kimi-k3z-ai/glm-5.220200,0500,0551,000
openai/gpt-5.6-solopenai/gpt-5.6-terra20200,0300,2951,000
openai/gpt-5.6-solqwen/qwen3.7-max20200,0900,0010,062
openai/gpt-5.6-solx-ai/grok-4.520200,0400,1411,000
openai/gpt-5.6-solz-ai/glm-5.220200,0600,0341,000
openai/gpt-5.6-terraqwen/qwen3.7-max20200,0600,0220,829
openai/gpt-5.6-terrax-ai/grok-4.520200,0100,6991,000
openai/gpt-5.6-terraz-ai/glm-5.220200,0300,2581,000
qwen/qwen3.7-maxx-ai/grok-4.52020-0,0500,0311,000
qwen/qwen3.7-maxz-ai/glm-5.22020-0,0300,2121,000
x-ai/grok-4.5z-ai/glm-5.220200,0200,4151,000
⌐ Yıldızlar: *** p<0,001 · ** p<0,01 · * p<0,05 · — anlamsız. Hücre n hedefin altındaysa yıldız basılmaz.
Tüccar Fayda analizi · pazarlık ve takas
Ham puan ve %95 bootstrap güven aralığı
ModelnHam ortalamaGA altGA üstStd. hatazHCB-100
anthropic/claude-opus-5200,4100,3730,4460,0190,652109,8
openai/gpt-5.6-sol200,3900,3520,4290,0200,448106,7
moonshotai/kimi-k3200,3700,3340,4070,0190,245103,7
openai/gpt-5.6-terra200,3600,3180,4020,0210,143102,1
x-ai/grok-4.5200,3500,3150,3860,0180,041100,6
anthropic/claude-sonnet-5200,3400,2980,3840,022-0,06199,1
z-ai/glm-5.2200,3200,2810,3590,020-0,26596,0
deepseek/deepseek-v4-flash200,3200,2780,3640,022-0,26596,0
google/gemini-3.6-flash200,3200,2730,3660,024-0,26596,0
qwen/qwen3.7-max200,2800,2440,3170,018-0,67289,9
Koltuk bazında ortalama — denge denetimi
ModelKoltuknHam ortalama
anthropic/claude-opus-5050,442
anthropic/claude-opus-5150,392
anthropic/claude-opus-5250,407
anthropic/claude-opus-5350,398
anthropic/claude-sonnet-5050,379
anthropic/claude-sonnet-5150,329
anthropic/claude-sonnet-5250,315
anthropic/claude-sonnet-5350,338
deepseek/deepseek-v4-flash050,306
deepseek/deepseek-v4-flash150,307
deepseek/deepseek-v4-flash250,321
deepseek/deepseek-v4-flash350,346
google/gemini-3.6-flash050,396
google/gemini-3.6-flash150,323
google/gemini-3.6-flash250,285
google/gemini-3.6-flash350,276
moonshotai/kimi-k3050,336
moonshotai/kimi-k3150,396
moonshotai/kimi-k3250,401
moonshotai/kimi-k3350,347
openai/gpt-5.6-sol050,419
openai/gpt-5.6-sol150,428
openai/gpt-5.6-sol250,357
openai/gpt-5.6-sol350,356
openai/gpt-5.6-terra050,408
openai/gpt-5.6-terra150,288
openai/gpt-5.6-terra250,371
openai/gpt-5.6-terra350,373
qwen/qwen3.7-max050,317
qwen/qwen3.7-max150,248
qwen/qwen3.7-max250,244
qwen/qwen3.7-max350,311
x-ai/grok-4.5050,351
x-ai/grok-4.5150,326
x-ai/grok-4.5250,372
x-ai/grok-4.5350,351
z-ai/glm-5.2050,318
z-ai/glm-5.2150,368
z-ai/glm-5.2250,332
z-ai/glm-5.2350,262
İkincil metrikler — model ortalaması
ModelZafer puanıTeklif sunduTeklifi kabul edildiTeklif kabul oranıKabul ettiği teklifTakas artığıBanka takasıBanka bağımlılığıÜretilen kaynakEn uzun yol
anthropic/claude-opus-55,95010,0005,800%58,23,8001,7690,5500,06115,4003,350
anthropic/claude-sonnet-55,45010,9003,850%35,63,6000,6351,6000,17716,4502,200
deepseek/deepseek-v4-flash5,20010,0502,650%26,33,2000,3012,5000,27816,5001,900
google/gemini-3.6-flash5,40010,5503,050%29,43,1500,3012,1000,24616,6502,050
moonshotai/kimi-k35,7509,9004,750%47,83,3501,1601,4000,14416,6502,950
openai/gpt-5.6-sol6,00011,4506,100%53,04,1001,4260,7000,06515,3503,250
openai/gpt-5.6-terra5,5009,8504,100%42,13,2500,9321,5000,17715,2502,650
qwen/qwen3.7-max4,95010,0501,550%15,53,000-0,3762,2500,31215,6001,650
x-ai/grok-4.55,40010,0003,650%36,63,5000,8221,6000,17815,5502,800
z-ai/glm-5.25,25011,0003,100%28,33,1000,3452,1000,25812,5502,050
İkili anlamlılık matrisi — permütasyon testi, Holm–Bonferroni
Model AModel Bn (A)n (B)Farkpp (Holm)Yeterli nAnlamlılık
anthropic/claude-opus-5anthropic/claude-sonnet-520200,0700,0240,833
anthropic/claude-opus-5deepseek/deepseek-v4-flash20200,0900,0050,197
anthropic/claude-opus-5google/gemini-3.6-flash20200,0900,0050,216
anthropic/claude-opus-5moonshotai/kimi-k320200,0400,1441,000
anthropic/claude-opus-5openai/gpt-5.6-sol20200,0200,4871,000
anthropic/claude-opus-5openai/gpt-5.6-terra20200,0500,0911,000
anthropic/claude-opus-5qwen/qwen3.7-max20200,1300,0000,009**
anthropic/claude-opus-5x-ai/grok-4.520200,0600,0311,000
anthropic/claude-opus-5z-ai/glm-5.220200,0900,0030,126
anthropic/claude-sonnet-5deepseek/deepseek-v4-flash20200,0200,5321,000
anthropic/claude-sonnet-5google/gemini-3.6-flash20200,0200,5521,000
anthropic/claude-sonnet-5moonshotai/kimi-k32020-0,0300,3261,000
anthropic/claude-sonnet-5openai/gpt-5.6-sol2020-0,0500,1081,000
anthropic/claude-sonnet-5openai/gpt-5.6-terra2020-0,0200,5311,000
anthropic/claude-sonnet-5qwen/qwen3.7-max20200,0600,0481,000
anthropic/claude-sonnet-5x-ai/grok-4.52020-0,0100,7511,000
anthropic/claude-sonnet-5z-ai/glm-5.220200,0200,5161,000
deepseek/deepseek-v4-flashgoogle/gemini-3.6-flash20200,0001,0001,000
deepseek/deepseek-v4-flashmoonshotai/kimi-k32020-0,0500,1071,000
deepseek/deepseek-v4-flashopenai/gpt-5.6-sol2020-0,0700,0230,828
deepseek/deepseek-v4-flashopenai/gpt-5.6-terra2020-0,0400,2011,000
deepseek/deepseek-v4-flashqwen/qwen3.7-max20200,0400,1811,000
deepseek/deepseek-v4-flashx-ai/grok-4.52020-0,0300,3051,000
deepseek/deepseek-v4-flashz-ai/glm-5.22020-0,0001,0001,000
google/gemini-3.6-flashmoonshotai/kimi-k32020-0,0500,1191,000
google/gemini-3.6-flashopenai/gpt-5.6-sol2020-0,0700,0321,000
google/gemini-3.6-flashopenai/gpt-5.6-terra2020-0,0400,2361,000
google/gemini-3.6-flashqwen/qwen3.7-max20200,0400,2001,000
google/gemini-3.6-flashx-ai/grok-4.52020-0,0300,3421,000
google/gemini-3.6-flashz-ai/glm-5.22020-0,0001,0001,000
moonshotai/kimi-k3openai/gpt-5.6-sol2020-0,0200,4721,000
moonshotai/kimi-k3openai/gpt-5.6-terra20200,0100,7271,000
moonshotai/kimi-k3qwen/qwen3.7-max20200,0900,0020,086
moonshotai/kimi-k3x-ai/grok-4.520200,0200,4381,000
moonshotai/kimi-k3z-ai/glm-5.220200,0500,0851,000
openai/gpt-5.6-solopenai/gpt-5.6-terra20200,0300,3041,000
openai/gpt-5.6-solqwen/qwen3.7-max20200,1100,0010,026*
openai/gpt-5.6-solx-ai/grok-4.520200,0400,1491,000
openai/gpt-5.6-solz-ai/glm-5.220200,0700,0200,740
openai/gpt-5.6-terraqwen/qwen3.7-max20200,0800,0080,328
openai/gpt-5.6-terrax-ai/grok-4.520200,0100,7361,000
openai/gpt-5.6-terraz-ai/glm-5.220200,0400,1981,000
qwen/qwen3.7-maxx-ai/grok-4.52020-0,0700,0120,456
qwen/qwen3.7-maxz-ai/glm-5.22020-0,0400,1531,000
x-ai/grok-4.5z-ai/glm-5.220200,0300,2891,000
⌐ Yıldızlar: *** p<0,001 · ** p<0,01 · * p<0,05 · — anlamsız. Hücre n hedefin altındaysa yıldız basılmaz.
Vezir Aldatma · tutarlılık ve çıkarım
Ham puan ve %95 bootstrap güven aralığı
ModelnHam ortalamaGA altGA üstStd. hatazHCB-100
anthropic/claude-opus-5200,5000,4530,5450,0230,688110,3
openai/gpt-5.6-sol200,4800,4350,5210,0220,499107,5
moonshotai/kimi-k3200,4600,4120,5080,0250,313104,7
openai/gpt-5.6-terra200,4400,3940,4840,0230,115101,7
x-ai/grok-4.5200,4300,3900,4710,0210,041100,6
anthropic/claude-sonnet-5200,4200,3730,4700,025-0,04899,3
z-ai/glm-5.2200,4000,3690,4300,016-0,24096,4
google/gemini-3.6-flash200,3900,3510,4280,020-0,32695,1
deepseek/deepseek-v4-flash200,3900,3360,4410,027-0,34394,9
qwen/qwen3.7-max200,3500,3040,3930,023-0,69989,5
Koltuk bazında ortalama — denge denetimi
ModelKoltuknHam ortalama
anthropic/claude-opus-5040,442
anthropic/claude-opus-5140,492
anthropic/claude-opus-5240,556
anthropic/claude-opus-5340,534
anthropic/claude-opus-5440,476
anthropic/claude-sonnet-5040,453
anthropic/claude-sonnet-5140,476
anthropic/claude-sonnet-5240,367
anthropic/claude-sonnet-5340,408
anthropic/claude-sonnet-5440,397
deepseek/deepseek-v4-flash040,383
deepseek/deepseek-v4-flash140,358
deepseek/deepseek-v4-flash240,488
deepseek/deepseek-v4-flash340,396
deepseek/deepseek-v4-flash440,325
google/gemini-3.6-flash040,461
google/gemini-3.6-flash140,317
google/gemini-3.6-flash240,428
google/gemini-3.6-flash340,353
google/gemini-3.6-flash440,392
moonshotai/kimi-k3040,587
moonshotai/kimi-k3140,455
moonshotai/kimi-k3240,396
moonshotai/kimi-k3340,455
moonshotai/kimi-k3440,408
openai/gpt-5.6-sol040,538
openai/gpt-5.6-sol140,443
openai/gpt-5.6-sol240,434
openai/gpt-5.6-sol340,450
openai/gpt-5.6-sol440,535
openai/gpt-5.6-terra040,404
openai/gpt-5.6-terra140,449
openai/gpt-5.6-terra240,506
openai/gpt-5.6-terra340,374
openai/gpt-5.6-terra440,467
qwen/qwen3.7-max040,349
qwen/qwen3.7-max140,313
qwen/qwen3.7-max240,373
qwen/qwen3.7-max340,365
qwen/qwen3.7-max440,349
x-ai/grok-4.5040,459
x-ai/grok-4.5140,497
x-ai/grok-4.5240,476
x-ai/grok-4.5340,372
x-ai/grok-4.5440,346
z-ai/glm-5.2040,387
z-ai/glm-5.2140,370
z-ai/glm-5.2240,432
z-ai/glm-5.2340,389
z-ai/glm-5.2440,421
Rol sınıfı bazında ortalama ve rol-içi z
ModelRol sınıfınHam ortalamazHCB-100
anthropic/claude-opus-5bilge40,5270,819112,3
anthropic/claude-opus-5iyi-sade80,4530,410106,2
anthropic/claude-opus-5kotu-sade40,5120,992114,9
anthropic/claude-opus-5suikastci40,5540,806112,1
anthropic/claude-sonnet-5bilge40,330-0,81287,8
anthropic/claude-sonnet-5iyi-sade80,4360,248103,7
anthropic/claude-sonnet-5kotu-sade40,344-0,72389,2
anthropic/claude-sonnet-5suikastci40,5530,800112,0
deepseek/deepseek-v4-flashbilge40,428-0,001100,0
deepseek/deepseek-v4-flashiyi-sade80,407-0,03299,5
deepseek/deepseek-v4-flashkotu-sade40,349-0,67589,9
deepseek/deepseek-v4-flashsuikastci40,359-0,97585,4
google/gemini-3.6-flashbilge40,347-0,67289,9
google/gemini-3.6-flashiyi-sade80,373-0,36694,5
google/gemini-3.6-flashkotu-sade40,397-0,18297,3
google/gemini-3.6-flashsuikastci40,461-0,04599,3
moonshotai/kimi-k3bilge40,4440,132102,0
moonshotai/kimi-k3iyi-sade80,4150,045100,7
moonshotai/kimi-k3kotu-sade40,4380,237103,6
moonshotai/kimi-k3suikastci40,5871,106116,6
openai/gpt-5.6-solbilge40,5380,907113,6
openai/gpt-5.6-soliyi-sade80,4580,454106,8
openai/gpt-5.6-solkotu-sade40,4870,730110,9
openai/gpt-5.6-solsuikastci40,460-0,05199,2
openai/gpt-5.6-terrabilge40,5210,767111,5
openai/gpt-5.6-terraiyi-sade80,385-0,25196,2
openai/gpt-5.6-terrakotu-sade40,4560,418106,3
openai/gpt-5.6-terrasuikastci40,454-0,10998,4
qwen/qwen3.7-maxbilge40,297-1,08083,8
qwen/qwen3.7-maxiyi-sade80,342-0,66090,1
qwen/qwen3.7-maxkotu-sade40,349-0,67089,9
qwen/qwen3.7-maxsuikastci40,419-0,42593,6
x-ai/grok-4.5bilge40,4610,277104,2
x-ai/grok-4.5iyi-sade80,4140,032100,5
x-ai/grok-4.5kotu-sade40,4590,452106,8
x-ai/grok-4.5suikastci40,401-0,58791,2
z-ai/glm-5.2bilge40,387-0,33794,9
z-ai/glm-5.2iyi-sade80,4230,119101,8
z-ai/glm-5.2kotu-sade40,358-0,57991,3
z-ai/glm-5.2suikastci40,409-0,52092,2
İkincil metrikler — model ortalaması
ModelOy sayısıÖnerdiği takımOy isabetiBilge sızıntısıAldatma başarısıBaşarısız kartSuikast isabeti
anthropic/claude-opus-591,9500,8410,2500,7620,8750,500
anthropic/claude-sonnet-591,6000,5710,2500,5261,7501,000
deepseek/deepseek-v4-flash91,4000,4850,7500,4341,7500,250
google/gemini-3.6-flash91,3000,4850,2500,4521,3750,500
moonshotai/kimi-k391,8000,7090,2500,6541,2500,750
openai/gpt-5.6-sol91,3000,7750,2500,7281,5000,500
openai/gpt-5.6-terra91,6000,6510,5000,5371,6250,750
qwen/qwen3.7-max91,3000,3330,5000,2530,7500,500
x-ai/grok-4.591,1500,6490,0000,5301,1250,750
z-ai/glm-5.2101,4500,5000,5000,4861,2500,500
İkili anlamlılık matrisi — permütasyon testi, Holm–Bonferroni
Model AModel Bn (A)n (B)Farkpp (Holm)Yeterli nAnlamlılık
anthropic/claude-opus-5anthropic/claude-sonnet-520200,0800,0260,870
anthropic/claude-opus-5deepseek/deepseek-v4-flash20200,1100,0040,160
anthropic/claude-opus-5google/gemini-3.6-flash20200,1100,0020,101
anthropic/claude-opus-5moonshotai/kimi-k320200,0400,2531,000
anthropic/claude-opus-5openai/gpt-5.6-sol20200,0200,5561,000
anthropic/claude-opus-5openai/gpt-5.6-terra20200,0600,0811,000
anthropic/claude-opus-5qwen/qwen3.7-max20200,1500,0000,018*
anthropic/claude-opus-5x-ai/grok-4.520200,0700,0391,000
anthropic/claude-opus-5z-ai/glm-5.220200,1000,0020,086
anthropic/claude-sonnet-5deepseek/deepseek-v4-flash20200,0300,4191,000
anthropic/claude-sonnet-5google/gemini-3.6-flash20200,0300,3561,000
anthropic/claude-sonnet-5moonshotai/kimi-k32020-0,0400,2691,000
anthropic/claude-sonnet-5openai/gpt-5.6-sol2020-0,0600,0821,000
anthropic/claude-sonnet-5openai/gpt-5.6-terra2020-0,0200,5641,000
anthropic/claude-sonnet-5qwen/qwen3.7-max20200,0700,0511,000
anthropic/claude-sonnet-5x-ai/grok-4.52020-0,0100,7591,000
anthropic/claude-sonnet-5z-ai/glm-5.220200,0200,5111,000
deepseek/deepseek-v4-flashgoogle/gemini-3.6-flash2020-0,0001,0001,000
deepseek/deepseek-v4-flashmoonshotai/kimi-k32020-0,0700,0651,000
deepseek/deepseek-v4-flashopenai/gpt-5.6-sol2020-0,0900,0180,616
deepseek/deepseek-v4-flashopenai/gpt-5.6-terra2020-0,0500,1781,000
deepseek/deepseek-v4-flashqwen/qwen3.7-max20200,0400,2651,000
deepseek/deepseek-v4-flashx-ai/grok-4.52020-0,0400,2471,000
deepseek/deepseek-v4-flashz-ai/glm-5.22020-0,0100,7521,000
google/gemini-3.6-flashmoonshotai/kimi-k32020-0,0700,0421,000
google/gemini-3.6-flashopenai/gpt-5.6-sol2020-0,0900,0040,172
google/gemini-3.6-flashopenai/gpt-5.6-terra2020-0,0500,1071,000
google/gemini-3.6-flashqwen/qwen3.7-max20200,0400,2161,000
google/gemini-3.6-flashx-ai/grok-4.52020-0,0400,1851,000
google/gemini-3.6-flashz-ai/glm-5.22020-0,0100,7071,000
moonshotai/kimi-k3openai/gpt-5.6-sol2020-0,0200,5531,000
moonshotai/kimi-k3openai/gpt-5.6-terra20200,0200,5721,000
moonshotai/kimi-k3qwen/qwen3.7-max20200,1100,0030,123
moonshotai/kimi-k3x-ai/grok-4.520200,0300,3751,000
moonshotai/kimi-k3z-ai/glm-5.220200,0600,0521,000
openai/gpt-5.6-solopenai/gpt-5.6-terra20200,0400,2261,000
openai/gpt-5.6-solqwen/qwen3.7-max20200,1300,0010,026*
openai/gpt-5.6-solx-ai/grok-4.520200,0500,1071,000
openai/gpt-5.6-solz-ai/glm-5.220200,0800,0080,289
openai/gpt-5.6-terraqwen/qwen3.7-max20200,0900,0100,377
openai/gpt-5.6-terrax-ai/grok-4.520200,0100,7551,000
openai/gpt-5.6-terraz-ai/glm-5.220200,0400,1721,000
qwen/qwen3.7-maxx-ai/grok-4.52020-0,0800,0150,554
qwen/qwen3.7-maxz-ai/glm-5.22020-0,0500,0861,000
x-ai/grok-4.5z-ai/glm-5.220200,0300,2641,000
⌐ Yıldızlar: *** p<0,001 · ** p<0,01 · * p<0,05 · — anlamsız. Hücre n hedefin altındaysa yıldız basılmaz.
Âşık Yaratıcı Türkçe · ölçü ve uyak
Ham puan ve %95 bootstrap güven aralığı
ModelnHam ortalamaGA altGA üstStd. hatazHCB-100
anthropic/claude-opus-5200,2700,2290,3120,0210,560108,4
openai/gpt-5.6-sol200,2500,2260,2750,0130,305104,6
moonshotai/kimi-k3200,2400,2020,2740,0180,178102,7
x-ai/grok-4.5200,2300,2040,2550,0130,051100,8
openai/gpt-5.6-terra200,2300,1970,2640,0170,051100,8
anthropic/claude-sonnet-5200,2200,1890,2500,016-0,07698,9
deepseek/deepseek-v4-flash200,2100,1700,2490,020-0,20396,9
google/gemini-3.6-flash200,2100,1800,2420,016-0,20396,9
z-ai/glm-5.2200,2100,1800,2400,016-0,20396,9
qwen/qwen3.7-max200,1900,1560,2250,018-0,45893,1
Koltuk bazında ortalama — denge denetimi
ModelKoltuknHam ortalama
anthropic/claude-opus-50100,279
anthropic/claude-opus-51100,261
anthropic/claude-sonnet-50100,206
anthropic/claude-sonnet-51100,234
deepseek/deepseek-v4-flash0100,227
deepseek/deepseek-v4-flash1100,193
google/gemini-3.6-flash0100,249
google/gemini-3.6-flash1100,171
moonshotai/kimi-k30100,248
moonshotai/kimi-k31100,232
openai/gpt-5.6-sol0100,247
openai/gpt-5.6-sol1100,253
openai/gpt-5.6-terra0100,234
openai/gpt-5.6-terra1100,226
qwen/qwen3.7-max0100,201
qwen/qwen3.7-max1100,179
x-ai/grok-4.50100,222
x-ai/grok-4.51100,238
z-ai/glm-5.20100,218
z-ai/glm-5.21100,202
İkincil metrikler — model ortalaması
ModelForm puanıHece isabetiUyak isabetiUyak derecesiBiçim ihlal oranıTam ölçülü dize oranıAçan ozanJüri puanı
anthropic/claude-opus-50,9120,9400,8942,050%5,5%89,90,5000,741
anthropic/claude-sonnet-50,5930,6410,5622,150%23,2%62,00,5000,457
deepseek/deepseek-v4-flash0,5590,6020,5311,950%27,6%57,90,5000,404
google/gemini-3.6-flash0,5460,5790,5232,150%27,2%56,40,5000,419
moonshotai/kimi-k30,7210,7520,7001,800%16,8%72,80,5000,584
openai/gpt-5.6-sol0,8120,8410,7932,050%11,3%81,00,5000,668
openai/gpt-5.6-terra0,6740,6930,6612,250%19,7%65,00,5000,535
qwen/qwen3.7-max0,4460,4740,4271,800%34,8%45,70,5000,309
x-ai/grok-4.50,6690,6820,6602,150%19,2%66,60,5000,536
z-ai/glm-5.20,5590,5900,5392,150%28,6%56,00,5000,416
İkili anlamlılık matrisi — permütasyon testi, Holm–Bonferroni
Model AModel Bn (A)n (B)Farkpp (Holm)Yeterli nAnlamlılık
anthropic/claude-opus-5anthropic/claude-sonnet-520200,0500,0711,000
anthropic/claude-opus-5deepseek/deepseek-v4-flash20200,0600,0571,000
anthropic/claude-opus-5google/gemini-3.6-flash20200,0600,0311,000
anthropic/claude-opus-5moonshotai/kimi-k320200,0300,2971,000
anthropic/claude-opus-5openai/gpt-5.6-sol20200,0200,4141,000
anthropic/claude-opus-5openai/gpt-5.6-terra20200,0400,1601,000
anthropic/claude-opus-5qwen/qwen3.7-max20200,0800,0050,225
anthropic/claude-opus-5x-ai/grok-4.520200,0400,1221,000
anthropic/claude-opus-5z-ai/glm-5.220200,0600,0351,000
anthropic/claude-sonnet-5deepseek/deepseek-v4-flash20200,0100,7051,000
anthropic/claude-sonnet-5google/gemini-3.6-flash20200,0100,6601,000
anthropic/claude-sonnet-5moonshotai/kimi-k32020-0,0200,4321,000
anthropic/claude-sonnet-5openai/gpt-5.6-sol2020-0,0300,1501,000
anthropic/claude-sonnet-5openai/gpt-5.6-terra2020-0,0100,6671,000
anthropic/claude-sonnet-5qwen/qwen3.7-max20200,0300,2271,000
anthropic/claude-sonnet-5x-ai/grok-4.52020-0,0100,6351,000
anthropic/claude-sonnet-5z-ai/glm-5.220200,0100,6551,000
deepseek/deepseek-v4-flashgoogle/gemini-3.6-flash20200,0001,0001,000
deepseek/deepseek-v4-flashmoonshotai/kimi-k32020-0,0300,3061,000
deepseek/deepseek-v4-flashopenai/gpt-5.6-sol2020-0,0400,1101,000
deepseek/deepseek-v4-flashopenai/gpt-5.6-terra2020-0,0200,4651,000
deepseek/deepseek-v4-flashqwen/qwen3.7-max20200,0200,4701,000
deepseek/deepseek-v4-flashx-ai/grok-4.52020-0,0200,4201,000
deepseek/deepseek-v4-flashz-ai/glm-5.220200,0001,0001,000
google/gemini-3.6-flashmoonshotai/kimi-k32020-0,0300,2541,000
google/gemini-3.6-flashopenai/gpt-5.6-sol2020-0,0400,0581,000
google/gemini-3.6-flashopenai/gpt-5.6-terra2020-0,0200,4061,000
google/gemini-3.6-flashqwen/qwen3.7-max20200,0200,4051,000
google/gemini-3.6-flashx-ai/grok-4.52020-0,0200,3411,000
google/gemini-3.6-flashz-ai/glm-5.220200,0001,0001,000
moonshotai/kimi-k3openai/gpt-5.6-sol2020-0,0100,6621,000
moonshotai/kimi-k3openai/gpt-5.6-terra20200,0100,7011,000
moonshotai/kimi-k3qwen/qwen3.7-max20200,0500,0631,000
moonshotai/kimi-k3x-ai/grok-4.520200,0100,6671,000
moonshotai/kimi-k3z-ai/glm-5.220200,0300,2301,000
openai/gpt-5.6-solopenai/gpt-5.6-terra20200,0200,3611,000
openai/gpt-5.6-solqwen/qwen3.7-max20200,0600,0100,440
openai/gpt-5.6-solx-ai/grok-4.520200,0200,2931,000
openai/gpt-5.6-solz-ai/glm-5.220200,0400,0631,000
openai/gpt-5.6-terraqwen/qwen3.7-max20200,0400,1141,000
openai/gpt-5.6-terrax-ai/grok-4.52020-0,0001,0001,000
openai/gpt-5.6-terraz-ai/glm-5.220200,0200,3881,000
qwen/qwen3.7-maxx-ai/grok-4.52020-0,0400,0841,000
qwen/qwen3.7-maxz-ai/glm-5.22020-0,0200,4051,000
x-ai/grok-4.5z-ai/glm-5.220200,0200,3421,000
⌐ Yıldızlar: *** p<0,001 · ** p<0,01 · * p<0,05 · — anlamsız. Hücre n hedefin altındaysa yıldız basılmaz.
Davranışsal Analitik oyun içi davranış örüntüleri
Diplomat — söz tutma ve ihanet
ModelSöz verdiSöz tutma oranıİhanet oranıMesaj sayısı
anthropic/claude-opus-54,600%82,3%18,98
anthropic/claude-sonnet-54,950%55,7%44,58
deepseek/deepseek-v4-flash4,650%39,1%60,58
google/gemini-3.6-flash4,350%41,6%59,38
moonshotai/kimi-k34,500%70,6%29,98
openai/gpt-5.6-sol4,700%77,2%21,98
openai/gpt-5.6-terra4,550%60,3%40,38
qwen/qwen3.7-max4,400%29,6%71,08
x-ai/grok-4.54,150%52,8%47,48
z-ai/glm-5.24,950%42,4%57,89
Tüccar — takas davranışı
ModelTeklif sunduTeklif kabul oranıTakas artığıBanka bağımlılığı
anthropic/claude-opus-510,000%58,21,7690,061
anthropic/claude-sonnet-510,900%35,60,6350,177
deepseek/deepseek-v4-flash10,050%26,30,3010,278
google/gemini-3.6-flash10,550%29,40,3010,246
moonshotai/kimi-k39,900%47,81,1600,144
openai/gpt-5.6-sol11,450%53,01,4260,065
openai/gpt-5.6-terra9,850%42,10,9320,177
qwen/qwen3.7-max10,050%15,5-0,3760,312
x-ai/grok-4.510,000%36,60,8220,178
z-ai/glm-5.211,000%28,30,3450,258
Vezir — oy isabeti ve aldatma
ModelOy isabetiAldatma başarısıBilge sızıntısıSuikast isabeti
anthropic/claude-opus-50,8410,7620,2500,500
anthropic/claude-sonnet-50,5710,5260,2501,000
deepseek/deepseek-v4-flash0,4850,4340,7500,250
google/gemini-3.6-flash0,4850,4520,2500,500
moonshotai/kimi-k30,7090,6540,2500,750
openai/gpt-5.6-sol0,7750,7280,2500,500
openai/gpt-5.6-terra0,6510,5370,5000,750
qwen/qwen3.7-max0,3330,2530,5000,500
x-ai/grok-4.50,6490,5300,0000,750
z-ai/glm-5.20,5000,4860,5000,500
Âşık — biçim ve jüri
ModelHece isabetiUyak isabetiBiçim ihlal oranıJüri puanı
anthropic/claude-opus-50,9400,894%5,50,741
anthropic/claude-sonnet-50,6410,562%23,20,457
deepseek/deepseek-v4-flash0,6020,531%27,60,404
google/gemini-3.6-flash0,5790,523%27,20,419
moonshotai/kimi-k30,7520,700%16,80,584
openai/gpt-5.6-sol0,8410,793%11,30,668
openai/gpt-5.6-terra0,6930,661%19,70,535
qwen/qwen3.7-max0,4740,427%34,80,309
x-ai/grok-4.50,6820,660%19,20,536
z-ai/glm-5.20,5900,539%28,60,416
⌐ Bu metrikler tarafsız özetleyici tarafından maç transkriptlerinden çıkarılır; puanlamaya girmez, yorumlayıcıdır (spec §6.1).
Hijyen Türkçe talimat izleme kapasitesi
ModelGeçersiz eylemDil ihlaliKararİstem tokenYanıt tokenMaliyet (USD)Bütçelenen tokenGeçersiz eylem oranıDil ihlali oranıOrt. istem tokenOrt. yanıt tokenOrt. bütçelenen token
anthropic/claude-opus-51641.7435.028.819473.233$36,974.309.928%0,9%0,22.8852722.473
anthropic/claude-sonnet-557101.7805.285.093478.339$23,034.577.888%3,2%0,62.9692692.572
deepseek/deepseek-v4-flash95141.8405.341.915479.470$1,814.612.150%5,2%0,82.9032612.507
google/gemini-3.6-flash59161.7455.073.077473.535$2,094.421.862%3,4%0,92.9072712.534
moonshotai/kimi-k33151.7895.066.842504.579$4,304.437.720%1,7%0,32.8322822.481
openai/gpt-5.6-sol2911.8515.195.626486.329$20,284.486.802%1,6%0,12.8072632.424
openai/gpt-5.6-terra42141.7665.254.284469.176$8,594.518.180%2,4%0,82.9752662.558
qwen/qwen3.7-max85191.6944.856.229437.032$5,284.207.021%5,0%1,12.8672582.483
x-ai/grok-4.550111.7214.907.182476.695$14,584.183.014%2,9%0,62.8512772.431
z-ai/glm-5.261121.8205.361.801481.921$2,924.622.203%3,4%0,72.9462652.540
⌐ Geçersiz eylem ve dil ihlali oranları, talimat izleme kapasitesinin doğrudan göstergeleridir (spec §4.2, §6). Bütçelenen istem tokenı, eşit bilişsel kaynak kuralının denetimidir: modeller arasında belirgin fark olmamalıdır.
Maliyet ve token kullanımı
ModelRolİstekİstem tokenYanıt tokenMaliyet (USD)
anthropic/claude-haiku-4.5jury100162.34635.160$0,0909
anthropic/claude-opus-5player1.7435.028.639473.240$36,97
anthropic/claude-sonnet-5player1.7805.285.181478.156$23,03
deepseek/deepseek-v4-flashplayer1.8405.341.972479.451$1,81
google/gemini-3.6-flashplayer1.7455.073.098473.535$2,09
google/gemini-3.6-flash-litesummarizer2.0855.181.930625.396$0,9345
moonshotai/kimi-k3player1.7895.066.915504.445$4,30
openai/gpt-5.6-minijury100158.09033.439$0,0875
openai/gpt-5.6-solplayer1.8515.195.665486.323$20,28
openai/gpt-5.6-terraplayer1.7665.254.311469.091$8,59
qwen/qwen3.7-maxplayer1.6944.856.121437.019$5,28
x-ai/grok-4.5player1.7214.907.141476.650$14,58
z-ai/glm-5.2player1.8205.361.806482.018$2,92
Yöntem Eki adalet tasarımı ve sınırlılıklar
Koşum adı
HCBfT-Games Tanıtım Koşumu
Ana tohum
20260801
Tarafsız özetleyici
google/gemini-3.6-flash-lite
Jüri modelleri
anthropic/claude-haiku-4.5, openai/gpt-5.6-mini
Sıcaklık
0.7
max_tokens
1024
Bağlam bütçesi
8000 token
Özet bütçesi
1200 token
Yakın pencere
2 faz
Tokenizer
cl100k_base
Eşzamanlılık
8 istek · 4 maç
Backend
openrouter
Toplam tur
4.236
Diplomat hedef n
20
Tüccar hedef n
20
Vezir hedef n
20
Âşık hedef n
20
Planlanan maçlar
Diplomat 50 · Tüccar 50 · Vezir 40 · Âşık 100
Adalet tasarımı. Koltuk ve rol rotasyonu blok temelli Latin kare ile üretildi; bir blok boyunca her model her koltuğu tam olarak bir kez işgal eder. Tahta tohumları farklı koltuk permütasyonlarıyla tekrar oynandı.

Eşit bilişsel kaynak. Tüm modeller için bağlam bütçesi 8000 token, özet bütçesi 1200 token, yakın pencere 2 faz; tek ortak tokenizer (cl100k_base) kullanıldı. Özetleme daima google/gemini-3.6-flash-lite tarafından yapıldı; hiçbir yarışmacı kendi geçmişini özetlemedi.

Örnekleme. Sıcaklık 0.7, max_tokens 1024.

Sınırlılıklar. HCB-100 mutlak bir yetenek ölçüsü değil, bu koşumdaki havuza göre göreli bir konumdur: havuz değişirse puanlar değişir. Az sayıda maçla üretilen puanlarda güven aralıkları geniştir ve sıralama farkları çoğunlukla anlamsızdır. Elo sıralaması yalnızca sağlamlık kontrolüdür.
Hücre sayıları — model × oyun
OyunModeln
Âşıkanthropic/claude-opus-520
Âşıkanthropic/claude-sonnet-520
Âşıkdeepseek/deepseek-v4-flash20
Âşıkgoogle/gemini-3.6-flash20
Âşıkmoonshotai/kimi-k320
Âşıkopenai/gpt-5.6-sol20
Âşıkopenai/gpt-5.6-terra20
Âşıkqwen/qwen3.7-max20
Âşıkx-ai/grok-4.520
Âşıkz-ai/glm-5.220
Diplomatanthropic/claude-opus-520
Diplomatanthropic/claude-sonnet-520
Diplomatdeepseek/deepseek-v4-flash20
Diplomatgoogle/gemini-3.6-flash20
Diplomatmoonshotai/kimi-k320
Diplomatopenai/gpt-5.6-sol20
Diplomatopenai/gpt-5.6-terra20
Diplomatqwen/qwen3.7-max20
Diplomatx-ai/grok-4.520
Diplomatz-ai/glm-5.220
Tüccaranthropic/claude-opus-520
Tüccaranthropic/claude-sonnet-520
Tüccardeepseek/deepseek-v4-flash20
Tüccargoogle/gemini-3.6-flash20
Tüccarmoonshotai/kimi-k320
Tüccaropenai/gpt-5.6-sol20
Tüccaropenai/gpt-5.6-terra20
Tüccarqwen/qwen3.7-max20
Tüccarx-ai/grok-4.520
Tüccarz-ai/glm-5.220
Veziranthropic/claude-opus-520
Veziranthropic/claude-sonnet-520
Vezirdeepseek/deepseek-v4-flash20
Vezirgoogle/gemini-3.6-flash20
Vezirmoonshotai/kimi-k320
Veziropenai/gpt-5.6-sol20
Veziropenai/gpt-5.6-terra20
Vezirqwen/qwen3.7-max20
Vezirx-ai/grok-4.520
Vezirz-ai/glm-5.220
Yapılandırma dökümü
{
  "run_name": "HCBfT-Games Tanıtım Koşumu",
  "models": [
    {
      "id": "anthropic/claude-opus-5",
      "temperature": null,
      "max_tokens": null,
      "price_in_per_mtok": null,
      "price_out_per_mtok": null
    },
    {
      "id": "openai/gpt-5.6-sol",
      "temperature": null,
      "max_tokens": null,
      "price_in_per_mtok": null,
      "price_out_per_mtok": null
    },
    {
      "id": "moonshotai/kimi-k3",
      "temperature": null,
      "max_tokens": null,
      "price_in_per_mtok": null,
      "price_out_per_mtok": null
    },
    {
      "id": "openai/gpt-5.6-terra",
      "temperature": null,
      "max_tokens": null,
      "price_in_per_mtok": null,
      "price_out_per_mtok": null
    },
    {
      "id": "x-ai/grok-4.5",
      "temperature": null,
      "max_tokens": null,
      "price_in_per_mtok": null,
      "price_out_per_mtok": null
    },
    {
      "id": "anthropic/claude-sonnet-5",
      "temperature": null,
      "max_tokens": null,
      "price_in_per_mtok": null,
      "price_out_per_mtok": null
    },
    {
      "id": "z-ai/glm-5.2",
      "temperature": null,
      "max_tokens": null,
      "price_in_per_mtok": null,
      "price_out_per_mtok": null
    },
    {
      "id": "google/gemini-3.6-flash",
      "temperature": null,
      "max_tokens": null,
      "price_in_per_mtok": null,
      "price_out_per_mtok": null
    },
    {
      "id": "deepseek/deepseek-v4-flash",
      "temperature": null,
      "max_tokens": null,
      "price_in_per_mtok": null,
      "price_out_per_mtok": null
    },
    {
      "id": "qwen/qwen3.7-max",
      "temperature": null,
      "max_tokens": null,
      "price_in_per_mtok": null,
      "price_out_per_mtok": null
    }
  ],
  "sampling": {
    "temperature": 0.7,
    "max_tokens": 1024,
    "top_p": null
  },
  "utility_model": "google/gemini-3.6-flash-lite",
  "jury_models": [
    "anthropic/claude-haiku-4.5",
    "openai/gpt-5.6-mini"
  ],
  "games": {
    "diplomat": {
      "enabled": true,
      "games_per_model": 20,
      "turn_cap": 16,
      "negotiation_rounds": 2,
      "messages_per_round": 2,
      "win_centers": 7,
      "solo_bonus": 0.25,
      "trade_rounds": null,
      "target_vp": 8,
      "win_bonus": 0.25,
      "discussion_cycles": 2,
      "assassin_bonus": 0.5,
      "bilge_bonus": 0.25,
      "exchanges": 3,
      "hece": 11,
      "form_weight": 0.4,
      "jury_weight": 0.6
    },
    "tuccar": {
      "enabled": true,
      "games_per_model": 20,
      "turn_cap": 60,
      "negotiation_rounds": 2,
      "messages_per_round": 2,
      "win_centers": 7,
      "solo_bonus": 0.25,
      "trade_rounds": 2,
      "target_vp": 8,
      "win_bonus": 0.25,
      "discussion_cycles": 2,
      "assassin_bonus": 0.5,
      "bilge_bonus": 0.25,
      "exchanges": 3,
      "hece": 11,
      "form_weight": 0.4,
      "jury_weight": 0.6
    },
    "vezir": {
      "enabled": true,
      "games_per_model": 20,
      "turn_cap": null,
      "negotiation_rounds": null,
      "messages_per_round": 2,
      "win_centers": 7,
      "solo_bonus": 0.25,
      "trade_rounds": null,
      "target_vp": 8,
      "win_bonus": 0.25,
      "discussion_cycles": 2,
      "assassin_bonus": 0.5,
      "bilge_bonus": 0.25,
      "exchanges": 3,
      "hece": 11,
      "form_weight": 0.4,
      "jury_weight": 0.6
    },
    "asik": {
      "enabled": true,
      "games_per_model": 20,
      "turn_cap": null,
      "negotiation_rounds": null,
      "messages_per_round": 2,
      "win_centers": 7,
      "solo_bonus": 0.25,
      "trade_rounds": null,
      "target_vp": 8,
      "win_bonus": 0.25,
      "discussion_cycles": 2,
      "assassin_bonus": 0.5,
      "bilge_bonus": 0.25,
      "exchanges": 3,
      "hece": 11,
      "form_weight": 0.4,
      "jury_weight": 0.6
    }
  },
  "context": {
    "context_budget_tokens": 8000,
    "digest_budget_tokens": 1200,
    "recent_window_phases": 2,
    "compress_every_phases": 1,
    "message_token_limit": 150,
    "scratchpad_token_limit": 200,
    "tokenizer": "cl100k_base"
  },
  "limits": {
    "concurrency": 8,
    "per_model_concurrency": 4,
    "max_run_cost_usd": 200.0,
    "invalid_action_retries": 2,
    "max_api_attempts": 8,
    "request_timeout_s": 180.0,
    "match_concurrency": 4
  },
  "jury": {
    "enabled": true,
    "identity_denylist": [
      "claude",
      "anthropic",
      "gpt",
      "openai",
      "chatgpt",
      "gemini",
      "google",
      "bard",
      "qwen",
      "alibaba",
      "llama",
      "meta ai",
      "mistral",
      "deepseek",
      "grok",
      "yapay zeka",
      "dil modeli",
      "language model"
    ]
  },
  "seed": 20260801,
  "output_dir": "runs",
  "mock": false,
  "behavioral_analytics": true,
  "language_check": true,
  "language_check_min_chars": 25,
  "composite_weights": {
    "diplomat": 1.0,
    "tuccar": 1.0,
    "vezir": 1.0,
    "asik": 1.0
  }
}