Digital Intelligence Institute Game Scorecard
Turkish Multi-Agent Game Evaluation for Large Language Models · 2026-08-01 16:21 UTC
HCBfT-Games
High-Cognitive Benchmark for Turkish  ·  A multi-agent evaluation measuring strategic reasoning, negotiation, resistance to deception, and creative Turkish generation across the Diplomat, Tüccar, Vezir, and Âşık games
Run demo-hcbft-games  ·  seed 20260801  ·  neutral summarizer google/gemini-3.6-flash-lite  ·  jury anthropic/claude-haiku-4.5, openai/gpt-5.6-mini
Competing Models
10
4 games · 240 matches
Leader
109.6
claude-opus-5
Composite Average
100.0
range 91.0–109.6 · scale 100 ± 15
Total Matches
240
23.5 h of total playtime
Invalid Actions
2.96%
highest: deepseek-v4-flash 5.16%
Total Cost
$120.98
20,034 requests · 62.3M tokens
Game Scorecards overall score and ranking per game
Diplomat
Diplomacy-inspired
Strategic communication · alliance and betrayal
110.0
Highest HCB-100 claude-opus-5
01 claude-opus-5 110.0
02 gpt-5.6-sol 106.6
03 kimi-k3 104.9
04 gpt-5.6-terra 101.5
05 claude-sonnet-5 99.8
06 grok-4.5 99.8
07 glm-5.2 96.4
08 deepseek-v4-flash 94.8
09 gemini-3.6-flash 94.8
10 qwen3.7-max 91.4
50 matches 10 models cell n 20/20 range 91.4–110.0
Tüccar
Catan-inspired
Utility analysis · bargaining and trade
109.8
Highest HCB-100 claude-opus-5
01 claude-opus-5 109.8
02 gpt-5.6-sol 106.7
03 kimi-k3 103.7
04 gpt-5.6-terra 102.1
05 grok-4.5 100.6
06 claude-sonnet-5 99.1
07 glm-5.2 96.0
08 deepseek-v4-flash 96.0
09 gemini-3.6-flash 96.0
10 qwen3.7-max 89.9
50 matches 10 models cell n 20/20 range 89.9–109.8
Vezir
Avalon-inspired
Deception · consistency and inference
110.3
Highest HCB-100 claude-opus-5
01 claude-opus-5 110.3
02 gpt-5.6-sol 107.5
03 kimi-k3 104.7
04 gpt-5.6-terra 101.7
05 grok-4.5 100.6
06 claude-sonnet-5 99.3
07 glm-5.2 96.4
08 gemini-3.6-flash 95.1
09 deepseek-v4-flash 94.9
10 qwen3.7-max 89.5
40 matches 10 models cell n 20/20 range 89.5–110.3
Âşık
Minstrel Duel
Creative Turkish · meter and rhyme
108.4
Highest HCB-100 claude-opus-5
01 claude-opus-5 108.4
02 gpt-5.6-sol 104.6
03 kimi-k3 102.7
04 grok-4.5 100.8
05 gpt-5.6-terra 100.8
06 claude-sonnet-5 98.9
07 deepseek-v4-flash 96.9
08 gemini-3.6-flash 96.9
09 glm-5.2 96.9
10 qwen3.7-max 93.1
100 matches 10 models cell n 20/20 range 93.1–108.4
⌐ The large number is the highest HCB-100 score in that game. The bars place models between the run's lowest and highest score; the 100 line is the pool average. In Vezir, standardization is done within each role class (spec §8.2).
Overall Ranking composite HCB-100
01
claude-opus-5 anthropic/claude-opus-5
109.6
GA 106.4–112.7
Diplomat 110.0Tüccar 109.8Vezir 110.3Âşık 108.4
02
gpt-5.6-sol openai/gpt-5.6-sol
106.3
GA 103.5–109.3
Diplomat 106.6Tüccar 106.7Vezir 107.5Âşık 104.6
03
kimi-k3 moonshotai/kimi-k3
104.0
GA 100.9–106.8
Diplomat 104.9Tüccar 103.7Vezir 104.7Âşık 102.7
04
gpt-5.6-terra openai/gpt-5.6-terra
101.5
GA 98.6–104.5
Diplomat 101.5Tüccar 102.1Vezir 101.7Âşık 100.8
05
grok-4.5 x-ai/grok-4.5
100.5
GA 97.7–103.0
Diplomat 99.8Tüccar 100.6Vezir 100.6Âşık 100.8
06
claude-sonnet-5 anthropic/claude-sonnet-5
99.3
GA 96.0–102.5
Diplomat 99.8Tüccar 99.1Vezir 99.3Âşık 98.9
07
glm-5.2 z-ai/glm-5.2
96.5
GA 93.6–99.2
Diplomat 96.4Tüccar 96.0Vezir 96.4Âşık 96.9
08
gemini-3.6-flash google/gemini-3.6-flash
95.7
GA 92.9–98.8
Diplomat 94.8Tüccar 96.0Vezir 95.1Âşık 96.9
09
deepseek-v4-flash deepseek/deepseek-v4-flash
95.6
GA 92.6–99.0
Diplomat 94.8Tüccar 96.0Vezir 94.9Âşık 96.9
10
qwen3.7-max qwen/qwen3.7-max
91.0
GA 88.3–93.8
Diplomat 91.4Tüccar 89.9Vezir 89.5Âşık 93.1
Elo — robustness check
GameModelElonWin rate
Diplomatanthropic/claude-opus-51,623.12055.0%
Diplomatmoonshotai/kimi-k31,581.42035.0%
Diplomatopenai/gpt-5.6-sol1,569.72030.0%
Diplomatanthropic/claude-sonnet-51,527.22030.0%
Diplomatx-ai/grok-4.51,498.22015.0%
Diplomatopenai/gpt-5.6-terra1,469.82015.0%
Diplomatdeepseek/deepseek-v4-flash1,463.52020.0%
Diplomatgoogle/gemini-3.6-flash1,437.82015.0%
Diplomatqwen/qwen3.7-max1,426.32010.0%
Diplomatz-ai/glm-5.21,403.02025.0%
Tüccaropenai/gpt-5.6-sol1,632.62040.0%
Tüccaranthropic/claude-opus-51,621.32035.0%
Tüccarmoonshotai/kimi-k31,561.22030.0%
Tüccarx-ai/grok-4.51,535.12020.0%
Tüccardeepseek/deepseek-v4-flash1,472.92025.0%
Tüccaropenai/gpt-5.6-terra1,460.22025.0%
Tüccarz-ai/glm-5.21,448.92025.0%
Tüccaranthropic/claude-sonnet-51,446.02025.0%
Tüccargoogle/gemini-3.6-flash1,432.92015.0%
Tüccarqwen/qwen3.7-max1,388.82010.0%
Veziranthropic/claude-opus-51,649.12070.0%
Veziropenai/gpt-5.6-sol1,619.02055.0%
Vezirmoonshotai/kimi-k31,543.12040.0%
Vezirx-ai/grok-4.51,541.52045.0%
Veziropenai/gpt-5.6-terra1,515.42045.0%
Veziranthropic/claude-sonnet-51,474.72035.0%
Vezirdeepseek/deepseek-v4-flash1,456.52040.0%
Vezirz-ai/glm-5.21,451.62020.0%
Vezirgoogle/gemini-3.6-flash1,436.82035.0%
Vezirqwen/qwen3.7-max1,312.32015.0%
Âşıkmoonshotai/kimi-k31,639.72070.0%
Âşıkopenai/gpt-5.6-sol1,605.22065.0%
Âşıkx-ai/grok-4.51,565.62060.0%
Âşıkanthropic/claude-opus-51,557.52055.0%
Âşıkopenai/gpt-5.6-terra1,533.62055.0%
Âşıkz-ai/glm-5.21,497.32050.0%
Âşıkdeepseek/deepseek-v4-flash1,470.82045.0%
Âşıkqwen/qwen3.7-max1,408.62040.0%
Âşıkgoogle/gemini-3.6-flash1,367.42030.0%
Âşıkanthropic/claude-sonnet-51,354.42030.0%
z-score and Elo rankings disagree: diplomat, tuccar, vezir, asik. This disagreement is worth reporting in its own right (spec §8.2).
Diplomat Strategic communication · alliance and betrayal
Raw score and 95% bootstrap confidence interval
ModelnRaw averageCI lowerCI upperStd. errorzHCB-100
anthropic/claude-opus-5200.3700.3290.4110.0210.665110.0
openai/gpt-5.6-sol200.3500.3100.3890.0200.440106.6
moonshotai/kimi-k3200.3400.3070.3730.0170.327104.9
openai/gpt-5.6-terra200.3200.2830.3570.0190.101101.5
anthropic/claude-sonnet-5200.3100.2670.3510.022-0.01199.8
x-ai/grok-4.5200.3100.2800.3420.016-0.01199.8
z-ai/glm-5.2200.2900.2540.3250.018-0.23796.4
deepseek/deepseek-v4-flash200.2800.2470.3140.017-0.35094.8
google/gemini-3.6-flash200.2800.2450.3160.018-0.35094.8
qwen/qwen3.7-max200.2600.2310.2890.015-0.57591.4
Seat-based average — balance check
ModelSeatnRaw average
anthropic/claude-opus-5050.421
anthropic/claude-opus-5150.299
anthropic/claude-opus-5250.384
anthropic/claude-opus-5350.377
anthropic/claude-sonnet-5050.363
anthropic/claude-sonnet-5150.356
anthropic/claude-sonnet-5250.300
anthropic/claude-sonnet-5350.220
deepseek/deepseek-v4-flash050.280
deepseek/deepseek-v4-flash150.320
deepseek/deepseek-v4-flash250.322
deepseek/deepseek-v4-flash350.198
google/gemini-3.6-flash050.324
google/gemini-3.6-flash150.252
google/gemini-3.6-flash250.270
google/gemini-3.6-flash350.273
moonshotai/kimi-k3050.318
moonshotai/kimi-k3150.335
moonshotai/kimi-k3250.343
moonshotai/kimi-k3350.364
openai/gpt-5.6-sol050.358
openai/gpt-5.6-sol150.359
openai/gpt-5.6-sol250.360
openai/gpt-5.6-sol350.324
openai/gpt-5.6-terra050.315
openai/gpt-5.6-terra150.369
openai/gpt-5.6-terra250.304
openai/gpt-5.6-terra350.293
qwen/qwen3.7-max050.286
qwen/qwen3.7-max150.257
qwen/qwen3.7-max250.217
qwen/qwen3.7-max350.280
x-ai/grok-4.5050.343
x-ai/grok-4.5150.322
x-ai/grok-4.5250.286
x-ai/grok-4.5350.289
z-ai/glm-5.2050.329
z-ai/glm-5.2150.321
z-ai/glm-5.2250.243
z-ai/glm-5.2350.267
Secondary metrics — model average
ModelCenterUnitMessage countAnnouncement countMove countSupport countMove success rateEliminatedPromises madePromise-keeping rateBetrayal rate
anthropic/claude-opus-55.5505.3008212572.8%0.0004.60082.3%18.9%
anthropic/claude-sonnet-55.1004.7508313450.7%0.0504.95055.7%44.5%
deepseek/deepseek-v4-flash5.0504.6008312442.5%0.0004.65039.1%60.5%
google/gemini-3.6-flash5.0504.8508213543.4%0.0004.35041.6%59.3%
moonshotai/kimi-k35.4504.9008212465.1%0.0004.50070.6%29.9%
openai/gpt-5.6-sol5.2005.2008212466.8%0.0004.70077.2%21.9%
openai/gpt-5.6-terra5.3505.1008313458.9%0.0004.55060.3%40.3%
qwen/qwen3.7-max4.9504.5008313434.6%0.0004.40029.6%71.0%
x-ai/grok-4.55.1505.0008212553.9%0.0004.15052.8%47.4%
z-ai/glm-5.25.0004.7009312443.9%0.0004.95042.4%57.8%
Pairwise significance matrix — permutation test, Holm–Bonferroni
Model AModel Bn (A)n (B)Differencepp (Holm)Sufficient nSignificance
anthropic/claude-opus-5anthropic/claude-sonnet-520200.0600.0581.000
anthropic/claude-opus-5deepseek/deepseek-v4-flash20200.0900.0030.118
anthropic/claude-opus-5google/gemini-3.6-flash20200.0900.0030.123
anthropic/claude-opus-5moonshotai/kimi-k320200.0300.2901.000
anthropic/claude-opus-5openai/gpt-5.6-sol20200.0200.5071.000
anthropic/claude-opus-5openai/gpt-5.6-terra20200.0500.0901.000
anthropic/claude-opus-5qwen/qwen3.7-max20200.1100.0010.027*
anthropic/claude-opus-5x-ai/grok-4.520200.0600.0351.000
anthropic/claude-opus-5z-ai/glm-5.220200.0800.0060.224
anthropic/claude-sonnet-5deepseek/deepseek-v4-flash20200.0300.2981.000
anthropic/claude-sonnet-5google/gemini-3.6-flash20200.0300.3091.000
anthropic/claude-sonnet-5moonshotai/kimi-k32020-0.0300.2971.000
anthropic/claude-sonnet-5openai/gpt-5.6-sol2020-0.0400.1981.000
anthropic/claude-sonnet-5openai/gpt-5.6-terra2020-0.0100.7331.000
anthropic/claude-sonnet-5qwen/qwen3.7-max20200.0500.0781.000
anthropic/claude-sonnet-5x-ai/grok-4.520200.0001.0001.000
anthropic/claude-sonnet-5z-ai/glm-5.220200.0200.4981.000
deepseek/deepseek-v4-flashgoogle/gemini-3.6-flash20200.0001.0001.000
deepseek/deepseek-v4-flashmoonshotai/kimi-k32020-0.0600.0220.829
deepseek/deepseek-v4-flashopenai/gpt-5.6-sol2020-0.0700.0150.577
deepseek/deepseek-v4-flashopenai/gpt-5.6-terra2020-0.0400.1341.000
deepseek/deepseek-v4-flashqwen/qwen3.7-max20200.0200.3951.000
deepseek/deepseek-v4-flashx-ai/grok-4.52020-0.0300.2081.000
deepseek/deepseek-v4-flashz-ai/glm-5.22020-0.0100.6891.000
google/gemini-3.6-flashmoonshotai/kimi-k32020-0.0600.0260.896
google/gemini-3.6-flashopenai/gpt-5.6-sol2020-0.0700.0170.638
google/gemini-3.6-flashopenai/gpt-5.6-terra2020-0.0400.1411.000
google/gemini-3.6-flashqwen/qwen3.7-max20200.0200.4171.000
google/gemini-3.6-flashx-ai/grok-4.52020-0.0300.2291.000
google/gemini-3.6-flashz-ai/glm-5.22020-0.0100.6991.000
moonshotai/kimi-k3openai/gpt-5.6-sol2020-0.0100.7151.000
moonshotai/kimi-k3openai/gpt-5.6-terra20200.0200.4451.000
moonshotai/kimi-k3qwen/qwen3.7-max20200.0800.0020.069
moonshotai/kimi-k3x-ai/grok-4.520200.0300.2171.000
moonshotai/kimi-k3z-ai/glm-5.220200.0500.0551.000
openai/gpt-5.6-solopenai/gpt-5.6-terra20200.0300.2951.000
openai/gpt-5.6-solqwen/qwen3.7-max20200.0900.0010.062
openai/gpt-5.6-solx-ai/grok-4.520200.0400.1411.000
openai/gpt-5.6-solz-ai/glm-5.220200.0600.0341.000
openai/gpt-5.6-terraqwen/qwen3.7-max20200.0600.0220.829
openai/gpt-5.6-terrax-ai/grok-4.520200.0100.6991.000
openai/gpt-5.6-terraz-ai/glm-5.220200.0300.2581.000
qwen/qwen3.7-maxx-ai/grok-4.52020-0.0500.0311.000
qwen/qwen3.7-maxz-ai/glm-5.22020-0.0300.2121.000
x-ai/grok-4.5z-ai/glm-5.220200.0200.4151.000
⌐ Stars: *** p<0.001 · ** p<0.01 · * p<0.05 · — not significant. No star is printed if the cell n is below target.
Tüccar Utility analysis · bargaining and trade
Raw score and 95% bootstrap confidence interval
ModelnRaw averageCI lowerCI upperStd. errorzHCB-100
anthropic/claude-opus-5200.4100.3730.4460.0190.652109.8
openai/gpt-5.6-sol200.3900.3520.4290.0200.448106.7
moonshotai/kimi-k3200.3700.3340.4070.0190.245103.7
openai/gpt-5.6-terra200.3600.3180.4020.0210.143102.1
x-ai/grok-4.5200.3500.3150.3860.0180.041100.6
anthropic/claude-sonnet-5200.3400.2980.3840.022-0.06199.1
z-ai/glm-5.2200.3200.2810.3590.020-0.26596.0
deepseek/deepseek-v4-flash200.3200.2780.3640.022-0.26596.0
google/gemini-3.6-flash200.3200.2730.3660.024-0.26596.0
qwen/qwen3.7-max200.2800.2440.3170.018-0.67289.9
Seat-based average — balance check
ModelSeatnRaw average
anthropic/claude-opus-5050.442
anthropic/claude-opus-5150.392
anthropic/claude-opus-5250.407
anthropic/claude-opus-5350.398
anthropic/claude-sonnet-5050.379
anthropic/claude-sonnet-5150.329
anthropic/claude-sonnet-5250.315
anthropic/claude-sonnet-5350.338
deepseek/deepseek-v4-flash050.306
deepseek/deepseek-v4-flash150.307
deepseek/deepseek-v4-flash250.321
deepseek/deepseek-v4-flash350.346
google/gemini-3.6-flash050.396
google/gemini-3.6-flash150.323
google/gemini-3.6-flash250.285
google/gemini-3.6-flash350.276
moonshotai/kimi-k3050.336
moonshotai/kimi-k3150.396
moonshotai/kimi-k3250.401
moonshotai/kimi-k3350.347
openai/gpt-5.6-sol050.419
openai/gpt-5.6-sol150.428
openai/gpt-5.6-sol250.357
openai/gpt-5.6-sol350.356
openai/gpt-5.6-terra050.408
openai/gpt-5.6-terra150.288
openai/gpt-5.6-terra250.371
openai/gpt-5.6-terra350.373
qwen/qwen3.7-max050.317
qwen/qwen3.7-max150.248
qwen/qwen3.7-max250.244
qwen/qwen3.7-max350.311
x-ai/grok-4.5050.351
x-ai/grok-4.5150.326
x-ai/grok-4.5250.372
x-ai/grok-4.5350.351
z-ai/glm-5.2050.318
z-ai/glm-5.2150.368
z-ai/glm-5.2250.332
z-ai/glm-5.2350.262
Secondary metrics — model average
ModelVictory pointsOffers madeOffers acceptedOffer acceptance rateOffers it acceptedTrade surplusBank tradesBank dependencyResources producedLongest road
anthropic/claude-opus-55.95010.0005.80058.2%3.8001.7690.5500.06115.4003.350
anthropic/claude-sonnet-55.45010.9003.85035.6%3.6000.6351.6000.17716.4502.200
deepseek/deepseek-v4-flash5.20010.0502.65026.3%3.2000.3012.5000.27816.5001.900
google/gemini-3.6-flash5.40010.5503.05029.4%3.1500.3012.1000.24616.6502.050
moonshotai/kimi-k35.7509.9004.75047.8%3.3501.1601.4000.14416.6502.950
openai/gpt-5.6-sol6.00011.4506.10053.0%4.1001.4260.7000.06515.3503.250
openai/gpt-5.6-terra5.5009.8504.10042.1%3.2500.9321.5000.17715.2502.650
qwen/qwen3.7-max4.95010.0501.55015.5%3.000-0.3762.2500.31215.6001.650
x-ai/grok-4.55.40010.0003.65036.6%3.5000.8221.6000.17815.5502.800
z-ai/glm-5.25.25011.0003.10028.3%3.1000.3452.1000.25812.5502.050
Pairwise significance matrix — permutation test, Holm–Bonferroni
Model AModel Bn (A)n (B)Differencepp (Holm)Sufficient nSignificance
anthropic/claude-opus-5anthropic/claude-sonnet-520200.0700.0240.833
anthropic/claude-opus-5deepseek/deepseek-v4-flash20200.0900.0050.197
anthropic/claude-opus-5google/gemini-3.6-flash20200.0900.0050.216
anthropic/claude-opus-5moonshotai/kimi-k320200.0400.1441.000
anthropic/claude-opus-5openai/gpt-5.6-sol20200.0200.4871.000
anthropic/claude-opus-5openai/gpt-5.6-terra20200.0500.0911.000
anthropic/claude-opus-5qwen/qwen3.7-max20200.1300.0000.009**
anthropic/claude-opus-5x-ai/grok-4.520200.0600.0311.000
anthropic/claude-opus-5z-ai/glm-5.220200.0900.0030.126
anthropic/claude-sonnet-5deepseek/deepseek-v4-flash20200.0200.5321.000
anthropic/claude-sonnet-5google/gemini-3.6-flash20200.0200.5521.000
anthropic/claude-sonnet-5moonshotai/kimi-k32020-0.0300.3261.000
anthropic/claude-sonnet-5openai/gpt-5.6-sol2020-0.0500.1081.000
anthropic/claude-sonnet-5openai/gpt-5.6-terra2020-0.0200.5311.000
anthropic/claude-sonnet-5qwen/qwen3.7-max20200.0600.0481.000
anthropic/claude-sonnet-5x-ai/grok-4.52020-0.0100.7511.000
anthropic/claude-sonnet-5z-ai/glm-5.220200.0200.5161.000
deepseek/deepseek-v4-flashgoogle/gemini-3.6-flash20200.0001.0001.000
deepseek/deepseek-v4-flashmoonshotai/kimi-k32020-0.0500.1071.000
deepseek/deepseek-v4-flashopenai/gpt-5.6-sol2020-0.0700.0230.828
deepseek/deepseek-v4-flashopenai/gpt-5.6-terra2020-0.0400.2011.000
deepseek/deepseek-v4-flashqwen/qwen3.7-max20200.0400.1811.000
deepseek/deepseek-v4-flashx-ai/grok-4.52020-0.0300.3051.000
deepseek/deepseek-v4-flashz-ai/glm-5.22020-0.0001.0001.000
google/gemini-3.6-flashmoonshotai/kimi-k32020-0.0500.1191.000
google/gemini-3.6-flashopenai/gpt-5.6-sol2020-0.0700.0321.000
google/gemini-3.6-flashopenai/gpt-5.6-terra2020-0.0400.2361.000
google/gemini-3.6-flashqwen/qwen3.7-max20200.0400.2001.000
google/gemini-3.6-flashx-ai/grok-4.52020-0.0300.3421.000
google/gemini-3.6-flashz-ai/glm-5.22020-0.0001.0001.000
moonshotai/kimi-k3openai/gpt-5.6-sol2020-0.0200.4721.000
moonshotai/kimi-k3openai/gpt-5.6-terra20200.0100.7271.000
moonshotai/kimi-k3qwen/qwen3.7-max20200.0900.0020.086
moonshotai/kimi-k3x-ai/grok-4.520200.0200.4381.000
moonshotai/kimi-k3z-ai/glm-5.220200.0500.0851.000
openai/gpt-5.6-solopenai/gpt-5.6-terra20200.0300.3041.000
openai/gpt-5.6-solqwen/qwen3.7-max20200.1100.0010.026*
openai/gpt-5.6-solx-ai/grok-4.520200.0400.1491.000
openai/gpt-5.6-solz-ai/glm-5.220200.0700.0200.740
openai/gpt-5.6-terraqwen/qwen3.7-max20200.0800.0080.328
openai/gpt-5.6-terrax-ai/grok-4.520200.0100.7361.000
openai/gpt-5.6-terraz-ai/glm-5.220200.0400.1981.000
qwen/qwen3.7-maxx-ai/grok-4.52020-0.0700.0120.456
qwen/qwen3.7-maxz-ai/glm-5.22020-0.0400.1531.000
x-ai/grok-4.5z-ai/glm-5.220200.0300.2891.000
⌐ Stars: *** p<0.001 · ** p<0.01 · * p<0.05 · — not significant. No star is printed if the cell n is below target.
Vezir Deception · consistency and inference
Raw score and 95% bootstrap confidence interval
ModelnRaw averageCI lowerCI upperStd. errorzHCB-100
anthropic/claude-opus-5200.5000.4530.5450.0230.688110.3
openai/gpt-5.6-sol200.4800.4350.5210.0220.499107.5
moonshotai/kimi-k3200.4600.4120.5080.0250.313104.7
openai/gpt-5.6-terra200.4400.3940.4840.0230.115101.7
x-ai/grok-4.5200.4300.3900.4710.0210.041100.6
anthropic/claude-sonnet-5200.4200.3730.4700.025-0.04899.3
z-ai/glm-5.2200.4000.3690.4300.016-0.24096.4
google/gemini-3.6-flash200.3900.3510.4280.020-0.32695.1
deepseek/deepseek-v4-flash200.3900.3360.4410.027-0.34394.9
qwen/qwen3.7-max200.3500.3040.3930.023-0.69989.5
Seat-based average — balance check
ModelSeatnRaw average
anthropic/claude-opus-5040.442
anthropic/claude-opus-5140.492
anthropic/claude-opus-5240.556
anthropic/claude-opus-5340.534
anthropic/claude-opus-5440.476
anthropic/claude-sonnet-5040.453
anthropic/claude-sonnet-5140.476
anthropic/claude-sonnet-5240.367
anthropic/claude-sonnet-5340.408
anthropic/claude-sonnet-5440.397
deepseek/deepseek-v4-flash040.383
deepseek/deepseek-v4-flash140.358
deepseek/deepseek-v4-flash240.488
deepseek/deepseek-v4-flash340.396
deepseek/deepseek-v4-flash440.325
google/gemini-3.6-flash040.461
google/gemini-3.6-flash140.317
google/gemini-3.6-flash240.428
google/gemini-3.6-flash340.353
google/gemini-3.6-flash440.392
moonshotai/kimi-k3040.587
moonshotai/kimi-k3140.455
moonshotai/kimi-k3240.396
moonshotai/kimi-k3340.455
moonshotai/kimi-k3440.408
openai/gpt-5.6-sol040.538
openai/gpt-5.6-sol140.443
openai/gpt-5.6-sol240.434
openai/gpt-5.6-sol340.450
openai/gpt-5.6-sol440.535
openai/gpt-5.6-terra040.404
openai/gpt-5.6-terra140.449
openai/gpt-5.6-terra240.506
openai/gpt-5.6-terra340.374
openai/gpt-5.6-terra440.467
qwen/qwen3.7-max040.349
qwen/qwen3.7-max140.313
qwen/qwen3.7-max240.373
qwen/qwen3.7-max340.365
qwen/qwen3.7-max440.349
x-ai/grok-4.5040.459
x-ai/grok-4.5140.497
x-ai/grok-4.5240.476
x-ai/grok-4.5340.372
x-ai/grok-4.5440.346
z-ai/glm-5.2040.387
z-ai/glm-5.2140.370
z-ai/glm-5.2240.432
z-ai/glm-5.2340.389
z-ai/glm-5.2440.421
Role-class average and within-role z
ModelRole classnRaw averagezHCB-100
anthropic/claude-opus-5sage40.5270.819112.3
anthropic/claude-opus-5good-plain80.4530.410106.2
anthropic/claude-opus-5evil-plain40.5120.992114.9
anthropic/claude-opus-5assassin40.5540.806112.1
anthropic/claude-sonnet-5sage40.330-0.81287.8
anthropic/claude-sonnet-5good-plain80.4360.248103.7
anthropic/claude-sonnet-5evil-plain40.344-0.72389.2
anthropic/claude-sonnet-5assassin40.5530.800112.0
deepseek/deepseek-v4-flashsage40.428-0.001100.0
deepseek/deepseek-v4-flashgood-plain80.407-0.03299.5
deepseek/deepseek-v4-flashevil-plain40.349-0.67589.9
deepseek/deepseek-v4-flashassassin40.359-0.97585.4
google/gemini-3.6-flashsage40.347-0.67289.9
google/gemini-3.6-flashgood-plain80.373-0.36694.5
google/gemini-3.6-flashevil-plain40.397-0.18297.3
google/gemini-3.6-flashassassin40.461-0.04599.3
moonshotai/kimi-k3sage40.4440.132102.0
moonshotai/kimi-k3good-plain80.4150.045100.7
moonshotai/kimi-k3evil-plain40.4380.237103.6
moonshotai/kimi-k3assassin40.5871.106116.6
openai/gpt-5.6-solsage40.5380.907113.6
openai/gpt-5.6-solgood-plain80.4580.454106.8
openai/gpt-5.6-solevil-plain40.4870.730110.9
openai/gpt-5.6-solassassin40.460-0.05199.2
openai/gpt-5.6-terrasage40.5210.767111.5
openai/gpt-5.6-terragood-plain80.385-0.25196.2
openai/gpt-5.6-terraevil-plain40.4560.418106.3
openai/gpt-5.6-terraassassin40.454-0.10998.4
qwen/qwen3.7-maxsage40.297-1.08083.8
qwen/qwen3.7-maxgood-plain80.342-0.66090.1
qwen/qwen3.7-maxevil-plain40.349-0.67089.9
qwen/qwen3.7-maxassassin40.419-0.42593.6
x-ai/grok-4.5sage40.4610.277104.2
x-ai/grok-4.5good-plain80.4140.032100.5
x-ai/grok-4.5evil-plain40.4590.452106.8
x-ai/grok-4.5assassin40.401-0.58791.2
z-ai/glm-5.2sage40.387-0.33794.9
z-ai/glm-5.2good-plain80.4230.119101.8
z-ai/glm-5.2evil-plain40.358-0.57991.3
z-ai/glm-5.2assassin40.409-0.52092.2
Secondary metrics — model average
ModelVote countTeam proposedVote accuracySage leakDeception successFailed cardAssassination accuracy
anthropic/claude-opus-591.9500.8410.2500.7620.8750.500
anthropic/claude-sonnet-591.6000.5710.2500.5261.7501.000
deepseek/deepseek-v4-flash91.4000.4850.7500.4341.7500.250
google/gemini-3.6-flash91.3000.4850.2500.4521.3750.500
moonshotai/kimi-k391.8000.7090.2500.6541.2500.750
openai/gpt-5.6-sol91.3000.7750.2500.7281.5000.500
openai/gpt-5.6-terra91.6000.6510.5000.5371.6250.750
qwen/qwen3.7-max91.3000.3330.5000.2530.7500.500
x-ai/grok-4.591.1500.6490.0000.5301.1250.750
z-ai/glm-5.2101.4500.5000.5000.4861.2500.500
Pairwise significance matrix — permutation test, Holm–Bonferroni
Model AModel Bn (A)n (B)Differencepp (Holm)Sufficient nSignificance
anthropic/claude-opus-5anthropic/claude-sonnet-520200.0800.0260.870
anthropic/claude-opus-5deepseek/deepseek-v4-flash20200.1100.0040.160
anthropic/claude-opus-5google/gemini-3.6-flash20200.1100.0020.101
anthropic/claude-opus-5moonshotai/kimi-k320200.0400.2531.000
anthropic/claude-opus-5openai/gpt-5.6-sol20200.0200.5561.000
anthropic/claude-opus-5openai/gpt-5.6-terra20200.0600.0811.000
anthropic/claude-opus-5qwen/qwen3.7-max20200.1500.0000.018*
anthropic/claude-opus-5x-ai/grok-4.520200.0700.0391.000
anthropic/claude-opus-5z-ai/glm-5.220200.1000.0020.086
anthropic/claude-sonnet-5deepseek/deepseek-v4-flash20200.0300.4191.000
anthropic/claude-sonnet-5google/gemini-3.6-flash20200.0300.3561.000
anthropic/claude-sonnet-5moonshotai/kimi-k32020-0.0400.2691.000
anthropic/claude-sonnet-5openai/gpt-5.6-sol2020-0.0600.0821.000
anthropic/claude-sonnet-5openai/gpt-5.6-terra2020-0.0200.5641.000
anthropic/claude-sonnet-5qwen/qwen3.7-max20200.0700.0511.000
anthropic/claude-sonnet-5x-ai/grok-4.52020-0.0100.7591.000
anthropic/claude-sonnet-5z-ai/glm-5.220200.0200.5111.000
deepseek/deepseek-v4-flashgoogle/gemini-3.6-flash2020-0.0001.0001.000
deepseek/deepseek-v4-flashmoonshotai/kimi-k32020-0.0700.0651.000
deepseek/deepseek-v4-flashopenai/gpt-5.6-sol2020-0.0900.0180.616
deepseek/deepseek-v4-flashopenai/gpt-5.6-terra2020-0.0500.1781.000
deepseek/deepseek-v4-flashqwen/qwen3.7-max20200.0400.2651.000
deepseek/deepseek-v4-flashx-ai/grok-4.52020-0.0400.2471.000
deepseek/deepseek-v4-flashz-ai/glm-5.22020-0.0100.7521.000
google/gemini-3.6-flashmoonshotai/kimi-k32020-0.0700.0421.000
google/gemini-3.6-flashopenai/gpt-5.6-sol2020-0.0900.0040.172
google/gemini-3.6-flashopenai/gpt-5.6-terra2020-0.0500.1071.000
google/gemini-3.6-flashqwen/qwen3.7-max20200.0400.2161.000
google/gemini-3.6-flashx-ai/grok-4.52020-0.0400.1851.000
google/gemini-3.6-flashz-ai/glm-5.22020-0.0100.7071.000
moonshotai/kimi-k3openai/gpt-5.6-sol2020-0.0200.5531.000
moonshotai/kimi-k3openai/gpt-5.6-terra20200.0200.5721.000
moonshotai/kimi-k3qwen/qwen3.7-max20200.1100.0030.123
moonshotai/kimi-k3x-ai/grok-4.520200.0300.3751.000
moonshotai/kimi-k3z-ai/glm-5.220200.0600.0521.000
openai/gpt-5.6-solopenai/gpt-5.6-terra20200.0400.2261.000
openai/gpt-5.6-solqwen/qwen3.7-max20200.1300.0010.026*
openai/gpt-5.6-solx-ai/grok-4.520200.0500.1071.000
openai/gpt-5.6-solz-ai/glm-5.220200.0800.0080.289
openai/gpt-5.6-terraqwen/qwen3.7-max20200.0900.0100.377
openai/gpt-5.6-terrax-ai/grok-4.520200.0100.7551.000
openai/gpt-5.6-terraz-ai/glm-5.220200.0400.1721.000
qwen/qwen3.7-maxx-ai/grok-4.52020-0.0800.0150.554
qwen/qwen3.7-maxz-ai/glm-5.22020-0.0500.0861.000
x-ai/grok-4.5z-ai/glm-5.220200.0300.2641.000
⌐ Stars: *** p<0.001 · ** p<0.01 · * p<0.05 · — not significant. No star is printed if the cell n is below target.
Âşık Creative Turkish · meter and rhyme
Raw score and 95% bootstrap confidence interval
ModelnRaw averageCI lowerCI upperStd. errorzHCB-100
anthropic/claude-opus-5200.2700.2290.3120.0210.560108.4
openai/gpt-5.6-sol200.2500.2260.2750.0130.305104.6
moonshotai/kimi-k3200.2400.2020.2740.0180.178102.7
x-ai/grok-4.5200.2300.2040.2550.0130.051100.8
openai/gpt-5.6-terra200.2300.1970.2640.0170.051100.8
anthropic/claude-sonnet-5200.2200.1890.2500.016-0.07698.9
deepseek/deepseek-v4-flash200.2100.1700.2490.020-0.20396.9
google/gemini-3.6-flash200.2100.1800.2420.016-0.20396.9
z-ai/glm-5.2200.2100.1800.2400.016-0.20396.9
qwen/qwen3.7-max200.1900.1560.2250.018-0.45893.1
Seat-based average — balance check
ModelSeatnRaw average
anthropic/claude-opus-50100.279
anthropic/claude-opus-51100.261
anthropic/claude-sonnet-50100.206
anthropic/claude-sonnet-51100.234
deepseek/deepseek-v4-flash0100.227
deepseek/deepseek-v4-flash1100.193
google/gemini-3.6-flash0100.249
google/gemini-3.6-flash1100.171
moonshotai/kimi-k30100.248
moonshotai/kimi-k31100.232
openai/gpt-5.6-sol0100.247
openai/gpt-5.6-sol1100.253
openai/gpt-5.6-terra0100.234
openai/gpt-5.6-terra1100.226
qwen/qwen3.7-max0100.201
qwen/qwen3.7-max1100.179
x-ai/grok-4.50100.222
x-ai/grok-4.51100.238
z-ai/glm-5.20100.218
z-ai/glm-5.21100.202
Secondary metrics — model average
ModelForm scoreSyllable accuracyRhyme accuracyRhyme gradeFormat violation rateFull-meter line rateOpening minstrelJury score
anthropic/claude-opus-50.9120.9400.8942.0505.5%89.9%0.5000.741
anthropic/claude-sonnet-50.5930.6410.5622.15023.2%62.0%0.5000.457
deepseek/deepseek-v4-flash0.5590.6020.5311.95027.6%57.9%0.5000.404
google/gemini-3.6-flash0.5460.5790.5232.15027.2%56.4%0.5000.419
moonshotai/kimi-k30.7210.7520.7001.80016.8%72.8%0.5000.584
openai/gpt-5.6-sol0.8120.8410.7932.05011.3%81.0%0.5000.668
openai/gpt-5.6-terra0.6740.6930.6612.25019.7%65.0%0.5000.535
qwen/qwen3.7-max0.4460.4740.4271.80034.8%45.7%0.5000.309
x-ai/grok-4.50.6690.6820.6602.15019.2%66.6%0.5000.536
z-ai/glm-5.20.5590.5900.5392.15028.6%56.0%0.5000.416
Pairwise significance matrix — permutation test, Holm–Bonferroni
Model AModel Bn (A)n (B)Differencepp (Holm)Sufficient nSignificance
anthropic/claude-opus-5anthropic/claude-sonnet-520200.0500.0711.000
anthropic/claude-opus-5deepseek/deepseek-v4-flash20200.0600.0571.000
anthropic/claude-opus-5google/gemini-3.6-flash20200.0600.0311.000
anthropic/claude-opus-5moonshotai/kimi-k320200.0300.2971.000
anthropic/claude-opus-5openai/gpt-5.6-sol20200.0200.4141.000
anthropic/claude-opus-5openai/gpt-5.6-terra20200.0400.1601.000
anthropic/claude-opus-5qwen/qwen3.7-max20200.0800.0050.225
anthropic/claude-opus-5x-ai/grok-4.520200.0400.1221.000
anthropic/claude-opus-5z-ai/glm-5.220200.0600.0351.000
anthropic/claude-sonnet-5deepseek/deepseek-v4-flash20200.0100.7051.000
anthropic/claude-sonnet-5google/gemini-3.6-flash20200.0100.6601.000
anthropic/claude-sonnet-5moonshotai/kimi-k32020-0.0200.4321.000
anthropic/claude-sonnet-5openai/gpt-5.6-sol2020-0.0300.1501.000
anthropic/claude-sonnet-5openai/gpt-5.6-terra2020-0.0100.6671.000
anthropic/claude-sonnet-5qwen/qwen3.7-max20200.0300.2271.000
anthropic/claude-sonnet-5x-ai/grok-4.52020-0.0100.6351.000
anthropic/claude-sonnet-5z-ai/glm-5.220200.0100.6551.000
deepseek/deepseek-v4-flashgoogle/gemini-3.6-flash20200.0001.0001.000
deepseek/deepseek-v4-flashmoonshotai/kimi-k32020-0.0300.3061.000
deepseek/deepseek-v4-flashopenai/gpt-5.6-sol2020-0.0400.1101.000
deepseek/deepseek-v4-flashopenai/gpt-5.6-terra2020-0.0200.4651.000
deepseek/deepseek-v4-flashqwen/qwen3.7-max20200.0200.4701.000
deepseek/deepseek-v4-flashx-ai/grok-4.52020-0.0200.4201.000
deepseek/deepseek-v4-flashz-ai/glm-5.220200.0001.0001.000
google/gemini-3.6-flashmoonshotai/kimi-k32020-0.0300.2541.000
google/gemini-3.6-flashopenai/gpt-5.6-sol2020-0.0400.0581.000
google/gemini-3.6-flashopenai/gpt-5.6-terra2020-0.0200.4061.000
google/gemini-3.6-flashqwen/qwen3.7-max20200.0200.4051.000
google/gemini-3.6-flashx-ai/grok-4.52020-0.0200.3411.000
google/gemini-3.6-flashz-ai/glm-5.220200.0001.0001.000
moonshotai/kimi-k3openai/gpt-5.6-sol2020-0.0100.6621.000
moonshotai/kimi-k3openai/gpt-5.6-terra20200.0100.7011.000
moonshotai/kimi-k3qwen/qwen3.7-max20200.0500.0631.000
moonshotai/kimi-k3x-ai/grok-4.520200.0100.6671.000
moonshotai/kimi-k3z-ai/glm-5.220200.0300.2301.000
openai/gpt-5.6-solopenai/gpt-5.6-terra20200.0200.3611.000
openai/gpt-5.6-solqwen/qwen3.7-max20200.0600.0100.440
openai/gpt-5.6-solx-ai/grok-4.520200.0200.2931.000
openai/gpt-5.6-solz-ai/glm-5.220200.0400.0631.000
openai/gpt-5.6-terraqwen/qwen3.7-max20200.0400.1141.000
openai/gpt-5.6-terrax-ai/grok-4.52020-0.0001.0001.000
openai/gpt-5.6-terraz-ai/glm-5.220200.0200.3881.000
qwen/qwen3.7-maxx-ai/grok-4.52020-0.0400.0841.000
qwen/qwen3.7-maxz-ai/glm-5.22020-0.0200.4051.000
x-ai/grok-4.5z-ai/glm-5.220200.0200.3421.000
⌐ Stars: *** p<0.001 · ** p<0.01 · * p<0.05 · — not significant. No star is printed if the cell n is below target.
Behavioral Analytics in-game behavior patterns
Diplomat — promise-keeping and betrayal
ModelPromises madePromise-keeping rateBetrayal rateMessage count
anthropic/claude-opus-54.60082.3%18.9%8
anthropic/claude-sonnet-54.95055.7%44.5%8
deepseek/deepseek-v4-flash4.65039.1%60.5%8
google/gemini-3.6-flash4.35041.6%59.3%8
moonshotai/kimi-k34.50070.6%29.9%8
openai/gpt-5.6-sol4.70077.2%21.9%8
openai/gpt-5.6-terra4.55060.3%40.3%8
qwen/qwen3.7-max4.40029.6%71.0%8
x-ai/grok-4.54.15052.8%47.4%8
z-ai/glm-5.24.95042.4%57.8%9
Tüccar — trading behavior
ModelOffers madeOffer acceptance rateTrade surplusBank dependency
anthropic/claude-opus-510.00058.2%1.7690.061
anthropic/claude-sonnet-510.90035.6%0.6350.177
deepseek/deepseek-v4-flash10.05026.3%0.3010.278
google/gemini-3.6-flash10.55029.4%0.3010.246
moonshotai/kimi-k39.90047.8%1.1600.144
openai/gpt-5.6-sol11.45053.0%1.4260.065
openai/gpt-5.6-terra9.85042.1%0.9320.177
qwen/qwen3.7-max10.05015.5%-0.3760.312
x-ai/grok-4.510.00036.6%0.8220.178
z-ai/glm-5.211.00028.3%0.3450.258
Vezir — vote accuracy and deception
ModelVote accuracyDeception successSage leakAssassination accuracy
anthropic/claude-opus-50.8410.7620.2500.500
anthropic/claude-sonnet-50.5710.5260.2501.000
deepseek/deepseek-v4-flash0.4850.4340.7500.250
google/gemini-3.6-flash0.4850.4520.2500.500
moonshotai/kimi-k30.7090.6540.2500.750
openai/gpt-5.6-sol0.7750.7280.2500.500
openai/gpt-5.6-terra0.6510.5370.5000.750
qwen/qwen3.7-max0.3330.2530.5000.500
x-ai/grok-4.50.6490.5300.0000.750
z-ai/glm-5.20.5000.4860.5000.500
Âşık — form and jury
ModelSyllable accuracyRhyme accuracyFormat violation rateJury score
anthropic/claude-opus-50.9400.8945.5%0.741
anthropic/claude-sonnet-50.6410.56223.2%0.457
deepseek/deepseek-v4-flash0.6020.53127.6%0.404
google/gemini-3.6-flash0.5790.52327.2%0.419
moonshotai/kimi-k30.7520.70016.8%0.584
openai/gpt-5.6-sol0.8410.79311.3%0.668
openai/gpt-5.6-terra0.6930.66119.7%0.535
qwen/qwen3.7-max0.4740.42734.8%0.309
x-ai/grok-4.50.6820.66019.2%0.536
z-ai/glm-5.20.5900.53928.6%0.416
⌐ These metrics are extracted from match transcripts by the neutral summarizer; they don't factor into scoring and are interpretive (spec §6.1).
Hygiene Turkish instruction-following capacity
ModelInvalid actionLanguage violationDecisionPrompt tokensResponse tokensCost (USD)Budgeted tokensInvalid action rateLanguage violation rateAvg. prompt tokensAvg. response tokensAvg. budgeted tokens
anthropic/claude-opus-51641,7435,028,819473,233$36.974,309,9280.9%0.2%2,8852722,473
anthropic/claude-sonnet-557101,7805,285,093478,339$23.034,577,8883.2%0.6%2,9692692,572
deepseek/deepseek-v4-flash95141,8405,341,915479,470$1.814,612,1505.2%0.8%2,9032612,507
google/gemini-3.6-flash59161,7455,073,077473,535$2.094,421,8623.4%0.9%2,9072712,534
moonshotai/kimi-k33151,7895,066,842504,579$4.304,437,7201.7%0.3%2,8322822,481
openai/gpt-5.6-sol2911,8515,195,626486,329$20.284,486,8021.6%0.1%2,8072632,424
openai/gpt-5.6-terra42141,7665,254,284469,176$8.594,518,1802.4%0.8%2,9752662,558
qwen/qwen3.7-max85191,6944,856,229437,032$5.284,207,0215.0%1.1%2,8672582,483
x-ai/grok-4.550111,7214,907,182476,695$14.584,183,0142.9%0.6%2,8512772,431
z-ai/glm-5.261121,8205,361,801481,921$2.924,622,2033.4%0.7%2,9462652,540
⌐ Invalid-action and language-violation rates are direct indicators of instruction-following capacity (spec §4.2, §6). Budgeted prompt tokens are an audit of the equal-cognitive-resource rule: there should be no notable difference between models.
Cost and token usage
ModelRoleRequestsPrompt tokensResponse tokensCost (USD)
anthropic/claude-haiku-4.5jury100162,34635,160$0.0909
anthropic/claude-opus-5player1,7435,028,639473,240$36.97
anthropic/claude-sonnet-5player1,7805,285,181478,156$23.03
deepseek/deepseek-v4-flashplayer1,8405,341,972479,451$1.81
google/gemini-3.6-flashplayer1,7455,073,098473,535$2.09
google/gemini-3.6-flash-litesummarizer2,0855,181,930625,396$0.9345
moonshotai/kimi-k3player1,7895,066,915504,445$4.30
openai/gpt-5.6-minijury100158,09033,439$0.0875
openai/gpt-5.6-solplayer1,8515,195,665486,323$20.28
openai/gpt-5.6-terraplayer1,7665,254,311469,091$8.59
qwen/qwen3.7-maxplayer1,6944,856,121437,019$5.28
x-ai/grok-4.5player1,7214,907,141476,650$14.58
z-ai/glm-5.2player1,8205,361,806482,018$2.92
Method Appendix fairness design and limitations
Run name
HCBfT-Games Demo Run
Master seed
20260801
Neutral summarizer
google/gemini-3.6-flash-lite
Jury models
anthropic/claude-haiku-4.5, openai/gpt-5.6-mini
Temperature
0.7
max_tokens
1024
Context budget
8000 tokens
Summary budget
1200 tokens
Recent window
2 phases
Tokenizer
cl100k_base
Concurrency
8 requests · 4 matches
Backend
openrouter
Total turns
4,236
Diplomat target n
20
Tüccar target n
20
Vezir target n
20
Âşık target n
20
Planned matches
Diplomat 50 · Tüccar 50 · Vezir 40 · Âşık 100
Fairness design. Seat and role rotation were generated with a block-based Latin square; over the course of one block, every model occupies every seat exactly once. Board seeds were replayed with different seat permutations.

Equal cognitive resources. Every model used the same context budget of 8000 tokens, summary budget of 1200 tokens, and recent window of 2 phases; a single shared tokenizer (cl100k_base) was used. Summarization was always done by google/gemini-3.6-flash-lite; no competitor summarized its own history.

Sampling. Temperature 0.7, max_tokens 1024.

Limitations. HCB-100 is not an absolute capability measure; it is a position relative to this run's pool — scores shift if the pool changes. Confidence intervals on scores produced from a small number of matches are wide, and most ranking differences are not meaningful. The Elo ranking is a robustness check only.
Cell counts — model × game
GameModeln
Âşıkanthropic/claude-opus-520
Âşıkanthropic/claude-sonnet-520
Âşıkdeepseek/deepseek-v4-flash20
Âşıkgoogle/gemini-3.6-flash20
Âşıkmoonshotai/kimi-k320
Âşıkopenai/gpt-5.6-sol20
Âşıkopenai/gpt-5.6-terra20
Âşıkqwen/qwen3.7-max20
Âşıkx-ai/grok-4.520
Âşıkz-ai/glm-5.220
Diplomatanthropic/claude-opus-520
Diplomatanthropic/claude-sonnet-520
Diplomatdeepseek/deepseek-v4-flash20
Diplomatgoogle/gemini-3.6-flash20
Diplomatmoonshotai/kimi-k320
Diplomatopenai/gpt-5.6-sol20
Diplomatopenai/gpt-5.6-terra20
Diplomatqwen/qwen3.7-max20
Diplomatx-ai/grok-4.520
Diplomatz-ai/glm-5.220
Tüccaranthropic/claude-opus-520
Tüccaranthropic/claude-sonnet-520
Tüccardeepseek/deepseek-v4-flash20
Tüccargoogle/gemini-3.6-flash20
Tüccarmoonshotai/kimi-k320
Tüccaropenai/gpt-5.6-sol20
Tüccaropenai/gpt-5.6-terra20
Tüccarqwen/qwen3.7-max20
Tüccarx-ai/grok-4.520
Tüccarz-ai/glm-5.220
Veziranthropic/claude-opus-520
Veziranthropic/claude-sonnet-520
Vezirdeepseek/deepseek-v4-flash20
Vezirgoogle/gemini-3.6-flash20
Vezirmoonshotai/kimi-k320
Veziropenai/gpt-5.6-sol20
Veziropenai/gpt-5.6-terra20
Vezirqwen/qwen3.7-max20
Vezirx-ai/grok-4.520
Vezirz-ai/glm-5.220
Configuration dump
{
  "run_name": "HCBfT-Games Demo Run",
  "models": [
    {
      "id": "anthropic/claude-opus-5",
      "temperature": null,
      "max_tokens": null,
      "price_in_per_mtok": null,
      "price_out_per_mtok": null
    },
    {
      "id": "openai/gpt-5.6-sol",
      "temperature": null,
      "max_tokens": null,
      "price_in_per_mtok": null,
      "price_out_per_mtok": null
    },
    {
      "id": "moonshotai/kimi-k3",
      "temperature": null,
      "max_tokens": null,
      "price_in_per_mtok": null,
      "price_out_per_mtok": null
    },
    {
      "id": "openai/gpt-5.6-terra",
      "temperature": null,
      "max_tokens": null,
      "price_in_per_mtok": null,
      "price_out_per_mtok": null
    },
    {
      "id": "x-ai/grok-4.5",
      "temperature": null,
      "max_tokens": null,
      "price_in_per_mtok": null,
      "price_out_per_mtok": null
    },
    {
      "id": "anthropic/claude-sonnet-5",
      "temperature": null,
      "max_tokens": null,
      "price_in_per_mtok": null,
      "price_out_per_mtok": null
    },
    {
      "id": "z-ai/glm-5.2",
      "temperature": null,
      "max_tokens": null,
      "price_in_per_mtok": null,
      "price_out_per_mtok": null
    },
    {
      "id": "google/gemini-3.6-flash",
      "temperature": null,
      "max_tokens": null,
      "price_in_per_mtok": null,
      "price_out_per_mtok": null
    },
    {
      "id": "deepseek/deepseek-v4-flash",
      "temperature": null,
      "max_tokens": null,
      "price_in_per_mtok": null,
      "price_out_per_mtok": null
    },
    {
      "id": "qwen/qwen3.7-max",
      "temperature": null,
      "max_tokens": null,
      "price_in_per_mtok": null,
      "price_out_per_mtok": null
    }
  ],
  "sampling": {
    "temperature": 0.7,
    "max_tokens": 1024,
    "top_p": null
  },
  "utility_model": "google/gemini-3.6-flash-lite",
  "jury_models": [
    "anthropic/claude-haiku-4.5",
    "openai/gpt-5.6-mini"
  ],
  "games": {
    "diplomat": {
      "enabled": true,
      "games_per_model": 20,
      "turn_cap": 16,
      "negotiation_rounds": 2,
      "messages_per_round": 2,
      "win_centers": 7,
      "solo_bonus": 0.25,
      "trade_rounds": null,
      "target_vp": 8,
      "win_bonus": 0.25,
      "discussion_cycles": 2,
      "assassin_bonus": 0.5,
      "bilge_bonus": 0.25,
      "exchanges": 3,
      "hece": 11,
      "form_weight": 0.4,
      "jury_weight": 0.6
    },
    "tuccar": {
      "enabled": true,
      "games_per_model": 20,
      "turn_cap": 60,
      "negotiation_rounds": 2,
      "messages_per_round": 2,
      "win_centers": 7,
      "solo_bonus": 0.25,
      "trade_rounds": 2,
      "target_vp": 8,
      "win_bonus": 0.25,
      "discussion_cycles": 2,
      "assassin_bonus": 0.5,
      "bilge_bonus": 0.25,
      "exchanges": 3,
      "hece": 11,
      "form_weight": 0.4,
      "jury_weight": 0.6
    },
    "vezir": {
      "enabled": true,
      "games_per_model": 20,
      "turn_cap": null,
      "negotiation_rounds": null,
      "messages_per_round": 2,
      "win_centers": 7,
      "solo_bonus": 0.25,
      "trade_rounds": null,
      "target_vp": 8,
      "win_bonus": 0.25,
      "discussion_cycles": 2,
      "assassin_bonus": 0.5,
      "bilge_bonus": 0.25,
      "exchanges": 3,
      "hece": 11,
      "form_weight": 0.4,
      "jury_weight": 0.6
    },
    "asik": {
      "enabled": true,
      "games_per_model": 20,
      "turn_cap": null,
      "negotiation_rounds": null,
      "messages_per_round": 2,
      "win_centers": 7,
      "solo_bonus": 0.25,
      "trade_rounds": null,
      "target_vp": 8,
      "win_bonus": 0.25,
      "discussion_cycles": 2,
      "assassin_bonus": 0.5,
      "bilge_bonus": 0.25,
      "exchanges": 3,
      "hece": 11,
      "form_weight": 0.4,
      "jury_weight": 0.6
    }
  },
  "context": {
    "context_budget_tokens": 8000,
    "digest_budget_tokens": 1200,
    "recent_window_phases": 2,
    "compress_every_phases": 1,
    "message_token_limit": 150,
    "scratchpad_token_limit": 200,
    "tokenizer": "cl100k_base"
  },
  "limits": {
    "concurrency": 8,
    "per_model_concurrency": 4,
    "max_run_cost_usd": 200.0,
    "invalid_action_retries": 2,
    "max_api_attempts": 8,
    "request_timeout_s": 180.0,
    "match_concurrency": 4
  },
  "jury": {
    "enabled": true,
    "identity_denylist": [
      "claude",
      "anthropic",
      "gpt",
      "openai",
      "chatgpt",
      "gemini",
      "google",
      "bard",
      "qwen",
      "alibaba",
      "llama",
      "meta ai",
      "mistral",
      "deepseek",
      "grok",
      "yapay zeka",
      "dil modeli",
      "language model"
    ]
  },
  "seed": 20260801,
  "output_dir": "runs",
  "mock": false,
  "behavioral_analytics": true,
  "language_check": true,
  "language_check_min_chars": 25,
  "composite_weights": {
    "diplomat": 1.0,
    "tuccar": 1.0,
    "vezir": 1.0,
    "asik": 1.0
  }
}