OpenRouter's Dev Rel Lead dropped 11 LLMs into a battle royale game for 30 rounds; Grok 4.1 Fast won 43% of matches. Claude Sonnet 4.6 showed cooperative tendencies, suggesting it may be more useful in real applications despite lower win rate.
OpenRouter's Dev Rel Lead Jacky ran an experiment dropping 11 LLMs into a 2D battle royale game for 30 rounds. Grok 4.1 Fast won 13 games (43%), followed by Claude Sonnet 4.6 with 5 wins. GPT 5.4 had the most kills (38) but only 2 wins. GPT 5.4-mini, DeepSeek 4 Flash, and Kimi K2.6 never won a single game.
Traditional benchmarks failed to predict model performance in this game environment. Grok 4.1 Fast was the most cost-effective at $0.97 per win, while Claude Sonnet 4.6 cost $26.78 per win—a 27x difference. Claude Sonnet 4.6 exhibited social behaviors like suggesting team-ups and revealing its location, but this hurt its winning chances.
This experiment highlights that model selection should consider real-task behavior and cost efficiency beyond standard benchmarks. While Grok excelled in win-oriented games, Claude's cooperative tendencies may be more suitable for applications requiring safety and collaboration.
Comments focus on suspicions that the article was LLM-generated (Claude/Grok), criticizing its artificial style and structure. Some argue the quality and coherence are poor regardless of authorship, while others find the LLM-detection meta-discussion tiresome and off-topic. Overall, skepticism about the article's authenticity and readability dominates.