Tren
dar
Dashboard
Papers
News
GitHub
Updates
KO
EN
Sign in
Updates
Papers
Dashboard
News
GitHub
“benchmark”
Papers, GitHub repos, and news related to this keyword, in one place.
Papers
12
All →
Semantic Scholar
ML Methods
65 citations
Deep-learning-based single-domain and multidomain protein structure prediction with D-I-TASSER
OpenAlex
NLP · LLMs
1.5K citations
A Survey of Large Language Models
Semantic Scholar
NLP · LLMs
1.7K citations
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Semantic Scholar
Multimodal
568 citations
Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Semantic Scholar
NLP · LLMs
554 citations
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
Semantic Scholar
NLP · LLMs
501 citations
LLM2Vec: Large Language Models Are Secretly Powerful Text Encoders
Semantic Scholar
NLP · LLMs
500 citations
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Semantic Scholar
Multimodal
489 citations
BLINK: Multimodal Large Language Models Can See but Not Perceive
Semantic Scholar
NLP · LLMs
395 citations
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
OpenAlex
ML Methods
338 citations
Axiom: A Householder-Parameterized Pure Unitary RNN for Long-Range Sequence Modeling
Semantic Scholar
NLP · LLMs
318 citations
Better & Faster Large Language Models via Multi-token Prediction
Semantic Scholar
NLP · LLMs
288 citations
HealthBench: Evaluating Large Language Models Towards Improved Human Health
GitHub
9
All →
Python
★ 30.5K
tirth8205/code-review-graph
Python
★ 3.1K
NVIDIA-NeMo/Switchyard
Python
★ 1.4K
uber/ADR
Python
★ 1.2K
harveyai/harvey-labs
Agents
Python
★ 1.8K
hexo-ai/sia
Agents
News
12
All →
Hacker News
Safety
▲ 631
GLM 5.2 beats Claude in our benchmarks
Hacker News
Product
▲ 5
GLM-5.2 is above GPT-5.5 in new agentic knowledge work eval
Hacker News
Agents
▲ 5
TypeScript
★ 23.2K
rohitg00/agentmemory
AI Infrastructure
Python
★ 5.6K
Andyyyy64/whichllm
Agents
Python
★ 55.5K
MemPalace/mempalace
Jupyter Notebook
★ 4.1K
FareedKhan-dev/all-agentic-architectures
FlowerBench: Benchmarking AI Agents on Real Enterprise Work
Hacker News
Industry
▲ 19
Why Weibo's tiny VibeThinker-3B has the AI world arguing over benchmarks again
Hacker News
Industry
▲ 5
LLM benchmarks are answering someone else's question
Hacker News
Product
▲ 5
Show HN: SOCBench – an open benchmark for AI on SoC tasks
OpenAI
Product
▲ 0
Introducing LifeSciBench
OpenAI
Product
▲ 0
Introducing GeneBench-Pro
Hacker News
Safety
▲ 7
Ask HN: Are there good security benchmarks for LLMs?
Hacker News
▲ 6
Insurance Agent Benchmark: 166 real-world cases for evaluating insurance AI
Hacker News
▲ 178
Real-SWE: Benchmarking AI models on private, real-world, enterprise codebases
Hacker News
▲ 37
Benchmark: CadQuery vs. OpenSCAD for agentic CAD work