Tren
dar
Dashboard
Papers
News
GitHub
Updates
KO
EN
Sign in
Updates
Papers
Dashboard
News
GitHub
“evaluation”
Papers, GitHub repos, and news related to this keyword, in one place.
Papers
12
All →
Semantic Scholar
NLP · LLMs
927 citations
Toward expert-level medical question answering with large language models
Semantic Scholar
NLP · LLMs
269 citations
Towards conversational diagnostic artificial intelligence
OpenAlex
NLP · LLMs
1.4K citations
A Survey of Large Language Models
Semantic Scholar
NLP · LLMs
1.7K citations
LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Semantic Scholar
NLP · LLMs
1.6K citations
ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Semantic Scholar
NLP · LLMs
554 citations
GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
Semantic Scholar
NLP · LLMs
500 citations
JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models
Semantic Scholar
Multimodal
489 citations
BLINK: Multimodal Large Language Models Can See but Not Perceive
Semantic Scholar
NLP · LLMs
395 citations
Stop Overthinking: A Survey on Efficient Reasoning for Large Language Models
Semantic Scholar
NLP · LLMs
288 citations
HealthBench: Evaluating Large Language Models Towards Improved Human Health
Semantic Scholar
Multimodal
170 citations
EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
Semantic Scholar
Multimodal
159 citations
M3D: Advancing 3D Medical Image Analysis with Multi-Modal Large Language Models
GitHub
2
All →
AI Infrastructure
Python
★ 2K
galilai-group/stable-worldmodel
AI Infrastructure
Python
★ 20.2K
comet-ml/opik
News
12
All →
Hacker News
Agents
▲ 4
A Deterministic Replacement for LLM-as-Judge in Stateful Agent Evaluation
OpenAI
Safety
▲ 0
Predicting model behavior before release by simulating deployment
Hacker News
Generative Models
▲ 41
DiffusionBench: Towards Holistic Evaluation of Generative Diffusion Transformers
OpenAI
Policy
▲ 0
Helping build shared standards for advanced AI
OpenAI
Product
▲ 0
Improving health intelligence in ChatGPT
Hacker News
▲ 52
Third-party cyber evaluations involving OpenAI models
OpenAI
▲ 0
Responding to the next frontier of critical cyber capabilities
OpenAI
▲ 0
Third-party cyber evaluations involving OpenAI models
techcrunch
▲ 0
Design Arena creators raise $7.9 million to bring taste to AI models
Anthropic
▲ 0
Investigating three real-world incidents in our cybersecurity evaluations - Anthropic
venturebeat
▲ 0
The agent evaluation gap: Enterprise AI organizations have a reality-alignment problem, not a coverage problem — and most are shipping to production anyway
OpenAI
▲ 0
Separating signal from noise in coding evaluations