Tren
dar
Dashboard
Papers
News
GitHub
Updates
KO
EN
Sign in
Updates
Papers
Dashboard
News
GitHub
“safety”
Papers, GitHub repos, and news related to this keyword, in one place.
Papers
12
All →
OpenAlex
NLP · LLMs
1.4K citations
A Survey of Large Language Models
Semantic Scholar
NLP · LLMs
288 citations
HealthBench: Evaluating Large Language Models Towards Improved Human Health
Semantic Scholar
NLP · LLMs
183 citations
Medical large language models are vulnerable to data-poisoning attacks
Semantic Scholar
NLP · LLMs
143 citations
Training large language models on narrow tasks can lead to broad misalignment
Semantic Scholar
Other
83 citations
A generative AI-discovered TNIK inhibitor for idiopathic pulmonary fibrosis: a randomized phase 2a trial
OpenAlex
ML Theory · Optimization
64 citations
Recursive Self-Improvement Stability under Endogenous Yardstick Drift
Semantic Scholar
ML Methods
8 citations
Qwen-Scope: Turning Sparse Features into Development Tools for Large Language Models
Semantic Scholar
NLP · LLMs
242 citations
SALAD-Bench: A Hierarchical and Comprehensive Safety Benchmark for Large Language Models
Semantic Scholar
NLP · LLMs
3 citations
Medical Reasoning with Large Language Models: A Survey and MR-Bench
Semantic Scholar
Safety
2 citations
MHSafeEval: Role-Aware Interaction-Level Evaluation of Mental Health Safety in Large Language Models
Semantic Scholar
Agents
122 citations
Agentic Large Language Models, a survey
Semantic Scholar
NLP · LLMs
1 citations
AIPatient Arena: EHR-grounded evaluation of large language models in end-to-end clinical consultation workflows
GitHub
1
All →
Go
★ 18.8K
alibaba/open-code-review
News
12
All →
Hacker News
Policy
▲ 211
Anthropic's Safety Superpower
Hacker News
Safety
▲ 8
Claude Fable 5 jailbroken to bypass Anthropic's new safety guardrails
Hacker News
Safety
▲ 4
LawZero: Safety from Honesty in a Disinterested AI Predictor
techcrunch
Policy
▲ 0
The White House is asking OpenAI to slow roll the release of its new model over safety concerns
techcrunch
Product
▲ 0
Anthropic launches Claude Sonnet 5 as a cheaper way to run agents
mit_tr
Safety
▲ 0
Google DeepMind is worried about what happens when millions of agents start to interact
techcrunch
Policy
▲ 0
Anthropic’s safety warnings may have just backfired — the government has pulled the plug on its most powerful AI
Hacker News
Safety
▲ 5
Anthropic's Fable Jailbreak (Circumvent safety nets)
Hacker News
Safety
▲ 5
LLMs use "safety" specific neuron layers to identify vulnerabilities in code
techcrunch
Safety
▲ 0
xAI fired an engineer who raised alarms about Grok safety, new lawsuit claims
OpenAI
Policy
▲ 0
A blueprint for democratic governance of frontier AI
OpenAI
Safety
▲ 0
Predicting model behavior before release by simulating deployment