Tren
dar
Dashboard
Papers
News
GitHub
Updates
KO
EN
Sign in
Updates
Papers
Dashboard
News
GitHub
“alignment”
Papers, GitHub repos, and news related to this keyword, in one place.
Papers
12
All →
OpenAlex
NLP · LLMs
1.5K citations
A Survey of Large Language Models
Semantic Scholar
NLP · LLMs
1.6K citations
ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Semantic Scholar
NLP · LLMs
143 citations
Training large language models on narrow tasks can lead to broad misalignment
Semantic Scholar
NLP · LLMs
105 citations
LLM Post-Training: A Deep Dive into Reasoning Large Language Models
Semantic Scholar
NLP · LLMs
25 citations
General-purpose large language models outperform specialized clinical AI tools on medical benchmarks
Semantic Scholar
Multimodal
8 citations
A Comprehensive Survey and Guide to Multimodal Large Language Models in Vision–Language Tasks
Semantic Scholar
NLP · LLMs
6 citations
Closing the Confidence-Faithfulness Gap in Large Language Models
Semantic Scholar
NLP · LLMs
4 citations
Spurious alignment between large language models and brains can emerge from non-robust methods and overlooked confounds
Semantic Scholar
NLP · LLMs
4 citations
SODA: Semi On-Policy Black-Box Distillation for Large Language Models
Semantic Scholar
Multimodal
4 citations
Do Audio-Visual Large Language Models Really See and Hear?
Semantic Scholar
NLP · LLMs
3 citations
Post-training makes large language models less human-like
Semantic Scholar
NLP · LLMs
2 citations
UniSD: Towards a Unified Self-Distillation Framework for Large Language Models
News
12
All →
mit_tr
Safety
▲ 0
Google DeepMind is worried about what happens when millions of agents start to interact
Hacker News
▲ 40
OpenAI's Misalignment Framework: A Tactical Bid to Preempt Global AI Governance
Anthropic
Safety
▲ 0
Diffuse AI Control on Fuzzy Tasks - Anthropic Alignment Science Blog
Hacker News
▲ 1.1K
A misalignment of AI in mathematics
Hacker News
▲ 60
A Misalignment of AI in Mathematics
Hacker News
▲ 29
A Stupid Idea for AI Alignment We Came with by Looking at Specification Gaming
Hacker News
▲ 8
Anthropic Has Some Alignment Problems
techcrunch
▲ 0
OpenAI caught its models leaving notes to successors to hide bad behavior
OpenAI
▲ 0
Our framework for reporting model misalignment
mit_tr
▲ 0
AI agents blew the whistle on their cheating colleagues
techcrunch
▲ 0
An Anthropic researcher’s doomsday warning comes at a very interesting time
techcrunch
▲ 0
OpenAI adds a prominent AI doomer to its board of directors