Tren
dar
대시보드
논문
뉴스
GitHub
AI 소식
KO
EN
로그인
AI 소식
논문
대시보드
뉴스
GitHub
“alignment”
이 키워드와 관련된 논문 · GitHub · 뉴스를 한곳에 모았습니다.
논문
12
전체 →
OpenAlex
자연어·LLM
인용 1.5K
대규모 언어 모델의 발전과 활용에 대한 종합적 조사
A Survey of Large Language Models
Semantic Scholar
자연어·LLM
인용 1.6K
ChatGLM-4, GPT-4에 필적하는 중국어·영어 LLM 시리즈
ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Semantic Scholar
자연어·LLM
인용 143
좁은 작업 파인튜닝이 광범위한 정렬 실패를 유발한다
Training large language models on narrow tasks can lead to broad misalignment
Semantic Scholar
자연어·LLM
인용 105
LLM 사후 훈련 방법론을 체계적으로 정리한 서베이
LLM Post-Training: A Deep Dive into Reasoning Large Language Models
Semantic Scholar
자연어·LLM
인용 25
범용 LLM, 의료 벤치마크에서 전문 임상 AI 도구를 능가하다
General-purpose large language models outperform specialized clinical AI tools on medical benchmarks
Semantic Scholar
멀티모달
인용 8
비전-언어 작업을 위한 멀티모달 대규모 언어 모델의 종합 서베이 및 가이드
A Comprehensive Survey and Guide to Multimodal Large Language Models in Vision–Language Tasks
Semantic Scholar
자연어·LLM
인용 6
LLM의 자신감과 정확도 간 괴리를 메우는 메커니즘 분석 및 조정 방법
Closing the Confidence-Faithfulness Gap in Large Language Models
Semantic Scholar
자연어·LLM
인용 4
LLM-뇌 정렬 연구의 비강건성 방법과 교란변수 문제를 대규모로 분석
Spurious alignment between large language models and brains can emerge from non-robust methods and overlooked confounds
Semantic Scholar
자연어·LLM
인용 4
반온정책 블랙박스 증류로 LLM 효율적 학습
SODA: Semi On-Policy Black-Box Distillation for Large Language Models
Semantic Scholar
멀티모달
인용 4
시청각 거대 언어 모델의 시각 편향을 메커니즘 분석으로 밝히다
Do Audio-Visual Large Language Models Really See and Hear?
Semantic Scholar
자연어·LLM
인용 3
사후 훈련이 대규모 언어 모델의 인간 행동 모델링 성능을 저하시킨다
Post-training makes large language models less human-like
Semantic Scholar
자연어·LLM
인용 2
자기 증류의 모든 것을 하나로 통합하는 LLM 학습 프레임워크
UniSD: Towards a Unified Self-Distillation Framework for Large Language Models
뉴스
12
전체 →
mit_tr
안전·보안
▲ 0
구글 딥마인드, 수백만 AI 에이전트 상호작용 위험 경고
Google DeepMind is worried about what happens when millions of agents start to interact
Anthropic
안전·보안
▲ 0
안트로픽, AI 모델의 퍼지 태스크 제어를 위한 확산 기반 방법 연구
Diffuse AI Control on Fuzzy Tasks - Anthropic Alignment Science Blog
Hacker News
▲ 1.1K
A misalignment of AI in mathematics
Hacker News
▲ 60
A Misalignment of AI in Mathematics
Hacker News
▲ 29
A Stupid Idea for AI Alignment We Came with by Looking at Specification Gaming
Hacker News
▲ 8
Anthropic Has Some Alignment Problems
OpenAI
▲ 0
Our framework for reporting model misalignment
mit_tr
▲ 0
AI agents blew the whistle on their cheating colleagues
techcrunch
▲ 0
An Anthropic researcher’s doomsday warning comes at a very interesting time
techcrunch
▲ 0
OpenAI adds a prominent AI doomer to its board of directors
Anthropic
▲ 0
An alignment assessment of recent cybersecurity incidents - Anthropic
OpenAI
▲ 0
Paul Christiano joins OpenAI Foundation Board