Tren
dar
대시보드
논문
뉴스
GitHub
AI 소식
KO
EN
로그인
AI 소식
논문
대시보드
뉴스
GitHub
“transformer”
이 키워드와 관련된 논문 · GitHub · 뉴스를 한곳에 모았습니다.
논문
12
전체 →
Semantic Scholar
자연어·LLM
인용 407
모든 가중치를 1.58비트로 표현하는 초저비용 LLM
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
Semantic Scholar
ML 방법론
인용 1.1K
3300만 개 세포로 학습한 단일세포 기초모델 scGPT
scGPT: toward building a foundation model for single-cell multi-omics using generative AI
Semantic Scholar
자연어·LLM
인용 382
트랜스포머 모델의 행과 열을 삭제하여 LLM을 압축하는 SliceGPT
SliceGPT: Compress Large Language Models by Deleting Rows and Columns
OpenAlex
ML 방법론
인용 329
토큰 임베딩이 어텐션의 기하학적 기반임을 증명하며 No-Q 어텐션 제안
The Embedding Hypothesis: From Fourier Circuits to No-Q Attention
Semantic Scholar
자연어·LLM
인용 7
대규모 언어 모델의 이해와 인지 능력에 대한 균형 잡힌 논의
Understanding Large Language Models
OpenAlex
ML 방법론
인용 178
시간을 나선형 좌표로 재정의해 시계열 예측 성능을 크게 향상
Spiral Time: A Geometric Reframing of Temporal Structure and Its Applications in Machine Learning
Semantic Scholar
자연어·LLM
인용 234
대규모 언어 모델에서 발견된 극단적 활성화 현상
Massive Activations in Large Language Models
Semantic Scholar
자연어·LLM
인용 3
잠에서 영감받은 기억 통합으로 LLM의 맥락 간섭 해결
Learning to Forget: Sleep-Inspired Memory Consolidation for Resolving Proactive Interference in Large Language Models
Semantic Scholar
ML 방법론
인용 2
기존 항생제 펩타이드를 최적화하는 생성 AI 방법
A generative artificial intelligence approach for peptide antibiotic optimization
Semantic Scholar
ML이론·최적화
인용 2
LLM 은닉 상태를 리만 다양체로 해석하는 수학적 프레임워크
Latent Semantic Manifolds in Large Language Models
Semantic Scholar
에이전트
인용 122
종양학 임상 의사결정을 위한 자율 AI 에이전트 개발 및 검증
Development and validation of an autonomous artificial intelligence agent for clinical decision-making in oncology
Semantic Scholar
멀티모달
인용 1
멀티모달 언어모델의 비전 토큰을 의미 기반 진화적 레이블링으로 압축하는 EvoComp
EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling
GitHub
2
전체 →
Python
★ 18.8K
kvcache-ai/ktransformers
생성모델
Python
★ 8.3K
고해상도 이미지 생성을 위한 선형 확산 트랜스포머
NVlabs/Sana
뉴스
6
전체 →
Hacker News
산업·기업
▲ 353
구글 출신 AI 연구자 Noam Shazeer, OpenAI 합류
Noam Shazeer Joins OpenAI
Hacker News
강화학습
▲ 82
단일 트랜스포머 레이어로 전체 파라미터 RL 학습과 동등한 성능 달성
Is One Layer Enough? A Single Transformer Layer Matches Full-Parameter RL Train
Hacker News
AI인프라
▲ 28
FPGA에서 80MHz로 초당 56k 토큰 처리하는 트랜스포머 가속기
GateGPT: 56k tokens per second Transformer (KV cache) on FPGA at 80 MHz
techcrunch
산업·기업
▲ 0
OpenAI, IPO 앞두고 주요 인재 대거 영입하며 내부 역량 강화
OpenAI is bringing on some big guns in the lead-up to its IPO
Hacker News
생성모델
▲ 41
DiffusionBench: 확산 트랜스포머의 종합적인 평가 벤치마크 제안
DiffusionBench: Towards Holistic Evaluation of Generative Diffusion Transformers
Hacker News
기타
▲ 60
900KB 트랜스포머를 과적합시켜 100MB CSV 파일을 7MB로 압축하는 실험
Overfitted a 900KB Transformer to Compress a 100MB CSV into 7MB