Tren
dar
Dashboard
Papers
News
GitHub
Updates
KO
EN
Sign in
Updates
Papers
Dashboard
News
GitHub
“transformer”
Papers, GitHub repos, and news related to this keyword, in one place.
Papers
12
All →
Semantic Scholar
NLP · LLMs
407 citations
The Era of 1-bit LLMs: All Large Language Models are in 1.58 Bits
Semantic Scholar
ML Methods
1.1K citations
scGPT: toward building a foundation model for single-cell multi-omics using generative AI
Semantic Scholar
NLP · LLMs
382 citations
SliceGPT: Compress Large Language Models by Deleting Rows and Columns
OpenAlex
ML Methods
331 citations
The Embedding Hypothesis: From Fourier Circuits to No-Q Attention
Semantic Scholar
NLP · LLMs
7 citations
Understanding Large Language Models
OpenAlex
ML Methods
179 citations
Spiral Time: A Geometric Reframing of Temporal Structure and Its Applications in Machine Learning
Semantic Scholar
NLP · LLMs
234 citations
Massive Activations in Large Language Models
Semantic Scholar
NLP · LLMs
3 citations
Learning to Forget: Sleep-Inspired Memory Consolidation for Resolving Proactive Interference in Large Language Models
Semantic Scholar
ML Methods
2 citations
A generative artificial intelligence approach for peptide antibiotic optimization
Semantic Scholar
ML Theory · Optimization
2 citations
Latent Semantic Manifolds in Large Language Models
Semantic Scholar
Agents
122 citations
Development and validation of an autonomous artificial intelligence agent for clinical decision-making in oncology
Semantic Scholar
Multimodal
1 citations
EvoComp: Learning Visual Token Compression for Multimodal Large Language Models via Semantic-Guided Evolutionary Labeling
GitHub
3
All →
Python
★ 166.3K
huggingface/transformers
Python
★ 18.8K
kvcache-ai/ktransformers
Generative Models
Python
★ 8.3K
NVlabs/Sana
News
11
All →
Hacker News
Industry
▲ 353
Noam Shazeer Joins OpenAI
Hacker News
Reinforcement Learning
▲ 82
Is One Layer Enough? A Single Transformer Layer Matches Full-Parameter RL Train
Hacker News
AI Infrastructure
▲ 28
GateGPT: 56k tokens per second Transformer (KV cache) on FPGA at 80 MHz
techcrunch
Industry
▲ 0
OpenAI is bringing on some big guns in the lead-up to its IPO
Hacker News
Generative Models
▲ 41
DiffusionBench: Towards Holistic Evaluation of Generative Diffusion Transformers
Hacker News
Other
▲ 60
Overfitted a 900KB Transformer to Compress a 100MB CSV into 7MB
Hacker News
▲ 93
A Mathematical Framework for Transformer Circuits (2021)
Hacker News
▲ 379
GPT-6 Astra, looped transformers, and hidden reasoning
Hacker News
▲ 14
Recurrent Looped Transformer
Hacker News
▲ 13
LLM Visualizer – Build a Transformer from Scratch
mit_tr
▲ 0
The Download: the next big thing in LLMs and how AI academic research is shifting