Reading
Books and papers that shape how I think.
Books
4 entries
- in progress
Build a Large Language Model (from Scratch)
Sebastian Raschka
80% - todo
Chip Huyen
- completed
Designing Machine Learning Systems
Chip Huyen
- completed
Papers
5 entries
- completed
Role of Bias Terms in Dot-Product Attention
Mahdi Namazifar, Devamanyu Hazarika, Dilek Hakkani-Tur
- completed
Ashish Vaswani et al.
- completed
An Image Is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy et al.
- completed
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Joshua Ainslie et al.
Library
35 entries
Papers, articles, and references bookmarked outside the core reading log.
Unknown author
webpageAntonio Lupetti
blogPost2024- webpage
Understanding the 4 Main Approaches to LLM Evaluation (From Scratch)
Sebastian Raschka PhD
webpage2025A Gentle Introduction to Graph Neural Networks
Benjamin Sanchez-Lengeling, Emily Reif, Adam Pearce +1 more
journalArticle2021Maarten Grootendorst
webpage2024Maarten Grootendorst
webpage2024Triton: an intermediate language and compiler for tiled neural network computations
Philippe Tillet, H. T. Kung, David Cox
conferencePaper2019Why do CPUs have multiple cache levels?
Unknown author
blogPost2016Quantization from the ground up
Sam Rose
webpage2026Tim Rocktaschel
webpage- webpage
Heterogeneous Graph Transformer
Ziniu Hu, Yuxiao Dong, Kuansan Wang +1 more
preprint2020Robust LLM Training Infrastructure at ByteDance
Borui Wan, Gaohong Liu, Zuquan Song +32 more
preprint2025DeepSeek's open-source week and why it's a big deal
PySpur-AI Agent Builder
webpage2025Reducing Activation Recomputation in Large Transformer Models
Vijay Korthikanti, Jared Casper, Sangkug Lym +4 more
preprint2022- webpage
Role of Bias Terms in Dot-Product Attention
Mahdi Namazifar, Devamanyu Hazarika, Dilek Hakkani-Tur
preprint2023GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Joshua Ainslie, James Lee-Thorp, Michiel de Jong +3 more
preprint2023Gemma Team, Aishwarya Kamath, Johan Ferret +213 more
preprint2025Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
Patrick Esser, Sumith Kulal, Andreas Blattmann +14 more
preprint2024Denoising Diffusion Probabilistic Models
Jonathan Ho, Ajay Jain, Pieter Abbeel
preprint2020V-VAE: A Variational Auto Encoding Framework Towards Fine-Grained Control over Human-Like Chat
Qi Lin, Weikai Xu, Lisi Chen +1 more
preprint2025Generative Adversarial Networks
Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza +5 more
preprint2014Auto-Encoding Variational Bayes
Diederik P. Kingma, Max Welling
preprint2022LoRA: Low-Rank Adaptation of Large Language Models
Edward J. Hu, Yelong Shen, Phillip Wallis +5 more
preprint2021Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
Patrick Lewis, Ethan Perez, Aleksandra Piktus +9 more
preprint2021Unknown author
webpage- webpage
- preprint2020
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov +9 more
preprint2021Papers with Code - Vision Transformer Explained
Unknown author
webpageThe State of LLM Reasoning Models
Sebastian Raschka PhD
webpage2025In Search of an Understandable Consensus Algorithm
Diego Ongaro, John Ousterhout
journalArticleYu A. Malkov, D. A. Yashunin
preprint2018