Deep Learning

Scaling Laws: Kaplan, Chinchilla, and Why Nobody Trains Chinchilla-Optimal

Scaling Laws: Kaplan, Chinchilla, and Why Nobody Trains Chinchilla-Optimal

Say you’ve got loss numbers from a 40M-parameter run and a 400M-parameter run, same data, same architecture family, same optimizer. …

Read More
Attention From Scratch: How Transformers Replaced a Bottleneck with a Lookup

Attention From Scratch: How Transformers Replaced a Bottleneck with a Lookup

Take a vanilla encoder-decoder RNN doing machine translation. It reads a source sentence token by token, updating one hidden-state vector as …

Read More