Lakshmanan LN

AI/ML Engineer — GenAI & RAG systems, MLOps, ML architecture

I write long-form, first-principles posts on the systems layer underneath modern ML: attention internals, inference engines, and the training dynamics behind why models behave the way they do.

Latest Posts

vLLM Internals: PagedAttention, Continuous Batching, and Where the 2-4x Comes From

vLLM Internals: PagedAttention, Continuous Batching, and Where the 2-4x Comes From

A decode step moves every weight in the model from HBM into the SMs to produce exactly one token per sequence. At batch size 1 that is a …

Read More
Scaling Laws: Kaplan, Chinchilla, and Why Nobody Trains Chinchilla-Optimal

Scaling Laws: Kaplan, Chinchilla, and Why Nobody Trains Chinchilla-Optimal

Say you’ve got loss numbers from a 40M-parameter run and a 400M-parameter run, same data, same architecture family, same optimizer. …

Read More
Attention From Scratch: How Transformers Replaced a Bottleneck with a Lookup

Attention From Scratch: How Transformers Replaced a Bottleneck with a Lookup

Take a vanilla encoder-decoder RNN doing machine translation. It reads a source sentence token by token, updating one hidden-state vector as …

Read More

Find Me Elsewhere

Shorter takes on X, video walkthroughs on YouTube.

Lakshmanan LN

@geekylax on X

Shorter and faster than the blog — half-formed ideas, quick reactions to papers, and whatever I'm currently stuck on.

Follow on X

Latest on YouTube