Welcome to my blog! I’m Lakshmanan LN, an AI/ML engineer working on GenAI and RAG systems, MLOps, and ML architecture — spending most of my time at the layer underneath the model card, where attention kernels, KV-caches, and scaling curves decide whether something actually ships.
This is where I write the long version of things I had to dig through papers, source code, and a fair amount of trial and error to actually understand. Expect posts that build ideas up from first principles rather than naming an optimization and moving on.
What I Write About
- Model internals — attention, transformers, and the mechanisms underneath the APIs
- Inference systems — how serving engines like vLLM get real-world throughput
- Scaling & training dynamics — what the scaling-law literature actually says
- Production ML — the parts of shipping models that papers don’t cover
Connect With Me
Feel free to reach out on LinkedIn for discussions and collaborations, follow along on X for shorter takes, or check YouTube for video walkthroughs. My latest posts on both are featured below.
(This blog is now published through an automated research-to-deploy pipeline.)