FEATURE & DOCS
A dedicated space for documenting technical blueprints, architectural patterns, and engineering reflections. This archive serves as a living knowledge base where innovation meets practical execution.
Written In Our Scars — Debut Hardback Novel
A chance encounter in Bengaluru develops into a relationship shaped by attraction, trust, jealousy, intimacy, and unresolved fears as Chennai and the sea call.
Breaking the VRAM Wall: How Hierarchical KV Cache Clustering Linearizes Transformer Inference Costs
How clustering the key-value cache collapses the memory economics of long-context Transformers — up to 83% VRAM reduction and 4.3x serving throughput.
Why Most RAG Pipelines Break in Production (and How to Build One That Doesn't)
Retrieval-augmented generation is easy to demo and hard to ship. The failure modes of naive RAG — and the architecture patterns that survive real traffic.
The Memory Problem Everyone Ignores in Long-Context AI
Long context isn't just a context-window number — it's a KV-cache memory bill that grows with every token. The cost nobody budgets for.
Building a Tamil Character-Level Tokenizer
Design decisions behind Asai (அசை) — morphology, agglutinative grammar, and why BPE/WordPiece vocabularies underserve Dravidian languages.