The Memory Problem Everyone Ignores in Long-Context AI

Long context isn't just a context-window number — it's a KV-cache memory bill that grows with every token. The cost nobody budgets for.

The Memory Problem Everyone Ignores in Long-Context AI
Written by
Mukesh Anand G
Published on2025-12-05

Read the Full Article

Long context isn't just a context-window number — it's a KV-cache memory bill that grows with every token. The cost nobody budgets for.

Read on Medium

From the blog

View all posts
Breaking the VRAM Wall: How Hierarchical KV Cache Clustering Linearizes Transformer Inference Costs
Applied AI

Breaking the VRAM Wall: How Hierarchical KV Cache Clustering Linearizes Transformer Inference Costs

How clustering the key-value cache collapses the memory economics of long-context Transformers — up to 83% VRAM reduction and 4.3x serving throughput.

Mukesh Anand G
Mukesh Anand G · 2026-02-10
Why Most RAG Pipelines Break in Production (and How to Build One That Doesn't)
Applied AI

Why Most RAG Pipelines Break in Production (and How to Build One That Doesn't)

Retrieval-augmented generation is easy to demo and hard to ship. The failure modes of naive RAG — and the architecture patterns that survive real traffic.

Mukesh Anand G
Mukesh Anand G · 2026-01-18