Why Most RAG Pipelines Break in Production (and How to Build One That Doesn't)

Retrieval-augmented generation is easy to demo and hard to ship. The failure modes of naive RAG — and the architecture patterns that survive real traffic.

Why Most RAG Pipelines Break in Production (and How to Build One That Doesn't)
Written by
Mukesh Anand G
Published on2026-01-18

Read the Full Article

Retrieval-augmented generation is easy to demo and hard to ship. The failure modes of naive RAG — and the architecture patterns that survive real traffic.

Read on Medium

From the blog

View all posts
Breaking the VRAM Wall: How Hierarchical KV Cache Clustering Linearizes Transformer Inference Costs
Applied AI

Breaking the VRAM Wall: How Hierarchical KV Cache Clustering Linearizes Transformer Inference Costs

How clustering the key-value cache collapses the memory economics of long-context Transformers — up to 83% VRAM reduction and 4.3x serving throughput.

Mukesh Anand G
Mukesh Anand G · 2026-02-10
The Memory Problem Everyone Ignores in Long-Context AI
Applied AI

The Memory Problem Everyone Ignores in Long-Context AI

Long context isn't just a context-window number — it's a KV-cache memory bill that grows with every token. The cost nobody budgets for.

Mukesh Anand G
Mukesh Anand G · 2025-12-05