Online compression framework based on localized processing regions and Progressive Similarity Association (PSA). Up to 83% memory reduction and 4.3x throughput.

A breakthrough research study on memory-efficient Transformer inference, specifically tackling the massive memory footprint of the Key-Value (KV) cache during long-context autoregressive generation. Proposes an online compression framework based on localized processing regions and Progressive Similarity Association (PSA). Experiments report up to 83% KV-cache memory reduction, up to 2x autoregressive inference speed improvement, and up to 4.3x serving-throughput improvement without significant perplexity degradation. Published on ResearchGate with DOI 10.13140/RG.2.2.11098.09927.
