Skip to content
bj.choi - notes
Posts
Tags
About
Archives
Search
Tags
KV Cache
1 posts
[LLM 추론] Attention 비용과 KV Cache·GQA·MLA 최적화
Updated:
2026년 8월 25일
Prefill과 Decode의 병목을 구분하고, KV Cache·GQA·MLA가 추론 비용을 줄이는 원리를 서빙 관점에서 정리합니다.