<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom"><channel><title>Papers</title><link>https://arxiv.yujunheo.me/</link><description>Recent content on Papers</description><generator>Hugo</generator><language>en</language><atom:link href="https://arxiv.yujunheo.me/index.xml" rel="self" type="application/rss+xml"/><item><title>Reflex: Real-Time Vision-Language-Action Control through Streaming Inference</title><link>https://arxiv.yujunheo.me/2607.14695/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://arxiv.yujunheo.me/2607.14695/</guid><description>&lt;p&gt;Yuanchun Guo (School of Computer Science, Beijing University of Posts and Telecommunications), Bingyan Liu (School of Computer Science, Beijing University of Posts and Telecommunications)&lt;/p&gt;
&lt;p&gt;교신 저자: Bingyan Liu. 원문: &lt;a href="https://arxiv.org/abs/2607.14695"&gt;arXiv:2607.14695&lt;/a&gt;.&lt;/p&gt;
&lt;blockquote class=''&gt;
&lt;p&gt;&lt;strong&gt;번역 및 라이선스 안내.&lt;/strong&gt; 이 페이지는 Yuanchun Guo, Bingyan Liu의 논문 &amp;ldquo;Reflex: Real-Time Vision-Language-Action Control through Streaming Inference&amp;rdquo;(arXiv:2607.14695)를 한국어로 번역하고 역주를 덧붙인 것이다. 원문은 &lt;a href="https://creativecommons.org/licenses/by-nc-sa/4.0/"&gt;CC BY-NC-SA 4.0&lt;/a&gt; 라이선스로 공개되어 있으며, 이 번역도 같은 라이선스(CC BY-NC-SA 4.0)를 따르고 비상업적 목적으로만 공유된다. 번역과 역주는 원저자가 검토하지 않았으며, 정확한 내용은 원문을 기준으로 한다.&lt;/p&gt;</description></item><item><title>Efficient Memory Management for Large Language Model Serving with PagedAttention</title><link>https://arxiv.yujunheo.me/2309.06180/</link><pubDate>Mon, 01 Jan 0001 00:00:00 +0000</pubDate><guid>https://arxiv.yujunheo.me/2309.06180/</guid><description>&lt;p&gt;Woosuk Kwon (UC Berkeley)*, Zhuohan Li (UC Berkeley)*, Siyuan Zhuang (UC Berkeley), Ying Sheng (UC Berkeley, Stanford University), Lianmin Zheng (UC Berkeley), Cody Hao Yu (Independent Researcher), Joseph E. Gonzalez (UC Berkeley), Hao Zhang (UC San Diego), Ion Stoica (UC Berkeley)&lt;/p&gt;
&lt;p&gt;* 공동 기여(Equal contribution). SOSP &amp;lsquo;23, October 23–26, 2023, Koblenz, Germany.&lt;/p&gt;
&lt;h2 id="abstract"&gt;Abstract&lt;a class="anchor" href="#abstract" aria-label="Link"&gt;&lt;/a&gt;&lt;/h2&gt;
&lt;p&gt;대규모 언어 모델(LLM)을 높은 처리량(throughput)으로 서빙하려면 한 번에 충분히 많은 요청을 배치(batch)로 묶어 처리해야 한다. 그러나 기존 시스템은 이를 잘 해내지 못한다. 요청마다 필요한 키-값 캐시(KV cache) 메모리가 매우 크고, 그 크기가 동적으로 커졌다 줄어들기 때문이다. 이 메모리를 비효율적으로 관리하면 단편화(fragmentation)와 중복 저장 때문에 상당량이 낭비되고, 배치 크기가 제한된다. 이 문제를 해결하기 위해 본 논문은 운영체제의 고전적인 가상 메모리와 페이징 기법에서 착안한 어텐션 알고리즘인 PagedAttention을 제안한다. 이를 기반으로 LLM 서빙 시스템 vLLM을 구축했다. vLLM은 (1) KV 캐시 메모리 낭비를 거의 0에 가깝게 줄이고, (2) 요청 내부와 요청 사이에서 KV 캐시를 유연하게 공유해 메모리 사용량을 더 줄인다. 평가 결과, vLLM은 FasterTransformer, Orca 같은 당시 최고 수준의 시스템과 같은 수준의 지연 시간(latency)에서 널리 쓰이는 LLM의 처리량을 2–4배 높였다. 이 향상 폭은 시퀀스가 길수록, 모델이 클수록, 디코딩 알고리즘이 복잡할수록 더 커진다. vLLM의 소스 코드는 &lt;a href="https://github.com/vllm-project/vllm"&gt;https://github.com/vllm-project/vllm&lt;/a&gt; 에 공개되어 있다.&lt;/p&gt;</description></item></channel></rss>