این اسکیل چه میکند؟
تسریع inference مدل LLM با speculative decoding، چند head در Medusa و lookahead decoding.
ORIGINAL DESCRIPTION
Accelerate LLM inference using speculative decoding, Medusa multiple heads, and lookahead decoding techniques.
چه زمانی به کار میآید؟
هنگام بهینهسازی سرعت inference (۱٫۵ تا ۳٫۶ برابر…
WHEN TO USE
when optimizing inference speed (1.5-3.6×…
بخشی از متن اصلی با «…» قطع شده است. ادامهٔ آن حدس زده یا تکمیل نشده است.