این اسکیل چه میکند؟
تولید ساختاریافته و سرویسدهی سریع LLM با RadixAttention prefix caching.
ORIGINAL DESCRIPTION
Fast structured generation and serving for LLMs with RadixAttention prefix caching.
چه زمانی به کار میآید؟
برای خروجی JSON/regex، constrained decoding، workflow ایجنتی با…
WHEN TO USE
for JSON/regex outputs, constrained decoding, agentic workflows with…
بخشی از متن اصلی با «…» قطع شده است. ادامهٔ آن حدس زده یا تکمیل نشده است.