PromptHub

Speculative Decoding

Quick Definition

Using draft model candidates verified by large model

Full Definition

Using a small draft model to generate candidates verified in parallel by the large target model.

Examples

inference acceleration, LLM serving, throughput optimization

Related Terms

inference-optimization kv-cache

More Large Language Models Terms