[ SPECIFICATION // SECURITY ]: Speculative Decoding speedup simulator and multi-tenant S-LoRA adapter GPU VRAM memory allocation calculator.
Standard: Speculative Execution & Memory Bandwidth Sizer [ VERIFIED // LOCAL EXECUTION ][ ZERO TELEMETRY ][ OFFLINE PWA ][ OPEN SOURCE // MIT ]

[ 1. DECODING & MODEL PARAMETERS ]

75%

Probability that target model accepts a draft token.

5 tokens

Number of candidate tokens speculated per iteration.

5 ms

Time taken by draft model to generate 1 token.

45 ms

Time taken by target model to verify 1 batch in parallel.

[ 2. REAL-TIME PERFORMANCE METRICS ]

Speedup Ratio (S)
2.34x
Expected Tokens / Step
3.16
Effective Latency
19.2 ms
Time Saved
57.3%
[ TOKEN VERIFICATION STREAM ]

[ ARCHITECTURE ] How Speculative Decoding Works

Speculative Decoding (Leviathan et al., Chen et al.) accelerates LLM inference by pairing a small, fast draft model with a large target model. The draft model speculatively generates $K$ tokens. The target model then verifies all $K$ tokens in a single parallel forward pass.

Mathematical Foundation

  • Expected Tokens per Step: \( E[T] = \frac{1 - \alpha^{K+1}}{1 - \alpha} \). Higher acceptance rate \(\alpha\) yields more accepted tokens per verification pass.
  • Step Execution Time: \( T_{step} = K \cdot t_{draft} + t_{target} \). Includes draft generation overhead plus 1 target verification pass.
  • Speedup Formula: \( S = \frac{E[T] \cdot t_{target}}{K \cdot t_{draft} + t_{target}} \). Maximum speedup occurs when \(t_{draft} \ll t_{target}\) and acceptance rate \(\alpha > 70\%\).
[ LOW-LATENCY INFERENCE ]

Speculative Decoding, Medusa & S-LoRA Multiplexing Architecture

Accelerate multi-token prediction using Medusa architecture and learn how to multiplex hundreds of LoRA adapters on a single GPU.

Frequently Asked Questions (FAQ)

What is Speculative Decoding in LLM inference serving?
Speculative Decoding is an inference acceleration method where a fast Draft Model produces candidate tokens that are subsequently validated in parallel by a larger Target Model in a single forward pass.
What is a realistic speedup ratio achieved through Speculative Decoding?
Speedup ratios typically range from 1.8x to 2.8x on a single GPU, depending on the token acceptance rate (alpha) between the draft and target models.
How does S-LoRA multiplexing optimize VRAM in multi-tenant serving?
S-LoRA retains a single unified base model in VRAM and dynamically swaps small LoRA adapters into GPU memory buffers upon incoming requests, saving up to 80% memory.
Does this simulator require a physical GPU to run?
No. This tool models memory bandwidth, arithmetic intensity, FLOPs, and latency mathematically using pure client-side JavaScript.
Copied to clipboard!