<?xml version="1.0" encoding="utf-8"?><feed xmlns="http://www.w3.org/2005/Atom" xml:lang="en"><generator uri="https://jekyllrb.com/" version="4.4.1">Jekyll</generator><link href="https://kissmetothemoon.github.io/feed.xml" rel="self" type="application/atom+xml"/><link href="https://kissmetothemoon.github.io/" rel="alternate" type="text/html" hreflang="en"/><updated>2026-10-10T09:13:28+00:00</updated><id>https://kissmetothemoon.github.io/feed.xml</id><title type="html">blank</title><subtitle>Personal homepage of Yuannuo Feng, AI infrastructure researcher at Beihang University, working on efficient LLM inference and compute-in-memory architectures. </subtitle><entry><title type="html">ASD: Trading a Bounded Regret Budget for Faster Speculative Decoding</title><link href="https://kissmetothemoon.github.io/blog/2026/asd-bounded-regret/" rel="alternate" type="text/html" title="ASD: Trading a Bounded Regret Budget for Faster Speculative Decoding"/><published>2026-09-02T02:00:00+00:00</published><updated>2026-09-02T02:00:00+00:00</updated><id>https://kissmetothemoon.github.io/blog/2026/asd-bounded-regret</id><content type="html" xml:base="https://kissmetothemoon.github.io/blog/2026/asd-bounded-regret/"><![CDATA[<p>Speculative decoding accelerates LLM inference by letting a small draft model propose tokens that a large target model verifies in parallel. But the standard acceptance rule is stricter than it needs to be — and the waste is measurable. This post summarizes the design of <strong>ASD (Approximate Speculative Decoding)</strong>, a bounded-regret acceptance policy, and what it buys in practice. Code: <a href="https://github.com/Kissmetothemoon/ASD">github.com/Kissmetothemoon/ASD</a> (Apache-2.0).</p> <h2 id="the-inefficiency-of-strict-verification">The inefficiency of strict verification</h2> <p>Under strict greedy verification, the first draft token that disagrees with the target model’s argmax discards the <em>entire</em> remaining draft suffix. In practice, many rejected tokens are nearly as good as the argmax — the target logit gap is tiny — yet they take the whole suffix down with them. The draft model’s work is wasted, and the mean accepted length (hence throughput) suffers.</p> <h2 id="a-bounded-regret-acceptance-policy">A bounded-regret acceptance policy</h2> <p>ASD relaxes the rule: a draft token is accepted when its <strong>regret</strong> — the gap between the target’s top logit and the target’s logit on the draft token — keeps the running total within a <strong>per-request bounded budget B</strong>. The budget is spent as accepted tokens accumulate regret, and clamps at zero.</p> <p>The key property: <strong>at B = 0, ASD exactly recovers strict verification</strong> — not approximately, but token-for-token identical (verified at 0% drift on 1,319 GSM8K requests). The default configuration is therefore lossless; any quality/speed trade-off is an explicit, quantified opt-in.</p> <h2 id="engineering-constraints">Engineering constraints</h2> <p>Three design decisions kept the integration honest:</p> <ul> <li><strong>Zero-intrusion adapter.</strong> The ASD adapter returns the same <code class="language-plaintext highlighter-rouge">(correct_len, bonus, cap_trim_lens)</code> tuple as the native greedy verifier. KV commit, finalization, metrics, and the next draft round are all unaware of the policy switch. With the flag unset, the code path is bit-identical to upstream.</li> <li><strong>No host synchronization on the hot path.</strong> The per-request regret budget lives as a device tensor, bound at the low-frequency prefill seam. The decode hot path never calls <code class="language-plaintext highlighter-rouge">.item()</code> — control flow stays on device.</li> <li><strong>Fail-loud optional dependency.</strong> The research package is lazily imported: missing package plus unset flag means fully native decoding; missing package plus set flag raises at startup, not mid-request. Unsupported runtimes (CUDA graph, overlap scheduling) are rejected at launch as well.</li> </ul> <h2 id="measured-results">Measured results</h2> <p>Setup: Qwen3-14B with a block7 draft model, 8×NVIDIA L20 (46 GB), torch 2.8.0+cu128, <code class="language-plaintext highlighter-rouge">temperature=0</code>, <code class="language-plaintext highlighter-rouge">num_speculative_tokens=7</code>, <code class="language-plaintext highlighter-rouge">max_new_tokens=256</code>, fixed seed.</p> <table> <thead> <tr> <th>Arm</th> <th style="text-align: right">GSM8K accuracy</th> <th style="text-align: right">TPS (tok/s)</th> <th style="text-align: right">Accept rate</th> <th style="text-align: right">Mean accept length</th> </tr> </thead> <tbody> <tr> <td>strict</td> <td style="text-align: right">79.68%</td> <td style="text-align: right">66.21</td> <td style="text-align: right">48.49%</td> <td style="text-align: right">3.40</td> </tr> <tr> <td>ASD, B=0</td> <td style="text-align: right">79.68% (identical)</td> <td style="text-align: right">72.60 (<strong>+9.7%</strong>)</td> <td style="text-align: right">48.49%</td> <td style="text-align: right">3.40</td> </tr> <tr> <td>ASD, q25 (B=2.06)</td> <td style="text-align: right">79.45% (<strong>−0.23 pp</strong>)</td> <td style="text-align: right">75.04 (<strong>+13.3%</strong>)</td> <td style="text-align: right">51.08%</td> <td style="text-align: right">3.58</td> </tr> </tbody> </table> <p>Two observations. First, even B=0 is faster than the stock strict path (+9.7%) — an engineering freebie from the adapter implementation. Second, spending a small regret budget (q25) buys another +3.4 points of throughput and lifts the accept rate by 2.6 pp, at a measured accuracy cost of 0.23 pp on GSM8K — well inside a 1.0 pp acceptance budget. Math500 shows the same direction (12.20% → 12.40%, 90.3 → 94.1 tok/s).</p> <p>Caveats: these are local research-evaluator numbers on fixed-length generation (<code class="language-plaintext highlighter-rouge">ignore_eos=true</code>), not CI results from any serving framework; server-side TTFT/ITL benchmarks remain to be run.</p> <h2 id="whats-next">What’s next</h2> <ul> <li>Server-path benchmarks (TTFT/ITL) once the target runtime supports the required CUDA kernels</li> <li>Budget schedules beyond fixed B (e.g., quantile-adaptive budgets per request phase)</li> <li>Applying the same bounded-regret idea to other parallel-decoding regimes</li> </ul> <p>If you work on speculative decoding or serving systems, I’d be glad to hear your thoughts — <a href="https://github.com/Kissmetothemoon/ASD">issues and PRs welcome</a>.</p>]]></content><author><name></name></author><category term="research"/><category term="llm-inference"/><category term="speculative-decoding"/><summary type="html"><![CDATA[Design notes on a bounded-regret acceptance policy for speculative decoding — +13.3% tokens/s on Qwen3-14B at a cost of 0.23 pp accuracy on GSM8K.]]></summary></entry></feed>