Speeding up Generation: From Speculative Decoding to DSpark (Speculative decoder, LLM, Inference) (Jul 04, 2026)