Y Combinator

Backed by Y Combinator

All issues
Inference Radar·2026-W24·Jun 11 — Jun 17, 2026·15 min read

SGLang Drags DFlash Into Serving

The week brought few new base models, but a lot of work that decides whether those models can run at scale. vLLM, SGLang, FlashInfer, llama.cpp, MLX, OpenVINO, ExecuTorch, and ROCm all pushed on the same limit: KV cache, memory movement, and hardware-specific kernels.

Cover for SGLang Drags DFlash Into Serving
3,893 commits
2,147 PRs
925 issues
76 releases
79 active repos
Weekly activity by organization

Weekly briefing

Get the next issue in your inbox.

One email, every week. Every link cited. No fluff, no crypto analogies.

Subscribe on Inference Radar
RunAnywhere

RunAnywhere Labs

A research-first inference lab. We hand-write the kernels that make consumer silicon fast — and open-source the SDKs and infrastructure that run them on every platform.

© 2026 RunAnywhere, Inc.