Y Combinator

Backed by Y Combinator

All issues
Inference Radar·2026-W30·Jul 23 — Jul 29, 2026·15 min read

vLLM And SGLang Spark MoE Land Grab

This week made one thing clear: the next inference stack is less a single server and more a set of coordinated engines, caches, routers, kernels, and device-specific paths. vLLM, SGLang, TensorRT-LLM, ROCm, Ollama, MLX, LiteRT, and ExecuTorch all moved in the same direction, faster support for new MoE models and tighter control over where each part of inference runs.

Cover for vLLM And SGLang Spark MoE Land Grab
4,807 commits
3,699 PRs
1,599 issues
84 releases
96 active repos
Weekly activity by organization

Weekly briefing

Get the next issue in your inbox.

One email, every week. Every link cited. No fluff, no crypto analogies.

Subscribe on Inference Radar
RunAnywhere

RunAnywhere Labs

A research-first inference lab. We hand-write the kernels that make consumer silicon fast — and open-source the SDKs and infrastructure that run them on every platform.

© 2026 RunAnywhere, Inc.