DailyUpdated 2026-09-19 12:21 UTC0 articles
Infra & inference
Serving, inference optimisation, GPUs, routing, cost and latency engineering
Live signals on Infra & inference
- @NVIDIAHPCDev: Introducing CUDA Rust! X · dev radar · 3 days ago · frontier 85%
- @OrcaRouter: PrismML team fit a 27B model into 5.9 GB. X · dev radar · 1 days ago · frontier 78%
- On-Demand Attention: Language Models Know When to Recall arXiv · 2 days ago · frontier 87%
- @AnthropicAI: Biologists use specialized open-source models for tasks like modeling the structure of… X · dev radar · 1 days ago · frontier 69%
- RISC-V and machine learning: a survey arXiv · 2 days ago · frontier 61%
- How GLM built its own inference infrastructure Hacker News · 2 days ago · frontier 63%
- DeepSeek-v4.1 Flash: Pushing the Limits of KV Cache Compression Hacker News · 2 days ago · frontier 78%
- Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL Hugging Face · 9 days ago · frontier 70%
- @thdxr: i had fable doing some inference optimization work, it was making a ton of progress X · vibecoding radar · 1 days ago · frontier 48%
- @typesafeai: Jev is now available on the @vercel AI Gateway https://vercel.com/ai-gateway/models/jev X · ai radar · 2 days ago · frontier 29%