Weekly AI & Engineering Digest — Sep 27, 2026
by Vamshi • 9/27/2026Jev's actual mechanics: SGLang scoring, Opik-as-judge, NVIDIA/Stanford's rival CLM. Plus MoE serving internals and this week's Anthropic/OpenAI price war.
Read PostHi, this is Vamshi. I am a full stack developer with deep expertise on Frontend. I have been working on Applied AI, augmented experiences and I want share my learnings through this blog. More details about me can be found on my about page.
Jev's actual mechanics: SGLang scoring, Opik-as-judge, NVIDIA/Stanford's rival CLM. Plus MoE serving internals and this week's Anthropic/OpenAI price war.
Read PostAgent harnesses can 2x inference cost for a 1pt accuracy gain. Plus RAG retrieval internals, LLM memory architecture, LLM-as-judge evals, GPU VRAM math, and GRPO fine-tuning.
Read PostDeepSeek's V4.1-Flash cuts KV cache 4x, OpenAI's 10K agents cracked Navier-Stokes, Anthropic found 151M Claude distillation exchanges. Plus 6 deep-dives on routing and serving.
Read PostGPT-6 Astra's benchmarks and safety classification, attention-mechanism internals, embedding compression, RAG retrieval failures, how Shopify beat GPT-5.6 with a 0.8B model.
Read PostvLLM vs SGLang vs Ollama, GraphRAG map-reduce search, a one-line Qdrant fix for multivector memory, and why agent skills work as runbooks, not facts (8,100 trials).
Read PostA cheaper model can double your turn cost, vLLM's 23x batching trick, Google's split TPU 8t/8i chips, and the 3-layer stack teams are shipping to sandbox AI agents.
Read PostKV cache economics, the LLM security threat map, 5 models on one GPU, and a 7B model beating a 32B via distillation — six deep dives from Daily Dose and ByteByteGo this week.
Read PostA scheduled agent reads 18 AI newsletters from Gmail every Monday and publishes a digest here. The five failure modes that quietly cost me content before I caught them.
Read PostLLMs are stateless by default, and that is fatal for autonomous agents. A look at tiered memory, declarative intent, and the evaluation metrics that actually predict agent survival.
Read Post