VMware Blogsblog

Split or Stay Whole? Disaggregated vs. Monolithic LLM Serving on VMware Cloud Foundation

Friday, July 31, 2026vmwareblogsView original

The particular combination of VMware Cloud Foundation (VCF), VMware vSphere Kubernetes Service (VKS), NVIDIA Run:ai, and NVIDIA Dynamo is not written down anywhere. So we wrote it down, ran a 120-billion-parameter Nemotron on it both ways, and found something we were not looking for. Every enterprise looking seriously at LLM infrastructure eventually reaches the same … Continued

The post Split or Stay Whole? Disaggregated vs. Monolithic LLM Serving on VMware Cloud Foundation appeared first on VMware Blogs.