Split or Stay Whole? Disaggregated vs. Monolithic LLM Serving on VMware Cloud Foundation

The particular combination of VMware Cloud Foundation (VCF), VMware vSphere Kubernetes Service (VKS), NVIDIA Run:ai, and NVIDIA Dynamo is not written down anywhere. So we wrote it down, ran a 120-billion-parameter Nemotron on it both ways, and found something we were not looking for. Every enterprise looking seriously at LLM infrastructure eventually reaches the same … Continued
The post Split or Stay Whole? Disaggregated vs. Monolithic LLM Serving on VMware Cloud Foundation appeared first on VMware Blogs.
