VCF Private AI Services (PAIS) ships with 3 x Model Engines: vLLM, LlamaCPP & Infinity
Just learned the default version of LlamaCPP is only for CPU-based models, doesn't include GPU accelerator. With that said, users can easily override engine w/server-cuda ๐
