Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
Allison Li, Kristjan Greenewald, Thomas Parnell, Navid Azizan
Action editor: Zachary Charles

Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA
Allison Li, Kristjan Greenewald, Thomas Parnell, Navid Azizan
Action editor: Zachary Charles
Running vLLM natively on Windows 10/11 with CUDA and an OpenAI-compatible server, no WSL or Docker, via the open-source vllm-windows-build patches and wheels. #vllm
Host your own lightning-fast AI server! This guide covers setting up vLLM to run LLM efficiently with minimal GPU memory.
If you want to see the coding part, view the full tutorial on our website! 👇
🔗 www.ctcservers.com/tutorials/ho...
TriCalRAG is a reproducible benchmark for on-premise log-anomaly detection and root-cause analysis, evaluating Qwen2.5-14B and Mistral-Small-22B on vLLM with one NVIDIA RTX PRO 6000 GPU across BGL, HDFS, Thunderbird, and OpenStack. It…
#TriCalRAG #AIOps #RAG #vLLM
https://arxiv.org/abs/2609.14762
Want to learn more about self-hosting your own models and agents? This Thursday at 10am ET I’m joining @hacktoberfest.com for some live demos and tips on using #vLLM 🎃
Bring your questions, I'll try to get to as many as I can :)
A serving setup for Qwen3.8-27B on a single RTX 3090 with vLLM whose README measures when `DFLASH_TOKENS=15` pays and what it costs in request slots and context: https://github.com/syv-ai/qwen38-27b-rtx3090
#Vllm #LocalLlm #SpeculativeDecoding
vLLM Scheduling: throughput vs latency in shared inference.
The GPU is shared; what decides how many fit is memory. Prefill, chunked prefill, KV cache.
📖 https://devopstales.github.io/ai/vllm-scheduling-throughput-latency/?utm_source=bluesky&utm_medium=social
#vllm #LLM #inference
#vLLM #Kubernetes #GPU #MLOps
Many teams want to host AI models in their own Kubernetes clusters to control spend and access. The usual Kubernetes answer to growing demand is more replicas, bigger nodes, and an autoscaler. With GPUs isn't always the best approach because
Alibaba's Qwen3.8-2.4T-A95B model is now open weights, ready for the most demanding tasks. Deploy on Amazon SageMaker HyperPod using vLLM for high performance and control. #Alibaba #Qwen38 #AmazonSageMaker #vLLM
New arXiv work introduces DLFP, a model-free vLLM controller that resizes prefill chunks based on observed decode-latency feedback to cut interference during concurrent inference on a single A100. Strong Qwen3-0.6B BF16 results, though…
#AI #LLMInference #vLLM #GPU
https://arxiv.org/abs/2609.38386
AI-discovered flaws are getting weaponized fast. Google Threat Intelligence Group says disclosures doubled in 2026, yet real-world exploitation stayed rare. Attackers are also targeting AI tools like Flowise and vLLM. #BeyondTrust #AI Security #vLLM
✍️ New blog post by xbill
Gemma 4 on Amazon SageMaker or a VM? The Same Model Server at 1.40x the Price
✍️ New blog post by xbill
Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers
✍️ New blog post by xbill
Gemma 4 on Amazon SageMaker: 4-Bit Embeddings Decode up to 1.39x Faster on One L4
ROCm vs Vulkan for AMD Local LLM Hosting: 2026 Guide
#LLM #SelfHosting #llamacpp #Ollama #vLLM #Docker #Linux
https://www.glukhov.org/llm-hosting/comparisons/amd-rocm-vs-vulkan-llm-hosting/
«Cómo servir 100 modelos afinados en una sola GPU»: la técnica vale la pena, pero no es cosa de Runpod. Es el multi-LoRA de #vLLM: una base con todos los adaptadores #LoRA. La demo usa 3; los 100 son extrapolación. Empieza por vLLM, no por el proveedor.
github.com/vllm-project/vllm
CVE-2026-94625: vLLM through 0.29.0 DoS in MooncakeConnector. Rejected prefill requests leak ownerless placeholders, exhausting sender task pools. Valid requests delayed up to 480s while health checks still pass. CVSS 5.3, https://www.valtersit.com/cve/CVE-2026-94625/ #CVE #infosec #vLLM
vLLM-OmniとSageMaker AIによる画像・動画生成:マルチモーダル推論の高度なデプロイ戦略
vLLM-OmniとSageMaker AIで画像・動画を生成する技術詳細。
Microarchitectural Memory Bandwidth Saturation, KV-Cache Paging Dynamics, and Time-to-First-Token Latency: A Comparative Benchmark of vLLM, TensorRT-LLM, and FlashAttention-3 on NVIDIA Hopper H100 versus AMD Instinct MI300X