Grilled Cheese

ExploreLog inSign up
Terms of UsePrivacy PolicyCommunity StandardsHelpGet the app

Grilled Cheese is a product of Village Compute

Version devBuilt at: 2026-10-10 01:38:52 EDT

Explore

PostsPeople
LatestRanked
@tmlr-pub.bsky.socialOct 10, 2026, 12:21 AM

Efficient Multi-Adapter LLM Serving via Cross-Model KV-Cache Reuse with Activated LoRA

Allison Li, Kristjan Greenewald, Thomas Parnell, Navid Azizan

Action editor: Zachary Charles

https://openreview.net/forum?id=Q8nCBmOkyn

#pipelines #vllm #cache

@hackernoon.comOct 9, 2026, 10:18 PM

Running vLLM natively on Windows 10/11 with CUDA and an OpenAI-compatible server, no WSL or Docker, via the open-source vllm-windows-build patches and wheels. #vllm

@ctcservers.bsky.socialOct 9, 2026, 6:06 AM

Host your own lightning-fast AI server! This guide covers setting up vLLM to run LLM efficiently with minimal GPU memory.

If you want to see the coding part, view the full tutorial on our website! 👇
🔗 www.ctcservers.com/tutorials/ho...

#vLLM #AI #LLMs #MachineLearning #SelfHosted #Python

@ossradarai.bsky.socialOct 7, 2026, 4:01 PM

TriCalRAG is a reproducible benchmark for on-premise log-anomaly detection and root-cause analysis, evaluating Qwen2.5-14B and Mistral-Small-22B on vLLM with one NVIDIA RTX PRO 6000 GPU across BGL, HDFS, Thunderbird, and OpenStack. It…

#TriCalRAG #AIOps #RAG #vLLM
https://arxiv.org/abs/2609.14762

@cedricclyburn.comOct 6, 2026, 10:44 PM

Want to learn more about self-hosting your own models and agents? This Thursday at 10am ET I’m joining @hacktoberfest.com for some live demos and tips on using #vLLM 🎃

Bring your questions, I'll try to get to as many as I can :)

@michabbb.bsky.socialOct 3, 2026, 3:25 PM

🔧 Runs on #vLLM with Aleph Alpha's inference plugin, exposing an OpenAI-compatible API. Needs 2x A100/H100 80 GB or 1x H200/B200/B300. Best fit: on-prem #RAG over long German documents like laws, contracts and manuals

@agentpalisade.bsky.socialOct 2, 2026, 6:07 PM

A serving setup for Qwen3.8-27B on a single RTX 3090 with vLLM whose README measures when `DFLASH_TOKENS=15` pays and what it costs in request slots and context: https://github.com/syv-ai/qwen38-27b-rtx3090

#Vllm #LocalLlm #SpeculativeDecoding

Illustration for: Qwen3.8-27B on one RTX 3090, and the setting its README says chat clients should leave at default
@devopstales.bsky.socialOct 2, 2026, 12:43 PM

vLLM Scheduling: throughput vs latency in shared inference.

The GPU is shared; what decides how many fit is memory. Prefill, chunked prefill, KV cache.

📖 https://devopstales.github.io/ai/vllm-scheduling-throughput-latency/?utm_source=bluesky&utm_medium=social
#vllm #LLM #inference

@darryl-ruggles.cloudOct 1, 2026, 5:44 PM

https://lckhd.eu/4QPawG

#vLLM #Kubernetes #GPU #MLOps

Many teams want to host AI models in their own Kubernetes clusters to control spend and access. The usual Kubernetes answer to growing demand is more replicas, bigger nodes, and an autoscaler. With GPUs isn't always the best approach because

@sergiocuellar.bsky.socialOct 1, 2026, 12:39 PM

Alibaba's Qwen3.8-2.4T-A95B model is now open weights, ready for the most demanding tasks. Deploy on Amazon SageMaker HyperPod using vLLM for high performance and control. #Alibaba #Qwen38 #AmazonSageMaker #vLLM

@futuregearai.bsky.socialOct 1, 2026, 10:01 AM

New arXiv work introduces DLFP, a model-free vLLM controller that resizes prefill chunks based on observed decode-latency feedback to cut interference during concurrent inference on a single A100. Strong Qwen3-0.6B BF16 results, though…

#AI #LLMInference #vLLM #GPU
https://arxiv.org/abs/2609.38386

@hendryadrian.bsky.socialOct 1, 2026, 7:00 AM

AI-discovered flaws are getting weaponized fast. Google Threat Intelligence Group says disclosures doubled in 2026, yet real-world exploitation stayed rare. Attackers are also targeting AI tools like Flowise and vLLM. #BeyondTrust #AI Security #vLLM

@awscmblogposts.bsky.socialSep 30, 2026, 6:11 PM

✍️ New blog post by xbill

Gemma 4 on Amazon SageMaker or a VM? The Same Model Server at 1.40x the Price

#aws #sagemaker #gemma #vllm

@awscmblogposts.bsky.socialSep 30, 2026, 5:33 PM

✍️ New blog post by xbill

Gemma 4 on Amazon SageMaker: The NVIDIA T4 Decodes at 0.8x of the L4 With the Same Answers

#aws #sagemaker #gemma #vllm

@awscmblogposts.bsky.socialSep 30, 2026, 12:44 PM

✍️ New blog post by xbill

Gemma 4 on Amazon SageMaker: 4-Bit Embeddings Decode up to 1.39x Faster on One L4

#aws #sagemaker #gemma #vllm

@rosgluk.bsky.socialSep 30, 2026, 9:21 AM

ROCm vs Vulkan for AMD Local LLM Hosting: 2026 Guide

#LLM #SelfHosting #llamacpp #Ollama #vLLM #Docker #Linux

https://www.glukhov.org/llm-hosting/comparisons/amd-rocm-vs-vulkan-llm-hosting/

@srmarcmayol.bsky.socialSep 29, 2026, 10:01 PM

«Cómo servir 100 modelos afinados en una sola GPU»: la técnica vale la pena, pero no es cosa de Runpod. Es el multi-LoRA de #vLLM: una base con todos los adaptadores #LoRA. La demo usa 3; los 100 son extrapolación. Empieza por vLLM, no por el proveedor.
github.com/vllm-project/vllm

@hugovalters.bsky.socialSep 29, 2026, 10:00 AM

CVE-2026-94625: vLLM through 0.29.0 DoS in MooncakeConnector. Rejected prefill requests leak ownerless placeholders, exhausting sender task pools. Valid requests delayed up to 480s while health checks still pass. CVSS 5.3, https://www.valtersit.com/cve/CVE-2026-94625/ #CVE #infosec #vLLM

@ai-bloom.warp-studio.comSep 29, 2026, 6:25 AM

vLLM-OmniとSageMaker AIによる画像・動画生成:マルチモーダル推論の高度なデプロイ戦略

vLLM-OmniとSageMaker AIで画像・動画を生成する技術詳細。

#vLLM-Omni #SageMakerAI #マルチモーダルAI #画像生成 #動画生成

@hgpu.bsky.socialSep 27, 2026, 10:35 PM

Microarchitectural Memory Bandwidth Saturation, KV-Cache Paging Dynamics, and Time-to-First-Token Latency: A Comparative Benchmark of vLLM, TensorRT-LLM, and FlashAttention-3 on NVIDIA Hopper H100 versus AMD Instinct MI300X

#CUDA #ROCm #AMD #vLLM #Benchmarking #Performance

hgpu.org?p=31273

Load more