vLLM + LiteLLM on AMD Strix Halo
Serving Qwen3.8-27B (INT4 AWQ) and Qwen3-4B (FP16) with vLLM on my two-node k3s cluster, fronted by a LiteLLM OpenAI-compatible router.
Popular topics
Showing posts from
Serving Qwen3.8-27B (INT4 AWQ) and Qwen3-4B (FP16) with vLLM on my two-node k3s cluster, fronted by a LiteLLM OpenAI-compatible router.
Testing and configuring opencode for usage with local llms
Strix Halo 128GB RAM, 100% local LLM agents, my tests