vLLM + LiteLLM on AMD Strix Halo
Serving Qwen3.8-27B (INT4 AWQ) and Qwen3-4B (FP16) with vLLM on my two-node k3s cluster, fronted by a LiteLLM OpenAI-compatible router.
Popular topics
Showing posts from
Serving Qwen3.8-27B (INT4 AWQ) and Qwen3-4B (FP16) with vLLM on my two-node k3s cluster, fronted by a LiteLLM OpenAI-compatible router.
A hands-on experiment building a self-managed at-home AI cluster with k3s, Ollama, and LiteLLM.
Testing and configuring opencode for usage with local llms
Strix Halo 128GB RAM, 100% local LLM agents, my tests