Files
nimmersky/guides-stack/llm-stack-architecture.md

3.2 KiB

SkyrimNet LLM Stack — LIVE

All four GPUs live. llama.cpp on dioscuri/comfy-dev. vLLM on theia.
Chat endpoint: /v1/chat/completions with {"chat_template_kwargs": {"enable_thinking": false}}

Instances

Instance Node GPU Port Model Quant Vision
alpha dioscuri Ada #1 31001 Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf-MTP Q4_K_M MTP
beta dioscuri Ada #2 31002 Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf-MTP Q4_K_M MTP
gamma comfy-dev RTX 3090 31003 Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf-MTP Q4_K_M MTP
delta theia Blackwell 6000 31001 Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf-MTP Q4_K_M MTP
epsilon theia Blackwell 6000 31002 Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf-MTP Q4_K_M MTP
zeta theia Blackwell 6000 31003 Huihui-Qwen3-VL-8B-Instruct-abliterated.Q8_0.gguf Q8_0 mmproj-Q8_0.gguf

Service Files under:

/nimmerverse/nimmersky/guides-stack/

llama-alpha.service
llama-beta.service
llama-delta.service
llama-epsilon.service
llama-gamma.service
llama-zeta.service

Theia specific

sudo nvidia-cuda-mps-control -d

This command starts the CUDA Multi-Process Service (MPS) control daemon in the background.

The control program used to manage the MPS server system.-d (Daemon mode): This flag launches the MPS control daemon in the background.What it does: Normally, if multiple Linux processes try to use the same GPU simultaneously, the GPU time-slices between them (hardware context switching), which adds overhead. MPS allows multiple different processes to share a single GPU context. This lets them execute kernels simultaneously on the same GPU, filling up underutilized GPU hardware and improving overall throughput.

sudo nvidia-smi -i 0 -c EXCLUSIVE_PROCESS

This command changes the Compute Mode of a specific GPU so that only one process can attach to it at a time.

    nvidia-smi: The NVIDIA System Management Interface utility.
    -i 0 (Index 0): Targets the specific GPU with an ID of 0.
    -c EXCLUSIVE_PROCESS (Compute Mode): Sets the compute mode to "Exclusive Process"

.What it does: It restricts the GPU so that only one single CUDA context (process) can access it at any given time. If a second process tries to use GPU 0 while the first is running, it will immediately throw an out-of-memory or initialization error.

💡 How They Work TogetherAt first glance, these two commands seem to contradict each other: one allows sharing (MPS), and the other forbids sharing (EXCLUSIVE_PROCESS).However, they are designed to be used together. When you set a GPU to EXCLUSIVE_PROCESS, the MPS server counts as that single allowed process.EXCLUSIVE_PROCESS blocks all regular user applications from touching GPU 0 directly.MPS daemon steps in as the sole owner of GPU 0.Your applications then talk to the MPS daemon, which safely multiplexes their workloads onto the GPU together.This specific combination is the recommended way to deploy MPS because it prevents rogue, non-MPS processes from accidentally hijacking the GPU and ruining performance.