# SkyrimNet LLM Stack — LIVE All four GPUs live. llama.cpp on dioscuri/comfy-dev. vLLM on theia. Chat endpoint: `/v1/chat/completions` with `{"chat_template_kwargs": {"enable_thinking": false}}` ## Instances | Instance | Node | GPU | Port | Model | Quant | Vision | |----------|------|-----|------|-------|-------|------| | **alpha** | dioscuri | Ada #1 | 31001 | Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf-MTP | Q4_K_M MTP | **beta** | dioscuri | Ada #2 | 31002 | Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf-MTP | Q4_K_M MTP | **gamma** | comfy-dev | RTX 3090 | 31003 | Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf-MTP | Q4_K_M MTP | **delta** | theia | Blackwell 6000 | 31001 | Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf-MTP | Q4_K_M MTP | **epsilon** | theia | Blackwell 6000 | 31002 | Gemma4-31B-QAT-Uncensored-HauhauCS-Balanced-Q4_K_M.gguf-MTP | Q4_K_M MTP | **zeta** | theia | Blackwell 6000 | 31003 | Huihui-Qwen3-VL-8B-Instruct-abliterated.Q8_0.gguf | Q8_0 | mmproj-Q8_0.gguf ## Service Files under: /nimmerverse/nimmersky/guides-stack/ llama-alpha.service llama-beta.service llama-delta.service llama-epsilon.service llama-gamma.service llama-zeta.service ## ## Theia specific sudo nvidia-cuda-mps-control -d This command starts the CUDA Multi-Process Service (MPS) control daemon in the background. The control program used to manage the MPS server system.-d (Daemon mode): This flag launches the MPS control daemon in the background.What it does: Normally, if multiple Linux processes try to use the same GPU simultaneously, the GPU time-slices between them (hardware context switching), which adds overhead. MPS allows multiple different processes to share a single GPU context. This lets them execute kernels simultaneously on the same GPU, filling up underutilized GPU hardware and improving overall throughput. ## sudo nvidia-smi -i 0 -c EXCLUSIVE_PROCESS This command changes the Compute Mode of a specific GPU so that only one process can attach to it at a time. nvidia-smi: The NVIDIA System Management Interface utility. -i 0 (Index 0): Targets the specific GPU with an ID of 0. -c EXCLUSIVE_PROCESS (Compute Mode): Sets the compute mode to "Exclusive Process" .What it does: It restricts the GPU so that only one single CUDA context (process) can access it at any given time. If a second process tries to use GPU 0 while the first is running, it will immediately throw an out-of-memory or initialization error. ## 💡 How They Work TogetherAt first glance, these two commands seem to contradict each other: one allows sharing (MPS), and the other forbids sharing (EXCLUSIVE_PROCESS).However, they are designed to be used together. When you set a GPU to EXCLUSIVE_PROCESS, the MPS server counts as that single allowed process.EXCLUSIVE_PROCESS blocks all regular user applications from touching GPU 0 directly.MPS daemon steps in as the sole owner of GPU 0.Your applications then talk to the MPS daemon, which safely multiplexes their workloads onto the GPU together.This specific combination is the recommended way to deploy MPS because it prevents rogue, non-MPS processes from accidentally hijacking the GPU and ruining performance.