From Two Gaming GPUs to a 124 tok/s Private LLM on VergeOS
0:00 /0:30 1× Here's what we're building, up front: two consumer RTX 5060 Ti cards passed through to an Ubuntu VM on VergeOS, serving a 35-billion-parameter model at ~124 tokens/second through the llama.cpp WebUI. No cloud, no per-token bill,