Everything we offer in private AI, we run for ourselves first. This platform is the foundation under our reference ecommerce business, our accounting platform, the assistants in our sales proposals and the agent organisation that writes our software.
The challenge
We wanted language, voice and image models to run on hardware we own, free from an outside provider’s pricing, rate limits and terms. The hard part wasn’t buying GPUs; it was running them as a production service. The model had to be available to many applications at once, and any badly scheduled job —a big download, a test run— could take the service down with it.
The hardware brings its own rules. NVIDIA DGX Spark systems use unified memory, so the memory Kubernetes’ scheduler sees isn’t the whole story, and under sustained load the chip needs a thermal cap to keep it from halting.
The solution
We built a seven-node Kubernetes cluster with a highly available control plane. The two DGX Spark units serve a single resident language model with tensor parallelism: the head on one machine, the worker on the other. Image generation runs on a separate GPU on another node, so image and language workloads live side by side without competing.
A single router sits in front of the models. Every application gets its own key, which only lets it call the models assigned to it; if one needs a fallback, it is set in configuration and documented. Every call is traced in our own Langfuse instance, and voice —speech-to-text with Whisper and speech synthesis, cloned voices included— is offered as an internal service.
A control panel publishes the active compute profile at all times, and an arbiter is the only thing allowed to change what holds each GPU. We learned the hard way why that matters: a download job once filled a node’s memory and stopped inference cold. Since then, no workload lands on the Sparks without first reading the arbiter’s state, and the thermal caps are active and monitored.
The outcome
The platform is in production and serves all of our systems. Every change goes in through git and is applied by ArgoCD; access to the dashboards goes through single sign-on with Keycloak. It’s the same architecture we propose to clients who need their data to stay in-house.