Sector: AI infrastructure · Client: In-house project · Status: Delivered

A private AI platform on two NVIDIA DGX Spark systems

The Kubernetes cluster every one of our systems runs on: a resident language model split across two DGX Spark units, image generation on a separate GPU, voice, memory and observability, all under GitOps.

Challenge
Serving AI to a wide range of in-house and client applications without relying solely on an external API, on hardware with unified memory and real thermal limits, without a badly scheduled job taking the live model down.
Solution
A seven-node k3s cluster with a highly available control plane, two NVIDIA DGX Spark systems serving the model with tensor parallelism, a single model router, self-hosted LLM observability and an arbiter that decides which workload may hold each GPU.

View solution: Private AI

Everything we offer in private AI, we run for ourselves first. This platform is the foundation under our reference ecommerce business, our accounting platform, the assistants in our sales proposals and the agent organisation that writes our software.

The challenge

We wanted language, voice and image models to run on hardware we own, free from an outside provider’s pricing, rate limits and terms. The hard part wasn’t buying GPUs; it was running them as a production service. The model had to be available to many applications at once, and any badly scheduled job —a big download, a test run— could take the service down with it.

The hardware brings its own rules. NVIDIA DGX Spark systems use unified memory, so the memory Kubernetes’ scheduler sees isn’t the whole story, and under sustained load the chip needs a thermal cap to keep it from halting.

The solution

We built a seven-node Kubernetes cluster with a highly available control plane. The two DGX Spark units serve a single resident language model with tensor parallelism: the head on one machine, the worker on the other. Image generation runs on a separate GPU on another node, so image and language workloads live side by side without competing.

A single router sits in front of the models. Every application gets its own key, which only lets it call the models assigned to it; if one needs a fallback, it is set in configuration and documented. Every call is traced in our own Langfuse instance, and voice —speech-to-text with Whisper and speech synthesis, cloned voices included— is offered as an internal service.

A control panel publishes the active compute profile at all times, and an arbiter is the only thing allowed to change what holds each GPU. We learned the hard way why that matters: a download job once filled a node’s memory and stopped inference cold. Since then, no workload lands on the Sparks without first reading the arbiter’s state, and the thermal caps are active and monitored.

The outcome

The platform is in production and serves all of our systems. Every change goes in through git and is applied by ArgoCD; access to the dashboards goes through single sign-on with Keycloak. It’s the same architecture we propose to clients who need their data to stay in-house.

Let's talk →Back