Private AI
Language, voice and image models running on hardware we control, so your data doesn't leave for third-party services unless you decide it should.
The problem
Commercial AI services are convenient, but every request travels to someone else’s servers, under their terms, their price changes and their usage caps. For many businesses that simply doesn’t work: client histories, health data, accounts or intellectual property shouldn’t hinge on a third party’s policy. And when everything runs through one external API, a change you don’t control can bring your operation to a halt.
How we work
We run our own inference platform: a Kubernetes cluster with two NVIDIA DGX Spark systems dedicated to the language model, a separate GPU for image generation, and voice, memory and observability services around them. Your applications talk to a single model router; each one has its own key and can only use what it has been granted.
The rule every system we design follows is simple: processing is local by default. If a cloud fallback makes sense in a specific case —so a public-facing assistant is never left without an answer, say— it is agreed in writing, limited to that use and written into the configuration. Systems that handle sensitive data don’t have one.
We apply the same operational discipline the founding team has practised for years on large-scale platforms: GitOps, monitoring, backups, thermal and memory limits under watch, and an arbiter that decides which workload may land on each GPU.
What’s included
- An assessment of your use cases and of which data must stay in-house.
- Model selection and capacity sizing.
- A managed service on our platform, or deployment on your own infrastructure.
- Integration with your applications through an OpenAI-compatible API.
- Monitoring, tracing, model updates and support.
Modules
Inference on our own hardware
Two NVIDIA DGX Spark systems serving a resident language model with tensor parallelism (vLLM, TP=2).
Model router
LiteLLM as the single gateway: every application gets its own key, restricted to the models it may use, with a fallback agreed up front.
LLM observability
Self-hosted Langfuse for traces, sessions, prompt debugging and evaluation, without shipping conversations elsewhere.
Voice
Speech-to-text with Whisper and speech synthesis, including cloned voices with the owner's consent.
Memory and retrieval (RAG)
Vector and graph indexes (Qdrant and FalkorDB) over your documents, your catalogue or your history.
GPU arbiter and control panel
A dashboard that publishes which workload holds each GPU, and a guard that blocks jobs that would knock the live model over.
Single sign-on
Keycloak and oauth2-proxy in front of every service: one identity per person, with groups and permissions.
Related cases
A private AI platform on two NVIDIA DGX Spark systems
The Kubernetes cluster every one of our systems runs on: a resident language model split across two DGX Spark units, image generation on a separate GPU, voice, memory and observability, all under GitOps.
View case →Sistema Galán: an interactive proposal for a bodybuilding coach
A proposal you can walk through, listen to and question: 23 modules to digitise the method of a coach who still runs it all by hand, with an AI assistant and screens for every phase.
View case →
Got a data or process problem?
Write to us: we reply with a concrete plan, no strings attached.
Let's talk →