Private AI GPU Cloud
AI in production needs GPU capacity you can plan against, and data that stays where you decide. We do not just supply GPUs: we design the GPU, storage and network architecture around your workload, install it, deploy your models on it, and operate it as one platform. It runs in our own cage space in Zurich, Frankfurt or Istanbul, or in your own data centre. Your data, your prompts and your model weights stay inside the boundary you choose.
It is built for production LLM, ML and image-processing workloads. That includes fintech, trading and life sciences, where data location, confidentiality and operational control matter. The cluster serves your organisation. Your teams share it under agreed quotas.
Tell us what you are training or serving, and on what data.
Boundary
Your data covers more than the training set. All of the following stay inside the boundary you choose, in the locations you choose.
This matters when your organisation has to state where processing happened, not estimate it. With a hosted service, retention, subprocessors and processing location follow that provider's terms. For many workloads those terms are acceptable. This platform is for the ones where they are not.
A fine-tuned model is derived from your data. Treat it with the same care as the source.
At inference time these can carry customer records, transaction details or clinical information.
What the model returns is derived from your data. Where it is stored or logged, that copy needs the same treatment.
An embedding is derived from the document it was computed from. Embeddings look like numbers, so they are easy to overlook.
Depending on how the application is configured, traces and logs can contain prompts and outputs.
Model independence
The platform is built around a class of workload, not around one model. Nothing in the storage layout, the serving layer or the operational tooling assumes a specific model. That gives you four things.
You can benchmark several candidate models on the same hardware before you choose.
You can run more than one at a time behind a routing layer. A larger model for the requests that need it, a smaller one for the rest.
You can replace a model later without rebuilding the platform underneath it.
The capacity is sized for a class of workload, so it stays useful when the model changes.
We do not sell a model. Our recommendation is not tied to our own revenue.
Benchmarking
We publish no performance figures for this platform. A number measured on someone else's model, data and concurrency tells you nothing about yours.
Benchmarking is part of the platform. We measure candidate models on the hardware you would use, with your data, against your evaluation set. The results decide which model you deploy.
What we measure
Quality on your evaluation set.
Behaviour as concurrent load rises.
Memory use at realistic context lengths.
Throughput per unit of hardware.
How the system behaves at its limit.
Cost per unit of work, for each candidate.
A comparison across candidate models on identical hardware. It shows the capacity each one implies, with a recommendation and the reasoning behind it.
Commercial model
In most cases there is no setup fee and no hardware investment. Where suitable capacity exists in our inventory, it is provided from there, and you pay monthly.
Accelerators are expensive and demand is high, so the specification your workload needs may not be in our inventory. Where dedicated hardware is required, that is established during design, before you commit. The investment model is then agreed with you.
The workload assessment comes before any commitment. It can show that a serving platform is enough where a training cluster was assumed. That changes the cost.
If you rent GPU capacity today, we can run a cloud cost analysis on those lines.
Workload fit
GPU demand that runs continuously rather than in short bursts.
Training or fine-tuning on data that has to stay inside your organisation.
Inference on prompts that carry customer, clinical or commercial records.
Several teams that need predictable, allocated capacity.
Production serving where latency and capacity have to be planned.
Retrieval systems built on internal documents, source code or pricing.
Regulated work with a stated processing location.
A large cluster for a few days a quarter, and nothing in between.
Work that is not yet defined, where training and serving volumes are unknown.
A product built on one provider's hosted model. A private platform cannot run it.
A requirement to be on the newest accelerator within weeks of its release.
A problem that is data quality or task definition. More GPUs produce the same results faster.
Next step
Tell us what you are training or serving, on what data, and at what concurrency. Tell us what is in the way today.
You get an architecture conversation with an engineer and a view of what the platform should look like. If renting capacity suits you better, we say so.
The assessment also answers whether you need a cluster at all. We reply within 2 business days.
Delivery scope
What you run, in what proportion, on what data.
GPU, CPU, memory, storage and network, designed together.
Specified against the workload. Where new hardware is needed, we source it.
Racking, power, cooling, cabling, firmware and drivers.
Cluster, GPU scheduling, namespaces, quotas, storage classes and ingress.
Your models brought up, with serving, routing and autoscaling around them.
Measured on your hardware, with your data.
Request distribution designed around how inference behaves.
Utilisation, memory, queues, job outcomes and the infrastructure underneath.
Two designs
Training takes all the capacity it is given, for as long as the job runs. It can wait in a queue.
Inference cannot. Work arrives when users send it. The number that matters is what happens to the slowest requests when the system is busy.
We design for both, and we keep them apart. Serving capacity is reserved. Batch training cannot take it. Where the workloads are large enough, they run in separate node pools.
Multi-tenancy
Each team gets its own namespace, with quotas covering GPUs, CPU, memory and storage. A quota sets what a team may consume, so no team can take the whole cluster by accident.
Batch work queues. Priority decides the order. Fair-share behaviour stops one team's backlog from holding the cluster.
Capacity for customer-facing inference is set aside for it. Batch work runs at lower priority and can be preempted. Where the workloads are large enough, serving and batch run in separate node pools.
Per-team GPU hours, utilisation, queue times and job outcomes are visible to you. That makes idle capacity easy to find and release.
Storage and network
A GPU that waits for data is capacity you pay for and do not use. So storage and network are designed together with the GPUs, not added afterwards.
Fast storage for training reads.
A local cache on each GPU node.
Separate space for checkpoints.
A store for model versions.
GPU-to-GPU traffic and storage traffic use separate network paths, so one does not slow the other when both are busy.
Responsibilities
The hardware, the Kubernetes environment and the model deployment layer. Uptime, performance, security, capacity and network are our job. Monitoring and maintenance do not stop.
Your models, your data and your data science. They stay with you or your partners.
Location and access
Zurich, Frankfurt, Istanbul or your own data centre. You choose before we build, and it is written into the agreement.
Nobody on our team. Running the infrastructure does not require reading your data or your application content, so we do not.
Named people only, with MFA, VPN and logging. Where access must be limited to one jurisdiction, we design for that and put it in the contract.
If you ask us to run an application as well, the access model is different and is set out in that agreement.
Questions
We built our first private cloud infrastructure in 2020. Our team also builds and operates a live crypto exchange platform, open to you on request. It is a different product, built and run by our own engineers.
We do not publish client names. Ask, and we will share relevant references under NDA.
Tell us what you are training or serving. If renting capacity suits you better, we say so. We reply within 2 business days.
Part of Cloud & Infrastructure. See also Managed HCI Private Cloud and DevOps-as-a-Service. More about how we work.