SERVICE 02 · DEDICATED INFERENCE

Critical inference at scale, on your dedicated endpoint.

You get your own endpoint where we deploy one or several AI models in the same environment: the open-source model you choose, a model of your own that you bring us, or a fine-tune of one of our base models served on your dedicated. If you do not know which model fits your case, we help you decide. Resources you share with nobody, ideal for your most critical processes, better use of the context cache, and a fixed monthly fee with unlimited tokens and predictable cost.

Reply within 48 hours.

Private zoneSingle tenantKIMI K3Endpoint in SpainZero data retentionFixed priceDedicated endpoint
The three tiers

A tier for each size, not for each workload.

Starter

Up to 40B parameters at full precision, up to 70B quantised.

  • Logotipo de Qwen3 14BQwen3 14B
  • Logotipo de Gemma 4 27BGemma 4 27B
  • Logotipo de DeepSeek-OCR 2DeepSeek-OCR 2
  • Logotipo de Qwen3-EmbeddingQwen3-Embedding

Best for Real-time voice, RAG, classification and high volume. The lowest cost per million tokens in the whole catalogue.

Pro

Up to 350B parameters at full precision, up to 700B quantised.

  • Logotipo de Qwen3-Coder-480B-A35BQwen3-Coder-480B-A35B
  • Logotipo de DeepSeek-V4-FlashDeepSeek-V4-Flash
  • Minimax

Best for Reasoning, agents, long documents and large MoE models.

Enterprise

No upper limit. Sized to fit your case.

  • Logotipo de Kimi K3Kimi K3
  • GLM-5.3
  • Logotipo de DeepSeek-V4 ProDeepSeek-V4 Pro

Best for Full sovereignty, a dedicated frontier model and on-premise deployment.

How to choose

One tier, the model you choose

You do not pick a tier by its specifications: you pick it by the size of the model your case needs. Inside your dedicated endpoint we deploy the model you decide on and, if you are not sure which, we recommend one during sizing.

RGPD

Sensitive data, with real control.

GDPR Article 9

If you process health data or other special categories under GDPR Article 9, you need to know exactly where it is processed and who can see it. Dedicated Inference, with a DPA, gives you that control: your workload shares infrastructure with nobody. A shared public API cannot.

Guarantees and price

What we sign in the contract. And what it costs per month.

Performance and guarantees

A fixed monthly fee and resources 100% dedicated to your workload, committed by contract. No one else’s spikes reaching you. The details of your case, from the model to the expected volume, are settled during sizing.

Price

A fixed monthly fee, paid up front. The longer the commitment, the lower the fee: terms of 6, 12, 18 and 24 months. Unlimited tokens within your capacity.

FAQ

The objections, answered.

Does my data leave Spain?

No. Your endpoint runs on our own infrastructure in Spain.

Can the US request access?

The CLOUD Act does not reach our infrastructure: we are a Spanish company with no US parent.

Do I share my endpoint with other customers?

No. Your endpoint is dedicated exclusively to your workload. If your security team needs the technical detail of how we guarantee that, we share it during sizing or on a call.

Do I need to know which model I need?

No. If you are not sure which one fits your case best, we recommend one during sizing.

Can I agree zero retention?

Yes: it is set out in the endpoint agreement itself.

How long until it is running?

Quote in 48 hours. Deployment within 5 working days of signing.

Can I bring my own weights?

Yes, including models fine-tuned here or anywhere else.

Shall we talk about your case?

A reply within 48 hours through the contact form, usually sooner.