Your inference runs in Europe. So does your provider.
Open models served from our own infrastructure in Europe, on an OpenAI- and Anthropic-compatible API. GDPR and EU AI Act covered. No US parent company: the CLOUD Act does not reach us.
No credit card · First integration in under 24 hours
from openai import OpenAI client = OpenAI( api_key="nb-...", base_url="https://api.nextbit256.com/v1", # the only change )
Measured in public, integrated where you already work












Production data measured by OpenRouter, not by us.
Source · OpenRouter · Jun 2026Why Nextbit. Not another inference provider.
Sovereignty
Spanish company, hardware in Spain. The CLOUD Act does not reach us.
Compliance
GDPR and the EU AI Act with documents, not with badges.
Proven scale
Partner of OpenRouter, the largest inference AI Gateway in the world.
Our own stack
We build our own stack: router, scheduler and separated compute phases.
Performance
Better than rivals on the same model, measured by a third party.
KV-cache
When your agent repeats context, we serve it from cache.
Prefill (reading your prompt) is the expensive part of compute, not decode. With prefix caching, an agent that iterates 10 times pays for the full prefill once and gets 10 light decodes. We lead cache hit rate on Gemma 4 among OpenRouter providers.
Putting AI in production in Europe has three exits today. All three charge a toll.
Proprietary APIs
The real difference is price. They charge between 70% and 80% more per token than we do, so switching alone cuts a big chunk of your bill. Most of them also process outside the EU and train their models on what you send them.



Hyperscalers
You depend on their GPU availability, and a reasonable price means a long-term rental. Deploying and operating the model yourself takes specialized knowledge, and you remain exposed to the CLOUD Act, since they are still US companies. We absorb that complexity.



US inference clouds
They are fast and good, but their compute runs outside the EU, in US data centers. And even if they opened a region in Europe, they would still be US companies: FISA and the CLOUD Act would reach them all the same, wherever their servers sit.



The model is 20% of the problem. We run the other 80%.
Downloading a model is the easy part. The hard part comes after: managing GPU memory without fragmenting it, not recomputing the same context thirty times, keeping a long prompt from blocking the short requests behind it, and holding the 99th latency percentile when load spikes on a Tuesday at 11:00. That is inference at scale, and it is the only problem we work on.
All of inference. One platform.
All four run on the same layer, the same team and the same data perimeter. Moving between them is one line of configuration.
Serverless inference
OpenAI- and Anthropic-compatible. Change the URL and you are serving from Europe, per million tokens.
$ curl api.nextbit256.com/v1/chat
▸ Aquítienesturespuesta,servidadesdeEuropa.
Enterprise plans also bill per token. That is the problem, not the solution.
A fixed price is just a volume discount: you still pay for every token. And with agents in production, volume is decided by the agent, not by your team.
Same workload, different pricing model
Text extraction · 2,800 tokens per request · 3 to 5 iterations · constant volume
10M tokens/day · 70% input / 30% output
We optimize every layer of inference. And deploy it wherever you say.
The same layer, the same team and the same optimizations run in three places, depending on what your case needs.
On our nodes
Our own infrastructure in Spain. You manage nothing: not hardware, not stack, not scaling.
In your cloud
We deploy and operate our stack inside your AWS, Azure or GCP account. You leverage your spend commitment and your network policies.






Final availability depends on the hyperscaler.
On your premises
Real on-premise: physical hardware in your datacenter, zero data egress. Operated end to end by us.
It is not enough for the server to be in Europe. The company has to be, too.
The US CLOUD Act compels any company incorporated under US law to hand over the data it controls, no matter where the servers physically are. An order addressed to a parent company in Delaware reaches a server in Paris.
We are a Spanish company
No US parent, subsidiary or controlling shareholder. The CLOUD Act and FISA 702 do not apply to us: it follows from where the company is incorporated.
Infrastructure in Spain
Our inference runs on our own hardware in Spain: company, hardware, location and jurisdiction, all European.
If anything runs outside, you will see it marked
Every model carries its origin badge in the catalog: before you integrate it, not after.
What we never do with your data
We never train, tune or evaluate any model on what you send us. On any plan, in any mode.
Compliance is not a badge. It is having the document when they ask for it.
DPA / Data processing agreement
Our processing agreement, with subprocessors, technical measures and a specific CLOUD Act / FISA 702 clause.
Data retention
We never train on your data. Operational data is deleted automatically after 90 days, or after 0 days if your Dedicated Inference agreement says so.
Model cards (EU AI Act Art. 53)
Verified license, training-data summary and systemic-risk assessment for every model in the catalog.
Subprocessors
The subprocessor list is documented in the DPA and available in due diligence: provider, service and country.
Don't take our word for it. OpenRouter measures it.
Two years serving production traffic in the same pool as the biggest providers in the world, with the same requests, measured by the same third party.
| Provider | Input $/M | Output $/M | Cache read $/M | Latency |
|---|---|---|---|---|
| NextbitEU | 1,55 | 3,00 | 0,13 | 1,18 s |
Together | 1,74 | 3,48 | 0,20 | 0,94 s |
Fireworks | 1,74 | 3,48 | 0,145 | 1,57 s |
Novita AI | 1,60 | 3,20 | 0,135 | 1,66 s |
| 1,60 | 3,135 | 0,135 | 1,78 s |
One of our engineers inside your project. Not a support ticket.
Inference at scale is not solved by reading docs. It is solved by looking at your real workload and tuning the stack to it. On dedicated and on-premise, an engineer is assigned to your project from sizing to production.
Technical session
We analyze your real workload and tell you what hardware and model you actually need. If the answer is "less than you thought", we say so.
PoC with success criteria
Latency, cost and quality, measurable and agreed in writing before starting. If they are not met, there is no contract.
Production
Deployment, engine calibration to your load profile, and integration with your observability (Grafana, Datadog, Prometheus).
Ongoing engineering
Every new kernel, quantization format or better-fitting model: we evaluate it and propose it. We do not wait for you to ask.
Response in 48 h · Deployment in 5 working days from signature
Built for regulated buyers
Healthcare
Patient data never leaves the perimeter
See solution →Finance & insurance
DORA: the second auditable European provider
See solution →Public sector
Spanish provider, hardware in Spain
See solution →Agents & voice
Flat cost and low latency from southern Europe
See solution →Media
Visual generation with assets staying in Europe
See solution →Where we are and how to buy us

NVIDIA Inception
Members of NVIDIA's program for companies building on its hardware: early access to architectures and engineering support.


Lanzadera + Angels Capital
Lanzadera, Juan Roig's accelerator, and Angels Capital, his investment firm. Nextbit is part of the Valencia program.

Buy with your AWS commitment
AWS Partner Network members, listed on AWS Marketplace. Buy Nextbit Serverless Inference with your AWS spend commitment or credits.
Available for Serverless Inference. Dedicated Inference is contracted directly with Nextbit on our own infrastructure in Spain.
Your first integration, in under 24 hours.
Change your OpenAI client's base URL, test with your real workload, and decide with your data, not ours.
Response in under 48 hours.


