NEW Anthropic-compatible API: connect Claude Code or Cursor

Your inference runs in Europe. So does your provider.

Open models served from our own infrastructure in Europe, on an OpenAI- and Anthropic-compatible API. GDPR and EU AI Act covered. No US parent company: the CLOUD Act does not reach us.

No credit card · First integration in under 24 hours

api.nextbit256.com
from openai import OpenAI

client = OpenAI(
    api_key="nb-...",
    base_url="https://api.nextbit256.com/v1",  # the only change
)
SERVED FROM THE EU·Qwen3.5-35B-A3B·TTFT 210 ms·0.00% tool-call errors

Measured in public, integrated where you already work

OpenRouterAWS MarketplaceAWS Partner NetworkNVIDIA Inception Program1 Million BotNeuroblockOnorato AIRespanConcentrateOpperLanzaderaAngels Capital
45.3%cache hit rate on Gemma 4, the highest in the OpenRouter pool
21.7%of tokens served, among the most-routed providers on OpenRouter in Europe
0.00%tool-call error rate, the only provider at zero
−11%cheaper than Together and Fireworks on the same model

Production data measured by OpenRouter, not by us.

Source · OpenRouter · Jun 2026
Why Nextbit

Why Nextbit. Not another inference provider.

SPAIN · EU

Sovereignty

Spanish company, hardware in Spain. The CLOUD Act does not reach us.

DPA · MODEL CARDS

Compliance

GDPR and the EU AI Act with documents, not with badges.

TOP 5 · DAILY VOLUME · EU

Proven scale

Partner of OpenRouter, the largest inference AI Gateway in the world.

CACHE-AWARE · SLA-AWARE

Our own stack

We build our own stack: router, scheduler and separated compute phases.

MEASURED BY OPENROUTER

Performance

Better than rivals on the same model, measured by a third party.

−95% REDUNDANT COMPUTE

KV-cache

When your agent repeats context, we serve it from cache.

Prefill (reading your prompt) is the expensive part of compute, not decode. With prefix caching, an agent that iterates 10 times pays for the full prefill once and gets 10 light decodes. We lead cache hit rate on Gemma 4 among OpenRouter providers.

The problem

Putting AI in production in Europe has three exits today. All three charge a toll.

Proprietary APIs

The real difference is price. They charge between 70% and 80% more per token than we do, so switching alone cuts a big chunk of your bill. Most of them also process outside the EU and train their models on what you send them.

OpenAIAnthropicGoogle

Hyperscalers

You depend on their GPU availability, and a reasonable price means a long-term rental. Deploying and operating the model yourself takes specialized knowledge, and you remain exposed to the CLOUD Act, since they are still US companies. We absorb that complexity.

AWSAzureGCP

US inference clouds

They are fast and good, but their compute runs outside the EU, in US data centers. And even if they opened a region in Europe, they would still be US companies: FISA and the CLOUD Act would reach them all the same, wherever their servers sit.

Together AIFireworks AIBaseten
No provider offers production-grade inference, below-market pricing, infrastructure in Spain and real physical deployment at once. That gap is where Nextbit exists.
The thesis

The model is 20% of the problem. We run the other 80%.

Downloading a model is the easy part. The hard part comes after: managing GPU memory without fragmenting it, not recomputing the same context thirty times, keeping a long prompt from blocking the short requests behind it, and holding the 99th latency percentile when load spikes on a Tuesday at 11:00. That is inference at scale, and it is the only problem we work on.

The model20 %
Running it80 %
One layer · four services

All of inference. One platform.

All four run on the same layer, the same team and the same data perimeter. Moving between them is one line of configuration.

Serverless inference

OpenAI- and Anthropic-compatible. Change the URL and you are serving from Europe, per million tokens.

Pay per useNo commitment

$ curl api.nextbit256.com/v1/chat

Aquítienesturespuesta,servidadesdeEuropa.

200 OKTTFT 210 msUE
Cost

Enterprise plans also bill per token. That is the problem, not the solution.

A fixed price is just a volume discount: you still pay for every token. And with agents in production, volume is decided by the agent, not by your team.

The difference is not the discount: it is the axis you are billed on. On Dedicated Inference you pay for capacity: loops do not move the bill, next month costs the same as this one, and every optimization we ship is your saving, not our margin.

Work out my case

Same workload, different pricing model

GPT-5.4 miniOpenAI · Per-token API
~4.000 €/mes
Qwen3.5-35B-A3BNextbit · Dedicated Inference
~1.199 €/mes

Text extraction · 2,800 tokens per request · 3 to 5 iterations · constant volume

10M tokens/day · 70% input / 30% output

Claude Opus 4.6Anthropic
~3.036 €/mes
GPT-5.4OpenAI
~1.725 €/mes
Kimi K2.5Nextbit
~75 €/mes
The stack

We optimize every layer of inference. And deploy it wherever you say.

The same layer, the same team and the same optimizations run in three places, depending on what your case needs.

CLIENTSOVEREIGN AI● ESPAÑA · UEYOUR CLOUDON-PREMISENON-EU

On our nodes

Our own infrastructure in Spain. You manage nothing: not hardware, not stack, not scaling.

api.nextbit256.com · Running
Regioneu-south · UE
Uptime 90 d99,9 %
Node local time--:--:--
Serverless · Dedicated

In your cloud

We deploy and operate our stack inside your AWS, Azure or GCP account. You leverage your spend commitment and your network policies.

AWSAzureGCP
Runs inYour account onAWSAzureGCP
RegionThe one you choose
eu-west-1us-east-1me-central-1

Final availability depends on the hyperscaler.

Managed Dedicated Inference

On your premises

Real on-premise: physical hardware in your datacenter, zero data egress. Operated end to end by us.

Sovereign deployment
The same API in all three. Starting in one and moving to another is a deployment decision, not a migration.
Sovereignty

It is not enough for the server to be in Europe. The company has to be, too.

The US CLOUD Act compels any company incorporated under US law to hand over the data it controls, no matter where the servers physically are. An order addressed to a parent company in Delaware reaches a server in Paris.

We are a Spanish company

No US parent, subsidiary or controlling shareholder. The CLOUD Act and FISA 702 do not apply to us: it follows from where the company is incorporated.

Infrastructure in Spain

Our inference runs on our own hardware in Spain: company, hardware, location and jurisdiction, all European.

If anything runs outside, you will see it marked

Every model carries its origin badge in the catalog: before you integrate it, not after.

What we never do with your data

We never train, tune or evaluate any model on what you send us. On any plan, in any mode.

EUSpanish companyNo US parent companyGDPR · EU AI Act
Compliance

Compliance is not a badge. It is having the document when they ask for it.

DPA / Data processing agreement

Our processing agreement, with subprocessors, technical measures and a specific CLOUD Act / FISA 702 clause.

DPA · CLOUD Act clause

Data retention

We never train on your data. Operational data is deleted automatically after 90 days, or after 0 days if your Dedicated Inference agreement says so.

Retention: 90 d → 0 d

Model cards (EU AI Act Art. 53)

Verified license, training-data summary and systemic-risk assessment for every model in the catalog.

Model card · Art. 53

Subprocessors

The subprocessor list is documented in the DPA and available in due diligence: provider, service and country.

Subprocessors in the DPA
Our customers sell to hospitals, banks and public administrations. Our compliance is a piece of theirs, so we document it as if we were the ones being audited.
Performance

Don't take our word for it. OpenRouter measures it.

Two years serving production traffic in the same pool as the biggest providers in the world, with the same requests, measured by the same third party.

DeepSeek V4 Pro on OpenRouter · same model and window for all providers
ProviderInput $/MOutput $/MCache read $/MLatency
NextbitEU1,553,000,131,18 s
Together1,743,480,200,94 s
Fireworks1,743,480,1451,57 s
Novita AI1,603,200,1351,66 s
SiliconFlow1,603,1350,1351,78 s
11% cheaper on input and 14% on output than Together and Fireworks, and the only provider on the list under a Spanish flag.
How we work

One of our engineers inside your project. Not a support ticket.

Inference at scale is not solved by reading docs. It is solved by looking at your real workload and tuning the stack to it. On dedicated and on-premise, an engineer is assigned to your project from sizing to production.

01

Technical session

We analyze your real workload and tell you what hardware and model you actually need. If the answer is "less than you thought", we say so.

02

PoC with success criteria

Latency, cost and quality, measurable and agreed in writing before starting. If they are not met, there is no contract.

03

Production

Deployment, engine calibration to your load profile, and integration with your observability (Grafana, Datadog, Prometheus).

04

Ongoing engineering

Every new kernel, quantization format or better-fitting model: we evaluate it and propose it. We do not wait for you to ask.

Response in 48 h · Deployment in 5 working days from signature

Programs & partners

Where we are and how to buy us

NVIDIA Inception

Members of NVIDIA's program for companies building on its hardware: early access to architectures and engineering support.

Lanzadera + Angels Capital

Lanzadera, Juan Roig's accelerator, and Angels Capital, his investment firm. Nextbit is part of the Valencia program.

Buy with your AWS commitment

AWS Partner Network members, listed on AWS Marketplace. Buy Nextbit Serverless Inference with your AWS spend commitment or credits.

Available for Serverless Inference. Dedicated Inference is contracted directly with Nextbit on our own infrastructure in Spain.

Your first integration, in under 24 hours.

Change your OpenAI client's base URL, test with your real workload, and decide with your data, not ours.

Response in under 48 hours.