SERVICE 03 · FINE-TUNING

A small model that knows your work beats a big one that knows everything.

We train a small model on your dataset from any base model, deploy it on a Dedicated Inference endpoint that is cheaper to run, and on your task it beats a large, expensive one. Faster to serve, and the weights are yours.

Reply within 48 hours.

GPUFine-tunedmodelFine-tunedmodelFine-tunedmodelZero extra infra cost1-click deployCUSTOMERDEDICATEDFine-tuned modelCUSTOMERDEDICATEDFine-tuned modelCUSTOMERDEDICATEDFine-tuned modelFine-tuning · Custom models served at scale
Specialised intelligence

A generalist knows a little about everything. Yours needs to know a lot about one thing.

Stop explaining your company in every prompt.

Frontier models are trained to answer on quantum physics, Japanese law and pastry. What you need is something that triages your tickets like your best engineer, writes in your industry’s vocabulary and returns exactly the format your system expects.

For that one task, a small model trained on your data can match or beat a much larger one, because it does not spread its capacity across a thousand domains — it concentrates it on yours.

Large and generic

Quality on your task
Good
Prompt needed
Long, context every time
Hardware to serve it
A lot
Cost on Dedicated Inference
High
Your model

Small and specialised

Quality on your task
Equal or better
Prompt needed
Short
Hardware to serve it
A fraction
Cost on Dedicated Inference
Much lower
Fine-tuning is not a cost added on top of inference: it is what makes Dedicated Inference cheap. Going from a 70B to a specialised 14B while keeping the quality on your task is the biggest saving lever in this industry.
Residency and fit

Where we train. And who it makes sense for.

Where it runs

On our infrastructure in Europe. The training dataset is the most sensitive data in the whole process, so it does not leave and it trains nothing but your model.

Who it is for

Narrow, repetitive tasks · Domain vocabulary (clinical, legal, industrial) · Very strict output formats · Moving from a large model to a small one without losing quality

Who it is not for

It makes no sense for adding factual knowledge — that is what RAG is for — nor while you are still exploring the use case. We will tell you before charging you for a training run.

The path

Four steps. And two of them are honesty.

01

We analyse your task

And we tell you whether fine-tuning is the answer or whether your case calls for RAG or a different base model. If it is not, we do not charge you for a training run.

02

We prepare the data

The training set, with you, built from your real examples.

03

We train in Europe

Your data does not leave, and it trains nothing else.

04

We evaluate against your current model

On your task and by your criteria. If it does not win, we say so.

The weights are yours

You take the weights, deploy them wherever you want, and depend neither on our roadmap nor on anyone deprecating a base version.

Your data trains nothing else

It is used only for your model. It does not feed base models, other customers’ models, or anything of ours.

To production without moving

The tuned model is served on the same layer and the same perimeter: serverless, Dedicated Inference or your own data centre.

FAQ

The objections, answered.

Does a tuned 14B really beat a 70B?

On your specific, well-defined task it can match or beat it. As a general claim it would be false. That is why step 4 evaluates against the model you use today, by your criteria.

Who keeps the weights?

You do. That is the difference from tuning a closed model on someone else’s API.

Where does training happen?

On our infrastructure in Europe. The training dataset is usually more sensitive than inference requests, and it is treated as such.

What if fine-tuning is not what I need?

We tell you in the initial analysis, before charging you anything.

Shall we talk about your case?

A reply within 48 hours through the contact form, usually sooner.