ML Env

Use casesOperations & logistics

Draft replies to routine supplier emails, in the supplier's language

Turns a one-line instruction ('chase PO 4521, firm but polite') into a ready-to-review reply in the supplier's language, drawn from the thread.

How it works
For everyday correspondence — chasing a late delivery, asking for a proforma, confirming quantities, requesting documents — you write a short instruction and the large model drafts the reply from the existing thread, so names, references and context carry over. It writes in the supplier's language: German, Italian, Polish, or plain international English. The draft sits in your outbox, not theirs — nothing is sent until you have read it.
Data you need
The email thread being replied to, and ideally a few examples of how your company writes so the tone matches. Order details (PO number, dates) come from the person writing the instruction or a pasted ERP snippet. All of this exists in any company that emails suppliers — no special dataset required.
What to expect
Solid for routine chasing and requests in major European languages. Fluent output can mask a factual slip — a wrong date or quantity reads perfectly, so check the numbers against the PO, not against the prose. In lower-resource languages the register can drift stiff or off-key, so judge it with a native reader first. And the model has no idea what your stock or cashflow actually allows — never let a draft commit you to anything.
Where people stay involved
Every draft is read and sent by a person. Anything touching price negotiation, contract terms or disputes deserves a careful human rewrite, not a skim. If nobody on your team reads the target language, treat drafts in it as unverified and keep those exchanges in a language you can check.

Which model, and what it costs to run

Qwen3-32B

This job answers while someone waits, so latency comes first: we hold it to a single small card and, within that, take the best independent score on structured output and tool calls.

Licence
Apache-2.0
Weights at 4-bit
18 GB
Context
32K tokens
Publisher
Alibaba (Qwen Team)

What the hardware costs

One 48 GB card holds it

Rent in the EU
$1.60/hrScaleway, Paris (PAR2)
Buy the card
$7,569new, one-off
Or rent it by the token
$0.12 / $0.12per M in / out · Nebius AI Studio · EU

Hardware only, third-party prices from 2026-07. The figure excludes the KV cache, which grows with context length and how many people use it at once — sized properly in a conversation, not guessed here. Renting by the token is cheaper up front; why our customers still self-host is below.

Structured output and tool calls

Berkeley Function Calling Leaderboard · v4 · 13 of 18 models measured

Measured in a sandbox, on somebody else's functions. It tells you which models are capable of the shape of the job, not which one will survive contact with your API.

GLM-4.672.38%Kimi K259.06%DeepSeek-V3.254.12%Qwen3-32B48.71%Qwen3-235B-A22B47.99%Qwen3-8B42.57%Qwen3-30B-A3B41.39%Qwen3-14B41.03%
Longer is betterFree to serveConditions apply⚠ answered under 95%

Independent measurement · UC Berkeley (Gorilla project) · board updated 2026-04-12

The API is cheaper per token. Here is why our customers don't use it.

We will not pretend otherwise: renting a model by the token from a serverless API costs less per million tokens than a card we run for you. We show that price on every use-case page. What it does not include is the part a shop with a customer database actually pays for.

  1. 01

    Your data never leaves hardware you can point at

    A serverless "we don't retain your data" is a clause in a contract. Running the model on a card in Amsterdam is a fact of architecture: your catalogue, tickets and customer records are never sent to a third party at all. For a GDPR audit, that is the difference between a promise and a floor plan.

  2. 02

    The price cannot move without your say-so

    A serverless rate card is somebody else's lever. The provider can raise the price, retire the model, or change the terms, and your cost moves with it. The same model on the same card costs the same next year — you own the number.

  3. 03

    The model cannot be taken away

    Hosted APIs deprecate models on their own schedule; the one you built on can be gone in a quarter. An open-weight model on your own hardware runs for as long as you keep the lights on. No vendor can end-of-life it out from under you.

And the price gap closes with volume: past a card you keep busy — very roughly four billion tokens a month — owning is cheaper outright, even before the three reasons above.

Getting a case like this one from a conversation to production takes about two months, and you can stop at the end of any phase.

Next

Tell us what your team does by hand

Describe the process that takes the most time. We will say plainly whether a model is the right tool for it.