How to Choose an LLM Finetuning Partner: Pre-Contract Checklist
01/10/2026

The first LLM finetuning contract I ever saw was two pages long. Not one line said how the model would be evaluated, who would keep the weights afterward, or how the customer's data would be handled. Both sides were excited, and six months later both were annoyed. This post is the list of questions I want customers to ask any company doing LLM finetuning, including us, before they sign.
Ask about the problem before the model
A trustworthy provider does not rush to quote. They ask you questions: who the users are, what the specific task is, how it is done today, what success looks like. If the first conversation circles only around model names and GPU counts, be careful.
Also ask whether finetuning is really necessary. Sometimes good prompting or RAG is enough. A partner willing to say you do not need finetuning yet is a good sign, because they are not selling you something you do not need.
Data: where many troubles begin
- What kind of data do they need, how much, in what format? A number that is too small or too large without explanation deserves a follow-up.
- Who cleans and labels it? You or them? Is that cost in the quote or billed separately?
- Where is your data stored, who can access it, and when is it deleted after the project?
- Will your data be used to train models for other customers? Get the answer in writing.
- If the data contains personal or medical information, how is it anonymized?
I have seen a project slip by three months simply because both sides assumed the other would clean the data. Put it on paper.
Evaluation: what separates real work from good talk
Ask: how do we measure success, on which test set? A serious partner will propose building the test set with you, from real user questions, kept separate from the training data. If the test set leaks into training data, the score looks great and means nothing.
Be wary of promises like "over ninety percent accuracy" without saying measured on what. A number only means something alongside the dataset, the scoring method and the task. The next question: how much better is it than the base model without finetuning? If they have no baseline, you do not know what you are buying.
Ask about failures too. Do they report the cases where the model still gets things wrong? Someone who is honest about limits is usually more trustworthy than someone who only shows strengths.
Ownership and the ability to leave
- When the work is done, who owns the model weights (or adapters)? Can you receive them and run them yourself?
- If the model runs on the partner's infrastructure, can you move it elsewhere later?
- Will source code, training scripts and configurations be handed over?
- What license does the base model carry, and does it allow commercial use?
The fourth gets forgotten. Some open models have terms restricting commercial use or requiring attribution. Choose wrongly and the whole project can end up with legal trouble. Have your legal team read the base model license, not just the leaderboard.
Infrastructure and security
Training consumes GPUs, and whoever does the work for you must have the compute or rent it somewhere. Ask directly: where does training run? If the data is sensitive, can training and deployment happen in your own environment, or at least an isolated zone? I have seen a project stall because patient records could not leave the hospital while the partner only offered a cloud option. This should surface in the first meeting. You can read more on this choice in posts from AIVISION.
Ask about access control, logging, incident procedures, and who is responsible if data leaks. If the contract says nothing about these, that is a gap to close before signing.
Domain and language expertise
Finetuning for Vietnamese is not the same as finetuning for English. Ask what the partner has done with Vietnamese: handling diacritics, regional words, unaccented writing, specialized terms. They do not need to name past customers, but you should hear them describe the kinds of problems they met and how they solved them. For fields like medicine or pharmacy, ask whether domain experts review the results, or whether engineers simply look at the output and decide it sounds reasonable.
AIVISION is an AI company in Vietnam that trains LLMs on a cluster of 24x NVIDIA H200 and 8x NVIDIA B300, has released the L1.0 LLM and the E1.0 speech-to-text model for Vietnamese, and does finetuning for domains like healthcare and pharmaceuticals. We say this not so you will default to us, but so you know the questions in this post apply to us as well.
Schedule, handover and maintenance
- What milestones does the project have, and what can you verify at each one?
- If the first results fall short, what is the iteration process and who bears the cost?
- After handover, who monitors quality as data and usage shift?
- Is there documentation so your team can operate it?
- When a better base model appears, can you upgrade, and at what price?
Models do not stand still. User data drifts, documents get updated, and quality erodes if nobody watches. A contract that only covers the first delivery has not thought about the hardest part.
Signs to slow down
- Promising specific results before seeing your data.
- Refusing to explain the method, offering only "proprietary technology."
- Declining a trial on a small slice of data before you commit to everything.
- A contract that does not spell out product ownership.
- Anyone who disparages all competitors and claims to always be the best.
The risk reducer I recommend most is a small pilot: a few weeks, one narrow task, with evaluation criteria written down in advance. You will learn a great deal about how the partner works in those weeks, far more cheaply than learning it after signing a large contract.