What fine-tuning an API actually means
Fine-tuning an API means taking a pre-trained model that a company has already built and adjusting it with your own data so it performs better on your specific task. Instead of using the model exactly as it comes out of the box, you feed it examples of the work you want it to do, and it learns from those examples to get better at that particular job.
Think of it like this: a general-purpose language model knows how to write in many styles and answer many kinds of questions. But if you want it to write product descriptions in your company's voice, or answer customer support questions the way your team does, you send it a batch of real examples from your own work. The model adjusts its internal weights based on those examples, so the next time you ask it something similar, it sounds more like you and less like a generic AI.
Most major API providers — OpenAI, Anthropic, Google, and others — now offer fine-tuning as a paid feature. You upload your training data, the provider's servers do the computational work, and you get back a custom version of the model that you can call just like the standard one, but with better results for your use case.
Key Takeaways
- Fine-tuning works by uploading examples of the task you want the model to do better at, and the API provider adjusts the model's weights based on those examples.
- You need at least 10 to 50 high-quality training examples for fine-tuning to show real improvement, though more examples usually produce better results.
- Fine-tuning costs money — typically a per-token charge for training plus higher per-token costs when you use the fine-tuned model — so it makes sense only when the improvement justifies the expense.
- The data you upload to fine-tune a model is usually retained by the API provider for safety and abuse prevention, so do not send proprietary information you cannot afford to have stored.
- Fine-tuning takes hours to days depending on the model size and your data volume, so plan ahead rather than expecting instant results.
Preparing your training data
The quality of your training data determines the quality of your fine-tuned model. A smaller set of really good examples beats a large set of mediocre ones. Each example should show the model exactly what you want it to do: the input you give it, and the output you want back.
Most API providers want your data in a specific format — usually JSONL (JSON Lines), where each line is a separate training example. For a customer support use case, one line might look like: {"messages": [{"role": "user", "content": "How do I reset my password?"}, {"role": "assistant", "content": "Go to the login page, click 'Forgot Password', enter your email, and follow the link we send you."}]}. Each example needs both the question and the answer you want the model to give.
Start with 10 to 50 examples if you are testing whether fine-tuning will help at all. If you see improvement, collect more — 100 to 500 examples usually produces noticeably better results. Beyond that, returns diminish unless your task is very complex. Make sure your examples are representative of the real inputs you will actually send the model, not edge cases or unusual scenarios.
Remove duplicates and examples where the output is wrong or unclear. If you are not sure what the right answer should be, the model will not be either. If your examples contradict each other — saying one thing in example 3 and the opposite in example 47 — the model will learn to average them rather than pick the right one.
Choosing which model to fine-tune
Not every model can be fine-tuned through every API. OpenAI currently allows fine-tuning on GPT-4o mini and GPT-3.5 Turbo, but not on their largest models. Google allows fine-tuning on Gemini 1.5 Flash. Anthropic's Claude models do not currently support fine-tuning through their API, though they offer it through a separate service for enterprise customers.
Smaller models are cheaper to fine-tune and faster to train, but they may not be smart enough for complex tasks. Larger models cost more but handle nuance better. Start with the smallest model that can do the job at all, fine-tune it, and see if the results are good enough. If not, try the next size up.
Check the provider's documentation for which models support fine-tuning right now — this changes frequently as companies add or remove the feature. Also check the pricing: some providers charge per token for training data, some charge a flat fee per fine-tuning job, and some charge both. Calculate the cost before you start, because a large training dataset on an expensive model can run into hundreds of dollars.
The fine-tuning process and what to expect
Once your data is formatted correctly and uploaded, the process is usually straightforward. You call the fine-tuning endpoint with your data file and a few settings — how many times the model should see your training data (called "epochs"), and sometimes a learning rate that controls how aggressively the model adjusts itself. Most of the time the default settings work fine, so you can leave those alone.
Training time depends on your data size and the model. Fine-tuning GPT-3.5 Turbo on 100 examples might take 30 minutes to an hour. Fine-tuning a larger model on 1,000 examples might take several hours. You can check the status of your job through the API or the provider's dashboard, and most providers send you a notification when training is done.
Once training finishes, you get back a model ID for your fine-tuned version. You use that ID exactly like you would use the standard model ID — you just swap it into your API calls. The fine-tuned model costs more per token than the base model, so you are paying for both the training and the ongoing use.
When fine-tuning actually makes sense
Fine-tuning is not always the right answer. If the base model already does what you need, fine-tuning wastes money. If you only need the model to follow a specific instruction once or twice, prompt engineering — writing a really detailed prompt that explains what you want — is cheaper and faster.
Fine-tuning makes sense when you have a repetitive task where the model's output is consistently off in the same way. If your customer support chatbot keeps giving answers that are too technical, and you have 50 examples of the simpler language you want, fine-tuning will fix that. If you are writing product descriptions and the model keeps missing your brand voice, fine-tuning on your actual descriptions will teach it your style.
It also makes sense when you need the model to know about specific information that was not in its training data — your company's policies, your product catalog, your internal terminology. Fine-tuning on examples that use that information teaches the model to use it correctly.
Fine-tuning does not make sense for one-off tasks, for tasks where the base model is already excellent, or when you need real-time updates (fine-tuning is a batch process, not something you can do continuously). It also does not work well if your task requires the model to know facts that change frequently — you would be fine-tuning constantly.
Data privacy and what happens to your training data
When you upload data to fine-tune an API model, that data goes to the provider's servers. Most major providers say they retain your data for safety and abuse prevention — to make sure nobody is using fine-tuning to create harmful outputs. They typically do not use your data to improve their base models, but the exact policy varies by provider.
If your training data contains confidential information — customer names, internal processes, proprietary formulas — you need to understand what the provider does with it. Read the fine-tuning section of their privacy policy and their data retention terms. Some providers offer data deletion on request, some keep it indefinitely, and some let you opt out of certain uses.
If you cannot afford to have your data stored by the provider, do not fine-tune through their API. You could instead run an open-source model on your own servers and fine-tune it locally, but that requires significant technical setup and computing power. For most teams, the trade-off is worth it — the convenience and quality of a fine-tuned API model outweighs the privacy cost — but you have to make that choice consciously.
Measuring whether fine-tuning worked
Before you fine-tune, decide how you will know if it worked. The best approach is to set aside some of your data as a test set — do not include it in training. After fine-tuning, run both the base model and your fine-tuned model on those test examples and compare the outputs side by side.
For some tasks, you can measure automatically: if you are fine-tuning a model to classify customer feedback as positive or negative, you can count how many it gets right. For others, you need human judgment: if you are fine-tuning for writing quality or tone, you or someone on your team has to read the outputs and decide if they are better.
If the fine-tuned model is not noticeably better than the base model, you have a few options. You can collect more training data and try again — sometimes 50 examples is not enough. You can adjust your training data to remove contradictions or unclear examples. Or you can accept that fine-tuning is not the right tool for this task and go back to prompt engineering or a different approach.
Frequently Asked Questions
How much does fine-tuning cost?
Costs vary by provider and model size. OpenAI charges per token for training data (usually $0.03 per 1 million tokens) plus higher per-token costs when you use the fine-tuned model. A fine-tuning job on 100 small examples might cost $5 to $20 total. Larger datasets or bigger models can cost $100 to $500 or more. Check the provider's pricing page for exact rates.
Can I fine-tune a model on data from a different provider?
No. You fine-tune through the provider's API using their infrastructure. If you want to fine-tune OpenAI's model, you use OpenAI's fine-tuning service. If you want to use Google's model, you use Google's service. You cannot take data to one provider and fine-tune a different provider's model.
What if I want to keep my training data completely private?
You would need to run an open-source model on your own servers and fine-tune it locally. Models like Llama 2 or Mistral can be fine-tuned on your own hardware, but this requires technical expertise and computing resources. For most teams, this is more expensive and complicated than using an API provider.
How long does fine-tuning take?
Training usually takes 30 minutes to several hours depending on your data size and the model. Smaller models and smaller datasets train faster. You can check the status through the provider's dashboard, and most send a notification when training is complete.
Can I fine-tune a model that is already fine-tuned?
Some providers allow it, but it is usually not necessary. If you want to add new examples to your training data, it is better to combine them with your original data and fine-tune from the base model again rather than fine-tuning on top of a fine-tuned version.