What fine-tuning does and when you need it
Fine-tuning means taking a pre-trained language model — one that already knows how to write, answer questions, and follow instructions — and training it further on your own specific data so it learns your domain, style, or task. You start with a model that works reasonably well at everything, then make it work much better at the one thing you actually need.
You do not need to fine-tune if you can get the results you want by writing better prompts to a model like ChatGPT or Claude. Fine-tuning costs time and money, and it makes sense only when the base model consistently misses what you need. Common reasons to fine-tune: your data is highly specialized (medical records, legal contracts, internal company documents), you need a consistent output format the model keeps getting wrong, or you want to reduce how often the model refuses your requests because it misunderstands your intent.
Fine-tuning also makes your model faster and cheaper to run at scale. A fine-tuned smaller model can often outperform a larger base model on your specific task, which means lower API costs and faster response times if you deploy it yourself.
Key Takeaways
- Fine-tuning works best when you have 100 to 1,000 examples of the exact input-output pairs you want the model to learn, formatted as JSON lines with "prompt" and "completion" fields.
- OpenAI's fine-tuning API, Hugging Face's Transformers library, and services like Modal or Together AI are the most common routes, each with different costs and setup complexity.
- The process involves preparing your data, uploading it, running training (which takes minutes to hours depending on data size), and then testing the fine-tuned model before using it in production.
- Fine-tuning a small model costs between $5 and $50 for a typical dataset, while fine-tuning larger models or using more data can run into hundreds of dollars.
- You will need to monitor your fine-tuned model's performance over time because it can overfit to your training data or drift if the real-world inputs change.
Preparing your training data in the right format
The quality of your fine-tuned model depends almost entirely on the quality of your training data. You need examples of the exact task you want the model to perform — not general examples, but real input-output pairs that show the model what success looks like for your use case.
Most fine-tuning services expect data in JSONL format (JSON Lines), where each line is a separate JSON object. For OpenAI's fine-tuning API, each line should contain a "prompt" field and a "completion" field. The prompt is what you will send to the model, and the completion is what you want it to output. For example:
{"prompt": "Classify this customer email as urgent or routine: I cannot access my account and my payment is due tomorrow.", "completion": " urgent"} {"prompt": "Classify this customer email as urgent or routine: Can you tell me more about your premium plan?", "completion": " routine"}
Start with at least 50 to 100 examples, though 500 to 1,000 is better if you want strong results. More data almost always improves performance, but you face diminishing returns after a few thousand examples. Make sure your examples are representative of the real inputs the model will see — if your training data is all formal emails but the model will receive casual Slack messages, it will not perform well on Slack.
Clean your data before uploading. Remove duplicates, fix obvious typos in the prompts (but keep them in completions if they are part of what you want the model to learn), and make sure every example is actually correct. A single incorrect example teaches the model the wrong thing, and bad training data produces a bad model.
Choosing a fine-tuning platform and model
OpenAI's fine-tuning API is the most straightforward entry point if you are already using ChatGPT or GPT-4. You upload your JSONL file through their dashboard or CLI, select a base model (GPT-3.5 Turbo is cheaper; GPT-4 is more capable), and training starts immediately. Costs run about $0.03 per 1,000 training tokens for GPT-3.5 Turbo and $0.12 per 1,000 tokens for GPT-4, plus usage costs when you run the fine-tuned model. Training usually completes in 30 minutes to a few hours.
Hugging Face's Transformers library is free and open-source, but requires you to write Python code and either run training on your own machine or rent cloud compute (AWS, Google Cloud, or Hugging Face's own infrastructure). This route gives you full control and no per-token costs, but demands more technical setup. It is the right choice if you want to run the model yourself rather than through an API, or if you need to fine-tune a model that OpenAI does not offer.
Specialized services like Modal, Together AI, or Replicate sit in the middle: they handle the infrastructure for you, support multiple base models, and charge per training hour rather than per token. These are useful if you want to fine-tune a model other than GPT (like Llama 2 or Mistral) without managing your own servers.
Choose based on what you already use and what you need to deploy. If you are already in the OpenAI ecosystem and want the simplest path, use their API. If you need to run the model locally or want to avoid per-token costs, use Hugging Face. If you want a middle ground with more model options, try a specialized service.
Running the fine-tuning job and monitoring training
Once your data is formatted and uploaded, the actual fine-tuning process is straightforward. With OpenAI, you upload your file, confirm the format is correct (they will validate it and report errors), and click start. The system shows you training progress in real time — you can watch the loss (a measure of how wrong the model is) decrease as training proceeds.
Training time depends on your data size and the model you chose. Fine-tuning GPT-3.5 Turbo on 500 examples typically takes 15 to 45 minutes. Larger models or larger datasets take longer. You do not need to stay logged in — training runs in the background and you get a notification when it is done.
While training runs, you will see metrics like training loss and validation loss. Training loss should decrease steadily; if it plateaus or increases, your model may be overfitting (memorizing your training data rather than learning the underlying pattern). Most platforms let you stop training early if you see this happening. Validation loss tells you how the model performs on data it has not seen before, which is a better indicator of real-world performance.
After training completes, you get a model ID or checkpoint you can use immediately. Test it on a few examples before deploying it to production. Ask it the same questions you asked the base model and compare the outputs. If the fine-tuned model is not better, the training data may not have been representative enough, or you may need more examples.
Testing and deploying your fine-tuned model
Before using your fine-tuned model in production, run it against a test set — examples you did not include in your training data. This tells you whether the model actually learned the task or just memorized your training examples. If performance on the test set is much worse than on the training set, you have overfitting and need either more training data or a smaller model.
Compare the fine-tuned model's output to the base model's output on the same inputs. You should see clear improvement on your specific task. If the improvement is marginal, fine-tuning may not have been worth the cost and complexity — you might get better results by refining your prompts instead.
Deployment depends on where you fine-tuned. If you used OpenAI, you call the fine-tuned model the same way you call the base model, just with a different model ID in your API request. If you used Hugging Face, you download the model weights and either run them locally using the Transformers library or deploy them to a service like Hugging Face Inference API or your own server. If you used a specialized service, they usually provide an API endpoint you can call directly.
Monitor your fine-tuned model in production. Track how often it produces the outputs you expect, and watch for performance drift — a gradual decline in quality over time. This often happens when real-world inputs start to differ from your training data. If drift occurs, collect new examples of the failures and retrain.
Understanding the costs and trade-offs
Fine-tuning costs money, and the total depends on your platform and model size. OpenAI charges for training tokens (the words in your training data) and then for every token the fine-tuned model generates when you use it. A typical fine-tuning job on 500 examples costs $5 to $15 in training fees, plus usage costs. If you run the model 10,000 times, usage costs can easily exceed training costs.
Hugging Face and open-source fine-tuning have no per-token costs, but you pay for compute time. Renting a GPU to fine-tune for a few hours costs $5 to $20 depending on the GPU type. Running the model yourself after that has no ongoing costs beyond your own infrastructure.
The trade-off is control versus simplicity. OpenAI is the simplest but most expensive at scale. Hugging Face gives you full control and lower long-term costs but requires more technical knowledge. Specialized services split the difference.
Before fine-tuning, estimate whether the improvement in model performance justifies the cost. If you are fine-tuning to reduce API costs, calculate how many times you will run the model and whether the savings outweigh the training expense. If you are fine-tuning to improve accuracy, measure the current error rate and decide whether a 10% or 20% improvement is worth the investment.
Common mistakes and how to avoid them
The most common mistake is training on too little data. If you have fewer than 50 examples, fine-tuning often makes the model worse, not better — it learns noise instead of signal. Start with at least 100 examples, and aim for 500 if possible.
The second mistake is including the same example in both training and test data. This makes your model look better than it actually is. Always hold back a separate test set that the model never sees during training.
The third mistake is fine-tuning when better prompting would work. Before you fine-tune, spend an hour writing detailed, specific prompts and testing them against the base model. Many tasks that seem to need fine-tuning actually just need clearer instructions.
The fourth mistake is forgetting that fine-tuned models can become outdated. If the base model gets updated, your fine-tuned version does not automatically improve. You may need to retrain on the new base model to get the benefits of its improvements.
Frequently Asked Questions
How much training data do I actually need?
Start with 50 to 100 examples if you are testing whether fine-tuning helps at all. For production use, aim for 500 to 1,000 examples. More data almost always improves results, but you see the biggest gains in the first few hundred examples. After 5,000 examples, improvements typically slow down.
Can I fine-tune a model to refuse certain requests or follow safety guidelines?
Yes, but it is tricky. If you include examples where the model refuses a request, it learns to refuse similar requests. However, fine-tuning alone is not a reliable safety mechanism — the base model's training still influences behavior. Use fine-tuning to reinforce guidelines, not as your only safety layer.
What happens if I fine-tune on data that contains errors or biases?
The model learns those errors and biases. If your training data is biased toward one outcome, the fine-tuned model will be too. Always audit your training data for obvious errors and imbalances before uploading. If your data is 90% positive examples and 10% negative, the model will lean toward positive.
Can I fine-tune a model that is already fine-tuned?
Yes, you can fine-tune a fine-tuned model, but results are often worse than fine-tuning the base model directly. Each round of fine-tuning increases the risk of overfitting. If you need to combine multiple datasets, merge them and fine-tune the base model once rather than fine-tuning sequentially.
How do I know if my fine-tuned model is actually better than the base model?
Test both on the same examples and compare outputs side by side. Measure a specific metric: accuracy on classification tasks, exact match on extraction tasks, or human ratings on generation tasks. If the fine-tuned model scores higher on your metric and the improvement is large enough to justify the cost, fine-tuning worked.