Backstory#
I hit a wall trying to set up local fine-tuning. I’d bought an AMD card for the price-to-performance ratio, and it’s fine for running models that are already trained — but trying to set up a local fine-tuning environment is where things fell apart:
- AMD’s software ecosystem lags well behind NVIDIA’s. CUDA is genuinely painless to get running, practically foolproof. ROCm, by contrast, threw error after error during install, and tracking down fixes was a slog — made worse by the fact I was also on a fresh Ubuntu 24.04 install, so there wasn’t much prior art to lean on.
- After finally fighting my way through the install, I discovered AMD’s own site doesn’t even list ROCm support for my consumer-grade card.
- So I looked at NVIDIA card prices again — and they’re well outside my budget.
After all that, I decided to fine-tune using cloud compute instead. Of the options out there, I went with Google Colab — mainly because it’s free, which settled every other consideration.
Giving up on local deployment turned out to be a relief: cloud-based fine-tuning meant I could still produce my own “customized” model, entirely for free — as long as your data doesn’t involve anything private or sensitive.
Deployment steps#
1. Open the unsloth project
https://github.com/unslothai/unsloth
Pick Llama 3 for training, and it walks you straight into running the process on Google Colab.
2. Pick a GPU type
The free tier’s T4 GPU is enough — 15GB of VRAM.

3. Follow unsloth’s steps one by one
At this point you need to swap in your own training set.

The training data needs to be in a specific question-answer format — you can use ChatGPT or a Python script to convert your own question bank or text into this shape. Once you’ve generated the JSON file, upload it to https://huggingface.co and swap that link into the Colab notebook.

Then just continue running the rest of the notebook.

Testing the result#
After training, I asked it “who are you?” — testing with a new model each time — and it could already handle the question in more than three languages.

You can see the fine-tuned model already handles domain-specific questions competently.

That’s a genuinely solid answer — I’d bet this model could pass a professional certification exam at this point.


