AI & Machine Learning
Local AI
Data Privacy
Small Business
RAG
Fine-Tuning

How to Train AI on Your Business Data Without Sharing It

Most small businesses never need to train an AI model. Here is how to make local AI know your business from your own documents and keep client data in house.

H
Hrishi Digital Solutions
13 min read
How to Train AI on Your Business Data Without Sharing It - Hero Image

Somebody in your business has probably pasted a client's details into a chatbot this month. They were not being careless. They had a quote to write or a long email to answer, and the tool was open in the next tab.

The usual response from owners is "we should train our own AI on our business data, and keep it in house." That instinct is right about the "in house" part. The "train" part is where most people get sent down an expensive path they do not need.

This is the second article in our series on local AI for Australian small businesses. Part 1 covered what local AI costs and which hardware suits it. This part covers how to make a local model useful for your business in particular: what "training" means in practice, which approach to start with, and what has to be true before you can call the setup private.


Client data is already leaving the building

A 2026 survey run by Wakefield Research for PagerDuty asked 1,250 office workers in Australia, Japan, the UK and the US how they use public AI tools at work. 34 per cent said they had entered customer data, and 31 per cent had entered financial information or confidential company documents. Those respondents worked at large companies with compliance teams and written policies. A ten-person business with neither is unlikely to be doing better.

The Australian privacy regulator has been plain about this. In its guidance on commercially available AI products, the OAIC recommends, as a matter of best practice, that organisations do not enter personal information into publicly available generative AI tools, and sensitive information in particular. It also tells businesses to check whether a product uses their inputs to train its models, and whether that can be switched off.

Banning the tools rarely works. People use them because they save real time. The fix that sticks is giving staff something just as handy that keeps the data where it belongs.


What "training AI on your data" really means

When people say they want to train AI on their business, they usually mean one of three different things. They differ a lot in cost and effort.

Approach In plain English Good for What you need
Instructions You tell the model who it is, how to write, and what rules to follow, every time it runs Tone, house rules, simple repeatable tasks An afternoon and someone who knows the job well
Retrieval (often called RAG) The model looks up your documents before it answers, like a new staff member checking the manual Anything that depends on your facts: prices, policies, past jobs, product codes Your documents in reasonable order, and a machine to index them
Fine-tuning You change the model itself by showing it hundreds of worked examples Getting the same format or style every time, at volume A clean set of examples, a capable GPU, and time to test

Only the third one is training in the technical sense. For most small businesses, the first two do nearly all the work.

The guides and research we reviewed for this article agree on the order: start with instructions and retrieval, and only reach for fine-tuning once you have shown the simpler setup falls short.


Start with retrieval: let the model look things up

Retrieval-augmented generation, or RAG, sounds more complicated than it is. Your documents are indexed on your own machine. When someone asks a question, the system finds the handful of pages most likely to hold the answer and hands them to the model along with the question. The model writes its answer from those pages.

Nothing about the model changes. It has simply been given the right paperwork at the right moment.

That design suits a small business for four practical reasons:

  • It stays current. Update the price list on Monday and the answers use the new prices on Monday. There is nothing to retrain.
  • It can show its working. A good setup tells you which document and page an answer came from, so staff can check it in seconds.
  • You can remove things. Delete a client's file from the index and the system can no longer draw on it. That matters if a client asks you to erase their records.
  • Access rules can carry over. The index can respect who is allowed to see which folders, so the apprentice does not get answers from the payroll files.

The catch is that retrieval is only as good as the documents behind it. If your current price list lives in three spreadsheets and nobody is sure which is right, the model will not know either. In our experience, tidying the source documents is most of the work, and it pays off whether or not you ever use AI.

If you want the technical detail on how the index works, we cover it in our guide to vector databases for mission-critical RAG.


When fine-tuning is worth the effort

Fine-tuning earns its place when the problem is how the model behaves, not what it knows.

Typical cases:

  • You need every output in exactly the same structure, such as a job sheet with fixed fields, and instructions alone keep drifting.
  • You sort thousands of similar items a month, such as inbound emails or support tickets, into your own categories.
  • You have a house writing style that is hard to describe but easy to recognise from examples.

The common method today is called LoRA, or QLoRA in its lighter form. Instead of retraining a whole model, you train a small add-on that sits on top of it. That brings the job within reach of a single workstation, where it once needed a rack of servers.

It still has costs that are easy to miss:

  • You need good examples, and enough of them. A few hundred carefully checked examples is a common starting point. Messy examples teach the model your mistakes.
  • It goes stale. Anything the model learned in training is frozen on that day. Fine-tune it on last year's prices and it will quote last year's prices with complete confidence.
  • It cannot cite a source. A fine-tuned model answers from memory, so there is no page to check.
  • It can memorise what you feed it. Research on these systems notes that a fine-tuned model can end up reproducing private text from its training data. If you train on real client records, strip out names, addresses and account details first.
  • You redo it when you change models. The add-on is tied to the base model it was trained on.

The practical rule: keep facts in documents, where retrieval can find them and you can change them. Use fine-tuning for format and style. Many working systems combine the two, a pattern described in the T-RAG paper from a team that built one for a large organisation.


What this looks like in a normal week

These are illustrations of how the pieces fit together, not client results.

Drafting quotes for a trades business. An enquiry email arrives. The system pulls the current rate card and the three most similar past jobs, then drafts a quote in your template. The estimator reads it, fixes what needs fixing, and sends it. The prices come from your rate card through retrieval. A person still approves every quote.

Checking invoices and delivery dockets. A supplier invoice arrives as a PDF. A local model reads the line items and the system compares them with the purchase order in Xero or MYOB. Anything that does not match is flagged for a person to look at. We explain the document-reading side in our article on intelligent document processing.

Answers on site with no signal. A field technician has a laptop with a small model and the equipment manuals indexed on it. They can ask how to reset a particular controller and get the relevant page, with no mobile coverage needed. This is one area where local AI does something cloud AI cannot.

In all three, the AI is one step inside a process that already exists. If you are still working out which process to start with, our piece on practical AI automation for Australian SMBs lists the ones that tend to pay back first.


Local does not automatically mean private

Running a model on your own hardware removes the biggest exposure, which is sending prompts and documents to an outside AI provider. It does not finish the job. Before you tell clients their data stays in house, check these:

  1. Which models are really local. Some local AI tools now list cloud-hosted models in the same menu as the ones that run on your machine. Make sure staff can only pick the local ones for sensitive work.
  2. Telemetry and updates. Check what the software sends back to its maker and turn off anything you do not need.
  3. Connected tools. If the model can call a web search, an email service or an outside API, data can leave through that door.
  4. Chat history and logs. Past conversations are stored somewhere. Decide who can read them and how long they are kept.
  5. Backups. If the machine backs up to an overseas cloud service, so do your indexed documents.
  6. Who can log in. Give each person their own account, and set folder permissions before you index anything.
  7. Physical access. A small computer on a shelf is easy to carry out the door. Encrypt the disk.

There is also a middle option. Several Australian providers now host open models in local data centres, which can suit a business that wants its data onshore but does not want to own hardware. If you go that way, ask where the data is stored, who operates the service, and who can access it. A data centre on Australian soil does not by itself tell you who controls the data.


Does the Privacy Act apply to my small business?

A date is doing the rounds that deserves a correction. Several articles this year have said the small business exemption from the Privacy Act ends on 10 December 2026. On our reading of the published material, that is not what happens on that date.

Here is where things stand as of October 2026. This is general information, not legal advice.

  • 10 December 2026 is when businesses covered by the Privacy Act must explain in their privacy policy if they use personal information in automated decisions that significantly affect people.
  • The small business exemption for most businesses turning over A$3 million or less is still in place. The government has agreed in principle to remove it eventually, but commentary on the draft legislation released on 31 August 2026 indicates the draft does not do so.
  • Some small businesses are covered anyway. Health service providers and businesses that trade in personal information have long been covered regardless of turnover. Privacy consultancy Helios Salinger reports the OAIC's estimate that anti-money laundering changes from 1 July 2026 bring more than 100,000 small businesses under the Act as well.

Even if you are exempt today, your clients are not exempt from caring. A bookkeeper, a clinic or a law practice that can say "your documents are never sent to an outside AI provider" has something its competitors may not.


A sensible order to do this in

  1. Find out what staff already use. Ask without blame. You need an honest list of tools and the kinds of data going into them.
  2. Pick one job. Choose something frequent, repetitive and document-heavy, such as quoting or answering questions about your own procedures.
  3. Tidy the documents for that job only. One current version of each, in one place.
  4. Set up retrieval on a local model and test it with real questions. Use questions your staff asked last month, and have the person who knows the answers mark the results.
  5. Consider fine-tuning last. Only if the answers are right but the format or style still will not hold.

Most businesses will stop at step 4 for a long while, and that is a perfectly good place to be.


Common questions

Do I need to train a model for it to know my business?

Usually not. Giving a local model access to your documents through retrieval covers most needs, and it is quicker to set up and easier to keep current than fine-tuning.

How many examples does fine-tuning need?

A few hundred well-checked examples is a common starting point for a narrow task. Quality matters more than the count. Ten sloppy examples do more harm than good.

Can a local model leak client data?

It can if the setup around it is loose. The model itself does not send data anywhere, but connected tools, backups, shared logins and cloud-hosted model options can. The checklist above covers the main gaps.

Is a local model good enough for this work?

For defined jobs such as looking up your own documents, drafting from templates and reading invoices, a suitable local model does well. The largest cloud models still lead on the hardest reasoning tasks. Part 1 of this series covers the hardware and model choices.


Key takeaways

  • "Training AI on your data" usually means retrieval, where the model looks up your documents. It rarely means changing the model.
  • Keep facts such as prices and policies in documents. Use fine-tuning only for consistent format and style.
  • Tidy source documents are the biggest factor in whether the answers are any good.
  • Local hardware removes the main privacy exposure, but backups, connected tools and logins still need checking.
  • The 10 December 2026 date is about disclosing automated decisions. The small business exemption remains in place for now.

Want to know if this would work on your documents?

Book a free 30-minute call and bring one job you would like off your plate. We will tell you whether retrieval on a local model is likely to handle it, what your documents would need first, and whether a cloud tool with the right settings would do the job for less. Our work is fixed scope and priced up front, so you know the cost before anything is built. You can also read how we approach this on our AI and workflow automation service page.

Local AI
Data Privacy
Small Business
RAG
Fine-Tuning
H

Hrishi Digital Solutions

Enterprise web application and AI automation specialists. We help Australian businesses and government agencies choose between cloud and on-premise AI on cost and compliance, not hype.

Contact Us →