Booking now for Q4 delivery Every engagement starts with the free assessment Human-centric AI for Main Street Denver · Phoenix · Remote
Main & Machine / The Field Guide / Where your data goes
The Field Guide / 09

Where does your AI data actually go?

Most owners sign the vendor agreement without reading the data clause. This guide is that clause, translated — cloud, on-premise, and the middle path, by data type.

Field GuideNo. 09
Reading time9 min
Paths comparedCloud · hybrid · on-prem
On-prem proofMARCUS · 14 agents
UpdatedJuly 2026

Where does your data go when you use cloud AI?

When you paste text into ChatGPT or send a document through an AI API, that content travels over the internet to the provider’s servers, gets processed there, and the response comes back. Your prompt, your document, and usually some metadata physically leave your building.

That is not a scandal; it is how cloud software has worked for twenty years. Your email, your accounting system, and your CRM already live on someone else’s servers. But AI feels different for a reason: the whole point of the tool is to read your material — contracts, customer lists, financials, the things you would not forward to a stranger. So the question is not “does my data leave?” It does. The questions that matter are what the provider is allowed to do with it once it arrives, and what is inside the documents you are sending. The Ampersand essay Where Your Data Goes walks the journey step by step; this guide is the decision framework built on top of it.

Does the provider train on your data?

On business and API tiers, usually not — the major providers’ standard terms say business data is not used to train their models. Consumer tiers are different: free and personal accounts often permit training use unless you find the setting and opt out.

This is the single most misunderstood distinction in small-business AI, and it lives in the pricing page, not the product. The $20 consumer subscription and the business API can run the exact same model with entirely different data terms. So check the tier you are actually on — not the one the vendor’s blog post describes. Then read the two other clauses that matter: retention (how long prompts are kept, and whether a zero-retention option exists) and subprocessors (who else touches the data in transit). “Contractually fine” is a real and reasonable standard for most business data. But contractual fine-ness has a boundary, and the boundary is what is in the documents.

One more risk hides outside the contract entirely: shadow AI. If your business has no sanctioned tool and no policy, employees are already pasting customer emails and spreadsheets into personal ChatGPT accounts — the consumer tier, with the consumer terms. The fix is not a ban; bans just push the pasting onto phones. The fix is giving people a business-tier tool and a one-page rule about what may go into it, which costs less than one month of pretending it is not happening.

What changes when the data is regulated?

Everything. Health records, lending files, Social Security numbers, and similar regulated data are governed by law, not just by your vendor agreement — and a clean contract with an AI provider does not by itself make you compliant.

If you run a medical or dental practice, HIPAA decides where patient data may travel, and sending records to an AI tool without a business associate agreement is a violation regardless of how good the provider’s security is. Lenders answer to examiners who will ask exactly where borrower files went and expect a documented answer. The pattern generalizes: for regulated data, “the provider promises not to train on it” is not the bar. The bar is being able to prove, to a regulator, where every record went and who read it. That is a different architecture — which is why the next two sections exist.

What does on-premise AI actually mean?

On-premise (or local) AI means the model itself runs on hardware you control — your server, your building or your private cloud tenancy — so documents are processed where they already live. Nothing goes to OpenAI, Anthropic, or anyone else, because the intelligence comes to the data instead of the data going to the intelligence.

This is not theoretical; it is how we built MARCUS for B:Side Capital, an SBA 504 / CDFI lender — a regulated business where borrower files simply do not get to leave. MARCUS runs 14 AI agents across 7 departments, entirely on-prem. Before any model reads a document, Microsoft Presidio strips the PII — names, SSNs, account numbers — so even the local model works on redacted text. Everything is encrypted at rest, every action lands in an append-only, tamper-evident audit log, and nothing sends, files, posts, or pays without human approval. That is what “your data never leaves” looks like when it has to survive an examiner. The complete architecture is public on our security page — read it even if you never hire us, because it is a useful checklist for interrogating anyone else’s claims.

Is there a middle path between cloud and on-prem?

Yes: strip the sensitive fields locally before anything goes to a cloud model. A PII-scrubbing layer such as Presidio runs on your side, replaces names, numbers, and identifiers with placeholders, and only the redacted text makes the trip.

The cloud model still does what it is good at — reading, summarizing, drafting — but the identifying details never left your network, and the response gets re-matched to the real records locally. This hybrid pattern covers a surprising share of real cases: a firm that wants cloud-grade models on customer correspondence without shipping the customer list, a clinic that wants help with documentation without transmitting patient identities. It costs less than full on-prem because you are not hosting the model, and it converts “we can’t use AI, our data is sensitive” into an engineering detail. Which model family sits behind it — open weights you host or a closed API you call — is its own trade-off, covered in Open Models, Closed Models.

Which deployment fits which data?

Match the deployment to the most sensitive data in the workflow, not to the average. One table covers the sensible defaults.

Data in the workflow Examples Sensible deployment Why
Public Marketing copy, published prices, job posts Any cloud tool, any tier Nothing to protect — optimize for cost and quality
Internal business SOPs, drafts, meeting notes, non-sensitive financials Cloud on a business/API tier No-training terms and retention controls are enough
Customer PII Names + contact details, order histories, correspondence Cloud with local PII stripping (hybrid) The model helps; the identifiers never travel
Regulated Health records, lending files, SSNs On-prem, or hybrid only with counsel’s sign-off Law sets the bar; you must prove where every record went

Defaults, not legal advice — a workflow inherits the requirements of the most sensitive data that touches it.

The inheritance rule is the one people miss. An invoice-drafting workflow sounds like internal business data — until you notice the invoices contain patient names, at which point the whole workflow is regulated. Classify by the worst thing in the pipe, not the average.

Does a 20-person business actually need on-prem?

Usually not. If your most sensitive asset is internal documents and everyday customer records, a business-tier cloud deployment with sane retention settings — plus PII stripping where customer data flows — is the right call, and on-prem would be paying enterprise money to solve a problem you do not have.

On-prem costs real money: hardware or dedicated hosting, maintenance, and models that trail the frontier. It earns that cost in exactly three situations — a regulator can ask where your data went; a contract (often with your own enterprise customers) forbids third-party processing; or a breach of one specific dataset would end the business. If none of those describes you, take the cloud dividend: better models, lower cost, and a running bill that typically lands at $50–$500 a month rather than an infrastructure line item. Anyone steering a 20-person e-commerce shop toward on-prem is selling complexity. Anyone waving a lender onto consumer-tier ChatGPT is selling negligence. The deployment should be an output of the diagnostic work, never a default — and it is one of the ten questions in our guide to vetting consultants.

What should you ask before signing any AI vendor agreement?

Five questions, answered in writing: Is our data used for training on this tier? How long are prompts and outputs retained, and can retention be set to zero? Who are the subprocessors? Where is the data processed geographically? And what happens to our data when we leave?

A serious vendor answers all five from documents they already have — a data processing addendum, a security page, a subprocessor list. Evasion on any of them is your answer, and so is the phrase “we take security very seriously” offered in place of a document. Keep the answers in the deal file; if a customer or regulator ever asks where their data went, that folder is the difference between an afternoon and a very bad month. If you want the same interrogation applied to your own setup — what data your workflows actually touch, and which deployment that implies — that is part of what an AI Readiness Audit maps, at the published prices on the pricing page. Or start smaller: book the free 30-minute assessment, tell us what data you handle, and we will give you a straight read on whether cloud, hybrid, or on-prem even belongs in your conversation. We reply within 24 hours.

Fair questions

Where AI data goes.

Want the numbers?

The full price list is published — audits, sprints, managed services, all on one page.

Read the price list
01Does ChatGPT use my business data to train its models?+

On business and API tiers, the major providers' standard terms say no; consumer and free tiers often permit training use unless you opt out. Check the tier you are actually on — the data terms live in the plan, not the model.

02Does a small business need on-premise AI?+

Usually not. On-prem earns its cost in three situations: a regulator can ask where your data went, a contract forbids third-party processing, or a breach of one dataset would end the business. Otherwise a business-tier cloud deployment, with PII stripping where customer data flows, is the right call.

03How do you use AI on regulated data like health or lending records?+

Run the model on hardware you control so records never leave, and strip identifying details before any model reads a document. That is how MARCUS runs at an SBA lender: 14 agents entirely on-prem, PII removed by Microsoft Presidio first, encrypted at rest, with an append-only tamper-evident audit log.

04What is the middle path between cloud AI and on-premise AI?+

Local PII stripping before cloud calls: software on your side replaces names, numbers, and identifiers with placeholders, and only the redacted text travels to the cloud model. You get cloud-grade models without shipping the sensitive details, at far less cost than hosting a model yourself.

Start here

Not sure which deployment your data needs?

Tell us what your workflows touch — patient records, borrower files, or plain business documents — and a senior advisor gives you a straight read in 30 minutes. Free, and we reply within 24 hours.