AI Data Residency Explained: Why Storage, Training and Inference Are Three Different Questions

Many enterprise AI procurements begin with the question: "Can our data stay in Australia?" That question contains at least three separate technical and commercial considerations. This article separates them and provides a framework for evaluating each one independently.

AI Data Residency Explained: Why Storage, Training and Inference Are Three Different Questions

One of the first questions organisations ask during an enterprise AI procurement is: "Can our data stay in Australia?"

It is an important question, but not a complete one.

For most enterprise AI procurements, it helps to separate the discussion into three distinct questions: where data is stored, whether data is used to train AI models, and where AI inference occurs. These are generally distinct concepts with different contractual implications, different commercial trade-offs, and in many cases, different answers from the same vendor for the same product.

A common pattern illustrates this:

Buyer: "Do you support Australian data residency?"

Vendor: "Yes."

Months later, the buyer discovers that documents are stored in Australia, but inference (the process of generating AI responses) occurs in a data centre overseas. Both statements may be technically accurate while still leaving important requirements unaddressed. The buyer asked a question; the vendor answered it. The difficulty is that the question was broad enough to allow an accurate answer that did not fully address the underlying requirement.

In many cases, vendors are answering the specific question they have been asked, within the scope of the product they are describing. The difficulty is that broad questions about data residency can legitimately produce answers that do not address storage, training, and inference separately. The procurement role in this context is not to catch vendors out. It is to ensure the right questions are being asked in the first place.

Marketing material and vendor documentation in this space frequently use "data residency," "data sovereignty," and "regional hosting" as interchangeable terms. Different vendors may define these concepts differently. Confirming how each vendor uses these terms at the outset of evaluation helps avoid miscommunication later in the process.

This article addresses these distinctions from a procurement evaluation perspective rather than a legal or regulatory one.

Why These Distinctions Matter for Procurement

Organisations evaluate data residency for a range of reasons: regulatory compliance obligations, privacy requirements, internal security policy, customer commitments, data sovereignty objectives, risk management, and latency. Each of these drivers may point to a different underlying technical requirement.

An organisation evaluating AI for a regulated industry use case may be focused primarily on where data is stored and whether contractual commitments prevent cross-border transfers. An organisation focused on security policy may be equally concerned with where inference occurs. An organisation with customer-facing data handling commitments may need clarity on all three questions before any platform selection.

Vendor evaluation becomes more productive once procurement has defined which specific requirements the business is trying to satisfy. A vendor response of "yes, we support Australian data residency" may be accurate in relation to one question while leaving the others entirely unanswered.

The Three Questions

Question 1: Where Is Data Stored?

For many enterprise AI products, data residency primarily refers to where customer data is stored at rest. However, vendors and regulators do not always define the term consistently, and some definitions extend to processing and backup locations as well. Confirming exactly what each vendor includes in their definition of data residency is a useful early step in evaluation. This includes uploaded documents, files, chat history, knowledge bases, embeddings and vector databases, logs, and backups.

A vendor commitment to Australian data residency typically means that this stored data remains on infrastructure located within Australia. It does not, in itself, indicate where AI processing occurs when that data is used.

What vendor evaluation at this layer typically covers. Where customer data is stored at rest, where backups are located, where logs are retained, whether data can leave Australia under any circumstances (including for support, incident response, or vendor infrastructure operations), and whether regional storage commitments are contractually guaranteed or subject to change without notice.

Question 2: Is Data Used to Train AI Models?

This question is entirely separate from geography. It concerns whether the content of customer interactions (prompts, documents, outputs) is used to improve or retrain the underlying foundation model.

The answer varies significantly by product type. Many consumer AI services retain the right to use interaction data for model improvement unless users opt out, although practices vary between providers and products. Many enterprise subscriptions and API-based services include contractual commitments preventing customer data from being used for foundation model training without explicit opt-in. The distinction between product tiers is material and worth confirming for each specific product under evaluation.

A clarification that frequently arises during vendor evaluation: Retrieval-Augmented Generation (RAG) uses customer data at inference time to generate responses. This is distinct from training or fine-tuning a foundation model, which modifies the model's parameters and capabilities. Using a customer document to answer a specific prompt is a retrieval operation. It is not model training in the technical sense, though confirming this distinction with each vendor is worthwhile, as terminology is not always used consistently in vendor documentation.

What vendor evaluation at this layer typically covers. Whether customer data is used for model training, whether this commitment varies by product tier or deployment configuration, whether opt-in is required for any training use, and whether the commitment is contractual rather than policy-based (policies can change; contractual terms create enforceable obligations).

Question 3: Where Does AI Inference Occur?

Inference is the process by which an AI model generates a response. When a user submits a query, the text is sent to a model, processed, and a response is returned. That processing occurs within one or more cloud regions, which may or may not be located in Australia. Some providers dynamically route requests, use multiple availability zones, or fail over between regions, meaning inference location is not always a single fixed answer.

Data may be stored in Australia while inference occurs in a data centre in the United States, Singapore, or elsewhere. Storage location and inference location are independent variables and do not necessarily align.

Flowchart illustrating how enterprise AI data flows: User → Australian Storage (documents, chats, logs) → Inference (which may occur in Australia, Singapore, the United States, or another region) → Foundation Model → Response. The diagram highlights that data storage location and AI inference location are separate and may occur in different geographic regions.

Storage location and inference location may differ. Evaluating them independently helps avoid incorrect assumptions during vendor assessments.

Inference location may be relevant for several reasons: cross-border data flow obligations under privacy legislation, government or regulatory policy requirements, security classification constraints, critical infrastructure considerations, and latency. The significance of inference location varies by organisation and use case.

Some platforms also perform safety filtering, orchestration, or request routing as part of the inference workflow, and these components can operate in different regions from the primary inference layer. Organisations with strict regional processing requirements may find it worthwhile to confirm where these components operate, not only where the model itself runs.

What vendor evaluation at this layer typically covers. Where inference occurs for each product under evaluation, whether this varies by model or deployment tier, whether regional inference is available and under what commercial structure, whether pre-inference processing components operate in the same region, and whether the deployment architecture changes the answer.

Why There Is Not One Answer

The answers to all three questions depend on deployment architecture, not simply on the vendor or the underlying model. The same foundation model may behave differently across deployment configurations.

A model accessed through a native SaaS platform may process inference in a different region than the same model accessed through Azure AI Foundry, Amazon Bedrock, or Google Vertex AI. A dedicated or private deployment may offer materially different residency and inference options than a standard multi-tenant subscription.

Vendors responding to data residency questions are not necessarily being misleading when different buyers receive different answers. They may be describing different deployment configurations, different product tiers, or accurately answering a narrower version of the question than the buyer intended to ask.

Understanding which deployment architecture is under discussion is therefore a prerequisite for interpreting vendor responses accurately.

Common Misconceptions

AssumptionA more accurate framing
Australian data residency means inference happens in Australia.Storage and inference are independent. A vendor may offer one without the other.
Enterprise AI platforms train on customer data by default.Most enterprise and API offerings include contractual commitments against training use. Consumer services differ.
One data residency answer applies to every deployment of a model.Deployment architecture often changes the answer materially.
Regional cloud infrastructure means every AI feature operates regionally.Feature availability varies between regions and is not guaranteed to be uniform across a vendor's product suite.

Commercial Trade-Offs

Regional deployments (where both storage and inference are confined to Australia or a specified geography) typically involve commercial trade-offs. These can include higher pricing, different licence structures, reduced model availability, slower access to new model versions and features, capacity constraints, and different service level agreements.

These trade-offs are not uniform across vendors. Some providers have invested more heavily in Australian infrastructure than others. Regional deployments may differ from global deployments in areas such as model availability, feature releases, and commercial terms.

Procurement processes that account for these trade-offs are better positioned to assess whether a regional deployment configuration meets the organisation's requirements at an acceptable commercial cost, or whether a non-regional deployment with specific contractual protections is more appropriate for the use case in question.

What Contracts Need to Cover

Technical architecture describes what a platform can do. The contract determines what the vendor is permitted to do over the life of the agreement. Both matter, and they are not always aligned.

Contract evaluation in this area commonly covers: data storage location and the conditions under which stored data may be transferred, the basis on which (if any) customer data may be used for model training, the regions in which processing (including inference) is permitted to occur, notification obligations if processing regions change, the identity and location of subprocessors, disaster recovery and backup locations, and audit or assurance rights where applicable.

A vendor may technically support Australian inference at the time of contracting while the contract itself permits processing location changes with limited notice. Enterprise AI governance frameworks that address data handling obligations commonly extend to contract review as well as technical architecture assessment.

Vendor Evaluation Questions

The following questions arise across most enterprise AI data residency evaluations and are worth working through for each product under consideration.

Data storage. Where is customer data stored at rest? Where are backups located? Where are logs retained? Under what circumstances, if any, can data leave the specified region?

Model training. Is customer data used to train or improve foundation models? Does this vary by product or deployment tier? Is any training use opt-in only? Is this commitment contractual?

Inference. Where does inference occur? Does this vary by model or deployment tier? Is regional inference available, and at what cost or commercial structure?

Commercial. Are regional deployments priced differently from standard deployments? Are all features and model versions available in regional configurations? What happens to regional commitments if the vendor changes its infrastructure footprint during the contract term?

Governance. How are changes to processing regions communicated to customers? Can the vendor change inference or storage locations without customer approval? What contractual protections govern changes to data handling arrangements over time?

When a vendor confirms support for Australian data residency, a productive follow-up distinguishes between where data is stored, whether customer data is used for model training, and where inference occurs. Each is a separate question. Each may have a different answer.

Good enterprise AI procurement is not about asking more vendor questions. It is about asking better ones.

This article provides general commercial, procurement, and technology commentary only. It does not constitute legal, privacy, regulatory, financial, or professional advice. Enterprise AI platforms, contractual terms, and regulatory requirements evolve over time. Organisations are encouraged to verify current vendor terms directly and to seek appropriate professional advice for their specific circumstances.