Balancing compliance and tenant isolation for hosted inference platforms

Hosted inference platform compliance is not solved by selecting a provider with strong certifications. Teams must decide what each tenant can share, where data can be processed, how evidence is retained, and when a tenant needs a dedicated stack rather than a logically isolated slice of a shared service.
The practical goal is to match the isolation model to the tenant’s risk and contractual obligations without making every deployment unnecessarily expensive or difficult to operate. That requires treating compliance controls, data lifecycle controls, and GPU or network isolation as related but separate design decisions.
Define hosted inference platform compliance as a layered responsibility
In a hosted inference service, a customer prompt may pass through an application tier, an identity layer, an API gateway, safety systems, logging pipelines, model-serving infrastructure, and GPU capacity. A compliance program that considers only the model API misses the controls that determine whether one tenant can access another tenant’s data, metadata, capacity, or audit trail.
Providers can supply important infrastructure assurances, but the platform operator still owns much of the system design. Google Cloud’s guidance for multi-tenant AI service providers, for example, places responsibility on the provider for logical and physical tenant isolation, per-tenant rate limiting, and encryption with distinct customer-managed keys. That is a useful reminder that using managed AI services does not remove architectural accountability.
Direct answer: Balance compliance and tenant isolation by classifying tenants by data sensitivity, residency, audit, performance, and network requirements; applying a shared baseline to all tenants; and moving only higher-risk tenants to dedicated keys, compute, networks, or full-stack silos when the shared controls cannot satisfy their requirements.
Three layers should be assessed independently:
Data governance:
retention, deletion, training use, residency, export, legal holds, and access to prompts, responses, files, and metadata.
Tenant isolation:
identity boundaries, authorization, tenant-scoped storage, encryption keys, network segmentation, workload scheduling, and protection against cross-tenant access.
Operational evidence:
logs, monitoring, alerting, investigation workflows, and records that demonstrate whether controls operated as intended.
These layers overlap, but none proves the others. Zero retention at an upstream API does not automatically isolate a platform’s own application logs. A dedicated Kubernetes namespace does not by itself meet a data residency commitment. And a security certification does not decide whether a regulated tenant needs its own inference deployment.
Start with tenant requirements, not a universal isolation model
Isolation should be a deliberate product and risk decision. Microsoft’s Azure OpenAI guidance explicitly advises multitenant system designers to determine the degree of tenant isolation required. It recognizes both single-tenant deployments for network-isolated requirements and multitenant designs with shared application tiers.
That choice is rarely binary. A platform can share a control plane while separating data planes, or it can retain a common application layer while assigning selected tenants dedicated storage, keys, node pools, and private network paths. The right endpoint depends on the specific promise made to the tenant.
Ask questions that produce an enforceable tier
What data enters inference?
Identify prompts, documents, attachments, retrieved content, generated outputs, identifiers, and operational metadata. Do not assume that only the prompt is sensitive.
What location commitment applies?
A residency requirement should identify both stored data and the region in which inference occurs. OpenAI states that eligible customers can choose data residency and in-region GPU inference in the United States or Europe on supported endpoints, but platform teams should confirm endpoint eligibility and their own downstream processing paths.
What evidence must the tenant receive?
Determine whether the tenant needs access logs, administrator activity, policy events, data exports, or a specific retention period for those records.
Is network isolation contractual or merely preferred?
If a tenant requires isolated network access, a shared application tier may be insufficient even when authorization is sound.
What performance boundary matters?
A tenant may need predictable capacity, rate limits, dedicated GPUs, or protection from noisy neighbors. These are operational requirements that can also influence compliance outcomes.
Convert answers into a small number of service tiers. For example, a standard tier might use a shared application and inference environment with strong logical controls. A regulated tier could add tenant-specific customer-managed encryption keys, dedicated node pools, and stricter export workflows. A highly restricted tier may need a separate account or subscription, private networking, and a full-stack silo.
The tier model must be specific enough to guide engineering. Labels such as “high compliance” or “enterprise isolation” are not controls. Document which resources are shared, which are tenant-scoped, who can administer them, how logs are separated, and which changes require approval.
Build a shared baseline for multi-tenant AI isolation
A shared environment can be appropriate when its boundaries are designed and tested rather than assumed. The baseline should make accidental cross-tenant access difficult at every layer, including the layers that are not usually described as AI infrastructure.
Make tenant identity a first-class control
Every request should carry a trusted tenant context established by authentication and authorization, not by a user-provided field in the request. The platform should enforce that context at the gateway, application service, retrieval layer, storage layer, logging pipeline, and administration interfaces.
This matters especially for retrieval-augmented inference. If document retrieval filters are optional or are applied after a broad search, a model request can be perfectly isolated at the API level while the retrieved context is not. Tenant-scoped indexes, storage paths, access policies, and test cases are all required.
Segment workloads, networks, and secrets
Google Cloud recommends dedicated node pools, namespaces, network policies, and customer-managed encryption keys for multi-tenant AI services. These mechanisms solve different problems, and using one should not be mistaken for using all of them.
Namespaces
provide an organizational boundary for workloads and policies, but need accompanying authorization controls.
Dedicated node pools
reduce co-location for selected workloads and can support stronger scheduling separation.
Network policies
constrain which services and workloads can communicate, reducing unnecessary lateral paths.
Distinct customer-managed keys
provide a tenant-specific cryptographic boundary and a clearer revocation and access-control model.
Per-tenant rate limiting
protects service availability and helps contain misuse or runaway workloads.
Secrets deserve the same rigor. API credentials, retrieval-store credentials, key-encryption-key permissions, and connector tokens should be scoped to the smallest practical tenant boundary. A central service identity with unrestricted access to every tenant’s data can undermine otherwise careful segmentation.
Understand the GPU boundary
GPU infrastructure can introduce a gap between logical tenancy and physical sharing. NVIDIA notes that, in self-managed deployments, multiple tenants can share the same physical GPU unless node-level isolation is configured. If the tenant requirement is explicit physical or node-level separation, namespace or application isolation alone does not establish it.
This does not mean every tenant needs an exclusive GPU. It means the platform must describe the actual boundary honestly and select node-level or dedicated infrastructure where the requirement calls for it. Capacity efficiency is a valid reason to share hardware; it is not evidence that shared hardware meets an isolation clause.
Choose when to use logical isolation, dedicated components, or a full-stack silo
The principal trade-off is straightforward: sharing improves utilization and simplifies operations, while dedicated resources can deliver stronger boundaries and clearer evidence. The implementation details are not straightforward, because an architecture can be shared in one layer and isolated in another.
AWS guidance on multi-tenant agentic AI recognizes this range. It describes compliance adaptation through tenant-specific configuration in a multi-tenant service, while also advising that tenants with compliance and performance requirements may be better served by a full-stack silo model.
Logical isolation is often the default
Logical isolation is a reasonable baseline for tenants that accept shared infrastructure and whose requirements can be satisfied through tenant-aware authorization, tenant-scoped data stores, encryption controls, rate limits, and monitoring. It can support efficient centralized patching, standard safety controls, and consistent releases.
Its limit is that it relies on correct configuration and enforcement across every shared layer. A platform cannot represent this as physical separation, private networking, or dedicated compute if those are not actually provided.
Dedicated components fit targeted requirements
Some tenants need more than shared defaults but not a separate platform. Dedicated node pools, tenant-specific keys, isolated retrieval indexes, private connectivity, or a dedicated inference endpoint can address a concrete requirement while retaining a shared control plane and operating model.
This model is often effective when a requirement is narrow: for example, a separate encryption key, a particular regional execution path, or an allocated capacity boundary. It also avoids converting every regulatory request into a separate application deployment.
Full-stack silos are justified by the complete requirement
A full-stack silo separates the tenant across the application, data, network, and inference layers. It is appropriate when contractual obligations, network-isolated requirements, high-assurance performance needs, or risk acceptance rules cannot be met credibly in a shared service.
The cost is real: more environments to patch, more policy deployments to validate, more monitoring configurations, and a greater chance that one tenant’s environment drifts from the approved baseline. A silo should therefore have automation, repeatable infrastructure definitions, and a documented lifecycle. Isolation that cannot be operated consistently is not a durable compliance strategy.
Align data retention, training controls, and residency with your platform flow
Data controls should be evaluated from the tenant application to every provider in the path. A platform can minimize provider retention while retaining full prompts in its own debugging store, or it can configure regional inference while exporting logs to another region. The effective policy is the combination of all those choices.
OpenAI says that it does not train on organization data by default and that enterprise and API customers retain ownership and control of business data. It also offers enhanced retention controls and, for eligible customers on supported endpoints, data residency and in-region GPU inference in the United States or Europe. These capabilities can be important inputs to a hosted inference platform design, but they must be paired with the operator’s own storage, observability, and support processes.
Retention settings change operational obligations
Under OpenAI’s February 2024 data processing addendum, API customer data is retained for a maximum of 30 days and then deleted, unless legal retention is required. OpenAI also describes Zero Data Retention (ZDR) as a mode in which prompts and responses are not retained after processing for eligible API customers.
OpenAI now pairs ZDR with Private Safety Processing for eligible API customers. According to OpenAI, that flow enables offline automated safety review without retaining customer prompts or responses. This is useful for teams that need lower retention while still accounting for safety operations, but it should not be read as a replacement for the platform’s own content governance and escalation process.
In particular, OpenAI’s API data-controls guidance says customers using Modified Abuse Monitoring or ZDR are responsible for user compliance with OpenAI policy and applicable moderation and reporting laws. Lower provider-side retention may reduce available investigative material, so the customer must decide what event data, redacted content, or policy decisions it needs to retain within its own lawful and documented program.
Do not treat data sharing as compatible with every privacy mode
OpenAI states that accounts with ZDR enabled cannot opt in to data-sharing mechanisms for feedback, evaluations, or fine-tuning data. That is a clear product-design trade-off. If a platform has a workflow that depends on sharing selected interactions for model improvement, it needs a separate, appropriately governed path rather than assuming it can coexist with a ZDR configuration.
Design the decision explicitly: which tenants use no retained provider data, which can contribute de-identified or approved evaluation material, who authorizes that choice, and how is it recorded? Consent language and product controls should match the technical setting.
Design compliance logging as evidence, not an afterthought
Compliance evidence has its own retention and access requirements. It should be isolated by tenant, protected from alteration, available to authorized investigators, and kept only as long as the applicable policy requires. It must also be useful enough to answer what happened without indiscriminately duplicating sensitive prompt content.
OpenAI’s Compliance Logs Platform retains data for 30 days and advises customers that need a longer history to continuously download logs. This creates a concrete operational requirement: a team that depends on those records for a longer investigation, retention schedule, or audit period needs an export pipeline and a destination governed under its own retention and residency policies.
Capture the decision trail
A useful tenant-aware audit record can include the tenant identifier, authenticated actor or workload identity, endpoint or model configuration, policy decision, timestamp, request correlation identifier, administrative action, and export event. Whether prompts or outputs should be included depends on the data classification and purpose. In many cases, storing a content hash, classification result, or policy outcome may be more appropriate than copying sensitive content into a broad observability system.
Separate operational logs from tenant-accessible audit records. Support engineers may need service health information, while a tenant may need evidence of access to its own data. The access model, filtering logic, and export process should preserve that distinction.
Integrate monitoring with the systems investigators use
OpenAI lists compliance-log integrations with partners including Concentric AI, CrowdStrike, Cyberhaven, Enkrypt AI, and Forcepoint. Such integrations can help teams route records into established DLP, eDiscovery, or SIEM workflows. They do not eliminate the need to define event ownership, retention, access rights, and response procedures.
Integration design should also account for product change. OpenAI states that its older stateful compliance route was deprecated and removed on June 5, 2026. Hosted inference platforms should avoid hard-coding assumptions about compliance endpoints and should maintain an inventory of integrations, versions, authentication methods, and fallback procedures. Change management for logging is a control, not merely an implementation task.
Map provider assurances to your actual compliance scope
Provider certifications and contractual support matter, but they cover defined services and boundaries. They should inform due diligence rather than end it. The key question is whether the provider control, your platform control, and the tenant’s required outcome line up.
OpenAI says the infrastructure supporting its API and enterprise services has been evaluated by an independent third-party auditor and aligns with ISO/IEC 27001:2022 and ISO/IEC 27701:2019 for security and confidentiality. It also says that a Business Associate Agreement is available for ChatGPT for Healthcare and API healthcare customers to support HIPAA compliance.
Those facts may help a healthcare or regulated-data assessment, but they do not make an entire hosted inference application compliant by themselves. The application’s identity architecture, minimum-necessary data handling, retrieval connectors, support access, logging, incident response, and customer contracts remain in scope for the platform operator.
Verify which service, endpoint, and region the assurance covers.
Confirm whether the required data type and intended workflow are permitted.
Document the shared-responsibility boundary in the tenant-facing service description.
Map each tenant obligation to a technical control, operating procedure, and evidence source.
Review changes to provider features, retention modes, and compliance integrations before relying on them in a commitment.
This mapping prevents a common failure mode: a sales or architecture statement promises an outcome such as “isolated,” “regional,” or “audit-ready,” while the engineering implementation delivers only a related feature. Precision is more valuable than broad assurances.
Operate the isolation model through testing and change control
Tenant isolation is not established at launch. It can be weakened by a new retrieval connector, a permissive support role, a logging schema change, a model-routing update, or a capacity workaround. The operating model must continuously verify the boundaries that the service claims.
Test negative authorization paths.
Attempt cross-tenant reads, writes, searches, log queries, export requests, and administrative actions. Test both user identities and service identities.
Validate network and workload placement.
Confirm that network policies, dedicated node pools, and scheduling rules enforce the intended design, especially for tenants that pay for stronger separation.
Trace representative requests.
Follow a request through ingress, inference, retrieval, safety processing, logs, exports, and deletion workflows. The trace should identify every storage and processing point.
Review data lifecycle events.
Verify retention timers, deletion behavior, legal-hold handling where applicable, and log-export continuity. Provider retention settings do not validate the platform’s copies.
Control configuration changes.
Treat changes to tenancy mapping, model endpoints, regions, keys, retention modes, and log integrations as reviewed changes with rollback plans.
Automated policy checks can help keep a tier model honest. For example, a deployment tagged as a dedicated tier should fail validation if it uses a shared key, lacks the expected node pool, or sends logs to an unapproved destination. The exact tooling will vary, but the principle is stable: stated isolation level should be machine-verifiable wherever possible.
Finally, expose the right information to tenants. A concise service description should state whether infrastructure is shared, which controls are tenant-specific, the applicable retention model, how long compliance logs remain available at the source, and what exports or customer actions are needed. Clear documentation reduces procurement friction and prevents security teams from inferring guarantees that the platform does not make.
Use a decision framework that preserves both assurance and efficiency
The most effective hosted inference platform compliance program is not the one with the maximum possible separation everywhere. It is the one that applies a defensible level of isolation to each tenant, keeps data controls consistent with the actual flow, and produces evidence that the design continues to work.
Start with a shared baseline built on tenant-aware authorization, segmented workloads and networks, scoped secrets, rate limits, encryption, and auditable operations. Add dedicated keys, node pools, private paths, or regional execution when a defined requirement needs them. Choose a full-stack silo when the complete combination of compliance, network, and performance needs cannot be met credibly in the shared model.
Keep retention, logging, and safety design in the same conversation as infrastructure isolation. Provider capabilities such as no-default-training treatment for organization data, configurable retention, ZDR, private safety processing, residency options, and compliance-log integrations are valuable only when they are connected to the platform’s own controls and the tenant’s documented obligations. That alignment is what turns hosted inference from a collection of features into a trustworthy service.