auraHub · GDPR and language models

The platform is not the problem.

If auraHub runs on your own infrastructure, your data stays with you. Who decides what happens to prompts and responses is determined by your choice of language model. Three levels for three use cases.

Oliver Range auraNexus.ai · May 28, 2026 9 min read
01 / Problem

When people say GDPR, they mean the platform. The risk sits in the model.

Definition

A GDPR-compliant AI platform refers to an AI application whose platform runs entirely within the company’s infrastructure and whose language model is integrated per use case within a tiered security architecture. The platform governs data storage, permissions, and auditability. The language model determines what happens to prompts and responses once they leave the infrastructure.

In reviews with IT security and data protection officers, one question always comes up first: Does it run on our servers? As soon as the answer is yes, the platform is considered solved and the discussion jumps to licensing. That is understandable, but this is where the real risk is overlooked.

An AI platform consists of two layers. One contains prompts, uploaded documents, user permissions, protocols, knowledge sources, and configuration. You can keep this layer entirely in-house, both technically and contractually. The other layer is the language model that generates the actual answer. This is where it is decided whether content leaves the company, which region it goes to, how long it is stored there, and whether a provider may train on it.

The productive question is therefore not whether it runs on our systems, but which model sees which data at which level. This is exactly the question that can be answered in a structured way with auraHub as a self-hosted AI platform, because platform and model are cleanly separated architecturally. The legal basis is provided by the General Data Protection Regulation, whose obligations apply differently depending on the level.

Modern workplace with the auraHub dashboard on a monitor, with Cologne Cathedral in the background. The platform runs in the company’s own infrastructure at the Frechen site near Cologne.
Workplace with the auraHub dashboard and a view of Cologne Cathedral. The platform runs in-house—you see it every day.
2 Layers
Platform and language model
3 Security levels
per use case
6 GDPR building blocks
at a glance
02 / Architecture

Two layers. A clear separation.

auraHub runs entirely in your own environment—either on-premises on your servers, in your private cloud, or air-gapped, i.e., without an internet connection. This means all platform-side data remains physically with you. At this layer, there is no transfer to a third country, no cloud lock-in, and no shift into external responsibility.

Servers in a data center with modern network infrastructure. Symbolic image for auraHub’s self-hosted platform layer within the company’s own infrastructure.
The physical platform layer: servers in your own infrastructure or private cloud infrastructure on which auraHub runs directly.

What stays with you

Specifically, the following remains in your infrastructure: prompts and inputs, uploaded documents, user and permission management, protocols and logs, connected knowledge sources, and the entire platform configuration. From a data protection perspective, this layer is the non-critical part. You have full control, no third-country transfer at platform level, and no cloud lock-in.

What the language model sees

As soon as a prompt needs to be processed, auraHub calls the configured language model. This is exactly where content may potentially leave the infrastructure—depending on the level: not at all (local model), to an EU region (Azure OpenAI Service with EU data residency), or to the US via Standard Contractual Clauses (OpenAI or Anthropic directly). This level can be set per use case, not as a blanket setting for the entire system. You may be familiar with a comparable architecture from our auraIR setup for Investor Relations, which follows the same logic.

auraHub-Architektur: Plattform und Sprachmodell getrennt Ihre Infrastruktur EBENE PLATTFORM On-premises, Private Cloud oder air-gapped Prompts und Eingaben Hochgeladene Dokumente Nutzer und Rechte Protokolle und Logs Wissensquellen Konfiguration BEWERTUNG Volle Kontrolle, kein Drittlandtransfer auf Plattformebene, kein Cloud-Zwang. API-AUFRUF nach gewählter Stufe Sprachmodell-Ebene EBENE MODELL Drei Sicherheitsstufen, pro Anwendungsfall einstellbar STUFE 1 Lokales Sprachmodell Höchste Sicherheit, Verarbeitung im Haus STUFE 2 Azure OpenAI EU EU-Datenresidenz, AVV über Microsoft STUFE 3 OpenAI oder Anthropic direkt SCC-abgesichert, kein Training auf API-Daten
Platform and language model are separated: what stays with you, and what is determined by the chosen level.
03 / Three levels

Local, EU, or safeguarded. You choose per use case.

The three levels do not differ in auraHub’s feature set, but in what happens outside your organization. The choice is made individually for each use case and can vary between use cases without rebuilding the platform.

Level 1 · Highest security
Local language model
vLLM, LM Studio, or Jan.ai. Data never leaves the infrastructure—no external transfer, no external storage. Suitable for highly sensitive client and patient data.
Level 2 · EU data residency
Azure OpenAI EU
EU region such as Sweden or France. Processing in EU data centers; data remains in the EU. DPA via Microsoft; no training on your data. Suitable for standard business data with an EU requirement.
Level 3 · SCC-safeguarded
OpenAI or Anthropic directly
Providers’ direct API. Processing in the US, safeguarded via Standard Contractual Clauses. No training on API data under OpenAI Enterprise Privacy and Anthropic Commercial Terms. Suitable for non-critical workloads.

Quick terminology clarification

DPA refers to the data processing agreement under Article 28 GDPR, in which the cloud provider commits to processing your data only on your behalf and according to your instructions. SCC stands for Standard Contractual Clauses, a set of contractual clauses issued by the European Commission to safeguard the transfer of personal data to third countries. With a local model, neither applies because no data leaves the organization.

Drei Sicherheitsstufen für das Sprachmodell EBENE SPRACHMODELL Drei Sicherheitsstufen zur Wahl STUFE 1 Höchste Sicherheit Lokales LLM vLLM · LM Studio · Jan.ai DATEN Verlassen die Infrastruktur nie, vollständig im Haus. TRANSFER Kein externer Transfer, keine API nach außen. SPEICHERUNG Keine externe Speicherung. EIGNUNG Hochsensible Mandanten- und Patientendaten STUFE 2 EU-Datenresidenz Azure OpenAI EU EU-Region · Sweden, France DATEN Verarbeitung in EU-Rechen- zentren, Daten bleiben in EU. VERTRAG AVV über Microsoft, kein Training auf Ihren Daten. SPEICHERUNG Bis 30 Tage, nur Missbrauch. Mit ZDR: keine. EIGNUNG Reguläre Geschäftsdaten mit EU-Pflicht STUFE 3 SCC-abgesichert OpenAI / Anthropic Direkt-API der Anbieter DATEN Verarbeitung in den USA, abgesichert über SCC. TRAINING Kein Training auf API-Daten. SPEICHERUNG OpenAI 30 Tage, Anthropic 7 Tage. Mit ZDR: keine. EIGNUNG Unkritische Workloads
Three security levels at the language model: local for highly sensitive data, EU for standard business data, SCC-safeguarded for non-critical workloads.
04 / Retention periods

How long your inputs are stored outside your organization.

Even if the language model is not trained on your data, inputs may be stored by the provider for a limited time, usually for abuse detection. Retention periods vary by provider and have been systematically decreasing in recent years. For particularly sensitive data, there is the premium option Zero Data Retention (ZDR), under which no storage takes place at all.

Option Content retention Training on content
Local LLM No external storage No
Azure OpenAI EU (standard) Up to 30 days, abuse detection only, in EU region No
OpenAI API (standard) Up to 30 days, then deleted No
Anthropic API (standard) 7 days (reduced from 30 since September 2025) No
With Zero Data Retention No storage after the response, in-memory only No

Premium option: Zero Data Retention

For particularly sensitive processing—such as finance and investor-relations data (as in our auraIR setup for listed companies)—Anthropic offers Zero Data Retention as part of an enterprise agreement. Inputs and outputs are then no longer stored after the response, are not written to abuse logs, and are not subjected to human review. Processing takes place purely in memory.

This level must be requested separately and is tied to a qualifying contract. Microsoft (Azure) and OpenAI offer ZDR in a similar way for enterprise or Microsoft customer agreements. A practical note: with Anthropic, the results of security classifiers are still retained even under ZDR—this should be considered in the risk model.

05 / GDPR building blocks

Six building blocks. A consistent review.

The following six building blocks cover the GDPR-relevant obligations when using auraHub and can be answered consistently in any audit question.

Article 28

Data processing agreement

Available with every cloud provider and must be concluded before use. Under Art. 28 GDPR, it governs that the provider processes your data only on your behalf.

Third country

Transfer safeguarded

For US providers via Standard Contractual Clauses. Not applicable at all with a local or EU model, because no third-country transfer takes place.

Training

No training on your data

Contractually assured for API and business plans. An important distinction from consumer apps, where inputs may be used for model training by default.

Retention

Data minimization

Short retention periods on the provider side, with optional Zero Data Retention for regulated data. On the platform side, you control retention through your own configuration.

Open Source

Auditability

The platform’s source code is accessible; data flows are auditable. Important for internal audit, external audits, and for trust with customers, investors, and regulators.

Access

Granular permissions

Role and group permissions, single sign-on via OpenID Connect (an open standard for centralized login), encrypted secrets. Who may see what is configurable down to the source level.

06 / Recommendation

One platform. Tiered model selection per use case.

The decisive lever for assessment in your organization is tiered model selection per use case. Highly sensitive cases run on a local model with no data leaving the organization. Standard business data runs via Azure in the EU. Non-critical workloads run via the direct API.

Where cloud models touch particularly sensitive data, Zero Data Retention safeguards processing without any storage. This allows each use case to be set to the appropriate protection level individually, without having to move the entire operation to the maximum security level. This reduces costs where they do not buy additional protection, and concentrates effort where it has an effect.

What that means for your assessment

Three questions are enough to arrive at the right level per use case. First: which data category is processed—for example personal data under Article 9 GDPR, trade secrets, or non-critical data? Second: which regulatory requirement applies—for example professional confidentiality, an industry-specific requirement, or GDPR in general? Third: how frequently does the use case run, and does the operating model justify a local model, or is the cloud option sufficient? If you need an initial assessment of your cases, we support this as part of our AI consulting.

07 / Workflow

How to arrive at the right level selection.

  • 01

    Collect use cases

    List between three and ten concrete use cases in which auraHub is to be used. Examples: summarizing press coverage, reviewing contract drafts, categorizing customer inquiries, pre-structuring patient letters.

  • 02

    Assign a data category

    For each use case: does it involve particularly sensitive data under Article 9 GDPR, standard personal data, trade secrets, or non-critical public data? This assignment directly determines the level selection.

  • 03

    Select the level

    Level 1 local for highly sensitive data and professional confidentiality. Level 2 Azure EU for standard business data with an EU requirement. Level 3 direct API for non-critical workloads. Enable ZDR as soon as sensitive data touches cloud models.

  • 04

    Conclude contracts

    Conclude a DPA with all relevant providers (Microsoft, OpenAI, Anthropic). Request ZDR explicitly where applicable. Review Standard Contractual Clauses where third-country transfer takes place.

  • 05

    Configuration and audit

    Configure auraHub per use case to the selected level. Assign roles and permissions according to the use case. Document data flows once in full, then review them regularly.

$ Request consulting

auraHub in-house.

We support you from the initial assessment of your use cases through to productive go-live—together with your data protection and IT stakeholders, in your infrastructure, with the right level selection per use case.

Schedule an appointment →
FAQ / Reference

Frequently asked questions

What does a self-hosted AI platform mean?

A self-hosted AI platform runs in your own infrastructure—either on-premises on your servers, in your private cloud, or air-gapped without an internet connection. All platform-side data such as prompts, documents, user permissions, and logs remains physically in-house. auraHub is designed as a self-hosted AI platform.

Which security levels does auraHub offer at the language model?

Three levels. Level 1 uses a local language model with no external data outflow. Level 2 uses Azure OpenAI in an EU region with a data processing agreement via Microsoft. Level 3 uses the direct APIs of OpenAI or Anthropic, safeguarded via Standard Contractual Clauses. The level is chosen per use case.

What is Zero Data Retention?

Zero Data Retention is a premium option for regulated data. Inputs and outputs are no longer stored after the response, are not written to abuse logs, and are not subjected to human review. Processing takes place purely in memory. Anthropic, Microsoft Azure, and OpenAI offer ZDR as part of enterprise contracts.

Is training done on our data?

No For API and business plans, it is contractually assured that no training takes place on the transmitted data. This is a key difference from consumer apps such as ChatGPT Free, where data may be used for model training by default.

How long are inputs stored outside the organization?

Not at all with a local LLM. Azure OpenAI EU stores data by default for up to 30 days for abuse detection in an EU region. The OpenAI API stores data by default for up to 30 days; Anthropic for 7 days (reduced from 30 since September 2025). With Zero Data Retention, any storage after the response is eliminated.

Which GDPR building blocks are covered?

Six building blocks: data processing agreement under Article 28; third-country transfer safeguarded via Standard Contractual Clauses (not applicable for local or EU); no training on your data; data minimization and short retention, with optional Zero Data Retention; open source and auditability of the platform source code; granular access control with roles, group permissions, and SSO via OpenID Connect.

Which level fits which use case?

Highly sensitive client or patient data runs on a local language model. Standard business data with an EU requirement runs via Azure OpenAI EU. Non-critical workloads run via the direct API of OpenAI or Anthropic. Where cloud models touch particularly sensitive data, Zero Data Retention safeguards processing without any storage.

OR
@oliverrange
Oliver Range
Founder auraNexus.ai · AI Manager (TÜV)

Founder of several digital companies, including Die Medialysten (exit to Linkfluence). Consultant for GDPR-compliant AI platforms in the DACH mid-market. Develops AI applications with auraNexus.ai for communications, tax advisory, law firms, healthcare, and manufacturing.

// read next