Insights · Privacy

Twelve privacy questions to answer before you ship an AI feature

Most AI privacy failures are not exotic. They are ordinary gaps that nobody wrote down: a log with no retention period, an index that ignored the permissions on the file share, a consumer account that reached production. Here are the twelve questions we ask, and what a good answer looks like.

Privacy engineeringPre-launch checklistNot legal advice
Quick answer

Before an AI feature ships, you should be able to answer twelve questions in writing: what personal data reaches the model, your lawful basis, whether a DPIA is triggered, prompt logging and retention, provider training terms, processing location and transfer mechanism, corpus contents, retrieval permissions, subject access, erasure, user disclosure and accountability. This is engineering guidance, not legal advice.

Context

Why a checklist, and why this one

Green Arrow Consultancy spent a decade running privacy, consent and analytics programmes for global consumer brands before it built AI systems for anyone. That order matters. The questions below are not derived from a regulation summary, they are the questions that actually caught things in review on real deployments, arranged so that the cheap ones come first.

This is not legal advice. It is engineering and governance guidance from a firm that builds these systems and then has to operate them. Data protection law is jurisdiction-specific and fact-specific, and the right answer for a UK employer processing staff data is not the right answer for a United States retailer with California customers. Use this to interrogate your own design and to brief your own advisers properly. Our own work in this area is described under privacy consulting.

One structural point before the list. A language model application is not a single processing operation, it is a chain: collection at the interface, transmission to a provider, inference, logging, retrieval against a corpus you assembled, and output storage. Each link has its own lawful basis question, its own retention question and its own location question. Teams that answer these at the level of “the AI feature” miss the link that matters, which is nearly always the logging.

The other structural point is that scale changes the answer. A prototype used by nine people in one department is a different risk proposition from the same code exposed to two hundred thousand customers. Answer these questions against the deployment you are about to ship, not the one you tested.

The list

The twelve questions

Each question has a reason it bites and a description of what a satisfactory answer looks like on paper.

  1. 01

    What personal data actually reaches the model?

    Why it bites. The answer is nearly always larger than the project team believes. A customer types a question, the application quietly attaches account context, the retrieval layer adds three passages from internal documents, and the system prompt carries an example that was copied from real correspondence. Four sources, one request, and typically only the first is on the diagram. A good answer. Somebody has traced one real request end to end and written down every field that left your estate, including the ones added automatically. That trace goes into the record of processing. Anything on it that is not needed for the answer gets removed before launch, because the cheapest data protection control is not sending the data.

  2. 02

    What is your lawful basis, and can you evidence it?

    Why it bites. Consent collected for a website is not consent for a language model, and “we have a privacy policy” is not a lawful basis. Where the data is special category, health, biometrics, trade union membership and the rest, you need a separate condition on top and the bar is considerably higher. A good answer. A named basis for each processing purpose, decided before launch rather than reverse-engineered afterwards. If the basis is legitimate interests, the balancing assessment is written and someone senior has signed it. If it is consent, the consent is specific to this use, freely given, and as easy to withdraw as it was to give.

  3. 03

    Does this trigger a data protection impact assessment?

    Why it bites. Regulators' published criteria for high-risk processing include innovative use of new technology, systematic and extensive evaluation of individuals, large-scale processing and the combination of datasets. A customer-facing assistant that reads account history ticks several at once. Teams skip the assessment because it feels like paperwork, then find that the paperwork was the only place the permission problem would have surfaced. A good answer. A completed DPIA signed off by whoever holds the data protection role, covering the whole chain rather than the interface, with residual risks stated plainly. Where you decide one is not required, the reasoning is written down and dated. The assessment is a design tool, and it is most useful when it is done early enough to change something.

  4. 04

    Are prompts and outputs logged, and for how long?

    Why it bites. This is the single most common gap we find. Logging is switched on during development for debugging, nobody switches it off, and eighteen months later there is an unindexed store containing everything every customer ever typed, sitting on a retention policy of forever, accessible to the whole engineering team. It is the highest-risk store in the system and usually the least governed. A good answer. A stated purpose for the logs, a retention period tied to that purpose and measured in weeks or months, automated deletion that actually runs, access control that a security reviewer would accept, and a named owner. Where you need history for quality evaluation, keep a redacted version on the longer clock and delete the raw one early.

  5. 05

    Does the provider train on your data, and how did you prove it?

    Why it bites. The default terms on a consumer tier are not the terms on an enterprise tier, and a proof of concept built on somebody's personal account has a way of becoming production. Assurance passed along in a meeting is not evidence. A good answer. The specific contractual clause, quoted, from the tier and account you are actually deployed on, with the date it was checked and a diary note to re-check at renewal. Alongside it, confirmation of what the provider retains for abuse monitoring and for how long, because that is a separate question from training and it is the one people forget to ask.

  6. 06

    Where does processing happen, and what covers the transfer?

    Why it bites. Inference, logging, abuse review and support access can each sit in a different country, and the marketing page for a region-pinned deployment does not always cover all four. Any hop out of the UK or the EEA needs a valid mechanism, whether that is adequacy, Standard Contractual Clauses, or the UK addendum or international data transfer agreement. A good answer. A diagram naming the country for every hop, the mechanism covering each transfer that needs one, and a transfer risk assessment where the destination warrants it. Region-pinned deployment where the workload allows it, because the simplest way to answer a transfer question is not to make the transfer.

  7. 07

    What is in the retrieval corpus, and should it be?

    Why it bites. An index built by pointing a crawler at a file share inherits every mistake in that file share: the spreadsheet of candidate interview notes, the exported customer list somebody saved in 2019, the folder of scanned occupational health referrals. None of that was a problem while it sat unfindable in a directory nobody opened. It becomes a problem the moment a search system makes it answerable. A good answer. A documented source inventory built by deliberate selection rather than by crawl. Special category and confidential material either excluded, or admitted on purpose with a recorded reason and a tighter access model. An owner for the corpus who reviews additions.

  8. 08

    Is the permission model of the source systems enforced in retrieval?

    Why it bites. An assistant that can read everything will eventually tell someone something they should not have seen, and it will do it in a form that is easy to screenshot. Filtering after generation is not a control, because by then the model has already read the material and may paraphrase it. A good answer. Entitlements applied to the candidate set before generation, so the model never sees a passage the asker is not allowed to read. A re-sync when source permissions change, because a stale access control list in an index is a leak on a delay. Testing done with a genuinely low-privilege account rather than an administrator, which is the mistake we find most often.

  9. 09

    How do you answer a subject access request?

    Why it bites. The right of access covers personal data wherever it sits, and prompt logs are not exempt for being awkward. Logs are usually indexed by session identifier and timestamp, not by person, so a request that takes minutes against the customer database can take days against the log store, inside a statutory deadline that does not move. A good answer. A written procedure that names every store in scope, including logs, caches and evaluation datasets. A decision on how a person is identified across those stores. A query that has been built and run at least once as a rehearsal, with the export format agreed and third-party data redaction thought through in advance.

  10. 10

    What happens on erasure?

    Why it bites. Deleting the source document does not remove it from the retrieval index, the embedding cache, the derived summary, or last week's prompt logs. Each of those is a separate deletion, and in most architectures they are owned by different people. Teams discover the gap when the system cheerfully answers a question using a record that was erased a fortnight earlier. A good answer. An erasure path that reaches the source system, the index, every cache, any derived artefact and the logs, with the order written down and the whole sequence tested once on a real record. A stated maximum lag between erasure at source and erasure everywhere else, so the exposure window is known rather than assumed.

  11. 11

    What do you tell users about the AI?

    Why it bites. People behave differently when they know a machine is answering, and several regimes now require you to say so. Burying it in a policy nobody opens satisfies neither the regulation nor the user. A good answer. Disclosure at the point of interaction in plain language, a short note on what happens to what they type and how long it is kept, a visible route to a human that does not require persistence, and honesty about limitations. If the system can be wrong, and it can, say so where people will read it rather than where lawyers will find it.

  12. 12

    Who is accountable when it is wrong?

    Why it bites. A system with no named owner gets defended rather than fixed. The first bad answer that reaches a senior person becomes an argument about whether the technology was a good idea, instead of a ticket. A good answer. A named accountable individual, not a committee. An escalation route that a front-line agent can actually use at half past four on a Friday. An incident log with categories, so patterns are visible. A published contact for complaints and rights requests. And a review cadence, because the system that was signed off in March is not the system running in September.

Output

The artefacts each question produces

If a question does not produce an artefact, it was not answered, it was discussed.

QuestionArtefact it should produceWho owns it
What data reaches the modelA request trace and an updated record of processingSolution architect
Lawful basisA basis per purpose, plus a legitimate interests assessment if relevantData protection lead
DPIA triggerA completed DPIA, or a dated reasoned decision that one is not requiredData protection lead
Prompt and output loggingA retention schedule with automated deletion and an access listEngineering owner
Provider training termsThe quoted clause, the tier it applies to, and a renewal diary noteCommercial or procurement
Processing locationA country-by-country flow diagram and the transfer mechanism for each hopData protection lead
Retrieval corpus contentsA source inventory with inclusion decisions recordedContent or knowledge owner
Retrieval permissionsA permission mapping design and a low-privilege test resultEngineering owner
Subject accessA written and rehearsed procedure covering the log storePrivacy operations
ErasureA tested deletion sequence with a stated maximum propagation lagEngineering owner
User disclosureInterface copy, a route to a human, and a limitations statementProduct owner
AccountabilityA named owner, an escalation route and an incident logAccountable executive

This is the pre-launch pack we assemble on client engagements. It is deliberately short enough to be finished.

Deliverable

The evidence pack we hand over

Eight documents. On a well-run project this is a fortnight of work, and it is the difference between a defensible system and a hopeful one.

Vocabulary

Terms worth being precise about

Precision here saves a great deal of argument later.

Record of processing
The register of what personal data you process, why, on what basis, where it goes and how long you keep it. An AI feature usually needs its own entry rather than an amendment to an existing one, because the chain introduces new recipients and a new store.
Data protection impact assessment
A structured assessment of the risk a processing operation poses to people, required where that risk is likely to be high. It is most valuable when it is done early enough that its findings can still change the design.
Retrieval corpus
The set of documents a retrieval system can search to answer a question. Its contents are a privacy decision as much as an engineering one, because indexing makes findable what was previously merely present.
Prompt log
The stored record of what users typed and what the system answered. Personal data in most consumer and employee-facing deployments, and the store most likely to have no retention period attached.
Transfer mechanism
The legal instrument that permits personal data to leave the UK or EEA: adequacy, Standard Contractual Clauses, the UK addendum or international data transfer agreement, or a narrow derogation. Needed for every hop, not just the obvious one.
Permission-aware retrieval
Applying the asker's entitlements to the candidate documents before generation, so the model cannot answer from material that person is not allowed to read. Filtering after the answer is written is not equivalent.
Governance

Who signs it off

The twelve questions produce artefacts. Artefacts need an owner, and the owner should be the person accountable for the business outcome, advised rather than replaced by whoever holds the data protection role. Privacy functions that own AI sign-off outright become a queue, and delivery teams route around queues, usually by describing the project as something else.

What works better is a short mandatory artefact list, a named reviewer, and a published service level for the review. Make the compliant route the fast route. If completing the pack takes a fortnight and going around it takes a quarter of an argument, people will complete the pack.

Above the individual feature sits the management system question. If you are shipping more than one AI feature, the twelve questions should be a template rather than a bespoke exercise each time, embedded in whatever framework you have adopted. ISO/IEC 42001 and the NIST AI Risk Management Framework both give you a structure for that, and the EU AI Act will supply obligations for some deployments regardless of which framework you prefer. We cover that layer under AI governance.

Two questions on this list have a security dimension that goes beyond privacy. What is in the retrieval corpus and what the system is permitted to do with it are also the questions that determine how much damage a prompt injection can cause, which is covered in our article on prompt injection. And the decision about whether knowledge lives in an index or in fine-tuned weights has direct consequences for erasure, which is one of the arguments in retrieval or fine-tuning.

Once more, plainly. This is not legal advice. It is a design checklist written by engineers and privacy practitioners who build these systems. Take it to your own advisers and let them tell you what your jurisdiction requires.

Questions

Frequently asked questions

More definitions in the glossary. More articles in insights.

Is a DPIA always required for an AI feature?

Not always, but the threshold is lower than most teams assume. Under UK and EU data protection law an assessment is required where processing is likely to result in a high risk to people, and the regulators' published criteria include innovative use of new technology, systematic and extensive evaluation of individuals, processing at scale, and matching or combining datasets. A customer-facing assistant reading account data ticks several of those at once. Where you conclude an assessment is not needed, write down the reasoning and keep it. An undocumented decision is indistinguishable from no decision at all.

Do the big model providers train on business data?

On enterprise and standard API tiers, the major providers commit not to train on customer inputs and outputs by default. On free and consumer tiers the position is often different, and it changes. The point is not to trust the summary, including this one. Find the clause in the terms that apply to the specific tier and account you are deployed on, quote it into your record of processing with the date, and re-check it at contract renewal. We have seen teams inherit a consumer account from a proof of concept and ship it to production without anyone re-reading the terms.

Are prompts and outputs personal data?

Frequently, yes, and they are easy to overlook because they do not live in a database schema. A support assistant's prompt log contains whatever the customer typed, which may include names, account numbers, health information and a great deal that nobody asked for. Retrieved passages placed into the context window are also personal data if they identify someone. Treat the prompt log as a personal data store with a retention period, an access control list, an owner and a deletion routine, because that is what it is.

How long should we keep prompt logs?

As long as you have a stated purpose for them, and no longer. Common purposes are abuse investigation, quality evaluation and debugging, and each of those has a natural horizon measured in weeks or a small number of months rather than years. Where you need conversation history for evaluation, consider keeping a redacted or aggregated version on a longer clock and deleting the raw log on a short one. Whatever you choose, automate the deletion, because a retention policy nobody enforces is a liability with a paper trail.

Can a person ask for their data out of an AI system?

Yes. The right of access applies to personal data wherever you hold it, and prompt logs are not exempt because they are inconvenient to search. The practical problem is that logs are usually indexed by session or timestamp rather than by data subject, so a request that would take minutes against the customer database takes days against the log store. Solve that before you receive the request: decide how a person is identified in the logs, build the query, and test the whole procedure end to end once.

What about erasure when the model has already seen the data?

Distinguish the model from the system. On the configurations we deploy, the provider does not retain inputs for training, so there is nothing to erase from the model itself. What has to be erased is everything in your estate: the source record, the retrieval index entry, any embedding cache, any derived summary, and the prompt logs. Those are ordinary deletion problems, they are just spread across more places than teams expect. Map them before launch and rehearse the sequence.

Do we have to tell people they are talking to an AI?

In most consumer contexts you should, and increasingly you must. The EU AI Act places transparency obligations on systems that interact with people, and fairness and transparency principles in data protection law point the same way regardless of jurisdiction. Beyond the legal position it is simply better practice. Disclosure at the point of interaction, a short plain-language note about what happens to what they type, and a visible route to a human. Hiding it is a reputational risk with no upside.

Who should own AI privacy sign-off internally?

The accountable owner should be the person who owns the business outcome, advised by whoever holds the data protection role. Privacy teams that own AI sign-off outright tend to become a queue, and delivery teams route around queues. What works better is a short, mandatory set of artefacts, a named reviewer, and a service level for the review. Make the compliant path the fast path and you will not have to police it.

Is this article legal advice?

No. This is engineering and governance guidance from a consultancy that builds and runs these systems, written to help you ask better questions of your own advisers. Green Arrow Consultancy is a member of the International Association of Privacy Professionals and holds ICO registration ZA822868, but nothing here is a substitute for advice from a qualified lawyer on your specific facts and jurisdictions.

Written and reviewed by the Green Arrow Consultancy team, led by Darren Tyler, founder and chief executive.

Green Arrow Consultancy Ltd, Cardiff, Wales. Company number 12491770. ICO registration ZA822868. Member of the International Association of Privacy Professionals. Last reviewed .

Run the twelve questions against your system

Send us the design and we will work through this list with your team, tell you which answers are missing, and write the artefacts that are not there yet.