Generative AI services such as ChatGPT, Gemini and Claude may produce information about a person from learned patterns, information supplied in a conversation or sources retrieved while answering. Public information may have been included in training data, but a particular page being online does not prove that a specific model used it.
A language model is not simply a searchable directory containing a verified profile of everyone. Its answers can combine accurate information with mistakes, including confusion between people with the same name. You can exercise applicable privacy rights and reduce unnecessary public exposure at its source. This guide explains the channels, risks and practical steps in a France/EU context.
At a glance
- AI systems can reproduce information from training or retrieve information from current sources, depending on the service.
- Public websites are one potential source. Providers also use other data, and their practices and safeguards differ.
- The GDPR applies where its scope is met, including to relevant personal data processing by AI providers.
- Removing information from a public source can reduce future access to that source. It does not guarantee removal from an already trained model.
Where can information in an AI answer come from?
Three channels matter:
- Training data. Large models learn from varied collections of information, which can include publicly available web content. A public page mentioning you may have been included, but this cannot be established just by asking the model. OpenAI describes the main sources used to develop its models, including public information, licensed or partnered sources, and material supplied or generated by people.
- Your conversations. Information you type or upload may be stored or used according to the service, account type and settings. Check the provider's controls before supplying sensitive material.
- Live web access or connected sources. Some services can retrieve pages while answering. Information still available on a website may therefore be found without having been part of the model's training.
These are distinct channels. Public-source removal addresses one part of exposure; account privacy controls address another. A free footprint scan can help identify some websites and records mentioning you.
How the GDPR applies
Personal data does not lose legal protection because an AI system processes it. The GDPR applies where the processing falls within its scope. See the CNIL's AI resources.
Depending on the circumstances, relevant rights include:
- Access: ask about the personal data processed about you.
- Rectification: request correction of inaccurate personal information.
- Erasure: request deletion where the conditions in GDPR Article 17 apply.
Correcting an output, removing a source, deleting account data and changing learned model behaviour are different technical actions. Removing a specific item from an already trained model can be complex. A rights request should clearly describe the information and desired outcome rather than assume every provider uses the same method.
Read personal data rights explained for the general framework.
Remember: report inaccurate personal information to the provider and address an exposed original source where possible. Neither step guarantees that every future AI answer will change immediately.
Privacy risks to consider
- Disclosure or confusion: an answer may mention your work, city or profile details, sometimes mixing them with a namesake's information.
- False statements: models may confidently generate inaccurate claims, often called hallucinations.
- Aggregation: tools that combine public records can make profiling easier. This overlaps with the wider data broker ecosystem.
- Sensitive uploads: pasting a confidential document into a service creates a new disclosure to that provider. If credentials were exposed, see what to do about a compromised password.
How to reduce exposure
- Map your footprint. Run a free scan covering public websites, brokers and known breaches.
- Address original sources. Request removal from directories, people search sites and brokers where applicable.
- Review search results. Personal-content removal and delisting may reduce discovery through Google, though the original page can remain accessible elsewhere. Read removing your name from Google.
- Review AI privacy settings. Where available, disable use of your conversations for model improvement and check retention controls. Avoid sharing sensitive information unnecessarily. OpenAI explains how content is used to improve model performance; consult each provider's own policy for its service.
- Check again over time. New sources and tools appear, so periodic review can be useful after the first cleanup.
Our method focuses on identifying exposed sources, sending appropriate removal requests and tracking the process. It does not promise to make an already trained model forget information.
See also: Privacy rights explained · Remove your traces online · Remove your name from Google · Who buys personal data



