Gen AI and LLM Data Privacy Ranking 2026

Each year, new competitors enter the AI race—and each year, Incogni’s researchers take a close look at the privacy practices behind them.

This year, Incogni’s researchers analyzed 13 AI platforms—ChatGPT, Claude, Gemini, Grok, Vibe (formerly Le Chat), Perplexity, Qwen, DeepSeek, Z.ai, Kimi, Meta AI, Pi, and Copilot—and scored them based on the privacy risks they pose.

These scores are based on 11 criteria, divided into three categories:

  • What happens to user data: whether conversations may be used for training, whether users can opt out, and what entities prompts may be shared with. 
  • How transparent platforms are: whether users can easily find and understand important privacy policy information. 
  • What data platforms collect and share: what personal data is gathered, where it comes from, and what parties it reaches.

Each provider received a weighted privacy-risk score, with lower scores indicating less privacy-invasive practices.

A note on terminology
This article refers to two kinds of names, depending on the context: platform names and provider names. ChatGPT, Copilot, and Gemini are platforms; OpenAI, Microsoft, and Google are their respective providers.This ranking covers platforms—the consumer-facing products with which people interact. Privacy policies, however, belong to the providers that write and manage them. For that reason, platform names are used when discussing the products themselves, while provider names are used once the discussion turns to privacy policies.
PlatformProvider
ChatGPTOpenAI
ClaudeAnthropic
GeminiGoogle
GrokxAI
Vibe (formerly Le Chat)Mistral AI
PerplexityPerplexity AI
QwenAlibaba
DeepSeekDeepSeek
Z.aiZ.ai
KimiMoonshot AI
Meta AIMeta
PiInflection AI
CopilotMicrosoft

Key insights:

  • The largest AI platforms turned out to be the most privacy-invasive, according to their privacy-risk scores. Gemini and Meta AI received some of the highest risk scores, while Vibe — less known to the general public—received the lowest risk score, making it the most privacy-friendly platform in this ranking.
  • ChatGPT had the most transparent privacy policy among the platforms assessed. OpenAI, the company behind ChatGPT, is the most transparent about how it uses user data—including messages shared within a chat—and its privacy policy is the easiest to understand among those reviewed.
  • Users are still told very little about what AI models were trained on. Most companies do not clearly explain whether training data may include personal information from websites, social media, or other public sources.
  • Privacy controls are often hidden from users’ view. ChatGPT, Claude, DeepSeek, Vibe, Gemini, Grok, Copilot, Perplexity AI, and Pi allow people to stop their conversations from being used for training with a simple toggle. However, that option is often hidden or even requires email communication or form-filling for other platforms.
  • 9 out of 13 investigated platforms maintain separate privacy disclosures for EU users, generally offering greater transparency or stronger controls than what applies elsewhere. Kimi, Z.ai, Qwen, and Meta AI have no equivalent EU-specific framework.
  • No platform allows users to have their data removed once it has been used to train a model. Opting out only stops future data from being used—it does not retroactively remove data already incorporated into a model’s training.

Vibe and ChatGPT lead the overall ranking with the lowest risk scores

Incogni’s researchers calculated a privacy score for each AI platform. Each score is based on 11 weighted criteria, grouped into three categories:

  • What happens to user data: whether conversations may be used for training, whether users can opt out, and which entities prompts may be shared with. 
  • How transparent platforms are: whether users can easily find and understand important privacy policy information. 
  • What data platforms collect and share: what personal data is gathered, where it comes from, and what parties it reaches.

Here are the results: 

There are three clusters of platforms that scored similarly in Incogni’s privacy ranking.

  1. The two lowest-scoring (least privacy-invasive) platforms, with strong privacy practices, with around 12 points each, with a clear lead over the rest.
  2. The middle group, comprising most platforms, with around 14 points each.
  3. The three highest-scoring (most privacy-invasive) platforms, which presented the greatest privacy risk with 16 points or more.

Vibe and ChatGPT received the lowest risk scores, corresponding to what Incogni’s researchers considered to be the least risky provider privacy policies in the study.

Vibe received the lowest overall risk score among the platforms evaluated, making it the most privacy-friendly platform in this ranking. This does not mean the provider posed the least risk in every individual category—only that its overall weighted score was the lowest among all the platforms under comparison. For example, Mistral AI—the company behind Vibe—placed fifth-lowest in the transparency category, and second-lowest in the data collection and sharing category. 

ChatGPT, while second overall, stood out particularly for transparency, offering the most approachable privacy policy among all analyzed platforms. OpenAI’s policy is easy to understand, even without a legal background, and gives users a clear picture of what happens to their data, mostly due to the extensive FAQ section.

By contrast, Copilot, Meta AI, and Kimi received the highest risk scores, pointing to greater privacy concerns according to the researchers’ criteria.

These platforms were considered the most privacy-invasive, mainly due to the complex data-handling approaches of the companies developing them, usually covering an extensive range of products without a clear delineation of which practices apply to which product.

Incogni’s researchers noted that many platforms suffered from similar issues, with one standing out significantly: a lack of transparency about what data goes into the training of the models.

What happens to a conversation after a message is sent

This section presents Incogni’s researchers’ findings on what happens to a user’s data after a message is sent to an AI platform. To determine this, Incogni’s researchers examined four aspects:

  1. Whether conversations users have with AI platforms may be used to train AI models.
  2. Whether users can opt out of having their chat history used for training.
  3. How much companies disclose about the data used to train their AI models.
  4. Whether prompts may be shared with third parties.

Each AI platform was assigned points based on how its provider handled information shared by users. A higher score indicates greater privacy risk in this category, with a maximum possible score of 9 points.

The results are as follows:

Vibe received the lowest score in this category, just below 4 points, indicating the strongest privacy practices among the assessed platforms.

Z.ai, Perplexity, ChatGPT, and DeepSeek followed closely, each scoring around 4 points.

Kimi received a substantially higher score than every other platform in this category, surpassing 7 points and representing the greatest privacy risk among the platforms assessed.

The four criteria behind these scores are detailed below, beginning with the question of whether AI platforms use conversations to train their models.

Do these platforms use conversations to train AI models?

For this criterion, Incogni’s researchers investigated whether the platforms use conversations to train AI models and how clearly they communicate such use cases to users.

The findings suggest that most, if not all, platforms use conversations and other user input to train their models to some degree. However, some platforms do not disclose this clearly enough to confirm it with certainty.

Most providers use a tiered approach: consumer accounts are typically included in training by default, while business, enterprise, and API tiers are excluded. OpenAI, Google, Microsoft, Perplexity AI, Mistral AI, and others follow this pattern. 

These providers are also straightforward about using users’ input to train their models, typically disclosing this within support or FAQ pages alongside clear steps for opting out.

Additional findings:

  • Anthropic used to be an exception to this pattern. Its documentation from March 2026 stated that conversations are used to improve Claude if a user chooses to allow it. A July 2026 update to the privacy policy changes that to explicitly stating that users’ input is used unless users opt out.
  • Meta states that direct conversations with Meta AI are used for training, but claims this does not apply when Meta AI is added to a group chat.
  • Z.ai confirms that it trains models, but its policy does not clearly state whether users’ conversations are included in that training.
  • Gemini ties this control to “Gemini Apps Activity,” a broader account setting rather than a dedicated training toggle, which may make the connection to AI training less obvious to users. Turning it off also stops chats from being saved to account history going forward—meaning privacy and convenience are bundled into a single decision.

Are users’ prompts shared with third parties?

Training a platform’s own models is one form of risk. Sharing user input with third parties is another, and this is the angle Incogni’s researchers examined next.

Specifically, researchers assessed whether prompts could be shared with third parties beyond those needed to operate, secure, or legally administer the service. Disclosures involving law enforcement, corporate transactions, and essential security providers were not counted against a platform, since these are standard practices generally required by law or necessary for the service to function.

One detail regarding law enforcement disclosures is worth noting, though. 

All platforms are legally required to cooperate with law enforcement requests in the countries where they operate, whether in the United States, the European Union, or China. What differs is not whether this cooperation happens, but how visible and reviewable the process is. 

In the United States and the European Union, such requests are generally subject to some form of judicial oversight, and companies can often disclose aggregate data about the requests they receive or challenge them in court. In China, comparable requests are not accompanied by independent judicial review or public reporting, leaving little opportunity for outside parties to assess how the process is used.

Some additional observations of note were:

  • Qwen states it may share personal data with analytics providers, search engine providers, and other third-party service providers.
  • xAI can share information with its related companies for purposes such as customer management and technical operations.
  • Perplexity AI may share data within its company group, including affiliates. Note that the lack of specificity in its policy means it doesn’t always make clear whether this specifically includes user prompts.

Are AI models trained on information available on the open web?

AI models are trained on content generated by real people—a substantial amount of it. As the volume of AI-generated content across the web increases rapidly, human-generated content is becoming scarcer. As a result, AI companies must draw on a wider range of sources to train their models. This raises the question of exactly which sources are being used.

This is why the details of model training were included as one of the factors considered in a platform’s privacy risk score. All but one of the platforms analyzed disclose that they train their models on “publicly available information.” On the internet, that typically means all accessible content, unless protected by a password, paywall, or encryption. 

In other words, most platforms collect information that’s freely accessible on websites across the internet, including social media data.

However, the fact that information is accessible online does not mean the person concerned intended it to be widely collected, repurposed, or used to train an AI model.

Incogni’s researchers highlighted the following:

  • Pi (developed by Inflection AI) is the only platform that openly admits to using social media profiles and content for AI-training purposes.
  • Kimi (developed by Moonshot AI) is the only platform that does not specify the source of its training material beyond user-generated content.
  • xAI discloses that Grok is trained in part on public posts shared on X (formerly Twitter), along with engagement data such as likes, reposts, and view counts.
  • Meta AI is trained on public content shared on Facebook and Instagram, excluding private messages unless a user chooses to share them directly with the AI model.
  • Several platforms, including Vibe, Qwen, Claude, and ChatGPT, also disclose training on private or licensed third-party datasets whose contents are undisclosed to anyone outside the companies involved. This makes it impossible for users to verify whether such datasets contain personal data or content that carries social or political bias.

Do the AI platforms make it easy to opt out of training?

Incogni’s researchers also checked how easy it is for ordinary users to stop their conversations from being used to train the respective models. The review focused on each platform’s consumer-facing tier, not its enterprise offerings.

Most platforms describe some form of opt-out mechanism in their privacy policy. However, these descriptions vary significantly. Some, like ChatGPT, provide clear instructions. Others describe the process only in general terms, without naming a specific setting or step. 

Incogni’s researchers categorized these platforms into four tiers, based on the ease of opting out.

Tier 1: A clear, single toggle, with no trade-off

  • ChatGPT, Claude, DeepSeek, Vibe, Grok, Copilot, Perplexity AI, and Pi all provide a simple toggle that users can activate—stopping these platforms from using chat history for model training—with no trade-off.

Tier 2: A real toggle exists, but it comes with a trade-off

  • Gemini lets users stop future conversations from being saved and used for training, but doing so also stops those conversations from appearing in chat history. 

Tier 3: The right exists on paper, but there’s no clean way to exercise it

Moonshot AI and Meta state that opting out is possible, but neither makes the process clear. 

  • Moonshot AI requires the user to email support and verify their identity. 
  • Qwen also requires users to submit an opt-out request
  • Meta’s policy states that users can object to AI training use, but never explains how.

Tier 4: No functioning opt-out found

  • Z.ai mentions only a general right to object to data processing, without addressing training specifically.  

Are users likely to understand what happens to their data?

In ranking the AI providers, Incogni’s researchers dedicated a whole category of criteria to transparency. Their focus was on how easy it is for a user to make an informed decision as to whether they want to use an AI platform, given the effect such use could have on their privacy. The researchers looked at conversation training, model-training data, and the accessibility of privacy policies and supporting resources.

Platforms could receive up to 8 privacy-risk points in this category (the lower the score, the better).

In this analysis, OpenAI came in first place based on its transparency surrounding privacy-impacting practices, achieving the lowest privacy risk score in this category. It performed exceptionally well given its low risk scores in the model training transparency and transparency of conversation use for training criteria, even though the privacy policy accessibility criterion constituted the largest share of its risk footprint.

The remaining providers—ranging from Perplexity AI to Anthropic—hovered in a mid-tier distribution, with overall risk scores escalating gradually from roughly 5 to 5.5. Conversely, Microsoft came in last, presenting the highest overall transparency risk, closely preceded by Meta and Z.ai. These worst-performing platforms suffered significantly from extremely poor scores in the privacy policy accessibility criterion, which represented the vast majority of their privacy risk profiles.

The first two criteria assess how clearly platforms explain whether conversations are used for training and what data their models were trained on, and how accessible that explanation is to users. Incogni’s researchers noted that a user’s ability to make an informed, privacy-conscious decision about sharing information with an AI model is hampered if they don’t know if that conversation is being used by the AI provider for any reason. Similarly, Incogni’s researchers noted that an inability to find information about what models are based on can lead people to use tools they would otherwise not support.

Incogni’s researchers noted that none of the providers revealed the exact datasets used to train their models. Even the most transparent providers relied on broad descriptions of the sources of their data, such as “publicly available information,” “private datasets,” and “social media data.” 

The other criterion in this category refers to the accessibility of the companies’ privacy policies, in which researchers considered both the linguistic complexity and less objective factors, like how information is presented, including whether FAQs or other dedicated privacy pages were used to explain key practices in simpler language. The researchers highlighted the following:

  • All of the investigated companies were found to have privacy policies estimated to require a college-graduate reading level, as per the Dale-Chall formula. However, some were noticeably easier to follow than others.
  • Google, Microsoft, and Meta have broad privacy policies covering many products, and have limited information that exclusively covers their AI offerings. In turn, their overall privacy policies allow for very extensive data collection and use, which may or may not apply to those who only use their AI products (like Gemini, Copilot, and Meta AI).
  • Providers like DeepSeek, Moonshot AI, Qwen, and Z.ai present only the most barebones, international versions of their privacy policies, without implementing accommodations that would make them easy to interpret from a CCPA or GDPR perspective. The types of personal data interacted with are ambiguous, and the purposes for sharing or collecting each type of data are limited or non-existent.

User data collection and sharing

This section looks beyond the prompts people type into AI tools. For this set of criteria, Incogni’s researchers focused on what data is collected by the AI providers through their sites, mobile apps, and other means, as well as what other entities could receive user data.

Platforms could receive up to 6 privacy-risk points across these criteria.

According to the privacy risk rankings, Inflection AI claimed the top spot as the best-performing provider overall, achieving the lowest cumulative privacy-risk score. It’s followed closely by Mistral AI and Z.ai, both of which maintained relatively low risk profiles across the analyzed metrics. For these top-performing platforms, restricted data sharing with third parties and highly limited app data collection practices played a major role in keeping their overall risk relatively low. 

Conversely, Microsoft and Meta stood out as the worst performers, landing at the bottom of the list with the highest privacy risk scores. These platforms suffered from significant risk accumulation, driven heavily by their reliance on expansive sources of user information and extensive third-party data sharing.

The first criterion in this category concerns where providers can obtain information about users—whether it’s just through users’ interactions with the providers, or whether they seek out additional data from publicly available sources, marketing companies, affiliates, data brokers, and so on. The more sources, the worse (higher) the score. 

One thing that stood out to the researchers here is that, again, Microsoft’s broad privacy policy makes it hard to determine the privacy implications of interacting with any given product are. Microsoft does get data from data brokers, meaning that even those who only use Microsoft’s AI offerings, such as Copilot, could have their data enriched via data brokers.

The researchers approached the second concern in this category, which third parties can receive personal information from the investigated AI providers, with the more “generous” sharers receiving a penalty. Highlighted findings include:

  • Microsoft may share user data with ad networks (where users had given consent by opting in).
  • Perplexity AI, Qwen, xAI, and OpenAI mention their respective corporate groups and affiliates as recipients of user data.

Incogni’s researchers also looked at the types of personal data collected by each platform, using the categories defined under the California Consumer Privacy Act (CCPA). Where providers used different terminology in their privacy policies, researchers mapped their descriptions to the closest CCPA category.

Key findings include:

  • Meta, Microsoft, Inflection AI, and Perplexity AI can derive inferences about their users based on other collected information. These inferences can be used to predict a person’s interests, preferences, behavior, or characteristics, including details they have not directly disclosed.

Researchers also checked what the platforms’ iPhone and Android apps say they collect and share. Since not every company offered a mobile app, this part of the research covered only 12 apps. Apps received higher privacy-risk scores when they collected or shared more personal data, especially sensitive information and details that could identify users.

Some findings of note were:

  • ChatGPT, Microsoft Copilot, and Meta AI iOS apps disclose that at least some user data is used to help third-party advertisers. 
  • Gemini and Meta AI iOS apps collect sensitive data (as described by Apple, this can include racial or ethnic data, sexual orientation, pregnancy status, as well as disabilities, religious beliefs, and political opinions).
  • 7 of the investigated 12 apps collected email addresses, and 5 collected phone numbers. 

This shows that the privacy risks of AI platforms extend far beyond conversations with chatbots. Depending on the service and device used, platforms may also collect contact details, sensitive characteristics, and information used to build broad and detailed profiles of individual users.

Conclusion

AI platforms’ privacy controls can create an illusion of control. Users may be offered a simple switch to stop conversations from being used for training, while the platforms continue to collect, combine, or share personal information far beyond the chat window itself.

This 2026 ranking shows that privacy strengths rarely extend across an entire platform. A provider may explain its training settings clearly but still collect extensive account, device, app, or marketing data. Another may limit some forms of data use while making its practices so difficult to understand that users cannot meaningfully judge the trade-off.

How does it compare to last year’s ranking?

Compared to last year’s ranking, Gen AI and LLM Data Privacy Ranking 2025, the same divide remains visible. Mistral AI and OpenAI again ranked near the top, although even the highest-ranked platforms still came with serious privacy trade-offs—while the largest AI providers continued to be held back by broad policies spanning many products. For someone using Copilot, Gemini, or Meta AI, finding out what applies specifically to that tool can require interpreting rules written for an entire corporate ecosystem.

The most consequential gap is also the least visible 

None of the providers disclosed their exact training datasets. Labels such as “publicly available information,” “private datasets,” and “social-media data” reveal little about whose personal information may have been absorbed, how it is or was used, or whether it can ever be removed. Users should therefore pay particular attention to vague descriptions of third parties.

The lack of transparency around pre-training data is even more consequential. These datasets may contain personal information gathered from websites, social media, commercial sources, or other poorly described collections. Once that information has been absorbed into a model, identifying it and having it corrected or removed is often extremely difficult, if not impossible.

The result is a privacy model that still puts most of the work on the user. People are expected to locate the right setting, interpret complex policies, understand app-store disclosures, and trace how information moves across products and third parties. AI platforms increasingly sell simplicity, but protecting personal data while using them remains anything but simple. 

That imbalance is far from accidental: the harder data practices are to understand or avoid, the easier it is for platforms to keep collecting, combining, and monetizing user information.

Platform overview

The following summaries provide highlights noted by the researchers regarding what each provider did comparatively well, where the main privacy concerns appeared, and what users should know before using its platform.

Mistral AI

  • Mistral AI has extensive, comparatively easy-to-understand privacy documentation.
  • Its platform, Vibe, can collect user information from publicly available sources, and third parties are only partly described.
  • Among providers with apps, Mistral AI’s mobile app is one of the most privacy-friendly ones.
  • Per the policy, user conversations are shared only with a minimal number of third parties.

OpenAI

  • OpenAI has an easy-to-digest privacy policy and extensive, accessible FAQs and privacy resources.
  • The provider can get information about users via interactions with the platform, security partners, marketing vendors, and publicly available information. 
  • Makes opting out of conversation use for model training easy to find and toggle.

Inflection AI

  • Inflection AI’s privacy policy is combined with its cookie policy and terms of service, making it harder to drill down to specific concerns. It does cover legal bases for processing, accommodating EU privacy entities, but lacks CCPA-themed information and presentation. 
  • Next to the usual list of entities, Inflection AI’s privacy policy describes potentially sharing user data with its corporate group and commercial and research partners.
  • Has one of the most privacy-friendly apps among the investigated platforms.

Perplexity AI

  • Has an easy-to-follow, although not very thorough, privacy policy. The layout allows users to understand what data is used and for what purposes. 
  • Makes opting out of conversation use for model training relatively easy to find and toggle.
  • Corporate group members and affiliated parties can receive Perplexity user data
  • Its Sonar model is based on models pretrained by other companies, such as Meta and DeepSeek. This means that Perplexity AI cannot control what data is inducted into the initial training datasets. 

Alibaba

  • Alibaba’s privacy policy describes what data is used and to what end, and the use of tables makes information easier to process. However, there is little information about which data is shared with what parties.
  • Understanding whether user conversations are used to train the models can be challenging, as this question is not explicitly covered in the privacy policy.
  • Unless self-hosted, the model can store user inputs for model improvement and other purposes.

DeepSeek

  • DeepSeek offers open-weight models, which can mean that user data doesn’t leave the user’s network. However, for those using the DeepSeek hosted version of the model, privacy considerations become more serious.
  • DeepSeek is one of the few AI providers that invite users to reach out if they spot (inaccurate) personal information and ask for corrections or removal.
  • DeepSeek has one of the richest descriptions of how its models are trained and created. 

Z.ai

  • Z.ai also offers open-weight models, which can mean that user data doesn’t leave the user’s network. However, for those using the Z.ai hosted version of the model, privacy considerations become more serious.
  • The privacy policy for Z.ai was found to be middle-of-the-pack in the sample, using approachable language, but not going into too much detail.
  • Leaves ambiguity in its privacy policy regarding the use of user conversations for model training and improvement. 

Google

  • Offers a partial privacy policy for its Gemini products. That is, in cases not described in the shorter, more limited document, it defaults to its general privacy policy. This means that getting answers to specific questions can be a real challenge. 
  • There have been news media reports describing Google failing to respect websites’ no-crawl signals, intended to limit data collection from them. 
  • Google does not rely on external sources to gather user data, although, through its wide suite of offerings, it already collects a lot of information and is able to make inferences about its users.

Anthropic

  • Anthropic had one of the more extensive privacy policies and included other privacy articles or resources, helping users make informed decisions about their data. 
  • Anthropic’s models are claimed not to share their data for the purpose of model training by default. This is given as an option to users on an opt-in basis. 
  • Anthropic may publish research using aggregated, de-identified user data.

xAI

  • xAI’s Grok is one of the most privacy-friendly iOS apps in the investigated sample.
  • Corporate group members and affiliated parties can receive xAI user data.
  • Unless opted out, xAI can train its models on users’ public conversations on X (formerly Twitter), and not only with its AI offering—Grok. 

Moonshot AI

  • At the time of writing, Moonshot AI has declared its intention to open-source its latest models’ weights, which would enable users to set up self-hosted versions of the model. However, for those who use the publicly hosted model, additional caution should be taken. 
  • As with other Chinese-owned platforms, there is a concern about it having to comply with state requests for user data. 
  • Moonshot AI’s app Kimi is one of the two (alongside Copilot) that uses data to track users.

Meta

  • Due to a very broad privacy policy, Meta creates difficulties in understanding the implications of just using its AI offerings.
  • Because of how the AI models are integrated into other services (e.g., WhatsApp), the use of people’s conversations for training purposes can become difficult to understand—the claim is that training takes place in private conversations but not group chats.
  • Has the most data-hungry apps in the investigated sample.

Microsoft

  • Due to a very broad privacy policy, Microsoft creates difficulties in understanding the implications of just using its AI offerings.
  • Buys data from data brokers.
  • Microsoft’s Copilot app shares data with third-party advertisers, being one of three that do so among the investigated iOS apps.

Methodology

To understand the privacy implications of the most popular AI platforms, Incogni’s researchers created a set of criteria according to which the platforms could be assessed. Each platform was scored from 0 (most privacy-friendly) to 1 (least privacy-friendly) for each criterion. Weights were applied to each of these criteria based on how important Incogni’s researchers and privacy experts determined them to be. The unweighted scores, as well as justifications for any weighting applied, are available in the public dataset.

The criteria were aimed at capturing three things considered important for user privacy when it comes to the use of advanced machine learning programs, like LLMs and multimodal AI platforms:

  • User data and model training: Incogni’s researchers investigated how the users’ data interacts with the base models and whether their prompts are shared with other entities.
  • Platform transparency: even if user data is used in a heavy-handed way, if the user can make an informed decision to continue interacting with an AI platform, at least they’re not in the dark about how their data is used. 
  • Data collection and sharing practices: researchers examined how much data is collected, what data is shared with which entities, and what sources of personal information the platforms draw upon.

A lot of these criteria required Incogni’s researchers to delve into the privacy policies and other legal resources provided by the platforms. The data was collected from June 15 to July 6. Note that some privacy policies might have been updated since then. 

In order to rate iOS and Android app data collection and sharing practices, Incogni’s researchers examined the data handling practices of the apps on the App and Play stores, respectively. For every data point collected or shared, points were added to the apps’ totals. Incogni’s researchers penalized the collection and especially the sharing of sensitive and personally identifiable data.

In order to attribute scores for data collection according to the privacy policies, Incogni’s researchers tried to match collected data to the categories of personal data defined in the CCPA. 

The public dataset with details about the criteria and findings can be found here:public data.

Notes on data

On the Google Play store, the developers of Microsoft’s Copilot app claim not to collect or share any user data, while admitting to doing both on Apple’s App Store. Incogni’s researchers, believing the Google Play store disclosure to be an error, gave Copilot the same score in the Android app ranking as it received in the Apple app ranking.

Many platforms (like Meta AI, Google Gemini, and Microsoft Copilot) lack specific privacy policies, leaving researchers with only broad, overarching privacy policies, which might overstate the extent of data collection. For example, Meta’s privacy policy details how Meta handles user information it could only have got from outside of Meta AI.

Another set of platforms (such as Kimi, Qwen, Z.ai) provides non-extensive and rather vague privacy policies. This creates issues in matching collected data to CCPA-defined categories of personal data, as well as in being fully certain about what data is used and how. 

In order to investigate how user data is treated, Incogni’s researchers had to rely on information specific to certain regions (e.g., the EU and California), choosing California-style privacy policies if available. This is a limitation of the research since practices specific to, for example, California might not apply to the European Union.

Incogni’s researchers generally tried to provide both provider names and the names of these providers’ most popular products or models (e.g., OpenAI and ChatGPT). Meta’s most popular model, Llama, does not feature prominently and instead frequently appears as Meta AI, whose offerings are based on the Llama models. The researchers decided to refer to the machine learning program offered by Meta as Meta AI and not Llama.

Is this article helpful?
YesNo
Scroll to Top