
What good is a promise not to train on your conversations if the chat application sends their titles to an analytics service?
The industry would prefer you debate the model. The plumbing deserves attention too.
[“Prompt like a Butterfly, Sting like a Tracker”](https://dspace.networks.imdea.org/handle/20.500.12761/2073) examines that plumbing across nine conversational AI services. Its findings are less sweeping than “everyone sells your chats”—and more useful. Apparently, reinventing computing did not require reconsidering the tracking stack.
Nine chatbots, one familiar mess
Researchers tested web and Android deployments of ChatGPT, Claude, Grok, DeepSeek, Perplexity, Gemini, Microsoft Copilot, Mistral’s Le Chat, and Meta AI in Spain during May 2026, using static and dynamic analysis across account and consent configurations.
The tally:
- 124 third-party domains representing 44 organizations, including 34 advertising/tracking services.
- Conversation artifacts disclosed by six web clients and three mobile clients.
- 80.8% of third-party trackers remained active after nonessential cookies were rejected.
- Free and premium accounts contacted largely identical third parties.
That last finding deserves a place next to the subscription button. Paying removes some product restrictions. It does not automatically remove the telemetry supply chain.
The manuscript became publicly available in September 2026; its repository lists it as peer-reviewed and in press for PETS 2027. These are measurements of particular deployments, not a permanent verdict on every version of every service.
A title can tell on you
The provider-specific findings matter. ChatGPT and Claude sent conversation identifiers to Datadog. Gemini sent conversation titles to Google Analytics. Grok’s shared-conversation pages sent the latest prompt to Meta and TikTok, and screenshots to TikTok. Several disclosures depended on accepting nonessential cookies.
Those are different exposures. Treating them as interchangeable makes for a punchier headline and a worse threat model.
<> A conversation title can disclose sensitive information without exposing the full transcript. Privacy has to cover the application, not merely the model./>
Imagine an automatically generated title such as “Questions about my HIV diagnosis.” That is a hypothetical example, not a measured payload. But it illustrates why “only metadata” is such a slippery reassurance: summarization can concentrate the sensitive part.
An identifier alone is not equivalent to that title. Its risk depends on what it can be linked to or used to access. Publicly accessible conversation URLs introduce another exposure pathway.
Precision matters. The research does not establish that every third-party connection contains a transcript, nor that observed disclosures prove sale, retention, or profiling.
Hot Take: “No training” is an incomplete privacy pitch
My controversial position: “We don’t train on your data” should never pass a procurement review by itself. It answers one question while leaving the application’s other exits unexamined.
OpenAI says it does not share conversations with advertisers or sell user data to them. That is a relevant policy statement, not a verified response to this paper—and it does not settle questions about observability vendors.
Meanwhile, OpenAI began U.S. ChatGPT advertising tests in February 2026 and reported a $1 billion annualized advertising revenue run rate in August. That is a company-reported pace, not a billion dollars already collected. Still, the commercial incentive is hardly subtle.
Advertising does not prove abuse. It does make independently verifiable boundaries more valuable than reassuring copy.
Audit the exits
For developers, the actionable work is refreshingly unglamorous:
1. Strip conversation content from telemetry. Include generated titles, URLs, error breadcrumbs, and session replay in that review.
2. Test rejection on the wire. A cookie banner is not evidence of what the browser stopped sending.
3. Constrain sharing. Use authenticated, revocable, expiring access where confidentiality matters.
4. Separate audit scopes. Consumer-app findings do not establish identical API behavior; your custom frontend can nevertheless recreate the problem.
The study has limits: each configuration was analyzed once after deterministic pilots, and consent did prevent some disclosures. Neither caveat makes the findings disappear.
The chatbot may be new. The need to inspect every outbound request is not.

