Who Actually Reads Your AI Companion Chats? Moderation Explained

Privacy & staying safe

Nearly everything you type to an AI companion is seen by no person at all. Some of it is, and the rules for when are usually tucked into a privacy policy. Here's how the layers stack up.

We may earn a commission from links on this page. It never changes a rating.

People say "moderation" as if it were a single thing. Inside a companion app it's at least four separate processes, each with its own speed and its own set of people involved.

The four layers

Four layers of moderation in AI companion apps: a filter on every message, automated flagging of chats, human review of flagged material, and legal or police requests

Going down the list, less gets seen, but more closely.

1. Instant screening. Software checks each message you send and each reply the model produces before it appears. It's hunting for prohibited categories: minors, some kinds of violence, hints of self-harm. When it finds one it can stop the message, soften the answer or display a notice. That's the source of those refusals and crisis cards, and our note on content filters covers why they occasionally interrupt a scene.

2. Pattern-spotting across chats. Other systems may rate entire conversations or accounts for warning signs: someone repeatedly probing a filter, harassment, a sign that a child is on an adult service. Being flagged isn't a penalty. It simply means a person may glance at it later.

3. People. Trust and safety teams, frequently employed by outside contractors, work through a portion of flagged chats, user complaints and appeals. Certain firms also pull samples of ordinary chats to check quality or train models. Privacy policies tend to cover this with wording such as "to improve our services" or "to enforce our terms".

4. Formal requests from outside. Courts, police and regulators can require data through legal orders, and the company answers according to its own policy and whichever laws bind it.

What tends to bring a human in

TriggerReason
Content about minorsLegally required in most places, and often passed on to authorities
Signs of self-harm or suicideCrisis handling is now expected or required in a growing number of jurisdictions
Threats aimed at real peopleRisk of genuine harm
A report or support ticket from youYou asked somebody to have a look
Persistent attempts to get past filtersEnforcement of the terms of service
Random spot checksProduct improvement, if the policy allows it

Where the law is heading

In Australia, the eSafety Commissioner's industry codes, in force from March 2026 for most obligations, expect companion chatbots to keep sexually explicit material away from children, either by checking ages or by not generating it, and to direct users to crisis and mental health help. Across the Pacific, New York (2025) and California (2026) have passed chatbot laws that oblige apps to spot signs of suicidal thinking and refer people to crisis services. In both cases the real-world result is more automated watching of your most sensitive chats, not less. See AI companion laws for the detail.

Reading a privacy policy with this in mind

Use your browser's find tool on the privacy policy and terms, hunting for:

  • "review", "moderate", "human" to learn whether staff read conversations and for what purpose.
  • "contractors" or "service providers" to see if the reviewers belong to another company.
  • "improve" or "train" to find out whether regular chats get sampled.
  • "law enforcement", "legal process", "court order" to see when data is surrendered and whether a court order is a precondition.

Good policies list the triggers and say staff open chats only when they must. A vague one giving staff access to "all content for any business purpose" is informative as well. The remaining parts of the document are covered by our four privacy checks. And if you're curious what a company holds on you, Australia's Privacy Principles give you a right to ask, as laid out in our data request guide.

Sensible conclusions

  • Type as though a reviewer could see it. In a small number of cases one will.
  • Keep other people's details out of it, particularly in explicit scenes, because a reviewer going through a flagged chat sees everything in it.
  • When a filter trips over fiction, a short out-of-character line normally fixes it. Arguing while in character can trigger further flags.
  • If privacy from staff matters more than anything, the one option where nobody else can read your chats is a companion run on your own machine. See running an AI companion locally.

Watching for rule-breaking isn't the same as spying on everyone. For nearly all conversations, no person and no process looks past the automatic screen. The exceptions, though, are written into the policy rather than into the app's friendly manner, so learn them before they matter.

Frequently asked questions

Do staff read my AI girlfriend chats?

Not routinely, as a rule, though most apps keep the right to. Staff usually see conversations that were flagged by automated systems, reported, tied to a support request or sampled for product improvement. The privacy policy should tell you which applies.

What happens if a chat gets flagged?

Mostly nothing you'd notice, or a blocked message or a warning. Repeated or serious breaches can get an account suspended, and material involving minors or genuine threats can be passed to authorities.

Can police get hold of my chats?

They can ask the company for them, and companies go along with valid legal orders. Mozilla's 2024 review found most romantic chatbot makers said they could hand data to authorities, sometimes without a court order. Assume anything stored could be disclosed.