People say "moderation" as if it were a single thing. Inside a companion app it's at least four separate processes, each with its own speed and its own set of people involved.
The four layers

Going down the list, less gets seen, but more closely.
1. Instant screening. Software checks each message you send and each reply the model produces before it appears. It's hunting for prohibited categories: minors, some kinds of violence, hints of self-harm. When it finds one it can stop the message, soften the answer or display a notice. That's the source of those refusals and crisis cards, and our note on content filters covers why they occasionally interrupt a scene.
2. Pattern-spotting across chats. Other systems may rate entire conversations or accounts for warning signs: someone repeatedly probing a filter, harassment, a sign that a child is on an adult service. Being flagged isn't a penalty. It simply means a person may glance at it later.
3. People. Trust and safety teams, frequently employed by outside contractors, work through a portion of flagged chats, user complaints and appeals. Certain firms also pull samples of ordinary chats to check quality or train models. Privacy policies tend to cover this with wording such as "to improve our services" or "to enforce our terms".
4. Formal requests from outside. Courts, police and regulators can require data through legal orders, and the company answers according to its own policy and whichever laws bind it.
What tends to bring a human in
| Trigger | Reason |
|---|---|
| Content about minors | Legally required in most places, and often passed on to authorities |
| Signs of self-harm or suicide | Crisis handling is now expected or required in a growing number of jurisdictions |
| Threats aimed at real people | Risk of genuine harm |
| A report or support ticket from you | You asked somebody to have a look |
| Persistent attempts to get past filters | Enforcement of the terms of service |
| Random spot checks | Product improvement, if the policy allows it |
Where the law is heading
In Australia, the eSafety Commissioner's industry codes, in force from March 2026 for most obligations, expect companion chatbots to keep sexually explicit material away from children, either by checking ages or by not generating it, and to direct users to crisis and mental health help. Across the Pacific, New York (2025) and California (2026) have passed chatbot laws that oblige apps to spot signs of suicidal thinking and refer people to crisis services. In both cases the real-world result is more automated watching of your most sensitive chats, not less. See AI companion laws for the detail.
Reading a privacy policy with this in mind
Use your browser's find tool on the privacy policy and terms, hunting for:
- "review", "moderate", "human" to learn whether staff read conversations and for what purpose.
- "contractors" or "service providers" to see if the reviewers belong to another company.
- "improve" or "train" to find out whether regular chats get sampled.
- "law enforcement", "legal process", "court order" to see when data is surrendered and whether a court order is a precondition.
Good policies list the triggers and say staff open chats only when they must. A vague one giving staff access to "all content for any business purpose" is informative as well. The remaining parts of the document are covered by our four privacy checks. And if you're curious what a company holds on you, Australia's Privacy Principles give you a right to ask, as laid out in our data request guide.
Sensible conclusions
- Type as though a reviewer could see it. In a small number of cases one will.
- Keep other people's details out of it, particularly in explicit scenes, because a reviewer going through a flagged chat sees everything in it.
- When a filter trips over fiction, a short out-of-character line normally fixes it. Arguing while in character can trigger further flags.
- If privacy from staff matters more than anything, the one option where nobody else can read your chats is a companion run on your own machine. See running an AI companion locally.
Watching for rule-breaking isn't the same as spying on everyone. For nearly all conversations, no person and no process looks past the automatic screen. The exceptions, though, are written into the policy rather than into the app's friendly manner, so learn them before they matter.