Nearly every gripe people have about these apps, whether it's repetition, drifting, or the same scene playing out differently on a second go, traces back to one mechanism. Ten minutes on it changes "this app is broken" into "this is what the app does, and here's the setting behind it."
Built piece by piece
The model doesn't draft a whole answer and then display it. It produces a single token (about a word, or a chunk of one), rereads everything including that new token, and produces the next.
Round and round it goes, with no blueprint, no skeleton and no proofreading. A reply that finishes beautifully had no clue where it was headed at the start.
You've probably met both of the results.
It can't take things back. Once the model has written "I've never been to Paris," everything afterwards leans on that line. It will construct a story to stay consistent with an accidental sentence before it will contradict itself.
Small slips grow. A slightly off word in the first sentence bends the second, which bends the third. That's why lengthy replies wander while short ones seldom do.
A dice roll for every word
At every step the model is looking at a league table of candidate words, each with a probability: say "good" 30%, "fine" 12%, "terrible" 3%, followed by a long list of outsiders.
Always choosing the favourite would make the character predictable to the point of tedium, with the same greeting and the same three jokes every time. So the app draws at random, weighted by those odds.
The knob that adjusts those odds is known as temperature. Low, and nearly all the weight sits on the safe choices, so she's steady, predictable and in time a bit boring. High, and the weight is spread thin, so she's livelier and more inventive but also more prone to saying something off-topic.
Hardly any app hands you that dial. The developers choose the balance, and that choice is a large part of why one app seems "sharper" than the next. Frequently it's not sharper at all, only warmer or cooler in temperament.
It's also the straight answer to "why did regenerating work?" Nothing got fixed. You rolled again and got a better result.
What the model has in front of it
Ahead of your message, the app quietly assembles a block of hidden text: her character profile, notes it has kept on you, a digest of your past and the last few exchanges. That bundle is the context window, and the reply is written as one continuous document following on from it.
Here's something most people never take advantage of: the model responds to the look of whatever is in front of it. Send three curt, colourless lines and curt, colourless lines come back, since it's extending the pattern. Send something rich and the whole register rises. There's no person to convince here. You're laying down a pattern, and patterns catch on from both sides.
That also explains a very common own goal. People slip into "hey", "how was your day", "what are you up to", then decide the app has gone downhill. It's carrying on the document you've been writing together.
One engine, many personalities
A lot of these apps sit on much the same base models but feel nothing alike. What separates them is the layer around the model:
- The system prompt: how the character is described, at what length, with what guidance on tone and pace.
- The sampling settings: where the temperature sits.
- What gets fetched: which saved facts and which stretch of history land in front of the model this turn.
- The filter: what's refused, and whether refusal is gracious or a brick wall.
Nomi stays consistent over weeks thanks to the third item, not a larger model. Candy AI puts chat, images and voice under a single subscription, which is a product decision rather than a modelling one. Neither difference would show in a benchmark, yet both appear within a fortnight of real use, which is how we score them.
Four ways to use this
Correct course early. How a reply opens decides where it goes. If a scene is heading wrong, stop and restate it rather than arguing with paragraph four.
Regenerate for a reason. A reply that's mostly good will lose its good part if you redraw it, while one that starts badly deserves an immediate redraw.
Type the way you want her to sound. It costs nothing and it's the quickest cure for a companion who has gone dull.
Let made-up facts go. Once the model has stated something untrue it tends to stick to it, because agreeing with its own text is its whole design. Correcting it in a fresh message works; demanding a confession doesn't. That topic has its own article: why AI companions invent things.
One last thing to keep in mind
Between your messages, nobody is waiting. She exists only while a reply is being produced, then gets reconstructed from text for the next one. The feeling of continuity is the work of the machinery around the model: saved facts, summaries and lookup. That machinery is what truly divides a US$10 app from a US$20 one.
That's the reason our ranking leans so heavily on memory and consistency, two things a sales page can't demonstrate.

