Pictures, Voice and Video: What Is Going On Behind the Scenes

How they work

Your companion's words and its pictures are made by unrelated systems that barely communicate. Knowing that clears up the shape-shifting face, and why every app charges for media before anything else.

We may earn a commission from links on this page. It never changes a rating.

Most of us imagine a companion app as a single smart mind that chats, sketches and talks. Under the bonnet it is closer to three or four unrelated tools bolted together, and the glitches with pictures and voice mostly happen where they meet.

Three tools behind one screen

The language model runs the conversation, and text is all it ever outputs.

The image model is another animal altogether, usually a diffusion model. Give it a written description and it renders a picture. Your chat history is a mystery to it.

The speech engine turns written lines into sound. It too works alone, aware of nothing except the sentence passed to it.

Here is what happens when you request a selfie. First the language model drafts a short brief for the shot. Next the app appends the character's stored look. The joint prompt goes off to the image model, which sends back a picture. Finally the chat replies to a photo it cannot actually see, relying on its own brief.

Almost every gripe about media in this category traces back to that relay.

Why her face keeps changing

A diffusion model never fetches your character. It makes up somebody who matches the words, and "long dark hair, green eyes, mid-twenties" describes an enormous crowd, so each attempt gives you a different person.

Apps tackle the problem to different degrees:

  • Stored look tags. One fixed phrase gets reused. Inexpensive, though only roughly steady.
  • A reference photo that steers each render. Much better, and the typical trick behind faces you can recognise.
  • A model tuned to one character. The steadiest by a distance and also the costliest, which is why you find it on top plans.

This is a real difference between products, and a free tier lets you test it at no cost. Ask for four shots of one character in four situations. Four unrelated women means no paid plan will fix it. Steady faces are a large part of why Candy AI heads our ranking, and why Secret Desires scores well: it develops looks, voice and personality together and adds short video.

Why media always gets a meter

Words are cheap: a chat reply costs the operator a sliver of a cent.

Each image is a good deal more expensive, voice accrues by the second, and video outstrips both.

That cost gap is the whole reason for the pattern on every pricing page: big message allowances, small image caps, voice minutes sold separately, video reserved for higher plans. It is not artificial scarcity to upsell you. It is the company's own invoice, handed on.

So when a listing says "unlimited", it almost always means unlimited text. Read it like that and the pricing pages click. The sums are set out in what image generation really costs you.

Voice: what it adds and what to check

Hearing a reply changes the experience more than most people expect. Reading a line and listening to it are not the same, and the difference is bigger than the technical description implies.

Before paying for voice, look at two things.

Delay. A pause of three seconds before each answer wrecks the illusion. Try it on a free tier or trial rather than trusting a demo video.

Call or clip. Voice notes, where the character reads its reply out, are common and inexpensive. Live calls, where you talk and it answers, are a separate product and usually sit on a separate plan. Nomi lists voice calls among its features, whereas plenty of apps write "voice" and mean clips.

What pictures still get wrong

Two limits to know about so you are not let down:

Hands, lettering and small details stay unreliable across every generator here. Nobody has fixed this, and an app that says otherwise is describing its best result, not its typical one.

The picture knows nothing about your scene. The chat model writes a one-line prompt, so anything outside that line is lost: the room you described, what she had on a few messages back, the time of day. Spelling things out in your request helps more than any setting.

Judging the media side in a single evening

Use up a whole free allowance in one sitting instead of one image per day. You are checking four things:

  1. Consistency. Four pictures, same character, different settings.
  2. Following instructions. Ask for something specific and see how much comes through.
  3. Speed. For both pictures and voice.
  4. What eats the allowance. Whether a failed or refused generation is still deducted.

That last one is the least documented and the most irritating to discover after paying. Along with everything else in what the free tiers actually include, it only comes to light when you use the thing properly.

Candy AI

4.6Rating: 4.6 out of 5

One subscription for chat, pictures and voice, with all three standing up to daily use.

Price
from US$12.99/month
Free tier
Yes
Australia
Available

Secret Desires

4.3Rating: 4.3 out of 5

Unfiltered roleplay that keeps its characters on track long after others lose the plot.

Free tier
Yes
Australia
Available

Frequently asked questions

Why does my companion look like someone new in every picture?

Each request makes the image model draw the character from scratch, using a text description. Apps that hold a face steady store a reference photo or a purpose-trained model. Without that, every image is a fresh throw of the dice.

Why do apps hand out so little voice time?

The company pays for speech by the second of audio produced. Writing a message costs vastly less, so text allowances run generous while voice time is rationed.

Can the image generator see what we have been chatting about?

Just via a brief instruction that the app composes. The conversation itself never reaches it, so a picture can leave out things that seemed obvious while you were chatting.