How to Get Better Results from a Photo
The input photo decides most of the outcome. Framing, lighting and a few things to avoid, plus what to do when a result comes back wrong.
Almost every disappointing result traces back to the input, not the model. An emoji has to read at 128 px, which means the model has very little room — it has to find the subject, simplify it hard, and keep whatever makes it recognisable. Give it a photo where those things are easy to find and the output improves immediately.
Framing
- Fill the frame. The subject should take up roughly the middle third or more. A face that occupies 5% of a wide holiday snap gives the model almost no pixels to work from.
- One subject. Two faces in frame usually produces a blend of both or an arbitrary choice between them. Crop to one before uploading.
- Head-on beats three-quarter, three-quarter beats profile. Emoji read as symmetrical shapes. A hard profile loses the second eye, and the result stops looking like the person.
- Leave a little margin. A subject cropped tight to the edge of the frame often comes back clipped.
Lighting
Soft, even, front-on light is what you want. Window light on an overcast day is close to ideal. What to avoid:
- Backlighting. A bright window behind the subject turns the face into a silhouette, and the model has nothing to restyle.
- Hard side light. Half the face in shadow tends to come back as half a face.
- Direct flash. Blown-out highlights on the forehead and nose read as missing detail, and the emoji looks flat.
- Strong colour casts. An orange sodium streetlight or a purple LED shifts skin tones, and the model carries the cast through into the result.
Resolution
Upload the original, not a screenshot of it. A photo that has been sent through a chat app, screenshotted, and cropped has already lost detail and gained compression artefacts, and those artefacts survive into the emoji as mush around the edges. Anything from about 800 px on the short edge upwards is plenty; JPG, PNG, WebP and HEIC all work.
Things that reliably cause trouble
- Sunglasses and heavy occlusion. Mirrored lenses hide the eyes, which are the strongest identity cue. Clear glasses are usually preserved fine.
- Hats with wide brims. They shadow the upper face and often get read as part of the head shape.
- Busy backgrounds. A patterned wall or foliage directly behind the head can bleed into the silhouette. A plain-ish background is worth the ten seconds it takes to move.
- Motion blur. Nothing recovers from this. Pick a different frame.
Which subjects work best
Pets are the strongest category — a dog or cat silhouette is instantly readable at small sizes, and small inaccuracies do not register the way they do on a human face. Faces are next. Objects with a distinct outline (a guitar, a coffee cup, a car) work well. Things that struggle: crowds, hands in complex poses, text and logos inside the photo, and anything whose identity depends on fine detail.
When a result comes back wrong
- Re-run the same photo before changing anything. Generation is not deterministic, and the second attempt is often materially better.
- Then crop tighter. This fixes more problems than any prompt change.
- Then add one style hint — see the prompt recipes. One hint, not five; stacked instructions tend to cancel each other out.
- Then change photo. If two good crops of the same image both fail, the image is the problem.
Because generation is not deterministic, it is worth running two or three attempts of a photo you care about and picking the best, rather than tuning a prompt against a single unlucky output.
Make an emoji now — no account needed for your first few each day. Or browse the library to see what other people have made.