Start with one face in each image
Choose two separate adult portraits. If you only have a group shot, crop it into two separate files before uploading. Make each person’s face large enough to inspect. The uploader accepts JPG, PNG or WebP, up to 10 MB, with at least 240 pixels on both sides.
Keep the lighting straightforward
Even daylight or a softly lit room works as a useful starting point. Avoid deep shadows over eyes, strong backlight, sunglasses hiding the face and aggressive beauty filters. Similar framing across the two photos makes the cast easier to compare.
If the people blend or swap faces
Check that each input contains only one person. Use images with distinct, unobstructed faces and swap their sides if you want a different arrangement. The template asks the model to keep the two identities separate, but that instruction is not a guarantee.
If the face looks different
A generated video interprets a reference photograph rather than copying every pixel. Try a front-facing image with better lighting. Increasing resolution makes the output sharper; it does not guarantee stronger likeness.
If the mouth or hands look unusual
Speech, gestures and faces are generated together. Some takes may have imperfect lip sync or hands. The original Hotel Lobby recording is not used, and exact original movement is not promised. Watch the samples with soundto judge the style before purchasing.
If uploading fails
Confirm the file type, dimensions and size, then try again on a stable connection. Uploaded files are inspected and re-encoded before they are sent to the video service. A rejected upload does not use generation credits.
If the video is still processing
Open My Videos to see the persisted task. Waiting longer in the browser does not create a second task. A confirmed generation failure returns credits; a task whose upstream submission is still being checked needs reconciliation before its credits can be settled.