Follow the reference scene
Use the clip’s background, lighting and microphone placement as scene references rather than rebuilding them from text alone.
Upload two photos. Replace the duo in the built-in clip while following its scene and performance.
The orange-booth reference is included. Upload one photo for each performer to replace their full appearance and outfits.
Create my Hotel Lobby videoHotel Lobby is a two-performer video trend. A reference clip guides the scene, framing and actions, while two photos supply each replacement performer’s face, hair, body appearance and outfit. Review the result for identity and motion consistency.
Use the clip’s background, lighting and microphone placement as scene references rather than rebuilding them from text alone.
Photo 1 supplies the left performer; Photo 2 supplies the right. Replace their full appearance and outfits, including more than just the faces.
Use the reference gestures, reactions and action order rather than inventing new choreography or splitting the clip into equal halves.
Follow the source orientation, shot size and camera behavior instead of forcing a vertical full-body composition.
No reference video upload needed. The built-in orange-booth clip guides the performance; your photos supply both replacement appearances and outfits.
The orange-booth performance reference is included. No video upload needed. Preview the reference
Edit the performance in @Video1 using @Image1 for the LEFT performer and @Image2 for the RIGHT performer. Image 1 is the LEFT subject, Image 2 is the RIGHT subject; these images are separate identity references, NOT first or last frames. Replace exactly these two performers completely: face, hair, skin tone, body appearance, clothing and accessories must come from their respective images. Two adult friends, each retaining the face, hair, clothing and body proportions of their own reference image. @Video1 controls the background, microphone, lighting, framing, camera movement, poses, gestures, interactions and performance timing. Follow its original action order and turn-taking; do not invent an equal-half split, new choreography, a different backdrop or a wider full-body composition. Preserve the reference composition and original camera behavior, whether horizontal or vertical. Keep both identities assigned to the same left/right roles, with natural anatomy, stable clothing and believable contact with the floor. Output 11 seconds to cover the reference clip; for any rounding remainder, hold the final pose without restarting the performance. No merged faces, identity swapping, extra people, duplicate limbs, wardrobe flicker, added titles, logos, cuts or background changes. Do not introduce an existing song or recording.
Before you click Generate, your photos stay in local previews. Generation sends both photos and the built-in clip to the configured provider. Use images you have permission to upload.
Watch the orange-booth performance for its scene, framing and actions. The generator uses the built-in orange-booth clip.
Choose one clear photo for each performer and a matching preset. Show the hair, clothing and body appearance you want in the result.
Review the quote, including reference duration, then compare the result with the source for identity, clothing, action order, background and camera behavior.
It is a two-performer video format inspired by the Hotel Lobby trend. This workflow uses a reference clip for the scene, camera and actions, with two photos supplying replacement appearances and outfits.
Each photo supplies one identity. Photo 1 is assigned to the left and photo 2 to the right. Group photos increase the chance of extra subjects or blended identities.
Yes. Choose the matching preset and upload one photo for each subject. The pet prompts retain fur markings and use small head and paw gestures. Animal anatomy and identity still depend on model output.
The reference clip contains audio, but whether the result follows that track depends on the model and has not been tested. MiniMax has no audio switch; Seedance offers generated audio. No separate music file is included.
Use separate, unobstructed photos and a matching preset. The built-in clip supplies the performance, while the prompt binds both identities and follows its timing. Review the result; a frame-for-frame match is not guaranteed.
The existing model pricing rules include output and reference video duration in the quote. Credits are reserved at submission, with failed tasks refunded through the existing flow.