Guide · Hotel Lobby AI
The Hotel Lobby AI Trend: Where It Came From and How to Make One
A 2022 live session with one orange wall and one hanging mic turned into the duo video everyone is remaking in 2026. This guide covers where it started, the prompt that holds the look together, and the models that do the rapping.
By the AI Rap Video teamUpdated October 3, 20269 min read
Short answer
The Hotel Lobby AI trend recreates Quavo and Takeoff’s June 2022 COLORS performance of “Hotel Lobby” (one flat orange room, one microphone hanging between two rappers) with two new people swapped in. It took off on TikTok and X in mid-September 2026. To make one, you need two clear front-facing photos, a prompt that pins down the booth, the hanging mic and the two-shot framing, and a video model that generates the rap audio and the lip-sync together, such as ByteDance’s Seedance 2.0.
Where the Hotel Lobby trend came from
The trend is built on a song and a single performance of it. “Hotel Lobby (Unc & Phew)” came out on May 20, 2022 as the lead single from Quavo and Takeoff’s joint album Only Built for Infinity Links (Quality Control Music and Motown). Murda Beatz, Keanu Beats and Fabio Aguilar produced it, and it reached No. 55 on the Billboard Hot 100.
The look everyone copies comes from June 17, 2022, when the COLORS channel published the duo’s session. Every COLORS show is filmed in a room painted one flat colour. Theirs was orange, with one microphone hanging from the ceiling between them. That upload has passed 100 million views.
Four years later the frame became a template. On September 16, 2026, the TikTok account @glorify_studio posted an AI edit that replaced the two rappers with two characters from the series Suits. It drew about 1.7 million views in twelve days, according to Know Your Meme. Within a week the same shot was hosting cats, historical figures and celebrity couples across TikTok and X.
On September 23, 2026 Quavo reposted the original performance with the caption “The Original #HotelLobby”. It is worth remembering why that matters: Takeoff died in November 2022, a few months after the session was filmed, and some fans find AI versions of the two of them hard to watch. The kindest version of this trend is the one it turned into, where people put themselves and their friends in the booth.
Why the shot works: the five things to copy
Part of why the trend spreads is that the original is easy to describe. Anyone recognises it from five elements, so a model only has to get those five right for the clip to read as “Hotel Lobby”, even with completely different people in it.
| Element | Why it matters | Words that get it into the prompt |
|---|---|---|
| A flat orange room | The colour is the signature; a busy background breaks the reference instantly | “seamless bright orange walls”, “small live-session recording booth” |
| One hanging mic between them | It anchors the composition and gives the two performers something to lean into | “a single silver studio microphone hanging from the ceiling between them” |
| Two people, waist up | A tight two-shot keeps both faces large enough to stay recognisable | “two friends perform together”, “left rapper” / “right rapper” |
| Trading lines | The back-and-forth is the energy of the original; one rapper vibes while the other raps | “they take turns rapping”, “whoever is not rapping vibes to the beat” |
| One continuous handheld shot | Cuts and camera tricks make it look like a different video | “one continuous shot”, “handheld camera with subtle movement” |
The Hotel Lobby AI prompt, explained
There are two ways people make these clips. Motion-transfer tools take the original COLORS video and rebuild it around your references, which copies its timing but starts from footage you do not own. Reference-to-video models start from your photos and a text prompt, and they generate the motion, the beat and the vocals from scratch. We use the second approach. It means every clip is original and you can rap about anything, so your post is less likely to be muted or claimed.
A good reference-to-video prompt for this trend has five parts, in this order: the format and the shot, who is who, the set, the performance and audio, then what to leave out. This is the template our generator sends to Seedance for the Orange Booth stage, filled in with an example:
Vertical 9:16 rap music video shot for phones, one continuous shot. Two friends perform together as a rap duo inside a small live-session recording booth with seamless bright orange walls, a single silver studio microphone hanging from the ceiling between them, soft even studio light. The rapper on the left is the person from image 1 and the rapper on the right is the person from image 2. Keep each face, hairstyle, skin tone and age exactly as in their photo. A punchy hip-hop beat with deep 808 bass and crisp hi-hats plays from start to finish. They take turns rapping straight to the camera with confident hand gestures and head nods, lips perfectly synced to every word; whoever is not rapping vibes to the beat and ad-libs. They rap in English with clear, rhythmic delivery. The left rapper raps: "You burned the rice, I could smell it down the hall" The left rapper raps: "Smoke alarm singing, that’s your dinner call" The right rapper raps: "Your soup’s so thin it’s water with a wish" The right rapper raps: "At least my kitchen made an actual dish" Handheld camera with subtle movement, realistic skin and lighting. No on-screen text, no captions, no subtitles, no logos.
- Say the format first. Models weigh the opening words heavily; “Vertical 9:16 … one continuous shot” stops them from cutting to a wide shot or a second angle.
- Tie each photo to a side. “The rapper on the left is the person from image 1” is what keeps the two identities from blending or swapping places halfway through.
- Describe the set in physical terms. “Seamless bright orange walls” and “hanging from the ceiling between them” work better than naming the song or the channel, which the model may not know and which you should not lean on anyway.
- Write the lyrics into the prompt. With audio-native models the words you quote are the words they rap. Give each line to a side so the mouths move on the right person.
- End with the negatives. Captions and logos are the most common unwanted extras; asking for none keeps the frame clean for your own captions later.
In our Hotel Lobby generator you never type this by hand. You pick the Orange Booth stage and add an optional topic, and the prompt is assembled for you. A language model writes the bars to fit the clip: about one line every 2.5 seconds, nine words or fewer each, traded in pairs, clean, and never quoting a real song.

Photos and lyrics that render well
Most bad results come from the inputs, not the model. Before you blame the prompt, check the photos and the lines against this list.
- One person per photo, facing the camera. A front-facing selfie with both eyes visible is the single biggest quality lever. Side profiles and group shots confuse the identity lock.
- Even light, no sunglasses. Harsh shadows and covered eyes get “repaired” by the model, and that is where faces drift.
- Shoulders in frame. The shot is waist up, so a photo that shows your shoulders and clothes gives the model something to continue.
- One photo of two people works too. If you only have a picture of the two of you together, use it, with the left person on the left.
- Short, rhythmic lines. Long sentences make lips race. Nine words or fewer per line keeps the sync tight.
- Your own words. Rapping the original lyrics invites copyright claims; a topic like “who’s the better cook” or “ten years of friendship” is funnier anyway.
| Problem | Usual cause | Fix |
|---|---|---|
| The faces drift or look like strangers | Low-light, angled or crowded photos | Use a clear front-facing photo per person; keep the identity sentence in the prompt |
| No mic, or two mics | The set is described loosely | Name exactly one mic and where it hangs: “a single … hanging from the ceiling between them” |
| Captions or a logo appear | The model imitates social video | Add “No on-screen text, no captions, no subtitles, no logos” |
| Lips fall out of sync | Lines too long for the clip | Fewer, shorter lines; about one line per 2.5 seconds |
| It cuts to another angle | No shot instruction up front | Open with “one continuous shot” and the aspect ratio |
The models behind Hotel Lobby AI videos
The trend arrived in the same year video models learned to sing. The breakthrough that makes Hotel Lobby clips possible is native audio-video generation: the model produces the beat, the vocals and the mouth movements in one pass, so the lips land on the words instead of being dubbed afterwards.
- Seedance 2.0 (ByteDance, released February 2026) generates audio and video jointly from text, images, audio and video references. Its platform accepts up to nine images, three video clips and three audio clips per request and returns clips of 4 to 15 seconds, natively at 480p or 720p. That is why most Hotel Lobby clips run around ten to fifteen seconds.
- Seedance 2.5 (July 31, 2026) doubles the length to 30 seconds in a single run, takes far more references (up to 30 images, 10 videos and 10 audio clips), and keeps characters and voices consistent across shots in one generation.
- Motion-transfer models and template apps work from the original clip instead. Higgsfield’s Genjutsu model, for example, has a Motion Transfer mode that keeps a reference video’s motion, camera and timing and rebuilds everything else from your references (Higgsfield offers Seedance 2.0 as well), and CapCut’s Dreamina publishes Hotel Lobby prompts and templates.
On AI Rap Video every quality tier is a Seedance model, and the price is per second of video, so you only pay for what you render:
| Tier | Model | Resolutions | Max length | A 12-second clip costs |
|---|---|---|---|---|
| Standard | Seedance 2.0 Mini | 480p, 720p | 15 s | 10 credits at 720p |
| Fast | Seedance 2.0 Fast | 480p, 720p | 15 s | 17 credits at 720p |
| Pro | Seedance 2.0 | 480p, 720p, 1080p | 15 s | 20 at 720p, 50 at 1080p |
| Ultra | Seedance 2.5 | 480p, 720p, 1080p | 30 s | 37 at 720p, 93 at 1080p |
Standard is what the trend looks like on most feeds. Pick Pro or Ultra when you want 1080p or a clip longer than 15 seconds. Failed renders are never charged. Current prices are on the pricing page.
How the video at the top was made: Claude, frame by frame
The explainer at the top of this page is the opposite of an AI video. Nothing in it was generated by a video model except the sample clip playing on the phone. It follows the method of pdoom-video, an open-source music video for the song “I’m Upping My P(doom)” that its author made with Claude in Claude Code and released under the MIT licence. The official three.js account shared it in late September 2026. Every frame of that video is a pure function of song time, so the live preview and the final export match exactly.
- An original song. We wrote the lyrics for this guide, explaining the trend itself, and generated a 58-second, 141 BPM track with Suno.
- Timing data. Word-level timestamps for every lyric, a beat grid and downbeats from the audio, and separate vocal and instrumental stems.
- The edit as code. Every cut is placed on the beat a lyric line starts on. The edit refers to lyric lines and bars, never to typed-in seconds, so if the song is re-timed the edit follows it.
- Scenes in TypeScript. The booth wall, the swinging mic (a damped pendulum), the selfie cards, the prompt that types itself and the phone are all drawn by code. The chorus runs the sample render through duotone, halftone and engraving looks in a three.js shader.
- Rendering. Headless Chrome renders each frame, averaging up to 108 sub-frames for real motion blur, and ffmpeg encodes the result with film grain.
Claude wrote the renderer, the scenes and the edit in conversation, checking rendered stills along the way. It is a good illustration of the trend’s two halves: the code-drawn diagram explains the format, and the generated clip on the phone shows how short the path from two selfies to a finished video has become.
Make your own Hotel Lobby video
- Open the Hotel Lobby AI generator and add two photos, one person each, or one photo of you both.
- Keep the Orange Booth stage selected.
- Optionally type what the two of you should rap about. Leave it blank and the bars are about being the best duo in the game.
- Render. A standard 12-second clip is ready in about three minutes, in HD with no watermark.
Sources
- Hotel Lobby (Unc & Phew), Wikipedia
- Quavo “Hotel Lobby” AI Trend, Know Your Meme
- Quavo brings back the original Hotel Lobby performance, The Source
- Seedance 2.0 official launch, ByteDance Seed
- ByteDance launches Seedance 2.5, TechNode
- Higgsfield Genjutsu
- Hotel Lobby AI video trend, CapCut Dreamina
- pdoom-video (MIT), GitHub




