AI Rap Video logoAI Rap Video

Guide · Hotel Lobby AI

The Hotel Lobby AI Trend: Where It Came From and How to Make One

A 2022 live session with one orange wall and one hanging mic turned into the duo video everyone is remaking in 2026. This guide covers where it started, the prompt that holds the look together, and the models that do the rapping.

By the AI Rap Video teamUpdated October 3, 20269 min read

This video was not made with a video model. Claude wrote the renderer and every scene in TypeScript and three.js, and each frame was computed from the song’s beat grid and word timings, the method of the open-source pdoom-video project. The song is original. The clip on the phone is a real 12-second render from our pipeline, and every face in it is AI-generated.

Short answer

The Hotel Lobby AI trend recreates Quavo and Takeoff’s June 2022 COLORS performance of “Hotel Lobby” (one flat orange room, one microphone hanging between two rappers) with two new people swapped in. It took off on TikTok and X in mid-September 2026. To make one, you need two clear front-facing photos, a prompt that pins down the booth, the hanging mic and the two-shot framing, and a video model that generates the rap audio and the lip-sync together, such as ByteDance’s Seedance 2.0.

Where the Hotel Lobby trend came from

The trend is built on a song and a single performance of it. “Hotel Lobby (Unc & Phew)” came out on May 20, 2022 as the lead single from Quavo and Takeoff’s joint album Only Built for Infinity Links (Quality Control Music and Motown). Murda Beatz, Keanu Beats and Fabio Aguilar produced it, and it reached No. 55 on the Billboard Hot 100.

The look everyone copies comes from June 17, 2022, when the COLORS channel published the duo’s session. Every COLORS show is filmed in a room painted one flat colour. Theirs was orange, with one microphone hanging from the ceiling between them. That upload has passed 100 million views.

Four years later the frame became a template. On September 16, 2026, the TikTok account @glorify_studio posted an AI edit that replaced the two rappers with two characters from the series Suits. It drew about 1.7 million views in twelve days, according to Know Your Meme. Within a week the same shot was hosting cats, historical figures and celebrity couples across TikTok and X.

On September 23, 2026 Quavo reposted the original performance with the caption “The Original #HotelLobby”. It is worth remembering why that matters: Takeoff died in November 2022, a few months after the session was filmed, and some fans find AI versions of the two of them hard to watch. The kindest version of this trend is the one it turned into, where people put themselves and their friends in the booth.

Why the shot works: the five things to copy

Part of why the trend spreads is that the original is easy to describe. Anyone recognises it from five elements, so a model only has to get those five right for the clip to read as “Hotel Lobby”, even with completely different people in it.

ElementWhy it mattersWords that get it into the prompt
A flat orange roomThe colour is the signature; a busy background breaks the reference instantly“seamless bright orange walls”, “small live-session recording booth”
One hanging mic between themIt anchors the composition and gives the two performers something to lean into“a single silver studio microphone hanging from the ceiling between them”
Two people, waist upA tight two-shot keeps both faces large enough to stay recognisable“two friends perform together”, “left rapper” / “right rapper”
Trading linesThe back-and-forth is the energy of the original; one rapper vibes while the other raps“they take turns rapping”, “whoever is not rapping vibes to the beat”
One continuous handheld shotCuts and camera tricks make it look like a different video“one continuous shot”, “handheld camera with subtle movement”

The Hotel Lobby AI prompt, explained

There are two ways people make these clips. Motion-transfer tools take the original COLORS video and rebuild it around your references, which copies its timing but starts from footage you do not own. Reference-to-video models start from your photos and a text prompt, and they generate the motion, the beat and the vocals from scratch. We use the second approach. It means every clip is original and you can rap about anything, so your post is less likely to be muted or claimed.

A good reference-to-video prompt for this trend has five parts, in this order: the format and the shot, who is who, the set, the performance and audio, then what to leave out. This is the template our generator sends to Seedance for the Orange Booth stage, filled in with an example:

Hotel Lobby AI prompt (reference-to-video)
Vertical 9:16 rap music video shot for phones, one continuous shot. Two friends perform together as a rap duo inside a small live-session recording booth with seamless bright orange walls, a single silver studio microphone hanging from the ceiling between them, soft even studio light.
The rapper on the left is the person from image 1 and the rapper on the right is the person from image 2. Keep each face, hairstyle, skin tone and age exactly as in their photo.
A punchy hip-hop beat with deep 808 bass and crisp hi-hats plays from start to finish. They take turns rapping straight to the camera with confident hand gestures and head nods, lips perfectly synced to every word; whoever is not rapping vibes to the beat and ad-libs. They rap in English with clear, rhythmic delivery.
The left rapper raps: "You burned the rice, I could smell it down the hall" The left rapper raps: "Smoke alarm singing, that’s your dinner call" The right rapper raps: "Your soup’s so thin it’s water with a wish" The right rapper raps: "At least my kitchen made an actual dish"
Handheld camera with subtle movement, realistic skin and lighting. No on-screen text, no captions, no subtitles, no logos.
  • Say the format first. Models weigh the opening words heavily; “Vertical 9:16 … one continuous shot” stops them from cutting to a wide shot or a second angle.
  • Tie each photo to a side. “The rapper on the left is the person from image 1” is what keeps the two identities from blending or swapping places halfway through.
  • Describe the set in physical terms. “Seamless bright orange walls” and “hanging from the ceiling between them” work better than naming the song or the channel, which the model may not know and which you should not lean on anyway.
  • Write the lyrics into the prompt. With audio-native models the words you quote are the words they rap. Give each line to a side so the mouths move on the right person.
  • End with the negatives. Captions and logos are the most common unwanted extras; asking for none keeps the frame clean for your own captions later.

In our Hotel Lobby generator you never type this by hand. You pick the Orange Booth stage and add an optional topic, and the prompt is assembled for you. A language model writes the bars to fit the clip: about one line every 2.5 seconds, nine words or fewer each, traded in pairs, clean, and never quoting a real song.

Two friends rapping at one silver microphone hanging in an orange booth, a 9:16 frame from an AI Rap Video render
A frame from a 12-second Orange Booth render (Seedance 2.0 Mini, 720p). Both people are AI-generated.

Photos and lyrics that render well

Most bad results come from the inputs, not the model. Before you blame the prompt, check the photos and the lines against this list.

  • One person per photo, facing the camera. A front-facing selfie with both eyes visible is the single biggest quality lever. Side profiles and group shots confuse the identity lock.
  • Even light, no sunglasses. Harsh shadows and covered eyes get “repaired” by the model, and that is where faces drift.
  • Shoulders in frame. The shot is waist up, so a photo that shows your shoulders and clothes gives the model something to continue.
  • One photo of two people works too. If you only have a picture of the two of you together, use it, with the left person on the left.
  • Short, rhythmic lines. Long sentences make lips race. Nine words or fewer per line keeps the sync tight.
  • Your own words. Rapping the original lyrics invites copyright claims; a topic like “who’s the better cook” or “ten years of friendship” is funnier anyway.
ProblemUsual causeFix
The faces drift or look like strangersLow-light, angled or crowded photosUse a clear front-facing photo per person; keep the identity sentence in the prompt
No mic, or two micsThe set is described looselyName exactly one mic and where it hangs: “a single … hanging from the ceiling between them”
Captions or a logo appearThe model imitates social videoAdd “No on-screen text, no captions, no subtitles, no logos”
Lips fall out of syncLines too long for the clipFewer, shorter lines; about one line per 2.5 seconds
It cuts to another angleNo shot instruction up frontOpen with “one continuous shot” and the aspect ratio

The models behind Hotel Lobby AI videos

The trend arrived in the same year video models learned to sing. The breakthrough that makes Hotel Lobby clips possible is native audio-video generation: the model produces the beat, the vocals and the mouth movements in one pass, so the lips land on the words instead of being dubbed afterwards.

  • Seedance 2.0 (ByteDance, released February 2026) generates audio and video jointly from text, images, audio and video references. Its platform accepts up to nine images, three video clips and three audio clips per request and returns clips of 4 to 15 seconds, natively at 480p or 720p. That is why most Hotel Lobby clips run around ten to fifteen seconds.
  • Seedance 2.5 (July 31, 2026) doubles the length to 30 seconds in a single run, takes far more references (up to 30 images, 10 videos and 10 audio clips), and keeps characters and voices consistent across shots in one generation.
  • Motion-transfer models and template apps work from the original clip instead. Higgsfield’s Genjutsu model, for example, has a Motion Transfer mode that keeps a reference video’s motion, camera and timing and rebuilds everything else from your references (Higgsfield offers Seedance 2.0 as well), and CapCut’s Dreamina publishes Hotel Lobby prompts and templates.

On AI Rap Video every quality tier is a Seedance model, and the price is per second of video, so you only pay for what you render:

TierModelResolutionsMax lengthA 12-second clip costs
StandardSeedance 2.0 Mini480p, 720p15 s10 credits at 720p
FastSeedance 2.0 Fast480p, 720p15 s17 credits at 720p
ProSeedance 2.0480p, 720p, 1080p15 s20 at 720p, 50 at 1080p
UltraSeedance 2.5480p, 720p, 1080p30 s37 at 720p, 93 at 1080p

Standard is what the trend looks like on most feeds. Pick Pro or Ultra when you want 1080p or a clip longer than 15 seconds. Failed renders are never charged. Current prices are on the pricing page.

How the video at the top was made: Claude, frame by frame

The explainer at the top of this page is the opposite of an AI video. Nothing in it was generated by a video model except the sample clip playing on the phone. It follows the method of pdoom-video, an open-source music video for the song “I’m Upping My P(doom)” that its author made with Claude in Claude Code and released under the MIT licence. The official three.js account shared it in late September 2026. Every frame of that video is a pure function of song time, so the live preview and the final export match exactly.

  1. An original song. We wrote the lyrics for this guide, explaining the trend itself, and generated a 58-second, 141 BPM track with Suno.
  2. Timing data. Word-level timestamps for every lyric, a beat grid and downbeats from the audio, and separate vocal and instrumental stems.
  3. The edit as code. Every cut is placed on the beat a lyric line starts on. The edit refers to lyric lines and bars, never to typed-in seconds, so if the song is re-timed the edit follows it.
  4. Scenes in TypeScript. The booth wall, the swinging mic (a damped pendulum), the selfie cards, the prompt that types itself and the phone are all drawn by code. The chorus runs the sample render through duotone, halftone and engraving looks in a three.js shader.
  5. Rendering. Headless Chrome renders each frame, averaging up to 108 sub-frames for real motion blur, and ffmpeg encodes the result with film grain.

Claude wrote the renderer, the scenes and the edit in conversation, checking rendered stills along the way. It is a good illustration of the trend’s two halves: the code-drawn diagram explains the format, and the generated clip on the phone shows how short the path from two selfies to a finished video has become.

Make your own Hotel Lobby video

  1. Open the Hotel Lobby AI generator and add two photos, one person each, or one photo of you both.
  2. Keep the Orange Booth stage selected.
  3. Optionally type what the two of you should rap about. Leave it blank and the bars are about being the best duo in the game.
  4. Render. A standard 12-second clip is ready in about three minutes, in HD with no watermark.

Sources

Questions people ask

What song is the Hotel Lobby AI trend based on?

“Hotel Lobby (Unc & Phew)” by Quavo and Takeoff, released in May 2022. The visual comes from their June 2022 COLORS performance: an orange room with one microphone hanging between them.

When did the Hotel Lobby AI trend start?

The first widely shared AI version was posted on TikTok on September 16, 2026, and the format spread across TikTok and X during the second half of September 2026.

What is the best prompt for a Hotel Lobby AI video?

State the format first (vertical 9:16, one continuous shot), tie each photo to the left or right rapper, describe seamless orange walls and a single silver mic hanging from the ceiling between them, write the lyrics into the prompt per rapper, and end with “no on-screen text, no captions, no logos”.

Which AI model makes Hotel Lobby videos?

Most of them come from audio-native video models such as ByteDance’s Seedance 2.0, which generates the beat, vocals and lip-sync together. Seedance 2.5 adds clips of up to 30 seconds. AI Rap Video uses Seedance 2.0 Mini, Fast and 2.0, plus Seedance 2.5.

Can I use the original Hotel Lobby song in my video?

The track is copyrighted. Some apps offer it through their licensed music libraries, but uploading it elsewhere can get a post muted or claimed. Generating an original rap with your own lyrics avoids the problem.

Can I make a Hotel Lobby video with my pet or with one photo?

Yes. One photo of two people works if they are side by side, and the pet mode keeps your animal’s breed, fur and markings while it raps.

Put your duo in the orange booth

Two selfies, one mic, about three minutes. HD, no watermark, failed renders never charged.