Funk Harbor Rum had a finished photo shoot and didn’t capture any video. We built them a 15-second vertical ad for Meta without putting a camera back on set.
Funk Harbor is a Jamaican rum founded by Aljamain Sterling, a three-time UFC champion whose parents emigrated from Jamaica to New York. The brand had a strong story, a gold medal from the Wine and Spirits Wholesalers of America, and a set of photographs that did the product justice.
What it did not have was motion. Meta’s cheapest reach right now sits in vertical video, and a still image campaign leaves most of that on the table.
Booking a video shoot means talent, location, crew and a day of everyone’s time. The photographs already contained everything the ad needed. The question was whether we could get them moving without the result looking obviously machine-made.
The source material. Four photographs, no video.
Generative video is cheap to run and expensive to run badly. The failure mode is obvious once you have seen it: someone types a prompt, gets something close enough, and ships it.
We wrote every shot out in full and sent it to the client for written approval before we generated anything. Camera move, what the subject does, how the light behaves, and an explicit list of what must not happen. A rejected shot description costs nothing. A rejected generation costs a generation and a round trip.
We also generated more than we needed and threw most of it away. One shot took three attempts before the read was right. That is the job. The model produces options and somebody has to have an opinion about them.
Video models rewrite packaging typography. Left alone, they will melt a wordmark into something that looks almost right, which is worse than something that looks clearly wrong.
Every shot description stated that the label stays sharp, legible and unchanged, and then said it again in the negative. We checked the output frame by frame at full resolution, and chose which section of each clip to use partly on which section held the type best.
Meta is strict about depicting alcohol consumption. Plenty of agencies discover this at the approval stage and go back to re-edit.
We built it in at the start. Nobody brings a glass to their lips at any point in the fifteen seconds. Not in the shot descriptions, not in the footage, not in the cut. The talent holds a glass, turns it in the light, sets it down.
We handled the voiceover the same way. A first-person script was the stronger piece of writing, but we didn’t use it: we generated the read rather than recording Aljamain, then shifted the narration to third person and found a voice that fit the scene.
Four clips at five seconds each is twenty seconds of footage for a fifteen second slot. The quick version trims every clip by the same amount and moves on.
We analysed the music first, found its tempo, and placed every cut on a beat. The shot changes land at an even musical interval rather than at arbitrary timecodes, which is most of the difference between an edit that feels standard and one that feels made.
Which part of each clip to use was decided individually. The back half here, the final second there, the head of the clip elsewhere. Different answers to the same question, because the footage served different purposes.
Most of a paid social feed plays muted. Captions are not an accessibility box to tick, they are how the ad gets read.
The type is set in a letterspaced serif chosen to echo the lettering on the bottle, so the words on screen belong to the same brand as the object in the frame. Nothing bounces or pops word by word. That style reads as amateur content and would undercut everything else about the spot.
Placement took some work. Platform interface covers a significant band at the top and bottom of a vertical frame, and clearing it is only half the problem, because the subject occupies the middle. Our rule ended up simple: captions go where the subject is not. On the shots featuring the founder, type sits at centre frame across his chest, clear of his face. On the product shots, it moves to the top, clear of the bottle and the box.
We verified every placement against the actual frames, at the start and end of each shot, because all four push in and the subject grows while they run.
Caption position changes with each shot so the type never covers the subject or the product.
The campaign is live and performance data is still coming in. We will update this page with results once there is enough spend behind it to say anything honest.
What we can report now is the production outcome. A finished, platform-ready vertical ad with voiceover, licensed music, captions and broadcast-standard audio levels, built entirely from photographs the client already owned. No shoot day, no travel, no crew.
If you have a library of photography and no video, you have more than you think. The constraint is rarely the source material. It is whether anyone is willing to make several hundred small decisions about it.
Generative tools did the rendering on this project. They did not decide where to cut, which frame to hold, what the ad should say, where the type belongs, or what using a real person’s voice was. That is the work, and it is the part worth paying for.
Ready to skip the agency BS and get results? We’re taking on new clients who want straight-shooting digital marketing that actually drives growth.
And hey, if you’re in Austin, TX and want to talk strategy over a pint, we’re down for that too.