Clipset Clipset
← All comparisons
Facts checked 9 min read

Hook, Body, CTA: The Modular Video Ad Framework (With Templates and Tools)

What is the Hook, Body, CTA framework?

Hook, Body, CTA is a modular video ad structure that splits an ad into three self-contained segments: a hook (the first one to three seconds, whose only job is to stop the scroll), a body (the case for the product: problem, demo, proof), and a CTA (the call to action, telling the viewer what to do next). Because each segment is recorded to stand on its own, any hook can be paired with any body and any CTA. That interchangeability is what lets you turn a small number of clips into a large number of testable ads.

Most people who run video ads already use something like this informally. The difference between “sort of modular” and actually modular is whether the segments swap cleanly. This page covers each segment, the swap rule, templates, a worked example, recording tips, testing, and the tools that assemble the combinations.

The three segments

Hook (0 to 3 seconds)

The hook exists to earn the next three seconds. It is not the pitch. It is a pattern interrupt, a claim, a question, or a visual that makes the viewer pause. Good hooks are specific and slightly uncomfortable: a price objection, a mistake the viewer is making, a result stated flatly.

One hard rule: the hook cannot set up anything the body has to pay off. “Here are three reasons” is a bad modular hook because it obligates the body to deliver three. “I stopped buying this after I found out what’s in it” works because any body can follow it. More in how to write video ad hooks that convert and video ad hooks that stop the scroll.

Body (roughly 10 to 30 seconds)

The body makes the case. Common shapes: problem-agitate-solve, demo, proof (before/after, testimonial), and objection handling (“I thought it would be X, but…”).

A body should open on a full sentence that does not depend on the hook, and end without hinting at a specific CTA. “So here’s what I’d do” is a bad ending because it implies a particular next step. Ending on the benefit (“…and I haven’t had that problem since”) lets any CTA follow.

CTA (3 to 6 seconds)

The CTA tells the viewer what to do: shop the link, use the code, try it free, follow for more. Modular CTAs vary in urgency and offer, not in product reference. Record them at the body’s energy level so the cut is not jarring.

Why modularity matters (the no-cross-reference rule)

A modular video ad structure only works if segments are interchangeable. That means:

  1. No backward references. A body cannot say “like I said” or “as you just saw.” It does not know which hook preceded it.
  2. No forward setups. A hook cannot promise “three tips” or “the shocking part at the end.” It does not know which body follows.
  3. No tonal cliffs. A whispered hook cut to a shouting body is technically modular and practically unusable.
  4. Consistent identity. Same person, framing, and product name across all segments.

Follow those and the arithmetic takes over. With H hooks, B bodies, and C CTAs, you get H × B × C finished ads. 3 × 2 × 2 is 12. 10 × 5 × 3 is 150. Clips grow linearly; testable ads grow multiplicatively. That is the economic argument for the framework, covered further in how to make 100 video ad variations.

Hook, Body, CTA template

Fill-in templates for each segment. The blanks are deliberately generic so results stay swappable. A fuller version lives in the modular UGC ad script template.

Hook templates

Type Template
Objection “I almost didn’t buy [product] because of [objection].”
Mistake “If you’re still [common behavior], stop.”
Result “[Specific outcome] in [timeframe]. Not kidding.”
Question “Why does nobody talk about [problem]?”
Contrast “[Old way] vs. [product]. Watch.”

Body templates

Shape Template
Problem-Solve “I used to deal with [problem] every [frequency]. [What I tried that failed]. Then I started using [product]. [How it works]. [Result].”
Demo “This is [product]. [Step one]. [Step two]. [What you see happen]. [Result or feeling].”
Proof “Before: [state]. After [time] with [product]: [state]. [One credible detail].”
Objection “I assumed [product] would be [assumption]. It’s actually [reality]. [Example].”

CTA templates

Type Template
Direct “Link’s below. Grab yours.”
Offer “Use code [CODE] for [discount] this week.”
Low-commitment “Try it free and see for yourself.”
Social “Follow for the full breakdown.”

Worked example: 3 hooks, 2 bodies, 2 CTAs = 12 ads

Illustrative product: a reusable coffee filter. Here are the seven clips.

Hooks

  • H1: “I stopped buying paper filters and my coffee got better. Not worse. Better.”
  • H2: “If you’re still throwing away a filter every single morning, watch this.”
  • H3: “Why does nobody talk about how much paper filters cost per year?”

Bodies

  • B1 (Problem-Solve): “I used to buy a box of filters every month and still run out on the worst possible morning. I tried a metal mesh one and it made sludge. This one is a fine steel weave, so it drains like paper but nothing goes in the trash. Rinse, dry, done.”
  • B2 (Demo): “This is the filter. Drops into any standard cone. Grounds go in, water goes over. See how clean the pour is? No grit in the cup. When it’s done, tip it out, rinse for five seconds, ready for tomorrow.”

CTAs

  • C1: “Link’s below. Grab yours.”
  • C2: “Use code MORNING for 15% off this week.”

Every hook promises nothing. Every body opens fresh and ends on a benefit. Every CTA is product-neutral. The twelve finished ads:

Hook Body CTAs Ads
H1 B1 C1, C2 2
H1 B2 C1, C2 2
H2 B1 C1, C2 2
H2 B2 C1, C2 2
H3 B1 C1, C2 2
H3 B2 C1, C2 2

Seven clips, maybe fifteen minutes of recording, twelve distinct ads. Add two hooks and one CTA and you are at 5 × 2 × 3 = 30 without touching the bodies.

Recording tips for clean swaps

The framework fails at the seams. These habits keep the cuts invisible:

  • Record each segment as its own take. Do not record a full ad and slice it later; the delivery carries over context.
  • Start and end on silence, then trim to the first and last word in post.
  • Reset your energy per segment. Hooks run hotter. Bring the first line of each body up to meet them.
  • Same setup for the whole session. Same lens, distance, light, mic, and background.
  • Same aspect ratio and resolution. Pick one (9:16 for paid social) and shoot everything in it. Mixed framing is the most common reason combinations look stitched.
  • Name files by slot and idea. hook-02-still-throwing-away.mp4 is searchable among fifty files and keeps your combinator’s manifest readable.
  • Overshoot hooks. They are cheap to record and vary most in performance.

How to test Hook, Body, CTA ads

The framework’s real payoff is clean testing. Change one slot at a time.

  1. Hold body and CTA constant, swap hooks. Launch every hook against your best-known body and CTA. The winner is your new control.
  2. Hold the winning hook, swap bodies. Now you learn whether demo beats proof for this audience.
  3. Hold hook and body, swap CTAs. Usually the smallest effect, so test it last.
  4. Recombine the winners into the baseline for the next round.

Because every ad is a known combination of parts, you are not guessing why one beat another. See how to test Facebook ad creative and how many hooks to test per ad.

Tools for assembling the combinations

Recording modular clips is the easy part. Rendering the combinations is where teams stall, because a timeline editor makes you build each ad by hand. Twelve is tolerable. One hundred fifty is not.

The tool category built for this is the video combinator: software that takes clips per slot and renders every permutation as a separate file. Two kinds exist.

Local desktop. Clipset is a macOS and Windows app that does exactly this. Drop hooks, bodies, and CTAs into their slots (or define custom sections), pick an aspect ratio, and it renders every combination as its own file with a CSV manifest of which parts are in each. $49 one-time, free trial, runs entirely on your machine, no render limits. It does not generate footage, add captions, or edit on a timeline.

Cloud. Sovran, HookScale, and InfiniteAds offer combinators as web apps. Sovran is a subscription ($99 to $399 per month for 50 to 450 renders per month) that also bundles AI voiceovers, captions, a timeline editor, and team seats. HookScale pairs its Video Combinator with an AI image generator and ZIP export; InfiniteAds offers a free first combination and automatic 9:16 formatting. Neither publishes pricing. See Clipset vs. Sovran and Clipset vs. HookScale.

Whichever you pick, the framework is the same. The tool just removes the assembly step.

Frequently Asked Questions

How long should each segment be?

Hooks: one to three seconds, up to five for a visual hook. Bodies: ten to thirty seconds depending on how much proof you need. CTAs: three to six seconds. Fifteen to forty seconds total is typical for paid social.

Can I use more than three segments?

Yes. Some teams add a “proof” slot between body and CTA, or split the body into “problem” and “solution.” Every slot must still be self-contained. Each added slot multiplies the count: 5 × 3 × 3 × 2 is 90 ads from 13 clips. Clipset supports custom sections for this.

Does Hook, Body, CTA work for non-UGC ads?

Yes. It works for any video with a beginning, middle, and end, including product b-roll, screen recordings, and motion graphics. No presenter required; what matters is a hard start and end with no cross-references.

What if a hook only makes sense with one body?

Then it is not a modular hook. Rewrite it so it works anywhere, or keep it out of the set. Every included hook gets paired with every body, so one that fits only a single body will produce awkward ads with the others.

How many combinations should I actually launch?

Render them all, launch a slice. A common first round is every hook against one body and one CTA, which isolates the hook effect and keeps spend concentrated. The rest of the rendered set is your queue for later rounds.