AIScriptwritingVideo

How to Script Short-Form Videos With AI (Reels, TikTok)

Prompt AI for the hook-value-CTA skeleton, then write the specifics and delivery yourself. A 15-second beat system for Reels, TikTok and Shorts.

Dan — Founder, SocialKit8 min read

AI-assisted short-form scripting means handing the model the structural work — the hook-value-CTA skeleton, the beat breakdown, three alternate angles on the same idea — and keeping the parts that make the video yours: the specific example, the number you can personally vouch for, the way the line sounds coming out of your mouth. The model is a fast outliner, not a writer of your lines. Used that way, a 45-second script that used to eat twenty minutes takes about four, which is what makes batch-scripting five videos in one sitting realistic.

This is deliberately the opposite approach to our guide on writing a YouTube script by hand, which is about long-form craft: earning context, pacing information density over eight minutes, planting re-engagement beats. Short-form video is a different animal. There is no room to earn anything. The structure is rigid enough that a model can produce a usable skeleton in seconds, and rigid enough that the skeleton is the least valuable part.

What AI is genuinely good at here — and what it wrecks

Split the job in two before you open a prompt box.

AI does the scaffolding well. Beat structure, the shape of a hook, ten angles on a topic you have looked at from one, condensing a rambling paragraph into a spoken sentence, and — the underrated one — telling you when your idea has too many ideas in it. Ask a model to compress a 60-second concept into 30 seconds and it will cut ruthlessly in a way most creators cannot cut their own work.

AI wrecks specificity and delivery. Left alone it produces the fully generic script: "In this video I'll show you three tips to grow your account." It invents statistics if you let it. It writes in balanced, tidy sentences nobody says out loud. And it has no access to the thing that actually makes short-form work — the client who called you at 11pm, the invoice you got wrong, the exact price you paid.

So the rule I hold to: the model proposes structure, you supply proof and voice. Every prompt below is built to keep that line clean.

Script in 15-second beats

Short-form retention is decided in blocks, not paragraphs. The practical unit is roughly 15 seconds — long enough to land one point, short enough that a viewer's attention doesn't wander mid-thought. Write to a beat grid and your scripts stop sprawling.

For a 45-second video:

TimecodeBeatJob
0:00–0:03HookName the problem or the payoff. No greeting, no logo, no "hey guys".
0:03–0:15SetupEstablish stakes and promise what's coming. One sentence of context, max.
0:15–0:30ValueThe actual thing. A demo, a number, a step, a before/after.
0:30–0:45Payoff + CTAClose the loop opened in the hook, then one specific ask.

Two rules hold this together. First, every beat has to end with a reason to stay — a partial answer, an unresolved comparison, a "but here's the part that broke". Second, the hook is not a beat you write last as decoration; the whole script exists to cash the promise the hook makes. Our breakdown of what has to happen in the first three seconds covers the mechanics, and the hook formulas library gives you the underlying patterns worth feeding into a prompt as named angles rather than hoping the model invents them.

If you want a 30-second video, drop the setup beat. If you want 60, add a second value beat — never a longer setup. Padding the front is the single most common way short-form scripts leak viewers, and the retention graph shows it plainly, as our guides to TikTok watch time and Reels retention both dig into.

The base prompt

Start every script with the same skeleton request, then swap in the format-specific block below it.

You are helping me script a vertical short-form video for [platform]. Topic: [one sentence]. Audience: [specific — not "small businesses" but "solo wedding photographers who book 15 weddings a year"]. My angle: [the non-obvious thing I believe].

Write a 45-second script on this beat grid: 0–3s hook, 3–15s setup, 15–30s value, 30–45s payoff and CTA. Rules:

  • Spoken language only. Short sentences. Contractions. No sentence longer than 14 words.
  • No invented statistics, studies, or results. Where a specific number belongs, write [NUMBER] and I'll fill it.
  • No greeting, no "in this video", no sign-off.
  • Mark each beat with its timecode and a rough word count.
  • Give me three different hooks for the same script, from three different angles.

Then list what you'd need from me to make it specific.

That last line does more work than the rest of the prompt combined. The model comes back asking for the client story, the price, the mistake — which is exactly the material you were going to skip. For a wider set of reusable prompt structures, our prompt frameworks for social media piece covers the same constrain-then-fill pattern applied to other formats.

A prompt block per format

Talking head

Add:

This is one continuous piece to camera. Write it to be read aloud by me, not performed. Include a hard cut every 4–6 seconds — mark each with [CUT] where the sentence naturally breaks, so the edit can remove pauses without breaking a sentence. Suggest one on-screen text line per beat, six words or fewer.

Talking head lives or dies on cut rhythm. Marking cuts in the script means you shoot in short takes and never have to rescue a rambling one in the edit.

Voiceover over b-roll

For narration laid over b-roll, add:

Write this as narration over footage. Output two columns: VO line, and the shot that plays under it. Each shot must be something I can film on a phone in under two minutes — no stock footage, no drone, no studio. Where the VO makes a claim, the shot must show it, not decorate it.

The "shows it, not decorates it" constraint is what stops you getting a script whose visual column reads "person typing on laptop, smiling" for all four beats.

Tutorial or screen recording

Add:

This is a screen recording of [tool/process]. Write the VO as numbered steps. Each step: one action, one reason. State the outcome in the hook before showing any steps. Flag the single step where people usually get it wrong and give it its own beat. Keep total steps to three or fewer.

Three steps maximum is the discipline. A five-step tutorial in 45 seconds is a blur nobody can follow, and it belongs in a carousel or a long-form video instead.

The human pass, in four minutes

Never film the model's draft. Run it through this:

  • Replace every [NUMBER] and [EXAMPLE] with something true. If you can't fill one honestly, cut the sentence rather than softening it into a vague claim.
  • Read it out loud once. Anywhere you stumble, rewrite. Anywhere you'd never say the word — "leverage", "utilize", "in order to" — swap it for the word you'd actually use.
  • Time it. Read at your real speaking pace with a stopwatch. Most 45-second AI drafts run 55 seconds when spoken. Cut from the setup, never the value.
  • Rewrite the first line yourself. Pick the strongest of the three generated hooks as raw material, then sharpen it by hand — the same filter-and-sharpen loop described in our AI hook generation workflow. The opening line is the one sentence in the script that is worth your full attention.
  • Check the CTA is singular and specific. "Comment SETUP and I'll send the checklist" beats "follow for more". One call to action, one verb.

Batch-script five, film once, schedule everywhere

The reason to make scripting fast is that it unlocks batching. A single 90-minute session realistically covers: 20 minutes generating and sharpening five scripts, 40 minutes filming all five back to back in one lighting setup and one outfit, 30 minutes cutting and scheduling. That's a fortnight of short-form from one block, which is the cadence question our short-form video strategy guide treats as the actual growth variable. The wider mechanics of working this way are in our batch content creation workflow.

Export one clean vertical master per video at the standard canvas — the specs are on our Reel size reference — and the distribution step gets cheap. In SocialKit you upload that master once and schedule it to Reels, TikTok and YouTube Shorts from a single composer, customizing the caption and hashtags per platform rather than posting identical text three times. Caption ceilings differ per network; ours are listed on the character limits reference and checked as you type.

To be straight about the boundary: SocialKit doesn't edit, trim or reframe video — that happens in your editor before the file reaches us. What it handles is everything after the export: one visual calendar for the batch, best-time-to-post recommendations per platform, auto-publish, and post analytics so you can see which of the five hooks earned the retention. All 11 platforms are on every plan, from €29/month on Solo (€17.40/month billed annually) as of September 2024, with a 7-day free trial.

Start here this week

  1. Pick five topics you can prove — five specific things you've done, priced, fixed or got wrong.
  2. Run the base prompt once per topic, with the format block that matches how you'll shoot it.
  3. Do the four-minute human pass on each: fill the blanks, read aloud, time it, rewrite the hook.
  4. Film all five in one sitting. Same setup, same lighting, script on a phone just under the lens.
  5. Cut, export one 9:16 master each, and schedule the batch across Reels, TikTok and Shorts.
  6. After a week, look at where each video's retention drops. If it's before five seconds, the hook is the problem. If it's mid-video, the setup beat is too long — and that's the exact instruction you tighten in the next round of prompts.

The prompt gets better every cycle because you're feeding it what your own retention data taught you. That loop, not the model, is the thing that compounds.