A TikTok script is a beat-by-beat plan for what gets said and shown in each second of a video — not a screenplay, and not a paragraph you read off a phone. For anything under a minute, the working form is four beats: hook, context, payoff, call to action, each with a word budget you can genuinely speak in the seconds allotted. Write to the budget and the video lands on length. Write prose first and you find out in the edit that your 30-second idea takes 51 seconds to say.
Plenty of tools will generate a script for you. Very few explain the craft underneath, which is what makes the difference between a draft you can film and a draft you rewrite three times. This guide is the structure layer. The opening line itself is covered in our breakdown of what has to happen in the first three seconds, and the narrative shapes that hold attention across a longer video are in the TikTok storytelling guide. Here we deal with the skeleton those two hang on.
Word budgets beat timecodes
Timecodes are how you plan. Word counts are how you actually hit the plan, because you can count words before you film and you cannot count seconds until after.
Measure your own pace once: record yourself talking naturally for thirty seconds about something you know well, then count the words. Double it. That is your per-minute rate, and it is more useful than any average, because delivery pace varies enormously between people and between formats — a demo voiceover runs slower than a rant.
For planning, work from a round number and scale to yours. At roughly 150 words per minute:
| Video length | Speakable words | Realistic sentence count |
|---|---|---|
| 15 seconds | ~37 | 4–5 short sentences |
| 30 seconds | ~75 | 8–10 short sentences |
| 60 seconds | ~150 | 16–20 short sentences |
Two things fall out of this table immediately. First, a 15-second video holds one idea — not one idea plus a caveat. Second, most scripts people write for 30 seconds are 30-second scripts plus an intro they did not budget for. If your draft is over budget, the fix is almost always cutting context, not speeding up delivery.
Which length to aim at in the first place is a separate decision with its own tradeoffs, and our TikTok video length guide walks through choosing based on goal rather than trend.
Word-for-word or bullet points
Both work. Choosing wrongly is what makes people sound either robotic or rambling.
| Script word-for-word when… | Use bullet points when… |
|---|---|
| The video is 20 seconds or less — no room to wander | You are demonstrating something with your hands or on screen |
| Precision matters: prices, specs, steps, anything you would hate to misstate | The material is a personal story you have already told out loud |
| You are filming in a second language | Your scripted reads come out stiff and your unscripted ones do not |
| The script doubles as the caption or as source text for another platform | You are shooting several takes and want variation to choose from |
The hybrid is what most working creators land on, and it is what I would recommend by default: write the first sentence and the last sentence word-for-word, bullet the middle. The hook decides whether anyone watches and the call to action decides whether the watch is worth anything, so those two get exact wording. The middle is where sounding human matters more than sounding precise.
This is close to the opposite of long-form. In our YouTube scripting guide the argument runs toward fuller scripting, because eight minutes of ad-lib drifts and information density has to be managed across the whole runtime. Short-form has no drift budget and no density problem — it has one point and forty seconds.
The four beats, and what each one owes the viewer
Hook (0–3s). Names the problem, the payoff, or the contradiction. No greeting, no logo, no "so I've been getting a lot of questions about". The hook is a promise, and the rest of the script exists to pay it.
Context (3–8s). The single sentence that makes the payoff land — who this is for, why the problem exists, or what you tried first. This is the beat that bloats. If you cannot say why a sentence of context is load-bearing, delete it.
Payoff (8–25s). The actual thing: the step, the number, the before-and-after, the demo. This is where your seconds should go. In a 30-second video, more than half the runtime belongs here or the script is misallocated.
CTA (last 3–5s). One ask, one verb, and it should follow logically from what you just showed. "Comment TEMPLATE and I'll send it" outperforms "follow for more" because it names an action and a reason in the same breath. TikTok gives you several surfaces for this beyond the spoken line — overlays, captions, pinned comments — and our TikTok CTA guide covers how they stack.
Beat budgets by length
| Beat | 15 seconds | 30 seconds | 60 seconds |
|---|---|---|---|
| Hook | 0:00–0:02 | 0:00–0:03 | 0:00–0:03 |
| Context | — (fold into hook) | 0:03–0:08 | 0:03–0:10 |
| Payoff | 0:02–0:12 | 0:08–0:25 | 0:10–0:28, then 0:31–0:50 |
| Reset line | — | — | ~0:28–0:31 |
| CTA | 0:12–0:15 | 0:25–0:30 | 0:50–1:00 |
The reset line in the 60-second version is the one piece of long-form technique that transfers: a short re-hook around the halfway mark — "and the second part is the one nobody does" — that gives a wavering viewer a reason to stay. Going from 30 to 60 seconds means adding a second payoff, never a longer setup.
Four templates worth reusing
Fill the brackets and you have a filmable draft in a few minutes.
1. Single tip, talking head (15–20s)
Hook: "If you [do common thing], stop — here's what to do instead." Payoff: "[The specific alternative.] I do it by [one concrete step]. Last time it meant [specific result you can vouch for]." CTA: "[One ask tied to the tip.]"
Works for any tip you can prove from your own work. It fails when the tip is generic advice anyone could give.
2. Mistake and fix (30s)
Hook: "This mistake cost me [specific consequence]." Context: "I was [situation], and I assumed [wrong assumption]." Payoff: "What actually happens is [mechanism]. So now I [new process], which takes [time] and means [outcome]." CTA: "[Ask that matches the fix.]"
The strongest of the four for cold audiences, because the confession buys attention that a tip has to earn.
3. Demo or screen walkthrough (45–60s)
Hook: "Here's how to [outcome] in [timeframe]." — state the outcome before showing any steps. Context: "You need [one prerequisite]. That's it." Payoff: three steps maximum, each one action plus one reason. Flag the step people get wrong and give it its own beat. Reset: "Step three is where most people quit." CTA: "[Ask.]"
Bullet-point this one and let your hands set the pace. Three steps is the ceiling in under a minute; a five-step process belongs in a carousel or a longer video.
4. Story to lesson (60s)
Hook: "[Specific moment, present tense.] It's 11pm and the client is on the phone." Context: "[What was at stake.]" Payoff A: "[What went wrong.]" Reset: "[The turn.] Then I noticed something." Payoff B: "[What you learned, stated as a transferable rule.]" CTA: "[Ask.]"
This is the template where narrative shape does the heavy lifting — the open loops and transformation arcs in the storytelling guide are the variations worth learning here.
Three passes before you film
- Read it aloud with a stopwatch. Not in your head. Drafts almost always run longer spoken than they feel on the page — that is what the timing pass is for.
- Cut from context, never from payoff. If you are over, the setup sentence goes first, then any adjective doing no work.
- Rewrite the last line. Scripts get written front to back and the CTA gets whatever attention is left. Give it one deliberate pass, and make sure it asks for one thing.
Keep the skeleton, not the script
Individual scripts are disposable. Script skeletons are the asset. When a video performs, strip out the specifics and save what remains — the beat pattern, the hook shape, the type of CTA — as a reusable template in your content bank alongside your evergreen ideas. After a month or two you stop writing from scratch and start filling in shapes that have already worked for your audience.
That is also what makes batching realistic: five skeletons filled in one sitting, filmed back to back in one lighting setup, cut and queued together. Our batch content creation workflow covers the production side of that rhythm.
Once the exports exist at the right vertical spec, scheduling the batch is the cheap part. In SocialKit you upload each master once and schedule it across TikTok and the other platforms you post to from a single composer, customizing the caption and hashtags per network rather than pasting identical text everywhere — ceilings differ, and ours are listed on the character limits reference. The visual calendar shows the whole batch at a glance, and post analytics tell you which script skeleton earned the watch time.
To be straight about the boundary: SocialKit is not a script generator, and it does not edit, trim or reframe video — that happens in your editor before the file reaches us. What it handles is everything after export: calendar, scheduling, auto-publish, best-time-to-post recommendations and analytics. All 11 platforms are included on every plan, from €29/month on Solo (€17.40/month billed annually) as of January 2025, with a 7-day free trial.
Start here this week
- Record thirty seconds of natural speech and count the words. Write your per-minute rate on a sticky note.
- Pick five topics you can prove from your own work — things you priced, fixed, or got wrong.
- Assign each one a template and a length. Aim for three at 30 seconds, two at 60.
- Draft with word budgets in view. Word-for-word on the first and last line, bullets in between.
- Read each aloud with a stopwatch and cut from context until it fits.
- Film all five in one session, export at the standard vertical spec, and schedule the batch.
- A week later, check where retention drops. Before five seconds means the hook; mid-video means the context beat is still too long. Fix that beat in the skeleton, not just in the next script.
The scripts get faster because the skeletons accumulate. That is the whole compounding mechanism — nothing about it requires a generator.