AIVideoShort-Form Video

Text-to-Video AI for Social: Sora 2, Veo 3 and Beyond

What Sora 2, Veo 3 and other text-to-video models actually produce for Reels, TikTok and Shorts — cost per usable clip, and where it reads as slop.

Dan — Founder, SocialKit10 min read

Text-to-video AI turns a written prompt into moving footage. As of July 2026, the strongest models — OpenAI's Sora 2, Google's Veo 3, Runway, Kling, Hailuo, and a widening set of open-weight alternatives you can run yourself — will reliably hand you a few seconds of convincing video at a time, several of them with synchronised audio baked in. For short-form social, that specific capability makes them a b-roll, hook and concept-teaser tool. It does not make them a "generate my Reel" tool, and the distance between those two ideas is where most of the wasted money and most of the embarrassing posts come from.

This is the hands-on companion to our overview of what AI video does and doesn't do well. That piece maps the category. This one is about the part you have to actually do: prompt, judge, count the cost, label it correctly, and get it onto a feed.

What Current Models Actually Produce

Set expectations against clips, not demos. The showreels are cherry-picked from hundreds of generations; your first afternoon will not look like them.

What you hope forWhat you get, as of July 2026
A 30-second finished videoA clip of a few seconds. Most models generate in roughly the 4–10 second range natively; longer runtimes come from extend features or from stitching shots in an editor
A consistent character across shotsDrift. Faces, clothing and lighting shift between generations unless you anchor with a reference image or a start frame
Usable on-screen textGarbled lettering. Add every word of text in your editor afterwards
Real audioGenuinely good ambience and sound effects on the audio-native models. Dialogue and lip sync are convincing in short bursts and give themselves away over longer takes
Your actual productSomething product-shaped that isn't yours. Generated footage cannot show a thing that exists in the world
Vertical outputDepends on the model and the surface. Some generate 9:16 natively, some are landscape-first and leave you cropping

Two of those deserve emphasis because they quietly decide your workflow.

Shot length dictates edit style. If a model gives you six seconds, your generated footage has to live as cutaways inside something else — a voiceover, a screen recording, a talking head. Trying to build a whole Reel out of eight stitched generations produces the stuttering, disconnected rhythm that viewers read as synthetic within a second or two.

Aspect ratio dictates composition. Prompting a landscape-first model and cropping to vertical will centre-punch your subject and lose the sides of every frame. Check what the model outputs before you write a shot list, and compose against the real target — our Instagram Reel size and TikTok video size references have the safe zones, or browse the full social media sizes list if you're publishing wider.

The Only Cost Number That Matters

Ignore the headline price per generation. The number that determines what generated footage costs you is your keep rate: how many clips you would actually publish out of every ten you generate.

The arithmetic is unforgiving. If one generation costs you a euro in credits and you keep one clip in six, that usable clip cost six euros — plus the fifteen minutes you spent watching the five rejects. Get the keep rate to one in two and the same clip costs two euros and almost no review time. Nothing else in this workflow moves the total as much.

Pricing shapes across the major tools are broadly similar as of July 2026: a consumer subscription tier in the region of twenty dollars a month with a monthly generation allowance, a substantially more expensive pro tier with higher resolution and longer clips, and per-second API pricing if you want to generate programmatically. Cheaper credit-based tools like Kling, Hailuo, Luma and Pika sit under the frontier models on price and, for simple cutaway shots, often close enough on quality that the price difference is the whole argument.

Before you commit to a subscription, run this test: write one real brief from a campaign you're actually shipping, generate ten clips against it, and count the keepers. Do that on two tools. The one with the better keep rate on your briefs is your tool, regardless of which model wins the benchmark threads.

Five Things That Raise a Keep Rate

  1. Prompt a shot, not a story. "Slow push-in on steam rising from a coffee cup on a wooden counter, morning window light, shallow depth of field" beats "a cosy cafe morning". Camera move, subject, light, lens feel. One idea per clip.
  2. Start from an image. Image-to-video, seeded with a still you generated or shot yourself, is dramatically more controllable than pure text-to-video and is the single biggest fix for character and brand drift across a series.
  3. Batch variants of one prompt rather than generating one clip each from ten different prompts. You are buying lottery tickets; buy them on the draw you care about.
  4. Stay away from the known failure modes. Hands doing fiddly things, legible signage, crowds, recognisable brands and logos, and sustained close-ups of a face all burn credits.
  5. Accept the first usable clip. Chasing the perfect version of a two-second cutaway is how a €2 clip becomes a €20 one. The audience is watching it for two seconds on a phone.

Where Generated Footage Genuinely Fits

B-roll under narration. The highest-value use, comfortably. Two- to three-second cutaways behind a voiceover break up a static talking head, and they're too brief for a viewer to lock onto any artefact. This is also the use where a mediocre generation is still fine.

The opening hook. A striking, impossible or arresting first frame buys you the attention you need for the actual content that follows. Generate the visual, write the words separately — our guide to writing hooks that earn the next three seconds covers the copy half of that.

Concept teasers. Announcing something that doesn't exist yet, or visualising an abstract idea — a metaphor, a before/after of a process, a scenario you can't film. Stock footage has always been the fallback here and generated footage is usually a better, more specific fallback.

Faceless and explainer formats. If you already produce faceless content, generated b-roll slots straight into an existing pipeline where nobody expected documentary footage in the first place.

Motion backgrounds for text-led posts. A slow, abstract, on-brand loop behind a quote or a stat is a low-risk, high-reuse asset. Generate a handful once, reuse them for months.

Where It Reads as Slop

The tell is almost never resolution. It's mismatch — footage that implies a claim the footage can't support.

  • Synthetic people talking to camera for thirty seconds. Short bursts pass. Sustained takes don't, and the audience's reaction when they clock it is worse than if you'd shot a wobbly phone video.
  • Anything presented as real. Generated "customer moments", "our team", "our workshop", testimonial-shaped footage. This is the category that costs you trust rather than just reach.
  • Your actual product. If people expect to see the thing they're buying, they need to see the thing they're buying.
  • Trend-reactive posts. Trend windows are measured in a day or two. A generation-and-assembly cycle is slower than picking up your phone.
  • Volume for its own sake. Platforms have tightened monetisation and distribution rules around mass-produced, repetitive content, and posting more of it is the fastest way to find out where the lines are. We went through what the evidence does and doesn't support in will AI content hurt your reach, and the short version is that low-effort output is the risk, not AI involvement as such.

A useful gut check before publishing: if a viewer knew exactly how this clip was made, would they feel misled? Cutaway b-roll under a real voice, no. A generated person recommending your service, yes.

Disclosure: Settle This Before You Publish

Every major platform now has a disclosure surface for realistic synthetic content, and every major model now stamps its output with provenance data whether you use the toggle or not. Assume the footage is identifiable and label it deliberately.

PlatformWhere you disclosePractical note
Instagram / FacebookAI-content toggle in the composer; Meta also applies its own labels when it detects industry-standard AI metadataExpect automatic labelling on generated footage even if you skip the toggle
TikTokAI-generated content toggle; TikTok also reads Content Credentials and labels accordinglyUndisclosed realistic AI content is a policy violation, not a style choice
YouTubeAltered-or-synthetic content disclosure in the upload flowRealism is the trigger. Obvious animation and cosmetic edits are out of scope; realistic depictions of people, places or events are in
LinkedInAutomatic labelling from embedded Content CredentialsLittle for you to do beyond not stripping metadata

Two things creators consistently get wrong. First, watermarks and provenance metadata are separate. Consumer tiers frequently burn a visible watermark into downloads while every tier embeds invisible provenance signals — Google's SynthID, C2PA Content Credentials — that survive re-encoding far better than people assume. Removing a watermark to make a clip look native breaches the tool's terms and doesn't remove the underlying signal. Second, disclosing does not tank your reach. The label is a label, not a penalty flag.

The regulatory direction reinforces this. The EU AI Act's transparency provisions for synthetic media are due to apply from August 2026, with the precise timing of some obligations still under discussion at EU level as of July 2026. Either way, the trend is toward more marking and more disclosure, not less. Our guide to AI content disclosure on social media has a policy template you can adopt as a team standard rather than deciding case by case at 4pm on a Friday.

Slotting It Into a Weekly Workflow

Generated footage is a step in a pipeline, not a pipeline. In practice, for a small team shipping vertical video weekly:

  1. Script and shot list first. Decide which shots you'll film and which you'll generate. Usually: film anything real, generate anything abstract or impossible.
  2. Generate in one batch session. Batching keeps you in the prompting rhythm and stops per-clip perfectionism. Our batch content creation workflow applies here nearly unchanged.
  3. Assemble, caption and colour-match in your editor. Generated clips rarely match your filmed footage out of the box. A quick grade pass makes the seams disappear.
  4. Export one vertical master. Then adapt per platform.
  5. Schedule everything from one calendar.

That last step is where SocialKit fits. Once the clip exists, you compose the post once, write a different caption and hashtag set for each network, and schedule it to Reels, TikTok and Shorts — plus the other eight platforms on the same plan — from one visual calendar, with auto-publish and best-time-to-post recommendations handling the timing. Every plan includes all 11 platforms and unlimited scheduled posts, from €29/month on Solo (€17.40/month billed annually) as of July 2026, with a 7-day free trial.

To be straight about the boundary: SocialKit doesn't reframe, trim or generate video for you. Your export needs to be the right aspect ratio and length before it reaches the scheduler. What it removes is the part after that — three apps, three uploads, three sets of copy retyped into three composers.

If you're new to shipping the same clip across all three vertical surfaces, how the three platforms differ in practice is worth reading before you write the captions, and the per-platform mechanics are covered in scheduling Instagram Reels and planning a YouTube Shorts posting schedule.

Start Here: A Two-Week Test

Rather than subscribing to three tools and hoping, run this:

  • Days 1–2. Take one real campaign brief. Write five generated-shot descriptions from it — camera move, subject, light, feel. No stories, just shots.
  • Days 3–4. Generate all five on two different tools, using the free or entry tier. Count keepers. That's your keep rate, and it's your buying decision.
  • Day 5. Cut two videos: one with generated b-roll under your voiceover, one filmed as normal. Same script, same length.
  • Week 2. Publish both across your vertical surfaces, spaced apart, and compare retention and sends — not likes. Fold the read into whatever short-form video strategy you're already running rather than treating it as a separate experiment.
  • Before every publish. Set the platform AI disclosure toggle. Make it a step in the checklist, not a judgement call.
  • Month 2. If the generated version held retention, build a small library of reusable cutaways and motion backgrounds. If it didn't, you've spent one subscription month learning something most people spend a year assuming.

The honest position as of mid-2026: text-to-video is a real production tool with a narrow, valuable job. It makes your b-roll cheaper and your hooks more striking. It does not make you a video studio, and the creators getting the most out of it are the ones who decided that early. If you want the wider view of where this sits among the other traps, common AI social media mistakes covers the ones that cost the most.

Key terms in this guide