Text and image to video

Turn a prompt or image into a short AI-generated video.

CapCut combines text-to-video, image animation, model choice and editing tools so a generated clip can move directly into a finished timeline.

Available models, output length, audio features and credit costs can change.

Prompt, generate, then finish the result instead of exporting a raw draft
Prompt1
Model2
Generated scene3
Edit4
Three-step workflow

Move from the starting point to a reviewed result.

1

Describe the scene

Write the subject, action, setting, camera and mood—or upload a reference image.

2

Generate a draft

Choose an available model and review motion, consistency and audio.

3

Refine in CapCut

Trim, caption, add music and correct weak sections before publishing.

Where it fits

Use the tool when the task matches the workflow.

Social concepts

Create short hooks, transitions and visual ideas.

Product storytelling

Animate a product image or build a brief scene around a campaign idea.

Creative previsualization

Test a visual direction before a larger production.

Practical check: CapCut changes features, plan benefits, models and limits over time. Use this guide to choose a workflow, then confirm the current details on the official page before paying or publishing.
Before you choose

Questions that matter for this workflow.

What inputs can the generator use?

CapCut presents text and image-based generation workflows.

How long are generated videos?

Limits depend on the current model and product version; verify them in the tool.

Does generation include editing?

The generated clip can continue into CapCut's broader editing environment.

Should generated content be published without review?

No. Check visual artifacts, claims, rights, audio and brand accuracy before use.