Grok Imagine AI Video Generator

Type a prompt or drop in your own photos, then turn them into a short video with sound — in seconds, right in your browser.

8s
10

See what Grok Imagine makes

Motion that obeys physics

A snowboarder carving hard through deep powder: the spray throws where the edge bites, the board flexes under the turn, and nothing melts or slides the way AI motion usually does.

Voice, sound and lip sync

The audio work is the standout. Lips track the speech across face angles and cuts, several speakers hold their own voices, and even the cat lands its meow on the frame - all from the prompt, with close adherence.

A whole sequence, not one shot

The Trojan horse at the gate, the city burning, a figure on the cliff, a girl among the pigs - four cinematic beats in one generation, with the light and the grade holding across every cut.

How it works

1

Describe or upload

Type what you want to see, or hit + to attach up to 7 reference photos.

2

Pick your settings

Resolution, length and aspect ratio — all right there in the prompt bar.

3

Generate & download

Your clip renders in under a minute and downloads as a ready-to-post MP4.

Frequently asked questions

Grok Imagine is the latest video generation model from Grok, xAI's AI assistant. It turns a text prompt — or your own reference photos — into a short video clip with synchronized audio. Everything runs from a single prompt bar: resolution, length and aspect ratio.
It is a short-form video model built for speed: it renders fast, supports both text-to-video and image-to-video, generates synchronized audio natively — sound effects, ambient noise and dialogue — simulates motion with more believable physics, and outputs clips of roughly 1 to 15 seconds at up to 1080p.
Anywhere from 1 to 15 seconds - drag the duration slider in the prompt bar. New clips default to 5 seconds.
Attach your photos, then type @ in the prompt box. A picker appears listing Image 1, Image 2 and so on - pick one to bind that part of your prompt to that photo.
Name the subject, what it is doing and where it is, then add one camera move and one lighting or style cue - for example: a red fox trotting through wet grass at dawn, slow push-in, soft backlight, shallow depth of field. One or two sentences beat a long list of adjectives, and verbs of motion matter most, since those are what the model has to animate.
Your first video is free - every new account gets one complimentary 480p clip of up to 5 seconds. After that each render costs credits by the second: 2 per second at 480p, 4 at 720p and 8 at 1080p, so a 5-second video starts at 10 credits.