Text-to-video workflow

H3 Max Text to Video — Prompt, Generate and Iterate Fast

Text-to-video is the fastest route from an idea to a moving shot. Start with a concrete subject and action, add one camera move, then describe the light, style, and sound that make the shot legible.

Direct a new scene

H3 Max Fast + MiniMax H3 2K

Sign in to view credits
Choose a model
0/2000
Enhance for H3 Max
Resolution
Duration
Aspect ratio
Generate variations

Single generation is available to every verified user.

This generation uses 40 credits

Native audio is included for both models.

Demo videoH3 Max · 768P · 10s · Native audio
No API keyPrivate by defaultFailed generations refundedCost shown before generation

A prompt structure that stays coherent

Use this order: subject and action; environment; camera and lens behavior; light and visual treatment; synchronized sound. The order is not magic, but it makes omissions and conflicts easier to spot.

  • One main action per short clip
  • Use 9:16 for vertical social work
  • State “locked camera” when no movement is wanted
  • Describe audio as part of the scene

Iterate before you upscale

Change one variable at a time. If motion is wrong, keep the visual style and revise the action. If framing is wrong, keep the action and revise the camera. Once the prompt is stable, switch to 768P or move the take to MiniMax H3 2K.

Frequently asked questions

Which aspect ratios can I choose?

Text mode offers 21:9, 16:9, 4:3, 1:1, 3:4, and 9:16.

Does H3 Max generate audio?

Yes. The H3 Max workflow includes synchronized native audio. Describe ambience, dialogue, sound effects, or silence directly in the prompt.

What durations and resolutions are available?

This generator offers 5, 10, or 15 seconds at 480P or 768P for H3 Max. MiniMax H3 remains available separately for 2K final output.

How should I start a new idea?

Begin with a five-second draft, compare a few variations, and only move the strongest direction to a higher resolution or 2K finish.