Audio workflow
Native audio works best when each requested sound has a visible or plausible source. Separate dialogue from ambience and keep spoken lines short enough for the clip.
- Use real task records
- Do not publish private prompts
- Separate queue time from inference time
- Prefer a short test before a long render
A practical workflow
Write sound in layers: the closest action sound, room or outdoor ambience, then dialogue if needed. State when the scene should be quiet. Avoid several long spoken lines in a five-second clip, because timing pressure competes with the visual action.
- Tie Foley to visible motion
- Keep dialogue short
- Name the acoustic space
- Use silence as an explicit direction
Frequently asked questions
Are these benchmark results?
No. This page is a feature and workflow comparison unless a measurement is explicitly labeled as a real task record.
Does H3 Max generate audio?
Yes. The H3 Max workflow includes synchronized native audio. Describe ambience, dialogue, sound effects, or silence directly in the prompt.
What durations and resolutions are available?
This generator offers 5, 10, or 15 seconds at 480P or 768P for H3 Max. MiniMax H3 remains available separately for 2K final output.
How should I start a new idea?
Begin with a five-second draft, compare a few variations, and only move the strongest direction to a higher resolution or 2K finish.