MiniMax H3 Prompt Guide:
Examples, Workflow & Best Practices
Structure high-performing briefs for MiniMax H3. Learn text-to-video, image-to-video identity preservation, native stereo audio cues, and artifact-free camera movement.
Build Your MiniMax H3 Brief
Structure of a Production-Ready Brief
MiniMax H3 interprets text, image, video, and audio signals in a single unified context window. Assigning explicit roles to each section prevents identity drift and unwanted morphing.
Define Invariants
Name clothing, facial structure, skin texture, material finish, or logo placement that must remain rock-solid across frames.
Block the Shot
Separate physical subject motion from camera movement (e.g. 'subject turns head right while camera pans left').
Spatial Sound Cues
H3 generates native stereo sound. Specify left/right spatial placement, ambience, and music rhythm cues.
Assign Reference Roles
When providing image/video files, explicitly state: 'Image 1 provides face identity; Video 1 provides camera velocity'.
Stability Rules
State features that must NOT change during motion, such as brand logos, eye color, background architecture.
Negative Filtering
List artifacts to suppress: flickering lighting, extra fingers, identity morphing, sudden camera jumps.
Tested MiniMax H3 Prompt Templates
Cyberpunk Neon Rain Alley
SUBJECT: Cyberpunk detective in dark trench coat standing under rain.
ACTION: Pulls glowing neon lit holographic visor down over eyes.
ENVIRONMENT: Wet asphalt street reflecting pink and cyan neon signs, heavy rain drops.
CAMERA: Low angle close-up pan forward with anamorphic lens glare.
AUDIO: Heavy raindrops splashing, ambient synth soundscape, crisp visor click.
PRESERVE: Trench coat color, facial silhouette, rain density.
AVOID: Face distortion, blurry neon reflections, sudden lighting jump.
Character Facial Motion Transfer
SUBJECT: Character identity defined by Image 1 reference input.
ACTION: Subject slowly smiles, turns head 45 degrees to the left, blinks naturally.
ENVIRONMENT: Studio lighting with soft rim light matching Image 1 color temperature.
CAMERA: Static medium portrait shot at 50mm f/1.8 depth of field.
AUDIO: Soft ambient breath sound, warm gentle acoustic melody.
REFERENCES: Image 1 provides exact facial identity, hair texture, and outfit.
PRESERVE: Face structure, eye color, skin tone, clothing pattern from Image 1.
AVOID: Morphing nose, flickering eyes, floating hair artifacts.
Matte Perfume Bottle Reveal
SUBJECT: Frosted glass perfume bottle with engraved gold logo.
ACTION: Condensation water drops bead and slide down cold glass surface smoothly.
ENVIRONMENT: Dark dramatic studio backdrop with golden spotlight rising from behind.
CAMERA: Slow macro zoom tracking from bottle base up to golden cap reveal.
AUDIO: Soft glass clink sound, ambient ocean breeze soundscape, warm chime beat.
PRESERVE: Gold logo typography alignment, glass transparency, drop physics.
AVOID: Logo blurring, glass distortion, unnatural liquid jump.
Desert Sunset Horizon Flyover
SUBJECT: Endless red sand dunes stretching toward glowing sun.
ACTION: Heat waves shimmer gently off sand ridges as wind sweeps dust clouds across dunes.
ENVIRONMENT: Golden hour sunset sky with deep orange and purple gradient clouds.
CAMERA: Smooth drone flyover moving forward at steady speed 2 meters above dunes.
AUDIO: Wind howling in stereo left-to-right, low cinematic bass rumble, subtle sand rustle.
PRESERVE: Sand dune ridge geometry, sun position on horizon, cloud gradient.
AVOID: Camera shaking, dune flickering, sudden color shifts.
Frequently Asked Questions
What is the best prompt structure for MiniMax H3?
The most reliable prompt structure separates instructions into clear labeled blocks: SUBJECT, ACTION, ENVIRONMENT, CAMERA, TIMING, AUDIO, REFERENCES, PRESERVE, and AVOID. Structuring prompts this way reduces hallucinations and keeps identity stable across 15-second clips.
How do I prevent face morphing in MiniMax H3 Image-to-Video?
Use an explicit REFERENCES block naming your reference image (e.g. 'Image 1 provides face identity'), state the exact facial features in PRESERVE, and add 'facial distortion, morphing eyes, nose shifts' under AVOID.
How does MiniMax H3 native stereo audio generation work?
Unlike video models that generate silent video, H3 supports native stereo sound generation within the same context window. You can describe spatial sound placement in the AUDIO block (e.g., 'footsteps panning left-to-right with low synth bass').
Is MiniMax H3 prompt syntax compatible with Hailuo 01?
While both MiniMax H3 and Hailuo 01 share underlying model lineage, H3 introduces expanded multimodal context supporting multi-reference roles and native audio directives in a single brief.