Limited-Time 30% OFF!Get Offer
Category: CreateSee two real reference-to-video results

MiniMax H3 Max Reference to Video

Feed up to 12 reference images, clips, and audio clips into MiniMax H3 Max, then write a prompt that carries their subject, style, motion, or voice into a new 5–15 second scene.

Try Agent modeNEW

Describe your idea in a conversation and keep refining it.

Try Agent
Reference images
Reference videos
Reference audio
Duration (seconds)5s
5s10s15s
Number of outputs

Production preview

The near man says flatly: "Did you hear about H3 Max?" A beat of silence, wind moving across the lot. The far man's mouth pulls into a slow half-smile and he answers, warm and certain: "Dude, I love it." The near man exhales and shakes his head once, almost smiling. Native audio, lip-synced English dialogue, distant traffic hum, wind. No music, no subtitles, no camera movement.
Two real MiniMax H3 Max Reference to Video results

Two solo reference photos, one scripted exchange

Both from this model's own example gallery — the same reference-to-video endpoint the generator above runs, and the same two references, rendered twice.

A parking-lot exchange, built from two solo reference photos

References: Image 1 is a solo photo of the near man; Image 2 is a solo photo of the far man — the two were never photographed together. Prompt: "The near man says flatly: 'Did you hear about H3 Max?' A beat of silence, wind moving across the lot. The far man's mouth pulls into a slow half-smile and he answers, warm and certain: 'Dude, I love it.' The near man exhales and shakes his head once, almost smiling. Native audio, lip-synced English dialogue, distant traffic hum, wind. No music, no subtitles, no camera movement." The result places both men in one generated scene, keeping each one's face and outfit from his own solo photo, and renders the dialogue as native, lip-synced audio in the same pass.

Run a reference set
Reference photo 1 of 2 (the near man, photographed alone) next to the MiniMax H3 Max reference-to-video result that combines him with a second solo photo

Same two references, a second take

Same two solo photos and the same prompt, published as a second take in this model's own example gallery. The framing and delivery both land a little differently — a useful check on how much a result can vary between attempts with identical references and prompt.

Run a reference set
Reference photo 2 of 2 (the far man, photographed alone) next to a second MiniMax H3 Max reference-to-video take built from the same two references
Controls

What a MiniMax H3 Max reference set can hold

  • Up to 9 reference images

    Subject, character, or style references — named in the prompt as Image 1, Image 2, and so on.

  • Up to 3 reference video clips

    Motion or camera references, 2–15 seconds each with a combined length under 15 seconds — named as Video 1, Video 2.

  • Up to 3 reference audio clips

    Voice or sound references, 2–15 seconds each — named as Audio 1, Audio 2. All three kinds are individually optional; mix and match up to 12 files total.

  • 5 to 15 seconds, adaptive framing by default

    Aspect ratio defaults to adaptive — it takes its shape from your references — or pick one of six fixed frames.

  • 480P, 768P, or 1080P

    Same three resolution tiers as this model's text-to-video and image-to-video modes — 480P for a quick check, 768P or 1080P for the keeper.

  • Up to 4 takes per run

    Generate more than one output from the same reference set and prompt on a paid plan; the first output is available on every plan.

Three steps

How to run a MiniMax H3 Max reference set

  1. 1

    Gather your references

    Up to 9 images, 3 video clips (2–15s each), and 3 audio clips (2–15s each) — 12 files at most, and none of the three kinds is required on its own.

  2. 2

    Write the prompt around them

    Name each reference by kind and order — Image 1, Video 1, Audio 1 — and describe the new action. The model reads the references through that ordering, not through file names.

  3. 3

    Set resolution and duration, then run

    480P, 768P, or 1080P; 5 to 15 seconds; adaptive aspect ratio by default. Watch the result, then adjust one reference or one line and run again.

How I use it

Where a reference set earns its slot

Jobs where the point is carrying something specific into a new shot, not starting from a blank prompt.

  • Keep one character consistent in a new scene

    Reference the same face and outfit as Image 1 and write the new action. Carrying the subject's identity over is a design goal, not a pixel-exact guarantee — check the result against your reference.

  • A product shot that has to stay on-model

    A packaging or product reference keeps color, shape, and label steady while the rest of the scene changes around it.

  • Reuse a camera move without reshooting it

    A short reference clip can donate its motion — a pan, a dolly, a spin — to a new subject and setting.

  • Carry a voice into a new line

    A reference audio clip can anchor the voice while the prompt writes new dialogue for the scene.

  • Combine a few anchors in one shot

    Mix a character image, a product image, and a short motion clip in the same prompt — up to 12 files total across all three kinds.

  • Reshoot the same exchange for another market

    Keep the references from a working scene and rewrite only the dialogue in the prompt for a new language pass.

FAQ

Before you build a reference set

Pricing

What a reference-to-video run costs here

480P is 7 credits a second, 768P is 15, and 1080P is 30 — the same table as this model's text-to-video and image-to-video modes. A 5-second 480P run is 35 credits.

Limited-time offer

Enjoy Limited-Time 30% OFF!

07DAY00HOUR00MIN00SEC
Loading…

Build the scene around what you already have.

Up to 12 images, clips, and audio in. One new video out, with your references' subject, style, motion, or voice carried through.