Wan 3.0 Spicy ref2v: A Character Consistency Guide

Wan 3.0 Spicy ref2v: A Character Consistency Guide
Wan 3.0 Spicy ref2v: A Character Consistency Guide

Run the same character through five separate Wan 3.0 Spicy generations and you'll usually get five slightly different faces. ref2v exists to fix that: instead of describing a character from scratch every time, you feed it reference material and it locks the appearance across the output. Here's how the inputs work, what a basic workflow looks like, and what it costs — since this endpoint bills differently from t2v/i2v.

Key Takeaways

  • ref2v accepts up to 10 reference images, 5 reference videos, and 5 reference audio files in a single call to anchor a character's look.
  • It uses a dedicated billing formula: input video duration (capped at 5 seconds) + output video duration, both charged at the same per-second rate as t2v/i2v.
  • prompt_expansion_enable defaults to true — your prompt gets rewritten unless you explicitly set it to false.
  • The endpoint is currently 33% off under an active Launch Week promotion, confirmed directly on the console on 2026-08-28 — deeper than the 15% off rate on record a few days earlier, so reconfirm before locking in a budget.

What Does ref2v Actually Take as Input?

ref2v takes three optional reference arrays on top of the standard prompt: up to 10 images, up to 5 videos, and up to 5 audios. Images are the main lever for locking a face or outfit — feed it two or three consistent stills of the same character from different angles, and the model has something concrete to match instead of re-deriving the look from text each time. Reference video adds motion or style cues stills can't carry; reference audio covers a specific voice or sound bed. None of the three are required — you can call ref2v with just a prompt — but skipping image references is what causes the drifting-face problem.

How Do You Lock a Character Across Shots?

Start with 3-5 stills of the character that already agree on face, outfit, and lighting — inconsistent references teach the model to average them. Load those into images, add a short clip in videos if you need a specific motion or camera style carried over, and set duration, size, and aspect_ratio as usual. Leave prompt_expansion_enable at its default true for a rough first pass; if later takes drift off-script, set it to false so the prompt runs unmodified and every variation traces back to something you actually wrote.

AI Influencer source:sendowl
AI Influencer source:sendowl

The ref2v Billing Formula, With an Example

t2v and i2v bill purely on output duration. ref2v doesn't: the console's Pricing block states it plainly — billable duration = input video duration (up to 5 seconds) + output video duration. Feed it an 8-second reference clip and generate a 10-second output at 720p, and you're billed for 15 seconds (5 capped input + 10 output), not 18 and not 10. A reference clip under 5 seconds costs nothing extra beyond the output; one over 5 seconds is capped, so trimming a longer reference before uploading it has no benefit.

What It Costs Right Now

As of 2026-08-28, all three Wan 3.0 Spicy endpoints — t2v, i2v, ref2v — are priced identically per second, confirmed against each endpoint's console Pricing block. They currently sit at 33% off under an active Launch Week promotion: 480p $0.045/s, 720p $0.09/s, 1080p $0.18/s — deeper than the 15% off rate on record a few days earlier. The console flags this as a live, time-limited promotion, so confirm the current rate on the console before locking in a budget.

AI Influencer consistency image:aiartistatalayserhans
AI Influencer consistency image:aiartistatalayserhans

When ref2v Isn't the Right Tool

If the project sits on a different model family, character consistency isn't unique to Wan 3.0 — Seedance 2.5 has its own reference mechanism, covered in how to keep character consistency in Seedance 2.5 fan edits. This guide only covers the Wan 3.0 Spicy path.

Developer Notes

ref2v sits behind the same request shape as t2v and i2v — swap the model string to alibaba/wan-3.0-ref2v-spicy and add the images/videos/audios arrays. Same API key, same endpoint family covered in the full Wan 3.0 Spicy overview.

Consistency across shots is a spec you feed the model, not a happy accident of a good prompt. Good references in, consistent character out.

Create your free Siray account and test ref2v on your own reference set before scaling up a multi-clip shoot.


Compliance Statement: Siray strictly prohibits all illegal content, with zero tolerance for CSAM. All content generated through Siray must depict legal adult (18+) individuals only. "Uncensored" means no content moderation beyond what the law requires — it never means permission for illegal content of any kind.