Wan 3.0 ref2v: A Consistent Brand Mascot Across Videos

Wan 3.0 ref2v: A Consistent Brand Mascot Across Videos
Wan 3.0 ref2v: A Consistent Brand Mascot Across Videos

A brand mascot that looks slightly different in every new video reads as unfinished, not flexible. The usual fix — a fresh art brief and a fresh shoot for every ad variant, seasonal refresh, or social clip — is slow, and drift creeps in anyway once several different sessions touch the same character. alibaba/wan-3.0-ref2v takes a set of reference images (plus optional reference video and audio) as identity anchors, so the same mascot can be re-generated into new scenes and actions while holding its look.

Key Takeaways

  • ref2v accepts up to 10 reference images, 5 reference videos, and 5 reference audio clips as identity anchors per request.
  • There is no negative_prompt field on this endpoint — steer the character through the reference set and a descriptive prompt, not exclusions.
  • prompt_expansion_enable defaults to true and rewrites the prompt before generation; turn it off once a mascot's look is locked, so re-writes do not drift the description video over video.
  • ref2v bills on its own formula: input video duration (capped at 5 seconds) plus output duration — separate from the flat per-second rate on t2v/i2v.
ref2v with Wan Spicy Image:A1 Art
ref2v with Wan Spicy Image:A1 Art

What ref2v Actually Anchors

A single portrait is not enough reference for a moving character. ref2v's three input arrays cover different parts of identity: images carry the face, costume, and color palette from multiple angles; an optional reference video carries how the mascot moves — a walk cycle, a gesture, a signature pose; an optional reference audio clip carries a voice or sound cue if the mascot speaks or has a jingle. None of the three is required beyond the image set, but stacking them narrows what the model has to guess.

Building a Reusable Reference Set

Treat the reference set as a brand asset, not a one-off upload:

  • Capture the mascot from front, three-quarter, and side angles, plus one or two signature poses — six to eight images is usually enough of the 10-image ceiling to leave room for a wardrobe variant.
  • Keep size and aspect_ratio consistent across every video generated from the same set; mixing ratios makes later edits harder to match.
  • Reuse the exact same image set for every new brief. Swapping even one reference image between videos is what reintroduces drift.
  • Write the prompt to describe the scene and action, not the character itself — the reference set is already doing that job.
Wan Spicy Image to video image:live 3d Vtuber maker
Wan Spicy Image to video image:live 3d Vtuber maker

What This Costs

Wan 3.0 ref2v is currently 30% off as part of Siray's Launch Month pricing: $0.04/s at 480p, $0.08/s at 720p, and $0.16/s at 1080p, down from $0.057, $0.1145, and $0.2285. The billing formula is specific to ref2v — input video duration, capped at 5 seconds, plus output duration — so a longer reference clip does not inflate cost past that cap. Launch pricing is time-limited; check the current rate on the model page before budgeting a full campaign.

How This Differs From Other ref2v Guides

This is a single-endpoint workflow: plain wan-3.0-ref2v alone, keeping one owned brand mascot consistent across marketing videos. It is not the two-model pipeline in Fan Edits Across Endpoints, which pairs Seedance 2.5 and Wan 3.0 ref2v to lock down fan-edit characters. The reference-anchoring mechanism is the same one used in Wan 3.0 Spicy ref2v: A Character Consistency Guide — that guide applies it to adult-content characters on the Spicy endpoint, while this one stays on the plain endpoint for brand marketing.

Put a Mascot on Wan 3.0 ref2v

Wan 3.0 ref2v runs through the same Siray API as every other model in the catalog — one key, one request format.

Create your free Siray account and generate the first video from a locked reference set today.