Wan 3.0 ref2v: A Consistent Brand Mascot Across Videos
A brand mascot that looks slightly different in every new video reads as unfinished, not flexible. The usual fix — a fresh art brief and a fresh shoot for every ad variant, seasonal refresh, or social clip — is slow, and drift creeps in anyway once several different sessions touch the same character. alibaba/wan-3.0-ref2v takes a set of reference images (plus optional reference video and audio) as identity anchors, so the same mascot can be re-generated into new scenes and actions while holding its look.
Key Takeaways
- ref2v accepts up to 10 reference images, 5 reference videos, and 5 reference audio clips as identity anchors per request.
- There is no negative_prompt field on this endpoint — steer the character through the reference set and a descriptive prompt, not exclusions.
- prompt_expansion_enable defaults to true and rewrites the prompt before generation; turn it off once a mascot's look is locked, so re-writes do not drift the description video over video.
- ref2v bills on its own formula: input video duration (capped at 5 seconds) plus output duration — separate from the flat per-second rate on t2v/i2v.

What ref2v Actually Anchors
A single portrait is not enough reference for a moving character. ref2v's three input arrays cover different parts of identity: images carry the face, costume, and color palette from multiple angles; an optional reference video carries how the mascot moves — a walk cycle, a gesture, a signature pose; an optional reference audio clip carries a voice or sound cue if the mascot speaks or has a jingle. None of the three is required beyond the image set, but stacking them narrows what the model has to guess.
Building a Reusable Reference Set
Treat the reference set as a brand asset, not a one-off upload:
- Capture the mascot from front, three-quarter, and side angles, plus one or two signature poses — six to eight images is usually enough of the 10-image ceiling to leave room for a wardrobe variant.
- Keep size and aspect_ratio consistent across every video generated from the same set; mixing ratios makes later edits harder to match.
- Reuse the exact same image set for every new brief. Swapping even one reference image between videos is what reintroduces drift.
- Write the prompt to describe the scene and action, not the character itself — the reference set is already doing that job.

What This Costs
Wan 3.0 ref2v is currently 30% off as part of Siray's Launch Month pricing: $0.04/s at 480p, $0.08/s at 720p, and $0.16/s at 1080p, down from $0.057, $0.1145, and $0.2285. The billing formula is specific to ref2v — input video duration, capped at 5 seconds, plus output duration — so a longer reference clip does not inflate cost past that cap. Launch pricing is time-limited; check the current rate on the model page before budgeting a full campaign.
How This Differs From Other ref2v Guides
This is a single-endpoint workflow: plain wan-3.0-ref2v alone, keeping one owned brand mascot consistent across marketing videos. It is not the two-model pipeline in Fan Edits Across Endpoints, which pairs Seedance 2.5 and Wan 3.0 ref2v to lock down fan-edit characters. The reference-anchoring mechanism is the same one used in Wan 3.0 Spicy ref2v: A Character Consistency Guide — that guide applies it to adult-content characters on the Spicy endpoint, while this one stays on the plain endpoint for brand marketing.
Put a Mascot on Wan 3.0 ref2v
Wan 3.0 ref2v runs through the same Siray API as every other model in the catalog — one key, one request format.
Create your free Siray account and generate the first video from a locked reference set today.