Skip to main content
Wan2.7 is Alibaba’s latest video generation model, now available in ComfyUI via Partner Nodes. It is a comprehensive upgrade over version 2.6 with significant improvements across image quality, audio, motion dynamics, stylization, and consistency. This release brings a fully upgraded multimodal video pipeline directly into your node graph, supporting text, image, audio, and video inputs across five task types.

What Wan2.7 is good at

  • Image-to-video: First-frame, first+last-frame, and audio-driven generation
  • Text-to-video: Pure text prompts with optional audio input and multi-shot narration
  • Video continuation: Extend an existing clip with new content guided by a text prompt
  • Reference-to-video: Reference both a subject’s visual appearance and vocal timbre; supports up to 5 real-person inputs and multi-character interactions
  • Video edit: Edit or replicate videos via text prompts, reference image, or style transfer

Highlights

  • Supports up to 5 real-person image inputs for multi-character scenes
  • Vocal timbre reference for consistent audio-visual identity
  • 3x3 grid-based image generation
  • Significant improvements in motion dynamics, stylization, and consistency over Wan2.6

Use it in ComfyUI

Wan2.7 workflows

Run the image-to-video, text-to-video, reference-to-video, and video edit workflows in ComfyUI, locally or on Comfy Cloud