What Wan2.7 is good at
- Image-to-video: First-frame, first+last-frame, and audio-driven generation
- Text-to-video: Pure text prompts with optional audio input and multi-shot narration
- Video continuation: Extend an existing clip with new content guided by a text prompt
- Reference-to-video: Reference both a subject’s visual appearance and vocal timbre; supports up to 5 real-person inputs and multi-character interactions
- Video edit: Edit or replicate videos via text prompts, reference image, or style transfer
Highlights
- Supports up to 5 real-person image inputs for multi-character scenes
- Vocal timbre reference for consistent audio-visual identity
- 3x3 grid-based image generation
- Significant improvements in motion dynamics, stylization, and consistency over Wan2.6
Use it in ComfyUI
Wan2.7 workflows
Run the image-to-video, text-to-video, reference-to-video, and video edit workflows in ComfyUI, locally or on Comfy Cloud