Animate a portrait from audio
Wan 2.2 A14B Speech-to-Video Turbo uses one reference image and one audio track to create a speaking-character video. The current AI Videoer integration exposes frame count, frame rate, resolution, inference steps, guidance scale, shift, negative prompt, seed, and safety-checker controls.
Output resolution can be set to 480p, 580p, or 720p. Duration is determined by frame count divided by frames per second. For example, the default 80 frames at 16 FPS produces a 5 second clip. Credits are calculated from this duration and the selected resolution before the task is submitted.
Use a clear portrait with an unobstructed face and an audio recording with minimal background noise. A fixed seed can help when iterating on settings, while the negative prompt can describe visual artifacts you want the model to avoid.
