
The Wan 2.5 API produces video and audio in one pass, so voice, sound, and lip-sync line up without a separate step. Built on Alibaba's Diffusion Transformer architecture, it covers text and image to video from 480p to 1080p at 5 or 10 seconds, with reliable sync even for Chinese prompts. Reach it through one key on Atlas Cloud, alongside 300+ models.
Wan 2.5 is developed by Alibaba. Atlas Cloud (operated by Atlas Cloud AI LLC) provides access to it and does not own it. All trademarks belong to their respective owners.
Atlas Cloud provides you with the latest industry-leading creative models.
Match each job to the right Wan 2.5 option: Fast tiers for quick, high-volume drafts, standard tiers for final audio-synced video, and image modes for stills, across text and image on Atlas Cloud.
| Variant | Description |
|---|---|
| Wan 2.5 T2V API (Text To Video) | Wan 2.5 T2V API turns a written prompt into a cinematic clip with synchronized voice, sound, and lip-sync generated in the same pass. It reads camera direction such as pans, tilts, and zooms, and outputs from 480p to 1080p at 5 or 10 seconds. This mode suits scripts, ad concepts, and any scene that needs sound without a separate audio step. |
| Wan 2.5 T2V Fast API (Text To Video Fast) | Wan 2.5 T2V Fast API prioritizes lower latency while keeping strong visual quality and synchronized audio. It trades a little polish for speed, delivering quick text-to-video results at a lower resolution tier when needed. This mode is well suited for drafts, prompt testing, and high-volume iteration before a final render. |
| Wan 2.5 I2V API (Image To Video) | Wan 2.5 I2V API animates a single still into motion while preserving the subject's identity, lighting, and style, and adds synchronized sound in the same generation. It holds facial features, proportions, and composition steady, making it a fit for portraits, product shots, and illustrations that need to move. This mode extends static visuals into short-form video with audio. |
| Wan 2.5 I2V Fast API (Image To Video Fast) | Wan 2.5 I2V Fast API accelerates image-to-video generation for time-sensitive work, keeping core subject identity and motion while optimizing inference speed. It returns animated visuals faster without a major quality sacrifice. This mode is well suited for previews at scale, batch animation, and rapid social content where speed leads. |
| Wan 2.5 Image Edit API (Image To Image) | Wan 2.5 Image Edit API refines and transforms stills from natural-language instructions, adjusting or recomposing an image while preserving the rest. It works on single or reference-guided edits for concept art and asset prep. This mode is a fit for polishing source frames before animating them through a video mode. |
| Wan 2.5 T2I API (Text To Image) | Wan 2.5 T2I API generates still images from text across photographic and artistic styles. It produces concept frames, keyframes, and reference stills that can feed the video modes. This mode suits ideation and building the visual starting points for an audio-synced clip. |
The Wan 2.5 API generates synchronized audio and video in one pass on Alibaba's Diffusion Transformer architecture, with voice, sound, and lip-sync built in, cinematic camera control, and output from 480p to 1080p on Atlas Cloud.
Wan 2.5 produces the picture and its soundtrack in the same step, aligning voice, sound effects, and lip movement without a separate audio pass. A single structured prompt returns a finished clip with sound already in place, removing the record-and-align stage from the workflow.
The Wan 2.5 API exposes 480p, 720p, and full 1080p output, so you can match resolution to budget and channel. Draft at 480p for speed and volume, then render the same request at 1080p for delivery, all without leaving one integration.
Wan 2.5 keeps voice and lip movement aligned across languages, and handles Chinese prompts for audio-visual generation where some competing models fall back to an unknown-language error. This makes it dependable for localized dialogue and cross-market content.
Beyond generated sound, the Wan 2.5 API accepts an uploaded audio track to drive a clip, and reads cinematic direction such as pans, tilts, zooms, and dolly moves from the prompt. This gives control over both what a scene sounds like and how the camera moves through it.
Built on a Diffusion Transformer architecture with an efficient video VAE, Wan 2.5 restores character appearance, expression, and movement style with frame-to-frame stability. Subjects stay recognizable and motion holds together across the clip, a foundation for coherent short scenes.
The Wan 2.5 API covers text to video and image to video, plus a speed-optimized Fast variant for lower-latency image to video. Prototype quickly on the Fast tier, then move to the standard model for the final render, switching by changing the model name.
The same prompt, generated by Wan 2.5 and other leading video models: Complete narrative and advertisement
Create a 10-second high-quality product ad. On the desktop lies a matte black wireless headphone case. The lid opens slowly, revealing the headphones slightly. There’s a gentle magnetic click as the lid opens. The camera smoothly pans around the product at close range, showcasing the matte finish, small charging indicator light, and minimalist industrial design. In the background, the phone screen lights up, displaying a simple pairing animation. There are no floating particles, no sci-fi-style lighting effects, and no excessive visual effects.
Wan 2.5
HappyHorse 1.1
Pixverse c1
Create a 10-second realistic short film. A small bookstore on a rainy afternoon. Shot 1: Raindrops fall from the window, with the sound of rain outside. Shot 2: A young man climbs a wooden ladder to reach a old book on the top shelf. The ladder creaks softly. Shot 3: A handwritten note falls out of the book and lands on the floor. Shot 4: He bends down to pick up the note, smiles in surprise after reading it. Shot 5: The camera slowly zooms in on the note, with the sound of rain continuing in the background. The character remains consistent, movements are natural, lighting is realistic. The overall style is cinematic but understated, with no fantasy effects.
Wan 2.5
Wan 2.2
Pixverse c1
From talking-head content and localized ads to product demos and social clips, the Wan 2.5 API turns Alibaba's one-pass audio-visual model into production features through one key on Atlas Cloud.
Produce presenters, explainers, and spokesperson clips where voice and lip movement have to match, without recording audio or aligning it by hand. Wan 2.5 generates the speech and lip-sync in the same pass, so a scripted talking-head arrives finished.
Reach several markets from one script using the Wan 2.5 API's reliable audio and lip-sync across languages, including Chinese. Regenerate a scene in a new language and voice while keeping the visuals intact, so a single production ships localized versions.
Build short commercial videos with synchronized voiceover, sound effects, and cinematic camera moves in one generation. Consistent characters and controlled pans, tilts, and zooms make Wan 2.5 a fit for product launches, promos, and brand marketing.
Draft high-volume vertical clips at 480p or 720p on the Wan 2.5 API's Fast tier, then render the winners at 1080p for release. The resolution tiers let social teams keep iteration cheap and publish quality where it counts.
Animate a portrait, product photo, or hero frame into motion with matching sound, preserving the subject's identity and style from the still. This turns existing image assets into audio-ready clips without switching to a separate workflow.
Feed an uploaded audio track to drive a clip, or wire the Wan 2.5 API into an async pipeline that turns scripts and images into finished, sound-synced video at scale. Submit jobs, poll for results, and produce many clips on a schedule through one integration.
See how the Wan 2.5 API lines up against other audio-capable video models on Atlas Cloud by provider, native audio, resolution range, and clip length, so you can match each project to the right model, all under one key.
| Model | Provider | Native Audio | Resolution Range | Max Clip Length |
|---|---|---|---|---|
| Wan 2.5 | Alibaba | Yes (one-pass voice, sound, lip-sync) | 480p to 1080p | 10s |
| Wan 2.6 | Alibaba | Yes | Up to 1080p | ~15s |
| Kling 3.0 | Kuaishou | Yes | Higher than 1080p | 10s |
| Hailuo 2.3 | MiniMax | No | Up to 1080p | 10s |
| Veo 3.1 | Yes | Higher than 1080p | ~8s |
Get started in minutes — follow these simple steps to integrate and deploy models through Atlas Cloud's platform.
Sign up at atlascloud.ai and complete verification. New users receive free credits to explore the platform and test models.
Combining the advanced Wan 2.5 models with Atlas Cloud's GPU-accelerated platform provides unmatched performance, scalability, and developer experience.
Low Latency:
GPU-optimized inference for real-time reasoning.
Unified API:
Run Wan 2.5, GPT, Gemini, and DeepSeek with one integration.
Transparent Pricing:
Predictable per-token billing with serverless options.
Developer Experience:
SDKs, analytics, fine-tuning tools, and templates.
Reliability:
99.99% uptime, RBAC, and compliance-ready logging.
Security & Compliance:
SOC 2 Type II, HIPAA alignment, data sovereignty in US.
The Wan 2.5 API gives developers Alibaba's Wan 2.5 video model on Atlas Cloud through one OpenAI-compatible key. Its defining trait is one-pass audio-visual generation: it produces the picture and a synchronized soundtrack, including voice, sound effects, and lip-sync, in a single step. Built on a Diffusion Transformer architecture, it runs alongside 300+ other models on the same account.
Yes. The Wan 2.5 API creates video and its audio together in one pass, aligning voice, ambient sound, and lip movement without a separate recording or manual sync step. A single structured prompt returns a finished clip with sound already in place, which removes the record-and-align stage that most video pipelines still need.
Wan 2.5 generates at 480p, 720p, and full 1080p, with clip lengths of 5 or 10 seconds. The tiered resolutions let you draft cheaply at 480p and deliver at 1080p from the same request. Exact resolution and duration options depend on the mode and configuration on Atlas Cloud, so confirm the current settings in the console before building around a specific output.
The Wan 2.5 API covers text to video and image to video, with a speed-optimized Fast variant for image to video and a text to image mode. Text to video builds a scene from a prompt, while image to video animates a still and preserves its identity and style. Switching modes is a change to the model name on one integration.
Yes. Alongside generated sound, the Wan 2.5 API accepts an uploaded audio file to drive a clip, so the video can follow a specific voiceover or track. It also reads cinematic direction from the prompt, such as pans, tilts, zooms, and dolly moves, giving control over both the soundtrack and the camera work in a scene.
Wan 2.5 keeps voice and lip movement aligned across languages and reliably processes Chinese prompts for audio-visual generation, an area where some competing models fall back to an unknown-language response. This makes it dependable for localized dialogue and cross-market content where the audio has to match the spoken language on screen.
The three generations serve different needs. Wan 2.2 is the open, Mixture-of-Experts model for those who want open weights; Wan 2.5 centers on one-pass audio-visual generation with 480p to 1080p tiers; and Wan 2.6 focuses on multi-shot storyboard storytelling with character consistency. Because all three sit behind one key on Atlas Cloud, you can pick the version that fits each job.
Fast is a speed-optimized Wan 2.5 variant for image to video, tuned for lower latency. Use it for drafts, batch generation, and prompt testing where quick turnaround matters more than maximum polish, then move approved shots to the standard model for the final render. It keeps iteration quick without changing your integration.
Yes. Video generated with Wan 2.5 can be used in commercial work such as ads, product demos, and published content. Review Atlas Cloud's terms of service for the specifics of your plan, and note the usual restrictions around generating content that depicts real, identifiable people without their consent.
Create an account on Atlas Cloud, generate an API key, and send a request to the Wan 2.5 model with your prompt or input image through the OpenAI-compatible endpoint, choosing your resolution and duration. Poll the prediction endpoint for the finished clip, then scale up as needed. The same key reaches 300+ other models, so you can test alternatives without extra setup.
Guides, tutorials, and product updates to help you get the most out of Atlas Cloud.