MiniMax H3 Developer Now Live — 60% Off, From $0.02 per Second
Wan 2.5 API: Video and Audio in One Pass

Wan 2.5 API: Video and Audio in One Pass

The Wan 2.5 API produces video and audio in one pass, so voice, sound, and lip-sync line up without a separate step. Built on Alibaba's Diffusion Transformer architecture, it covers text and image to video from 480p to 1080p at 5 or 10 seconds, with reliable sync even for Chinese prompts. Reach it through one key on Atlas Cloud, alongside 300+ models.

Wan 2.5 is developed by Alibaba. Atlas Cloud (operated by Atlas Cloud AI LLC) provides access to it and does not own it. All trademarks belong to their respective owners.

Explore the Leading Wan 2.5

Atlas Cloud provides you with the latest industry-leading creative models.

Compare the Wan 2.5 API Variants

Match each job to the right Wan 2.5 option: Fast tiers for quick, high-volume drafts, standard tiers for final audio-synced video, and image modes for stills, across text and image on Atlas Cloud.

VariantDescription
Wan 2.5 T2V API (Text To Video)Wan 2.5 T2V API turns a written prompt into a cinematic clip with synchronized voice, sound, and lip-sync generated in the same pass. It reads camera direction such as pans, tilts, and zooms, and outputs from 480p to 1080p at 5 or 10 seconds. This mode suits scripts, ad concepts, and any scene that needs sound without a separate audio step.
Wan 2.5 T2V Fast API (Text To Video Fast)Wan 2.5 T2V Fast API prioritizes lower latency while keeping strong visual quality and synchronized audio. It trades a little polish for speed, delivering quick text-to-video results at a lower resolution tier when needed. This mode is well suited for drafts, prompt testing, and high-volume iteration before a final render.
Wan 2.5 I2V API (Image To Video)Wan 2.5 I2V API animates a single still into motion while preserving the subject's identity, lighting, and style, and adds synchronized sound in the same generation. It holds facial features, proportions, and composition steady, making it a fit for portraits, product shots, and illustrations that need to move. This mode extends static visuals into short-form video with audio.
Wan 2.5 I2V Fast API (Image To Video Fast)Wan 2.5 I2V Fast API accelerates image-to-video generation for time-sensitive work, keeping core subject identity and motion while optimizing inference speed. It returns animated visuals faster without a major quality sacrifice. This mode is well suited for previews at scale, batch animation, and rapid social content where speed leads.
Wan 2.5 Image Edit API (Image To Image)Wan 2.5 Image Edit API refines and transforms stills from natural-language instructions, adjusting or recomposing an image while preserving the rest. It works on single or reference-guided edits for concept art and asset prep. This mode is a fit for polishing source frames before animating them through a video mode.
Wan 2.5 T2I API (Text To Image)Wan 2.5 T2I API generates still images from text across photographic and artistic styles. It produces concept frames, keyframes, and reference stills that can feed the video modes. This mode suits ideation and building the visual starting points for an audio-synced clip.

Wan 2.5 API Features

The Wan 2.5 API generates synchronized audio and video in one pass on Alibaba's Diffusion Transformer architecture, with voice, sound, and lip-sync built in, cinematic camera control, and output from 480p to 1080p on Atlas Cloud.

One-Pass Audio-Visual Generation

Wan 2.5 produces the picture and its soundtrack in the same step, aligning voice, sound effects, and lip movement without a separate audio pass. A single structured prompt returns a finished clip with sound already in place, removing the record-and-align stage from the workflow.

Resolution Tiers from 480p to 1080p with the Wan 2.5 API

The Wan 2.5 API exposes 480p, 720p, and full 1080p output, so you can match resolution to budget and channel. Draft at 480p for speed and volume, then render the same request at 1080p for delivery, all without leaving one integration.

Reliable Multilingual and Chinese Sync

Wan 2.5 keeps voice and lip movement aligned across languages, and handles Chinese prompts for audio-visual generation where some competing models fall back to an unknown-language error. This makes it dependable for localized dialogue and cross-market content.

Custom Audio and Camera Control with the Wan 2.5 API

Beyond generated sound, the Wan 2.5 API accepts an uploaded audio track to drive a clip, and reads cinematic direction such as pans, tilts, zooms, and dolly moves from the prompt. This gives control over both what a scene sounds like and how the camera moves through it.

Character Consistency and Motion Stability

Built on a Diffusion Transformer architecture with an efficient video VAE, Wan 2.5 restores character appearance, expression, and movement style with frame-to-frame stability. Subjects stay recognizable and motion holds together across the clip, a foundation for coherent short scenes.

Text, Image, and a Fast Tier with the Wan 2.5 API

The Wan 2.5 API covers text to video and image to video, plus a speed-optimized Fast variant for lower-latency image to video. Prototype quickly on the Fast tier, then move to the standard model for the final render, switching by changing the model name.

Wan 2.5 vs Other Models - One Prompt

The same prompt, generated by Wan 2.5 and other leading video models: Complete narrative and advertisement

Prompt

Create a 10-second high-quality product ad. On the desktop lies a matte black wireless headphone case. The lid opens slowly, revealing the headphones slightly. There’s a gentle magnetic click as the lid opens. The camera smoothly pans around the product at close range, showcasing the matte finish, small charging indicator light, and minimalist industrial design. In the background, the phone screen lights up, displaying a simple pairing animation. There are no floating particles, no sci-fi-style lighting effects, and no excessive visual effects.

Wan 2.5

HappyHorse 1.1

Pixverse c1

Prompt

Create a 10-second realistic short film. A small bookstore on a rainy afternoon. Shot 1: Raindrops fall from the window, with the sound of rain outside. Shot 2: A young man climbs a wooden ladder to reach a old book on the top shelf. The ladder creaks softly. Shot 3: A handwritten note falls out of the book and lands on the floor. Shot 4: He bends down to pick up the note, smiles in surprise after reading it. Shot 5: The camera slowly zooms in on the note, with the sound of rain continuing in the background. The character remains consistent, movements are natural, lighting is realistic. The overall style is cinematic but understated, with no fantasy effects.

Wan 2.5

Wan 2.2

Pixverse c1

What You Can Build with the Wan 2.5 API

From talking-head content and localized ads to product demos and social clips, the Wan 2.5 API turns Alibaba's one-pass audio-visual model into production features through one key on Atlas Cloud.

Talking-Head and Dialogue Video

Produce presenters, explainers, and spokesperson clips where voice and lip movement have to match, without recording audio or aligning it by hand. Wan 2.5 generates the speech and lip-sync in the same pass, so a scripted talking-head arrives finished.

Localized and Multilingual Campaigns with the Wan 2.5 API

Reach several markets from one script using the Wan 2.5 API's reliable audio and lip-sync across languages, including Chinese. Regenerate a scene in a new language and voice while keeping the visuals intact, so a single production ships localized versions.

Product Demos and Ad Spots

Build short commercial videos with synchronized voiceover, sound effects, and cinematic camera moves in one generation. Consistent characters and controlled pans, tilts, and zooms make Wan 2.5 a fit for product launches, promos, and brand marketing.

Cost-Tiered Social and Short-Form Video with the Wan 2.5 API

Draft high-volume vertical clips at 480p or 720p on the Wan 2.5 API's Fast tier, then render the winners at 1080p for release. The resolution tiers let social teams keep iteration cheap and publish quality where it counts.

Image-Led Animation with Audio

Animate a portrait, product photo, or hero frame into motion with matching sound, preserving the subject's identity and style from the still. This turns existing image assets into audio-ready clips without switching to a separate workflow.

Audio-Driven and Pipeline Video with the Wan 2.5 API

Feed an uploaded audio track to drive a clip, or wire the Wan 2.5 API into an async pipeline that turns scripts and images into finished, sound-synced video at scale. Submit jobs, poll for results, and produce many clips on a schedule through one integration.

How the Wan 2.5 API Compares

See how the Wan 2.5 API lines up against other audio-capable video models on Atlas Cloud by provider, native audio, resolution range, and clip length, so you can match each project to the right model, all under one key.

ModelProviderNative AudioResolution RangeMax Clip Length
Wan 2.5AlibabaYes (one-pass voice, sound, lip-sync)480p to 1080p10s
Wan 2.6AlibabaYesUp to 1080p~15s
Kling 3.0KuaishouYesHigher than 1080p10s
Hailuo 2.3MiniMaxNoUp to 1080p10s
Veo 3.1GoogleYesHigher than 1080p~8s

How to Use Wan 2.5 on Atlas Cloud

Get started in minutes — follow these simple steps to integrate and deploy models through Atlas Cloud's platform.

Create an Atlas Cloud Account

Sign up at atlascloud.ai and complete verification. New users receive free credits to explore the platform and test models.

Why Use Wan 2.5 on Atlas Cloud

Combining the advanced Wan 2.5 models with Atlas Cloud's GPU-accelerated platform provides unmatched performance, scalability, and developer experience.

Performance & flexibility

Low Latency:
GPU-optimized inference for real-time reasoning.

Unified API:
Run Wan 2.5, GPT, Gemini, and DeepSeek with one integration.

Transparent Pricing:
Predictable per-token billing with serverless options.

Enterprise & Scale

Developer Experience:
SDKs, analytics, fine-tuning tools, and templates.

Reliability:
99.99% uptime, RBAC, and compliance-ready logging.

Security & Compliance:
SOC 2 Type II, HIPAA alignment, data sovereignty in US.

Common Questions About the Wan 2.5 API

The Wan 2.5 API gives developers Alibaba's Wan 2.5 video model on Atlas Cloud through one OpenAI-compatible key. Its defining trait is one-pass audio-visual generation: it produces the picture and a synchronized soundtrack, including voice, sound effects, and lip-sync, in a single step. Built on a Diffusion Transformer architecture, it runs alongside 300+ other models on the same account.

Yes. The Wan 2.5 API creates video and its audio together in one pass, aligning voice, ambient sound, and lip movement without a separate recording or manual sync step. A single structured prompt returns a finished clip with sound already in place, which removes the record-and-align stage that most video pipelines still need.

Wan 2.5 generates at 480p, 720p, and full 1080p, with clip lengths of 5 or 10 seconds. The tiered resolutions let you draft cheaply at 480p and deliver at 1080p from the same request. Exact resolution and duration options depend on the mode and configuration on Atlas Cloud, so confirm the current settings in the console before building around a specific output.

The Wan 2.5 API covers text to video and image to video, with a speed-optimized Fast variant for image to video and a text to image mode. Text to video builds a scene from a prompt, while image to video animates a still and preserves its identity and style. Switching modes is a change to the model name on one integration.

Yes. Alongside generated sound, the Wan 2.5 API accepts an uploaded audio file to drive a clip, so the video can follow a specific voiceover or track. It also reads cinematic direction from the prompt, such as pans, tilts, zooms, and dolly moves, giving control over both the soundtrack and the camera work in a scene.

Wan 2.5 keeps voice and lip movement aligned across languages and reliably processes Chinese prompts for audio-visual generation, an area where some competing models fall back to an unknown-language response. This makes it dependable for localized dialogue and cross-market content where the audio has to match the spoken language on screen.

The three generations serve different needs. Wan 2.2 is the open, Mixture-of-Experts model for those who want open weights; Wan 2.5 centers on one-pass audio-visual generation with 480p to 1080p tiers; and Wan 2.6 focuses on multi-shot storyboard storytelling with character consistency. Because all three sit behind one key on Atlas Cloud, you can pick the version that fits each job.

Fast is a speed-optimized Wan 2.5 variant for image to video, tuned for lower latency. Use it for drafts, batch generation, and prompt testing where quick turnaround matters more than maximum polish, then move approved shots to the standard model for the final render. It keeps iteration quick without changing your integration.

Yes. Video generated with Wan 2.5 can be used in commercial work such as ads, product demos, and published content. Review Atlas Cloud's terms of service for the specifics of your plan, and note the usual restrictions around generating content that depicts real, identifiable people without their consent.

Create an account on Atlas Cloud, generate an API key, and send a request to the Wan 2.5 model with your prompt or input image through the OpenAI-compatible endpoint, choosing your resolution and duration. Poll the prediction endpoint for the finished clip, then scale up as needed. The same key reaches 300+ other models, so you can test alternatives without extra setup.

Explore More Families

Seedance 2.5

Seedance 2.5 API is now available on Atlas Cloud! It gives developers ByteDance's newest video model. It generates up to 30 seconds of native video in a single pass from text, a single image, or as many as 50 multimodal references, with synchronized audio and in-frame multilingual text. On Atlas Cloud you reach it through one key, with subject consistency and improved physics keeping long shots coherent. (Update: Seedance 2.5 1080P API Is Available NOW!)

View Family

Wan 3.0

Wan 3.0 API is the next generation of Alibaba's Wan video family, built to push long-form generation, multi-reference control, and audiovisual quality to new heights. Atlas Cloud already hosts Wan 2.7, 2.6, and 2.5, and Wan 3.0 runs on the same unified key with no separate setup. Start building today. Scroll down to the showcase to see what Wan 3.0 can create.

View Family

MiniMax H3

MiniMax H3 is MiniMax’s video model family, spanning H3, H3 Max, and H3 Developer. Create from text, animate a first frame with an optional last frame, or preserve subjects from references. H3 and H3 Developer reach 2K, while H3 Max supports 480P and 768P clips lasting 5 to 15 seconds. Atlas Cloud adds OpenAI-compatible access and transparent pay-as-you-go pricing from $0.05 per second. Start building today.

View Family

Seedream 5.0 Pro

Seedream 5.0 Pro API gives developers ByteDance's controllable image editing model on Atlas Cloud. It places edits precisely with anchors and coordinates, separates images into editable layers, fuses multiple references, and matches exact colors and materials, with multilingual text at 2K and 3K. On Atlas Cloud you reach it through one key!

View Family

Seedance 2.0

The Seedance 2.0 API gives you production access to ByteDance's multimodal video model — quad-modal inputs (text, image, video, audio) and an industry-leading "Universal Reference" system that locks composition, camera movement, and character actions across shots. Integrate director-level control with one API call, a flat $0.09/s, instant key, and no waitlist — backed by enterprise-grade uptime and compliance. Seedance 2.0 Native 4K is now live!

View Family

GPT Image 2

The GPT Image 2 API gives developers access to OpenAI's latest image model, the successor to GPT Image 1.5. It generates and edits images with accurate text rendering across Latin and CJK scripts, plus strong composition for posters, mockups, and infographics. On Atlas Cloud you reach it through one unified API alongside 300+ models, with free credits, 99.99% uptime, and no OpenAI organization verification required.

View Family

Gemini Omni Flash

The gemini omni API brings Google DeepMind's natively multimodal Gemini Omni Flash family, including Gemini Omni 1.1 Flash, to developers. Create cinematic video with synchronized native audio, animate still images with precise start and end frame control, or revise existing footage through text guided edits that preserve untouched content. Atlas Cloud provides one OpenAI-compatible key, unified access, and transparent pay-as-you-go pricing. Start building today.

View Family

Grok Imagine

The Grok Imagine API covers xAI's image, video, and speech models, from Image 2.0 to Video 1.5 and xAI TTS v1. Render 1K or 2K stills across 14 aspect ratios, push a scene to 15 seconds of 1080p motion, steer shots with up to 7 reference images, or narrate them in 20 languages. Atlas Cloud runs every mode on one endpoint, priced pay-as-you-go from $0.02 per image and $0.05 per second. Start building today.

View Family

Google

Google's most powerful creative models are all available on Atlas Cloud. Veo 3.1 delivers cinematic video generation, Nano Banana 2 powers high-fidelity image creation, and Gemini brings multimodal intelligence to every workflow. Access the full Google model suite through one API key with Day-0 availability and pay-as-you-go pricing.

View Family

Seedance 2.0 Mini

The Seedance 2.0 Mini API is the lightest, lowest-cost tier of ByteDance's Seedance video line, built for teams where throughput and unit cost matter more than maximum polish. Use it for batch generation, rapid prototyping, and draft passes, all through one OpenAI-compatible key on Atlas Cloud.

View Family

ByteDance

From cinematic video generation to high-fidelity image creation, ByteDance's most powerful models are live on Atlas Cloud. Run Seedance and Seedream at scale with the lowest inference pricing and zero infrastructure overhead.

View Family

Alibaba

Atlas Cloud brings together Alibaba's full model lineup under one API: Qwen for language and image tasks, Wan for video generation up to 1080p. Access every model pay-as-you-go with no subscriptions. The Alibaba API is available via a single base URL using your existing OpenAI-compatible client.

View Family

One API for All Media AI.

Explore all models