MiniMax H3 Developer Now Live — 60% Off, From $0.02 per Second

Wan 3.0 API is Now Live on Atlas Cloud!

Wan 3.0 API is the next generation of Alibaba's Wan video family, built to push long-form generation, multi-reference control, and audiovisual quality to new heights. Atlas Cloud already hosts Wan 2.7, 2.6, and 2.5, and Wan 3.0 runs on the same unified key with no separate setup. Start building today. Scroll down to the showcase to see what Wan 3.0 can create.

Wan 3.0 is developed by Alibaba. Atlas Cloud (operated by Atlas Cloud AI LLC) provides access to it and does not own it. All trademarks belong to their respective owners.

Explore the Leading Wan 3.0

Atlas Cloud provides you with the latest industry-leading creative models.

Text, Image, or Reference: Which Wan 3.0 API Endpoint Fits

Three Wan 3.0 API endpoints share one Atlas Cloud key, so the only real question is what you already have to feed the model.

ModalityDescription
Wan 3.0 T2V API (Text to Video)A written prompt of up to 20000 characters becomes a finished clip with a native audio track, rendered at 480P, 720P, or 1080P and running anywhere from 2 to 30 seconds. Set the aspect ratio to 16:9, 9:16, 4:3, 3:4, or 1:1, or leave both ratio and duration adaptive and let the model pick the framing and the pacing. It fits concept films, ad boards, and social spots that begin with nothing but a script.
Wan 3.0 I2V API (Image to Video)Feed a first frame and the clip is animated out of it; add an optional last frame and the motion between the two is interpolated for you. Source stills accept jpeg, jpg, png, bmp, and webp between 240 and 8000 pixels and under 20MB, while output reaches 1080P with the audio track on by default. Reach for this endpoint when a product photo, a poster, or a storyboard pair already fixes how the scene has to open and close.
Wan 3.0 R2V API (Reference to Video)Up to 20 reference materials ride along with the prompt: 10 images, 5 videos totaling 15 seconds, and 5 audio clips totaling 15 seconds, aligned at pixel level for identity, voice, and space. Deep thinking mode opens a second door, letting a document of up to 100MB and 50 pages (pdf, docx, pptx, xlsx, txt, md, and more) or a live webpage act as the source. Reserve it for work where a recurring character, an exact brand asset, or a full report must survive intact into the video.

Wan 3.0 in Action

6 flagship capabilities that show what one Wan 3.0 request can produce, from multi-reference brand films to 30-second single takes.

Multi-Reference Brand Films

Omni mode ingests up to 20 reference materials, 10 images, 5 clips, and 5 audio tracks, plus one source document or webpage, so a full brand kit becomes a single on-brand video in one pass.

30-Second Single-Take Stories

Native 30-second output gives a narrative room to breathe, holding continuous camera moves and one-shot sequences that shorter clips cannot carry. Intelligent duration lets the model pace each scene on its own.

Pixel-Level Reference Consistency

Faces, products, and fine details are reproduced with pixel-level fidelity across every frame. Production teams gain the delivery certainty that repeat characters and exact brand assets demand.

Immersive Sound and Realism

Realism, image texture, and native audio all step up together, so a generated scene arrives with matching sound instead of silent footage. The result lands with real audiovisual impact.

Instruction and Reference Editing

Precise editing reworks existing footage from a written instruction or a reference clip, changing a subject, style, or moment without regenerating the whole video.

Documents and Webpages to Video

Beyond text, image, audio, and video, Wan 3.0 reads doc, xls, ppt, pdf, and md files as well as live webpages, turning a report or slide deck straight into a finished clip.

Video Comparisons in Same Prompt

See how video present generated by Wan 3.0, Seedance 2.5 and Wan 2.7 model.

Prompt

Generating a 30-second, 16:9 widescreen adult sci-fi war animated short film. Visual style: Ultra-photorealistic CGI Biomechanical structures Fusion of gray-white bone and black metal Insect-like joints Extremely fine bone texturing Dried blood streaks, ceramic cracks, and industrial oil grime Abandoned futuristic city Cold gray skylight Sparse toxic-green organisms High-contrast cinematic lighting Low-saturation cool color palette Minimal orange-red used only for flames and alarms Realistic physical destruction Oppressive, heavy, unheroic tone Characters no longer have complete flesh bodies. They are composed primarily of: Human spine Ribs Skull Black hydraulic muscle Metal joints Neural fiber optics Weapon interface ports The skeletons are not zombies — they are "legacy hardware" still being operated by military-industrial systems. Narrative & Dynamic Timeline 0–4 sec | Still Fighting After Death The shot opens on a soldier's dog tag lying in a pool of standing water. Want me to continue translating the rest of the timeline if you have more seconds/beats written out? This looks like it's cut off after the 0-4 second beat.

Generated by Wan 3.0

Generated by Seedance 2.5 on Altas Cloud

Generated by Wan 2.7 on Altas Cloud

Prompt

Generating a 15-second, 16:9 ultra-widescreen plush-universe adventure clip. A spaceship made of plush fabric, buttons, zippers, and stuffing cotton is fleeing at high speed through a universe made of giant yarn planets, fuzzy nebulae, and rag-doll megafauna. The ship is piloted by three small plush animals: a rabbit, a fox, and a crow. A giant interstellar whale-beast covered in deep-blue long plush fur, with eyes like two glass buttons, chases the ship out of the nebula. Overall style: ultra-fine plush textures, cinematic space scale, a strong contrast between cute appearance and intense action, exaggerated FPV camera work, and realistic soft-body physics. **15-Second Dynamic Timeline** **0–3 sec | Zipper Ship Ejection Launch** The shot is inside the plush ship, with a cockpit made of stitching, buttons, and fabric instrument panels. A red warning button flashes rapidly. The rabbit pulls down a giant zipper-lever control, and the ship's forward hatch splits open like a cloth sack. The camera pulls back rapidly from the cockpit, passes through the zipper hatch, and emerges outside the ship. The ship ejects and launches out of the knitted crater of a yarn planet. **3–6 sec | High-Speed Orbit Skimming the Yarn Planet** The camera hugs the ship's tail, flying low over the planet's surface. The ground is woven from giant strands of yarn: - Knitted mountain ranges streak past below at high speed - Loose thread-ends sway like a forest - The ship dodges sharply left and right - The camera continuously banks and rolls - The thrusters emit a trailing wake of white cotton fluff **6–8.5 sec | The Giant Plush Whale Appears** The fuzzy nebula suddenly parts to either side as a giant deep-blue plush whale bursts out from behind. The long fur on the whale's body is blown backward by the high-speed airflow, and its glass-button eyes reflect the planet's light. The camera whips halfway around the ship, switching to a rear-facing (reverse) view, so the whale continuously grows larger in the frame behind the ship.

Generated by Wan 3.0

Generated by Seedance 2.5 on Altas Cloud

Generated by Wan 2.7 on Altas Cloud

Wan 3.0 API - Built for Real Creative Pipelines

Wan 3.0's upgrades trace back to real industry demands, folding straight into the production workflows below, from short drama to corporate film.

Film, TV, Short Drama & MV

Native 30-second takes and real-scene restoration let AI film, short drama, and music MVs skip much of the physical shoot. Pixel consistency holds recurring leads, and native audio arrives matched to every cut.

Animation, Brand & IP

All-style control moves from realistic to anime within one model, while pixel consistency locks a mascot, logo, or IP character across every frame. Full brand animations render without the tedious production chain.

Advertising & Corporate Promotion

Omni creation pushes past text-and-image limits for home appliance, beauty, auto, FMCG, and apparel brands. Feed a product page or company report straight in to build everything from product demos to corporate brand films.

Software & Product Design

Feed a product UI walkthrough, a feature animation, or a data visualization, and Wan 3.0 keeps the original aesthetic while turning it into a story-driven motion piece. The tedious animation step disappears.

Education

Document input reads a pdf, ppt, or webpage directly, so a lesson or research report becomes a clear explainer video with no separate script. Complex material turns visual in a single step.

Cultural Tourism

Real-scene restoration recreates landscapes, heritage sites, and local cuisine without a large-scale shoot, so city films and cultural digitization finish at low cost. Distant places arrive right in front of viewers.

How the Wan 3.0 API Stacks Up Against Other Video Models

Put the Wan 3.0 API side by side with the video models it shares one Atlas Cloud key with, and check where duration, reference capacity, audio, and per second cost actually differ.

ModelMax DurationResolutionReference & File InputsNative AudioPrice per Second
Wan-3.0 Reference-to-video2 to 30s, or -1 for smart duration480P, 720P, 1080PUp to 20 mixed items: 10 images, 5 videos, 5 audio clips, plus one document or webpage√ (on by default)$0.05
Wan-3.0 Text-to-video2 to 30s, or -1 for smart duration480P, 720P, 1080PPrompt only, up to 20,000 characters√ (on by default)$0.05
Wan-3.0 Image-to-video2 to 30s, or -1 for smart duration480P, 720P, 1080PFirst frame, with an optional last frame√ (on by default)$0.05
Seedance 2.5 Reference-to-Video4 to 30s, or -1 for automatic length480p, 720p, 1080p native, higher tiers via enhancementUp to 30 images, 10 videos, 10 audio clips, no document input√ (on by default)$0.134
Kling V3.0 Turbo Image-to-Video3 to 15s, default 5s720p, 1080pOne source image, no multi reference mode√ (optional, off by default)$0.112
Veo3.1 Reference-to-video8s fixed720p, 1080p, 4K1 to 3 reference images√ (optional, off by default)$0.2
MiniMax H3 Reference-to-Video4 to 15s768P, 2KMixed image, video, and audio references, at least one visual reference required-$0.1

How to Use Wan 3.0 on Atlas Cloud

Get started in minutes — follow these simple steps to integrate and deploy models through Atlas Cloud's platform.

Create an Atlas Cloud Account

Sign up at atlascloud.ai and complete verification. New users receive free credits to explore the platform and test models.

Why Use Wan 3.0 on Atlas Cloud

Combining the advanced Wan 3.0 models with Atlas Cloud's GPU-accelerated platform provides unmatched performance, scalability, and developer experience.

Performance & flexibility

Low Latency:
GPU-optimized inference for real-time reasoning.

Unified API:
Run Wan 3.0, GPT, Gemini, and DeepSeek with one integration.

Transparent Pricing:
Predictable per-token billing with serverless options.

Enterprise & Scale

Developer Experience:
SDKs, analytics, fine-tuning tools, and templates.

Reliability:
99.99% uptime, RBAC, and compliance-ready logging.

Security & Compliance:
SOC 2 Type II, HIPAA alignment, data sovereignty in US.

Wan 3.0 API FAQ

Wan 3.0 is Alibaba's next-generation all-in-one video model, now in public beta. One request can turn text, images, video, audio, documents, or webpages into a video up to 30 seconds long, and the model pairs that broad input with pixel-level consistency, precise editing, and natively generated audio.

Wan 3.0 is the broadest release in the line. It lifts native duration to 30 seconds, accepts up to 20 mixed reference materials plus documents and webpages in omni-reference mode, adds a first-and-last-frame mode, and pushes toward pixel-level consistency with native audio. The earlier Wan versions on Atlas Cloud stay strong picks for shorter, more focused text-to-video and image-to-video work.

Wan 3.0 treats far more than a prompt as source material. One request can draw on up to 20 mixed references, 10 images, 5 video clips totaling 15 seconds, and 5 audio clips totaling 15 seconds, alongside a single document (doc, xls, ppt, pdf, or md) or webpage, with the text prompt itself running as long as 20000 characters.

Clips run natively up to 30 seconds. A text or image request can land anywhere from 2 to 30 seconds, while a request that includes video keeps input and output within a combined 30 seconds, and intelligent duration can set the length for you. Resolution spans 480p, 720p, and 1080p, with 16:9, 9:16, 4:3, 3:4, 1:1, and intelligent ratios available.

Consistency is handled at pixel level, so a face, an outfit, a product, or a fine reference detail holds its look as scenes cut and subjects move. For production runs that reuse the same character or an exact brand asset, that steadiness is what makes each delivery predictable.

Yes, through both instruction and reference. A written note or a sample clip can restyle a subject, shift lighting, or replace a single segment, and the elements you want left alone stay in place across the edit.

Yes. Atlas Cloud already serves Wan 2.7, 2.6, and 2.5 under one unified key, and Wan 3.0 runs on that same key with no new account, no endpoint reshuffle, and no separate setup. Teams already calling the Wan family can reach Wan 3.0 by swapping the model name.

Explore More Families

Seedance 2.5

Seedance 2.5 API is now available on Atlas Cloud! It gives developers ByteDance's newest video model. It generates up to 30 seconds of native video in a single pass from text, a single image, or as many as 50 multimodal references, with synchronized audio and in-frame multilingual text. On Atlas Cloud you reach it through one key, with subject consistency and improved physics keeping long shots coherent. (Update: Seedance 2.5 1080P API Is Available NOW!)

View Family

Wan 3.0

Wan 3.0 API is the next generation of Alibaba's Wan video family, built to push long-form generation, multi-reference control, and audiovisual quality to new heights. Atlas Cloud already hosts Wan 2.7, 2.6, and 2.5, and Wan 3.0 runs on the same unified key with no separate setup. Start building today. Scroll down to the showcase to see what Wan 3.0 can create.

View Family

MiniMax H3

MiniMax H3 is MiniMax’s video model family, spanning H3, H3 Max, and H3 Developer. Create from text, animate a first frame with an optional last frame, or preserve subjects from references. H3 and H3 Developer reach 2K, while H3 Max supports 480P and 768P clips lasting 5 to 15 seconds. Atlas Cloud adds OpenAI-compatible access and transparent pay-as-you-go pricing from $0.05 per second. Start building today.

View Family

Seedream 5.0 Pro

Seedream 5.0 Pro API gives developers ByteDance's controllable image editing model on Atlas Cloud. It places edits precisely with anchors and coordinates, separates images into editable layers, fuses multiple references, and matches exact colors and materials, with multilingual text at 2K and 3K. On Atlas Cloud you reach it through one key!

View Family

Seedance 2.0

The Seedance 2.0 API gives you production access to ByteDance's multimodal video model — quad-modal inputs (text, image, video, audio) and an industry-leading "Universal Reference" system that locks composition, camera movement, and character actions across shots. Integrate director-level control with one API call, a flat $0.09/s, instant key, and no waitlist — backed by enterprise-grade uptime and compliance. Seedance 2.0 Native 4K is now live!

View Family

GPT Image 2

The GPT Image 2 API gives developers access to OpenAI's latest image model, the successor to GPT Image 1.5. It generates and edits images with accurate text rendering across Latin and CJK scripts, plus strong composition for posters, mockups, and infographics. On Atlas Cloud you reach it through one unified API alongside 300+ models, with free credits, 99.99% uptime, and no OpenAI organization verification required.

View Family

Gemini Omni Flash

The gemini omni API brings Google DeepMind's natively multimodal Gemini Omni Flash family, including Gemini Omni 1.1 Flash, to developers. Create cinematic video with synchronized native audio, animate still images with precise start and end frame control, or revise existing footage through text guided edits that preserve untouched content. Atlas Cloud provides one OpenAI-compatible key, unified access, and transparent pay-as-you-go pricing. Start building today.

View Family

Grok Imagine

The Grok Imagine API covers xAI's image, video, and speech models, from Image 2.0 to Video 1.5 and xAI TTS v1. Render 1K or 2K stills across 14 aspect ratios, push a scene to 15 seconds of 1080p motion, steer shots with up to 7 reference images, or narrate them in 20 languages. Atlas Cloud runs every mode on one endpoint, priced pay-as-you-go from $0.02 per image and $0.05 per second. Start building today.

View Family

Google

Google's most powerful creative models are all available on Atlas Cloud. Veo 3.1 delivers cinematic video generation, Nano Banana 2 powers high-fidelity image creation, and Gemini brings multimodal intelligence to every workflow. Access the full Google model suite through one API key with Day-0 availability and pay-as-you-go pricing.

View Family

Seedance 2.0 Mini

The Seedance 2.0 Mini API is the lightest, lowest-cost tier of ByteDance's Seedance video line, built for teams where throughput and unit cost matter more than maximum polish. Use it for batch generation, rapid prototyping, and draft passes, all through one OpenAI-compatible key on Atlas Cloud.

View Family

ByteDance

From cinematic video generation to high-fidelity image creation, ByteDance's most powerful models are live on Atlas Cloud. Run Seedance and Seedream at scale with the lowest inference pricing and zero infrastructure overhead.

View Family

Alibaba

Atlas Cloud brings together Alibaba's full model lineup under one API: Qwen for language and image tasks, Wan for video generation up to 1080p. Access every model pay-as-you-go with no subscriptions. The Alibaba API is available via a single base URL using your existing OpenAI-compatible client.

View Family

One API for All Media AI.

Explore all models