MiniMax H3 Developer Now Live — 60% Off, From $0.02 per Second
Hero background 1Hero background 2

MiniMax H3 API for 2K Video

MiniMax H3 is MiniMax’s video model family, spanning H3, H3 Max, and H3 Developer. Create from text, animate a first frame with an optional last frame, or preserve subjects from references. H3 and H3 Developer reach 2K, while H3 Max supports 480P and 768P clips lasting 5 to 15 seconds. Atlas Cloud adds OpenAI-compatible access and transparent pay-as-you-go pricing from $0.05 per second. Start building today.

MiniMax H3 is developed by MiniMax. Atlas Cloud (operated by Atlas Cloud AI LLC) provides access to it and does not own it. All trademarks belong to their respective owners.

Explore the Leading MiniMax H3

Atlas Cloud provides you with the latest industry-leading creative models.

Choose the Right MiniMax H3 Endpoint for Each Video Input

Compare text, image, and reference driven MiniMax H3 modes across H3 Max, H3 Developer, and H3 by resolution, duration, aspect ratio, audio support, and standard pricing.

ModalityDescription
MiniMax H3 Max T2V API (Text to Video)Turn a text prompt into a cinematic video at 480P or 768P, with durations from 5 to 15 seconds. Select 16:9, 9:16, 1:1, or adaptive framing for trailers, advertisements, and social video at the standard price of $0.05 per second.
MiniMax H3 Max I2V API (Image to Video)Starting from a first frame, the model generates a prompt directed video and can optionally use a final frame to guide the ending. Its 480P and 768P output supports 5 to 15 second product animations, storyboard shots, and image based narratives at a standard $0.05 per second.
MiniMax H3 Developer T2V API (Text to Video)Built as a self hosted text to video model, this endpoint converts prompts into videos with generated audio at 480P, 768P, or 2K. Choose 16:9, 9:16, or 1:1 for narrative scenes, campaign concepts, and social content at the standard price of $0.05 per second.
MiniMax H3 Developer I2V API (Image to Video)Provide an opening image, an optional last frame, and a text prompt to create motion with generated audio. Output at 480P, 768P, or 2K makes the endpoint suitable for animating key art, product stills, and storyboard sequences, with standard pricing of $0.05 per second.
MiniMax H3 Developer Reference to Video APIWhen subject continuity matters, one or more reference images or videos can guide a new prompt driven video with generated audio. Available 480P, 768P, and 2K output supports recurring characters, branded subjects, and connected visual sequences at the standard price of $0.05 per second.
MiniMax H3 T2V API (Text to Video)A text prompt becomes a cinematic 2K video lasting from 5 to 15 seconds. Frame the result in 16:9, 9:16, 1:1, or an adaptive aspect ratio for polished trailers, vertical stories, and square campaign assets at the standard price of $0.08 per second.
MiniMax H3 I2V API (Image to Video)Use a first frame and text prompt to produce a 2K video, adding an optional last frame when the final composition needs direction. The 5 to 15 second range fits product reveals, character motion, and storyboard transitions, with standard pricing of $0.08 per second.
MiniMax H3 Reference to Video APIReference driven generation preserves the subject from an input image while a text prompt directs the resulting 2K video. Clips can run from 5 to 15 seconds, making this endpoint useful for character focused shots and consistent product visuals at the standard price of $0.08 per second.

Sight, Sound and Editing in One MiniMax H3 API Call

Every clip the MiniMax H3 API returns runs up to fifteen seconds at 1440p with native stereo sound, built from any mix of text, image, video and voice references, and the same call can also edit characters, scenes and dialogue inside footage you already have.

Twelve Reference Files, One MiniMax H3 API Call

One MiniMax H3 API request accepts up to nine reference images, three video clips and three audio tracks, capped at twelve files. Rather than reading them as separate slots, the model treats text, picture and sound as one context, pulling a character from a photo, a camera move from a clip and a mood from a track into the same scene. Teams with existing material get a direct route to a finished shot.

Native Stereo Sound in Every Clip

Every result ships with sound. Dialogue, ambience and music are produced in native stereo during the same pass that renders the picture at 24 frames per second, so nothing has to be scored or dubbed afterwards. Because timing is decided while the shot is generated, footsteps, speech and cuts stay locked to the action. Short ads and social spots come out ready to publish.

Prompt-Level Edits Through the MiniMax H3 API

Need a cat swapped for a dog, a green screen replaced, or one line of dialogue rewritten? The MiniMax H3 API applies edits like these to video you supply, covering characters, objects, backgrounds, lighting and effects, while parts you did not mention stay close to the original. Prompts run to 7000 characters, so a dozen changes can be stacked into one pass. That makes iterating on an approved cut practical.

Voice Timbre Transfer and Dialogue Swaps

Send a voice sample alongside your images or footage and the generated character speaks in that timbre. Up to three audio references are allowed per request, each between two and fifteen seconds, and audio must always accompany a visual input rather than arrive on its own. Existing dialogue can also be replaced and the performance adjusted to match. Series work keeps one recognizable voice across every episode.

1440p and Six Aspect Ratios in the MiniMax H3 API

Clips run from five to fifteen seconds, and the 1440p mode puts 1440 pixels on the short side between 16:9 and 9:16, or roughly 3.7 megapixels at wider ratios such as 2976 by 1248 for 21:9. Six ratios are selectable, from cinematic 21:9 to vertical 9:16, and the MiniMax H3 API can also choose one for you. That range covers a trailer, a product loop and a vertical drama without switching models.

Production Formats In, No Conversion Step

If your source files come straight off a camera or an editing timeline, they go in as they are: H.264 and H.265 video, JPG, PNG, WEBP, HEIC and HEIF stills, plus WAV and MP3 audio. Per-file limits sit at 50MB for video, 30MB for images and 15MB for audio, while passing assets by URL keeps requests within the 64MB body limit. On Atlas Cloud the whole set runs through one OpenAI-compatible key with pay-as-you-go billing.

Same Prompt, Three Engines: MiniMax H3 API Head to Head

Every clip in this set comes from one identical prompt sent to the MiniMax H3 API and two other video models hosted on Atlas Cloud, so motion, sound, and instruction fidelity can be compared without changing a single word.

Prompt

15 seconds, 16:9 landscape short video. Live-action footage of a late-night self-service laundromat, blended with hand-drawn glowing animation into a mixed-media image. A small self-service laundromat, its fluorescent lights faintly flickering; inside are running washing machines, plastic laundry baskets, and an old bench, with a single sock lying on the floor. The whole space is quiet, carrying a faint, nostalgic mood. It has the texture of one-handed handheld phone footage, with noticeable camera shake; the white fluorescent light causes the exposure to fluctuate between bright and dim; glass surfaces carry ambient reflections; there's a focus lag when the lens moves close to objects. The image should not be as polished and orderly as a commercial ad — the overall feel should be like a genuine documentary snapshot, as if you stumbled in by chance late at night and grabbed the shot while chasing some strange, dreamlike vision.

Generated with MiniMax H3 on Atlas Cloud

Generated with Seedance 2.0 on Atlas Cloud

Generated with Wan-2.7 on Atlas Cloud

Prompt

First-person perspective · eye-level height · handheld gaming camera Scene: The shot simulates a player operating a modern-warfare FPS game, both hands holding an assault rifle while slowly advancing along the outer perimeter of a military base. The player moves forward along a road beside cover, the crosshair sweeping across the passage ahead; after a brief pause, they fire a few rounds toward a distant objective, then continue pushing forward — like the live gameplay footage of an ordinary player. Lighting: The cool-toned natural light of a modern military base interweaves with smoke and muzzle fire. The image is realistic and crisp, with the metallic weapon and the battlefield dust and haze carrying a AAA-game quality. Camera work: The camera has a slight handheld sway as the player moves — first advancing slowly, then making small left-and-right sweeps to observe, with a subtle recoil shake when firing, before finally continuing to push steadily forward.

Generated with MiniMax H3 on Atlas Cloud

Generated with Seedance 2.0 on Atlas Cloud

Generated with Wan-2.7 on Atlas Cloud

MiniMax H3 Production Paths from Draft to Delivery

Choose MiniMax H3 Max for rapid text or image driven drafts, MiniMax H3 for 2K campaign and reference led work, or H3 Developer when generated audio belongs in the same production flow.

Rapid Drafting with MiniMax H3 Max

MiniMax H3 Max turns text prompts into 5 to 15 second clips at 480P or 768P. Product teams can test story beats and visual directions before selecting a final treatment.

First and Last Frame Product Motion

Start with a product image and optionally set the last frame through MiniMax H3 image to video. Merchandising teams can create launch teasers, catalog motion, and paid social variations for multiple placements.

2K Campaign Previews with MiniMax H3

Use H3 Developer to generate 2K video with audio from text or images for campaign previews. Agencies gain social spots, pitch sequences, and character moments that already include sound.

Recurring Characters Across Short Series

When identity matters across episodes, MiniMax H3 reference to video preserves the subject from supplied visual references. Studios can build recurring character shorts, episodic promos, and connected campaign scenes with continuity.

One Brief for Every Screen

Need landscape, square, and vertical deliverables? MiniMax H3 text to video supports 16:9, 1:1, 9:16, and adaptive framing, helping social teams tailor one concept for several placements without changing model families.

Prompt Based Footage Revisions

Revise existing MiniMax H3 footage by replacing a subject, background, or spoken line through prompt based editing. Post production teams can explore alternate cuts and targeted corrections without rebuilding every scene.

MiniMax H3 Models Compared with Leading Video APIs

Compare MiniMax H3 Max, H3 Developer, and H3 with leading text to video APIs across resolution, clip length, native audio, and standard per second pricing on Atlas Cloud.

ModelResolution CeilingClip LengthNative AudioStandard Price
MiniMax H3 Max Text-to-Video768P5 to 15s-$0.05/sec
MiniMax H3-Developer Text-to-Video2K-$0.05/sec
MiniMax H3 Text-to-Video2K5 to 15s-$0.08/sec
Wan-3.0-Prime Text-to-video1080P2 to 30s$0.068/sec
Seedance 2.5 Text-to-Video4K enhancedUp to 30s$0.167/sec
Kling V3.0 Turbo Text-to-Video1080p3 to 15s$0.112/sec

How to Use MiniMax H3 on Atlas Cloud

Get started in minutes — follow these simple steps to integrate and deploy models through Atlas Cloud's platform.

Create an Atlas Cloud Account

Sign up at atlascloud.ai and complete verification. New users receive free credits to explore the platform and test models.

Why Use MiniMax H3 on Atlas Cloud

Combining the advanced MiniMax H3 models with Atlas Cloud's GPU-accelerated platform provides unmatched performance, scalability, and developer experience.

Performance & flexibility

Low Latency:
GPU-optimized inference for real-time reasoning.

Unified API:
Run MiniMax H3, GPT, Gemini, and DeepSeek with one integration.

Transparent Pricing:
Predictable per-token billing with serverless options.

Enterprise & Scale

Developer Experience:
SDKs, analytics, fine-tuning tools, and templates.

Reliability:
99.99% uptime, RBAC, and compliance-ready logging.

Security & Compliance:
SOC 2 Type II, HIPAA alignment, data sovereignty in US.

MiniMax H3 API Choices, Limits, and Setup

MiniMax H3 is MiniMax's general purpose multimodal video model family for generating video from text, images, and reference media. Atlas Cloud provides standard H3, H3 Max, and H3 Developer variants through a unified API, allowing developers to select a route by modality, resolution, and deployment profile.

Start from a text prompt to create a cinematic clip, animate a first frame with an optional last frame, or use a supported reference route to preserve a subject from source media. These workflows cover product motion, short scenes, storyboard previews, marketing assets, and other video generation tasks.

Create an Atlas Cloud account and API key, select the exact model ID for your workflow, and construct the request using that model's current parameter schema. Submit it through the OpenAI-compatible Atlas Cloud API and process the returned video result in your application. Start building today.

Choose text to video when the scene starts from a written prompt, or image to video when an opening image and optional final frame should guide the result. Use a reference to video endpoint when source media must preserve a subject or influence the generated clip. H3 Max currently provides text and image routes on Atlas Cloud, while standard H3 and H3 Developer also include reference routes.

H3 Max text and image models support 480P or 768P clips lasting 5 to 15 seconds, while its text model offers 16:9, 9:16, 1:1, and adaptive aspect ratios. Standard H3 text, image, and reference models support 2K clips lasting 5 to 15 seconds, while H3 Developer covers 480P, 768P, and 2K. Check the selected model schema for route-specific ratio controls.

Yes. The image to video routes can animate a required first-frame image and optionally accept a last frame to guide how the clip concludes. This workflow is available through the H3 Max, standard H3, and H3 Developer image to video model IDs on Atlas Cloud.

Generated audio is explicitly supported by all three H3 Developer endpoints in the verified Atlas Cloud backend profiles. Reference inputs are available through standard H3 and H3 Developer reference to video routes, while H3 Max currently lists only text and image routes. Do not depend on audio from H3 Max unless its current schema or model page confirms that behavior.

The standard listed price is $0.05 per generated second for H3 Max and H3 Developer, and $0.08 per generated second for standard H3. Atlas Cloud uses pay-as-you-go billing, so the total depends on the selected model and requested output duration. Review the live model page before scheduling a large production batch.

H3 Max offers 480P and 768P text to video and image to video routes, while standard H3 provides 2K text, image, and reference workflows. H3 Developer is an Atlas Cloud self-hosted option spanning 480P, 768P, and 2K across all three workflows, with generated audio explicitly supported. Select the variant according to required resolution, input modality, audio needs, and standard per-second price.

Validate the JSON against the exact model schema and confirm the model identifier, because accepted parameters differ among variants. Next, check the API key, account balance, media URLs, and whether the selected endpoint exposes the requested capability. Correct validation or moderation issues before resubmitting, and retry temporary service failures cautiously.

Explore More Families

Seedance 2.5

Seedance 2.5 API is now available on Atlas Cloud! It gives developers ByteDance's newest video model. It generates up to 30 seconds of native video in a single pass from text, a single image, or as many as 50 multimodal references, with synchronized audio and in-frame multilingual text. On Atlas Cloud you reach it through one key, with subject consistency and improved physics keeping long shots coherent. (Update: Seedance 2.5 1080P API Is Available NOW!)

View Family

Wan 3.0

Wan 3.0 API is the next generation of Alibaba's Wan video family, built to push long-form generation, multi-reference control, and audiovisual quality to new heights. Atlas Cloud already hosts Wan 2.7, 2.6, and 2.5, and Wan 3.0 runs on the same unified key with no separate setup. Start building today. Scroll down to the showcase to see what Wan 3.0 can create.

View Family

MiniMax H3

MiniMax H3 is MiniMax’s video model family, spanning H3, H3 Max, and H3 Developer. Create from text, animate a first frame with an optional last frame, or preserve subjects from references. H3 and H3 Developer reach 2K, while H3 Max supports 480P and 768P clips lasting 5 to 15 seconds. Atlas Cloud adds OpenAI-compatible access and transparent pay-as-you-go pricing from $0.05 per second. Start building today.

View Family

Seedream 5.0 Pro

Seedream 5.0 Pro API gives developers ByteDance's controllable image editing model on Atlas Cloud. It places edits precisely with anchors and coordinates, separates images into editable layers, fuses multiple references, and matches exact colors and materials, with multilingual text at 2K and 3K. On Atlas Cloud you reach it through one key!

View Family

Seedance 2.0

The Seedance 2.0 API gives you production access to ByteDance's multimodal video model — quad-modal inputs (text, image, video, audio) and an industry-leading "Universal Reference" system that locks composition, camera movement, and character actions across shots. Integrate director-level control with one API call, a flat $0.09/s, instant key, and no waitlist — backed by enterprise-grade uptime and compliance. Seedance 2.0 Native 4K is now live!

View Family

GPT Image 2

The GPT Image 2 API gives developers access to OpenAI's latest image model, the successor to GPT Image 1.5. It generates and edits images with accurate text rendering across Latin and CJK scripts, plus strong composition for posters, mockups, and infographics. On Atlas Cloud you reach it through one unified API alongside 300+ models, with free credits, 99.99% uptime, and no OpenAI organization verification required.

View Family

Gemini Omni Flash

The gemini omni API brings Google DeepMind's natively multimodal Gemini Omni Flash family, including Gemini Omni 1.1 Flash, to developers. Create cinematic video with synchronized native audio, animate still images with precise start and end frame control, or revise existing footage through text guided edits that preserve untouched content. Atlas Cloud provides one OpenAI-compatible key, unified access, and transparent pay-as-you-go pricing. Start building today.

View Family

Grok Imagine

The Grok Imagine API covers xAI's image, video, and speech models, from Image 2.0 to Video 1.5 and xAI TTS v1. Render 1K or 2K stills across 14 aspect ratios, push a scene to 15 seconds of 1080p motion, steer shots with up to 7 reference images, or narrate them in 20 languages. Atlas Cloud runs every mode on one endpoint, priced pay-as-you-go from $0.02 per image and $0.05 per second. Start building today.

View Family

Google

Google's most powerful creative models are all available on Atlas Cloud. Veo 3.1 delivers cinematic video generation, Nano Banana 2 powers high-fidelity image creation, and Gemini brings multimodal intelligence to every workflow. Access the full Google model suite through one API key with Day-0 availability and pay-as-you-go pricing.

View Family

Seedance 2.0 Mini

The Seedance 2.0 Mini API is the lightest, lowest-cost tier of ByteDance's Seedance video line, built for teams where throughput and unit cost matter more than maximum polish. Use it for batch generation, rapid prototyping, and draft passes, all through one OpenAI-compatible key on Atlas Cloud.

View Family

ByteDance

From cinematic video generation to high-fidelity image creation, ByteDance's most powerful models are live on Atlas Cloud. Run Seedance and Seedream at scale with the lowest inference pricing and zero infrastructure overhead.

View Family

Alibaba

Atlas Cloud brings together Alibaba's full model lineup under one API: Qwen for language and image tasks, Wan for video generation up to 1080p. Access every model pay-as-you-go with no subscriptions. The Alibaba API is available via a single base URL using your existing OpenAI-compatible client.

View Family

One API for All Media AI.

Explore all models