

MiniMax H3 is MiniMax’s video model family, spanning H3, H3 Max, and H3 Developer. Create from text, animate a first frame with an optional last frame, or preserve subjects from references. H3 and H3 Developer reach 2K, while H3 Max supports 480P and 768P clips lasting 5 to 15 seconds. Atlas Cloud adds OpenAI-compatible access and transparent pay-as-you-go pricing from $0.05 per second. Start building today.
MiniMax H3 is developed by MiniMax. Atlas Cloud (operated by Atlas Cloud AI LLC) provides access to it and does not own it. All trademarks belong to their respective owners.
Atlas Cloud provides you with the latest industry-leading creative models.
Compare text, image, and reference driven MiniMax H3 modes across H3 Max, H3 Developer, and H3 by resolution, duration, aspect ratio, audio support, and standard pricing.
| Modality | Description |
|---|---|
| MiniMax H3 Max T2V API (Text to Video) | Turn a text prompt into a cinematic video at 480P or 768P, with durations from 5 to 15 seconds. Select 16:9, 9:16, 1:1, or adaptive framing for trailers, advertisements, and social video at the standard price of $0.05 per second. |
| MiniMax H3 Max I2V API (Image to Video) | Starting from a first frame, the model generates a prompt directed video and can optionally use a final frame to guide the ending. Its 480P and 768P output supports 5 to 15 second product animations, storyboard shots, and image based narratives at a standard $0.05 per second. |
| MiniMax H3 Developer T2V API (Text to Video) | Built as a self hosted text to video model, this endpoint converts prompts into videos with generated audio at 480P, 768P, or 2K. Choose 16:9, 9:16, or 1:1 for narrative scenes, campaign concepts, and social content at the standard price of $0.05 per second. |
| MiniMax H3 Developer I2V API (Image to Video) | Provide an opening image, an optional last frame, and a text prompt to create motion with generated audio. Output at 480P, 768P, or 2K makes the endpoint suitable for animating key art, product stills, and storyboard sequences, with standard pricing of $0.05 per second. |
| MiniMax H3 Developer Reference to Video API | When subject continuity matters, one or more reference images or videos can guide a new prompt driven video with generated audio. Available 480P, 768P, and 2K output supports recurring characters, branded subjects, and connected visual sequences at the standard price of $0.05 per second. |
| MiniMax H3 T2V API (Text to Video) | A text prompt becomes a cinematic 2K video lasting from 5 to 15 seconds. Frame the result in 16:9, 9:16, 1:1, or an adaptive aspect ratio for polished trailers, vertical stories, and square campaign assets at the standard price of $0.08 per second. |
| MiniMax H3 I2V API (Image to Video) | Use a first frame and text prompt to produce a 2K video, adding an optional last frame when the final composition needs direction. The 5 to 15 second range fits product reveals, character motion, and storyboard transitions, with standard pricing of $0.08 per second. |
| MiniMax H3 Reference to Video API | Reference driven generation preserves the subject from an input image while a text prompt directs the resulting 2K video. Clips can run from 5 to 15 seconds, making this endpoint useful for character focused shots and consistent product visuals at the standard price of $0.08 per second. |
Every clip the MiniMax H3 API returns runs up to fifteen seconds at 1440p with native stereo sound, built from any mix of text, image, video and voice references, and the same call can also edit characters, scenes and dialogue inside footage you already have.
One MiniMax H3 API request accepts up to nine reference images, three video clips and three audio tracks, capped at twelve files. Rather than reading them as separate slots, the model treats text, picture and sound as one context, pulling a character from a photo, a camera move from a clip and a mood from a track into the same scene. Teams with existing material get a direct route to a finished shot.
Every result ships with sound. Dialogue, ambience and music are produced in native stereo during the same pass that renders the picture at 24 frames per second, so nothing has to be scored or dubbed afterwards. Because timing is decided while the shot is generated, footsteps, speech and cuts stay locked to the action. Short ads and social spots come out ready to publish.
Need a cat swapped for a dog, a green screen replaced, or one line of dialogue rewritten? The MiniMax H3 API applies edits like these to video you supply, covering characters, objects, backgrounds, lighting and effects, while parts you did not mention stay close to the original. Prompts run to 7000 characters, so a dozen changes can be stacked into one pass. That makes iterating on an approved cut practical.
Send a voice sample alongside your images or footage and the generated character speaks in that timbre. Up to three audio references are allowed per request, each between two and fifteen seconds, and audio must always accompany a visual input rather than arrive on its own. Existing dialogue can also be replaced and the performance adjusted to match. Series work keeps one recognizable voice across every episode.
Clips run from five to fifteen seconds, and the 1440p mode puts 1440 pixels on the short side between 16:9 and 9:16, or roughly 3.7 megapixels at wider ratios such as 2976 by 1248 for 21:9. Six ratios are selectable, from cinematic 21:9 to vertical 9:16, and the MiniMax H3 API can also choose one for you. That range covers a trailer, a product loop and a vertical drama without switching models.
If your source files come straight off a camera or an editing timeline, they go in as they are: H.264 and H.265 video, JPG, PNG, WEBP, HEIC and HEIF stills, plus WAV and MP3 audio. Per-file limits sit at 50MB for video, 30MB for images and 15MB for audio, while passing assets by URL keeps requests within the 64MB body limit. On Atlas Cloud the whole set runs through one OpenAI-compatible key with pay-as-you-go billing.
Every clip in this set comes from one identical prompt sent to the MiniMax H3 API and two other video models hosted on Atlas Cloud, so motion, sound, and instruction fidelity can be compared without changing a single word.
15 seconds, 16:9 landscape short video. Live-action footage of a late-night self-service laundromat, blended with hand-drawn glowing animation into a mixed-media image. A small self-service laundromat, its fluorescent lights faintly flickering; inside are running washing machines, plastic laundry baskets, and an old bench, with a single sock lying on the floor. The whole space is quiet, carrying a faint, nostalgic mood. It has the texture of one-handed handheld phone footage, with noticeable camera shake; the white fluorescent light causes the exposure to fluctuate between bright and dim; glass surfaces carry ambient reflections; there's a focus lag when the lens moves close to objects. The image should not be as polished and orderly as a commercial ad — the overall feel should be like a genuine documentary snapshot, as if you stumbled in by chance late at night and grabbed the shot while chasing some strange, dreamlike vision.
Generated with MiniMax H3 on Atlas Cloud
Generated with Seedance 2.0 on Atlas Cloud
Generated with Wan-2.7 on Atlas Cloud
First-person perspective · eye-level height · handheld gaming camera Scene: The shot simulates a player operating a modern-warfare FPS game, both hands holding an assault rifle while slowly advancing along the outer perimeter of a military base. The player moves forward along a road beside cover, the crosshair sweeping across the passage ahead; after a brief pause, they fire a few rounds toward a distant objective, then continue pushing forward — like the live gameplay footage of an ordinary player. Lighting: The cool-toned natural light of a modern military base interweaves with smoke and muzzle fire. The image is realistic and crisp, with the metallic weapon and the battlefield dust and haze carrying a AAA-game quality. Camera work: The camera has a slight handheld sway as the player moves — first advancing slowly, then making small left-and-right sweeps to observe, with a subtle recoil shake when firing, before finally continuing to push steadily forward.
Generated with MiniMax H3 on Atlas Cloud
Generated with Seedance 2.0 on Atlas Cloud
Generated with Wan-2.7 on Atlas Cloud
Choose MiniMax H3 Max for rapid text or image driven drafts, MiniMax H3 for 2K campaign and reference led work, or H3 Developer when generated audio belongs in the same production flow.
MiniMax H3 Max turns text prompts into 5 to 15 second clips at 480P or 768P. Product teams can test story beats and visual directions before selecting a final treatment.
Start with a product image and optionally set the last frame through MiniMax H3 image to video. Merchandising teams can create launch teasers, catalog motion, and paid social variations for multiple placements.
Use H3 Developer to generate 2K video with audio from text or images for campaign previews. Agencies gain social spots, pitch sequences, and character moments that already include sound.
When identity matters across episodes, MiniMax H3 reference to video preserves the subject from supplied visual references. Studios can build recurring character shorts, episodic promos, and connected campaign scenes with continuity.
Need landscape, square, and vertical deliverables? MiniMax H3 text to video supports 16:9, 1:1, 9:16, and adaptive framing, helping social teams tailor one concept for several placements without changing model families.
Revise existing MiniMax H3 footage by replacing a subject, background, or spoken line through prompt based editing. Post production teams can explore alternate cuts and targeted corrections without rebuilding every scene.
Compare MiniMax H3 Max, H3 Developer, and H3 with leading text to video APIs across resolution, clip length, native audio, and standard per second pricing on Atlas Cloud.
| Model | Resolution Ceiling | Clip Length | Native Audio | Standard Price |
|---|---|---|---|---|
| MiniMax H3 Max Text-to-Video | 768P | 5 to 15s | - | $0.05/sec |
| MiniMax H3-Developer Text-to-Video | 2K | - | √ | $0.05/sec |
| MiniMax H3 Text-to-Video | 2K | 5 to 15s | - | $0.08/sec |
| Wan-3.0-Prime Text-to-video | 1080P | 2 to 30s | √ | $0.068/sec |
| Seedance 2.5 Text-to-Video | 4K enhanced | Up to 30s | √ | $0.167/sec |
| Kling V3.0 Turbo Text-to-Video | 1080p | 3 to 15s | √ | $0.112/sec |
Get started in minutes — follow these simple steps to integrate and deploy models through Atlas Cloud's platform.
Sign up at atlascloud.ai and complete verification. New users receive free credits to explore the platform and test models.
Combining the advanced MiniMax H3 models with Atlas Cloud's GPU-accelerated platform provides unmatched performance, scalability, and developer experience.
Low Latency:
GPU-optimized inference for real-time reasoning.
Unified API:
Run MiniMax H3, GPT, Gemini, and DeepSeek with one integration.
Transparent Pricing:
Predictable per-token billing with serverless options.
Developer Experience:
SDKs, analytics, fine-tuning tools, and templates.
Reliability:
99.99% uptime, RBAC, and compliance-ready logging.
Security & Compliance:
SOC 2 Type II, HIPAA alignment, data sovereignty in US.
MiniMax H3 is MiniMax's general purpose multimodal video model family for generating video from text, images, and reference media. Atlas Cloud provides standard H3, H3 Max, and H3 Developer variants through a unified API, allowing developers to select a route by modality, resolution, and deployment profile.
Start from a text prompt to create a cinematic clip, animate a first frame with an optional last frame, or use a supported reference route to preserve a subject from source media. These workflows cover product motion, short scenes, storyboard previews, marketing assets, and other video generation tasks.
Create an Atlas Cloud account and API key, select the exact model ID for your workflow, and construct the request using that model's current parameter schema. Submit it through the OpenAI-compatible Atlas Cloud API and process the returned video result in your application. Start building today.
Choose text to video when the scene starts from a written prompt, or image to video when an opening image and optional final frame should guide the result. Use a reference to video endpoint when source media must preserve a subject or influence the generated clip. H3 Max currently provides text and image routes on Atlas Cloud, while standard H3 and H3 Developer also include reference routes.
H3 Max text and image models support 480P or 768P clips lasting 5 to 15 seconds, while its text model offers 16:9, 9:16, 1:1, and adaptive aspect ratios. Standard H3 text, image, and reference models support 2K clips lasting 5 to 15 seconds, while H3 Developer covers 480P, 768P, and 2K. Check the selected model schema for route-specific ratio controls.
Yes. The image to video routes can animate a required first-frame image and optionally accept a last frame to guide how the clip concludes. This workflow is available through the H3 Max, standard H3, and H3 Developer image to video model IDs on Atlas Cloud.
Generated audio is explicitly supported by all three H3 Developer endpoints in the verified Atlas Cloud backend profiles. Reference inputs are available through standard H3 and H3 Developer reference to video routes, while H3 Max currently lists only text and image routes. Do not depend on audio from H3 Max unless its current schema or model page confirms that behavior.
The standard listed price is $0.05 per generated second for H3 Max and H3 Developer, and $0.08 per generated second for standard H3. Atlas Cloud uses pay-as-you-go billing, so the total depends on the selected model and requested output duration. Review the live model page before scheduling a large production batch.
H3 Max offers 480P and 768P text to video and image to video routes, while standard H3 provides 2K text, image, and reference workflows. H3 Developer is an Atlas Cloud self-hosted option spanning 480P, 768P, and 2K across all three workflows, with generated audio explicitly supported. Select the variant according to required resolution, input modality, audio needs, and standard per-second price.
Validate the JSON against the exact model schema and confirm the model identifier, because accepted parameters differ among variants. Next, check the API key, account balance, media URLs, and whether the selected endpoint exposes the requested capability. Correct validation or moderation issues before resubmitting, and retry temporary service failures cautiously.
Guides, tutorials, and product updates to help you get the most out of Atlas Cloud.