Image and Video Generation API Reference
AIAPIAIAPI unified asynchronous image and video generation API documentation
Endpoints
Image: POST /v1/images/generations/async
Video: POST /v1/video/generations
Authentication
Authorization: Bearer <AIAPIAIAPI_API_KEY>
Flow
For image tasks, use the returned task_id with GET /v1/images/generations/async/{task_id}; read outputs when status is SUCCESS.
For video tasks, use the returned id with GET /v1/video/generations/{task_id}; read outputs[0].url when status is completed.
Error Responses
Image APIs return code/message/data; video APIs return error.type/error.code/error.message/error.param.
Model API Reference
| Model | Type | Task | Description |
|---|---|---|---|
alibaba/qwen-image-edit-spicy | image | image-to-image | Text-guided image editing for adding, removing, or modifying elements, style transfer, and background changes. |
alibaba/wan-2.2-i2v-spicy | video | image-to-video | Image-to-video generation from a required first-frame image, with optional last-frame interpolation. |
alibaba/wan-2.7-i2v | video | image-to-video | Image-to-video generation with optional end-frame control. Resolution adapts to the input image, so aspect_ratio is not configurable. |
alibaba/wan-2.7-i2v-spicy | video | image-to-video | Image-to-video generation. Animates a single source image from a text prompt, with optional end-frame control. |
alibaba/wan-2.7-i2v-uncensored | video | image-to-video | Image-to-video generation with the same parameters as wan-2.7-i2v-spicy, without content moderation. |
alibaba/wan-2.7-ref2v-uncensored | video | reference-to-video | Reference-to-video generation from up to 3 images and 3 videos. |
alibaba/wan-2.7-t2v | video | text-to-video | Text-to-video generation with no source image required. |
alibaba/wan-3.0-i2v-spicy | video | image-to-video | Image-to-video generation with a wider duration range, more resolutions, and configurable aspect ratio. |
alibaba/wan-3.0-prime-i2v-spicy | video | image-to-video | Image-to-video generation with a wider duration range, more resolutions, and configurable aspect ratio. |
alibaba/wan-3.0-prime-ref2v-spicy | video | reference-to-video | Reference-to-video generation from images, videos, and/or audio references, with configurable duration/resolution/aspect ratio. |
alibaba/wan-3.0-ref2v-spicy | video | reference-to-video | Reference-to-video generation from images, videos, and/or audio references, with configurable duration/resolution/aspect ratio. |
alibaba/wan-3.0-video | video | text-image-reference-to-video | One endpoint selects text-to-video, keyframe image-to-video, or reference-to-video from the media fields you send. Keyframe and reference inputs cannot be mixed. |
alibaba/z-image-spicy | image | text-to-image | Text-to-image generation with custom dimensions, reproducible seeds, and optional intelligent prompt rewriting. |
alibaba/z-image-spicy-pro | image | text-to-image | Text-to-image generation with dimensions up to 2560 pixels, reproducible seeds, and optional prompt rewriting. |
bytedance/seedance-2.0-i2v-spicy | video | image-to-video | Image-to-video generation with resolutions up to 4K and 7 aspect ratio presets, including 21:9. |
bytedance/seedance-2.0-mini-i2v-spicy | video | image-to-video | Image-to-video generation from a required first-frame image, with optional last-frame control, at 720p or 1080p for 4–15 seconds. |
bytedance/seedance-2.0-mini-ref2v-spicy | video | reference-to-video | Reference-to-video generation from up to 9 images, 3 videos and 3 audio clips, at 720p or 1080p for 4–15 seconds. |
bytedance/seedance-2.0-ref2v-spicy | video | reference-to-video | Reference-to-video generation from images, videos, and/or audio references, with resolutions up to 4K. |
bytedance/seedance-2.5-i2v-spicy | video | image-to-video | Image-to-video generation with an extended duration range of up to 30 seconds. |
bytedance/seedance-2.5-ref2v-spicy | video | reference-to-video | Reference-to-video generation from up to 30 images, 10 videos, and 10 audio clips, with an auto-duration mode. |
minimax-h3-enhanced-reference-to-video | video | reference-to-video | Reference-to-video generation from up to 9 images, 9 videos, and 9 audio clips at resolutions up to 4K. |
minimax-h3-image-to-video | video | image-to-video | Image-to-video generation with 5-15 second output at 768p or 2K resolution. |
minimax-h3-max-image-to-video | video | image-to-video | MiniMax H3 Max image-to-video generation with 5-15 second output at 480p or 768p resolution. |
minimax-h3-max-reference-to-video | video | reference-to-video | MiniMax H3 Max reference-to-video generation through Venice. Model-specific duration, resolution, and reference limits are validated by the upstream provider. |
minimax-h3-max-turbo-image-to-video | video | image-to-video | MiniMax H3 Max Turbo image-to-video generation through Venice. Model-specific duration and resolution limits are validated by the upstream provider. |
minimax-h3-reference-to-video | video | reference-to-video | Reference-to-video generation from up to 9 images, 9 videos, and 9 audio clips at 768p or 2K resolution. |
minimax/minimax-h3-i2v-spicy | video | image-to-video | Image-to-video generation with 5-15 second output at 768p, 2K, or 4K resolution. |
minimax/minimax-h3-ref2v-spicy | video | reference-to-video | Reference-to-video generation from up to 9 images, 9 videos, and 9 audio clips at resolutions up to 4K. |
openai/gpt-image-2.5-flare/edit | image | image-to-image | Image editing from 1 to 16 reference images, with an optional mask and transparent-background output. |
openai/gpt-image-2.5-flare/text-to-image | image | text-to-image | Text-to-image generation with custom dimensions, quality tiers, and transparent backgrounds. |
openai/gpt-image-2.5-sunburst/edit | image | image-to-image | Image editing from 1 to 16 reference images, with an optional mask and transparent-background output. |
openai/gpt-image-2.5-sunburst/text-to-image | image | text-to-image | Text-to-image generation with custom dimensions, quality tiers, and transparent backgrounds. |
wan-3-0-prime-image-to-video | video | image-to-video | Image-to-video generation with a wider duration range, more resolutions, and configurable aspect ratio. |
wan-3-0-prime-reference-to-video | video | reference-to-video | Reference-to-video generation from images, videos, and/or audio references, with configurable duration/resolution/aspect ratio. |
Documentation Index
Fetch the complete documentation index at:
https://aiapiaiapi.com/llms.txt
Use this file to discover all available API documentation.