Our flagship video model, with incredible cinematic motion. Clips up to 30 seconds at 1080p with audio, from text, up to 50 references or Characters, or a source video to edit.
Our most powerful video model, now up to 30 seconds
Our flagship video model, with incredible cinematic motion. Clips up to 30 seconds at 1080p with audio, from text, up to 50 references or Characters, or a source video to edit.
Our most capable image model. Keeps a character identical across a whole set, edits exactly what you ask and nothing else, and merges up to 10 references or Characters into one scene, at up to 2K.
Our photorealism model at its very best, for portraits, fashion, and product shots with true-to-life skin, fabric, and light. Takes up to 3 references or Characters, up to 2K.
OpenAI's quick, precise image model that follows the prompt to the letter and renders readable text. Generates or edits from up to 15 reference images. OpenAI's safety filter applies.
Google's image model for graphics and marketing content with legible text. Generates or edits from up to 14 reference images at up to 4K. Google's safety filter applies.
A fast open video model that generates picture and stereo audio together. Start and end frames or up to 8 references (longer clips take fewer), 5 to 15 seconds at up to 768p.
Speech, dialogue, music, and sound effects from a single prompt, as an MP3 of exactly the length you ask, 5 seconds to 2 minutes. Match a voice from up to 3 audio clips.
xAI's Grok Imagine image model in three quality tiers. Generates or edits from up to 3 reference images while keeping the source's details intact. xAI's safety filter applies.
Video that stays faithful to every reference you give it. A start frame, up to 9 references or Characters, or a source video to edit; 3 to 15 seconds at up to 1080p.
Our anime video model, as Melon and Melon Pro. Start and end frames or up to 7 references or Characters, optional audio, 3 to 16 seconds at up to 1080p.