Explore and test the latest AI models for image generation, video creation, audio synthesis, and 3D modelling. Try FLUX, Kling, Veo, Seedance, Nano Banana, Wan, Suno, and 500+ other models directly in your browser — pay per generation, no subscription required.
All models are accessible via the MuAPI REST API with a single API key. No per-provider accounts needed. Pay per generation and switch models freely.
FLUX 3 Image-to-Video animates a still image into a cinematic clip with optional native synchronized audio, using Black Forest Labs' unified image/video/audio architecture. Motion stays physically grounded and consistent with the source frame, making it suited for product animation, portrait bring-to-life effects, and scene extension.
FLUX 3 Text-to-Video generates cinematic video clips with optional native synchronized audio from a single unified model — the same architecture Black Forest Labs uses for FLUX 3's action-prediction research. Expect coherent motion, strong physical plausibility, and scene-appropriate ambient sound baked directly into generation.

FLUX 3 Dev is the faster, lower-cost variant of Black Forest Labs' FLUX 3 frontier model, planned for open-weight release. It trades a small amount of peak fidelity for significantly reduced latency and price, making it well suited for rapid iteration, prototyping, and high-volume text-to-image generation.

FLUX 3 Image-to-Image edits and restyles existing images using a text instruction plus up to several reference images. Built on Black Forest Labs' unified multimodal architecture, it preserves subject identity and scene structure while applying precise, prompt-driven edits — ideal for product retouching, style transfer, and character-consistent edits.

FLUX 3 Text-to-Image is Black Forest Labs' next-generation multimodal frontier model, jointly trained across image, video, and audio for outputs that are truer to life in every style. It generates highly photorealistic and stylistically flexible images from text prompts, with sharper detail, more coherent composition, and stronger prompt adherence than the FLUX.2 generation.
Video Background Remover automatically removes the background from any video, producing a clean cutout of the subject with a transparent or solid-color backdrop. It handles hair, edges, and fine detail with frame-accurate matting, supports videos up to 60 seconds, and can output transparent WebM/MOV, standard MP4, or animated GIF while optionally preserving the original audio.
Gemini 2.5 Pro TTS is Google's premium text-to-speech model for studio-quality, high-fidelity multi-speaker audio with expressive control over voice, accent, emotional style, and pace.
Gemini 3.1 Flash TTS turns written dialogue into expressive, natural multi-speaker speech with fine-grained control over voice, accent, emotional style, and pace. Ideal for fast, affordable voiceovers, character dialogue, and narration.
Seedance 2 Mini Spicy Image-to-Video is the fastest, lowest-cost Spicy-tier image animation, with reduced content-safety filtering on top of Seedance 2 Mini's speed and pricing.
Seedance 2 Mini Spicy Text-to-Video is the fastest, lowest-cost Spicy-tier text-to-video generation, with reduced content-safety filtering on top of Seedance 2 Mini's speed and pricing.
Seedance 2 Spicy Image-to-Video Fast by ByteDance. The quickest Spicy-tier image animation, with reduced content-safety filtering and the same fast queue as Seedance 2 VIP Fast.
Seedance 2 Spicy Image-to-Video by ByteDance. Animates a start frame into cinematic video with VIP-tier priority routing and up to 2K resolution, with reduced content-safety filtering for more creative freedom.
Seedance 2 Spicy Text-to-Video Fast by ByteDance. The quickest Spicy-tier text-to-video generation, with reduced content-safety filtering and the same fast queue as Seedance 2 VIP Fast.
Seedance 2 Spicy Text-to-Video by ByteDance. Same VIP-tier priority routing, native audio-visual sync, and up to 2K resolution as Seedance 2 VIP, with reduced content-safety filtering for more creative freedom.
Seedance 2.5 Spicy Text-to-Video is the relaxed-moderation variant of the Seedance 2.5 flagship model, generating photorealistic 4K cinematic video directly from a text prompt with reduced content-safety filtering and more dramatic, higher-contrast motion than the standard tier — while retaining native audio synthesis and extended clip durations.
Seedance 2.5 Spicy Image-to-Video is the relaxed-moderation variant of the Seedance 2.5 flagship model. It animates a single image into photorealistic 4K video with reduced content-safety filtering and more dramatic, higher-contrast motion than the standard tier — while keeping native audio generation and precise camera trajectory control.

Gemini Audio Vision uses Google Gemini's native audio understanding to analyze and describe audio content in detail — speech, tone, background sounds, speaker changes, and more. Upload an audio URL and a prompt, and Gemini returns a detailed text analysis. Token-based pricing.

Gemini Video Vision uses Google Gemini's native video understanding to analyze and describe video content in detail — motion, composition, subjects, on-screen text, and more. Upload a video URL and a prompt, and Gemini returns a detailed text analysis. Token-based pricing.

Update the title, description, tags, category, privacy, or made-for-kids status of a video on a connected YouTube account.

Set or replace the thumbnail image of a video on a connected YouTube account.
Drive a video's lip movements to match a target audio track, producing a lip-synced video output.

Detect and extract text fragments and their positions from an image using local OCR.

GPT 5.6 Sol is OpenAI's flagship reasoning model, optimized for complex math, programming, and scientific research. Supports mixed text and image inputs, adjustable reasoning effort, and integrated web search tools. Pricing: $10.00/M input tokens, $60.00/M output tokens.

GPT 5.6 Terra is OpenAI's balanced multimodal reasoning model for general business and analytical tasks. Supports mixed text and image inputs, adjustable reasoning effort, and integrated web search tools. Pricing: $5.00/M input tokens, $30.00/M output tokens.