Wan-AI_Wan2.1-I2V-14B-720P API Usage Guide
Overview
Wan-AI_Wan2.1-I2V-14B-720P is a comprehensive and open video foundation model that transforms static images into dynamic videos. After thousands of rounds of human evaluations, this model has outperformed both closed-source and open-source alternatives, achieving state-of-the-art performance.Key Features:
- SOTA Performance: Consistently outperforms existing open-source models and state-of-the-art commercial solutions across multiple benchmarks
- High Definition: Creates high-quality videos at 720P resolution
- Visual Text Generation: Capable of generating both Chinese and English text with robust text generation capabilities
- Multi-GPU Support: Optimized for both single and multi-GPU inference using FSDP + xDiT USP technology
Use Cases:
- Animating static images
- Creating dynamic presentations from still photos
- Bringing artwork and photographs to life
- Marketing and advertising content
Authentication
All API requests require authentication using an API key. Include your API key in the Authorization header:Submit Video Generation Request
Endpoint
Request Format
Request Parameters
Response
Check Request Status
Endpoint
Example
Response
Request Status Values
List Your Requests
Endpoint
Example
Get Model Information
Endpoint
Example
List Available Models
Endpoint
Example
Response
Pricing
- Pricing Type: Video length based pricing
- Price: $0.60 per second
- Unit: Second
- 5-second video: 5 × 3.00
- 10-second video: 10 × 6.00
Video Specifications
- Duration: 5-10 seconds
- Resolution: Fixed at 720P (1280x720)
- Quality: High definition with excellent detail and clarity
Tips for Better Results
- High-Quality Input Images: Use clear, well-lit images with good composition
- Detailed Prompts: Provide comprehensive descriptions of desired motion and style
- Prompt Extension: Enable for enhanced results with more detailed scene generation
- Seed Usage:
- Use -1 for random generation (default)
- Use specific values (0-2147483647) for reproducible results
- Video Length:
- 5 seconds: Quick animations, social media content
- 10 seconds: More complex motion sequences, detailed storytelling
- Image Requirements:
- Use high-quality images with clear subjects
- Avoid overly complex or cluttered images
- Ensure good lighting and contrast
- High Definition Benefits:
- Better detail preservation
- Sharper motion sequences
- Professional quality output
Parameter Examples
High-Quality Landscape Animation
Professional Portrait Animation
Artistic Transformation
Product Animation
Nature Animation
Advanced Features
Prompt Extension
Enable advanced prompt extension for enhanced scene generation:Reproducible Results
Use specific seeds for consistent video generation:Model Architecture Details
- Architecture: 14B parameter Diffusion Transformer
- Dimension: 5120
- Heads: 40
- Layers: 40
- Resolution: 720P (1280x720)
- VRAM Requirement: 24GB+ recommended (supports multi-GPU inference)
Best Practices
- Image Quality: Use high-resolution, well-composed images for best results
- Motion Description: Be specific about the type of motion you want to see
- Prompt Extension: Enable for complex scenes requiring detailed generation
- Seed Management: Use consistent seeds for series of related videos
- High Definition: This 720P model provides superior quality for professional use
- Hardware Requirements: Ensure adequate GPU resources for optimal performance
Image Input Guidelines
Image Requirements
- Quality: High-resolution images work best
- Composition: Clear subjects with good framing
- Lighting: Well-lit images with good contrast
- Content: Avoid overly complex or cluttered scenes