Skip to main content
Model ID
Calling method: async

Wan-AI_Wan2.1-I2V-14B-720P API Usage Guide

Overview

Wan-AI_Wan2.1-I2V-14B-720P is a comprehensive and open video foundation model that transforms static images into dynamic videos. After thousands of rounds of human evaluations, this model has outperformed both closed-source and open-source alternatives, achieving state-of-the-art performance.

Key Features:

  • SOTA Performance: Consistently outperforms existing open-source models and state-of-the-art commercial solutions across multiple benchmarks
  • High Definition: Creates high-quality videos at 720P resolution
  • Visual Text Generation: Capable of generating both Chinese and English text with robust text generation capabilities
  • Multi-GPU Support: Optimized for both single and multi-GPU inference using FSDP + xDiT USP technology

Use Cases:

  • Animating static images
  • Creating dynamic presentations from still photos
  • Bringing artwork and photographs to life
  • Marketing and advertising content

Authentication

All API requests require authentication using an API key. Include your API key in the Authorization header:

Submit Video Generation Request

Endpoint

Request Format

Request Parameters

Response

Check Request Status

Endpoint

Example

Response

Request Status Values

List Your Requests

Endpoint

Example

Get Model Information

Endpoint

Example

List Available Models

Endpoint

Example

Response

Pricing

  • Pricing Type: Video length based pricing
  • Price: $0.60 per second
  • Unit: Second
Example cost calculation:
  • 5-second video: 5 × 0.60=0.60 = 3.00
  • 10-second video: 10 × 0.60=0.60 = 6.00

Video Specifications

  • Duration: 5-10 seconds
  • Resolution: Fixed at 720P (1280x720)
  • Quality: High definition with excellent detail and clarity

Tips for Better Results

  1. High-Quality Input Images: Use clear, well-lit images with good composition
  2. Detailed Prompts: Provide comprehensive descriptions of desired motion and style
  3. Prompt Extension: Enable for enhanced results with more detailed scene generation
  4. Seed Usage:
    • Use -1 for random generation (default)
    • Use specific values (0-2147483647) for reproducible results
  5. Video Length:
    • 5 seconds: Quick animations, social media content
    • 10 seconds: More complex motion sequences, detailed storytelling
  6. Image Requirements:
    • Use high-quality images with clear subjects
    • Avoid overly complex or cluttered images
    • Ensure good lighting and contrast
  7. High Definition Benefits:
    • Better detail preservation
    • Sharper motion sequences
    • Professional quality output

Parameter Examples

High-Quality Landscape Animation

Professional Portrait Animation

Artistic Transformation

Product Animation

Nature Animation

Advanced Features

Prompt Extension

Enable advanced prompt extension for enhanced scene generation:

Reproducible Results

Use specific seeds for consistent video generation:

Model Architecture Details

  • Architecture: 14B parameter Diffusion Transformer
  • Dimension: 5120
  • Heads: 40
  • Layers: 40
  • Resolution: 720P (1280x720)
  • VRAM Requirement: 24GB+ recommended (supports multi-GPU inference)

Best Practices

  1. Image Quality: Use high-resolution, well-composed images for best results
  2. Motion Description: Be specific about the type of motion you want to see
  3. Prompt Extension: Enable for complex scenes requiring detailed generation
  4. Seed Management: Use consistent seeds for series of related videos
  5. High Definition: This 720P model provides superior quality for professional use
  6. Hardware Requirements: Ensure adequate GPU resources for optimal performance

Image Input Guidelines

Image Requirements

  • Quality: High-resolution images work best
  • Composition: Clear subjects with good framing
  • Lighting: Well-lit images with good contrast
  • Content: Avoid overly complex or cluttered scenes