Happy Horse 1.1
An Alibaba model that turns a single image into a video with synchronized audio and multilingual lip-sync.
Happy Horse 1.1 is part of Alibaba's image-to-video model family and can turn a single frame into a clip with synchronized audio, up to 1080p.
The model's standout strength is generating natural mouth movement for the character in the source image, with lip-sync support across multiple languages. That makes it a strong pick for talking-head content, presenter characters, or promotional videos.
Inside Syntina, Happy Horse works as a standard image-to-video model in both the studio workflow and the video presets on the canvas. The default setting is 720p / 5 seconds; you can raise it to 1080p and up to 15 seconds if you need to.
Why Happy Horse 1.1?
Why is Happy Horse 1.1 #1?
Three things set the model apart on the leaderboard: leadership in real user preference, single-pass audio-video generation, and fast inference.
Text-to-Video (silent)
Image-to-Video (silent)
Text-to-Video (audio)
Image-to-Video (audio)
Source: Artificial Analysis Video Arena, April 2026. Scores are based on early vote counts and may change as more votes come in.
Ranking Performance
#1 in blind preference testing
Happy Horse 1.1 holds the highest Elo score on the Artificial Analysis Video Arena in both the Text-to-Video and Image-to-Video (silent) categories. The rankings are based on blind preference votes from real users who don't know which model produced the output.
Synchronized Audio + Video
Voiced video in a single pass
The model generates video and audio together in a single 40-layer self-attention Transformer, with no cross-attention module. That means the output arrives with audio synchronized to the video, with no separate post-production step needed.
Inference Speed
1080p video in under 40 seconds
According to figures shared by the team, a 1080p output is generated in about 38 seconds on a single NVIDIA H100 GPU; at 256p, a 5-second clip drops to roughly 2 seconds. That's a clear speed advantage over the current alternatives.
The Team Behind the Model
Who makes Happy Horse?
What is Alibaba Token Hub (ATH)?
Alibaba Token Hub is a senior unit that brings together Alibaba's AI expertise under one roof, from research labs to real-world software. Led by CEO Eddie Wu, it focuses on turning advanced models like Qwen into everyday tools, placing the "token" at the center as the core fuel of the modern AI economy.
Led by Zhang Di
Zhang Di is an AI engineer with 15+ years of field experience. He held a director role at Alibaba Group from 2010–2022, then became Vice President at Kuaishou, where he was the technical architect behind Kling AI. He returned to Alibaba in late 2025, founded the Taotian Future Life Lab under ATH, and shipped Happy Horse 1.1 within a few months.
Open-source status
While some sources have claimed Happy Horse 1.1 would be open-sourced, the model is closed-source. It won't be licensed or opened up — access is available only through the official API.
Sample Outputs
What can Happy Horse 1.1 produce?
Sample outputs from the Artificial Analysis Video Arena and community shares. Lip-sync, camera motion, and native audio all arrive together in these scenes.
Use Cases
Which jobs is it a good fit for?
- Talking-head videos and presentations
- Character or mascot animations
- User/presenter frames placed inside an architectural scene
- Short, voiced promo clips for social media
- Scenario tests that need lip-sync
Pricing
Credit table
Credit usage on Happy Horse is determined by duration × resolution. 720p uses fewer credits than 1080p, which is why Syntina opens with 720p as the default.
| Configuration | Credit cost |
|---|---|
| 5s @ 720p (default) | 140 credits |
| 5s @ 1080p | 180 credits |
| 10s @ 720p | 280 credits |
| 10s @ 1080p | 360 credits |
| 15s @ 720p | 420 credits |
| 15s @ 1080p | 540 credits |
The credit cost of a generation is shown clearly in the panel before you run it, so you can check before sending the scene.
Model Inputs
What parameters are available?
All inputs shown in the Syntina interface, with their default values.
| Parameter | Type | Default | Description |
|---|---|---|---|
| image_url | string (required) | — | Source image URL. Min 400px short side, max 10 MB. |
| prompt | string | — | Optional. The motion or speech text wanted in the scene (max 2,500 characters). |
| resolution | enum | 720p | Output resolution. |
| duration | number | 5 | Output duration (3 – 15 seconds). |
| aspect_ratio | enum | auto | Output ratio: 16:9, 9:16, 1:1, 4:3, 3:4, or auto. |
| enable_safety_checker | boolean | true | Whether the safety checker is active. |
| seed | number | — | Optional seed to make generation reproducible. |
FAQ
Frequently asked questions
- What is Happy Horse 1.1?
- Happy Horse 1.1 is an AI video generation model that appeared on the Artificial Analysis Video Arena on April 7, 2026, and instantly took the #1 spot in both the Text-to-Video and Image-to-Video (silent) categories. The ranking is based on blind preference tests, where real users compare outputs without knowing which model produced them.
- Who built Happy Horse 1.1?
- According to the model's technical notes, it comes from the Future Life Lab team under Alibaba's Taotian Group. The team is led by Zhang Di, former Vice President at Kuaishou and technical lead behind Kling AI.
- What are the technical specs?
- In their own materials, the team cites 15 billion parameters and a unified 40-layer self-attention Transformer architecture (no cross-attention) that generates video and audio in a single pass. Inference speed is given as roughly 38 seconds for a 1080p clip on a single NVIDIA H100 GPU.
- How do I use Happy Horse 1.1 in Syntina?
- Pick any scene on the canvas and turn on the Happy Horse model under the image-to-video preset. The default setting is 720p / 5 seconds; raise it to 1080p and up to 15 seconds if you need to.
- How many credits does Happy Horse 1.1 cost?
- Credit usage varies by duration × resolution. The default 5s / 720p generation costs 140 credits; switching to 1080p doubles the per-second credit cost. The credit cost of a generation is shown clearly in the panel before you run it, so you can check before sending the scene.
- Which languages does lip-sync support?
- The model currently supports native lip-sync in seven languages: Mandarin, Cantonese, English, Japanese, Korean, German, and French. Turkish isn't among the natively supported languages yet.
- Is it open-source?
- No. Happy Horse 1.1 is closed-source; it isn't licensed and you can't run the model on your own servers. Access is only possible through the official fal.ai API — Syntina integrates that API for you.
- Where are generated videos stored?
- All outputs are saved to the library in your Syntina account; you can download them, reopen them, or pull them back in as a reference on the canvas whenever you want.
Start creating with Happy Horse 1.1 now
Open the canvas — the default settings are already there, and the credit cost is shown clearly in the panel before you generate.