AI Models
Image to VideoProvider: AlibabaNew

Happy Horse 1.1

An Alibaba model that turns a single image into a video with synchronized audio and multilingual lip-sync.

Happy Horse 1.1 is part of Alibaba's image-to-video model family and can turn a single frame into a clip with synchronized audio, up to 1080p.

The model's standout strength is generating natural mouth movement for the character in the source image, with lip-sync support across multiple languages. That makes it a strong pick for talking-head content, presenter characters, or promotional videos.

Inside Syntina, Happy Horse works as a standard image-to-video model in both the studio workflow and the video presets on the canvas. The default setting is 720p / 5 seconds; you can raise it to 1080p and up to 15 seconds if you need to.

Why Happy Horse 1.1?

Why is Happy Horse 1.1 #1?

Three things set the model apart on the leaderboard: leadership in real user preference, single-pass audio-video generation, and fast inference.

Text-to-Video (silent)

#1Elo 1333

Image-to-Video (silent)

#1Elo 1392

Text-to-Video (audio)

#2Elo 1205

Image-to-Video (audio)

#2Elo 1161

Source: Artificial Analysis Video Arena, April 2026. Scores are based on early vote counts and may change as more votes come in.

Ranking Performance

#1 in blind preference testing

Happy Horse 1.1 holds the highest Elo score on the Artificial Analysis Video Arena in both the Text-to-Video and Image-to-Video (silent) categories. The rankings are based on blind preference votes from real users who don't know which model produced the output.

Synchronized Audio + Video

Voiced video in a single pass

The model generates video and audio together in a single 40-layer self-attention Transformer, with no cross-attention module. That means the output arrives with audio synchronized to the video, with no separate post-production step needed.

Inference Speed

1080p video in under 40 seconds

According to figures shared by the team, a 1080p output is generated in about 38 seconds on a single NVIDIA H100 GPU; at 256p, a 5-second clip drops to roughly 2 seconds. That's a clear speed advantage over the current alternatives.

The Team Behind the Model

Who makes Happy Horse?

What is Alibaba Token Hub (ATH)?

Alibaba Token Hub is a senior unit that brings together Alibaba's AI expertise under one roof, from research labs to real-world software. Led by CEO Eddie Wu, it focuses on turning advanced models like Qwen into everyday tools, placing the "token" at the center as the core fuel of the modern AI economy.

Led by Zhang Di

Zhang Di is an AI engineer with 15+ years of field experience. He held a director role at Alibaba Group from 2010–2022, then became Vice President at Kuaishou, where he was the technical architect behind Kling AI. He returned to Alibaba in late 2025, founded the Taotian Future Life Lab under ATH, and shipped Happy Horse 1.1 within a few months.

Open-source status

While some sources have claimed Happy Horse 1.1 would be open-sourced, the model is closed-source. It won't be licensed or opened up — access is available only through the official API.

Sample Outputs

What can Happy Horse 1.1 produce?

Sample outputs from the Artificial Analysis Video Arena and community shares. Lip-sync, camera motion, and native audio all arrive together in these scenes.

Use Cases

Which jobs is it a good fit for?

  • Talking-head videos and presentations
  • Character or mascot animations
  • User/presenter frames placed inside an architectural scene
  • Short, voiced promo clips for social media
  • Scenario tests that need lip-sync

Pricing

Credit table

Credit usage on Happy Horse is determined by duration × resolution. 720p uses fewer credits than 1080p, which is why Syntina opens with 720p as the default.

ConfigurationCredit cost
5s @ 720p (default)140 credits
5s @ 1080p180 credits
10s @ 720p280 credits
10s @ 1080p360 credits
15s @ 720p420 credits
15s @ 1080p540 credits

The credit cost of a generation is shown clearly in the panel before you run it, so you can check before sending the scene.

Model Inputs

What parameters are available?

All inputs shown in the Syntina interface, with their default values.

ParameterTypeDefaultDescription
image_urlstring (required)Source image URL. Min 400px short side, max 10 MB.
promptstringOptional. The motion or speech text wanted in the scene (max 2,500 characters).
resolutionenum720pOutput resolution.
durationnumber5Output duration (3 – 15 seconds).
aspect_ratioenumautoOutput ratio: 16:9, 9:16, 1:1, 4:3, 3:4, or auto.
enable_safety_checkerbooleantrueWhether the safety checker is active.
seednumberOptional seed to make generation reproducible.

FAQ

Frequently asked questions

What is Happy Horse 1.1?
Happy Horse 1.1 is an AI video generation model that appeared on the Artificial Analysis Video Arena on April 7, 2026, and instantly took the #1 spot in both the Text-to-Video and Image-to-Video (silent) categories. The ranking is based on blind preference tests, where real users compare outputs without knowing which model produced them.
Who built Happy Horse 1.1?
According to the model's technical notes, it comes from the Future Life Lab team under Alibaba's Taotian Group. The team is led by Zhang Di, former Vice President at Kuaishou and technical lead behind Kling AI.
What are the technical specs?
In their own materials, the team cites 15 billion parameters and a unified 40-layer self-attention Transformer architecture (no cross-attention) that generates video and audio in a single pass. Inference speed is given as roughly 38 seconds for a 1080p clip on a single NVIDIA H100 GPU.
How do I use Happy Horse 1.1 in Syntina?
Pick any scene on the canvas and turn on the Happy Horse model under the image-to-video preset. The default setting is 720p / 5 seconds; raise it to 1080p and up to 15 seconds if you need to.
How many credits does Happy Horse 1.1 cost?
Credit usage varies by duration × resolution. The default 5s / 720p generation costs 140 credits; switching to 1080p doubles the per-second credit cost. The credit cost of a generation is shown clearly in the panel before you run it, so you can check before sending the scene.
Which languages does lip-sync support?
The model currently supports native lip-sync in seven languages: Mandarin, Cantonese, English, Japanese, Korean, German, and French. Turkish isn't among the natively supported languages yet.
Is it open-source?
No. Happy Horse 1.1 is closed-source; it isn't licensed and you can't run the model on your own servers. Access is only possible through the official fal.ai API — Syntina integrates that API for you.
Where are generated videos stored?
All outputs are saved to the library in your Syntina account; you can download them, reopen them, or pull them back in as a reference on the canvas whenever you want.

Start creating with Happy Horse 1.1 now

Open the canvas — the default settings are already there, and the credit cost is shown clearly in the panel before you generate.