Back to Discover

Vidu says its new AI video model can keep characters' looks and voices consistent across a scene

The early version takes up to 15 pictures and three voice clips as a guide, and it is open now on Vidu's website

A woman in a red outfit stands on a grassy cliff looking out at a city of golden towers and bridges floating above the clouds, with airships overhead, a scene made with Vidu Q4

Vidu, the AI video tool from ShengShu Technology, has opened an early version of its next model, Q4 Preview. It makes short clips with sound from pictures and a written description, and a user can give it up to 15 pictures and three voice recordings, which Vidu says help characters keep the same look and voice across a scene. It is open now on Vidu's website, and in a simpler form in Perfect Corp.'s YouCam apps, but the examples shown so far all come from Vidu and Perfect Corp., and Vidu has not said when the finished model will arrive.

Vidu, the AI video tool made by ShengShu Technology, released Q4 Preview on Oct. 7, the first public preview of its next top model, Q4. It turns pictures and a written description into short videos with sound. Vidu says the new version improves how characters act, how the camera follows the action, and how effects such as explosions, smoke and fireworks fit into a scene.

The model's main pitch is control over who appears on screen. Instead of describing a character only in words, a user can give it up to 15 pictures of the people, clothes, props and places they want, plus up to three voice recordings. Vidu says it copies each character's voice from those recordings, with lips moving in time, which it says keeps how characters look and sound consistent across a scene. A simpler mode turns a single photo into a moving clip and keeps the photo's shape, whether tall, wide or square.

Clips run up to 16 seconds and can hold several shots with moving cameras. Vidu says the model is meant for independent creators and small studios, who often need more than one try to get a shot right, and that it makes each extra try cheaper. "Advanced video models should prove themselves beyond carefully selected demos. Every improvement in efficiency should ultimately give creators the freedom to try one more shot or make one more version," said Yihang Luo, ShengShu's co-founder and chief executive.

It is open now on Vidu's website, and developers can build it into their own apps through Vidu's service for app makers. Vidu says its launch prices are a promotion, and that final prices and features may vary by plan and region. Perfect Corp., which makes the YouCam photo and video apps, has also added the model to its apps and its web editor, where people turn one photo into a short video with sound, paid for with in-app coins.

Vidu calls this a preview: it is releasing it ahead of the full Q4 model so creators can use it in real projects and help shape the final version. The examples shown so far come from Vidu and Perfect Corp., so it is not yet clear how often an ordinary request turns out that well, what the finished model will cost, or when it will arrive.

Sources

About this video

  • Available now: “Vidu Q4 Preview is now available through Vidu's web product and API platform.” Source
  • The official sources don't say whether the video is sped up or whether the footage is real, simulated or AI-generated.