Vidu, the AI video tool made by ShengShu Technology, released Q4 Preview on Oct. 7, the first public preview of its next top model, Q4. It turns pictures and a written description into short videos with sound. Vidu says the new version improves how characters act, how the camera follows the action, and how effects such as explosions, smoke and fireworks fit into a scene.
The model's main pitch is control over who appears on screen. Instead of describing a character only in words, a user can give it up to 15 pictures of the people, clothes, props and places they want, plus up to three voice recordings. Vidu says it copies each character's voice from those recordings, with lips moving in time, which it says keeps how characters look and sound consistent across a scene. A simpler mode turns a single photo into a moving clip and keeps the photo's shape, whether tall, wide or square.
Clips run up to 16 seconds and can hold several shots with moving cameras. Vidu says the model is meant for independent creators and small studios, who often need more than one try to get a shot right, and that it makes each extra try cheaper. "Advanced video models should prove themselves beyond carefully selected demos. Every improvement in efficiency should ultimately give creators the freedom to try one more shot or make one more version," said Yihang Luo, ShengShu's co-founder and chief executive.
It is open now on Vidu's website, and developers can build it into their own apps through Vidu's service for app makers. Vidu says its launch prices are a promotion, and that final prices and features may vary by plan and region. Perfect Corp., which makes the YouCam photo and video apps, has also added the model to its apps and its web editor, where people turn one photo into a short video with sound, paid for with in-app coins.
Vidu calls this a preview: it is releasing it ahead of the full Q4 model so creators can use it in real projects and help shape the final version. The examples shown so far come from Vidu and Perfect Corp., so it is not yet clear how often an ordinary request turns out that well, what the finished model will cost, or when it will arrive.
