Kling AI, which makes a tool that turns written descriptions and pictures into video, announced Kling 4.0 on Sept. 29. The company says the new model can make a scene up to 30 seconds long in a single go, long enough for one continuous take or a short story with several moments in it. That is twice the 15 seconds of Kling's previous version, according to fal, one of the services that offers Kling's models.
The scenes come with sound. Kling says the model makes stereo sound, and that characters' lips now match their speech more closely. Characters can talk in languages including English, Spanish, Japanese, Korean and Chinese, in a range of accents, the company says. Kling also says the model can draw readable text and logos that stay steady as the camera moves.
Kling's main pitch is control over the film. Besides a written description, a user can give the model up to 10 still images that set what each key moment should look like, such as a character's pose or a change of scene, and the video moves through them in order. Users can also add their own pictures, clips and characters to a request, and Kling says a character can keep the same look, and even the same voice, from shot to shot. Kling says users can also change one part of an existing clip, such as an expression or the camera angle, while the rest of the scene holds.
Most of it is not open to everyone yet. Kling says Kling 4.0 is in early access, with a wider rollout in October. A faster, simpler version called Kling 4.0 Flash opened on Sept. 28 to people on Kling's Ultra yearly plan, but it makes clips of up to 20 seconds at lower picture quality and Kling's list of what it can do leaves out the planning images. Kling says sharper, richer-color video and a way to stretch a clip to two minutes are coming soon.
Every example so far was made and chosen by Kling itself. It is not yet clear how often an ordinary request turns out that well, what the full model will cost, or exactly when in October everyone can use it.
