Shortly after launching an AI program that can turn doodles into fine art, Nvidia is back with another: an AI model that converts a single 2D image of a person into a ‘talking head‘ on video.
Imagine how much more relaxing video conferencing would be if you could just upload a single image of your face, with no one the wiser that you’re still lounging around in your pyjamas.
Known as the ‘Vid2Vid Cameo’, this new deep-learning program helps hide the chaos behind a camera. For example, you could roll out of bed disheveled, but upload an image of yourself well-dressed, and the AI will map your facial movements via the camera to the reference image, so you appear in your well-groomed state to others in the chat.
Surprisingly, the AI can also adjust to show you looking directly at the screen, even if your eyes are distracted by the latest soap playing in the background. While this model sounds like the perfect way to trick your co-workers into thinking you’ve got your life together, it’s unclear just how much of a mess the program can cover up.
According toTNW, the program is powered by generative adversarial networks (GANs), which produces the video shown to your co-workers by pitting two neural networks against each other. One being a generator that creates realistic-looking samples, and another that attempts to figure out if they’re real or fake.
In turn, these two networks are able to synthesize realistic-looking videos from just one image of the user. During the video call itself, the camera will then capture the user’s real-time motion and expressions, applying it to the uploaded image.
Apart from its cool AI-powered features, the model is also said to reduce bandwidth significantly. Nvidia claims that the program can cut the bandwidth needed for video conferences by up to 10 times.
Vid2Vid Cameo will soon be available in Nvidia Maxine SDK and Nvidia Video Code SDK, but users can first try out a demo here.