Google is giving the AI image generation game a little blitz with its latest experiment, Whisk—a tool that’s more about mixing pixels than penning prompts. Unlike many existing tools that rely on text-based descriptions, this one primarily lets users upload images to dictate the subject, scene, and style of a newly crafted picture.
When users drag and drop images into the tool—powered by Google’s Gemini AI and its newest image generation model, Imagen 3—Whisk generates captions to analyze key features. These descriptors become the building blocks for Imagen 3 to create a unique, blended image that captures the essence of the input without replicating it outright.
The tool is designed with brainstorming and experimentation in mind, rather than picture-perfect precision. Users can tweak results by editing underlying text descriptions or swapping out input images, toying around in a playground for rapid visual ideation.
Currently, Whisk is available as part of Google Labs in the US. As with most experimental tools, it’s a work in progress, and the company is looking to gain feedback from early adopters.