Google Unveils Genie 2, an AI Tool Capable of Creating Immersive 3D Worlds from a Single Image
Google's Genie 2 AI model is a groundbreaking foundation world model that can produce a vast array of playable, action-controllable 3D environments from a single prompt image. It has the capability to create various perspectives, including first-person views, isometric views, and third-person driving videos, as well as intricate 3D visual scenes featuring interactive objects. The model also supports rapid prototyping of physics effects like smoke, gravity, lighting, and reflections, which can be controlled by humans or AI agents using a keyboard and mouse. According to a recent report, this technology enables artists and designers to quickly prototype ideas, thereby accelerating the creative process for environment design and research. The report highlights that Genie 2's generalization capabilities allow concept art and drawings to be transformed into fully interactive environments, facilitating swift prototyping and bootstrapping the creative process. Although this research is still in its early stages, with significant room for improvement in agent and environment generation capabilities, Genie 2 is believed to be a crucial step towards addressing the challenges of training embodied agents safely while progressing towards Artificial General Intelligence (AGI). The detailed report, complete with examples, is available on Google's Deepmind website. In related news, Future, a UK-based specialist media publisher, has recently partnered with OpenAI to integrate its ChatGPT tool across various business functions, including sales, marketing, and editorial operations.