Yup, people have been using local video models like Wan2.2 to generate stills, finding that for some things like human anatomy, it can outperform image generation models. Very cool how moving training data helps build spatial understanding that is applicable even to still images.