Hands play a crucial role in how we perceive and interact with the world. The accurate representation of human hands in AI-generated images is a significant challenge. It can seriously detract from the believability and cohesiveness of an AI-generated scene when hands are present. AI image generators like Stable Diffusion often fail to model hands correctly. Supreeth Narasimhaswamy, a researcher at Adobe Applied Research (GenAI), decided to address this challenge with a two-stage diffusion approach, HanDiffuser. This AI model leverages topological and geometric human priors to generate realistic-looking hands.
Table of Contents
Introducing HanDiffuser: Text-to-Image Generation With Realistic Hand Appearances
HanDiffuser introduces a novel two-stage diffusion that incorporates detailed human hand anatomical knowledge into the image generation process. As Narasimhaswamy notes, “Users get tired of seeing weird hands in generated images.” With HanDiffuser, his goal was to introduce anatomy-guided diffusion specific to the hands during diffusion to make generative models more practical, robust and trustworthy for users worldwide.
The Development of HanDiffuser
Supreeth developed this two-step process during an internship at Adobe Applied Research. He collaborated closely with many other researchers, including his PhD advisor Minh Hoai, on this impactful research.
Supreeth’s paper “HanDiffuser: Text-to-Image Generation With Realistic Hand Appearances” was recently accepted to the prestigious IEEE Conference on Computer Vision and Pattern Recognition (CVPR) in 2024. While the paper has yet to be published, Supreeth shared some initial information that provides insight into this groundbreaking new research.
(Accepted to #CVPR2024) Are you tired of looking at weird human hands in AI-generated images? We introduce a two-stage diffusion process that leverages topological and geometric human priors to generate realistic-looking hands. #GenAI, #AI, #machinelearning, #computervision pic.twitter.com/JZA3tL9NID
— Supreeth Narasimhaswamy (@nsupreeth5) February 27, 2024
How HanDiffuser Works
This involves two main steps: Text-to-Hand Params (T2H) Diffusion and Text-Guided Hand Params-to-image (T-H21) Diffusion.
1. T2H Diffusion
First, it uses text prompts and human hand parameters like pose and topology to generate an initial set of hand features through a “T2H Diffusion” model. This captures key aspects of hand shape and positioning based on the requested prompt.
2. T-H21 Diffusion
Next, those generated hand parameters are fed into a second diffusion model along with the original text prompt in a “T-H21 Diffusion” process. This allows the hand details from the first step to guide highly realistic rendering of hands within the full image context specified by the prompts.
By leveraging human hand anatomical knowledge at the parameter level before image generation, HanDiffuser is able to produce vastly more plausible and accurately depicted hands compared to standard diffusion models.
HanDiffuser vs Stable Diffusion
Stable Diffusion still struggles with generating plausible hands, often creating disoriented hands with distorted or missing fingers. Creator Supreeth Narasimhaswamy offered an early glimpse at how it compares to Stable Diffusion with some generated images from both models using the same text prompt.
The results showed that HanDiffuser excelled at generating anatomically accurate hands aligned seamlessly with the scene context described. In contrast, Stable Diffusion continued struggling with the prompt, frequently warping or distorting hands and misaligning fingers.
What Can We Expect From HanDiffuser?
This two-stage anatomical approach can become a major step forward for generative hand modelling. The ability to reliably generate naturalistic hands aligned with scene context and poses described in prompts represents a significant improvement over legacy diffusion baselines. Using human anatomy, HanDiffuser shows that diffusion models can faithfully capture the complex articulations and nuanced details of real hands. This can potentially achieve a new standard of photorealism for hands-in AI imagery across domains like fashion, design, and more.
Conclusion
Supreeth latest research on generating realistic hand appearances through HanDiffuser can advance text-to-image generation. Supreeth and his collaborators should be commended for their novel approach. Their technique represents an exciting advancement that will likely inspire further research into incorporating more human structure and knowledge into generative models. The computer vision and AI community eagerly await further results from HanDiffuser in the forthcoming CVPR paper. Once again, congratulations to Supreeth Narasimhaswamy on this groundbreaking research!
- Forget Towers: Verizon and AST SpaceMobile Are Launching Cellular Service From Space

- This $1,600 Graphics Card Can Now Run $30,000 AI Models, Thanks to Huawei

- The Global AI Safety Train Leaves the Station: Is the U.S. Already Too Late?

- The AI Breakthrough That Solves Sparse Data: Meet the Interpolating Neural Network

- The AI Advantage: Why Defenders Must Adopt Claude to Secure Digital Infrastructure







