In the world of animation, creating lifelike human characters has always been a challenge. Recent advances in generative diffusion models have enabled compelling human image animation. However, existing techniques often struggle with shape alignment and motion guidance. To address these challenges, researchers from Nanjing University, Fudan University and Alibaba Group have introduced Champ. In this article, we will delve into its details
Table of Contents
What is Champ?
Champ is a novel framework for human image animation that leverages 3D parametric modelling to enhance shape representation and motion guidance. This approach utilizes the SMPL model, a popular 3D parametric human model. Moreover, SMPL represents the human form using shape and pose parameters in a low-dimensional space. By establishing correspondence between reference and driving motion through SMPL, Champ can accurately capture intricate geometry and motion characteristics while aligning 3D shapes.
Example Human Animations Generated by Champ
Key Methodology of CHAMP
Its methodology consists of three key components:
1. SMPL Motion Extraction
The SMPL model fitted to reference and driving video extracts pose parameter sequences. These are projected to generate multi-layer maps carrying 3D structural information.
2. Parametric Shape Alignment
SMPL models from reference and driving video are aligned in parameter space to transfer shapes accurately across identities during animation.
3. Multi-Layer Motion Guidance
Generated maps alongside skeleton guidance are encoded with self-attention and then fused as multi-level embeddings, conditioning the diffusion process for precise animation.
Combining CHAMP with T2I
CHAMP integrates techniques from text-to-image generation models like Stable Diffusion by conditioning on CLIP embeddings of the reference image. This allows animations to be controlled not just with video motion but also with natural language descriptions.
Training Details
The training dataset includes over 5,000 high-fidelity human videos sourced from popular sources with over 1 million frames. This dataset features large variations to train a robust model. The model was trained using 8 NVIDIA A100 GPUs in two phases – first processing individual frames and then adding a temporal layer for coherent video generation.
Performance Evaluation of CHAMP
1. Comparisons with Existing Approaches
Extensive experiments compare CHAMP against state-of-the-art methods like DisCo, MagicAnimate and AnimateAnyone. Significantly, it surpasses these diffusion-based techniques in metrics like PSNR, SSIM and LPIPS, demonstrating visually sharper and more consistent results.
2. Animation on TikTok Dataset
When tested on the standard TikTok dataset, it outperforms baselines in all evaluated metrics, from L1 loss to FID and FVD scores. Fine-tuning it on this data further improves performance.
3. Unseen Domain Animation
To assess generalization, CHAMP animates a new and more diverse “wild” dataset. It continues producing high-quality videos and exceeds competitors in terms of LPIPS and FVD, showing strong cross-domain abilities.
4. Cross-ID Animation
It can also generate animations where the reference image person differs from those in the driving videos. Aligning shapes parametrically maintains identity coherence better than pose-only approaches.
Conclusion
Champ by Alibaba researchers introduces a novel approach for generating lifelike human animation leveraging 3D parametric modeling. It addresses current limitations and produces controllable, consistent animations of human characters. Last but not least, CHAMP sheds new light on unlocking the full potential of human image animation. For code and installation details, you can visit the GitHub repository at fudan-generative-vision/champ. Additionally, to learn more technical details, please visit the project arXiV paper by Alibaba!
| Also Read Latest From Us
- Forget Towers: Verizon and AST SpaceMobile Are Launching Cellular Service From Space

- This $1,600 Graphics Card Can Now Run $30,000 AI Models, Thanks to Huawei

- The Global AI Safety Train Leaves the Station: Is the U.S. Already Too Late?

- The AI Breakthrough That Solves Sparse Data: Meet the Interpolating Neural Network

- The AI Advantage: Why Defenders Must Adopt Claude to Secure Digital Infrastructure







