MSCOCO
Human Pose Estimation (HPE) aims to estimate the position of each joint point of the human body in a given image. HPE tasks support a wide range of downstream tasks such as activity recognition, motion capture, etc. Recently with the ViT model being proven effective on many visual tasks, many transformer-based methods have achieved excellent performance on HPE tasks.
BibTex: