Vision Foundation Model - Groups

DINOv2

The dataset used in the paper is DINOv2, a vision foundation model trained on a large-scale dataset.
- Dataset
- JSON
CLIP

The CLIP model and its variants are becoming the de facto backbone in many applications. However, training a CLIP model from hundreds of millions of image-text pairs can be...
- Dataset
- JSON

Before browse our site, please accept our cookies policy

2 datasets found