常用图像数据集大全（分类，跟踪，分割，检测等）

最新推荐文章于 2025-03-16 17:27:50 发布

曼陀罗彼岸花

最新推荐文章于 2025-03-16 17:27:50 发布

阅读量4.5w

点赞数 19

分类专栏：机器视觉

本文链接：https://blog.csdn.net/tiandijun/article/details/44539387

版权

这篇博客汇总了大量用于图像处理的公开数据集，包括搜狗实验室的280万张图片数据集、IMAGECLEF、NUS-WIDE等，涵盖了分类、检测、跟踪、分割等多个领域。此外，还提供了多个专注于特定任务如人脸识别、动作识别、医学图像等的数据集链接，是进行可复现研究的重要资源。

摘要生成于 C知道，由 DeepSeek-R1 满血版支持，前往体验 >

常用图像数据集大全（分类，跟踪，分割，检测等）

1.搜狗实验室数据集：

http://www.sogou.com/labs/dl/p.html

互联网图片库来自sogou图片搜索所索引的部分数据。其中收集了包括人物、动物、建筑、机械、风景、运动等类别，总数高达2,836,535张图片。对于每张图片，数据集中给出了图片的原图、缩略图、所在网页以及所在网页中的相关文本。200多G

http://www.imageclef.org/

IMAGECLEF致力于位图片相关领域提供一个基准（检索、分类、标注等等） Cross Language Evaluation Forum (CLEF) 。从2003年开始每年举行一次比赛.

http://staff.science.uva.nl/~xirong/index.php?n=Main.Dataset

Xiaorong Li 维护的数据集。PhD ,Intelligent Systems Lab Amsterdam.research on video and image retrieval.

Flickr-3.5M: A collection of 3.5 million social-tagged images.
Social20: A ground-truth set for tag-based social image retrieval.
Biconcepts2012test: A ground-truth set for retrieving bi-concepts (concept pairs) in unlabeled images.
neg4free: A set of negative examples automatically harvested from social-tagged images for 20 PASCAL VOC concepts.

wikipedia featured articles 函数图片（以及特征）以及对应的wiki文本。可以看看文章A New Approach to Cross-Modal Multimedia Retrieval，还有一批文章On the Role of Correlation and Abstraction in Cross-Modal Multimedia Retrieval不过还没有下载链接

http://www.svcl.ucsd.edu/projects/crossmodal/

http://lms.comp.nus.edu.sg/research/NUS-WIDE.htm

To our knowledge, this is the largest real-world web image dataset comprising over 269,000 images with over 5,000 user-provided tags, and ground-truth of 81 concepts for the entire dataset. The dataset is much larger than the popularly available Corel and Caltech 101 datasets. Though some datasets comprise over 3 million images, they only have ground-truth for a small fraction of images. Our proposed NUS-WIDE dataset has the ground-truth for the entire dataset.

http://www.cs.washington.edu/research/imagedatabase/

http://lear.inrialpes.fr/~jegou/data.php

Jegou的数据集，不过Jegou是专门做CBIR的，图像有ground truth，没有标注。

http://www.robots.ox.ac.uk/~vgg/data/oxbuildings/

vgg的osford building dataset。也是专门CBIR的数据。

http://acmmm13.org/submissions/call-for-multimedia-grand-challenge-solutions/msr-bing-grand-challenge-on-image-retrieval-scientific-track/

The dataset for the Microsoft Image Grand Challenge on Image Retrieval

另外介绍cvpaper上的整理的数据集

http://www.cvpapers.com/index.html

Participate in Reproducible Research

Detection

PASCAL VOC 2009 dataset

Classification/Detection Competitions, Segmentation Competition, Person Layout Taster Competition datasets

LabelMe dataset

LabelMe is a web-based image annotation tool that allows researchers to label images and share the annotations with the rest of the community. If you use the database, we only ask that you contribute to it, from time to time, by using the labeling tool.

BioID Face Detection Database

1521 images with human faces, recorded under natural conditions, i.e. varying illumination and complex background. The eye positions have been set manually.

CMU/VASC & PIE Face dataset

Yale Face dataset

Caltech

Cars, Motorcycles, Airplanes, Faces, Leaves, Backgrounds

Caltech 101

Pictures of objects belonging to 101 categories

Caltech 256

Pictures of objects belonging to 256 categories

Daimler Pedestrian Detection Benchmark

15,560 pedestrian and non-pedestrian samples (image cut-outs) and 6744 additional full images not containing pedestrians for bootstrapping. The test set contains more than 21,790 images with 56,492 pedestrian labels (fully visible or partially occluded), captured from a vehicle in urban traffic.

MIT Pedestrian dataset

CVC Pedestrian Datasets

CBCL Pedestrian Database

MIT Face dataset

CBCL Face Database

MIT Car dataset

CBCL Car Database

MIT Street dataset

CBCL Street Database

INRIA Person Data Set

A large set of marked up images of standing or walking people

INRIA car dataset

A set of car and non-car images taken in a parking lot nearby INRIA

INRIA horse dataset

A set of horse and non-horse images

H3D Dataset

3D skeletons and segmented regions for 1000 people in images

HRI RoadTraffic dataset

A large-scale vehicle detection dataset

BelgaLogos

10000 images of natural scenes, with 37 different logos, and 2695 logos instances, annotated with a bounding box.

FlickrBelgaLogos

10000 images of natural scenes grabbed on Flickr, with 2695 logos instances cut and pasted from the BelgaLogos dataset.

FlickrLogos-32

The dataset FlickrLogos-32 contains photos depicting logos and is meant for the evaluation of multi-class logo detection/recognition as well as logo retrieval methods on real-world images. It consists of 8240 images downloaded from Flickr.

TME Motorway Dataset

30000+ frames with vehicle rear annotation and classification (car and trucks) on motorway/highway sequences. Annotation semi-automatically generated using laser-scanner data. Distance estimation and consistent target ID over time available.

PHOS (Color Image Database for illumination invariant feature selection)

Phos is a color image database of 15 scenes captured under different illumination conditions. More particularly, every scene of the database contains 15 different images: 9 images captured under various strengths of uniform illumination, and 6 images under different degrees of non-uniform illumination. The images contain objects of different shape, color and texture and can be used for illumination invariant feature detection and selection.

CaliforniaND: An Annotated Dataset For Near-Duplicate Detection In Personal Photo Collections

California-ND contains 701 photos taken directly from a real user's personal photo collection, including many challenging non-identical near-duplicate cases, without the use of artificial image transformations. The dataset is annotated by 10 different subjects, including the photographer, regarding near duplicates.

Classification

PASCAL VOC 2009 dataset

Classification/Detection Competitions, Segmentation Competition, Person Layout Taster Competition datasets

Caltech

Cars, Motorcycles, Airplanes, Faces, Leaves, Backgrounds

Caltech 101

Pictures of objects belonging to 101 categories

Caltech 256

Pictures of objects belonging to 256 categories

ETHZ Shape Classes

A dataset for testing object class detection algorithms. It contains 255 test images and features five diverse shape-based classes (apple logos, bottles, giraffes, mugs, and swans).

Flower classification data sets

17 Flower Category Dataset

Animals with attributes

A dataset for Attribute Based Classification. It consists of 30475 images of 50 animals classes with six pre-extracted feature representations for each image.

Stanford Dogs Dataset

Dataset of 20,580 images of 120 dog breeds with bounding-box annotation, for fine-grained image categorization.

Recognition

Face and Gesture Recognition Working Group FGnet

Feret

Face and Gesture Recognition Working Group FGnet

PUT face

9971 images of 100 people

Labeled Faces in the Wild

A database of face photographs designed for studying the problem of unconstrained face recognition

Urban scene recognition

Traffic Lights Recognition, Lara's public benchmarks.

PubFig: Public Figures Face Database

The PubFig database is a large, real-world face dataset consisting of 58,797 images of 200 people collected from the internet. Unlike most other existing face datasets, these images are taken in completely uncontrolled situations with non-cooperative subjects.

YouTube Faces

The data set contains 3,425 videos of 1,595 different people. The shortest clip duration is 48 frames, the longest clip is 6,070 frames, and the average length of a video clip is 181.3 frames.

MSRC-12: Kinect gesture data set

The Microsoft Research Cambridge-12 Kinect gesture data set consists of sequences of human movements, represented as body-part locations, and the associated gesture to be recognized by the system.

QMUL underGround Re-IDentification (GRID) Dataset

This dataset contains 250 pedestrian image pairs + 775 additional images captured in a busy underground station for the research on person re-identification.

Person identification in TV series

Face tracks, features and shot boundaries from our latest CVPR 2013 paper. It is obtained from 6 episodes of Buffy the Vampire Slayer and 6 episodes of Big Bang Theory.

ChokePoint Dataset

ChokePoint is a video dataset designed for experiments in person identific