[转] CV Datasets on the web
转自:CVPapers
This material is presented to ensure timely dissemination of scholarly and technical work. Copyright and all rights therein are retained by authors or by other copyright holders. All persons copying this information are expected to adhere to the terms and constraints invoked by each authors copyright.
Participate in Reproducible Research
Detection
- PASCAL VOC 2009 dataset
- Classification/Detection Competitions, Segmentation Competition, Person Layout Taster Competition datasets
- LabelMe dataset
- LabelMe is a web-based image annotation tool that allows researchers to label images and share the annotations with the rest of the community. If you use the database, we only ask that you contribute to it, from time to time, by using the labeling tool.
- BioID Face Detection Database
- 1521 images with human faces, recorded under natural conditions, i.e. varying illumination and complex background. The eye positions have been set manually.
- CMU/VASC & PIE Face dataset
- Yale Face dataset
- Caltech
- Cars, Motorcycles, Airplanes, Faces, Leaves, Backgrounds
- Caltech 101
- Pictures of objects belonging to 101 categories
- Caltech 256
- Pictures of objects belonging to 256 categories
- Daimler Pedestrian Detection Benchmark
- 15,560 pedestrian and non-pedestrian samples (image cut-outs) and 6744 additional full images not containing pedestrians for bootstrapping. The test set contains more than 21,790 images with 56,492 pedestrian labels (fully visible or partially occluded), captured from a vehicle in urban traffic.
- MIT Pedestrian dataset
- CVC Pedestrian Datasets
- CVC Pedestrian Datasets
- CBCL Pedestrian Database
- MIT Face dataset
- CBCL Face Database
- MIT Car dataset
- CBCL Car Database
- MIT Street dataset
- CBCL Street Database
- INRIA Person Data Set
- A large set of marked up images of standing or walking people
- INRIA car dataset
- A set of car and non-car images taken in a parking lot nearby INRIA
- INRIA horse dataset
- A set of horse and non-horse images
- H3D Dataset
- 3D skeletons and segmented regions for 1000 people in images
- HRI RoadTraffic dataset
- A large-scale vehicle detection dataset
- BelgaLogos
- 10000 images of natural scenes, with 37 different logos, and 2695 logos instances, annotated with a bounding box.
- FlickrBelgaLogos
- 10000 images of natural scenes grabbed on Flickr, with 2695 logos instances cut and pasted from the BelgaLogos dataset.
- FlickrLogos-32
- The dataset FlickrLogos-32 contains photos depicting logos and is meant for the evaluation of multi-class logo detection/recognition as well as logo retrieval methods on real-world images. It consists of 8240 images downloaded from Flickr.
- TME Motorway Dataset
- 30000+ frames with vehicle rear annotation and classification (car and trucks) on motorway/highway sequences. Annotation semi-automatically generated using laser-scanner data. Distance estimation and consistent target ID over time available.
- PHOS (Color Image Database for illumination invariant feature selection)
- Phos is a color image database of 15 scenes captured under different illumination conditions. More particularly, every scene of the database contains 15 different images: 9 images captured under various strengths of uniform illumination, and 6 images under different degrees of non-uniform illumination. The images contain objects of different shape, color and texture and can be used for illumination invariant feature detection and selection.
- CaliforniaND: An Annotated Dataset For Near-Duplicate Detection In Personal Photo Collections
- California-ND contains 701 photos taken directly from a real user's personal photo collection, including many challenging non-identical near-duplicate cases, without the use of artificial image transformations. The dataset is annotated by 10 different subjects, including the photographer, regarding near duplicates.
- USPTO Algorithm Challenge, Detecting Figures and Part Labels in Patents
- Contains drawing pages from US patents with manually labeled figure and part labels.
- Abnormal Objects Dataset
- Contains 6 object categories similar to object categories in Pascal VOC that are suitable for studying the abnormalities stemming from objects.
Classification
- PASCAL VOC 2009 dataset
- Classification/Detection Competitions, Segmentation Competition, Person Layout Taster Competition datasets
- Caltech
- Cars, Motorcycles, Airplanes, Faces, Leaves, Backgrounds
- Caltech 101
- Pictures of objects belonging to 101 categories
- Caltech 256
- Pictures of objects belonging to 256 categories
- ETHZ Shape Classes
- A dataset for testing object class detection algorithms. It contains 255 test images and features five diverse shape-based classes (apple logos, bottles, giraffes, mugs, and swans).
- Flower classification data sets
- 17 Flower Category Dataset
- Animals with attributes
- A dataset for Attribute Based Classification. It consists of 30475 images of 50 animals classes with six pre-extracted feature representations for each image.
- Stanford Dogs Dataset
- Dataset of 20,580 images of 120 dog breeds with bounding-box annotation, for fine-grained image categorization.
- Video classification USAA dataset
- The USAA dataset includes 8 different semantic class videos which are home videos of social occassions which feature activities of group of people. It contains around 100 videos for training and testing respectively. Each video is labeled by 69 attributes. The 69 attributes can be broken down into five broad classes: actions, objects, scenes, sounds, and camera movement.
Recognition
- Face and Gesture Recognition Working Group FGnet
- Face and Gesture Recognition Working Group FGnet
- Feret
- Face and Gesture Recognition Working Group FGnet
- PUT face
- 9971 images of 100 people
- Labeled Faces in the Wild
- A database of face photographs designed for studying the problem of unconstrained face recognition
- Urban scene recognition
- Traffic Lights Recognition, Lara's public benchmarks.
- PubFig: Public Figures Face Database
- The PubFig database is a large, real-world face dataset consisting of 58,797 images of 200 people collected from the internet. Unlike most other existing face datasets, these images are taken in completely uncontrolled situations with non-cooperative subjects.
- YouTube Faces
- The data set contains 3,425 videos of 1,595 different people. The shortest clip duration is 48 frames, the longest clip is 6,070 frames, and the average length of a video clip is 181.3 frames.
- MSRC-12: Kinect gesture data set
- The Microsoft Research Cambridge-12 Kinect gesture data set consists of sequences of human movements, represented as body-part locations, and the associated gesture to be recognized by the system.
- QMUL underGround Re-IDentification (GRID) Dataset
- This dataset contains 250 pedestrian image pairs + 775 additional images captured in a busy underground station for the research on person re-identification.
- Person identification in TV series
- Face tracks, features and shot boundaries from our latest CVPR 2013 paper. It is obtained from 6 episodes of Buffy the Vampire Slayer and 6 episodes of Big Bang Theory.
- ChokePoint Dataset
- ChokePoint is a video dataset designed for experiments in person identification/verification under real-world surveillance conditions. The dataset consists of 25 subjects (19 male and 6 female) in portal 1 and 29 subjects (23 male and 6 female) in portal 2.
- Hieroglyph Dataset
- Ancient Egyptian Hieroglyph Dataset.
Tracking
- BIWI Walking Pedestrians dataset
- Walking pedestrians in busy scenarios from a bird eye view
- "Central" Pedestrian Crossing Sequences
- Three pedestrian crossing sequences
- Pedestrian Mobile Scene Analysis
- The set was recorded in Zurich, using a pair of cameras mounted on a mobile platform. It contains 12'298 annotated pedestrians in roughly 2'000 frames.
- Head tracking
- BMP image sequences.
- KIT AIS Dataset
- Data sets for tracking vehicles and people in aerial image sequences.
- MIT Traffic Data Set
- MIT traffic data set is for research on activity analysis and crowded scenes. It includes a traffic video sequence of 90 minutes long. It is recorded by a stationary camera.
Segmentation
- Image Segmentation with A Bounding Box Prior dataset
- Ground truth database of 50 images with: Data, Segmentation, Labelling - Lasso, Labelling - Rectangle
- PASCAL VOC 2009 dataset
- Classification/Detection Competitions, Segmentation Competition, Person Layout Taster Competition datasets
- Motion Segmentation and OBJCUT data
- Cows for object segmentation, Five video sequences for motion segmentation
- Geometric Context Dataset
- Geometric Context Dataset: pixel labels for seven geometric classes for 300 images
- Crowd Segmentation Dataset
- This dataset contains videos of crowds and other high density moving objects. The videos are collected mainly from the BBC Motion Gallery and Getty Images website. The videos are shared only for the research purposes. Please consult the terms and conditions of use of these videos from the respective websites.
- CMU-Cornell iCoseg Dataset
- Contains hand-labelled pixel annotations for 38 groups of images, each group containing a common foreground. Approximately 17 images per group, 643 images total.
- Segmentation evaluation database
- 200 gray level images along with ground truth segmentations
- The Berkeley Segmentation Dataset and Benchmark
- Image segmentation and boundary detection. Grayscale and color segmentations for 300 images, the images are divided into a training set of 200 images, and a test set of 100 images.
- Weizmann horses
- 328 side-view color images of horses that were manually segmented. The images were randomly collected from the WWW.
- Saliency-based video segmentation with sequentially updated priors
- 10 videos as inputs, and segmented image sequences as ground-truth
Foreground/Background
- Wallflower Dataset
- For evaluating background modelling algorithms
- Foreground/Background Microsoft Cambridge Dataset
- Foreground/Background segmentation and Stereo dataset from Microsoft Cambridge
- Stuttgart Artificial Background Subtraction Dataset
- The SABS (Stuttgart Artificial Background Subtraction) dataset is an artificial dataset for pixel-wise evaluation of background models.
Saliency Detection (source)
- AIM
- 120 Images / 20 Observers (Neil D. B. Bruce and John K. Tsotsos 2005).
- LeMeur
- 27 Images / 40 Observers (O. Le Meur, P. Le Callet, D. Barba and D. Thoreau 2006).
- Kootstra
- 100 Images / 31 Observers (Kootstra, G., Nederveen, A. and de Boer, B. 2008).
- DOVES
- 101 Images / 29 Observers (van der Linde, I., Rajashekar, U., Bovik, A.C., Cormack, L.K. 2009).
- Ehinger
- 912 Images / 14 Observers (Krista A. Ehinger, Barbara Hidalgo-Sotelo, Antonio Torralba and Aude Oliva 2009).
- NUSEF
- 758 Images / 75 Observers (R. Subramanian, H. Katti, N. Sebe1, M. Kankanhalli and T-S. Chua 2010).
- JianLi
- 235 Images / 19 Observers (Jian Li, Martin D. Levine, Xiangjing An and Hangen He 2011).
- Extended Complex Scene Saliency Dataset (ECSSD)
- ECSSD contains 1000 natural images with complex foreground or background. For each image, the ground truth mask of salient object(s) is provided.
Video Surveillance
- CAVIAR
- For the CAVIAR project a number of video clips were recorded acting out the different scenarios of interest. These include people walking alone, meeting with others, window shopping, entering and exitting shops, fighting and passing out and last, but not least, leaving a package in a public place.
- ViSOR
- ViSOR contains a large set of multimedia data and the corresponding annotations.
Multiview
- 3D Photography Dataset
- Multiview stereo data sets: a set of images
- Multi-view Visual Geometry group's data set
- Dinosaur, Model House, Corridor, Aerial views, Valbonne Church, Raglan Castle, Kapel sequence
- Oxford reconstruction data set (building reconstruction)
- Oxford colleges
- Multi-View Stereo dataset (Vision Middlebury)
- Temple, Dino
- Multi-View Stereo for Community Photo Collections
- Venus de Milo, Duomo in Pisa, Notre Dame de Paris
- IS-3D Data
- Dataset provided by Center for Machine Perception
- CVLab dataset
- CVLab dense multi-view stereo image database
- 3D Objects on Turntable
- Objects viewed from 144 calibrated viewpoints under 3 different lighting conditions
- Object Recognition in Probabilistic 3D Scenes
- Images from 19 sites collected from a helicopter flying around Providence, RI. USA. The imagery contains approximately a full circle around each site.
- Multiple cameras fall dataset
- 24 scenarios recorded with 8 IP video cameras. The first 22 first scenarios contain a fall and confounding events, the last 2 ones contain only confounding events.
- CMP Extreme View Dataset
- 15 wide baseline stereo image pairs with large viewpoint change, provided ground truth homographies.
Action
- UCF Sports Action Dataset
- This dataset consists of a set of actions collected from various sports which are typically featured on broadcast television channels such as the BBC and ESPN. The video sequences were obtained from a wide range of stock footage websites including BBC Motion gallery, and GettyImages.
- UCF Aerial Action Dataset
- This dataset features video sequences that were obtained using a R/C-controlled blimp equipped with an HD camera mounted on a gimbal.The collection represents a diverse pool of actions featured at different heights and aerial viewpoints. Multiple instances of each action were recorded at different flying altitudes which ranged from 400-450 feet and were performed by different actors.
- UCF YouTube Action Dataset
- It contains 11 action categories collected from YouTube.
- Weizmann action recognition
- Walk, Run, Jump, Gallop sideways, Bend, One-hand wave, Two-hands wave, Jump in place, Jumping Jack, Skip.
- UCF50
- UCF50 is an action recognition dataset with 50 action categories, consisting of realistic videos taken from YouTube.
- ASLAN
- The Action Similarity Labeling (ASLAN) Challenge.
- MSR Action Recognition Datasets
- The dataset was captured by a Kinect device. There are 12 dynamic American Sign Language (ASL) gestures, and 10 people. Each person performs each gesture 2-3 times.
- KTH Recognition of human actions
- Contains six types of human actions (walking, jogging, running, boxing, hand waving and hand clapping) performed several times by 25 subjects in four different scenarios: outdoors, outdoors with scale variation, outdoors with different clothes and indoors.
- Hollywood-2 Human Actions and Scenes dataset
- Hollywood-2 datset contains 12 classes of human actions and 10 classes of scenes distributed over 3669 video clips and approximately 20.1 hours of video in total.
- Collective Activity Dataset
- This dataset contains 5 different collective activities : crossing, walking, waiting, talking, and queueing and 44 short video sequences some of which were recorded by consumer hand-held digital camera with varying view point.
- Olympic Sports Dataset
- The Olympic Sports Dataset contains YouTube videos of athletes practicing different sports.
- SDHA 2010
- Surveillance-type videos
- VIRAT Video Dataset
- The dataset is designed to be realistic, natural and challenging for video surveillance domains in terms of its resolution, background clutter, diversity in scenes, and human activity/event categories than existing action recognition datasets.
- HMDB: A Large Video Database for Human Motion Recognition
- Collected from various sources, mostly from movies, and a small proportion from public databases, YouTube and Google videos. The dataset contains 6849 clips divided into 51 action categories, each containing a minimum of 101 clips.
- Stanford 40 Actions Dataset
- Dataset of 9,532 images of humans performing 40 different actions, annotated with bounding-boxes.
- 50Salads dataset
- Fully annotated dataset of RGB-D video data and data from accelerometers attached to kitchen objects capturing 25 people preparing two mixed salads each (4.5h of annotated data). Annotated activities correspond to steps in the recipe and include phase (pre-/ core-/ post) and the ingredient acted upon.
- Penn Sports Action
- The dataset contains 2326 video sequences of 15 different sport actions and human body joint annotations for all sequences.
Human pose/Expression
- AFEW (Acted Facial Expressions In The Wild)/SFEW (Static Facial Expressions In The Wild)
- Dynamic temporal facial expressions data corpus consisting of close to real world environment extracted from movies.
- ETHZ CALVIN Dataset
Image stitching
- IPM Vision Group Image Stitching datasets
- Images and parameters for registeration
Medical
- VIP Laparoscopic / Endoscopic Dataset
- Collection of endoscopic and laparoscopic (mono/stereo) videos and images
Misc
- Zurich Buildings Database
- ZuBuD Image Database contains over 1005 images about Zurich city building.
- Color Name Data Sets
- Mall dataset
- The mall dataset was collected from a publicly accessible webcam for crowd counting and activity profiling research.
- QMUL Junction Dataset
- A busy traffic dataset for research on activity analysis and behaviour understanding.
[转] CV Datasets on the web的更多相关文章
- CV code references
转:http://www.sigvc.org/bbs/thread-72-1-1.html 一.特征提取Feature Extraction: SIFT [1] [Demo program][SI ...
- CV codes代码分类整理合集 《转》
from:http://www.sigvc.org/bbs/thread-72-1-1.html 一.特征提取Feature Extraction: SIFT [1] [Demo program] ...
- paper 28 :一些常见常用数据库的下载网站集锦
做图像处理+模式识别的童鞋怎么可以没有数据库呢? 但是,如果自己做一个数据库,费时费力费钱先不说,关键是建立的数据库的公信力一般不会高,做出的算法也别人也不好比较,所以呢,下载比较权威的公共数据库还是 ...
- paper 15 :整理的CV代码合集
这篇blog,原来是西弗吉利亚大学的Li xin整理的,CV代码相当的全,不知道要经过多长时间的积累才会有这么丰富的资源,在此谢谢LI Xin .我现在分享给大家,希望可以共同进步!还有,我需要说一下 ...
- 图像库---Image Datasets---OpenSift源代码---openSurf源代码
1.Computer Vision Datasets on the web http://www.cvpapers.com/datasets.html 2.Dataset Reference http ...
- tornado之模板扩展
当我们有多个模板的时候,很多模板之间其实相似度很高.我们期望可以重用部分网页代码.这在tornado中可以通过extends语句来实现.为了扩展一个已经存在的模板,你只需要在新的模板文件的顶部放上一句 ...
- java web学习总结(五) -------------------servlet开发(一)
一.Servlet简介 Servlet是sun公司提供的一门用于开发动态web资源的技术. Sun公司在其API中提供了一个servlet接口,用户若想用发一个动态web资源(即开发一个Java程序向 ...
- java web学习总结(三十一) -------------------EL表达式
一.EL表达式简介 EL 全名为Expression Language.EL主要作用: 1.获取数据 EL表达式主要用于替换JSP页面中的脚本表达式,以从各种类型的web域 中检索java对象.获取数 ...
- java web学习总结(三十) -------------------JSTL表达式
一.JSTL标签库介绍 JSTL标签库的使用是为弥补html标签的不足,规范自定义标签的使用而诞生的.使用JSLT标签的目的就是不希望在jsp页面中出现java逻辑代码 二.JSTL标签库的分类 核心 ...
随机推荐
- self和super的区别
(1)self调用自己方法,super调用父类方法 (2)self是类,super是预编译指令 (3)[self class]和[super class]输出是一样的 ①当使用 self 调用方法时, ...
- 【转】Go语言入门教程(一)Linux下安装Go
说明 系统是Ubuntu. 关于安装 下载安装包 当前官方下载地址是https://golang.org/dl/,如果不能访问,请自行FQ,FQ是技术工作者的必备技能. 安装 tar -xzvf go ...
- Python中使用SQLite
参考原文 廖雪峰Python教程 使用SQLite SQLite是一种嵌入式数据库,它的数据库就是一个文件.由于SQLite本身是用C写的,而且体积很小,所以经常被集成到各种应用程序中,甚至在IOS和 ...
- 如何在MONO 3D寻找最短路路径
前段时间有个客户说他们想在我们的3D的机房中找从A点到B点的最短路径,然而在2D中确实有很多成熟的寻路算法,其中A*是最为常见的,而这个Demo也是用的A*算法,以下计算的是从左上角到右下角的最短路径 ...
- Extjs杂记录
1,页面跳转到另外一个页面 这段话的意思:取得恢复密码窗口,关闭这个窗口,页面跳转到Login页面 2,keypecial 当与导航相关的键(如箭头.tab键.Enter键.ESC键等)按下时,该事件 ...
- UVA - 11212 Editing a Book (IDA*搜索)
题目: 给出n(1<n<10)个数字组成的序列,每次操作可以选取一段连续的区间将这个区间之中的数字放到其他任意位置.问最少经过多少次操作可将序列变为1,2,3……n. 思路: 利用IDA* ...
- python_ 学习笔记(hello world)
python中的循环语句 循环语句均可以尾随一个else语句块,该块再条件为false后执行一次 如果使用break跳出则不执行. for it in [1,2,3,4]: print(it,end= ...
- python黑科技库:FuckIt.py,让你代码从此远离bug
今天给你推荐的这个库叫 “FuckIt.py”,名字一看就是很黄很暴力的那种,作者是这样介绍它的: FuckIt.py uses state-of-the-art technology to make ...
- C. Day at the Beach
codeforces 599c C. Day at the Beach One day Squidward, Spongebob and Patrick decided to go to the be ...
- 【codeforces 3C】Tic-tac-toe
[链接] 我是链接,点我呀:) [题意] 题意 [题解] 写一个函数判断当前局面是否有人赢. 然后枚举上一个人的棋子下在哪个地方. 然后把他撤回 看看撤回前是不是没人赢然后没撤回之前是不是有人赢了. ...