state-of-the-art implementations related to visual recognition and search

http://rogerioferis.com/VisualRecognitionAndSearch2014/Resources.html

Source Code

Non-exhaustive list of state-of-the-art implementations related to visual recognition and search. There is no warranty for the source code links below – use them at your own risk!

Feature Detection and Description

General Libraries:

VLFeat – Implementation of various feature descriptors (including SIFT, HOG, and LBP) and covariant feature detectors (including DoG, Hessian, Harris
Laplace, Hessian Laplace, Multiscale Hessian, Multiscale Harris). Easy-to-use Matlab interface. See
a=v&pid=sites&srcid=ZGVmYXVsdGRvbWFpbnxlY2N2MTJmZWF0dXJlc3xneDo3ZDllMzVhMDA4YzEzNmU2" style="color:rgb(165,88,88)">Modern
features: Software – Slides providing a demonstration of VLFeat and also links to other software. Check also VLFeat hands-on
session training
OpenCV – Various implementations of modern feature detectors and descriptors (SIFT, SURF, FAST, BRIEF, ORB, FREAK, etc.)

Fast Keypoint Detectors for Real-time Applications:

FAST – High-speed corner detector implementation for a wide variety of platforms
AGAST – Even faster than the FAST corner detector. A multi-scale version of this method is used for the BRISK descriptor (ECCV
2010).

Binary Descriptors for Real-Time Applications:

BRIEF – C++ code for a fast and accurate interest point descriptor (not invariant to rotations and scale) (ECCV 2010)
ORB – OpenCV implementation of the Oriented-Brief (ORB) descriptor (invariant to rotations,
but not scale)
BRISK – Efficient Binary descriptor invariant to rotations and scale. It includes a Matlab mex interface. (ICCV 2011)
FREAK – Faster than BRISK (invariant to rotations and scale) (CVPR 2012)

SIFT and SURF Implementations:

SIFT: VLFeat, OpenCV, Original
code by David Lowe, GPU implementation, OpenSIFT
SURF: Herbert Bay’s code, OpenCV, GPU-SURF

Other Local Feature Detectors and Descriptors:

VGG Affine Covariant features – Oxford code for various affine covariant feature detectors and descriptors.
LIOP descriptor – Source code for the Local Intensity order Pattern (LIOP) descriptor (ICCV 2011).
Local Symmetry Features – Source code for matching of local symmetry features under large variations in lighting, age, and
rendering style (CVPR 2012).

Global Image Descriptors:

GIST – Matlab code for the GIST descriptor
CENTRIST – Global visual descriptor for scene categorization and object detection (PAMI 2011)

Feature Coding and Pooling

VGG Feature Encoding Toolkit – Source code for various state-of-the-art feature encoding methods – including
Standard hard encoding, Kernel codebook encoding, Locality-constrained linear encoding, and Fisher kernel encoding.
Spatial Pyramid Matching – Source code for feature pooling based on spatial pyramid matching (widely used for image classification)

Convolutional Nets and Deep Learning

Caffe – Fast C++ implementation of deep convolutional networks (GPU / CPU / ImageNet 2013 demonstration).
id=software:overfeat:start" style="color:rgb(165,88,88)">OverFeat – C++ library for integrated classification and localization of objects.
EBLearn – C++ Library for Energy-Based Learning. It includes several demos and step-by-step instructions to train classifiers based on
convolutional neural networks.
Torch7 – Provides a matlab-like environment for state-of-the-art machine learning algorithms, including a fast implementation of convolutional neural
networks.
Deep Learning - Various links for deep learning software.

Facial Feature Detection and Tracking

IntraFace – Very accurate detection and tracking of facial features (C++/Matlab API).

Part-Based Models

Deformable Part-based Detector – Library provided by the authors of the original paper (state-of-the-art in PASCAL VOC detection
task)
Efficient Deformable Part-Based Detector – Branch-and-Bound implementation for a deformable part-based detector.
Accelerated Deformable Part Model – Efficient implementation of a method that achieves the exact same performance of deformable
part-based detectors but with significant acceleration (ECCV 2012).
Coarse-to-Fine Deformable Part Model – Fast approach for deformable object detection (CVPR 2011).
Poselets – C++ and Matlab versions for object detection based on poselets.
Part-based Face Detector and Pose Estimation – Implementation of a unified approach for face detection, pose estimation, and landmark
localization (CVPR 2012).

Attributes and Semantic Features

Relative Attributes – Modified implementation of RankSVM to train Relative Attributes (ICCV 2011).
Object Bank – Implementation of object bank semantic features (NIPS 2010). See also ActionBank
Classemes, Picodes, and Meta-class features – Software for extracting high-level image descriptors
(ECCV 2010, NIPS 2011, CVPR 2012).

Large-Scale Learning

Additive Kernels – Source code for fast additive kernel SVM classifiers (PAMI 2013).
LIBLINEAR – Library for large-scale linear SVM classification.
VLFeat – Implementation for Pegasos SVM and Homogeneous Kernel map.

Fast Indexing and Image Retrieval

FLANN – Library for performing fast approximate nearest neighbor.
Kernelized LSH – Source code for Kernelized Locality-Sensitive Hashing (ICCV 2009).
ITQ Binary codes – Code for generation of small binary codes using Iterative Quantization and other baselines such as Locality-Sensitive-Hashing
(CVPR 2011).
INRIA Image Retrieval – Efficient code for state-of-the-art large-scale image retrieval (CVPR 2011).

Object Detection

See Part-based Models and Convolutional
Nets above.
Pedestrian Detection at 100fps – Very fast and accurate pedestrian detector (CVPR 2012).
Caltech Pedestrian Detection Benchmark – Excellent resource for pedestrian detection, with various links
for state-of-the-art implementations.
OpenCV – Enhanced implementation of Viola&Jones real-time object
detector, with trained models for face detection.
Efficient Subwindow Search – Source code for branch-and-bound optimization for efficient object localization (CVPR
2008).

3D Recognition

Point-Cloud Library – Library for 3D image and point cloud processing.

Action Recognition

ActionBank – Source code for action recognition based on the ActionBank representation (CVPR 2012).
STIP Features – software for computing space-time interest point descriptors
Independent Subspace Analysis – Look for Stacked ISA for Videos (CVPR 2011)
Velocity Histories of Tracked Keypoints - C++ code for activity recognition using the velocity histories of tracked keypoints
(ICCV 2009)

Datasets

Attributes

Animals with Attributes – 30,475 images of 50 animals classes with 6 pre-extracted feature representations for each image.
aYahoo and aPascal – Attribute annotations for images collected from Yahoo and Pascal VOC 2008.
FaceTracer – 15,000 faces annotated with 10 attributes and fiducial points.
PubFig – 58,797 face images of 200 people with 73 attribute classifier outputs.
LFW – 13,233 face images of 5,749 people with 73 attribute classifier outputs.
Human Attributes – 8,000 people with annotated attributes. Check also this link for
another dataset of human attributes.
SUN Attribute Database – Large-scale scene attribute database with a taxonomy of 102 attributes.
ImageNet Attributes – Variety of attribute labels for the ImageNet dataset.
Relative attributes – Data for OSR and a subset of PubFig datasets. Check also this link for
the WhittleSearch data.
Attribute Discovery Dataset – Images of shopping categories associated with textual descriptions.

Fine-grained Visual Categorization

Caltech-UCSD Birds Dataset – Hundreds of bird categories with annotated parts and attributes.
Stanford Dogs Dataset – 20,000 images of 120 breeds of dogs from around the world.
Oxford-IIIT Pet Dataset – 37 category pet dataset with roughly 200 images for each class. Pixel level trimap segmentation is
included.
Leeds Butterfly Dataset – 832 images of 10 species of butterflies.
Oxford Flower Dataset – Hundreds of flower categories.

Face Detection

FDDB – UMass face detection dataset and benchmark (5,000+ faces)
CMU/MIT – Classical face detection dataset.

Face Recognition

Face Recognition Homepage – Large collection of face recognition datasets.
LFW – UMass unconstrained face recognition dataset (13,000+ face images).
NIST Face Homepage – includes face recognition grand challenge (FRGC), vendor tests (FRVT) and others.
CMU Multi-PIE – contains more than 750,000 images of 337 people, with 15 different views and 19 lighting conditions.
FERET – Classical face recognition dataset.
Deng Cai’s face dataset in Matlab Format – Easy to use if you want play with simple face datasets including Yale,
ORL, PIE, and Extended Yale B.
SCFace – Low-resolution face dataset captured from surveillance cameras.

Handwritten Digits

MNIST – large dataset containing a training set of 60,000 examples, and a test set of 10,000 examples.

Pedestrian Detection

Caltech Pedestrian Detection Benchmark – 10 hours of video taken from a vehicle,350K bounding boxes for
about 2.3K unique pedestrians.
INRIA Person Dataset – Currently one of the most popular pedestrian detection datasets.
ETH Pedestrian Dataset – Urban dataset captured from a stereo rig mounted on a stroller.
TUD-Brussels Pedestrian Dataset – Dataset with image pairs recorded in an crowded urban setting with an onboard camera.
PASCAL Human Detection – One of 20 categories in PASCAL VOC detection challenges.
USC Pedestrian Dataset – Small dataset captured from surveillance cameras.

Generic Object Recognition

ImageNet – Currently the largest visual recognition dataset in terms of number of categories and images.
Tiny Images – 80 million 32x32 low resolution images.
Pascal VOC – One of the most influential visual recognition datasets.
Caltech 101 / Caltech
256 – Popular image datasets containing 101 and 256 object categories, respectively.
MIT LabelMe – Online annotation tool for building computer vision databases.

Scene Recognition

MIT SUN Dataset – MIT scene understanding dataset.
UIUC Fifteen Scene Categories – Dataset of 15 natural scene categories.

Feature Detection and Description

VGG Affine Dataset – Widely used dataset for measuring performance of feature detection and description. CheckVLBenchmarksfor
an evaluation framework.

Action Recognition

Benchmarking Activity Recognition – CVPR 2012 tutorial covering various datasets
for action recognition.

RGBD Recognition

RGB-D Object Dataset – Dataset containing 300 common household objects

state-of-the-art implementations related to visual recognition and search的更多相关文章

Image Processing and Analysis_8_Edge Detection：Edge and line oriented contour detection State of the art ——2011
此主要讨论图像处理与分析.虽然计算机视觉部分的有些内容比如特征提取等也可以归结到图像分析中来,但鉴于它们与计算机视觉的紧密联系,以及它们的出处,没有把它们纳入到图像处理与分析中来.同样,这里面也有 ...
Convolutional Neural Networks for Visual Recognition
http://cs231n.github.io/ 里面有很多相当好的文章 http://cs231n.github.io/convolutional-networks/ Table of Cont ...
大规模视觉识别挑战赛ILSVRC2015各团队结果和方法 Large Scale Visual Recognition Challenge 2015
Large Scale Visual Recognition Challenge 2015 (ILSVRC2015) Legend: Yellow background = winner in thi ...
论文笔记之： Bilinear CNN Models for Fine-grained Visual Recognition
Bilinear CNN Models for Fine-grained Visual Recognition CVPR 2015 本文提出了一种双线性模型( bilinear models),一种识 ...
CNN for Visual Recognition (01)
CS231n: Convolutional Neural Networks for Visual Recognitionhttp://vision.stanford.edu/teaching/cs23 ...
【论文阅读】Deep Mixture of Diverse Experts for Large-Scale Visual Recognition
导读: 本文为论文<Deep Mixture of Diverse Experts for Large-Scale Visual Recognition>的阅读总结.目的是做大规模图像分类 ...
目标检测--Spatial pyramid pooling in deep convolutional networks for visual recognition(PAMI, 2015)
Spatial pyramid pooling in deep convolutional networks for visual recognition 作者: Kaiming He, Xiangy ...
A Theoretical Analysis of Feature Pooling in Visual Recognition
这篇是10年ICML的论文,但是它是从原理上来分析池化的原因,因为池化的好坏的确会影响到结果,比如有除了最大池化和均值池化,还有随机池化等等,在eccv14中海油在顶层加个空间金字塔池化的方法.可谓多 ...
Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition
Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition Kaiming He, Xiangyu Zh ...

随机推荐

JAVA 读取图片储存至本地
需求:serlvet经过处理通过报表工具返回一张报表图(柱状图折线图). 现在需要把这个图存储到本地以便随时查看 // 构造URL URL url = new URL(endStr); // 打开 ...
Windows下文件或文件夹不能删除时的解决办法
windows在删除文件或文件夹时,提示文件或文件夹被占用而无法删除解决办法:win7: winxp:需要借助第三方工具Unlocker.360.Process Explorer(这个是微软支持的) ...
PV FV PMT
Android 性能优化五性能分析工具dumpsys的使用
Android提供的dumpsys工具能够用于查看感兴趣的系统服务信息与状态,手机连接电脑后能够直接命令行运行adb shell dumpsys 查看全部支持的Service可是这样输出的太多,能够通 ...
android com.handmark.pulltorefresh 使用技巧
近期使用android com.handmark.pulltorefresh 遇到一些小问题.如今总结一些: 集体使用教程见: http://blog.csdn.net/harvic880925/ar ...
从零开始学Xamarin.Forms(一) 概述
原文:从零开始学Xamarin.Forms(一) 概述 Xamarin 读 "ˈzæmərin",是一个基于开源项目mono的能够使用C#开发的收费的跨平台(iOS.And ...
网络协议——IP
IPv4地址不论什么网络设备能够经过一个网络接口卡(NIC)接入网,假定该设备要能够访问的其它设备,然后该卡必须有一个唯一的地址.候接入多个网络,相应地该设备就有多个地址.假设这个设备是主机的话.一 ...
关于mysql主从复制的概述与分类（转）
一.概述: 按照MySQL的同步复制特点,大体上可以分为三种类别: 1.异步复制: 2.半同步复制: 3.完全同步的复制: -------------------------------------- ...
sql server事物控制
一.多个数据库 1.存储过程 2.Commit写在 Try...Catch后面 protected void Button1_Click(object sender, EventArgs e) ...
Blend4精选案例图解教程（二）：找张图片玩特效
原文:Blend4精选案例图解教程(二):找张图片玩特效 Blend中的特效给了我们在处理资源时更多的想象空间,合理地运用特效往往会得到梦幻般效果,本次教程展示对图片应用特效的常规操作,当然特效不仅限 ...

state-of-the-art implementations related to visual recognition and search

state-of-the-art implementations related to visual recognition and search的更多相关文章

随机推荐

热门专题