Showing posts with label machine learning algorithms. Show all posts
Showing posts with label machine learning algorithms. Show all posts

Tuesday, August 25, 2015

Monday, September 29, 2014

Corner Detector

Detecting corners is often a good first step in computer vision.  If you can match corners from two images you are well on your way to figuring out how they fit together for example.

Corner detection is an approach used within computer vision systems to extract certain kinds of features and infer the contents of an image. Corner detection is frequently used in motion detection, image registration, video tracking, image mosaicing, panorama stitching, 3D modelling and object recognition. Corner detection overlaps with the topic of interest point detection.

http://en.wikipedia.org/wiki/Corner_detection


Lecture slide decks:


http://www.cse.psu.edu/~rcollins/CSE486/lecture06.pdf

http://courses.cs.washington.edu/courses/cse577/05sp/notes/harris.pdf

Tutorials:

Python/OpenCV

http://opencv-python-tutroals.readthedocs.org/en/latest/py_tutorials/py_feature2d/py_features_harris/py_features_harris.html

YouTube video:

https://www.youtube.com/watch?v=vkWdzWeRfC4

Matlab code:

http://www.mathworks.com/matlabcentral/fileexchange/9272-harris-corner-detector


Deep Learning

Deep Learning is a new area of Machine Learning research, which has been introduced with the objective of moving Machine Learning closer to one of its original goals: Artificial Intelligence.

Website of organization dedicated to all things deep learning:

http://deeplearning.net/

Multiple tutorials and a wealth of other info:

http://deeplearning.net/tutorial/intro.html


Peer review paper:

Theoretical results suggest that in order to learn the kind of complicated functions that can represent high- level abstractions (e.g. in vision, language, and other AI-level tasks), one may need deep architectures. Deep architectures are composed of multiple levels of non-linear operations, such as in neural nets with many hidden layers or in complicated propositional formulae re-using many sub-formulae. Searching the parameter space of deep architectures is a difficult task, but learning algorithms such as those for Deep Belief Networks have recently been proposed to tackle this problem with notable success, beating the state-of-the-art in certain areas. This paper discusses the motivations and principles regarding learning algorithms for deep architectures, in particular those exploiting as building blocks unsupervised learning of single-layer models such as Restricted Boltzmann Machines, used to construct deeper models such as Deep Belief Networks.

http://www.iro.umontreal.ca/~bengioy/papers/ftml.pdf


Deep learning implementations in many languages:

http://deeplearning.net/software_links/

Restricted Boltzmann machine

Learning to use RBM's is on my todo list...I'll update when I get around to it.  RBM's are just one technique for deep learning.

The Restricted Boltzmann Machine (RBM) has become increasingly popular of late after its success in the Netflix prize competition and other competitions. Most of the inventive work behind RBMs was done by Geoffrey Hinton. In particular the training of RBMs using an algorithm called "Contrastive Divergence" (CD). CD is very similar to gradient descent. A good consequence of the CD is its ability to "dream". Of the various machine learning methods out there, the RBM is the only one which has this capacity baked in implicitly.

http://bayesianthink.blogspot.com/2013/05/the-restricted-boltzmann-machine-rbm.html#.VCnWzikijjI

This is some Matlab code a guy made of a class he was taking.  It is probably not great but if you are working in Matlab it is probably better than starting from scratch:

https://code.google.com/p/matrbm/

RBM tutorial:

http://deeplearning.net/tutorial/rbm.html#rbm


RBM in scikit-learn:

http://scikit-learn.org/stable/modules/neural_networks.html



A Practical Guide to Training Restricted Boltzmann Machines:


http://www.cs.toronto.edu/~hinton/absps/guideTR.pdf

Multi-scale Oriented Patches (MOPS)

MOPS is a useful tool for many computer vision and related applications. They key is that it gives you a way to generate features that are (mostly) invariant to rotation and scale to use as input for machine learning/artificial intelligence algorithms.

http://www.cs.bath.ac.uk/brown/mops/mops.html


Technical report from Microsoft Research:

http://research.microsoft.com/pubs/70120/tr-2004-133.pdf

Nice slide deck on MOPS:

http://www.csie.ntu.edu.tw/~cyy/courses/vfx/08spring/lectures/handouts/lec06_feature2_4up.pdf


Example app from Microsoft Research:

http://research.microsoft.com/en-us/um/redmond/groups/ivm/PhotoTours/

AdaBoost

Note: AdaBoost is extremely sensitive to mislabeled samples in your data. For example, if you are trying to classify transactions as either "fraud" or "not fraud" if you have even one mislabeled, then the classifier will over learn that one bad sample and be useless.  There are other versions of boosting algorithms that try to overcome this but if you have data for which you can not be sure of the labels then consider using some other method.


AdaBoost, short for "Adaptive Boosting", is a machine learning meta-algorithm formulated by Yoav Freund and Robert Schapire who won the prestigious "Gödel Prize" in 2003 for their work. It can be used in conjunction with many other types of learning algorithms to improve their performance. The output of the other learning algorithms ('weak learners') is combined into a weighted sum that represents the final output of the boosted classifier. AdaBoost is adaptive in the sense that subsequent weak learners are tweaked in favor of those instances misclassified by previous classifiers. AdaBoost is sensitive to noisy data and outliers. In some problems, however, it can be less susceptible to the overfitting problem than other learning algorithms. The individual learners can be weak, but as long as the performance of each one is slightly better than random guessing (i.e., their error rate is smaller than 0.5 for binary classification), the final model can be proven to converge to a strong learner.

While every learning algorithm will tend to suit some problem types better than others, and will typically have many different parameters and configurations to be adjusted before achieving optimal performance on a dataset, AdaBoost (with decision trees as the weak learners) is often referred to as the best out-of-the-box classifier. When used with decision tree learning, information gathered at each stage of the AdaBoost algorithm about the relative 'hardness' of each training sample is fed into the tree growing algorithm such that later trees tend to focus on harder to classify examples.

http://en.wikipedia.org/wiki/AdaBoost

Very nice AdaBoost slide deck:

http://cmp.felk.cvut.cz/~sochmj1/adaboost_talk.pdf

Matlab and C++ implementations:

http://graphics.cs.msu.ru/en/science/research/machinelearning/adaboosttoolbox

Viola Jones object detection framework

The Viola–Jones object detection framework is the first object detection framework to provide competitive object detection rates in real-time proposed in 2001 by Paul Viola and Michael Jones. Although it can be trained to detect a variety of object classes, it was motivated primarily by the problem of face detection. This algorithm is implemented in OpenCV as cvHaarDetectObjects().

YouTube video explaining Viola Jones face detection:

https://www.youtube.com/watch?v=WfdYYNamHZ8

This is a slide deck explaining Viola Jones face detection:

http://www.slideshare.net/wolf/avihu-efrats-viola-and-jones-face-detection-slides/


Haar features are not an AI or ML algorithm themselves, but instead are often a useful tool for transforming data into a format that an AI or ML algorithm can use.

The Wikipedia article only talks about them with respect to object recognition in visible light images.  However, they can be used with images from any spectrum or even any type of data that can be represented as X,Y,Z such as a digital elevation map of terrain.

http://en.wikipedia.org/wiki/Haar-like_features

This pdf has a good explanation of how to use Haar features:

http://nichol.as/papers/Wilson/Facial%20feature%20detection%20using%20Haar.pdf


Friday, September 19, 2014

Genetic Algorithms: Cool Name & Damn Simple

Nice GA tutorial.

Genetic algorithms are a mysterious sounding technique in mysterious sounding field--artificial intelligence. This is the problem with naming things appropriately. When the field was labeled artificial intelligence, it meant using mathematics to artificially create the semblance of intelligence, but self-engrandizing researchers and Isaac Asimov redefined it as robots.

The name genetic algorithms does sound complex and has a faintly magical ring to it, but it turns out that they are one of the simplest and most-intuitive concepts you'll encounter in A.I.




Genetic Algorithms: Cool Name & Damn Simple - Irrational Exuberance

Curve fitting with Pyevolve

This is a very nice tutorial for genetic algorithms.  It uses pyevolve but the tutorial part is useful even if you are using a different language/implementation for GA.

A Coder's Musings: Curve fitting with Pyevolve

Genetic Algorithms tutorial

Great tutorial and introduction to genetic algorithms.  There are java applets that you can play with to see how GA's work.

These pages introduce some fundamentals of genetic algorithms. Pages are intended to be used for learning about genetic algorithms without any previous knowledge from this area. Only some knowledge of computer programming is assumed. You can find here several interactive Java applets demonstrating work of genetic algorithms.

As the area of genetic algorithms is very wide, it is not possible to cover everything in these pages. But you should get some idea, what the genetic algorithms are and what they could be useful for. Do not expect any sophisticated mathematics theories here.

Main page - Introduction to Genetic Algorithms - Tutorial with Interactive Java Applets

Pyevolve genetic algorithm python software

I have used this software to successfully create a genetic algorithm python script that I use to tune parameters on extra tree classifiers and RDF classifiers.  It is pretty easy to use and you can make almost any type of GA with it.  It is open source so you can go in a tinker with it.

Welcome to Pyevolve documentation ! — Pyevolve v0.5 documentation

There is a great pyevolve tutorial here:

A Coder's Musings: Curve fitting with Pyevolve

Neural Networks for Machine Learning

This is an online course from the University of Toronto.  If you can spend about 8 hours a week for 8 weeks you should be thoroughly familiar with ANN's.

Learn about artificial neural networks and how they're being used for machine learning, as applied to speech and object recognition, image segmentation, modeling language and human motion, etc. We'll emphasize both the basic algorithms and the practical tricks needed to get them to work well.

Neural Networks for Machine Learning | Coursera

Basic Neural Network Tutorial : C++ Implementation and Source Code

This tutorial is in two parts, one is the theory of ANN and the other part is a C++ implementation with hints on how to modify it efficiently.  If you are new to neural networks and you are a C++ programmer this is a great place to start.

Basic Neural Network Tutorial – Theory | Taking Initiative

Basic Neural Network Tutorial : C++ Implementation and Source Code | Taking Initiative

Wednesday, September 17, 2014

Theoretical Machine Learning

Introductory lecture notes from a class taught by Professor Rob Schapire who is a leader in the field.  This is a good place to start for a beginner in machine learning.

www.cs.princeton.edu/courses/archive/spr08/cos511/scribe_notes/0204.pdf

Decision Forests for Classification, Regression, Density Estimation, Manifold Learning and Semi-Supervised Learning

This technical report from Microsoft Research is an A-Z tutorial on how decision tree machine learning algorithms work.  It includes in depth explanations of random forests, extra tree classifiers, random ferns and other variations for both classification and regression.

It is in report format and compares decision forests to other types of machine learning algorithms such as SVM.  Some simple toy problems give the basics and some real life applications such as body position recognition and medical image are included.

There is also an accompanying PowerPoint with some nice animations.

http://research.microsoft.com/pubs/155552/decisionForests_MSR_TR_2011_114.pdf is not available

Alternatives to support vector machines in neuroimaging ensembles of decision trees for classification and information mapping with predictive models

This is a nice tutorial for using random decision forests for classifying medical images.  There is a comparison with some other methods, especially SVM's.

http://web.stanford.edu/~richiard/slides/PRNI2013Tutorial_export.pdf is not available

One-Class Support Vector Machines: Methods and Applications

This tutorial on SVM's is a decent one for a beginner.  It is a slide deck from a presentation.

isites.harvard.edu/fs/docs/icb.topic274302.files/Dan.Nick.pdf

A Comparison of Methods for Multiclass Support Vector Machines

Support vector machines (SVMs) were originally designed for binary classification. How to effectively extend it for multiclass classification is still an ongoing research issue. Several methods have been proposed where typically we construct a multiclass classifier by combining several binary classifiers. Some authors also proposed methods that consider all classes at once. As it is computationally more expensive to solve multiclass problems, comparisons of these methods using large-scale problems have not been seriously conducted. Especially for methods solving multiclass SVM in one step, a much larger optimization problem is required so up to now experiments are limited to small data sets. In this paper we give decomposition implementations for two such “all-together” methods. We then compare their performance with three methods based on binary classifications: “one-against-all,” “one-against-one,” and directed acyclic graph SVM (DAGSVM). Our experiments indicate that the “one-against-one” and DAG methods are more suitable for practical use than the other methods. Results also show that for large problems methods by considering all data at once in general need fewer support vectors.


cs.ecs.baylor.edu/~hamerly/courses/5325_11s/papers/svm/hsu2001multiclass.pdf

Nonlinear regression in environmental sciences by support vector machines combined with evolutionary strategy

A hybrid algorithm combining support vector regression with evolutionary strategy (SVR-ES) is proposed for predictive models in the environmental sciences. SVR-ES uses uncorrelated mutation with p step sizes to find the optimal SVR hyper-parameters. Three environmental forecast datasets used in the WCCI-2006 contest – surface air temperature, precipitation and sulphur dioxide concentration – were tested. We used multiple linear regression (MLR) as benchmark and a variety of machine learning techniques including bootstrap-aggregated ensemble artificial neural network (ANN), SVR-ES, SVR with hyper-parameters given by the Cherkassky–Ma estimate, the M5 regression tree, and random forest (RF). We also tested all techniques using stepwise linear regression (SLR) first to screen out irrelevant predictors. We concluded that SVR-ES is an attractive approach because it tends to outperform the other techniques and can also be implemented in an almost automatic way. The Cherkassky–Ma estimate is a useful approach for minimizing the mean absolute error and saving computational time related to the hyper-parameter search. The ANN and RF are also good options to outperform multiple linear regression (MLR). Finally, the use of SLR for predictor selection can dramatically reduce computational time and often help to enhance accuracy.

Nonlinear regression in environmental sciences by support vector machines combined with evolutionary strategy