There are many useful articles at this site. It is useful for everyone from novice to advanced:
http://fastml.com/
Showing posts with label machine learning algorithms. Show all posts
Showing posts with label machine learning algorithms. Show all posts
Tuesday, August 25, 2015
Thursday, December 4, 2014
11 open source tools for making the most of machine learning
A summary of open source machine learning libraries.
http://www.networkworld.com/article/2855100/opensource-subnet/11-open-source-tools-machine-learning.html
http://www.networkworld.com/article/2855100/opensource-subnet/11-open-source-tools-machine-learning.html
Monday, September 29, 2014
Corner Detector
Detecting corners is often a good first step in computer vision. If you can match corners from two images you are well on your way to figuring out how they fit together for example.
Corner detection is an approach used within computer vision systems to extract certain kinds of features and infer the contents of an image. Corner detection is frequently used in motion detection, image registration, video tracking, image mosaicing, panorama stitching, 3D modelling and object recognition. Corner detection overlaps with the topic of interest point detection.
http://en.wikipedia.org/wiki/Corner_detection
Lecture slide decks:
http://www.cse.psu.edu/~rcollins/CSE486/lecture06.pdf
http://courses.cs.washington.edu/courses/cse577/05sp/notes/harris.pdf
Tutorials:
Python/OpenCV
http://opencv-python-tutroals.readthedocs.org/en/latest/py_tutorials/py_feature2d/py_features_harris/py_features_harris.html
YouTube video:
https://www.youtube.com/watch?v=vkWdzWeRfC4
Matlab code:
http://www.mathworks.com/matlabcentral/fileexchange/9272-harris-corner-detector
Corner detection is an approach used within computer vision systems to extract certain kinds of features and infer the contents of an image. Corner detection is frequently used in motion detection, image registration, video tracking, image mosaicing, panorama stitching, 3D modelling and object recognition. Corner detection overlaps with the topic of interest point detection.
http://en.wikipedia.org/wiki/Corner_detection
Lecture slide decks:
http://www.cse.psu.edu/~rcollins/CSE486/lecture06.pdf
http://courses.cs.washington.edu/courses/cse577/05sp/notes/harris.pdf
Tutorials:
Python/OpenCV
http://opencv-python-tutroals.readthedocs.org/en/latest/py_tutorials/py_feature2d/py_features_harris/py_features_harris.html
YouTube video:
https://www.youtube.com/watch?v=vkWdzWeRfC4
Matlab code:
http://www.mathworks.com/matlabcentral/fileexchange/9272-harris-corner-detector
Deep Learning
Deep Learning is a new area of Machine Learning research, which has been introduced with the objective of moving Machine Learning closer to one of its original goals: Artificial Intelligence.
Website of organization dedicated to all things deep learning:
http://deeplearning.net/
Multiple tutorials and a wealth of other info:
http://deeplearning.net/tutorial/intro.html
Peer review paper:
Theoretical results suggest that in order to learn the kind of complicated functions that can represent high- level abstractions (e.g. in vision, language, and other AI-level tasks), one may need deep architectures. Deep architectures are composed of multiple levels of non-linear operations, such as in neural nets with many hidden layers or in complicated propositional formulae re-using many sub-formulae. Searching the parameter space of deep architectures is a difficult task, but learning algorithms such as those for Deep Belief Networks have recently been proposed to tackle this problem with notable success, beating the state-of-the-art in certain areas. This paper discusses the motivations and principles regarding learning algorithms for deep architectures, in particular those exploiting as building blocks unsupervised learning of single-layer models such as Restricted Boltzmann Machines, used to construct deeper models such as Deep Belief Networks.
http://www.iro.umontreal.ca/~bengioy/papers/ftml.pdf
Deep learning implementations in many languages:
http://deeplearning.net/software_links/
Website of organization dedicated to all things deep learning:
http://deeplearning.net/
Multiple tutorials and a wealth of other info:
http://deeplearning.net/tutorial/intro.html
Peer review paper:
Theoretical results suggest that in order to learn the kind of complicated functions that can represent high- level abstractions (e.g. in vision, language, and other AI-level tasks), one may need deep architectures. Deep architectures are composed of multiple levels of non-linear operations, such as in neural nets with many hidden layers or in complicated propositional formulae re-using many sub-formulae. Searching the parameter space of deep architectures is a difficult task, but learning algorithms such as those for Deep Belief Networks have recently been proposed to tackle this problem with notable success, beating the state-of-the-art in certain areas. This paper discusses the motivations and principles regarding learning algorithms for deep architectures, in particular those exploiting as building blocks unsupervised learning of single-layer models such as Restricted Boltzmann Machines, used to construct deeper models such as Deep Belief Networks.
http://www.iro.umontreal.ca/~bengioy/papers/ftml.pdf
Deep learning implementations in many languages:
http://deeplearning.net/software_links/
Restricted Boltzmann machine
Learning to use RBM's is on my todo list...I'll update when I get around to it. RBM's are just one technique for deep learning.
The Restricted Boltzmann Machine (RBM) has become increasingly popular of late after its success in the Netflix prize competition and other competitions. Most of the inventive work behind RBMs was done by Geoffrey Hinton. In particular the training of RBMs using an algorithm called "Contrastive Divergence" (CD). CD is very similar to gradient descent. A good consequence of the CD is its ability to "dream". Of the various machine learning methods out there, the RBM is the only one which has this capacity baked in implicitly.
http://bayesianthink.blogspot.com/2013/05/the-restricted-boltzmann-machine-rbm.html#.VCnWzikijjI
This is some Matlab code a guy made of a class he was taking. It is probably not great but if you are working in Matlab it is probably better than starting from scratch:
https://code.google.com/p/matrbm/
RBM tutorial:
http://deeplearning.net/tutorial/rbm.html#rbm
RBM in scikit-learn:
http://scikit-learn.org/stable/modules/neural_networks.html
A Practical Guide to Training Restricted Boltzmann Machines:
http://www.cs.toronto.edu/~hinton/absps/guideTR.pdf
The Restricted Boltzmann Machine (RBM) has become increasingly popular of late after its success in the Netflix prize competition and other competitions. Most of the inventive work behind RBMs was done by Geoffrey Hinton. In particular the training of RBMs using an algorithm called "Contrastive Divergence" (CD). CD is very similar to gradient descent. A good consequence of the CD is its ability to "dream". Of the various machine learning methods out there, the RBM is the only one which has this capacity baked in implicitly.
http://bayesianthink.blogspot.com/2013/05/the-restricted-boltzmann-machine-rbm.html#.VCnWzikijjI
This is some Matlab code a guy made of a class he was taking. It is probably not great but if you are working in Matlab it is probably better than starting from scratch:
https://code.google.com/p/matrbm/
RBM tutorial:
http://deeplearning.net/tutorial/rbm.html#rbm
RBM in scikit-learn:
http://scikit-learn.org/stable/modules/neural_networks.html
A Practical Guide to Training Restricted Boltzmann Machines:
http://www.cs.toronto.edu/~hinton/absps/guideTR.pdf
Multi-scale Oriented Patches (MOPS)
MOPS is a useful tool for many computer vision and related applications. They key is that it gives you a way to generate features that are (mostly) invariant to rotation and scale to use as input for machine learning/artificial intelligence algorithms.
http://www.cs.bath.ac.uk/brown/mops/mops.html
Technical report from Microsoft Research:
http://research.microsoft.com/pubs/70120/tr-2004-133.pdf
Nice slide deck on MOPS:
http://www.csie.ntu.edu.tw/~cyy/courses/vfx/08spring/lectures/handouts/lec06_feature2_4up.pdf
Example app from Microsoft Research:
http://research.microsoft.com/en-us/um/redmond/groups/ivm/PhotoTours/
http://www.cs.bath.ac.uk/brown/mops/mops.html
Technical report from Microsoft Research:
http://research.microsoft.com/pubs/70120/tr-2004-133.pdf
Nice slide deck on MOPS:
http://www.csie.ntu.edu.tw/~cyy/courses/vfx/08spring/lectures/handouts/lec06_feature2_4up.pdf
Example app from Microsoft Research:
http://research.microsoft.com/en-us/um/redmond/groups/ivm/PhotoTours/
AdaBoost
Note: AdaBoost is extremely sensitive to mislabeled samples in your data. For example, if you are trying to classify transactions as either "fraud" or "not fraud" if you have even one mislabeled, then the classifier will over learn that one bad sample and be useless. There are other versions of boosting algorithms that try to overcome this but if you have data for which you can not be sure of the labels then consider using some other method.
AdaBoost, short for "Adaptive Boosting", is a machine learning meta-algorithm formulated by Yoav Freund and Robert Schapire who won the prestigious "Gödel Prize" in 2003 for their work. It can be used in conjunction with many other types of learning algorithms to improve their performance. The output of the other learning algorithms ('weak learners') is combined into a weighted sum that represents the final output of the boosted classifier. AdaBoost is adaptive in the sense that subsequent weak learners are tweaked in favor of those instances misclassified by previous classifiers. AdaBoost is sensitive to noisy data and outliers. In some problems, however, it can be less susceptible to the overfitting problem than other learning algorithms. The individual learners can be weak, but as long as the performance of each one is slightly better than random guessing (i.e., their error rate is smaller than 0.5 for binary classification), the final model can be proven to converge to a strong learner.
While every learning algorithm will tend to suit some problem types better than others, and will typically have many different parameters and configurations to be adjusted before achieving optimal performance on a dataset, AdaBoost (with decision trees as the weak learners) is often referred to as the best out-of-the-box classifier. When used with decision tree learning, information gathered at each stage of the AdaBoost algorithm about the relative 'hardness' of each training sample is fed into the tree growing algorithm such that later trees tend to focus on harder to classify examples.
http://en.wikipedia.org/wiki/AdaBoost
Very nice AdaBoost slide deck:
http://cmp.felk.cvut.cz/~sochmj1/adaboost_talk.pdf
Matlab and C++ implementations:
http://graphics.cs.msu.ru/en/science/research/machinelearning/adaboosttoolbox
AdaBoost, short for "Adaptive Boosting", is a machine learning meta-algorithm formulated by Yoav Freund and Robert Schapire who won the prestigious "Gödel Prize" in 2003 for their work. It can be used in conjunction with many other types of learning algorithms to improve their performance. The output of the other learning algorithms ('weak learners') is combined into a weighted sum that represents the final output of the boosted classifier. AdaBoost is adaptive in the sense that subsequent weak learners are tweaked in favor of those instances misclassified by previous classifiers. AdaBoost is sensitive to noisy data and outliers. In some problems, however, it can be less susceptible to the overfitting problem than other learning algorithms. The individual learners can be weak, but as long as the performance of each one is slightly better than random guessing (i.e., their error rate is smaller than 0.5 for binary classification), the final model can be proven to converge to a strong learner.
While every learning algorithm will tend to suit some problem types better than others, and will typically have many different parameters and configurations to be adjusted before achieving optimal performance on a dataset, AdaBoost (with decision trees as the weak learners) is often referred to as the best out-of-the-box classifier. When used with decision tree learning, information gathered at each stage of the AdaBoost algorithm about the relative 'hardness' of each training sample is fed into the tree growing algorithm such that later trees tend to focus on harder to classify examples.
http://en.wikipedia.org/wiki/AdaBoost
Very nice AdaBoost slide deck:
http://cmp.felk.cvut.cz/~sochmj1/adaboost_talk.pdf
Matlab and C++ implementations:
http://graphics.cs.msu.ru/en/science/research/machinelearning/adaboosttoolbox
Viola Jones object detection framework
The Viola–Jones object detection framework is the first object detection framework to provide competitive object detection rates in real-time proposed in 2001 by Paul Viola and Michael Jones. Although it can be trained to detect a variety of object classes, it was motivated primarily by the problem of face detection. This algorithm is implemented in OpenCV as cvHaarDetectObjects().
YouTube video explaining Viola Jones face detection:
https://www.youtube.com/watch?v=WfdYYNamHZ8
This is a slide deck explaining Viola Jones face detection:
http://www.slideshare.net/wolf/avihu-efrats-viola-and-jones-face-detection-slides/
Haar features are not an AI or ML algorithm themselves, but instead are often a useful tool for transforming data into a format that an AI or ML algorithm can use.
The Wikipedia article only talks about them with respect to object recognition in visible light images. However, they can be used with images from any spectrum or even any type of data that can be represented as X,Y,Z such as a digital elevation map of terrain.
http://en.wikipedia.org/wiki/Haar-like_features
This pdf has a good explanation of how to use Haar features:
http://nichol.as/papers/Wilson/Facial%20feature%20detection%20using%20Haar.pdf
YouTube video explaining Viola Jones face detection:
https://www.youtube.com/watch?v=WfdYYNamHZ8
This is a slide deck explaining Viola Jones face detection:
http://www.slideshare.net/wolf/avihu-efrats-viola-and-jones-face-detection-slides/
Haar features are not an AI or ML algorithm themselves, but instead are often a useful tool for transforming data into a format that an AI or ML algorithm can use.
The Wikipedia article only talks about them with respect to object recognition in visible light images. However, they can be used with images from any spectrum or even any type of data that can be represented as X,Y,Z such as a digital elevation map of terrain.
http://en.wikipedia.org/wiki/Haar-like_features
This pdf has a good explanation of how to use Haar features:
http://nichol.as/papers/Wilson/Facial%20feature%20detection%20using%20Haar.pdf
Friday, September 19, 2014
Genetic Algorithms: Cool Name & Damn Simple
Nice GA tutorial.
Genetic algorithms are a mysterious sounding technique in mysterious sounding field--artificial intelligence. This is the problem with naming things appropriately. When the field was labeled artificial intelligence, it meant using mathematics to artificially create the semblance of intelligence, but self-engrandizing researchers and Isaac Asimov redefined it as robots.
The name genetic algorithms does sound complex and has a faintly magical ring to it, but it turns out that they are one of the simplest and most-intuitive concepts you'll encounter in A.I.
Genetic Algorithms: Cool Name & Damn Simple - Irrational Exuberance
Genetic algorithms are a mysterious sounding technique in mysterious sounding field--artificial intelligence. This is the problem with naming things appropriately. When the field was labeled artificial intelligence, it meant using mathematics to artificially create the semblance of intelligence, but self-engrandizing researchers and Isaac Asimov redefined it as robots.
The name genetic algorithms does sound complex and has a faintly magical ring to it, but it turns out that they are one of the simplest and most-intuitive concepts you'll encounter in A.I.
Genetic Algorithms: Cool Name & Damn Simple - Irrational Exuberance
Curve fitting with Pyevolve
This is a very nice tutorial for genetic algorithms. It uses pyevolve but the tutorial part is useful even if you are using a different language/implementation for GA.
A Coder's Musings: Curve fitting with Pyevolve
A Coder's Musings: Curve fitting with Pyevolve
Genetic Algorithms tutorial
Great tutorial and introduction to genetic algorithms. There are java applets that you can play with to see how GA's work.
These pages introduce some fundamentals of genetic algorithms. Pages are intended to be used for learning about genetic algorithms without any previous knowledge from this area. Only some knowledge of computer programming is assumed. You can find here several interactive Java applets demonstrating work of genetic algorithms.
As the area of genetic algorithms is very wide, it is not possible to cover everything in these pages. But you should get some idea, what the genetic algorithms are and what they could be useful for. Do not expect any sophisticated mathematics theories here.
Main page - Introduction to Genetic Algorithms - Tutorial with Interactive Java Applets
These pages introduce some fundamentals of genetic algorithms. Pages are intended to be used for learning about genetic algorithms without any previous knowledge from this area. Only some knowledge of computer programming is assumed. You can find here several interactive Java applets demonstrating work of genetic algorithms.
As the area of genetic algorithms is very wide, it is not possible to cover everything in these pages. But you should get some idea, what the genetic algorithms are and what they could be useful for. Do not expect any sophisticated mathematics theories here.
Main page - Introduction to Genetic Algorithms - Tutorial with Interactive Java Applets
Pyevolve genetic algorithm python software
I have used this software to successfully create a genetic algorithm python script that I use to tune parameters on extra tree classifiers and RDF classifiers. It is pretty easy to use and you can make almost any type of GA with it. It is open source so you can go in a tinker with it.
Welcome to Pyevolve documentation ! — Pyevolve v0.5 documentation
There is a great pyevolve tutorial here:
A Coder's Musings: Curve fitting with Pyevolve
Welcome to Pyevolve documentation ! — Pyevolve v0.5 documentation
There is a great pyevolve tutorial here:
A Coder's Musings: Curve fitting with Pyevolve
Neural Networks for Machine Learning
This is an online course from the University of Toronto. If you can spend about 8 hours a week for 8 weeks you should be thoroughly familiar with ANN's.
Learn about artificial neural networks and how they're being used for machine learning, as applied to speech and object recognition, image segmentation, modeling language and human motion, etc. We'll emphasize both the basic algorithms and the practical tricks needed to get them to work well.
Neural Networks for Machine Learning | Coursera
Learn about artificial neural networks and how they're being used for machine learning, as applied to speech and object recognition, image segmentation, modeling language and human motion, etc. We'll emphasize both the basic algorithms and the practical tricks needed to get them to work well.
Neural Networks for Machine Learning | Coursera
Basic Neural Network Tutorial : C++ Implementation and Source Code
This tutorial is in two parts, one is the theory of ANN and the other part is a C++ implementation with hints on how to modify it efficiently. If you are new to neural networks and you are a C++ programmer this is a great place to start.
Basic Neural Network Tutorial – Theory | Taking Initiative
Basic Neural Network Tutorial : C++ Implementation and Source Code | Taking Initiative
Basic Neural Network Tutorial – Theory | Taking Initiative
Basic Neural Network Tutorial : C++ Implementation and Source Code | Taking Initiative
Wednesday, September 17, 2014
Theoretical Machine Learning
Introductory lecture notes from a class taught by Professor Rob Schapire who is a leader in the field. This is a good place to start for a beginner in machine learning.
www.cs.princeton.edu/courses/archive/spr08/cos511/scribe_notes/0204.pdf
www.cs.princeton.edu/courses/archive/spr08/cos511/scribe_notes/0204.pdf
Decision Forests for Classification, Regression, Density Estimation, Manifold Learning and Semi-Supervised Learning
This technical report from Microsoft Research is an A-Z tutorial on how decision tree machine learning algorithms work. It includes in depth explanations of random forests, extra tree classifiers, random ferns and other variations for both classification and regression.
It is in report format and compares decision forests to other types of machine learning algorithms such as SVM. Some simple toy problems give the basics and some real life applications such as body position recognition and medical image are included.
There is also an accompanying PowerPoint with some nice animations.
http://research.microsoft.com/pubs/155552/decisionForests_MSR_TR_2011_114.pdf is not available
It is in report format and compares decision forests to other types of machine learning algorithms such as SVM. Some simple toy problems give the basics and some real life applications such as body position recognition and medical image are included.
There is also an accompanying PowerPoint with some nice animations.
http://research.microsoft.com/pubs/155552/decisionForests_MSR_TR_2011_114.pdf is not available
Alternatives to support vector machines in neuroimaging ensembles of decision trees for classification and information mapping with predictive models
This is a nice tutorial for using random decision forests for classifying medical images. There is a comparison with some other methods, especially SVM's.
http://web.stanford.edu/~richiard/slides/PRNI2013Tutorial_export.pdf is not available
http://web.stanford.edu/~richiard/slides/PRNI2013Tutorial_export.pdf is not available
One-Class Support Vector Machines: Methods and Applications
This tutorial on SVM's is a decent one for a beginner. It is a slide deck from a presentation.
isites.harvard.edu/fs/docs/icb.topic274302.files/Dan.Nick.pdf
isites.harvard.edu/fs/docs/icb.topic274302.files/Dan.Nick.pdf
A Comparison of Methods for Multiclass Support Vector Machines
Support vector machines (SVMs) were originally designed for binary classification. How to effectively extend it for multiclass classification is still an ongoing research issue. Several methods have been proposed where typically we construct a multiclass classifier by combining several binary classifiers. Some authors also proposed methods that consider all classes at once. As it is computationally more expensive to solve multiclass problems, comparisons of these methods using large-scale problems have not been seriously conducted. Especially for methods solving multiclass SVM in one step, a much larger optimization problem is required so up to now experiments are limited to small data sets. In this paper we give decomposition implementations for two such “all-together” methods. We then compare their performance with three methods based on binary classifications: “one-against-all,” “one-against-one,” and directed acyclic graph SVM (DAGSVM). Our experiments indicate that the “one-against-one” and DAG methods are more suitable for practical use than the other methods. Results also show that for large problems methods by considering all data at once in general need fewer support vectors.
cs.ecs.baylor.edu/~hamerly/courses/5325_11s/papers/svm/hsu2001multiclass.pdf
cs.ecs.baylor.edu/~hamerly/courses/5325_11s/papers/svm/hsu2001multiclass.pdf
Nonlinear regression in environmental sciences by support vector machines combined with evolutionary strategy
A hybrid algorithm combining support vector regression with evolutionary strategy (SVR-ES) is proposed for predictive models in the environmental sciences. SVR-ES uses uncorrelated mutation with p step sizes to find the optimal SVR hyper-parameters. Three environmental forecast datasets used in the WCCI-2006 contest – surface air temperature, precipitation and sulphur dioxide concentration – were tested. We used multiple linear regression (MLR) as benchmark and a variety of machine learning techniques including bootstrap-aggregated ensemble artificial neural network (ANN), SVR-ES, SVR with hyper-parameters given by the Cherkassky–Ma estimate, the M5 regression tree, and random forest (RF). We also tested all techniques using stepwise linear regression (SLR) first to screen out irrelevant predictors. We concluded that SVR-ES is an attractive approach because it tends to outperform the other techniques and can also be implemented in an almost automatic way. The Cherkassky–Ma estimate is a useful approach for minimizing the mean absolute error and saving computational time related to the hyper-parameter search. The ANN and RF are also good options to outperform multiple linear regression (MLR). Finally, the use of SLR for predictor selection can dramatically reduce computational time and often help to enhance accuracy.
Nonlinear regression in environmental sciences by support vector machines combined with evolutionary strategy
Nonlinear regression in environmental sciences by support vector machines combined with evolutionary strategy
Subscribe to:
Posts (Atom)