Probably approximately correct learning

Machine learning and data mining

Problems Classification Clustering Regression Anomaly detection Association rules Reinforcement learning Structured prediction Feature engineering Feature learning Online learning Semi-supervised learning Unsupervised learning Learning to rank Grammar induction
Supervised learning (classification • regression) Decision trees Ensembles (Bagging, Boosting, Random forest) k-NN Linear regression Naive Bayes Neural networks Logistic regression Perceptron Relevance vector machine (RVM) Support vector machine (SVM)
Clustering BIRCH Hierarchical k-means Expectation-maximization (EM) DBSCAN OPTICS Mean-shift
Dimensionality reduction Factor analysis CCA ICA LDA NMF PCA t-SNE
Structured prediction Graphical models (Bayes net, CRF, HMM)
Anomaly detection k-NN Local outlier factor
Neural nets Autoencoder Deep learning Multilayer perceptron RNN Restricted Boltzmann machine SOM Convolutional neural network
Reinforcement Learning Q-Learning SARSA Temporal Difference (TD)
Theory Bias-variance dilemma Computational learning theory Empirical risk minimization Occam learning PAC learning Statistical learning VC theory
Machine learning venues NIPS ICML ML JMLR ArXiv:cs.LG
Related articles List of datasets for machine learning research Outline of machine learning
Machine learning portal

In computational learning theory, probably approximately correct learning (PAC learning) is a framework for mathematical analysis of machine learning. It was proposed in 1984 by Leslie Valiant.^[1]

In this framework, the learner receives samples and must select a generalization function (called the hypothesis) from a certain class of possible functions. The goal is that, with high probability (the "probably" part), the selected function will have low generalization error (the "approximately correct" part). The learner must be able to learn the concept given any arbitrary approximation ratio, probability of success, or distribution of the samples.

The model was later extended to treat noise (misclassified samples).

An important innovation of the PAC framework is the introduction of computational complexity theory concepts to machine learning. In particular, the learner is expected to find efficient functions (time and space requirements bounded to a polynomial of the example size), and the learner itself must implement an efficient procedure (requiring an example count bounded to a polynomial of the concept size, modified by the approximation and likelihood bounds).

Definitions and terminology

In order to give the definition for something that is PAC-learnable, we first have to introduce some terminology.^[2]^[3]

For the following definitions, two examples will be used. The first is the problem of character recognition given an array of $n$ bits encoding a binary-valued image. The other example is the problem of finding an interval that will correctly classify points within the interval as positive and the points outside of the range as negative.

Let $X$ be a set called the instance space or the encoding of all the samples. In the character recognition problem, the instance space is $X=\{0,1\}^{n}$ . In the interval problem the instance space, $X$ , is the set of all bounded intervals in $\mathbb {R}$ , where $\mathbb {R}$ denotes the set of all real numbers.

A concept is a subset $c\subset X$ . One concept is the set of all patterns of bits in $X=\{0,1\}^{n}$ that encode a picture of the letter "P". An example concept from the second example is the set of open intervals, $\{(a,b)\mid 0\leq a\leq \pi /2,\pi \leq b\leq {\sqrt {13}}\}$ , each of which contain only the positive points. A concept class $C$ is a set of concepts over $X$ . This could be the set of all subsets of the array of bits that are skeletonized 4-connected (width of the font is 1).

Let $EX(c,D)$ be a procedure that draws an example, $x$ , using a probability distribution $D$ and gives the correct label $c(x)$ , that is 1 if $x\in c$ and 0 otherwise.

Now, given $0<\epsilon ,\delta <1$ , assume there is an algorithm $A$ and a polynomial $p$ in $1/\epsilon ,1/\delta$ (and other relevant parameters of the class $C$ ) such that, given a sample of size p drawn according to $EX(c,D)$ , then, with probability of at least $1-\delta$ , $A$ outputs a hypothesis $h\in C$ that has an average error less than or equal to $\epsilon$ on $X$ with the same distribution $D$ . Further if the above statement for algorithm $A$ is true for every concept $c\in C$ and for every distribution $D$ over $X$ , and for all $0<\epsilon ,\delta <1$ then $C$ is (efficiently) PAC learnable (or distribution-free PAC learnable). We can also say that $A$ is a PAC learning algorithm for $C$ .

Equivalence

Under some regularity conditions these three conditions are equivalent:

The concept class C is PAC learnable.
The VC dimension of C is finite.
C is a uniform Glivenko-Cantelli class.

References

↑ L. Valiant. A theory of the learnable. Communications of the ACM, 27, 1984.
↑ Kearns and Vazirani, pg. 1-12,
↑ Balas Kausik Natarajan, Machine Learning , A Theoretical Approach, Morgan Kaufmann Publishers, 1991

Probably approximately correct learning

Definitions and terminology

Equivalence

See also

References

Further reading