Sequential tests for large-scale learning

Anoop Korattikara,Yutian Chen,Max Welling

Sequential tests for large-scale learning

2016

Anoop Korattikara
Yutian Chen
Max Welling

We argue that when faced with big data sets, learning and inference algorithms should compute updates using only subsets of data items. We introduce algorithms that use sequential hypothesis tests to adaptively select such a subset of data points. The statistical properties of this subsampling process can be used to control the efficiency and accuracy of learning or inference. In the context of learning by optimization, we test for the probability that the update direction is no more than 90 degrees in the wrong direction. In the context of posterior inference using Markov chain Monte Carlo, we test for the probability that our decision to accept or reject a sample is wrong. We experimentally evaluate our algorithms on a number of models and data sets.

Keywords:

Statistical hypothesis testing
Big data
Mathematics
Artificial intelligence
Machine learning
Data point
Data set
Markov chain Monte Carlo
Inference
Text mining
Wrong direction

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations