Risk Bounds for Positive-Unlabeled Learning Under the Selected At Random Assumption

Olivier Coudray; Christine Keribin; Pascal Massart; Patrick Pamphile

Positive-Unlabeled learning (PU learning) is a special case of semi-supervised binary classification where only a fraction of positive examples is labeled. The challenge is then to find the correct classifier despite this lack of information. Recently, new methodologies have been introduced to address the case where the probability of being labeled may depend on the covariates. In this paper, we are interested in establishing risk bounds for PU learning under this general assumption. In addition, we quantify the impact of label noise on PU learning compared to the standard classification setting. Finally, we provide a lower bound on the minimax risk proving that the upper bound is almost optimal.

Risk Bounds for Positive-Unlabeled Learning Under the Selected At Random Assumption

Abstract