Using regression makes extraction of shared variation in multiple datasets easy

作者:Jussi Korpela, Andreas Henelius, Lauri Ahonen, Arto Klami, Kai Puolamäki

摘要

In many data analysis tasks it is important to understand the relationships between different datasets. Several methods exist for this task but many of them are limited to two datasets and linear relationships. In this paper, we propose a new efficient algorithm, termed cocoreg, for the extraction of variation common to all datasets in a given collection of arbitrary size. cocoreg extends redundancy analysis to more than two datasets, utilizing chains of regression functions to extract the shared variation in the original data space. The algorithm can be used with any linear or non-linear regression function, which makes it robust, straightforward, fast, and easy to implement and use. We empirically demonstrate the efficacy of shared variation extraction using the cocoreg algorithm on five artificial and three real datasets.

论文关键词:Shared variation, Multiple regression, Regression chains

论文评审过程:

论文官网地址:https://doi.org/10.1007/s10618-016-0465-y