Description
Abstract:With the rapid expansion of big data in all science and engineering domains, the potential that lays behind these massive data is undoubtedly significant. Leveraging such information in an optimal way requires innovative strategies of thinking and novel learning techniques to address the various accompanying challenges. Traditional machine learning techniques have been left behind during the past technological advances in complex data-intensive fields such as medicine, astronomy, and biology. The sizeable volume of high scale, noisy, and heterogeneous data being collected from complex domains has deprived the traditional learning approaches from significant capabilities for mining hidden information. To orderly cope with these critical limitations, a broad taxonomy for pivotal issues of machine learning techniques has been proposed where large scale data analysis, heterogeneity modeling, and uncertainty quantification represent the most prevalent classes in practice. For large scale data analysis, the assumption that the collected observations need to be completely loaded into memory for centralized processing makes the traditional learning techniques incapacitated in the absence of exascale computing resources. Heterogeneity presents another level of complexity, as the efficacy of the learning process would be contingent on the capacity of handling the discrepancy in sources and types of the available data. As for uncertainty, it is increasingly appreciated as an epistemic situation of imperfect knowledge that requires a proper modeling and quantification. From this perspective, we investigate in this work the class of data-based integrative approaches with a special focus on heterogeneous data analysis and uncertainty modeling. Towards this goal, we provide a set of innovative solutions to real-world applications that fall within scientific machine learning preoccupations. First, we investigate the problem of data clustering with different layers of information where we propose a novel hybrid clustering algorithm that successfully combines two different types of data to conduct an efficient clustering. Additionally, we study the problem of biomarker discovery and disease prognosis for complex disorders, and we provide an innovative approach that integrates network prior knowledge with gene expression data to derive compact representations by graph embedding which enables an accurate identification of robust disease markers. Ultimately, we explore the problem of error estimation based on transfer learning within a Bayesian paradigm under uncertainty. We introduce a novel class of Bayesian minimum mean square error estimators that can leverage different source domains to accurately estimate the classification error in a target domain of interest. In summary, our work collectively demonstrates the strong potential of integrative approaches in leveraging network and cross-domain knowledge for an effective learning from data. The electronic version of this dissertation is accessible from https://hdl.handle.net/1969.1/198481
Item Description:"Major Subject: Electrical Engineering"
Includes vita.
Physical Description:1 online resource.
Bibliography:Includes bibliographical references.