Search engine for discovering works of Art, research articles, and books related to Art and Culture
ShareThis
Javascript must be enabled to continue!

Neural Dimensionality Reduction for Data Visualization

View through CrossRef
Information is a crucial resource for humankind, as it allows us to achieve previously unimaginable goals. The bottleneck for the development of complex systems is often the ability to collect high-quality and highly descriptive data characterizing phenomena of interest. While simple phenomena might be characterizable by measuring a few attributes, this stops being the case as our objects of study become complex. Modeling complex events requires collecting complex data, sometimes containing thousands of attributes, which we call high-dimensional data. In this thesis, we study the visualization of high-dimensional data using Dimensionality Reduction: techniques that cast the problem of finding patterns directly in the data to that of inspecting visual patterns in an alternative representation. We notice how several properties from Neural Networks - a class of Machine Learning models - synergyze with Dimensionality Reduction: such models possess high expressive power, generalize to unseen data, support a variety of data types, and can express useful inductive biases through their architecture. We call the interplay of these two kinds of techniques Neural Dimensionality Reduction. We first study how to approximate classic Dimensionality Reduction algorithms through the use of Neural Networks. Such approximations allow us to build fast and accurate projections of datasets comprised of millions of data points - a threshold at which traditional Dimensionality Reduction methods hit scalability ceilings. Other techniques exist to address this goal, but their limitations impede their widespread adoption. We lift these limitations, proposing a new approximate projection algorithm. Further, we are motivated by a curious aspect of Dimensionality Reduction algorithms: They tend to reproduce, sometimes irrespective of dataset, recognizable visual patterns. These patterns are artifacts of the algorithm's design, and might mislead users into inferences unsupported by the data. We investigate whether these patterns can be placed under the user's direct control, so that their provenance is known ahead of time and not a confounder. Our study shows that indeed, with an appropriate network architecture, this is feasible. Next, we study how the output of Dimensionality Reduction algorithms is evaluated. The standard approach is to use Projection Quality Metrics: functions that measure the degree of pattern preservation in a projection. We show an important flaw of such metrics: they behave akin to summary statistics, in that they cannot fully identify whether a projection is showing true patterns existing in the data. Finally, we turn our attention to Explainable AI, a set of techniques aimed at shedding light onto the outputs of a Machine Learning model. Since Dimensionality Reduction is our object of study, we take a visualization-based approach. We propose extensions to a technique called the Decision Boundary Map, without which we argue the visualization over-simplifies the behavior of the model. Our contributions show, across a wide spectrum of applications, how Neural Networks can power Dimensionality Reduction techniques: Their versatility allows them to be used to approximate, create, evaluate, and invert projections. Perhaps more importantly, their use effectively imbues DR algorithms with novel, useful capabilities, a sample of which is described in this thesis.
Utrecht University Library
Title: Neural Dimensionality Reduction for Data Visualization
Description:
Information is a crucial resource for humankind, as it allows us to achieve previously unimaginable goals.
The bottleneck for the development of complex systems is often the ability to collect high-quality and highly descriptive data characterizing phenomena of interest.
While simple phenomena might be characterizable by measuring a few attributes, this stops being the case as our objects of study become complex.
Modeling complex events requires collecting complex data, sometimes containing thousands of attributes, which we call high-dimensional data.
In this thesis, we study the visualization of high-dimensional data using Dimensionality Reduction: techniques that cast the problem of finding patterns directly in the data to that of inspecting visual patterns in an alternative representation.
We notice how several properties from Neural Networks - a class of Machine Learning models - synergyze with Dimensionality Reduction: such models possess high expressive power, generalize to unseen data, support a variety of data types, and can express useful inductive biases through their architecture.
We call the interplay of these two kinds of techniques Neural Dimensionality Reduction.
We first study how to approximate classic Dimensionality Reduction algorithms through the use of Neural Networks.
Such approximations allow us to build fast and accurate projections of datasets comprised of millions of data points - a threshold at which traditional Dimensionality Reduction methods hit scalability ceilings.
Other techniques exist to address this goal, but their limitations impede their widespread adoption.
We lift these limitations, proposing a new approximate projection algorithm.
Further, we are motivated by a curious aspect of Dimensionality Reduction algorithms: They tend to reproduce, sometimes irrespective of dataset, recognizable visual patterns.
These patterns are artifacts of the algorithm's design, and might mislead users into inferences unsupported by the data.
We investigate whether these patterns can be placed under the user's direct control, so that their provenance is known ahead of time and not a confounder.
Our study shows that indeed, with an appropriate network architecture, this is feasible.
Next, we study how the output of Dimensionality Reduction algorithms is evaluated.
The standard approach is to use Projection Quality Metrics: functions that measure the degree of pattern preservation in a projection.
We show an important flaw of such metrics: they behave akin to summary statistics, in that they cannot fully identify whether a projection is showing true patterns existing in the data.
Finally, we turn our attention to Explainable AI, a set of techniques aimed at shedding light onto the outputs of a Machine Learning model.
Since Dimensionality Reduction is our object of study, we take a visualization-based approach.
We propose extensions to a technique called the Decision Boundary Map, without which we argue the visualization over-simplifies the behavior of the model.
Our contributions show, across a wide spectrum of applications, how Neural Networks can power Dimensionality Reduction techniques: Their versatility allows them to be used to approximate, create, evaluate, and invert projections.
Perhaps more importantly, their use effectively imbues DR algorithms with novel, useful capabilities, a sample of which is described in this thesis.

Related Results

A high-dimensionality-trait-driven learning paradigm for high dimensional credit classification
A high-dimensionality-trait-driven learning paradigm for high dimensional credit classification
AbstractTo solve the high-dimensionality issue and improve its accuracy in credit risk assessment, a high-dimensionality-trait-driven learning paradigm is proposed for feature extr...
New Perspectives for 3D Visualization of Dynamic Reservoir Uncertainty
New Perspectives for 3D Visualization of Dynamic Reservoir Uncertainty
This reference is for an abstract only. A full paper was not submitted for this conference. Abstract 1 Int...
Dimensionality Reduction Techniques for IoT Based Data
Dimensionality Reduction Techniques for IoT Based Data
Background: Internet of Things (IoT) plays a vital role by connecting several heterogeneous devices seamlessly via the Internet through new services. Every second, the scale of IoT...
Between Cluster Analysis: Supervised Dimensionality Reduction for Trajectory Inference
Between Cluster Analysis: Supervised Dimensionality Reduction for Trajectory Inference
Abstract Motivation Single-cell RNA sequencing (scRNA-seq) measures the transcriptional state of individual cells, enabling more...
3D Reservoir Visualization
3D Reservoir Visualization
Summary This paper shows how some simple 3D graphics tools can be combined to provide efficient soft-ware for visualizing and analyzing data obtained from reservo...
Neural stemness contributes to cell tumorigenicity
Neural stemness contributes to cell tumorigenicity
Abstract Background: Previous studies demonstrated the dependence of cancer on nerve. Recently, a growing number of studies reveal that cancer cells share the property and ...
Reduction of Dimensionality
Reduction of Dimensionality
Abstract In data analysis, the expression “Reduction of dimensionality”, or “Dimensionality reduction”, refers to the process of mapping a set of high‐dimensional statist...

Back to Top