Javascript must be enabled to continue!
A Study on Gene Selection and Classification Algorithms for Classification of Microarray Gene Expression Data
View through CrossRef
Pembangunan teknologi microarray membenarkan penyelidik untuk meneliti tahap ekspresi gen dalam sel. Salah satu aplikasi teknologi microarray adalah pengkelasan sampel tisu kepada tisu kanser atau tisu biasa. Pemilihan gen memainkan peranan yang penting sebelum pengkelasan. Dalam makalah ini, beberapa kombinasi teknik pemilihan gen dan teknik pengkelasan yang berlainan untuk pengkelasan data expresi gen microarray telah dikaji. Teknik pemilihan gen terdiri dari Fisher Criterion, Golub Signal–to–Noise, traditional t–test dan Mann–Whitney rank sum statistic. Teknik pengkelasan terdiri dari support vector machines (SVMs) dengan pelbagai kernel dan k–nearest neighoor (k–nn). Prestasi kombinasi teknik–teknik yang dikaji disahkan dengan menggunakan teknik leave–one–out cross validation (LOOCV) dan receiver operating characteristic (ROC) digunakan untuk menganalisa prestasi kombinasi teknik–teknik yang dikaji. Kajian yang telah dijalankan dalam eksperimen ini menunjukkan bahawa pemilihan gen sebelum pengkelasan adalah penting untuk memperolehi prestasi pengkelasan yang lebih baik. Kombinasi yang menghasilkan prestasi tertinggi adalah dengan menggunakan Mann–Whitney rank sum statistic dan SVMs. Nilai ROC tertinggi yang dicapai oleh kombinasi ini adalah 0.91. Ini adalah penting bagi tujuan rawatan dan kajian biologi seterusnya.
Kata kunci: Data expresi gen microarray, pemilihan gen, kaedah statistik, algoritma pengkelasan, Support Vector Machines, k-nearest neighbor
The development of microarray technology allows researchers to monitor the expression of genes on a genomic scale. One of the main applications of microarray technology is the classification of tissue samples into tumor or normal tissue. Gene selection plays an important role prior to tissue classification. In this paper, a study on numerous combinations of gene selection techniques and classifcation algorithms for classification of microarray gene expression data is presented. The gene selection techniques include Fisher Criterion, Golub Signal–to–Noise, traditional t–test and Mann–whitney rank sum statistic. The classification algorithms include support vector machines (SVMs) with several kernels and k–nearest neighbor (k–nn). The performance of the combined techniques is validated by using leave–one–out cross validabon (LOOCV) technique and receiver operating characteristic (ROC) is used to analyze the results. The study demonstrated that selecting genes prior to tissue classification plays an important role for a better classification performance. The best combination is obtained by using Mann–Whitney Rank Sum Statistic and SVMs. The best ROC score achieved for this combination is at 0.91. This should be of significant value for diagnostic purposes as well as for guiding further exploration of the underlying biology.
Key words: Microarray gene expression data, gene selection, statistical methods, classification algorithms, support vector machines, k-nearest neighbor
Title: A Study on Gene Selection and Classification Algorithms for Classification of Microarray Gene Expression Data
Description:
Pembangunan teknologi microarray membenarkan penyelidik untuk meneliti tahap ekspresi gen dalam sel.
Salah satu aplikasi teknologi microarray adalah pengkelasan sampel tisu kepada tisu kanser atau tisu biasa.
Pemilihan gen memainkan peranan yang penting sebelum pengkelasan.
Dalam makalah ini, beberapa kombinasi teknik pemilihan gen dan teknik pengkelasan yang berlainan untuk pengkelasan data expresi gen microarray telah dikaji.
Teknik pemilihan gen terdiri dari Fisher Criterion, Golub Signal–to–Noise, traditional t–test dan Mann–Whitney rank sum statistic.
Teknik pengkelasan terdiri dari support vector machines (SVMs) dengan pelbagai kernel dan k–nearest neighoor (k–nn).
Prestasi kombinasi teknik–teknik yang dikaji disahkan dengan menggunakan teknik leave–one–out cross validation (LOOCV) dan receiver operating characteristic (ROC) digunakan untuk menganalisa prestasi kombinasi teknik–teknik yang dikaji.
Kajian yang telah dijalankan dalam eksperimen ini menunjukkan bahawa pemilihan gen sebelum pengkelasan adalah penting untuk memperolehi prestasi pengkelasan yang lebih baik.
Kombinasi yang menghasilkan prestasi tertinggi adalah dengan menggunakan Mann–Whitney rank sum statistic dan SVMs.
Nilai ROC tertinggi yang dicapai oleh kombinasi ini adalah 0.
91.
Ini adalah penting bagi tujuan rawatan dan kajian biologi seterusnya.
Kata kunci: Data expresi gen microarray, pemilihan gen, kaedah statistik, algoritma pengkelasan, Support Vector Machines, k-nearest neighbor
The development of microarray technology allows researchers to monitor the expression of genes on a genomic scale.
One of the main applications of microarray technology is the classification of tissue samples into tumor or normal tissue.
Gene selection plays an important role prior to tissue classification.
In this paper, a study on numerous combinations of gene selection techniques and classifcation algorithms for classification of microarray gene expression data is presented.
The gene selection techniques include Fisher Criterion, Golub Signal–to–Noise, traditional t–test and Mann–whitney rank sum statistic.
The classification algorithms include support vector machines (SVMs) with several kernels and k–nearest neighbor (k–nn).
The performance of the combined techniques is validated by using leave–one–out cross validabon (LOOCV) technique and receiver operating characteristic (ROC) is used to analyze the results.
The study demonstrated that selecting genes prior to tissue classification plays an important role for a better classification performance.
The best combination is obtained by using Mann–Whitney Rank Sum Statistic and SVMs.
The best ROC score achieved for this combination is at 0.
91.
This should be of significant value for diagnostic purposes as well as for guiding further exploration of the underlying biology.
Key words: Microarray gene expression data, gene selection, statistical methods, classification algorithms, support vector machines, k-nearest neighbor.
Related Results
Selection Gradients
Selection Gradients
Natural selection and sexual selection are important evolutionary processes that can shape the phenotypic distributions of natural populations and, consequently, a primary goal of ...
Hierarchical information representation and efficient classification of gene expression microarray data
Hierarchical information representation and efficient classification of gene expression microarray data
In the field of computational biology, microarryas are used to measure the activity of thousands of genes at once and create a global picture of cellular function. Microarrays allo...
Microrna Regulation of Nodule Zone-Specific Gene Expression In Soybean
Microrna Regulation of Nodule Zone-Specific Gene Expression In Soybean
Nitrogen is a paramount important essential element for all living organisms. It has been found to bea crucial structural component of proteins, nucleic acids, enzymes and other ce...
Enhancing Cancerous Gene Selection and Classification for High-Dimensional Microarray Data Using a Novel Hybrid Filter and Differential Evolutionary Feature Selection
Enhancing Cancerous Gene Selection and Classification for High-Dimensional Microarray Data Using a Novel Hybrid Filter and Differential Evolutionary Feature Selection
Background: In recent years, microarray datasets have been used to store information about human genes and methods used to express the genes in order to successfully diagnose cance...
The Prognostic Impact of High MEL1 Gene Expression in Pediatric Acute Myeloid Leukemia
The Prognostic Impact of High MEL1 Gene Expression in Pediatric Acute Myeloid Leukemia
Abstract
Background
Acute myeloid leukemia (AML) is a complex disease caused by mutations, epigenetic modifications, and deregulated expression of gen...
Binary Puma Optimizer: A Metaheuristic Approach for Gene Selection in Bioinformatics with Machine Learning Classification
Binary Puma Optimizer: A Metaheuristic Approach for Gene Selection in Bioinformatics with Machine Learning Classification
Abstract
Gene expression data is a matrix derived from DNA microarray analysis; a single chip can contain a thousand genetic instructions via DNA microarray technol...
Microarray Image Analysis Using Genetic Algorithm
Microarray Image Analysis Using Genetic Algorithm
<p>Microarray technology allows the simultaneous monitoring of thousands of genes. Based on the gene expression measurements, microarray technology have proven powerful in ge...

