Javascript must be enabled to continue!
Celeste : A cloud-based genomics infrastructure with variant-calling pipeline suited for population-scale sequencing projects
View through CrossRef
Abstract
Background
The
All of Us
Research Program (
All of Us
) is one of the world’s largest sequencing efforts that will generate genetic data for over one million individuals from diverse backgrounds. This historic megaproject will create novel research platforms that integrate an unprecedented amount of genetic data with longitudinal health information. Here, we describe the design of
Celeste
, a resilient, open-source cloud architecture for implementing genomics workflows that has successfully analyzed petabytes of participant genomic information for
All of Us
– thereby enabling other large-scale sequencing efforts with a comprehensive set of tools to power analysis. The
Celeste
infrastructure is tremendously scalable and has routinely processed fluctuating workloads of up to 9,000 whole-genome sequencing (WGS) samples for
All of Us
, monthly. It also lends itself to multiple projects. Serverless technology and container orchestration form the basis of
Celeste
’s system for managing this volume of data.
Results
In 12 months of production (within a single Amazon Web Services (AWS) Region), around 200 million serverless functions and over 20 million messages coordinated the analysis of 1.8 million bioinformatics, quality control, and clinical reporting jobs. Adapting WGS analysis to clinical projects requires adaptation of variant-calling methods to enrich the reliable detection of variants with known clinical importance. Thus, we also share the process by which we tuned the variant-calling pipeline in use by the multiple genome centers supporting
All of Us
to maximize precision and accuracy for low fraction variant calls with clinical significance.
Conclusions
When combined with hardware-accelerated implementations for genomic analysis, Celeste had far-reaching, positive implications for turn-around time, dynamic scalability, security, and storage of analysis for one hundred-thousand whole-genome samples and counting. Other groups may align their sequencing workflows to this harmonized pipeline standard, included within the
Celeste
framework, to meet clinical requisites for population-scale sequencing efforts.
Celeste
is available as an Amazon Web Services (AWS) deployment in GitHub, and includes command-line parameters and software containers.
Title: Celeste
: A cloud-based genomics infrastructure with variant-calling pipeline suited for population-scale sequencing projects
Description:
Abstract
Background
The
All of Us
Research Program (
All of Us
) is one of the world’s largest sequencing efforts that will generate genetic data for over one million individuals from diverse backgrounds.
This historic megaproject will create novel research platforms that integrate an unprecedented amount of genetic data with longitudinal health information.
Here, we describe the design of
Celeste
, a resilient, open-source cloud architecture for implementing genomics workflows that has successfully analyzed petabytes of participant genomic information for
All of Us
– thereby enabling other large-scale sequencing efforts with a comprehensive set of tools to power analysis.
The
Celeste
infrastructure is tremendously scalable and has routinely processed fluctuating workloads of up to 9,000 whole-genome sequencing (WGS) samples for
All of Us
, monthly.
It also lends itself to multiple projects.
Serverless technology and container orchestration form the basis of
Celeste
’s system for managing this volume of data.
Results
In 12 months of production (within a single Amazon Web Services (AWS) Region), around 200 million serverless functions and over 20 million messages coordinated the analysis of 1.
8 million bioinformatics, quality control, and clinical reporting jobs.
Adapting WGS analysis to clinical projects requires adaptation of variant-calling methods to enrich the reliable detection of variants with known clinical importance.
Thus, we also share the process by which we tuned the variant-calling pipeline in use by the multiple genome centers supporting
All of Us
to maximize precision and accuracy for low fraction variant calls with clinical significance.
Conclusions
When combined with hardware-accelerated implementations for genomic analysis, Celeste had far-reaching, positive implications for turn-around time, dynamic scalability, security, and storage of analysis for one hundred-thousand whole-genome samples and counting.
Other groups may align their sequencing workflows to this harmonized pipeline standard, included within the
Celeste
framework, to meet clinical requisites for population-scale sequencing efforts.
Celeste
is available as an Amazon Web Services (AWS) deployment in GitHub, and includes command-line parameters and software containers.
Related Results
CLOUD COMPUTING - NAVIGATING THE DIGITAL SKY
CLOUD COMPUTING - NAVIGATING THE DIGITAL SKY
“Cloud Computing – Navigating the Digital Sky” is an extensive guide designed to provide a thorough understanding of cloud computing, an essential technology in today’s digital age...
ANALISIS PERTIMBANGAN MAHKAMAH AGUNG DALAM MENGABULKAN KASASI TERDAKWA (STUDI PUTUSAN NOMOR 2959/K/PID.SUS/2022)
ANALISIS PERTIMBANGAN MAHKAMAH AGUNG DALAM MENGABULKAN KASASI TERDAKWA (STUDI PUTUSAN NOMOR 2959/K/PID.SUS/2022)
<p><em><span class="markedContent"><span style="left: calc(var(--scale-factor)*195.53px); top: calc(var(--scale-factor)*496.87px); font-size: calc(var(--scale-...
Frequency of Common Chromosomal Abnormalities in Patients with Idiopathic Acquired Aplastic Anemia
Frequency of Common Chromosomal Abnormalities in Patients with Idiopathic Acquired Aplastic Anemia
Objective: To determine the frequency of common chromosomal aberrations in local population idiopathic determine the frequency of common chromosomal aberrations in local population...
MARS-seq2.0: an experimental and analytical pipeline for indexed sorting combined with single-cell RNA sequencing v1
MARS-seq2.0: an experimental and analytical pipeline for indexed sorting combined with single-cell RNA sequencing v1
Human tissues comprise trillions of cells that populate a complex space of molecular phenotypes and functions and that vary in abundance by 4–9 orders of magnitude. Relying solely ...
KEDUDUKAN AHLI BAHASA DALAM PEMBUKTIAN PERKARA PENCEMARAN NAMA BAIK (STUDI PUTUSAN NOMOR: 47/PID.SUS/2019/PN. MGT)
KEDUDUKAN AHLI BAHASA DALAM PEMBUKTIAN PERKARA PENCEMARAN NAMA BAIK (STUDI PUTUSAN NOMOR: 47/PID.SUS/2019/PN. MGT)
<em><span id="page3R_mcid52" class="markedContent"><span style="left: calc(var(--scale-factor)*125.30px); top: calc(var(--scale-factor)*539.11px); font-size: calc(va...
Abstract P1-05-23: Utilities and challenges of RNA-Seq based expression and variant calling in a clinical setting
Abstract P1-05-23: Utilities and challenges of RNA-Seq based expression and variant calling in a clinical setting
Abstract
Introduction
Variant calling based on DNA samples has been the gold standard of clinical testing since the advent of Sanger sequencing. The u...
First Arctic Subsea Pipelines Moving to Reality
First Arctic Subsea Pipelines Moving to Reality
Abstract
Two offshore development projects which involve subsea arctic pipelines are being proposed by British Petroleum Exploration (BP). Both projects are locat...
Installation Analysis of Matterhorn Pipeline Replacement
Installation Analysis of Matterhorn Pipeline Replacement
Abstract
The paper describes the installation analysis for the Matterhorn field pipeline replacement, located in water depths between 800-ft to 1200-ft in the Gul...

