Javascript must be enabled to continue!
Long-term Reproducibility for Jupyter Notebook
View through CrossRef
Computational notebooks (e.g. Jupyter notebook) are a popular choice for interactive scientific computing to convey descriptive information together with executable source code. The user can annotate the scientific development of the work, the methods applied, describe ancillary data or the analysis of results, with text, illustrations, figures, and equations. Such ‘executable’ documents provide a paradigm shift in scientific writing, where not only the science is described, but the actual computation and source code are openly available and can be reproduced and validated.Therefore, it is of paramount importance to preserve these documents. A unique and persistent identification (PID) is essential together with providing enough information to execute the source code. Generating a PID for a Jupyter notebook is not technically challenging. We can automatically collect system and run-time information and, with a guided workflow for the user, assemble a rich set of metadata. The collected information allows us to recreate the computational environment and run the source code, which in return (theoretically) should produce the same results as published.The importance of providing a rich set of metadata for all digital objects in a human readable and machine actionable form is well understood and widely accepted as necessity for reproducibility, traceability, and provenance. This is reflected in the FAIR principles (Wilkinson, https://doi.org/10.1038/sdata.2016.18) which are regarded as gold standard by many scientific communities.Pimentel et al. (https://doi.org/10.1109/MSR.2019.00077) analysed over 800’000 Jupyter notebooks from GitHub. 24 % executed without errors and only 4 % produced the same results. The likelihood to successfully compile and run a decade old source code is slim. Long term support for well established operating systems varies between 5 to 10 years, user software support is usually shorter and looking at free and open-source repositories there is often no support (or best effort) offered.We present an approach to safely reproduce the computational environment in the future with a focus on long-term availability. Instead of trying to reinstall the computational environment based on the stored metadata, we propose to archive the docker image, the user space (user installed packages) and finally the source code. Recreating the system in this way is more like restoring a backup, where backup is the equivalent of an entire computer system. It does not solve all the problems but removes a great deal of complexity and uncertainty.Though there are shortcomings in our approach, we believe our solution will lower the threshold for scientists to provide rich meta data, code and results attached to a publication that can be reproduced in the far future.
Title: Long-term Reproducibility for Jupyter Notebook
Description:
Computational notebooks (e.
g.
Jupyter notebook) are a popular choice for interactive scientific computing to convey descriptive information together with executable source code.
The user can annotate the scientific development of the work, the methods applied, describe ancillary data or the analysis of results, with text, illustrations, figures, and equations.
Such ‘executable’ documents provide a paradigm shift in scientific writing, where not only the science is described, but the actual computation and source code are openly available and can be reproduced and validated.
Therefore, it is of paramount importance to preserve these documents.
A unique and persistent identification (PID) is essential together with providing enough information to execute the source code.
Generating a PID for a Jupyter notebook is not technically challenging.
We can automatically collect system and run-time information and, with a guided workflow for the user, assemble a rich set of metadata.
The collected information allows us to recreate the computational environment and run the source code, which in return (theoretically) should produce the same results as published.
The importance of providing a rich set of metadata for all digital objects in a human readable and machine actionable form is well understood and widely accepted as necessity for reproducibility, traceability, and provenance.
This is reflected in the FAIR principles (Wilkinson, https://doi.
org/10.
1038/sdata.
2016.
18) which are regarded as gold standard by many scientific communities.
Pimentel et al.
(https://doi.
org/10.
1109/MSR.
2019.
00077) analysed over 800’000 Jupyter notebooks from GitHub.
24 % executed without errors and only 4 % produced the same results.
The likelihood to successfully compile and run a decade old source code is slim.
Long term support for well established operating systems varies between 5 to 10 years, user software support is usually shorter and looking at free and open-source repositories there is often no support (or best effort) offered.
We present an approach to safely reproduce the computational environment in the future with a focus on long-term availability.
Instead of trying to reinstall the computational environment based on the stored metadata, we propose to archive the docker image, the user space (user installed packages) and finally the source code.
Recreating the system in this way is more like restoring a backup, where backup is the equivalent of an entire computer system.
It does not solve all the problems but removes a great deal of complexity and uncertainty.
Though there are shortcomings in our approach, we believe our solution will lower the threshold for scientists to provide rich meta data, code and results attached to a publication that can be reproduced in the far future.
Related Results
Enhancing Learning About Epidemiological Data Analysis Using R for Graduate Students in Medical Fields With Jupyter Notebook: Classroom Action Research (Preprint)
Enhancing Learning About Epidemiological Data Analysis Using R for Graduate Students in Medical Fields With Jupyter Notebook: Classroom Action Research (Preprint)
BACKGROUND
Graduate students in medical fields must learn about epidemiology and data analysis to conduct their research. R is a software environment used t...
Enhancing Learning About Epidemiological Data Analysis Using R for Graduate Students in Medical Fields With Jupyter Notebook: Classroom Action Research
Enhancing Learning About Epidemiological Data Analysis Using R for Graduate Students in Medical Fields With Jupyter Notebook: Classroom Action Research
Background
Graduate students in medical fields must learn about epidemiology and data analysis to conduct their research. R is a software environment used to de...
Jupyter notebooks in science gateways
Jupyter notebooks in science gateways
Jupyter Notebooks empower scientists to create executable documents that include text, equations, code and figures. Notebooks are a simple way to create reproducible and shareable ...
Jupyter notebooks in science gateways
Jupyter notebooks in science gateways
Jupyter Notebooks empower scientists to create executable documents that include text, equations, code and figures. Notebooks are a simple way to create reproducible and shareable ...
Understanding reproducibility of bioinformatics workflows
Understanding reproducibility of bioinformatics workflows
Reproducibility is an essential factor in establishing the reliability of results reported in the scientific literature and ultimately in the knowledge that is generated. This know...
Short- and long-term reproducibility of peripheral superficial vein depth and diameter measurements using ultrasound imaging
Short- and long-term reproducibility of peripheral superficial vein depth and diameter measurements using ultrasound imaging
Abstract
Background
Ultrasound imaging is used for diagnosis, treatment, and blood vessel visualization during venous cat...
Improving Reproducibility in AI Research: Four Mechanisms Adopted by JAIR
Improving Reproducibility in AI Research: Four Mechanisms Adopted by JAIR
Background: Lately, the reproducibility of scientific results has become an increasing worry in the scientific community. Several studies show that artificial intelligence research...
Fostering Computational Skills in Secondary Education Earth Sciences through Jupyter Notebooks
Fostering Computational Skills in Secondary Education Earth Sciences through Jupyter Notebooks
As a multi-year coding in the classroom initiative, the Earth Science
Information Partners’ education committee and community members have
been exploring how to integrate coding sk...

