<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom"><title>Jupyter Blog - Lindsey Heagy</title><link href="https://jasongrout.github.io/medium-archive/pelican/" rel="alternate"/><link href="https://jasongrout.github.io/medium-archive/pelican/feeds/author-lindsey-heagy.atom.xml" rel="self"/><id>https://jasongrout.github.io/medium-archive/pelican/</id><updated>2020-09-11T00:32:00+00:00</updated><subtitle>The Project Jupyter blog: news, releases, and community stories, archived from blog.jupyter.org.</subtitle><entry><title>Jupyter meets the Earth: EarthCube Community Meeting</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2020/jupyter-meets-the-earth-earthcube-community-meeting/" rel="alternate"/><published>2020-08-17T19:42:00+00:00</published><updated>2020-09-11T00:32:00+00:00</updated><author><name>Lindsey Heagy</name></author><id>tag:jasongrout.github.io,2020-08-17:/medium-archive/pelican/posts/2020/jupyter-meets-the-earth-earthcube-community-meeting/</id><summary type="html">&lt;p&gt;Summary of the EarthCube community meeting on July 27, 2020&lt;/p&gt;
</summary><content type="html">&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2020/jupyter-meets-the-earth-earthcube-community-meeting/images/001-0_h1CnyqyQk9K9crjW.jpg" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;By: Lindsey Heagy, Fernando Pérez, Joe Hamman and the Jupyter meets the Earth team&lt;/em&gt; (cross-posted on the &lt;a href="https://medium.com/pangeo/jupyter-meets-the-earth-earthcube-community-meeting-ab32f5c91caf"&gt;Pangeo Blog&lt;/a&gt;)&lt;/p&gt;
&lt;p&gt;As a part of the &lt;a href="https://www.earthcube.org/EC2020"&gt;2020 EarthCube annual meeting&lt;/a&gt;, we held a &lt;a href="/posts/2019/jupyter-meets-the-earth/"&gt;&lt;em&gt;Jupyter meets the Earth&lt;/em&gt;&lt;/a&gt; community discussion session on July 27. The Jupyter meets the Earth project is an EarthCube funded effort that combines research use cases in geosciences with technical developments within the Jupyter and Pangeo ecosystems. In this model of equal partners, scientific questions help drive software infrastructure development, and new technologies expand the horizons of viable research. This online workshop was an opportunity to gather members of the community, welcome newcomers, provide updates on the Jupyter and Pangeo ecosystems, and have time for discussion.&lt;/p&gt;
&lt;p&gt;The goals for the meeting were to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Provide an overview of the Jupyter &amp;amp; Pangeo ecosystems for researchers from the EarthCube community.&lt;/li&gt;
&lt;li&gt;Outline avenues for getting involved.&lt;/li&gt;
&lt;li&gt;Gather input for what advancements would best serve your research needs.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Over 100 participants registered, and we had contributions from 9 speakers. The meeting was a mix of presentations and Q&amp;amp;A from the community. The full recording of the meeting is available on &lt;a href="https://youtu.be/Zj3Gm4LNfwo"&gt;youtube&lt;/a&gt;, and we encourage continued discussion on the &lt;a href="https://discourse.pangeo.io/t/jupyter-meets-the-earth-earthcube-meeting-july-27/689"&gt;associated discourse post&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="presentations-google-drive-folder"&gt;Presentations (&lt;a href="https://drive.google.com/drive/folders/1lyIJcqHKqhstrnQU5ZWEgSsTRbVjpkZn?usp=sharing"&gt;google drive folder&lt;/a&gt;)&lt;/h2&gt;
&lt;p&gt;Fernando Pérez (&lt;a href="https://docs.google.com/presentation/d/1oR-LmqSkUsFUZBWH4qz1TDnRzd2oWHxUXgYxqjUb3LY/edit?usp=sharing"&gt;slides&lt;/a&gt;) started off the meeting by introducing the &lt;em&gt;Jupyter meets the Earth&lt;/em&gt; project — an effort aimed at driving forward technological developments in the Jupyter and Pangeo ecosystems in partnership with researchers in the geosciences. The motivation is to advance research and the software that supports it by combining domain expertise with methods in data science, software &amp;amp; data engineering practices. He provided an overview of Project Jupyter, highlighting the interplay between software and content, services, standards, community and governance that is necessary for broad-impact scientific open source software projects. He presented the extensible &lt;a href="https://jupyterlab.readthedocs.io/en/stable/"&gt;JupyterLab&lt;/a&gt; platform, that can be adapted to domain-specific needs as illustrated by the &lt;a href="http://www.bionet.ee.columbia.edu/research/ffbo/fbl"&gt;FlyBrainLab&lt;/a&gt; and &lt;a href="/posts/2020/jupyterlab-ros/"&gt;Cloud Robotics Command Station&lt;/a&gt; efforts. The &lt;em&gt;Jupyter meets the Earth&lt;/em&gt; team aims to similarly develop tools and extensions that will support interactive computing workflows in the geosciences.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Overview of Jupyter and Jupyter meets the Earth from Fernando Pérez" src="https://jasongrout.github.io/medium-archive/pelican/posts/2020/jupyter-meets-the-earth-earthcube-community-meeting/images/002-1_SmsFWNwWDW9-8jtDOa6_RQ.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Overview of Jupyter and Jupyter meets the Earth from Fernando Pérez&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Next up, Scott Henderson (&lt;a href="https://docs.google.com/presentation/d/1pdlUGRrX46kJYZHTkGA6HuOYHQ_Yt2NNxQlLpRFB9jo/edit?usp=sharing"&gt;slides&lt;/a&gt;) provided an overview of Pangeo and associated community events, including Hackweeks. “Pangeo is first and foremost a community promoting open, reproducible, and scalable science.” In terms of technology, this involves developing fully open source tools that can be deployed on shared computational infrastructure, such as HPC centers or the cloud, and hosting several forums to foster communication between scientists and software developers. Software is an important avenue for connection, but the overarching goals are a rallying point for a community. The critical mass of enthusiastic people has been key to the success of the Pangeo model.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Overview of Pangeo and Hackweeks from Scott Henderson" src="https://jasongrout.github.io/medium-archive/pelican/posts/2020/jupyter-meets-the-earth-earthcube-community-meeting/images/003-1_VOlTcY7mi7kIKT22Kc2_bw.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Overview of Pangeo and Hackweeks from Scott Henderson&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The Pangeo model is intended for use both on the cloud and on High Performance Computing (HPC) infrastructure. Kevin Paul gave the third talk on Pangeo on HPC (&lt;a href="https://docs.google.com/presentation/d/1eKDCK25jxjSFixwQX9GW56w84ZpaxRywEvMMYuX3Tg8/edit?usp=sharing"&gt;slides&lt;/a&gt;, &lt;a href="https://binder.pangeo.io/v2/gh/pangeo-data/pangeo-tutorial-agu-2018/master?filepath=notebooks%2Fgmet_ensemble.ipynb"&gt;notebook&lt;/a&gt;). HPC and cloud computing environments present technical differences in terms of usage patterns, file access, and resource allocations, however, the goals of Jupyter and Pangeo are similar in both cases — to enable interactive computing and simplify the user experience on both. Tools such as &lt;a href="https://dask.org/"&gt;dask&lt;/a&gt; and &lt;a href="https://kubernetes.dask.org/en/latest/"&gt;dask-kubernetes&lt;/a&gt;/&lt;a href="https://jobqueue.dask.org/en/latest/index.html"&gt;dask-jobqueue&lt;/a&gt; are targeted at enabling parallel computing on both infrastructures.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Kevin Paul giving us a demo of Pangeo on the Cheyenne supercomputer" src="https://jasongrout.github.io/medium-archive/pelican/posts/2020/jupyter-meets-the-earth-earthcube-community-meeting/images/004-1_ofmEc1YhrAqvCs3FXhuI2g.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Kevin Paul giving us a demo of Pangeo on the Cheyenne supercomputer&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;After a Q&amp;amp;A session that included questions on the computational cost of running Pangeo Infrastructure and efficient use of tools including Zarr, we moved on to a series of lightning talks.&lt;/p&gt;
&lt;h2 id="lightning-talks"&gt;Lightning Talks&lt;/h2&gt;
&lt;p&gt;Six speakers presented short lightning talks on aspects of the Jupyter and Pangeo ecosystems ranging from technologies to scientific applications to opportunities to engage with the Pangeo community.&lt;/p&gt;
&lt;p&gt;Anderson Banihirwe (&lt;a href="https://gist.github.com/andersy005/e08891883d91c01ab0ce963046d86343#file-intake-jupyter-meets-earth-ipynb"&gt;notebook&lt;/a&gt; and details in &lt;a href="https://github.com/earthcube2020/ec20_banihirwe_etal"&gt;intake-esm&lt;/a&gt;) kicked off the lightning talks by giving a demo and overview of Intake — a project to streamline loading and sharing of data. He showed a demo that included both Optimum Interpolation Sea Surface Temperature (OISST) data, as well as data from the Coupled Model Intercomparison Project (CMIP) running interactively on Cheyenne, the supercomputer at NCAR and using Dask for distributing the workload across nodes, with real-time diagnostics of the distributed computation provided by &lt;a href="https://github.com/dask/dask-labextension"&gt;Dask’s JupyterLab extension&lt;/a&gt;.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Demo from Anderson Banihirwe using intake to access OSSIT and CMIP data" src="https://jasongrout.github.io/medium-archive/pelican/posts/2020/jupyter-meets-the-earth-earthcube-community-meeting/images/005-0_cCp6QLc_60xWGsAN.jpg" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Demo from Anderson Banihirwe using intake to access OSSIT and CMIP data&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Next up, Scott Dale Peckham (&lt;a href="https://github.com/peckhams/balto_gui"&gt;notebook&lt;/a&gt;) gave a presentation that demonstrated the use of ipywidgets and ipyleaflet to create an interactive interface for fast access to geoscience data on servers that support the OpenDAP protocol. The project he champions is called &lt;a href="https://cires.colorado.edu/research/research-groups/project/balto-earthcube-brokered-alignment-long-tail-observations"&gt;BALTO, the Brokered Alignment of Long Tail Observations&lt;/a&gt; (also a famous Siberian Husky and sled dog). These graphical interface elements can be used in a programmatic workflow such as a Jupyter Notebook, but they conveniently encapsulate many details of accessing the data and resources provided by BALTO. This allows the scientists to focus on their research questions, without having to break their workflow to access data with external tools.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="BALTO GUI demo from Scott Peckham" src="https://jasongrout.github.io/medium-archive/pelican/posts/2020/jupyter-meets-the-earth-earthcube-community-meeting/images/006-1_Esr4Ec_RKwqx2la7G9eDFg.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;BALTO GUI demo from Scott Peckham&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;We then had a talk from Edom Moges (&lt;a href="https://docs.google.com/presentation/d/1QUdRZEI84jq9PoBEucnrdHkReghYrEg-jjGJZWzVvJA/edit?usp=sharing"&gt;slides&lt;/a&gt;) who presented work he is conducting with Laurel Larsen’s research group in hydrology as one use case in the Jupyter meets the Earth project. The presentation focused on a data synthesis work that aims to build a Jupyter based interactive platform that transforms raw hydrometeorological data to a gap-filled ready to use data for several intensively monitored watersheds across the US. The platform will be a basis for future community initiatives to benchmark data processing approaches, support comparative hydrological studies and comprehensive data-driven forecasts.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Lightning talk from Edom Moges and Laurel Larsen on the hydrology use-case in the Jupyter meets the Earth project" src="https://jasongrout.github.io/medium-archive/pelican/posts/2020/jupyter-meets-the-earth-earthcube-community-meeting/images/007-1_HPOqGuk9hRBV2WCl5QATIg.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Lightning talk from Edom Moges and Laurel Larsen on the hydrology use-case in the Jupyter meets the Earth project&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Georgiana Dolocan (&lt;a href="https://drive.google.com/file/d/1yu7gnRkHXkNhefVoBiho_Sl-C_7f0l85/view"&gt;video&lt;/a&gt;) impressed us next with an animated video accompanied with her narration to explain JupyterHub, its components (authenticator, spawner, proxy) as well as deployment options. The littlest JupyterHub (TLJH) is designed to make it simple to deploy multi-user Jupyter infrastructure on a single machine, and the more sophisticated Zero 2 JupyterHub Kubernetes (Z2JH) option is meant to scale to many users and large computational needs.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Animations from Georgiana Dolocan on JupyterHub" src="https://jasongrout.github.io/medium-archive/pelican/posts/2020/jupyter-meets-the-earth-earthcube-community-meeting/images/008-1_c1JTc_WrA5t1oUgnvvtcKA.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Animations from Georgiana Dolocan on JupyterHub&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Presenting from the perspective of an enthusiastic user, Erik Sundell (&lt;a href="https://docs.google.com/presentation/d/1TafZRXouz57SRBonHGt6bogwQp5_kVrTP1RkpkO0xB8/edit?usp=sharing"&gt;slides&lt;/a&gt;) gave us an overview of &lt;a href="http://jupyterbook.org"&gt;Jupyter Book&lt;/a&gt;: a tool to quickly create beautiful websites from notebooks and markdown. He walked through how to host them for free online in a time efficient way, and highlighted features including connections to Binder, which enable users to run content interactively.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Overview of JupyterBook from Erik Sundell" src="https://jasongrout.github.io/medium-archive/pelican/posts/2020/jupyter-meets-the-earth-earthcube-community-meeting/images/009-1_YfKDEmbHLgt-3eJlRqClaw.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Overview of JupyterBook from Erik Sundell&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Joe Hamman (&lt;a href="https://docs.google.com/presentation/d/1GKVLUsa971FHZkPDhdTSgxZQralWyy-292aIg1tFY64/edit?usp=sharing"&gt;slides&lt;/a&gt;) finished off our lightning talk session by outlining avenues for connecting with the Pangeo community. These include day-to-day communication on GitHub, Gitter, discourse and twitter, as well as more recent coffee-breaks. Depending on your topic of interest, there are also working groups that you can join on topics including data, machine learning, education, cloud computing, or you can suggest your own!&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Connecting with the Pangeo community — an overview from Joe Hamman" src="https://jasongrout.github.io/medium-archive/pelican/posts/2020/jupyter-meets-the-earth-earthcube-community-meeting/images/010-1_OOkrUZhvWq5ywRPlaQ8QDA.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Connecting with the Pangeo community — an overview from Joe Hamman&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="follow-up-and-further-discussion"&gt;Follow up and further discussion&lt;/h2&gt;
&lt;p&gt;To continue the discussion afterwards, we posed (&lt;a href="https://docs.google.com/presentation/d/1UqRd34zeOa5cW3aXFsjjh3TprgEHDl1cf4n1_BZezls/edit?usp=sharing"&gt;slides&lt;/a&gt;) a few questions where we hope to learn from the community’s needs, such as:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;What does your interactive computing workflow look like today? What do you envision it will be in 5 years?&lt;/li&gt;
&lt;li&gt;How would you like to publish and share your computational research and where can improvements be made?&lt;/li&gt;
&lt;li&gt;How do you stay up to date with the evolving open-source ecosystem? How would you like to be keeping up-to-date?&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;We are looking for your input and ideas! Please add your thoughts to the &lt;a href="https://discourse.pangeo.io/t/jupyter-meets-the-earth-earthcube-meeting-july-27/689"&gt;discourse post&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="thanks"&gt;Thanks&lt;/h2&gt;
&lt;p&gt;Thank you to the participants, speakers, and especially Lynne Schreiber and Ouida Meier from the EarthCube office for all of their support and work (even with very last-minute requests!).&lt;/p&gt;
&lt;p&gt;This work is part of the &lt;em&gt;Jupyter meets the Earth&lt;/em&gt; project, supported by the NSF EarthCube program under awards &lt;a href="https://www.nsf.gov/awardsearch/showAward?AWD_ID=1928406"&gt;1928406&lt;/a&gt;, &lt;a href="https://www.nsf.gov/awardsearch/showAward?AWD_ID=1928374"&gt;1928374&lt;/a&gt;.&lt;/p&gt;
&lt;iframe src="https://www.youtube-nocookie.com/embed/Zj3Gm4LNfwo" title="Jupyter Meets the Earth - Community Forum" width="560" height="315" style="aspect-ratio: 560 / 315" loading="lazy" allow="accelerometer; clipboard-write; encrypted-media; gyroscope; picture-in-picture" referrerpolicy="strict-origin-when-cross-origin" allowfullscreen&gt;&lt;/iframe&gt;
</content><category term="geoscience"/><category term="open science"/><category term="science"/></entry><entry><title>Jupyter meets the Earth</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2019/jupyter-meets-the-earth/" rel="alternate"/><published>2019-09-09T17:48:00+00:00</published><updated>2020-08-07T17:34:00+00:00</updated><author><name>Lindsey Heagy</name></author><id>tag:jasongrout.github.io,2019-09-09:/medium-archive/pelican/posts/2019/jupyter-meets-the-earth/</id><summary type="html">&lt;p&gt;By Lindsey Heagy and Fernando Pérez&lt;/p&gt;
</summary><content type="html">&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/jupyter-meets-the-earth/images/001-1_s3i12gpdCYyM0srHQkPZnw.webp" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;By Lindsey Heagy and Fernando Pérez&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;We are thrilled to announce that the NSF is funding our EarthCube proposal &lt;em&gt;“Jupyter meets the Earth: Enabling discovery in geoscience through interactive computing at scale”&lt;/em&gt; (&lt;a href="https://doi.org/10.5281/zenodo.3369938"&gt;pdf&lt;/a&gt;). The team working on this project consists of Fernando Pérez [1, 2, 3], Joe Hamman [4], Laurel Larsen [5], Kevin Paul [6], Lindsey Heagy [1], Chris Holdgraf [1, 2] and Yuvi Panda [7]. Our project team includes members from the &lt;a href="https://jupyter.org"&gt;Jupyter&lt;/a&gt; and &lt;a href="http://pangeo.io/"&gt;Pangeo&lt;/a&gt; communities, with representation across the geosciences including climate modeling, water resource applications, and geophysics. Three active research projects, one in each domain, will motivate developments in the Jupyter and Pangeo ecosystems. Each of these research applications demonstrates aspects of a research workflow which requires scalable, interactive computational tools.&lt;/p&gt;
&lt;p&gt;In this project we intend to follow the patterns that have made Jupyter an effective and successful platform: we will drive the development of computational machinery by concrete use cases from our own experience and research needs, and then find the appropriate points for extension, abstraction, and generalization. We are motivated to advancing research of contemporary importance in geoscience, and are equally committed to producing work that leads to broad impact, general use infrastructure that benefits scientists, educators, industry, and the general community.&lt;/p&gt;
&lt;p&gt;The adoption of open languages such as Python and the coalescence of communities of practice around open-source tools, is visible in nearly every domain of science. This is a fundamental shift in how science is conducted and shared. In recent years, there have been several high-profile examples in which open tools from the Python and Jupyter ecosystems played an integral role in the research, from data analysis to the dissemination of results. These include the first image of a black hole from the &lt;a href="https://eventhorizontelescope.org/"&gt;Event Horizon Telescope team&lt;/a&gt; and the detection of gravitational waves by the &lt;a href="https://www.caltech.edu/about/news/gravitational-waves-detected-100-years-after-einstein-s-prediction-49777"&gt;LIGO collaboration&lt;/a&gt;. The utility of open-source software in projects like these and the success of open communities such as Pangeo, provide evidence of the force-multiplying impact of investing in an ecosystem of open, community-driven tools. Through this project, we will advance this open paradigm in geoscience research, while strengthening and improving the infrastructure that supports it. We made this argument when discussing the intended impacts of our proposal, and we are pleased that the NSF is investing in this vision.&lt;/p&gt;
&lt;h2 id="geoscience-use-cases"&gt;Geoscience use cases&lt;/h2&gt;
&lt;p&gt;Given our project’s aims and approach, participating actively in domain research is crucial to our success. The following descriptions are meant to offer a flavour of the research questions we are tackling, each led by a geoscientist in the team.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;CMIP6 climate data analysis (Hamman).&lt;/strong&gt; The &lt;a href="https://www.wcrp-climate.org/wgcm-cmip"&gt;World Climate Research Program’s Coupled Model Intercomparison Project&lt;/a&gt; is now in its sixth phase and is expected to provide the most comprehensive and robust projections of future climate predictions. When complete, the archive is expected to exceed 18 PB in size. In the coming years, this collection of climate model experiments will form the basis for fundamental research, climate adaptation studies, and policy initiatives. While the CMIP6 dataset is likely to hold new answers to many pressing climate questions, the sheer volume of data is likely to present significant challenges to researchers. Indeed, new tools for scalable data analysis, machine learning, and inference are required to make the most out of these data.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Large-Scale Hydrologic Modeling (Larsen).&lt;/strong&gt; Streamflow forecasts are a valuable tool for flood mitigation and water management. Creating these forecasts requires that a variety of data types be brought together including model-generated streamflow estimates, sensor-based observations of water discharge, and hydrometeorological forcing factors, such as precipitation, temperature, relative humidity, and snow-water equivalent. The integration of simulated and observed data over disparate spatial and temporal scales creates new avenues for exploring data science techniques for effective water management.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Geophysical inversions (Heagy).&lt;/strong&gt; Geophysical inversions construct models of the subsurface by combining simulations of the governing physics with optimization techniques. These models are critical tools for locating and managing natural resources, such as groundwater, or for assessing the risk from natural hazards, such as volcanoes. Today, we need models applicable to increasingly complex scenarios, such as the socially delicate task of developing groundwater management policies in water-limited regions. This will require the development of new techniques for combining multiple geophysical data sets in a joint inversion, as well as the use of statistical and data science methods for including geologic and hydrologic data in the construction of 3D models.&lt;/p&gt;
&lt;p&gt;These scientific problems exhibit, each with its own flavour, similar technical challenges with respect to handling large volumes of data, performing expensive computations, and doing both of these as a part of the interactive, exploratory workflow that is necessary for scientific discovery.&lt;/p&gt;
&lt;h2 id="jupyter-pangeo-empowering-scientists"&gt;Jupyter &amp;amp; Pangeo: empowering scientists&lt;/h2&gt;
&lt;p&gt;Jupyter and Pangeo are both open communities that share the goal of developing tools and practices for interactive scientific workflows; these tools aim to deliver practical value to scientists who, in the course of everyday research, face a combination of big data and large-scale computing needs.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Project Jupyter&lt;/strong&gt; creates open-source tools and standards for interactive computing. These span the spectrum from low-level protocols for running code interactively up to &lt;a href="/posts/2018/jupyterlab-is-ready-for-users/"&gt;the web-based JupyterLab interface&lt;/a&gt; that a researcher uses. Jupyter is agnostic of programming language: over &lt;a href="https://github.com/jupyter/jupyter/wiki/Jupyter-kernels"&gt;130 different Jupyter kernels exist&lt;/a&gt;, and they provide support for most programming languages in widespread use today. Jupyter can be run on a laptop, in an HPC center, or in the cloud. Shared-infrastructure deployments (e.g. HPC / cloud) are enabled by JupyterHub, a component in the Jupyter toolbox that supports the deployment and management of Jupyter sessions for multiple users. The development process of tools in the Jupyter ecosystem is community-oriented and includes a diverse set of stakeholders across research, education, and industry. The project has a strong tradition of building tools that are first designed to solve specific problems, and then generalized to other users and applications.&lt;/p&gt;
&lt;p&gt;A &lt;strong&gt;Pangeo Platform&lt;/strong&gt; is a modular composition of open, community-driven projects, tailored to the scientific needs of a specific scientific domain. In its simplest form, it is based on the following generic components: a browser-based user interface (&lt;a href="https://jupyter.org"&gt;Jupyter&lt;/a&gt;), a data model and analytics toolkit (&lt;a href="http://xarray.pydata.org"&gt;Xarray&lt;/a&gt;), a parallel job distribution system (&lt;a href="https://dask.org/"&gt;Dask&lt;/a&gt;), a resource management system (either &lt;a href="https://kubernetes.dask.org/en/latest/"&gt;Kubernetes&lt;/a&gt; or a job queuing system such as &lt;a href="https://jobqueue.dask.org/"&gt;PBS&lt;/a&gt;), and a storage system (either cloud object store or traditional HPC file system). These are complemented by problem- and domain-specific libraries.&lt;/p&gt;
&lt;p&gt;This modular design allows for individual components to be readily exchanged and the system to be applied in new use cases. &lt;a href="https://medium.com/pangeo/announcing-pangeo-earthcube-award-fefbe54acbec"&gt;Pangeo was created by, and for, geoscientists&lt;/a&gt; faced with large-scale data and computation challenges, but such problems are now common in science. Researchers in a variety of disciplines including neuroscience and astrophysics are working to adapt the Pangeo design pattern for their communities. Beyond the initial Pangeo deployments supported by the NSF EarthCube grant for Pangeo, the platform has been adopted internationally, including by the &lt;a href="https://medium.com/pangeo/whats-so-cool-about-pangeo-974598f4bafc"&gt;UK Met office&lt;/a&gt;. It has also supported new research and education efforts from other federal agencies, such as NASA’s HackWeek focused on the analysis of &lt;a href="https://medium.com/pangeo/icesat-2-hackweek-mix-70-scientists-and-1-pangeo-jupyterhub-for-5-days-and-what-do-you-get-85f5267a4dfa"&gt;ICESat-2 data&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="user-centered-development"&gt;User-Centered Development&lt;/h2&gt;
&lt;p&gt;Pushing the boundaries of any toolset unveils areas for improvement and opportunities for developments that streamline and upgrade the user experience. We aim to take a holistic view of the scientific discovery process, from initial data acquisition through computational analysis to the dissemination of findings. Our development efforts will make technological improvements within the Jupyter and Pangeo ecosystems in order to reduce pain-points along the discovery lifecycle and to advance the infrastructure that serves scientists. Following established patterns in Jupyter’s development, we take a user-first, needs-driven approach and then generalize these ideas to work across related fields. Broadly, there are 4 areas along the research lifecycle where we will invest development efforts, which we discuss next.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Data discovery.&lt;/strong&gt; An early step in the research process is locating and acquiring data of interest. Data catalogs provide a way to expose datasets to the community in a way that is structured. Within the geosciences, there are a number of emerging community standards for data catalogs (e.g. THREDDS, STAC). To streamline access to such data sets, we plan to develop JupyterLab extensions which provide a user-interface that exposes these catalogs to researchers. This work will build upon the &lt;a href="https://github.com/jupyterlab/jupyterlab/issues/5548"&gt;JupyterLab Data Registry,&lt;/a&gt; which will provide a consistent set of standards for how data can be consumed and displayed by extensions in the Jupyter ecosystem, as well as &lt;a href="https://intake.readthedocs.io/en/latest/index.html"&gt;Intake&lt;/a&gt;, a lightweight library for finding, loading, and sharing data which is already serving the Pangeo community. Our communities are already &lt;a href="https://github.com/jupyterlab/jupyterlab-data-explorer/issues/51"&gt;discussing potential avenues for integration&lt;/a&gt; between these tools.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Scientific discovery through interactive computing.&lt;/strong&gt; The Jupyter Notebook has been adopted by many scientists because it supports an iterative, exploratory workflow combining code and narrative. Beyond code, text, and images, Jupyter supports the creation of Graphical User Interfaces (GUIs) with minimal programming effort on the part of the scientist. The &lt;a href="https://jupyter.org/widgets"&gt;Jupyter widgets framework&lt;/a&gt; lets scientists create a “Research GUI” that combines scientific code with interactive elements such as sliders, buttons and menus in just a single line of code, while still allowing for extensive customization and more complex interfaces when required. In this project, we will develop custom widgets tailored at the specific scientific needs of each of our driving use cases.&lt;/p&gt;
&lt;p&gt;Beyond their utility in the exploratory phase of research, interactive interfaces, or “dashboards” provide a mechanism for delivering custom scientific displays to collaborators, stakeholders, and students for whom the details of the code may not be pertinent. &lt;a href="/posts/2019/and-voila/"&gt;Voilà&lt;/a&gt; is a project, led by the &lt;a href="http://quantstack.net"&gt;QuantStack&lt;/a&gt; team, that enables dashboards to be generated from Jupyter notebooks. We plan to develop interactive dashboards using Voilà for our geoscience use-cases, contribute generic improvements to the Voilà codebase, and provide a demonstration of how researchers can deploy dashboards to share their research.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Research GUIs to explore Maxwell’s equations in research and education. Photo credit: SEOGI KANG" src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/jupyter-meets-the-earth/images/002-1_djLNdz13Z4kickGSQFCNUQ.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;&lt;em&gt;Research GUIs to explore Maxwell’s equations in research and education. Photo credit: SEOGI KANG&lt;/em&gt;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;&lt;strong&gt;Established tools and data visualization.&lt;/strong&gt; Many widely-used tools, particularly for visualization (e.g. Ncview, Paraview), are desktop-based applications and therefore cannot easily be used in cloud or HPC workflows. In some cases, modern, open-source alternatives are available. But often for specialized tasks, modern tools may not yet have functionality equivalent to the desktop version. JupyterHub can readily serve non-Jupyter web-native software applications such as RStudio, Shiny applications, and Stencila to users; under this project we aim to extend JupyterHub to also be able to serve desktop-native applications.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Using and managing shared computational infrastructure.&lt;/strong&gt; JupyterHub makes it possible to manage computing resources, user accounts, and provide access to computational environments online. Currently, JupyterHubs in the Pangeo project are deployed and maintained using the &lt;a href="http://z2jh.jupyter.org/en/latest/"&gt;Zero to JupyterHub guide&lt;/a&gt; alongside the &lt;a href="https://hubploy.readthedocs.io/en/latest/"&gt;HubPloy library&lt;/a&gt;. Together, these libraries have simplified the initial setup and automated upgrades to the Hubs. There are still many improvements that can make JupyterHub more suitable for larger, more complex deployments, both in terms of managing users and efficiently allocating resources. Under this project, we plan to build tools which collect metrics such as CPU and memory usage and expose these to both users and administrators so they can make more efficient use of the Hub. For shared deployments, we will improve user management so that user-groups can be used to manage permissions and allocations, and so that usage can be tracked and appropriately billed to the relevant grants. Within the HubPloy library, we also plan to make improvements to streamline continuous deployments so that installation and upgrade processes are repeatable and reliable.&lt;/p&gt;
&lt;h2 id="an-opportunity-for-meaningful-impact-join-us"&gt;An opportunity for meaningful impact — join us!&lt;/h2&gt;
&lt;p&gt;The impacts of climate change and the need for data-driven management of resources are some of the most critical and complex challenges facing society today. In recent years, we have experienced severe droughts in California and are in the midst of water management crises in the Central Valley, while devastating wildfires have destroyed entire communities. Through this partnership between researchers and Jupyter developers, we hope to contribute to the advancement of science-based solutions to these challenges, both by contributing directly to the research and by improving the open ecosystem of tools available to researchers and the stakeholders impacted by these issues.&lt;/p&gt;
&lt;p&gt;If developing open-source tech to advance research in geoscience and beyond excites you, then please &lt;a href="mailto:lheagy@berkeley.edu;fernando.perez@berkeley.edu"&gt;get in touch&lt;/a&gt;! At UC Berkeley, we will be hiring in 2 positions: a dev-ops position focussed on JupyterHub and shared infrastructure deployments, and a JupyterLab-oriented role focussed on extensions, dashboards, and interactivity. There will also be a position opening up at NCAR for a software engineer targeting improvements to improving the user experience of Xarray and Dask workflows.&lt;/p&gt;
&lt;p&gt;Even if you aren’t looking for a new job, there are other ways to get involved with both the Jupyter and Pangeo communities. We welcome new participants to the &lt;a href="https://pangeo.io/meeting-notes.html"&gt;weekly Pangeo meetings&lt;/a&gt; (on Wednesdays alternating between 4p GMT and 8p GMT) and there are &lt;a href="https://discourse.jupyter.org/t/jupyter-community-calls/668"&gt;monthly Jupyter community calls&lt;/a&gt;, which are open and meant to be accessible to a wide audience. Outside of calls, general Jupyter conversations happen on the &lt;a href="https://discourse.jupyter.org"&gt;Jupyter discourse&lt;/a&gt; and Pangeo conversations are typically on the &lt;a href="https://github.com/pangeo-data/pangeo"&gt;Pangeo GitHub&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="in-closing-a-step-toward-sustainable-open-science"&gt;In closing: a step toward sustainable open science&lt;/h2&gt;
&lt;p&gt;This project provides our team with $2 Million in funding over 3 years as a part of the NSF &lt;a href="https://earthcube.org/"&gt;EarthCube&lt;/a&gt; program. It also represents the first time federal funding is being allocated for the development of core Jupyter infrastructure.&lt;/p&gt;
&lt;p&gt;The open source ecosystem that Jupyter and Pangeo belong to has become part of the backbone that supports much of today’s computation in science, from astronomy and cosmology to microbiology and subatomic physics. Such broad usage represents a victory for this open and collaborative model of building scientific tools. Much of this success has come through the efforts of scientists and engineers who are committed to an open model of science, but who have had to work with little direct funding, minimal institutional support, and few viable career paths within science.&lt;/p&gt;
&lt;p&gt;There is real strategic risk to continuing with the implicit assumption that scientific open-source tools can be developed and maintained “for free.” If open, community-driven tools are to sustainably grow into the computational backbone of science, we need to recognize this role and support those who create them as regular members of the scientific community (we recently talked about this in more detail in a &lt;a href="http://www.tvworldwide.com/events/nsf/190815"&gt;talk at NSF headquarters&lt;/a&gt;). Projects like ours, where funding and resources are explicitly allocated toward this goal, are a step in the right direction. We hope that our experiences will contribute to ongoing conversations in the scientific community around these complex issues.&lt;/p&gt;
&lt;p&gt;In the past, we have tried to maintain a close relationship between domain problems and software development in Jupyter. However, this has typically been done in an ad-hoc manner, either by “hiding” the software development under the cover of science or by having funding to Jupyter alone. This is the first project where we explicitly partner with a team of domain scientists to simultaneously drive forward domain research and the development of Jupyter infrastructure.&lt;/p&gt;
&lt;p&gt;We are excited about this opportunity and hope to be able to demonstrate that investing in open tools can be a force-multiplier of resources. As always, our work will be done openly, transparently, and with constant community engagement. We look forward to your critiques, ideas, and contributions to make this effort as successful as possible.&lt;/p&gt;
&lt;h2 id="acknowledgments"&gt;Acknowledgments&lt;/h2&gt;
&lt;p&gt;Thanks to Joe Hamman, Chris Holdgraf, and Doug Oldenburg for constructive feedback and edits on this blog post.&lt;/p&gt;
&lt;p&gt;Many thanks to &lt;a href="https://www.ldeo.columbia.edu/user/rpa"&gt;Ryan Abernathey (Columbia)&lt;/a&gt;, &lt;a href="https://www.usgs.gov/staff-profiles/paul-a-bedrosian?qt-staff_profile_science_products=3#qt-staff_profile_science_products"&gt;Paul Bedrosian (USGS)&lt;/a&gt;, &lt;a href="https://quantstack.net/sylvain.html"&gt;Sylvain Corlay (QuantStack)&lt;/a&gt;, &lt;a href="https://www.usgs.gov/staff-profiles/richard-p-signell?qt-staff_profile_science_products=0#qt-staff_profile_science_products"&gt;Rich Signell (USGS)&lt;/a&gt;, and &lt;a href="https://www.nersc.gov/about/nersc-staff/data-analytics-services/rollin-thomas/"&gt;Rollin Thomas (NERSC)&lt;/a&gt;, who provided us with letters of support for this project; we look forward to working with you all! We are also grateful to &lt;a href="https://www.nsf.gov/staff/staff_bio.jsp?lan=sumishra"&gt;Shree Mishra&lt;/a&gt;, our NSF Program Director on this project, and to Dave Stuart for their support as we move forward with the project. This project is a part of the to the EarthCube program and we look forward to engaging with its working group.&lt;/p&gt;
&lt;p&gt;Finally, the sustained growth of Jupyter to the large-scale project that it has become would not have happened without the generous support of the Alfred P. Sloan, the Gordon and Betty Moore, and the Helmsley Foundations, as well as the leadership of Josh Greenberg and Chris Mentzel respectively at Sloan and Moore.&lt;/p&gt;
&lt;p&gt;This work is supported by the NSF EarthCube program under awards 1928406, 1928374&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;strong&gt;Affiliations&lt;/strong&gt;&lt;/p&gt;
&lt;p&gt;[1] UC Berkeley, Statistics Department&lt;/p&gt;
&lt;p&gt;[2] UC Berkeley, Berkeley Institute for Data Science&lt;/p&gt;
&lt;p&gt;[3] Lawrence Berkeley National Lab, Computational Research Division&lt;/p&gt;
&lt;p&gt;[4] National Center for Atmospheric Research, Climate and Global Dynamics Laboratory&lt;/p&gt;
&lt;p&gt;[5] UC Berkeley, Department of Geography&lt;/p&gt;
&lt;p&gt;[6] National Center for Atmospheric Research, Computational Information Systems Laboratory&lt;/p&gt;
&lt;p&gt;[7] UC Berkeley, Division of Data Sciences&lt;/p&gt;
</content><category term="geoscience"/><category term="open science"/><category term="science"/></entry></feed>