<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom"><title>Jupyter Blog - Kubernetes</title><link href="https://jasongrout.github.io/medium-archive/pelican/" rel="alternate"/><link href="https://jasongrout.github.io/medium-archive/pelican/feeds/tag-kubernetes.atom.xml" rel="self"/><id>https://jasongrout.github.io/medium-archive/pelican/</id><updated>2020-03-19T18:26:00+00:00</updated><subtitle>The Project Jupyter blog: news, releases, and community stories, archived from blog.jupyter.org.</subtitle><entry><title>The superheroes and the magic wand</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2020/the-superheros-and-the-magic-wand/" rel="alternate"/><published>2020-03-19T18:23:00+00:00</published><updated>2020-03-19T18:26:00+00:00</updated><author><name>Georgiana Dolocan</name></author><id>tag:jasongrout.github.io,2020-03-19:/medium-archive/pelican/posts/2020/the-superheros-and-the-magic-wand/</id><summary type="html">&lt;p&gt;In a place far, far away, on a planet called Jupyter, magic happens every day. This land is special because it’s full of magical tools…&lt;/p&gt;
</summary><content type="html">&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2020/the-superheros-and-the-magic-wand/images/001-1_v9JBClFvN4V-YZOQgwNxig.mp4" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;This is the first of a series of posts describing tools in the JupyterHub ecosystem, written by our wonderful &lt;em&gt;Contributor in Residence&lt;/em&gt;, Georgiana. For our first post, we’ll share some lore of JupyterHub, and tell you a story of how it all began…&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;In a place far, far away, on a planet called Jupyter, magic happens every day. This land is special because it’s full of magical tools with all kind of powers, devoted to one common purpose: to help people. Everyone sympathizing with this goal either advocates, uses, or cares for these tools. So, in the blink of an eye, these people with sometimes nothing else in common than the same drive to help others, gathered together and formed a community.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;This special group is known in the galaxy as “&lt;em&gt;&lt;strong&gt;&lt;/em&gt;The Jovyans&lt;/strong&gt;&lt;/em&gt;*” and new recruits join every day.* 🚀&lt;/p&gt;
&lt;/blockquote&gt;
&lt;figure&gt;
&lt;img alt="This image was created by Scriberia for The Turing Way community and is used under a CC-BY licence. Zenodo record." src="https://jasongrout.github.io/medium-archive/pelican/posts/2020/the-superheros-and-the-magic-wand/images/002-1_lz2yH1jyAlILFF9hZNtFPg.jpeg" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;This image was created by &lt;a href="http://www.scriberia.co.uk/"&gt;Scriberia&lt;/a&gt; for &lt;a href="https://github.com/alan-turing-institute/the-turing-way"&gt;&lt;strong&gt;The Turing Way&lt;/strong&gt;&lt;/a&gt; community and is used under a CC-BY licence. &lt;a href="https://zenodo.org/record/3695300"&gt;Zenodo record&lt;/a&gt;.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;But somebody needed to take care of these magical tools. So, the first Jovyans decided that from then on, they will become tool-keepers. Because their greatest responsibility was to teach new Jovyans how to use the magic, shortly after, they created a set of guiding laws. Some call this &lt;em&gt;“the&lt;/em&gt; &lt;em&gt;&lt;strong&gt;Documentation&lt;/strong&gt;&lt;/em&gt;”.&lt;/p&gt;
&lt;h2 id="the-magic-wand"&gt;The magic wand&lt;/h2&gt;
&lt;p&gt;One greatly cherished tool on planet Jupyter is the magical wand. This wand’s very special power is to help people work together as a team and find solutions to important problems.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;The wand is called “&lt;em&gt;&lt;strong&gt;&lt;/em&gt;JupyterHub&lt;/strong&gt;&lt;/em&gt;”.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2020/the-superheros-and-the-magic-wand/images/003-1_3nDwvUQiUfa54OsDPqKNag.mp4" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;p&gt;But the wand is of vast complexity and because of this, so are the guiding laws that control it.&lt;/p&gt;
&lt;p&gt;As a consequence, people trying to benefit from the magic of JupyterHub, spent a lot of time reading and understanding the instructions.&lt;/p&gt;
&lt;h2 id="the-superheros"&gt;The superheros&lt;/h2&gt;
&lt;p&gt;The tool-keepers noticed that most of the people coming to planet Jupyter to become Jovyans were from a big city in the cloud 🌤 called &lt;em&gt;Kubernetes&lt;/em&gt;. So they gathered together and debated what is the best way to help the people of Kubernetes.&lt;/p&gt;
&lt;p&gt;After 3 days and 3 nights of intense discussions (also lots of pizza breaks of course) they decided that one of them needed to get special training, rent a house in Kubernetes and teach the people there the wonders of the JupyterHub magic wand.&lt;/p&gt;
&lt;p&gt;The Chosen One gained the people’s trust, and it got better and better at anticipating and understanding their needs. So the locals started seeing it as a superhero and they even gave it a name, “&lt;strong&gt;Z2JH&lt;/strong&gt;”. 👓&lt;/p&gt;
&lt;p&gt;The news of these events started to spread far and wide, and people living in little towns outside of Kubernetes felt that they deserved the support of a superhero too. They also had great ideas and little time and needed to use the wand’s magic to do good.&lt;/p&gt;
&lt;p&gt;The tool-keepers were inspired by the magnificent accomplishments of these little towns, taking place even without a superhero around. So the littlest of them all, volunteered to go into superhero training to help the people of the little towns do good, faster. Though it was little, in no time, it grew to become the helper these small groups needed.&lt;/p&gt;
&lt;p&gt;The residents called it “&lt;strong&gt;The Littlest JupyterHub” or “TLJH”&lt;/strong&gt; because of its stature. But everybody knew that its stature didn’t reflect its enormous tenacity, speed and ability to do great things. This ended up inspiring people all over to believe that not all superheros wear capes, nor do they need to know how to fly in the cloud to be cool.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2020/the-superheros-and-the-magic-wand/images/004-1_egrzbBcwyaxm7-793O7uiQ.mp4" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Nowadays, people from all over join their forces and help the superheros learn new skills. Thanks to them, TLJH and Z2JH get stronger and better and have more special powers than they had when first created by the tool-keepers. Join them, become a Jovyan! ツ&lt;/p&gt;
&lt;/blockquote&gt;
</content><category term="JupyterHub"/><category term="Kubernetes"/></entry><entry><title>nbviewer has a new host: OVHcloud</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2020/nbviewer-has-a-new-host-ovhcloud/" rel="alternate"/><published>2020-03-06T19:46:00+00:00</published><updated>2020-03-06T19:46:00+00:00</updated><author><name>Min RK</name></author><id>tag:jasongrout.github.io,2020-03-06:/medium-archive/pelican/posts/2020/nbviewer-has-a-new-host-ovhcloud/</id><summary type="html">&lt;p&gt;nbviewer has moved from Rackspace to OVHcloud. That means moving from Docker to Kubernetes+Helm&lt;/p&gt;
</summary><content type="html">&lt;p&gt;For several years, &lt;a href="https://nbviewer.jupyter.org"&gt;nbviewer&lt;/a&gt; has been generously hosted by Rackspace. That sponsorship program appears to be ending, so nbviewer needed a new home; it has found one in &lt;a href="https://ovhcloud.com"&gt;OVHcloud&lt;/a&gt;. We are extremely grateful to OVHcloud for their support in keeping nbviewer running, building on their existing participation in the &lt;a href="/posts/2019/the-international-binder-federation/"&gt;Binder Federation&lt;/a&gt; (literally—nbviewer is now running on the same kubernetes cluster as ovh.mybinder.org).&lt;/p&gt;
&lt;h2 id="how-we-moved"&gt;How we moved&lt;/h2&gt;
&lt;h3 id="where-were-we-before"&gt;Where were we before?&lt;/h3&gt;
&lt;p&gt;nbviewer was previously deployed using a private repo (because it contained credentials) and various commands using &lt;a href="http://www.pyinvoke.org"&gt;invoke&lt;/a&gt;. It was a mixture of custom steps, using the openstack Python API and docker machine to:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;allocate two VMs&lt;/li&gt;
&lt;li&gt;build docker images&lt;/li&gt;
&lt;li&gt;deploy two nbviewer instances per node&lt;/li&gt;
&lt;li&gt;deploy memcached via nbcache on each node&lt;/li&gt;
&lt;li&gt;deploy statuspage publisher as a separate step&lt;/li&gt;
&lt;li&gt;update fastly to point to the running nbviewer instances&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The upside was that we had a repo that could immediately deploy nbviewer to anywhere with docker. With this repo, we moved our deployment strategy from a CoreOS cluster &lt;a href="https://github.com/jupyter/nbviewer.org-deploy/commit/2bf42cdfd54552d2669549bda002a8bf5f3df60d"&gt;to&lt;/a&gt; Rackspace’s short-lived &lt;a href="https://www.rackspace.com/newsroom/carina-by-rackspace-simplifies-containers-with-easy-to-use-instant-on-native-container-environment"&gt;Carina service&lt;/a&gt; to &lt;a href="https://github.com/jupyter/nbviewer.org-deploy/commit/22f54fc971941764ec3e14a68174823397a56163"&gt;deploying VMs ourselves&lt;/a&gt; with docker-machine. Migrating to a new source of VMs would not have been hard, but it wouldn’t have solved any of our challenges.&lt;/p&gt;
&lt;p&gt;Known downsides of this deployment that we’ve experienced over the years:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;only a few folks ever knew how to use it and could thus deploy updates to nbviewer&lt;/li&gt;
&lt;li&gt;it was private, contributing to above. It’s hard to onboard folks in an open community to a private repo!&lt;/li&gt;
&lt;li&gt;independent machines meant cache was not shared (minor, but contributes to our consumption of the GitHub API rate limit)&lt;/li&gt;
&lt;li&gt;no automatic recovery based on health monitoring, so when a container had issues, some humans got automated emails but no action was automatically taken. Extra frustrating because the fix was ~always to restart the container.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The credentials used by this repo have been revoked and an archived version of the repo is &lt;a href="https://github.com/jupyter/nbviewer.org-deploy/tree/archive-rackspace"&gt;available&lt;/a&gt; (with credentials redacted from history) in the new, &lt;a href="https://github.com/jupyter/nbviewer.org-deploy"&gt;public nbviewer.org-deploy repo&lt;/a&gt;.&lt;/p&gt;
&lt;h3 id="what-have-we-learned"&gt;What have we learned?&lt;/h3&gt;
&lt;p&gt;We’ve learned a lot about open, automatic, and sustainable deployments, and can now comfortably address all of the downsides above. Most of this has been learned from the communities participating in the JupyterHub and Binder projects, as seen in the &lt;a href="https://github.com/jupyterhub/mybinder.org-deploy"&gt;mybinder.org-deploy&lt;/a&gt; repo. mybinder.org-deploy is a public repo that automatically deploys and tests updates to at least four different Kubernetes clusters at the push of a button (the Big Green Merge Button, to be precise). Some things we have learned in the years since we set up our nbviewer deployment:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;kubernetes and helm are great :)&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/AGWA/git-crypt"&gt;git-crypt&lt;/a&gt; allows us to have public repos with some secret contents, so deployment repos don’t need to be fully private just to protect a couple api keys.&lt;/li&gt;
&lt;li&gt;Continuous Deployment via services like Travis or Circle lowers the bar to adding maintainers on a given deployment since all that’s needed is to press the Big Green Button.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;The OVH sponsorship came in the form of a slice of a kubernetes cluster; that meant we had to migrate the nbviewer deployment tools from using docker-machine to kubernetes, which for us means helm.&lt;/p&gt;
&lt;h3 id="step-1-helm-chart-for-nbviewer"&gt;Step 1: helm chart for nbviewer&lt;/h3&gt;
&lt;p&gt;The first step was to create a helm chart for nbviewer, which is done &lt;a href="https://github.com/jupyter/nbviewer/pull/905"&gt;here&lt;/a&gt;. Before, nbviewer was two docker containers running nbviewer and one running memcached per server. To turn this into a helm chart we need:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;dependency on &lt;a href="https://github.com/helm/charts/tree/master/stable/memcached"&gt;memcached helm chart&lt;/a&gt; for easy deployment of the cache. All the nbviewer instances will talk to this memcache, which should improve our GitHub rate limit consumption, since the cache will be shared. This means the deprecation of our own &lt;a href="https://github.com/jupyter/nbcache"&gt;nbcache repo&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;Deployment&lt;/code&gt; for nbviewer. This is (one of) the kubernetes wrappers around containers. It has a nice ‘replicas’ field to easily scale nbviewer up and down. We are currently running with 3 replicas.&lt;/li&gt;
&lt;li&gt;&lt;code&gt;Service&lt;/code&gt; to expose nbviewer to the Internet (likely also need an ingress in the future)&lt;/li&gt;
&lt;li&gt;&lt;code&gt;Deployment&lt;/code&gt; for the statuspage publisher, which updates &lt;a href="https://status.jupyter.org"&gt;https://status.jupyter.org&lt;/a&gt; with the remaining github rate limit available&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;We also needed helm charts and docker images for some other components, such as cdn.jupyter.org and the nbviewer statuspage data source.&lt;/p&gt;
&lt;h3 id="step-1b-helm-chart-for-cdnjupyterorg"&gt;Step 1b: helm chart for cdn.jupyter.org&lt;/h3&gt;
&lt;p&gt;nbviewer and some other services interact with cdn.jupyter.org, which is a lightweight nginx configuration to serve the classic notebook’s javascript and css as static files(this is no longer needed for npm-based jupyterlab, which can use &lt;a href="https://unpkg.com"&gt;unpkg&lt;/a&gt; as a CDN). This used to run on a multipurpose server VM, but that is also being retired for the same reason. The result was creating a docker image and helm chart for serving the contents of cdn.jupyter.org. The scripts used to run the existing CDN were already in &lt;a href="http://github.com/jupyter/cdn.jupyter.org/"&gt;a repo&lt;/a&gt;, so it was a small amount of work to adapt this to run in a container instead of on a server. We may retire cdn.jupyter.org in the future, so please don’t rely on it :)&lt;/p&gt;
&lt;p&gt;This is where we are right now — nbviewer.jupyter.org and cdn.jupyter.org are being served by OVH and the Rackspace machines are being retired.&lt;/p&gt;
&lt;h3 id="step-2-automatic-deployment"&gt;Step 2: automatic deployment&lt;/h3&gt;
&lt;p&gt;It would have been ideal for this to be step 1, but that’s now how it happened. Sometimes you need to get it done quickly before you get it done right. The task here will be to make a new, public nbviewer-deploy repo following the patterns we have learned in mybinder.org-deploy. That will mean:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;create public nbviewer.org-deploy repo with configuration and git-crypt encrypted secrets for deploying nbviewer on the ovh cluster (&lt;a href="https://github.com/jupyter/nbviewer.org-deploy"&gt;done&lt;/a&gt;)&lt;/li&gt;
&lt;li&gt;adopting &lt;a href="https://github.com/jupyterhub/chartpress"&gt;chartpress&lt;/a&gt; to publish our helm charts for nbviewer and version-tagged images from the nbviewer repo&lt;/li&gt;
&lt;li&gt;configure Travis-CI or other CI service to automatically deploy updates with helm, so merging a PR is all we need to do to deploy updates to nbviewer&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Thanks to OVHCloud for their continued support of Jupyter and Binder, and the JupyterHub and Binder communities for teaching me how to operate services in the open with kubernetes and helm.&lt;/p&gt;
</content><category term="Kubernetes"/><category term="nbviewer"/></entry><entry><title>A 2019 retrospective from the Binder Project</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2020/a-2019-retrospective-from-the-binder-project/" rel="alternate"/><published>2020-01-15T17:48:00+00:00</published><updated>2020-01-15T18:10:00+00:00</updated><author><name>Chris Holdgraf</name></author><id>tag:jasongrout.github.io,2020-01-15:/medium-archive/pelican/posts/2020/a-2019-retrospective-from-the-binder-project/</id><summary type="html">&lt;p&gt;2019 was a busy year for the Binder and JupyterHub projects — each saw growth in both their community and technology. Now that the year…&lt;/p&gt;
</summary><content type="html">&lt;p&gt;2019 was a busy year for the Binder and JupyterHub projects — each saw growth in both their community and technology. Now that the year has wrapped up, it is a good time to reflect on some of the highlights from the year.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2020/a-2019-retrospective-from-the-binder-project/images/001-1_ZeTyAGGLPlXs0UlQfKsr6g.webp" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;p&gt;Overall, 2019 was about improving the robustness, stability, and team dynamics around the JupyterHub and Binder projects, as well as connecting these projects with other tools and services in the open source community. Here are a few things that we are most-excited about.&lt;/p&gt;
&lt;h2 id="more-people-are-using-mybinderorg"&gt;More people are using mybinder.org&lt;/h2&gt;
&lt;p&gt;The Binder Federation is a collection of BinderHubs accessible from mybinder.org. This deployment is run as a public service and a demonstration of BinderHub, the underlying technology of the Binder project. This deployment is run on a volunteer basis by Binder community members, and is supported through grants and donations in infrastructure from project stakeholders. In 2019, the user base of mybinder.org grew from around 70,000 users per week to around 100,000 users (a growth of nearly 40%). mybinder.org is being used for teaching classes, sharing reproducible analyses, creating interactive documentation and narratives, and much more. We’re astonished at the rapid growth of Binder-ready repositories, and we’re excited to see what the community creates next.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Weekly user sessions at mybinder.org. Here you can see a typical pattern of activity over the course of a year. There are dips in activity over the summer and winter months, reflecting reduced activity from academic institutions." src="https://jasongrout.github.io/medium-archive/pelican/posts/2020/a-2019-retrospective-from-the-binder-project/images/002-0_0_V6rIgF_oAW5VS-.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Weekly user sessions at mybinder.org. Here you can see a typical pattern of activity over the course of a year. There are dips in activity over the summer and winter months, reflecting reduced activity from academic institutions.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="jupyterhub-reaches-10"&gt;JupyterHub reaches 1.0&lt;/h2&gt;
&lt;p&gt;JupyterHub, the underlying technology that provides interactive computing sessions to multiple users, &lt;a href="/posts/2019/announcing-jupyterhub-1-0/"&gt;is now at 1.0 status&lt;/a&gt;. The JupyterHub Python application was written several years ago, and reaching 1.0 reflects the work of dozens of open source contributors over time. JupyterHub is now a robust and stable application, having been used at smaller scales (think 5–10 people running on a single VM) as well as much larger scales (think 5,000 students running Jupyter sessions for a class).&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="The JupyterHub logo" src="https://jasongrout.github.io/medium-archive/pelican/posts/2020/a-2019-retrospective-from-the-binder-project/images/003-1_m2PN2PR-a6X_J3M602tgog.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;The JupyterHub logo&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="binderhub-is-out-of-beta"&gt;BinderHub is out of beta&lt;/h2&gt;
&lt;p&gt;BinderHub, the kubernetes-based technology that powers mybinder.org, also came out of Beta this year. This reflects the fact that BinderHub is a battle-hardened application that can provide a stable service over time. mybinder.org runs nearly 100,000 sessions a week, and requires minimal maintenance time from the Binder’s projects team of volunteer operators.&lt;/p&gt;
&lt;h2 id="the-binderhub-federation-is-launched"&gt;The BinderHub federation is launched&lt;/h2&gt;
&lt;p&gt;The Binder Project envisions a world in which technology can be used in vendor-agnostic and decentralized ways. BinderHub runs on Kubernetes, which can be deployed on a variety of cloud and local infrastructure. While the Binder team runs one BinderHub deployment at mybinder.org, our goal has always been to see &lt;em&gt;other&lt;/em&gt; organizations running their own BinderHubs. This year, we went one step beyond this by &lt;a href="/posts/2019/the-international-binder-federation/"&gt;launching the &lt;strong&gt;BinderHub Federation&lt;/strong&gt;&lt;/a&gt;. This is a collection of research and technology organizations that combine their expertise and computational resources to power mybinder.org. When users visit mybinder.org, they are now directed to one of several BinderHub instances. This makes mybinder.org more robust, and grows the number of organizations that utilize the project’s technology for their communities. A BIG THANKS goes out to Google, OVH, GESIS, and the Turing Institute for supporting the Binder Federation.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="The (rough) location of each BinderHub deployment in the mybinder.org federation" src="https://jasongrout.github.io/medium-archive/pelican/posts/2020/a-2019-retrospective-from-the-binder-project/images/004-1_KU35naJhl1LSDxKY8gog4g.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;The (rough) location of each BinderHub deployment in the mybinder.org federation&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="binder-connects-with-open-science-services"&gt;Binder connects with open science services&lt;/h2&gt;
&lt;p&gt;Another goal of Binder is to be a &lt;em&gt;part&lt;/em&gt; of the solution to more transparent, sharable, reproducible computational work. This means &lt;a href="/posts/2019/binder-with-zenodo/"&gt;plugging in to other ecosystems and projects&lt;/a&gt; in order to leverage the broader open science community. This year we saw a number of new connections with other services. BinderHub now supports links that point directly to Zenodo and Dataverse repositores, and we are working on a few other integrations in the coming months. This means that projects utilizing these resources will be able to share reproducible and interactive links to their work out-of-the-box.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Binder now works with Zenodo repositories!" src="https://jasongrout.github.io/medium-archive/pelican/posts/2020/a-2019-retrospective-from-the-binder-project/images/005-1_WVSV4_bwWlGWnT0k8xrlWg.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Binder now works with Zenodo repositories!&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="the-binder-community-grows"&gt;The Binder community grows&lt;/h2&gt;
&lt;p&gt;The Binder project’s most important asset is its people — this is a collection of volunteers spread across the world and from a variety of organizations. Binder community members do a variety of things — from working on technology, to teaching others how to make their work more reproducible, to participating in community discussions, to maintaining and debugging Binder tech. There is also a “core team” of Binder members that dedicates a significant part of their time to supporting the project. In 2019, we saw several new members join the core team, as well as a general growth in the Binder community. Welcome to all of our new team members! You can find a &lt;a href="https://jupyterhub-team-compass.readthedocs.io/en/latest/team.html"&gt;list of our current team members here&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="we-welcome-our-first-contributor-in-residence"&gt;We welcome our first Contributor in Residence&lt;/h2&gt;
&lt;p&gt;One challenge with running large, open projects is that resources tend to be scarce. The Binder Project has no formal project funding, and must find ways to both grow its technology as well as run mybinder.org on resources that are donated from its community. One thing that often suffers as a result is the maintenance and general improvement of our open source technology. This work is often under-appreciated, difficult, and unlikely to happen with purely volunteer labor.&lt;/p&gt;
&lt;p&gt;For this reason, the Binder project decided to &lt;a href="/posts/2019/the-jupyterhub-and-binder-contributor-in-residence/"&gt;apply for the CZI Essential Open Source grant series&lt;/a&gt;. We proposed the creation of the “Binder Contributor in Residence” position — an annual contractor position that pays a member of the Binder community to do many of the daily things that are crucial for the project’s growth. We are excited to have &lt;a href="https://github.com/GeorgianaElena"&gt;Georgiana Dolocan&lt;/a&gt; as our first contributor in residence, and look forward to where this project will go in 2020.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Many thanks to CZI for their support of the Binder and JupyterHub projects in 2020!" src="https://jasongrout.github.io/medium-archive/pelican/posts/2020/a-2019-retrospective-from-the-binder-project/images/006-0_pzbKC79Svh5Xrtev.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Many thanks to CZI for their support of the Binder and JupyterHub projects in 2020!&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="there-are-more-binderhub-deployments"&gt;There are more BinderHub deployments&lt;/h2&gt;
&lt;p&gt;The Binder Project aims to create technology that is deployable anywhere so that other organizations can support their communities with Binder infrastructure. In 2019 we ran a number of training sessions for how other groups can run their own BinderHub. In particular, the Turing Institute ran several workshops that had attendees up-and-running with their own functioning BinderHubs.&lt;/p&gt;
&lt;h2 id="thanks-to-our-community"&gt;Thanks to our community&lt;/h2&gt;
&lt;p&gt;As you can see, 2019 was a busy and exciting year for the Binder community. As a final note, we want to say thanks to all of you who have supported Binder in one form or another over the years. Binder is a project run by the community, for the community. It wouldn’t be possible without all of your hard work and friendly faces, thanks! We look forward to what’s coming next in 2020!&lt;/p&gt;
</content><category term="Binder"/><category term="Kubernetes"/><category term="reproducibility"/></entry><entry><title>The International Binder Federation</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2019/the-international-binder-federation/" rel="alternate"/><published>2019-07-01T09:11:00+00:00</published><updated>2019-07-01T09:11:00+00:00</updated><author><name>Chris Holdgraf</name></author><id>tag:jasongrout.github.io,2019-07-01:/medium-archive/pelican/posts/2019/the-international-binder-federation/</id><summary type="html">&lt;p&gt;We are happy to announce that mybinder.org is now backed by two clusters hosted by two different cloud providers: Google Cloud and OVH.&lt;/p&gt;
</summary><content type="html">&lt;p&gt;About two years ago, the Binder project evolved into the community led project that it is today. The deployment at &lt;a href="http://mybinder.org/"&gt;mybinder.org&lt;/a&gt; was upgraded to use&lt;br&gt;
BinderHub, a scalable open-source web application that runs on Kubernetes and provides free, sharable, interactive computing environments to people&lt;br&gt;
all around the world.&lt;/p&gt;
&lt;p&gt;In the ensuing years, the Binder community has grown considerably.&lt;br&gt;
Now people use the public deployment at &lt;a href="http://mybinder.org/"&gt;mybinder.org&lt;/a&gt; around 100,000 times each week, and there are around &lt;a href="https://github.com/betatim/binderlyzer/blob/master/binder-launches.ipynb"&gt;8,000 unique repositories&lt;/a&gt; compatible with Binder. Binder links now work with multiple repository providers, such as GitHub, GitLab, BitBucket, and even Zenodo! This means that providing enough compute power for everyone is a challenge.&lt;/p&gt;
&lt;p&gt;This is why today we are happy to announce that mybinder.org is now backed by two clusters hosted by two different cloud providers: Google Cloud and &lt;a href="https://www.ovh.com"&gt;OVH&lt;/a&gt;. This is the beginning of the International Binder Federation.&lt;/p&gt;
&lt;p&gt;How did we get here? Alongside mybinder.org’s growth we’ve seen growth in another part of the Binder ecosystem: people deploying a BinderHub for their own groups. In the last year, we have seen BinderHubs deployed &lt;a href="http://pangeo.io/"&gt;for large-scale earth analytics with the Pangeo project&lt;/a&gt;, for &lt;a href="https://the-turing-way.netlify.com/introduction/introduction"&gt;communities of best-practices in open science with The Turing Way&lt;/a&gt;, and for &lt;a href="https://notebooks.gesis.org/binder/"&gt;specific domains such as the social sciences like GESIS&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;As more BinderHubs are deployed, we realized that there was an opportunity to leverage the strengths and resources of the community to improve the large, public BinderHub deployment at &lt;a href="http://mybinder.org/"&gt;mybinder.org&lt;/a&gt;. Instead of a single BinderHub run by the Binder team, we could build a &lt;strong&gt;network&lt;/strong&gt; of BinderHubs that shares the load and keeps &lt;code&gt;mybinder.org&lt;/code&gt; stable and quick.&lt;/p&gt;
&lt;h2 id="ovh-joins-the-mybinderorg-federation"&gt;OVH Joins the &lt;a href="http://mybinder.org/"&gt;mybinder.org&lt;/a&gt; Federation&lt;/h2&gt;
&lt;p&gt;Today, we are thrilled to announce that the Binder Project now has &lt;strong&gt;a world-wide federation of BinderHubs powering&lt;/strong&gt; &lt;a href="http://mybinder.org/"&gt;&lt;strong&gt;mybinder.org&lt;/strong&gt;&lt;/a&gt;. We’ve partnered with OVH, a cloud hosting company based in Europe that is supportive of open projects such as Jupyter and Binder.&lt;/p&gt;
&lt;p&gt;Through the partnership with OVH, all traffic to &lt;code&gt;mybinder.org&lt;/code&gt; will now be split between &lt;strong&gt;two&lt;/strong&gt; BinderHubs - one run by the Binder team, and another run by a team of open-source advocates at OVH. They have generously offered their resources and computing time to allow Binder to serve the scientific and educational communities.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="The mybinder.org federation currently has two BinderHubs, but this network can (and will!) grow as more groups offer to connect their own BinderHubs to the network." src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/the-international-binder-federation/images/001-1_D0lhgUpJeWhb6K7igl4DvQ.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;The mybinder.org federation currently has two BinderHubs, but this network can (and will!) grow as more groups offer to connect their own BinderHubs to the network.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="what-does-this-mean"&gt;What does this mean?&lt;/h2&gt;
&lt;p&gt;So what has changed for you, the user? Probably not much. The&lt;br&gt;
biggest difference you’ll notice is that landing at &lt;code&gt;mybinder.org&lt;/code&gt;&lt;br&gt;
will now &lt;em&gt;redirect&lt;/em&gt; you to one of two places:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;code&gt;gke.mybinder.org&lt;/code&gt; is the BinderHub hosted on the Google Cloud Platform&lt;/li&gt;
&lt;li&gt;&lt;code&gt;ovh.mybinder.org&lt;/code&gt; is the BinderHub hosted on the OVH platform in France.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Other than this, your experience should be the same. You might even notice an improvement in speed and load times because the BinderHub you’re using has a little bit less traffic on it 😀.&lt;/p&gt;
&lt;p&gt;As always, we are pushing this out as soon as we think it is useful. You can help us to scale this model and make it rock solid by reporting weird things you notice. Expect tweaks over the next few weeks based on feedback from users. Let us know about the things you like, dislike or have questions about at &lt;a href="https://discourse.jupyter.org/t/the-binder-federation/1286"&gt;https://discourse.jupyter.org/t/the-binder-federation/1286&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="why-is-this-a-big-deal"&gt;Why is this a big deal?&lt;/h2&gt;
&lt;p&gt;We think that having a federation of BinderHubs behind mybinder.org is pretty cool, for a few different reasons. First, it demonstrates the dedication that the Binder community has towards building tools that anybody can deploy — we certainly don’t want to be the only ones running a BinderHub, and having other BinderHubs behind mybinder.org is a great example of this. Second, a network of BinderHubs for mybinder.org means that we can eventually do some clever things like use geolocation to distribute users, which should improve the performance that you all experience. Third, having fewer users per BinderHub means that we can optimize the resources available to each hub — resulting in improvements in speed and stability. Finally, having a federation of BinderHubs makes the mybinder.org service significantly more robust, and less-dependent on a single team, deployment, platform, and funding source. We are now more confident than ever that mybinder.org will continue to be available as a stable, free, public service that keeps growing.&lt;/p&gt;
&lt;h2 id="whats-next"&gt;What’s next?&lt;/h2&gt;
&lt;p&gt;Now that we’ve reached N=2 BinderHubs powering &lt;code&gt;mybinder.org&lt;/code&gt;,&lt;br&gt;
it is straightforward for us to grow this network to N=3 and beyond. Over the summer we will continue our work on approaching universities, research councils, and cloud hosting companies.&lt;/p&gt;
&lt;p&gt;If you are interested in helping out: we’d love to see other community members come forward to offer their time or resources to provide more nodes in the international BinderHub federation. If you’re interested in doing so, please &lt;a href="https://github.com/jupyterhub/team-compass/issues"&gt;open an issue in the JupyterHub team compass repository&lt;/a&gt;&lt;br&gt;
to discuss the possibilities of working together!&lt;/p&gt;
&lt;p&gt;We’re excited about the ability for the Binder community continuing&lt;br&gt;
to grow, and to build more tools for distributed, community-run&lt;br&gt;
infrastructure for open science and education. This is all possible&lt;br&gt;
because of the hard work of many people in the community, so a big&lt;br&gt;
thank you to those who spend their time working on open tools in&lt;br&gt;
the Jupyter and Binder ecosystems. We’re excited to see what comes&lt;br&gt;
next!&lt;/p&gt;
</content><category term="Binder"/><category term="cloud computing"/><category term="Kubernetes"/></entry><entry><title>BinderHub is out of Beta!</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2019/binderhub-is-out-of-beta/" rel="alternate"/><published>2019-04-26T16:49:00+00:00</published><updated>2019-04-26T18:41:00+00:00</updated><author><name>Tim Head</name></author><id>tag:jasongrout.github.io,2019-04-26:/medium-archive/pelican/posts/2019/binderhub-is-out-of-beta/</id><summary type="html">&lt;p&gt;Today we are proud to announce that we now consider BinderHub a stable tool with a track record of operating in production.&lt;/p&gt;
</summary><content type="html">&lt;p&gt;Nearly two years ago, the Binder Project &lt;a href="/posts/2017/binder-2-0-a-tech-guide-2017/"&gt;released the beta version&lt;/a&gt;&lt;br&gt;
of &lt;a href="https://binderhub.readthedocs.io/"&gt;BinderHub&lt;/a&gt;, the technology behind &lt;a href="https://mybinder.org"&gt;mybinder.org&lt;/a&gt;. Since then, mybinder.org has grown to serve nearly 90,000 launches each week and hit the &lt;a href="/posts/2018/mybinder-org-serves-two-million-launches/"&gt;two million launches in a year milestone last year&lt;/a&gt;. Over these two years BinderHub has matured as technology and community. Several new organizations now run their own BinderHubs and have joined the community of maintainers.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="A year of Binder sessions served from mybinder.org. Darker colors show users in countries that have launched more Binder sessions. The pins show the approximate locations of public BinderHubs that we know of. Where is mybinder.org? It is a global effort so there is no one location associated with it or the team that runs it." src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/binderhub-is-out-of-beta/images/001-1_zGQcuph2mrh7ry_tH_QJYQ.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;A year of Binder sessions served from mybinder.org. Darker colors show users in countries that have launched more Binder sessions. The pins show the approximate locations of public BinderHubs that we know of. Where is mybinder.org? It is a global effort so there is no one location associated with it or the team that runs it.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Today we are proud to announce that we’ve removed the “beta” label from all pages served by BinderHub. BinderHub is now a tool with a track record of working well in production (both at mybinder.org and in other deployments), its general features and API have stabilized, and &lt;a href="https://jupyterhub.readthedocs.io/en/stable/"&gt;JupyterHub&lt;/a&gt;, the underlying technology that BinderHub uses, is close to its 1.0 release.&lt;/p&gt;
&lt;p&gt;A big thank you to all those who have used, commented, advocated, advertised, contributed, maintained and funded this journey. The deployment at mybinder.org is funded with grants from the &lt;a href="https://www.moore.org/grant-detail?grantId=GBMF6865"&gt;Moore Foundation&lt;/a&gt; and the Google Cloud Platform.&lt;/p&gt;
&lt;h3 id="who-else-has-deployed-a-binderhub"&gt;Who else has deployed a BinderHub?&lt;/h3&gt;
&lt;p&gt;A goal of Project Binder is to build modular, open-source tools that&lt;br&gt;
others can deploy for their own communities. BinderHub, the core technology behind mybinder.org, runs on Kubernetes this means it can be deployed on many cloud providers or even on your own hardware. Over the years, we have seen many new organizations deploy their own BinderHubs. Here are our highlights of other organizations who have joined the BinderHub community.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;GESIS&lt;/strong&gt; were the first to deploy a public BinderHub that is operated independently of mybinder.org. The &lt;a href="https://notebooks.gesis.org/binder/"&gt;GESIS BinderHub instance&lt;/a&gt; went live&lt;br&gt;
in December 2017 and has been running ever since. They are frequent&lt;br&gt;
contributors to the upstream project. Their BinderHub runs on a bare metal&lt;br&gt;
Kubernetes cluster.&lt;/p&gt;
&lt;p&gt;The &lt;strong&gt;Pangeo Project&lt;/strong&gt; instance was the next to come online, launching &lt;a href="https://binder.pangeo.org"&gt;their&lt;br&gt;
public BinderHub instance&lt;/a&gt; in September 2018 (more details in their &lt;a href="https://medium.com/pangeo/pangeo-meets-binder-2ea923feb34f"&gt;blog post&lt;/a&gt;). They provide additional compute resources and have customized their setup to provide on-demand &lt;a href="https://dask.org/"&gt;dask&lt;/a&gt; clusters to users. Try out &lt;a href="http://binder.pangeo.io/v2/gh/pangeo-data/pangeo_ocean_examples/master"&gt;one&lt;br&gt;
of their examples using ocean data&lt;/a&gt;. The Pangeo cluster is hosted on Google Kubernetes Engine.&lt;/p&gt;
&lt;p&gt;Recently, the &lt;a href="https://www.turing.ac.uk/research/research-projects/turing-way-handbook-reproducible-data-science"&gt;&lt;strong&gt;Turing Way project&lt;/strong&gt;&lt;/a&gt; has been working on deploying a BinderHub at the &lt;a href="https://www.turing.ac.uk/"&gt;Turing Institute&lt;/a&gt; in the UK. This BinderHub will serve both internal and external users with the goal of making it easier to share data science projects. They have run several workshops including one that teaches scientists and research software engineers how to deploy their own BinderHub instance. Recently they led a workshop in which ten academics and IT staff deployed their own BinderHub on the Microsoft Azure cloud! Sarah Gibson from their team has recently joined the team that operates mybinder.org.&lt;/p&gt;
&lt;p&gt;Finally, the &lt;a href="https://www.pims.math.ca/"&gt;&lt;strong&gt;Pacific Institute for the Mathematical Sciences&lt;/strong&gt;&lt;/a&gt; (PIMS) runs a service &lt;a href="https://intro.syzygy.ca/"&gt;called Syzygy&lt;/a&gt;. It deploys several JupyterHubs and BinderHubs for scientific organizations around Canada. A Binder team recently held a tutorial on deploying Binder and JupyterHub at the PEARC (&lt;a href="https://docs.google.com/presentation/d/1ELYepgptS7LcpBwjrotPRcQENr_WlxX0eQaBvzZonxs/edit?usp=sharing"&gt;you can find slides for the talk here&lt;/a&gt;). They have also &lt;a href="https://github.com/etiennedub/terraform-binderhub"&gt;provided a script for deploying JupyterHub&lt;/a&gt; (and Binder) on Terraform.&lt;/p&gt;
&lt;h3 id="whats-next"&gt;What’s next?&lt;/h3&gt;
&lt;p&gt;Now that BinderHub is not in beta anymore, what is next? As Project Binder we will focus on adoption, training, and growing the Binder community. This means growing the creation of &lt;strong&gt;Binder-ready repositories&lt;/strong&gt; (repositories that have the necessary structure for Binder to create the environment needed to run the repository’s code). We will increase our outreach and marketing efforts to make sure a diverse audience everywhere around the world knows about Binder.&lt;/p&gt;
&lt;p&gt;We will also work on making it easier to setup and operate a public BinderHub no matter what cloud vendor you are using. We are excited to see more organizations deploy their own BinderHubs for their communities. In the coming months, our goal is to create a &lt;strong&gt;federation of public BinderHubs&lt;/strong&gt;&lt;br&gt;
that operate in unison to serve the global user base of mybinder.org.&lt;/p&gt;
&lt;p&gt;If this caught your attention consider joining the Binder community, or contributing to Project Binder! A good place to start is the &lt;a href="https://discourse.jupyter.org"&gt;Jupyter Community Forum&lt;/a&gt; or dive straight into &lt;a href="http://github.com/jupyterhub/"&gt;the code on GitHub&lt;/a&gt;.&lt;/p&gt;
</content><category term="Binder"/><category term="Kubernetes"/><category term="releases"/></entry><entry><title>Introducing TraefikProxy — a scalable and highly available proxy for JupyterHub</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2019/introducing-traefikproxy-a-new-jupyterhub-proxy-based/" rel="alternate"/><published>2019-03-06T11:58:00+00:00</published><updated>2019-03-06T11:58:00+00:00</updated><author><name>Georgiana Dolocan</name></author><id>tag:jasongrout.github.io,2019-03-06:/medium-archive/pelican/posts/2019/introducing-traefikproxy-a-new-jupyterhub-proxy-based/</id><summary type="html">&lt;p&gt;Removing the single point of failure from your JupyterHub infrastructure with traefik, etcd and Outreachy&lt;/p&gt;
</summary><content type="html">&lt;p&gt;In the JupyterHub context, the proxy is the unit in charge of directing the user requests to their notebook servers.&lt;/p&gt;
&lt;p&gt;The proxy manages a list of &lt;strong&gt;[user : notebook]&lt;/strong&gt; mappings (the proxy routing table) in order to decide which request is sent where. The routing table must be continuously updated as users start and stop their servers without disrupting the requests being processed. The following drawing illustrates the proxy functionality in a JupyterHub deployment.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/introducing-traefikproxy-a-new-jupyterhub-proxy-based/images/001-1_cy9fESaJhd0iDhH2v-LySQ.jpeg" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;Why the need for a new proxy?&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Currently, the &lt;strong&gt;default&lt;/strong&gt; proxy implementation for JupyterHub is &lt;a href="https://github.com/jupyterhub/configurable-http-proxy"&gt;&lt;em&gt;configurable-http-proxy&lt;/em&gt;&lt;/a&gt; &lt;em&gt;(CHP)&lt;/em&gt;, which is a single-process nodejs proxy, that stores the routing table in-memory. &lt;em&gt;CHP&lt;/em&gt; is easy to install and run, and thus in most of the cases it’s a fine option. However, because you can only run a single copy of the proxy at a time, it has its limitations when used in dynamic, large scale systems.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;What makes this new proxy special?&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;&lt;strong&gt;JupyterHub 0.8.&lt;/strong&gt; opened the way towards allowing users to &lt;a href="https://jupyterhub.readthedocs.io/en/stable/reference/proxy.html"&gt;create custom proxy implementations&lt;/a&gt; based on their deployment needs. &lt;a href="https://github.com/jupyterhub/traefik-proxy"&gt;&lt;em&gt;JupyterHub Traefik Proxy&lt;/em&gt;&lt;/a&gt; leverages this feature to offer an alternative to the default proxy. It is an implementation of the JupyterHub Proxy API based on &lt;a href="https://traefik.io"&gt;traefik&lt;/a&gt;, an extremely lightweight, portable reverse proxy implementation, that supports load balancing and can configure itself automatically and dynamically. JupyterHub Traefik Proxy comes in two flavors, depending on how traefik stores the routing table:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;TraefikTomlProxy&lt;/strong&gt; — &lt;em&gt;for&lt;/em&gt; smaller, &lt;em&gt;single-node deployments&lt;/em&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;TraefikEtcdProxy&lt;/strong&gt; — &lt;em&gt;for&lt;/em&gt; distributed &lt;em&gt;setups&lt;/em&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;How does it work?&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Both &lt;em&gt;TraefikTomlProxy&lt;/em&gt; and &lt;em&gt;TraefikEtcdProxy&lt;/em&gt; use a &lt;em&gt;toml&lt;/em&gt; file for the global configuration. This file contains information about how to set up the connections to the routing table provider (the unit that stores the routing rules like “/user/mary” should be sent to Mary’s server at http://10.0.1.5:12345) and to the network entry points into Traefik (listening port, SSL, traffic redirection). However, the two proxies go in different directions when it comes to the provider used for storing the routing table.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;TraefikTomlProxy&lt;/strong&gt;&lt;/em&gt; uses a &lt;em&gt;toml&lt;/em&gt; file to store the routes and keeps an in-memory copy of it for a faster access to the routes. This is appropriate for smaller-scale deployments. For example, &lt;a href="https://github.com/jupyterhub/the-littlest-jupyterhub"&gt;&lt;em&gt;the Littlest Jupyterhub&lt;/em&gt;&lt;/a&gt; (the single-node JupyterHub distribution, for a small number of users) just switched from using two proxies (traefik as an edge proxy with letsencrypt support and CHP for routing) to a configuration with just &lt;em&gt;TraefikTomlProxy&lt;/em&gt; that serves both requirements. ❤&lt;/p&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;TraefikEtcdProxy&lt;/strong&gt;&lt;/em&gt; uses &lt;a href="https://coreos.com/etcd/"&gt;&lt;em&gt;etcd&lt;/em&gt;&lt;/a&gt;, a distributed key-value store to persist the routing table. This implementation aims to benefit &lt;a href="http://z2jh.jupyter.org/en/stable/"&gt;&lt;em&gt;Zero to JupyterHub with Kubernetes&lt;/em&gt;&lt;/a&gt; because it allows having multiple proxy replicas, making the proxy highly available and thus improving the scalability and stability of the system.&lt;/p&gt;
&lt;p&gt;Adding the information about TraefikProxy to the first diagram , the drawing below presents the two proxy flavors with their &lt;strong&gt;common&lt;/strong&gt; parts (the global configuration file — &lt;em&gt;traefik.toml&lt;/em&gt;) and their &lt;strong&gt;differences&lt;/strong&gt; (the mechanism used for storing the routing table — a toml file vs. etcd)&lt;/p&gt;
&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/introducing-traefikproxy-a-new-jupyterhub-proxy-based/images/002-1_A--P2mM21Cj6gfWImInnww.webp" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;p&gt;Another cool Traefik feature is the &lt;em&gt;&lt;strong&gt;Web UI dashboard&lt;/strong&gt;&lt;/em&gt; which lists all of the registered frontends (the set of rules that determine how incoming requests are forwarded) and backends (the notebooks), the routing rules, some useful metrics, and other configuration elements. The port on which TraefikProxy’s api will run, as well as the username and password used for authenticating, are all configurable.&lt;/p&gt;
&lt;p&gt;Here’s what the dashboard looks like:&lt;/p&gt;
&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/introducing-traefikproxy-a-new-jupyterhub-proxy-based/images/003-1_DQjZmUX2a6nzPJzFdZJqTw.jpeg" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;em&gt;&lt;strong&gt;How to enable Traefik Proxy?&lt;/strong&gt;&lt;/em&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;These instructions help you enable one of the Traefik proxies on your JupyterHub.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Install it:&lt;/li&gt;
&lt;/ol&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;python3 -m pip install jupyterhub-traefik-proxy
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;ol start="2"&gt;
&lt;li&gt;Install traefik and etcd:&lt;/li&gt;
&lt;/ol&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;$&lt;span class="w"&gt; &lt;/span&gt;python3&lt;span class="w"&gt; &lt;/span&gt;-m&lt;span class="w"&gt; &lt;/span&gt;jupyterhub_traefik_proxy.install&lt;span class="w"&gt; &lt;/span&gt;--output&lt;span class="o"&gt;=&lt;/span&gt;/usr/local/bin
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This will install the default versions of traefik and etcd, namely &lt;code&gt;traefik-1.7.5&lt;/code&gt; and &lt;code&gt;etcd-3.3.10&lt;/code&gt; to &lt;code&gt;/usr/local/bin&lt;/code&gt; specified through the &lt;code&gt;--output&lt;/code&gt; option.&lt;/p&gt;
&lt;ol start="3"&gt;
&lt;li&gt;Configure JupyterHub to run with TraefikProxy through &lt;em&gt;jupyterhub_config.py,&lt;/em&gt; using the &lt;em&gt;proxy_class&lt;/em&gt; config option.&lt;/li&gt;
&lt;/ol&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;jupyterhub_traefik_proxy&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;TraefikEtcdlProxy&lt;/span&gt;
&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;JupyterHub&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;proxy_class&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;TraefikEtcdProxy&lt;/span&gt;
&lt;span class="c1"&gt;# will configure JupyterHub to run with TraefikEtcdProxy&lt;/span&gt;

&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;jupyterhub_traefik_proxy&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;TraefikTomlProxy&lt;/span&gt;
&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;JupyterHub&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;proxy_class&lt;/span&gt; &lt;span class="o"&gt;=&lt;/span&gt; &lt;span class="n"&gt;TraefikTomlProxy&lt;/span&gt;
&lt;span class="c1"&gt;# will configure JupyterHub to run with TraefikTomlProxy&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;As there are many other proxy configuration options, please check out the project &lt;a href="https://jupyterhub-traefik-proxy.readthedocs.io/en/latest/"&gt;documentation&lt;/a&gt; for more info and some example configurations.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;What’s next?&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;TraefikEtcdProxy is soon to be integrated into &lt;a href="https://github.com/jupyterhub/zero-to-jupyterhub-k8s"&gt;&lt;strong&gt;zero-to-jupyterhub-k8s&lt;/strong&gt;&lt;/a&gt;, the Helm Chart for deploying JupyterHub on Kubernetes. By using the Traefik proxy with etcd, we can eliminate the downtime while the proxy restarts or upgrades, as well as have an all-in-one proxy that can be used for both routing and &lt;em&gt;HTTPS (Let’s Encrypt)&lt;/em&gt; support.&lt;/p&gt;
&lt;p&gt;Also, there’s a performance analysis of all the three proxies on the way that will help us understand the advantages and limitations of this new JupyterHub Proxy implementation.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;strong&gt;The story behind&lt;/strong&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;JupyterHub Traefik Proxy was developed as part of my &lt;a href="https://www.outreachy.org/"&gt;Outreachy&lt;/a&gt; internship. I am very grateful to have had the chance to work with the Jupyter community and I’m happy and proud to say that I’ve learned a lot and the experience has been invaluable. There were a lot of questions and a ton of fears along the way, but the encouragements and guidance I received, helped me move forward and finish this great project. So, a big &lt;strong&gt;thank you to&lt;/strong&gt; everyone for their support and to &lt;a href="https://numfocus.org/"&gt;&lt;em&gt;NumFocus&lt;/em&gt;&lt;/a&gt; &lt;em&gt;and&lt;/em&gt; &lt;a href="https://bids.berkeley.edu/"&gt;&lt;em&gt;Berkeley Institute for Data Science&lt;/em&gt;&lt;/a&gt; &lt;em&gt;for sponsoring this Outreachy round&lt;/em&gt;! ❤&lt;/p&gt;
&lt;/blockquote&gt;
</content><category term="JupyterHub"/><category term="Kubernetes"/><category term="Outreachy"/></entry><entry><title>Zero to JupyterHub helm chart 0.8</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2019/zero-to-jupyterhub-helm-chart-0-8/" rel="alternate"/><published>2019-02-21T18:34:00+00:00</published><updated>2019-02-21T18:34:00+00:00</updated><author><name>Min RK</name></author><id>tag:jasongrout.github.io,2019-02-21:/medium-archive/pelican/posts/2019/zero-to-jupyterhub-helm-chart-0-8/</id><summary type="html">&lt;p&gt;We’ve just released version 0.8 of the jupyterhub helm chart.&lt;/p&gt;
</summary><content type="html">&lt;p&gt;We’ve just released version 0.8 of the jupyterhub helm chart.&lt;/p&gt;
&lt;p&gt;For those who may not know, &lt;a href="https://zero-to-jupyterhub.readthedocs.io"&gt;Zero to JupyterHub&lt;/a&gt; is a guide and &lt;a href="https://helm.sh"&gt;helm chart&lt;/a&gt; for deploying &lt;a href="https://jupyterhub.readthedocs.io"&gt;JupyterHub&lt;/a&gt; on &lt;a href="https://kubernetes.io"&gt;Kubernetes&lt;/a&gt;. Helm is like a package manager for Kubernetes, which aims to make it easy to install and manage applications on your cluster.&lt;/p&gt;
&lt;p&gt;There are loads of bugfixes and improvements to this release, and we encourage you to check out the &lt;a href="https://github.com/jupyterhub/zero-to-jupyterhub-k8s/blob/master/CHANGELOG.md"&gt;change log&lt;/a&gt;. Here, we’ll discuss two highlights from the new features: profiles and autoscaling.&lt;/p&gt;
&lt;h2 id="profiles"&gt;&lt;strong&gt;Profiles&lt;/strong&gt;&lt;/h2&gt;
&lt;p&gt;By adopting the profiles pattern established in &lt;a href="https://github.com/jupyterhub/wrapspawner"&gt;WrapSpawner&lt;/a&gt;, Kubernetes users can now pick from a selection of “profiles” or pre-set configurations, defined by the operator of a cluster. For example, defining small, medium, and large resource requests, or special profiles for GPUs.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/zero-to-jupyterhub-helm-chart-0-8/images/001-1_XRxmQ4Mmd48iE7ST-zdcUg.webp" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;p&gt;The above sample is created from the following snippet in your values.yaml:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="n"&gt;singleuser&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="n"&gt;profileList&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;display_name&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Small: default&amp;quot;&lt;/span&gt;
&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="n"&gt;A&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;small&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;job&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;CPU&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;no&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;GPU&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;This&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;is&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;the&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;default&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;
&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="k"&gt;default&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;True&lt;/span&gt;
&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="n"&gt;kubespawner_override&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="n"&gt;cpu_limit&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="n"&gt;cpu_guarantee&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;1&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="n"&gt;mem_limit&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;1G&amp;quot;&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="n"&gt;mem_guarantee&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;512M&amp;quot;&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="n"&gt;extra_resource_limits&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;{}&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;display_name&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;Big: 8 CPUs&amp;quot;&lt;/span&gt;
&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="n"&gt;A&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;big&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;job&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;CPUs&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;no&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;GPU&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;and&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;64&lt;/span&gt;&lt;span class="n"&gt;GB&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;of&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;RAM&lt;/span&gt;
&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="n"&gt;kubespawner_override&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="n"&gt;cpu_limit&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="n"&gt;cpu_guarantee&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;8&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="n"&gt;mem_limit&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;64G&amp;quot;&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="n"&gt;mem_guaranttee&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;64G&amp;quot;&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="n"&gt;extra_resource_limits&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;{}&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;display_name&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;GPU job&amp;quot;&lt;/span&gt;
&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="n"&gt;description&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="n"&gt;This&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;configuration&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;gives&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;you&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;CPUs&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;16&lt;/span&gt;&lt;span class="n"&gt;GB&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;of&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;RAM&lt;/span&gt;&lt;span class="o"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;and&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;a&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;GPU&lt;/span&gt;
&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="n"&gt;kubespawner_override&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="n"&gt;image&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;consideratio&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;singleuser&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;gpu&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="n"&gt;v0&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="mf"&gt;3.0&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="n"&gt;cpu_limit&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="n"&gt;cpu_guarantee&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;4&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="n"&gt;mem_limit&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;16G&amp;quot;&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="n"&gt;mem_guarantee&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;16G&amp;quot;&lt;/span&gt;
&lt;span class="w"&gt;        &lt;/span&gt;&lt;span class="n"&gt;extra_resource_limits&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;          &lt;/span&gt;&lt;span class="n"&gt;nvidia&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;com&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;gpu&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;1&amp;quot;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;h2 id="autoscaling"&gt;Autoscaling&lt;/h2&gt;
&lt;p&gt;Autoscaling is a big focus of this release, and many of the new features in 0.8 of the chart relate to scaling in some way. At the center is a new ‘user scheduler’ a custom kubernetes scheduler responsible for assigning user pods to nodes. To enable the user scheduler:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="n"&gt;scheduling&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="n"&gt;userScheduler&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;enabled&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;There are two major pain-points for autoscaling with JupyterHub. First, is &lt;strong&gt;scaling up&lt;/strong&gt;. Kubernetes scale-up is very basic: If scheduling a pod fails due insufficient resources, a new node is requested. This isn’t a great fit for JupyterHub. If you happen to be the unlucky user who is the first to need a new node, launching your server can take several minutes as the new node is requested and warming up. To prevent this, the chart adds the notion of “placeholder pods,” which are pods that request the same resources as user pods, but have lower priority and will thus be evicted immediately if a user pod needs their resources. This allows a cluster to always have a certain amount of “headroom” of vacant slots available for users so that requesting a new node occurs not when there are 0 slots available for users, but instead when there are 5 or 10 (you choose) slots available. See &lt;a href="https://discourse.jupyter.org/t/planning-placeholders-with-jupyterhub-helm-chart-0-8-tested-on-mybinder-org/213"&gt;this exploration&lt;/a&gt; of how mybinder.org arrived at 25 placeholders for its deployment. To enable 10 user placeholders:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="n"&gt;scheduling&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="n"&gt;podPriority&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;enabled&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="n"&gt;userPlaceholder&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;enabled&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;replicas&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;10&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The second major improvement to autoscaling that comes with the user scheduler is in &lt;strong&gt;scaling down&lt;/strong&gt;. If your cluster is automatically scaling up, you probably want it to scale down as well when you no longer need the added capacity. Unfortunately, because user pods cannot be kicked out, Kubernetes’ default scheduler doesn’t manage to drain nodes that are no longer needed without some manual intervention. The user scheduler adds this feature, ensuring that, over time, your newest nodes eventually drain when they are no longer needed. It accomplishes this by prioritizing the node with the most other user pods when picking a node to start new users. If there is an extra node, over time eventually all the users on that node will stop and the node will become idle and culled by the cluster autoscaler. This is the default behavior . There’s loads more scheduling optimizations available &lt;a href="https://zero-to-jupyterhub.readthedocs.io/en/latest/optimization.html#optimizations"&gt;in the updated documentation&lt;/a&gt;, including dedicating certain nodes to users so that user pods never run on the same node as others, etc.&lt;/p&gt;
&lt;p&gt;And finally, a huge thanks to the numerous contributors who helped make this release, be it via documentation contributions, testing, discussion, or code.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://zero-to-jupyterhub.readthedocs.io/"&gt;Check it out&lt;/a&gt;&lt;/p&gt;
</content><category term="JupyterHub"/><category term="Kubernetes"/></entry><entry><title>MyBinder.org serves two million launches</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2018/mybinder-org-serves-two-million-launches/" rel="alternate"/><published>2018-11-13T12:35:00+00:00</published><updated>2018-11-13T12:35:00+00:00</updated><author><name>Tim Head</name></author><id>tag:jasongrout.github.io,2018-11-13:/medium-archive/pelican/posts/2018/mybinder-org-serves-two-million-launches/</id><summary type="html">&lt;p&gt;by the Binder Team&lt;/p&gt;
</summary><content type="html">&lt;p&gt;by the &lt;a href="https://twitter.com/mybinderteam"&gt;Binder Team&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;Since the beginning of 2018, the Binder community has been hosting a &lt;a href="https://github.com/jupyterhub/binderhub"&gt;BinderHub&lt;/a&gt; at &lt;a href="https://mybinder.org"&gt;https://mybinder.org&lt;/a&gt; as a free public service. Today, we are proud to announce that this hub has served over two million &lt;a href="https://mybinder.readthedocs.io/en/latest/introduction.html"&gt;Binders&lt;/a&gt;. To mark this milestone we would like to say a huge Thank You! to the large community of people who use, &lt;a href="https://github.com/jupyterhub/binder#binder"&gt;build&lt;/a&gt;, and &lt;a href="https://www.moore.org/"&gt;fund&lt;/a&gt; the project. Without you this important public infrastructure would not be the user-friendly, reliable, well-supported and documented resource that we enjoy today. mybinder.org has enabled people from almost every country in the world to learn, participate and share countless projects, ideas and stories. Here’s to two million more!&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="A huge thank you to all those who help with building, using, and operating ." src="https://jasongrout.github.io/medium-archive/pelican/posts/2018/mybinder-org-serves-two-million-launches/images/001-1_Rs3YZOC_GtnOO37TRI7pQA.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;A huge thank you to all those who help with building, using, and operating &lt;a href="https://mybinder.org"&gt;https://mybinder.org&lt;/a&gt;.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="what-is-mybinderorg"&gt;What is mybinder.org?&lt;/h2&gt;
&lt;p&gt;mybinder.org let’s you take a repository full of Jupyter notebooks or RMarkdown and turn it into a collection of interactive notebooks. You can share your work with anyone by sending them a simple link (&lt;a href="https://mybinder.org/v2/gh/norvig/pytudes/master?filepath=ipynb%2FMaze.ipynb"&gt;like this one&lt;/a&gt;). All they have to do is open the link in a web browser and they can run those notebooks from anywhere in the world without having to install anything.&lt;/p&gt;
&lt;h2 id="who-is-using-mybinderorg"&gt;Who is using mybinder.org?&lt;/h2&gt;
&lt;p&gt;Currently about 70–80,000 &lt;a href="https://mybinder.readthedocs.io/en/latest/introduction.html"&gt;Binders&lt;/a&gt; are launched every week. A lot of those are people who are looking for a quick and easy way to launch a &lt;a href="https://mybinder.org/v2/gh/ipython/ipython-in-depth/master?filepath=binder/Index.ipynb"&gt;Python&lt;/a&gt; or &lt;a href="http://beta.mybinder.org/v2/gh/binder-examples/r/master?urlpath=rstudio"&gt;RStudio&lt;/a&gt; environment. However in the last week a notebook showing off &lt;a href="https://mybinder.org/v2/gh/quasiben/fiftyfizzbuzzes/master?filepath=Fifty%20Fizzbuzzes.ipynb"&gt;fifty ways to solve Fizz Buzz&lt;/a&gt; has been getting a lot of love. Beyond those heavy hitters and short-lived audience favorites there is a long tail of over 400 unique repositories that get launched every week. It would take the rest of the post to list them all!&lt;/p&gt;
&lt;p&gt;One last statistic that we are particularly proud of: over the last 80 days we have had users from almost all around the globe! Binder was started as a way to make computational research easier to share and reuse. We have been amazed at how many people around the world have used Binder for teaching classes, reproducing results, sharing interactive analyses, and making their work more accessible to others. We are particularly proud that this includes people from around the entire world.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Countries from which has received visitors between 22 August 2018 and 10 November 2018." src="https://jasongrout.github.io/medium-archive/pelican/posts/2018/mybinder-org-serves-two-million-launches/images/002-1_t6_W35g1Kl1Yr8oGPVlyIA.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Countries from which &lt;a href="https://mybinder.org"&gt;https://mybinder.org&lt;/a&gt; has received visitors between 22 August 2018 and 10 November 2018.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="data-set-of-all-launches-on-mybinderorg"&gt;Data set of all launches on mybinder.org&lt;/h2&gt;
&lt;p&gt;mybinder.org is operated as a public infrastructure that is transparent, open, and inclusive. We &lt;a href="https://gitter.im/jupyterhub/binder"&gt;chat&lt;/a&gt;, &lt;a href="https://discourse.jupyter.org/"&gt;discuss&lt;/a&gt;, and &lt;a href="https://github.com/jupyterhub/binder"&gt;work&lt;/a&gt; in the open. This is why we are now publishing a new data set: a continuously updated log of every launch that happens on mybinder.org!&lt;/p&gt;
&lt;p&gt;&lt;a href="https://archive.analytics.mybinder.org/"&gt;MyBinder.org Events Archive&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;We would love to see people explore this data set as a public resource that describes the kinds of repositories being shared and launched on the public mybinder.org deployment.&lt;/p&gt;
&lt;h2 id="a-new-badge"&gt;A new badge!&lt;/h2&gt;
&lt;p&gt;One more thing … we thought now is a good time to give the trusty “Launch Binder” badge an overhaul. To improve the badge we put together some suggestions, &lt;a href="https://discourse.jupyter.org/t/help-us-choose-an-updated-launch-binder-badge/100"&gt;reached out to the community&lt;/a&gt;, and within a few days received a lot of feedback and new ideas. After combining all the inputs, our new badge went live earlier this week. We present to you our new badge:&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="The new “Launch Binder” badge. Binder blue instead of bright red!" src="https://jasongrout.github.io/medium-archive/pelican/posts/2018/mybinder-org-serves-two-million-launches/images/003-1_gHu0NhZgTuRlvnD7eDI6Og.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;The new “Launch Binder” badge. Binder blue instead of bright red!&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;We hope you like it as much as we do. If you are in love with the old design or not quite ready to switch yet, do not worry! The old badge is not going anywhere. If you are using the old badge in your README it will continue to look the same as it always has.&lt;/p&gt;
&lt;p&gt;If you do want to change a previously generated link to the new badge, edit the name of the SVG in the link from the old:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="n"&gt;![Binder&lt;/span&gt;&lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nl"&gt;http&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="o"&gt;//&lt;/span&gt;&lt;span class="n"&gt;mybinder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;org&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;badge&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;svg&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="err"&gt;]&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nl"&gt;http&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="o"&gt;//&lt;/span&gt;&lt;span class="n"&gt;mybinder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;org&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;v2&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;gh&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;binder&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;examples&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;master&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;to the new:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="n"&gt;![Binder&lt;/span&gt;&lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nl"&gt;http&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="o"&gt;//&lt;/span&gt;&lt;span class="n"&gt;mybinder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;org&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;badge_logo&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;svg&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="err"&gt;]&lt;/span&gt;&lt;span class="p"&gt;(&lt;/span&gt;&lt;span class="nl"&gt;http&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="o"&gt;//&lt;/span&gt;&lt;span class="n"&gt;mybinder&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;org&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;v2&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;gh&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;binder&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;examples&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;r&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;master&lt;/span&gt;&lt;span class="p"&gt;)&lt;/span&gt;&lt;span class="n"&gt;Outro&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;h2 id="open-infrastructure-in-the-cloud"&gt;Open infrastructure in the cloud&lt;/h2&gt;
&lt;p&gt;The Binder project is a community-driven experiment in radically-open infrastructure. BinderHub, the underlying technology that powers a Binder deployment, is an open project and can be deployed in many other cloud environments. For example, see the &lt;a href="http://binder.pangeo.io/"&gt;Pangeo Binder deployment&lt;/a&gt; for geospatial analytics, or the &lt;a href="https://notebooks.gesis.org/binder/"&gt;Gesis Binder deployment for social sciences&lt;/a&gt;. We are excited to see the project head in new directions as we continue to grow the technology and the community around Binder.&lt;/p&gt;
&lt;p&gt;Finally, we could not have done any of this without a ton of support from the Binder community. First, many thanks to &lt;a href="https://www.moore.org/"&gt;the Moore Foundation&lt;/a&gt; for funding initial development of Binder’s underlying tech, and for helping us finance running the deployment at mybinder.org. Second, many thanks to the &lt;a href="https://jupyterhub-team-compass.readthedocs.io/en/latest/team.html"&gt;Binder project core team&lt;/a&gt; for fostering excellent technology and a great community. Finally, thanks to everybody in the Binder community — whether you’ve launched repositories, shared your Binders, participated in discussions, built features, &lt;a href="https://github.com/jupyterhub/team-compass/issues/67"&gt;helped design this post’s banner image&lt;/a&gt; (❤), or gave us some critical feedback. Binder’s purpose is to serve the community, and you all have made it so worth it!&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;&lt;em&gt;If you’d like to get involved in the Binder project, here are a few helpful links:&lt;/em&gt;&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;em&gt;To learn about Binder, see the Binder documentation:&lt;/em&gt; &lt;a href="http://docs.mybinder.org"&gt;&lt;em&gt;docs.mybinder.org&lt;/em&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;em&gt;For information on how to deploy your own BinderHub, see the BinderHub deployment docs:&lt;/em&gt; &lt;a href="http://binderhub.readthedocs.io"&gt;&lt;em&gt;binderhub.readthedocs.io&lt;/em&gt;&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;em&gt;To participate in conversations with the Binder community, say hello on the Binder gitter channel (&lt;/em&gt;&lt;a href="https://gitter.im/jupyterhub/binder"&gt;&lt;em&gt;https://gitter.im/jupyterhub/binder&lt;/em&gt;&lt;/a&gt;&lt;em&gt;) or the JupyterHub/Binder Discourse forum (&lt;/em&gt;&lt;a href="https://discourse.jupyter.org"&gt;&lt;em&gt;discourse.jupyter.org&lt;/em&gt;&lt;/a&gt;).&lt;/li&gt;
&lt;li&gt;&lt;em&gt;If you’d like to see the code itself, see the three main open projects that make up a Binder deployment: BinderHub (&lt;/em&gt;&lt;a href="http://github.com/jupyterhub/binderhub"&gt;&lt;em&gt;github.com/jupyterhub/binderhub&lt;/em&gt;&lt;/a&gt;&lt;em&gt;), repo2docker (&lt;/em&gt;&lt;a href="http://github.com/jupyter/repo2docker"&gt;&lt;em&gt;github.com/jupyter/repo2docker&lt;/em&gt;&lt;/a&gt;&lt;em&gt;), and JupyterHub (&lt;/em&gt;&lt;a href="http://github.com/jupyterhub/jupyterhub"&gt;&lt;em&gt;github.com/jupyterhub/jupyterhub&lt;/em&gt;&lt;/a&gt;)&lt;/li&gt;
&lt;/ul&gt;
</content><category term="Binder"/><category term="Kubernetes"/></entry><entry><title>On-demand Notebooks with JupyterHub, Jupyter Enterprise Gateway and Kubernetes</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2018/on-demand-notebooks-with-jupyterhub-jupyter-enterprise/" rel="alternate"/><published>2018-10-16T15:03:00+00:00</published><updated>2018-11-01T20:29:00+00:00</updated><author><name>Luciano Resende</name></author><id>tag:jasongrout.github.io,2018-10-16:/medium-archive/pelican/posts/2018/on-demand-notebooks-with-jupyterhub-jupyter-enterprise/</id><summary type="html">&lt;p&gt;by: Luciano Resende, Kevin Bates, Alan Chin&lt;/p&gt;
</summary><content type="html">&lt;p&gt;by: &lt;a href="https://twitter.com/lresende1975"&gt;Luciano Resende&lt;/a&gt;, &lt;a href="https://twitter.com/kbates4"&gt;Kevin Bates&lt;/a&gt;, Alan Chin&lt;/p&gt;
&lt;p&gt;&lt;a href="https://jupyter-notebook.readthedocs.io/en/stable/"&gt;&lt;strong&gt;Jupyter Notebook&lt;/strong&gt;&lt;/a&gt; has become the “de facto” platform used by data scientists to build interactive applications and to tackle big data and AI problems.&lt;/p&gt;
&lt;p&gt;With the increased adoption of Machine Learning and AI by enterprises, we have seen more and more requirement to build analytics platforms that provide on-demand notebooks for data scientists and data engineers in general.&lt;/p&gt;
&lt;p&gt;This article describes how to deploy multiple components from the Jupyter Notebook stack to provide an on-demand analytics platform powered by JupyterHub and Jupyter Enterprise Gateway on a Kubernetes cluster.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Image 1 — Deployment Architecture" src="https://jasongrout.github.io/medium-archive/pelican/posts/2018/on-demand-notebooks-with-jupyterhub-jupyter-enterprise/images/001-1_F_jJ1nDSQgBhrkbsEXB93Q.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Image 1 — Deployment Architecture&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="on-demand-notebooks-infrastructure"&gt;On-Demand Notebooks Infrastructure&lt;/h2&gt;
&lt;p&gt;Below are the main components we are going to use to build our solution, and its high-level description:&lt;/p&gt;
&lt;p&gt;&lt;a href="https://jupyterhub.readthedocs.io/"&gt;JupyterHub&lt;/a&gt; enables the creation of a multi-user Hub which spawns, manages, and proxies multiple instances of the single-user &lt;a href="https://jupyter-notebook.readthedocs.io/"&gt;Jupyter Notebook&lt;/a&gt; server providing the ‘as a service’ feeling we are looking for.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://jupyter-enterprise-gateway.readthedocs.io/en/latest/"&gt;Jupyter Enterprise Gateway&lt;/a&gt; provides optimal resource allocations by enabling kernels to be launched in its own pod enabling notebook pods to have minimal resources while kernel specific resources are allocated/deallocated accordingly to its lifecycle. It also enables the base image of the kernel to become a choice.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://kubernetes.io/"&gt;Kubernetes&lt;/a&gt; enables easy management of containerized applications and resources with the benefit of Elasticity and multiple other quality of services.&lt;/p&gt;
&lt;h2 id="jupyterhub-deployment"&gt;JupyterHub Deployment&lt;/h2&gt;
&lt;p&gt;JupyterHub is the entry point for our solution, it will manage user authorization and provisioning of individual Notebook servers for each user.&lt;/p&gt;
&lt;p&gt;JupyterHub configuration is done via a config.yaml, and the following settings are required:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Enable custom notebook configuration (coming from the customized user image).&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="n"&gt;hub&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="n"&gt;extraConfig&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|-&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;/etc/jupyter/jupyter_notebook_config.py&amp;#39;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;Define the docker image to be used when instantiating the notebook server for each user&lt;/li&gt;
&lt;li&gt;Define custom environment variables used to connect the Notebook server with Jupyter Enterprise Gateway to enable support for remote kernels&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="n"&gt;singleuser&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="n"&gt;image&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;elyra&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;nb2kg&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;tag&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;dev&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="n"&gt;storage&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="kd"&gt;dynamic&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="n"&gt;storageClass&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;nfs&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="kd"&gt;dynamic&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="n"&gt;extraEnv&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;KG_URL&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;FQDN&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;of&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;Gateway&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;Endpoint&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;KG_HTTP_USER&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;jovyan&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;KERNEL_USERNAME&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;jovyan&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;KG_REQUEST_TIMEOUT&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The complete config.yaml would look like the one below:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="n"&gt;hub&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="n"&gt;db&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;type&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;sqlite&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;memory&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="n"&gt;extraConfig&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;|-&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;config&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;/etc/jupyter/jupyter_notebook_config.py&amp;#39;&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;c&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;Spawner&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;cmd&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="s1"&gt;&amp;#39;jupyter-labhub&amp;#39;&lt;/span&gt;&lt;span class="o"&gt;]&lt;/span&gt;
&lt;span class="n"&gt;proxy&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="n"&gt;secretToken&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;xxx&amp;quot;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="n"&gt;ingress&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="n"&gt;enabled&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="n"&gt;hosts&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;FQDN&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;Kubernetes&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;Master&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="n"&gt;singleuser&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="n"&gt;defaultUrl&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;/lab&amp;quot;&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="n"&gt;image&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;name&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;elyra&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;nb2kg&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;tag&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;2.0&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="mi"&gt;0&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="na"&gt;dev0&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="n"&gt;storage&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="kd"&gt;dynamic&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;      &lt;/span&gt;&lt;span class="n"&gt;storageClass&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;nfs&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="kd"&gt;dynamic&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="n"&gt;extraEnv&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;KG_URL&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;&amp;lt;&lt;/span&gt;&lt;span class="n"&gt;FQDN&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;of&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;Gateway&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;Endpoint&lt;/span&gt;&lt;span class="o"&gt;&amp;gt;&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;KG_HTTP_USER&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;jovyan&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;KERNEL_USERNAME&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;jovyan&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;KG_REQUEST_TIMEOUT&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mi"&gt;60&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="n"&gt;rbac&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="n"&gt;enabled&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="n"&gt;debug&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;
&lt;span class="w"&gt;  &lt;/span&gt;&lt;span class="n"&gt;enabled&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kc"&gt;true&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Detailed deployment instructions for JupyterHub can be found at &lt;a href="https://zero-to-jupyterhub.readthedocs.io/en/stable/"&gt;Zero to JupyterHub for Kubernetes&lt;/a&gt;, but the command below would deploy it into a Kubernetes environment.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;helm upgrade --install --force hub jupyterhub/jupyterhub --namespace hub --version 0.7.0 --values jupyterhub-config.yaml
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;h2 id="custom-jupyterhub-user-image"&gt;Custom JupyterHub user image&lt;/h2&gt;
&lt;p&gt;By default, JupyterHub would deploy a vanilla Notebook Server image which will require that all resources ever used by the image to be allocated when the Kubernetes image is instantiated.&lt;/p&gt;
&lt;p&gt;Our custom image will enable kernels to be started in its own pod, promoting a better resource allocation as resources can be allocated and freed up as needed. This also gives us the flexibility of supporting different frameworks for different notebooks (e.g. a notebook using Python and TensorFlow, while another is using Python and Caffe2).&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Dockerfile for &lt;a href="https://github.com/jupyter/enterprise_gateway/tree/master/etc/docker/nb2kg"&gt;elyra-nb2kg custom image&lt;/a&gt;:&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;FROM jupyterhub/k8s-singleuser-sample:0.7.0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="gh"&gt;#&lt;/span&gt; Do the pip installs as the unprivileged notebook user
USER $NB_USER
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;ADD jupyter_notebook_config.py /etc/jupyter/jupyter_notebook_config.py
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;# Install NB2KG
RUN pip install --upgrade nb2kg &amp;amp;&amp;amp; \
    jupyter serverextension enable --py nb2kg --sys-prefix
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;Jupyter Notebook custom configuration to override Notebook handlers with the ones from NB2KG that will enable the notebook to connect with the Enterprise Gateway that enables remote kernels.&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="kn"&gt;from&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;jupyter_core.paths&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="kn"&gt;import&lt;/span&gt; &lt;span class="n"&gt;jupyter_data_dir&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;subprocess&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;os&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;errno&lt;/span&gt;
&lt;span class="kn"&gt;import&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nn"&gt;stat&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;c = get_config()
c.NotebookApp.ip = &amp;#39;*&amp;#39;
c.NotebookApp.port = 8888
c.NotebookApp.open_browser = False
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;c.NotebookApp.session_manager_class = &amp;#39;nb2kg.managers.SessionManager&amp;#39;
c.NotebookApp.kernel_manager_class = &amp;#39;nb2kg.managers.RemoteKernelManager&amp;#39;
c.NotebookApp.kernel_spec_manager_class = &amp;#39;nb2kg.managers.RemoteKernelSpecManager&amp;#39;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="gh"&gt;#&lt;/span&gt; https://github.com/jupyter/notebook/issues/3130
c.FileContentsManager.delete_to_trash = False
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="gh"&gt;#&lt;/span&gt; Generate a self-signed certificate
if &amp;#39;GEN_CERT&amp;#39; in os.environ:
    dir_name = jupyter_data_dir()
    pem_file = os.path.join(dir_name, &amp;#39;notebook.pem&amp;#39;)
    try:
        os.makedirs(dir_name)
    except OSError as exc:  # Python &amp;gt;2.5
        if exc.errno == errno.EEXIST and os.path.isdir(dir_name):
            pass
        else:
            raise
    # Generate a certificate if one doesn&amp;#39;t exist on disk
    subprocess.check_call([&amp;#39;openssl&amp;#39;, &amp;#39;req&amp;#39;, &amp;#39;-new&amp;#39;,
                           &amp;#39;-newkey&amp;#39;, &amp;#39;rsa:2048&amp;#39;,
                           &amp;#39;-days&amp;#39;, &amp;#39;365&amp;#39;,
                           &amp;#39;-nodes&amp;#39;, &amp;#39;-x509&amp;#39;,
                           &amp;#39;-subj&amp;#39;, &amp;#39;/C=XX/ST=XX/L=XX/O=generated/CN=generated&amp;#39;,
                           &amp;#39;-keyout&amp;#39;, pem_file,
                           &amp;#39;-out&amp;#39;, pem_file])
    # Restrict access to the file
    os.chmod(pem_file, stat.S_IRUSR | stat.S_IWUSR)
    c.NotebookApp.certfile = pem_file
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Note that the document above was generated by &lt;code&gt;jupyter notebook --generate-config&lt;/code&gt; and then updated with the required handlers override:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;c.NotebookApp.session_manager_class = &amp;#39;nb2kg.managers.SessionManager&amp;#39;
c.NotebookApp.kernel_manager_class = &amp;#39;nb2kg.managers.RemoteKernelManager&amp;#39;
c.NotebookApp.kernel_spec_manager_class = &amp;#39;nb2kg.managers.RemoteKernelSpecManager&amp;#39;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;h2 id="jupyter-enterprise-gateway-deployment"&gt;Jupyter Enterprise Gateway deployment&lt;/h2&gt;
&lt;p&gt;Jupyter Enterprise Gateway enables Jupyter Notebook to launch and manage remote kernels in a distributed cluster, including Kubernetes cluster.&lt;/p&gt;
&lt;p&gt;Enterprise Gateway provides a Kubernetes deployment descriptor that makes it simple to deploy it on a Kubernetes environment with the command below:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;kubectl apply -f https://raw.githubusercontent.com/jupyter-incubator/enterprise_gateway/master/etc/kubernetes/enterprise-gateway.yaml
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;We also recommend that the kernel images be downloaded on all nodes of the Kubernetes cluster to avoid delays/timeouts when launching kernels for the first time on these nodes.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;docker pull elyra/enterprise-gateway:dev
docker pull elyra/kernel-py:dev
docker pull elyra/kernel-tf-py:dev
docker pull elyra/kernel-r:dev
docker pull elyra/kernel-scala:dev
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;h2 id="automated-one-click-deployment-using-ansible"&gt;Automated One-Click Deployment using Ansible&lt;/h2&gt;
&lt;p&gt;If you are eager to get started and try this in a few machines, we have published an&lt;a href="https://github.com/lresende/ansible-kubernetes-cluster"&gt;&lt;code&gt;ansible script&lt;/code&gt;&lt;/a&gt; that deploys the full set of components described above on vanilla RHEL machines/VMs.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;ansible-playbook --verbose setup-kubernetes.yml -c paramiko -i hosts-fyre-kubernetes
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;h2 id="conclusion"&gt;Conclusion&lt;/h2&gt;
&lt;p&gt;Jupyter Enterprise Gateway provides remote kernel management to Jupyter Notebooks. In a JupyterHub/Kubernetes environment, it enables hub to launch tiny Jupyter Notebook pods and only allocate large kernel resources when these are created as independent pods. This approach also allows for easy sharing of expensive resources as GPUs, etc&lt;/p&gt;
&lt;h2 id="special-thanks"&gt;Special Thanks&lt;/h2&gt;
&lt;p&gt;Special thanks to &lt;a href="https://twitter.com/e_sundell"&gt;Erik Sundell&lt;/a&gt; and &lt;a href="https://twitter.com/minrk"&gt;Min RK&lt;/a&gt; from the JupyterHub team for the support and initial discussions around JupyterHub.&lt;/p&gt;
</content><category term="Jupyter Enterprise Gateway"/><category term="JupyterHub"/><category term="Kubernetes"/></entry><entry><title>Deploying JupyterHub with Kubernetes on OpenStack</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2018/how-to-deploy-jupyterhub-with-kubernetes-on-openstack/" rel="alternate"/><published>2018-10-15T13:11:00+00:00</published><updated>2018-10-15T15:04:00+00:00</updated><author><name>Loïc Gouarin</name></author><id>tag:jasongrout.github.io,2018-10-15:/medium-archive/pelican/posts/2018/how-to-deploy-jupyterhub-with-kubernetes-on-openstack/</id><summary type="html">&lt;p&gt;Jupyter is now widely used for teaching and research. The use of Kubernetes for JupyterHub deployment enabled reliable setups scaling to…&lt;/p&gt;
</summary><content type="html">&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2018/how-to-deploy-jupyterhub-with-kubernetes-on-openstack/images/001-1_cbBQmPCCtGh-j4a45vfcmA.webp" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Jupyter is now widely used for teaching and research. The use of Kubernetes for deploying a JupyterHub has enabled reliable setups scaling to thousands of users.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;There are many cloud computing vendors (Google, Amazon, …) and the first attempts to use JupyterHub with Kubernetes is based on them. But relying on vendor clouds increases the risk of vendor lock-in.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;In addition, there are many pre-existing academic clouds managed by people with a high level of expertise and a thorough knowledge of their infrastructure and associated tools. These are often more cost-effective for research and education. Could we build upon these academic cloud computing to provide scalable and high-quality infrastructure for education and research?&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;In this post, we will focus on how to deploy JupyterHub with Kubernetes on OpenStack. A first attempt to create academic cloud computing in France.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;This post is split into two parts (see links below).&lt;/p&gt;
&lt;p&gt;&lt;a href="/posts/2018/how-to-deploy-jupyterhub-with-kubernetes-on-openstack/#bcd0"&gt;&lt;strong&gt;Why to deploy JupyterHub on OpenStack&lt;/strong&gt;&lt;/a&gt; is a high-level description of our problem, and our steps to solve it. It explains why we want to deploy a JupyterHub on OpenStack, what difficulties we have encountered, and what we want to do in a near future.&lt;/p&gt;
&lt;p&gt;&lt;a href="/posts/2018/how-to-deploy-jupyterhub-with-kubernetes-on-openstack/#4157"&gt;&lt;strong&gt;A technical guide to deploying JupyterHub on OpenStack&lt;/strong&gt;&lt;/a&gt; is an in-depth guide that you may follow in order to deploy your own JupyterHub on Kubernetes on OpenStack. It’s designed for any person interested in how to replicate our deployment on their own infrastructure.&lt;/p&gt;
&lt;h2 id="our-story-our-difficulties-and-our-plans"&gt;Our story, our difficulties and our plans&lt;/h2&gt;
&lt;h3 id="why-deploy-jupyterhub-on-openstack"&gt;Why deploy JupyterHub on OpenStack?&lt;/h3&gt;
&lt;p&gt;JupyterHub, the multi-user Jupyter server, has been actively developed since 2014 and has seen a rapidly growing adoption in the past year.&lt;/p&gt;
&lt;p&gt;You may know about &lt;a href="https://zero-to-jupyterhub.readthedocs.io"&gt;Zero to JupyterHub&lt;/a&gt;, which provides step-by-step instructions for installing JupyterHub using a vendor-managed Kubernetes cluster. In the guide, you can also find how to set up a Kubernetes cluster on many vendor clouds such as AWS, Azure, and more recently on OpenShift. But what about other cloud infrastructures based on open-source infrastructure, such as OpenStack? While cloud vendors often provide you with many tools that make your life easier, OpenStack requires more explicit configuration and setup.&lt;/p&gt;
&lt;p&gt;Earlier this year, we set up a working group across several France universities to explore how to easily set up JupyterHub for teaching and research in our academic cloud infrastructures. It turns out that the technology used across these academic clouds is OpenStack. One of the objectives of this working group is to make just as easy to deploy JupyterHub on OpenStack as compared to following &lt;em&gt;Zero to JupyterHub&lt;/em&gt; and using vendor infrastructure.&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Note: We are not the first to work on this problem, and we should also mention the work done in Canada through &lt;a href="http://intro.syzygy.ca/"&gt;Syzygy.ca&lt;/a&gt; which is a project of PIMS, Compute Canada, and Cybera. They have developed their own deployment tools using terraform and ansible scripts. Our approach differs in that while we use the same technological stack, we prefer not to build a custom deployment tool that we would need to maintain over time.&lt;/p&gt;
&lt;/blockquote&gt;
&lt;h3 id="issues-we-encountered"&gt;Issues we encountered&lt;/h3&gt;
&lt;p&gt;We started a deployment using Kubespray in January of this year and have had a bumpy path since then. To begin, we looked at what existed already in the OpenStack world. We came across &lt;a href="https://github.com/kubernetes-incubator/kubespray"&gt;Kubespray&lt;/a&gt;, which offers a great facility and a lot of flexibility when you want to deploy a Kubernetes cluster. An interesting fact about Kubespray is that it’s not dedicated to OpenStack infrastructures, so you should be able to follow the same procedure for other deployments such as a baremetal cluster.&lt;/p&gt;
&lt;p&gt;Using Kubespray, we very quickly had a Kubernetes cluster on OpenStack. However, we ran into network problems, and would lose network packets that made the JupyterHub completely unusable. It took us a long time to realize that we had &lt;a href="https://en.wikipedia.org/wiki/Maximum_transmission_unit"&gt;MTU issues&lt;/a&gt; and even longer to solve it. To make things harder, we used a production platform which made it very difficult to update. We finally solved the problem by using a test platform where we could have more freedom.&lt;/p&gt;
&lt;p&gt;In Kubespray, there are various CNIs (&lt;em&gt;Container Network Interface&lt;/em&gt;) and one of them (&lt;em&gt;weave&lt;/em&gt;) allows to modify the MTU. We tried to configure it carefully on the production platform, but we continued facing the same problem. Trying a new version of OpenStack on the test platform, we were able to solve the problem. That means that something bad had also happened with the LoadBalancer. For more explanation, see the &lt;a href="/posts/2018/how-to-deploy-jupyterhub-with-kubernetes-on-openstack/#5dbc"&gt;MTU section&lt;/a&gt; in the technical description below.&lt;/p&gt;
&lt;p&gt;We thought we could deploy JupyterHub with Kubernetes on OpenStack in a few weeks but as you can see, that’s not what happened at all. That’s why it was important for us to share our experience in the hopes that it makes the process easier for others. In the last section, we’ll cover more of the technical details for our deployment&lt;/p&gt;
&lt;h3 id="whats-next"&gt;What’s next ?&lt;/h3&gt;
&lt;p&gt;For us, the installation of JupyterHub on OpenStack was just the first step of a long journey. We are able now to offer to our researchers and our students a JupyterHub but we want more. Here’s a short wish-list our deployments.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;On-demand environments.&lt;/strong&gt; Imagine offering researchers and teachers an even more flexible platform where they can create their work environment and distribute them without needing to use central IT for the installation of their packages. As you may have guessed, we are more interested in what BinderHub has to offer.&lt;/p&gt;
&lt;p&gt;The steps described above also work for the installation of BinderHub. We deployed a BinderHub on OpenStack alongside a DockerHub registry. Kubespray also offers the possibility to deploy a private registry and we would like to test it with BinderHub.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Cluster monitoring.&lt;/strong&gt; It would also be great to have monitoring of the Kubernetes cluster. This would allow us to inspect the usage rates and resources available on the deployment. In the &lt;a href="https://github.com/kubernetes-incubator/kubespray/blob/master/docs/roadmap.md"&gt;roadmap&lt;/a&gt; of Kubespray, it is planned to add Grafana and Prometheus installations.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Authentication for BinderHub.&lt;/strong&gt; Currently BinderHub does not support authentication for users. However, note that a recent pull request on this subject was merged in BinderHub (see &lt;a href="https://github.com/jupyterhub/binderhub/pull/666"&gt;https://github.com/jupyterhub/binderhub/pull/666&lt;/a&gt;).&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Persistent storage in BinderHub.&lt;/strong&gt; It is also currently not possible to persist storage across BinderHub sessions. Once authentication is possible in BinderHub, we’d also like to connect user accounts to their storage so that they can keep their work over time. This will also require being able to mount the home directory of each user.&lt;/p&gt;
&lt;p&gt;We will work on all these items in the next months.&lt;/p&gt;
&lt;h2 id="the-technical-details"&gt;The Technical Details&lt;/h2&gt;
&lt;p&gt;This part details the set up of a Kubernetes cluster and JupyterHub using a bare OpenStack infrastructure. To make it as reproducible as possible, we will start by listing the versions we have used.&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;OpenStack&lt;/strong&gt;: Pike&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Kubespray&lt;/strong&gt;: commit &lt;a href="https://github.com/kubernetes-incubator/kubespray/commit/36322901a6c057a6a1f6a157abab63b165b2a0a8"&gt;3632290&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Kubernetes&lt;/strong&gt;: 1.11.3&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Helm&lt;/strong&gt;: 2.9.1&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;JupyterHub&lt;/strong&gt;: 0.7.0&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Now that we’ve described the components and the versions used, let’s start to deploy our JupyterHub on OpenStack !!&lt;/p&gt;
&lt;p&gt;The deployment steps are the following&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Connect to our OpenStack infrastructure&lt;/li&gt;
&lt;li&gt;Download Kubespray&lt;/li&gt;
&lt;li&gt;Create your infrastructure using terraform&lt;/li&gt;
&lt;li&gt;Deploy your Kubernetes cluster using ansible&lt;/li&gt;
&lt;li&gt;Deploy your JupyterHub using Helm chart&lt;/li&gt;
&lt;li&gt;Enjoy!&lt;/li&gt;
&lt;/ul&gt;
&lt;h3 id="connect-to-openstack"&gt;Connect to OpenStack&lt;/h3&gt;
&lt;p&gt;Kubespray needs a access to your OpenStack infrastructure in order to create all the instances needed for your Kubernetes cluster using the OpenStack CLI (&lt;em&gt;Command-Line Interface&lt;/em&gt;). When you log in to your OpenStack dashboard, you can download all the environment variables to use the CLI.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2018/how-to-deploy-jupyterhub-with-kubernetes-on-openstack/images/002-1_amhpJittHOQA_6dYHOWlCg.webp" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;p&gt;We chose to download the &lt;strong&gt;OpenStack RC File V3&lt;/strong&gt;. You should obtain something like this:&lt;/p&gt;
&lt;figure&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="ch"&gt;#!/usr/bin/env bash&lt;/span&gt;

&lt;span class="nb"&gt;export&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;OS_AUTH_URL&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;https://keystone.xxxxxx:5000/v3
&lt;span class="nb"&gt;export&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;OS_PROJECT_ID&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
&lt;span class="nb"&gt;export&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;OS_PROJECT_NAME&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;xxxxxxxx&amp;quot;&lt;/span&gt;
&lt;span class="nb"&gt;export&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;OS_USER_DOMAIN_NAME&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;xxxxxxxx&amp;quot;&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;-z&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;&lt;/span&gt;&lt;span class="nv"&gt;$OS_USER_DOMAIN_NAME&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;then&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;unset&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;OS_USER_DOMAIN_NAME&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;fi&lt;/span&gt;

&lt;span class="c1"&gt;# unset v2.0 items in case set&lt;/span&gt;
&lt;span class="nb"&gt;unset&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;OS_TENANT_ID
&lt;span class="nb"&gt;unset&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;OS_TENANT_NAME

&lt;span class="nb"&gt;export&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;OS_USERNAME&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;xxxxxxxx&amp;quot;&lt;/span&gt;
&lt;span class="nb"&gt;export&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;OS_REGION_NAME&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;xxxxxxxx&amp;quot;&lt;/span&gt;
&lt;span class="c1"&gt;# Don&amp;#39;t leave a blank variable, unset it if it was empty&lt;/span&gt;
&lt;span class="k"&gt;if&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;[&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;-z&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;&lt;/span&gt;&lt;span class="nv"&gt;$OS_REGION_NAME&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;]&lt;/span&gt;&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;then&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nb"&gt;unset&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;OS_REGION_NAME&lt;span class="p"&gt;;&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="k"&gt;fi&lt;/span&gt;

&lt;span class="nb"&gt;export&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;OS_INTERFACE&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;public
&lt;span class="nb"&gt;export&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;OS_IDENTITY_API_VERSION&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="m"&gt;3&lt;/span&gt;

&lt;span class="c1"&gt;# to be added&lt;/span&gt;
&lt;span class="nb"&gt;export&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;OS_PASSWORD&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;xxxxxxxx&amp;quot;&lt;/span&gt;
&lt;span class="nb"&gt;export&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;OS_CLOUD&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;xxxxxxxx
&lt;span class="nb"&gt;export&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nv"&gt;OS_CACERT&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;/home/loic/.certs/openstack_cacert.pem
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;figcaption&gt;
&lt;p&gt;Example of rc_file&lt;/p&gt;
&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Note that we’ve removed the lines which ask for your password when you use the CLI and add it to “never ask again”. You also have to provide OS_CLOUD and OS_CACERT (even if you don’t have a certificate to access to your OpenStack infrastructure, you must provide one but you can keep it blank).&lt;/p&gt;
&lt;p&gt;Now, you can install the OpenStack CLI with the command line. We’ll show two ways to do this below:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;with &lt;strong&gt;virtualenv&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;virtualenv ~/openstack
source ~/openstack/bin/activate
pip install python-openstackclient
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;ul&gt;
&lt;li&gt;with &lt;strong&gt;conda&lt;/strong&gt;&lt;/li&gt;
&lt;/ul&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;conda create -n openstack python=3.6
source activate openstack
pip install python-openstackclient
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Next, source your rc file to activate it&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;source rc_file
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;and test your connection&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;openstack project list
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;You should be able to see your projects listed.&lt;/p&gt;
&lt;p&gt;Once you have access, you will need some information in order to use terraform with Kubespray. You should find the following things (we have highlighted them in the images below):&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;The name of the image you want to deploy. To list it, run the following command:&lt;/li&gt;
&lt;/ol&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;openstack image list
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;figure&gt;
&lt;img alt="List of OpenStack images" src="https://jasongrout.github.io/medium-archive/pelican/posts/2018/how-to-deploy-jupyterhub-with-kubernetes-on-openstack/images/003-1_eHy-VQtYqNo-myHQKqkKFQ.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;List of OpenStack images&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;ol start="2"&gt;
&lt;li&gt;The id of the flavor describing the type of machine you want to deploy (the flavor must be aUUID and not an integer ID). To find it, run this command:&lt;/li&gt;
&lt;/ol&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;openstack flavor list
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;figure&gt;
&lt;img alt="List of OpenStack flavors" src="https://jasongrout.github.io/medium-archive/pelican/posts/2018/how-to-deploy-jupyterhub-with-kubernetes-on-openstack/images/004-1_iP9duMZM1rjBamzc8hl1Gg.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;List of OpenStack flavors&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Once you’ve got this information, it’s time to install Kubespray.&lt;/p&gt;
&lt;h3 id="install-kubespray"&gt;Install Kubespray&lt;/h3&gt;
&lt;p&gt;Because Kubespray is simply a GitHub repository, we don’t “install” it in a traditional sense, we only clone the repository to our machine. Since Kubespray is a project that evolves quickly, we’ll list the commit that we used for this post. Run the following command to get Kubespray:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;git clone https://github.com/kubernetes-incubator/kubespray.git
cd kubespray
git checkout 3632290
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Next, prepare all the files describing your Kubernetes cluster. We’ll follow &lt;a href="https://github.com/kubernetes-incubator/kubespray/tree/master/contrib/terraform/openstack"&gt;the documentation given by Kubespray&lt;/a&gt; and will just change some flags. We encourage you to follow the procedure described below as the documentation seems to have some errors.&lt;/p&gt;
&lt;p&gt;Kubespray uses terraform and ansible to deploy your Kubernetes cluster. ansible needs an inventory file which describes your cluster in order to execute the playbook roles on it. Kubespray provides a skeleton dedicated to OpenStack platform to provision your cluster using terraform and create the inventory file for ansible accordingly. To use the skeleton provided by Kubespray, the steps are the following&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;cp -LRp contrib/terraform/openstack/sample-inventory inventory/jhub
cd inventory/jhub
ln -s ../../contrib/terraform/openstack/hosts
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;&lt;strong&gt;jhub&lt;/strong&gt; is the name directory we choose to store our inventory but you can choose what you want.&lt;/p&gt;
&lt;p&gt;If you look at the &lt;strong&gt;inventory/jhub&lt;/strong&gt; directory you will see&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;cluster.tf&lt;/strong&gt;: the terraform file describing your inventory.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;group_vars&lt;/strong&gt;: the directory where we set all the variables used by ansible scripts provided by Kubespray.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;Let’s describe our inventory.&lt;/p&gt;
&lt;h3 id="initialize-terraform"&gt;Initialize Terraform&lt;/h3&gt;
&lt;p&gt;In the &lt;strong&gt;cluster.tf&lt;/strong&gt; file, you can specify different kinds of Kubernetes clusters with floating IP for each VM. &lt;em&gt;floating ip&lt;/em&gt; means that you ask to OpenStack to give you a public IP address in order to connect to the VM from an external network. You can also have a &lt;em&gt;bastion&lt;/em&gt; where you have to log before reaching your Kubernetes cluster.&lt;/p&gt;
&lt;p&gt;In the following, we only choose to have a master VM and two nodes for our Kubernetes cluster. Another important part is to specify a GlusterFS to have some storage resources for JupyterHub (database and home directories). Our inventory file &lt;strong&gt;cluster.tf&lt;/strong&gt; looks like this&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="c1"&gt;# your Kubernetes cluster name here&lt;/span&gt;
&lt;span class="na"&gt;cluster_name&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;jhub&amp;quot;&lt;/span&gt;

&lt;span class="c1"&gt;# SSH key to use for access to nodes&lt;/span&gt;
&lt;span class="na"&gt;public_key_path&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;~/.ssh/id_rsa.pub&amp;quot;&lt;/span&gt;

&lt;span class="c1"&gt;# image to use for bastion, masters, standalone etcd instances, and nodes&lt;/span&gt;
&lt;span class="na"&gt;image&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;CentOS-7-x86_64-GenericCloud-20180108.qcow2&amp;quot;&lt;/span&gt;
&lt;span class="c1"&gt;# user on the node (ex. core on Container Linux, ubuntu on Ubuntu, etc.)&lt;/span&gt;
&lt;span class="na"&gt;ssh_user&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;centos&amp;quot;&lt;/span&gt;

&lt;span class="c1"&gt;# 0|1 bastion nodes&lt;/span&gt;
&lt;span class="na"&gt;number_of_bastions&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;0&lt;/span&gt;
&lt;span class="c1"&gt;# standalone etcds&lt;/span&gt;
&lt;span class="na"&gt;number_of_etcd&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;0&lt;/span&gt;

&lt;span class="c1"&gt;# masters&lt;/span&gt;
&lt;span class="na"&gt;number_of_k8s_masters&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;1&lt;/span&gt;
&lt;span class="na"&gt;number_of_k8s_masters_no_etcd&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;0&lt;/span&gt;
&lt;span class="na"&gt;number_of_k8s_masters_no_floating_ip&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;0&lt;/span&gt;
&lt;span class="na"&gt;number_of_k8s_masters_no_floating_ip_no_etcd&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;0&lt;/span&gt;
&lt;span class="na"&gt;flavor_k8s_master&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;5e879c45-e709-4671-a38b-45ed0573dc38&amp;quot;&lt;/span&gt;

&lt;span class="c1"&gt;# nodes&lt;/span&gt;
&lt;span class="na"&gt;number_of_k8s_nodes&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;0&lt;/span&gt;
&lt;span class="na"&gt;number_of_k8s_nodes_no_floating_ip&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;2&lt;/span&gt;
&lt;span class="na"&gt;flavor_k8s_node&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;5e879c45-e709-4671-a38b-45ed0573dc38&amp;quot;&lt;/span&gt;

&lt;span class="c1"&gt;# GlusterFS&lt;/span&gt;
&lt;span class="c1"&gt;# either 0 or more than one&lt;/span&gt;
&lt;span class="na"&gt;number_of_gfs_nodes_no_floating_ip&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;2&lt;/span&gt;
&lt;span class="na"&gt;gfs_volume_size_in_gb&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;100&lt;/span&gt;
&lt;span class="c1"&gt;# Container Linux does not support GlusterFS&lt;/span&gt;
&lt;span class="na"&gt;image_gfs&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;CentOS-7-x86_64-GenericCloud-20180108.qcow2&amp;quot;&lt;/span&gt;
&lt;span class="c1"&gt;# May be different from other nodes&lt;/span&gt;
&lt;span class="na"&gt;ssh_user_gfs&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;centos&amp;quot;&lt;/span&gt;
&lt;span class="na"&gt;flavor_gfs_node&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;5e879c45-e709-4671-a38b-45ed0573dc38&amp;quot;&lt;/span&gt;

&lt;span class="c1"&gt;# networking&lt;/span&gt;
&lt;span class="c1"&gt;#network_name = &amp;quot;&amp;lt;network&amp;gt;&amp;quot;&lt;/span&gt;
&lt;span class="na"&gt;external_net&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;6cd08271-...&amp;quot;&lt;/span&gt;
&lt;span class="c1"&gt;#subnet_cidr = &amp;quot;&amp;lt;cidr&amp;gt;&amp;quot;&lt;/span&gt;
&lt;span class="na"&gt;floatingip_pool&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;public&amp;quot;&lt;/span&gt;

&lt;span class="na"&gt;dns_nameservers&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;8.8.8.8&amp;quot;, &amp;quot;8.8.4.4&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The flavors are the same for each master, node, and GlusterFS but you can do what you want. We also added &lt;strong&gt;dns_nameservers&lt;/strong&gt; to be sure that we have a correct DNS on each nodes. We will check in future experiments if it’s really necessary.&lt;/p&gt;
&lt;p&gt;The ID of the external network and the name of the &lt;strong&gt;floatingip_pool&lt;/strong&gt; can be obtained with the following command:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;openstack network list
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;figure&gt;
&lt;img alt="List of OpenStack networks" src="https://jasongrout.github.io/medium-archive/pelican/posts/2018/how-to-deploy-jupyterhub-with-kubernetes-on-openstack/images/005-1_xPXJTn5YDLAHJ6yWlInnuQ.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;List of OpenStack networks&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;You need ssh keys to be able to connect to the nodes. From the documentation of Kubespray:&lt;/p&gt;
&lt;blockquote&gt;
&lt;p&gt;Ensure your local ssh-agent is running and your ssh key has been added. This step is required by the terraform provisioner:&lt;/p&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;p&gt;&lt;code&gt;eval $(ssh-agent -s)   ssh-add ~/.ssh/id_rsa&lt;/code&gt;&lt;/p&gt;
&lt;/blockquote&gt;
&lt;p&gt;Now, it’s time to initialize terraform. It’s important for the next steps to be run from the root directory of Kubespray.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;terraform init contrib/terraform/openstack
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Now, create your VMs!&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="n"&gt;terraform&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;apply&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;inventory&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;jhub&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;terraform&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tfstate&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="k"&gt;var&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;file&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;inventory&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;jhub&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;cluster&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tf&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;contrib&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;terraform&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;openstack&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;At the end of this process, you can see your instances in the dashboard of OpenStack. It’s important to keep the information given at the end of the output.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2018/how-to-deploy-jupyterhub-with-kubernetes-on-openstack/images/006-1_1kFAkM2Y6HAaA8fOs7b3aw.webp" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;p&gt;At this stage, you’ve just created several VMs with the images given in the &lt;strong&gt;cluster.tf&lt;/strong&gt; file. You don’t have a Kubernetes cluster up and running yet. It’s the next step!&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Note:&lt;/strong&gt; If you want to destroy all that you’ve done, run this command:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="n"&gt;terraform&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;destroy&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;state&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;inventory&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;jhub&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;terraform&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tfstate&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="k"&gt;var&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;file&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;inventory&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;jhub&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;cluster&lt;/span&gt;&lt;span class="o"&gt;.&lt;/span&gt;&lt;span class="n"&gt;tf&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;contrib&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;terraform&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;openstack&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;h3 id="configure-your-kubernetes-cluster"&gt;Configure your Kubernetes cluster&lt;/h3&gt;
&lt;p&gt;Again, Kubespray lets you configure your Kubernetes cluster with a lot of possibilities. We will show you one setup, but once you understand the procedure, you should be able to make your own choices.&lt;/p&gt;
&lt;p&gt;Let’s start to see if we can ping our VMs. You have to add the following script that we called &lt;strong&gt;ssh-nodes.conf&lt;/strong&gt; in your &lt;strong&gt;inventory/jhub&lt;/strong&gt; directory&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="n"&gt;Host&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="mf"&gt;10.0.0&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="o"&gt;*&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;User&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;centos&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;UserKnownHostsFile&lt;/span&gt;&lt;span class="o"&gt;=/&lt;/span&gt;&lt;span class="n"&gt;dev&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;null&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;StrictHostKeyChecking&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;no&lt;/span&gt;
&lt;span class="w"&gt;    &lt;/span&gt;&lt;span class="n"&gt;ProxyCommand&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;ssh&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;o&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;UserKnownHostsFile&lt;/span&gt;&lt;span class="o"&gt;=/&lt;/span&gt;&lt;span class="n"&gt;dev&lt;/span&gt;&lt;span class="o"&gt;/&lt;/span&gt;&lt;span class="n"&gt;null&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;o&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;StrictHostKeyChecking&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="n"&gt;no&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="n"&gt;W&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;%&lt;/span&gt;&lt;span class="n"&gt;h&lt;/span&gt;&lt;span class="o"&gt;:%&lt;/span&gt;&lt;span class="n"&gt;p&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;centos&lt;/span&gt;&lt;span class="mf"&gt;@134.&lt;/span&gt;&lt;span class="n"&gt;xx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;xx&lt;/span&gt;&lt;span class="p"&gt;.&lt;/span&gt;&lt;span class="n"&gt;xx&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;You also have to modify the &lt;strong&gt;ssh_args&lt;/strong&gt; variable in &lt;strong&gt;ansible.cfg&lt;/strong&gt; script in the root directory of Kubespray accordingly&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;ssh_args = -F inventory/jhub/ssh-nodes.conf -o ControlMaster=auto -o ControlPersist=30m -o ConnectionAttempts=100 -o UserKnownHostsFile=/dev/null -o StrictHostKeyChecking=no
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Just pay attention that you give the right external address and that your internal network is &lt;strong&gt;10.0.0.*&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;To check that everything is configured correctly, this command:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;ansible -i inventory/jhub/hosts -m ping all
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;should have an output like this&lt;/p&gt;
&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2018/how-to-deploy-jupyterhub-with-kubernetes-on-openstack/images/007-1_NFZ3Rd-JNp0W8rxZK6EJ-Q.webp" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;p&gt;You can now install your Kubernetes cluster with the ansible scripts provided by Kubespray. To do that, you will edit the files found in the &lt;strong&gt;group_vars&lt;/strong&gt; directory in &lt;strong&gt;inventory/jhub&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;You need several things to have a JupyterHub up and running&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A CNI (Container Network Interface).&lt;/strong&gt; Kubespray offers different CNI for your Kubernetes cluster: cilium, calico, contiv, weave or flannel. We will choose &lt;strong&gt;calico&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Storage for the data.&lt;/strong&gt; We’ll deploy a GlusterFS and add storage on the Kubernetes cluster to have access to it.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;A LoadBalancer&lt;/strong&gt; to have access to the service from the external network. You have two kinds of LoadBalancer on Openstack: Neutron or Octavia. You can use both with Kubespray. We will choose &lt;strong&gt;Neutron&lt;/strong&gt; but it will be preferable to use Octavia in the future.&lt;/p&gt;
&lt;p&gt;So how do we configure all these items?&lt;/p&gt;
&lt;p&gt;First, open the file &lt;strong&gt;inventory/jhub/group_vars/all/all.yml&lt;/strong&gt; and modify the following entries&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="n"&gt;bootstrap_os&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;centos&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;upstream_dns_servers:
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;    - 8.8.8.8
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;    - 8.8.4.4
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="n"&gt;cloud_provider&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;openstack&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Note that the dns address is specific to our infrastructure.&lt;/p&gt;
&lt;p&gt;Now, open the file &lt;strong&gt;inventory/jhub/group_vars/all/openstack.yml&lt;/strong&gt; and configure the LoadBalancer&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="n"&gt;openstack_lbaas_enabled&lt;/span&gt;&lt;span class="o"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="n"&gt;True&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;openstack_lbaas_subnet_id: &amp;quot;48ec8433-...&amp;quot;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;openstack_lbaas_floating_network_id: &amp;quot;6cd08271-...&amp;quot;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;The two IDs are those given at the end of the &lt;strong&gt;terraform apply&lt;/strong&gt; step.&lt;/p&gt;
&lt;p&gt;Open the file &lt;strong&gt;inventory/jhub/group_vars/k8s-cluster/k8s-cluster.yml&lt;/strong&gt; and set &lt;strong&gt;persistent_volumes_enabled&lt;/strong&gt; to &lt;strong&gt;true&lt;/strong&gt; and &lt;strong&gt;resolvconf_mode&lt;/strong&gt; to &lt;strong&gt;host_resolvconf&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Our OpenStack cloud infrastructure is configured with a VXLAN tunnel where the header size is 50 bytes. We use Calico with Kubernetes which also uses a VXLAN tunnel with a header of 50 bytes. Then, for a default MTU of 1500 bytes, we already have 100 bytes taken by the headers. So, we need to configure carefully the MTU of calico in order to be sure that the packet size (headers included) doesn’t exceed the 1500 bytes.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2018/how-to-deploy-jupyterhub-with-kubernetes-on-openstack/images/008-1_kz9Duiwop3_9J-eEmBhViA.webp" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;p&gt;To configure the MTU of calico, we have to edit the file &lt;strong&gt;inventory/jhub/group_vars/k8s-cluster/k8s-net-calico.yml&lt;/strong&gt; and set the &lt;strong&gt;calico_mtu&lt;/strong&gt; flag to &lt;strong&gt;1400&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;It’s important to notice that setting MTU had no effect for OpenStack versions earlier than Pike. The LoadBalancer didn’t work correctly.&lt;/p&gt;
&lt;p&gt;The last file to modify is &lt;strong&gt;inventory/jhub/group_vars/k8s-cluster/addons.yml&lt;/strong&gt;. JupyterHub uses Helm charts to deploy all that you need on the Kubernetes cluster and Kubespray can install Helm for you. So just set the &lt;strong&gt;helm_enabled&lt;/strong&gt; flag to &lt;strong&gt;true&lt;/strong&gt;.&lt;/p&gt;
&lt;p&gt;Now we can run ansible playbook&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;ansible-playbook --become -i inventory/jhub/hosts cluster.yml
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;You can take a coffee break because it takes time to install all the stuff. At the end of this process, you have a Kubernetes cluster up and running.&lt;/p&gt;
&lt;p&gt;To be sure, log in on the master nodes (the external address given by terraform) and enter the command&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;kubectl -n kube-system get pods
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;You should be able to see all pods of the kube-system namespace running.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2018/how-to-deploy-jupyterhub-with-kubernetes-on-openstack/images/009-1_iozJBhJS2Dh6CTEopZKzdA.webp" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;p&gt;The last step is to install the persistent volume from our GlusterFS.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;ansible-playbook --become -i inventory/jhub/hosts ./contrib/network-storage/glusterfs/glusterfs.yml
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;If you log in again to the master of your Kubernetes cluster and enter the following command&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;kubectl get pv
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;you will see your GlusterFS storage connected to your Kubernetes cluster.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2018/how-to-deploy-jupyterhub-with-kubernetes-on-openstack/images/010-1_qTxenvjzUpkvnt4SLCYyYQ.webp" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;h3 id="install-jupyterhub"&gt;Install JupyterHub&lt;/h3&gt;
&lt;p&gt;Now that you have a Kubernetes cluster running, the procedure to install JupyterHub is exactly the same as the one described in &lt;a href="https://zero-to-jupyterhub.readthedocs.io"&gt;Zero to JupyterHub&lt;/a&gt;. The only difference is that you don’t have to install Helm, since Kubespray did it for you. We’ll post the commands below, and you can go to the Zero to JupyterHub website for more information.&lt;/p&gt;
&lt;p&gt;The first step is to log in to the master node of your Kubernetes cluster. Then, initialize Helm.&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;helm init --service-account tiller
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="nx"&gt;kubectl&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;patch&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;deployment&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="nx"&gt;tiller&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;deploy&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;--&lt;/span&gt;&lt;span class="kn"&gt;namespace&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="nx"&gt;kube&lt;/span&gt;&lt;span class="o"&gt;-&lt;/span&gt;&lt;span class="nx"&gt;system&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;--&lt;/span&gt;&lt;span class="k"&gt;type&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="nx"&gt;json&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="o"&gt;--&lt;/span&gt;&lt;span class="nx"&gt;patch&lt;/span&gt;&lt;span class="p"&gt;=&lt;/span&gt;&lt;span class="err"&gt;&amp;#39;&lt;/span&gt;&lt;span class="p"&gt;[{&lt;/span&gt;&lt;span class="s"&gt;&amp;quot;op&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;&amp;quot;add&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;&amp;quot;path&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;&amp;quot;/spec/template/spec/containers/0/command&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;&amp;quot;value&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;:&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="p"&gt;[&lt;/span&gt;&lt;span class="s"&gt;&amp;quot;/tiller&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;,&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="s"&gt;&amp;quot;--listen=localhost:44134&amp;quot;&lt;/span&gt;&lt;span class="p"&gt;]}]&lt;/span&gt;&lt;span class="err"&gt;&amp;#39;&lt;/span&gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Next, follow the procedure described here&lt;/p&gt;
&lt;p&gt;&lt;a href="https://zero-to-jupyterhub.readthedocs.io/en/stable/setup-jupyterhub.html"&gt;Setting up JupyterHub - Zero to JupyterHub with Kubernetes 0.7.0 documentation&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;If the LoadBalancer did its job, you should be able to see the external IP to connect to your JupyterHub&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;kubectl -n jhub  get svc
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;If you enter this address in your web browser. If everything worked, you will see the JupyterHub login page:&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="JupyterHub login page" src="https://jasongrout.github.io/medium-archive/pelican/posts/2018/how-to-deploy-jupyterhub-with-kubernetes-on-openstack/images/011-1_X6gU8TYrMV9v-GarkSWvaQ.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;JupyterHub login page&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h3 id="wrapping-up-and-feedback"&gt;Wrapping up and feedback&lt;/h3&gt;
&lt;p&gt;The steps above described our attempts at running a JupyterHub on Kubernetes using OpenStack. There are likely many other ways to accomplish the same thing, and we’d love to hear feedback on the best procedure to install JupyterHub or BinderHub on OpenStack infrastructure. If you have encountered any issues, please leave a comment or ping us on the gitter channel of &lt;a href="https://gitter.im/jupyterhub/jupyterhub"&gt;JupyterHub&lt;/a&gt; or &lt;a href="https://gitter.im/binder-project/binder"&gt;Binder&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Thanks to the Project Jupyter team for their review and helpful comments and especially to Sylvain Corlay and Chris Holdgraf.&lt;/em&gt;&lt;/p&gt;
&lt;h3 id="about-the-authors-alphabetical-order"&gt;About the Authors (alphabetical order)&lt;/h3&gt;
&lt;ul&gt;
&lt;li&gt;&lt;strong&gt;David Delavennat&lt;/strong&gt;, Research Engineer in Scientific Infrastructures at CMLS (Polytechnique/CNRS) and INSMI (CNRS)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Loïc Gouarin&lt;/strong&gt;, Research Engineer in Scientific Computing at CMAP (Polytechnique/CNRS)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Guillaume Philippon&lt;/strong&gt;, Research Engineer in Scientific Infrastructures at LAL (IN2P3/CNRS)&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2018/how-to-deploy-jupyterhub-with-kubernetes-on-openstack/images/012-1_BvlZVRfsREg9GLQxZsvVCA.webp" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
</content><category term="JupyterHub"/><category term="Kubernetes"/></entry><entry><title>Introducing Jupyter Enterprise Gateway</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2018/introducing-jupyter-enterprise-gateway/" rel="alternate"/><published>2018-09-17T20:19:00+00:00</published><updated>2018-09-17T20:19:00+00:00</updated><author><name>Luciano Resende</name></author><id>tag:jasongrout.github.io,2018-09-17:/medium-archive/pelican/posts/2018/introducing-jupyter-enterprise-gateway/</id><summary type="html">&lt;p&gt;by Luciano Resende, Kevin Bates, Alan Chin&lt;/p&gt;
</summary><content type="html">&lt;p&gt;by &lt;a href="https://twitter.com/lresende1975"&gt;Luciano Resende&lt;/a&gt;, &lt;a href="https://twitter.com/kbates4"&gt;Kevin Bates&lt;/a&gt;, Alan Chin&lt;/p&gt;
&lt;p&gt;Yesterday, the Jupyter Steering Council voted to make Jupyter Enterprise Gateway a &lt;a href="https://github.com/jupyter/enhancement-proposals/blob/master/jupyter-enterprise-gateway-incorporation/jupyter-enterprise-gateway-incorporation.md"&gt;top-level Jupyter Project&lt;/a&gt;. I want to thank everyone for their contributions so far — code from my teammates at IBM and the community in general; advice from the Jupyter development team and mentors; and questions, issues, and requirements from end users.&lt;/p&gt;
&lt;p&gt;As we become an official Jupyter project, I would like to take the opportunity to give an update on the project’s progress during our incubation period.&lt;/p&gt;
&lt;h2 id="what-is-jupyter-enterprise-gateway"&gt;What is Jupyter Enterprise Gateway?&lt;/h2&gt;
&lt;p&gt;Jupyter Enterprise Gateway enables Jupyter Notebook to launch remote kernels in a distributed cluster, including Apache Spark managed by YARN, IBM Spectrum Conductor or Kubernetes.&lt;/p&gt;
&lt;p&gt;Although Enterprise Gateway is mostly kernel agnostic, it provides out of the box configuration examples for the following kernels:&lt;/p&gt;
&lt;p&gt;· Python using &lt;a href="https://ipython.org/"&gt;IPython&lt;/a&gt; kernel&lt;/p&gt;
&lt;p&gt;· R using &lt;a href="https://github.com/IRkernel/IRkernel"&gt;IRkernel&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;· Scala using &lt;a href="https://toree.incubator.apache.org/"&gt;Apache Toree&lt;/a&gt; kernel&lt;/p&gt;
&lt;p&gt;Jupyter Enterprise Gateway does not manage multiple Jupyter Notebook deployments, for that you should look for &lt;a href="https://github.com/jupyterhub/jupyterhub"&gt;JupyterHub&lt;/a&gt;. Having said that, Enterprise Gateway can enable &lt;a href="https://github.com/jupyterhub/jupyterhub"&gt;JupyterHub&lt;/a&gt; to launch remote kernels as individual Kubernetes pods, providing better resource allocation and enabling better environment management as each pod can be based on different images (e.g. TensorFlow, Anaconda, etc)&lt;/p&gt;
&lt;h2 id="supported-platforms"&gt;Supported Platforms&lt;/h2&gt;
&lt;p&gt;Jupyter Enterprise Gateway currently enables remote kernels in the following platforms:&lt;/p&gt;
&lt;h3 id="distributed-kernels-in-apache-spark"&gt;Distributed Kernels in Apache Spark&lt;/h3&gt;
&lt;p&gt;Jupyter Enterprise Gateway leverages different resource managers to enable distributed kernels in Apache Spark clusters. One example shown below describes kernels being launched in YARN cluster mode across all nodes of a cluster.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Jupyter Enterprise Gateway leverages Apache Spark resource managers to distribute kernels" src="https://jasongrout.github.io/medium-archive/pelican/posts/2018/introducing-jupyter-enterprise-gateway/images/001-1_oKl3bDSanz-SFqsgWsAPBw.mp4" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;&lt;em&gt;Jupyter Enterprise Gateway leverages Apache Spark resource managers to distribute kernels&lt;/em&gt;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Note that, Jupyter Enterprise Gateway also provides some other value-added capabilities such as enhanced security and multiuser support with user impersonation.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Jupyter Enterprise Gateway provides Enhanced Security and Multiuser support with user Impersonation" src="https://jasongrout.github.io/medium-archive/pelican/posts/2018/introducing-jupyter-enterprise-gateway/images/002-1_ihpHPqvgzXKepRAVZc7EIA.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Jupyter Enterprise Gateway provides Enhanced Security and Multiuser support with user Impersonation&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h3 id="distributed-kernels-in-kubernetes"&gt;Distributed Kernels in Kubernetes&lt;/h3&gt;
&lt;p&gt;Jupyter Enterprise Gateway support for Kubernetes enables decoupling the Jupyter Notebook Server and its kernels into multiple pods. This enables running Notebook server pods with minimally necessary resources based on the workload being processed.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Jupyter Enterprise Gateway enable remote kernels on Kubernetes cluster" src="https://jasongrout.github.io/medium-archive/pelican/posts/2018/introducing-jupyter-enterprise-gateway/images/003-1__R0tS0CZLy__LmL7o5b6vg.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Jupyter Enterprise Gateway enable remote kernels on Kubernetes cluster&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="jupyter-enterprise-gateway-and-jupyterhub"&gt;Jupyter Enterprise Gateway and JupyterHub&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://github.com/jupyterhub/jupyterhub"&gt;JupyterHub&lt;/a&gt; is a multi-user server that manages and proxies multiple instances of the single-user Jupyter notebook server. Particularly in a Kubernetes environment, Jupyter Enterprise Gateway can enable &lt;a href="https://github.com/jupyterhub/jupyterhub"&gt;JupyterHub&lt;/a&gt; to launch remote kernels as individual Kubernetes pods, providing better resource allocation and enabling better environment management as each pod can be based on different images (e.g. TensorFlow, Anaconda, etc). This has proven to be very desired, particularly when working on Deep Learning related Notebooks.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="JupyterHub and Jupyter Enterprise Gateway together in a Kubernetes cluster" src="https://jasongrout.github.io/medium-archive/pelican/posts/2018/introducing-jupyter-enterprise-gateway/images/004-1_9QJMPJLTZ04CFLc30aapTg.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;JupyterHub and Jupyter Enterprise Gateway together in a Kubernetes cluster&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="some-project-metrics"&gt;Some project metrics&lt;/h2&gt;
&lt;p&gt;The following stats have been collected from the Jupyter Enterprise Gateway GitHub repository &lt;strong&gt;during the incubation period&lt;/strong&gt;:&lt;/p&gt;
&lt;blockquote&gt;
&lt;ul&gt;
&lt;li&gt;10 releases&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;ul&gt;
&lt;li&gt;12 individual contributors&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;ul&gt;
&lt;li&gt;90 Stars&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;
&lt;blockquote&gt;
&lt;ul&gt;
&lt;li&gt;34 Forks&lt;/li&gt;
&lt;/ul&gt;
&lt;/blockquote&gt;
&lt;h2 id="source-code-documentation-and-other-community-resources"&gt;Source code, documentation, and other community resources&lt;/h2&gt;
&lt;p&gt;The Jupyter Enterprise Gateway community provides multiple resources that both users and contributors can use:&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Source Code available at GitHub&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://github.com/jupyter/enterprise_gateway"&gt;https://github.com/jupyter/enterprise_gateway&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Documentation available at ReadTheDocs&lt;/strong&gt;&lt;br&gt;
&lt;a href="http://jupyter-enterprise-gateway.readthedocs.io/en/latest/"&gt;http://jupyter-enterprise-gateway.readthedocs.io/en/latest/&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Automated builds available at Travis.CI&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://travis-ci.org/jupyter-incubator/enterprise_gateway"&gt;https://travis-ci.org/jupyter/enterprise_gateway&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Releases available at PyPi.org and Conda Forge&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://pypi.org/project/jupyter_enterprise_gateway/"&gt;https://pypi.org/project/jupyter_enterprise_gateway/&lt;/a&gt;&lt;br&gt;
&lt;a href="https://github.com/conda-forge/jupyter_enterprise_gateway-feedstock"&gt;https://github.com/conda-forge/jupyter_enterprise_gateway-feedstock&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;Related Docker Images available at Elyra organization at DockerHub&lt;/strong&gt;&lt;br&gt;
&lt;a href="https://hub.docker.com/u/elyra/dashboard/"&gt;https://hub.docker.com/u/elyra/dashboard/&lt;/a&gt;&lt;/p&gt;
&lt;h2 id="whats-next"&gt;What’s next?&lt;/h2&gt;
&lt;p&gt;We are eager to build an even greater community around the project, and tailor the project roadmap based on community advise.&lt;/p&gt;
&lt;p&gt;Currently, we are busy working on advancing our Kubernetes support and integration with JupyterHub.&lt;/p&gt;
&lt;p&gt;As always, we welcome questions, comments, and suggestions from users and the community in general.&lt;/p&gt;
</content><category term="kernels"/><category term="Kubernetes"/></entry><entry><title>Announcing the JupyterHub Helm Chart v0.5</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2017/announcing-the-jupyterhub-helm-chart-v0-5/" rel="alternate"/><published>2017-12-20T16:24:00+00:00</published><updated>2018-02-01T10:33:00+00:00</updated><author><name>Chris Holdgraf</name></author><id>tag:jasongrout.github.io,2017-12-20:/medium-archive/pelican/posts/2017/announcing-the-jupyterhub-helm-chart-v0-5/</id><summary type="html">&lt;p&gt;JupyterHub makes it possible to serve Jupyter instances to multiple users. The JupyterHub Helm Chart makes it possible to run this setup on…&lt;/p&gt;
</summary><content type="html">&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2017/announcing-the-jupyterhub-helm-chart-v0-5/images/001-1_el1BFE8ImVEackH13RyyWw.webp" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;p&gt;JupyterHub makes it possible to serve Jupyter instances to multiple users. The JupyterHub Helm Chart makes it possible to run this setup on &lt;a href="https://kubernetes.io/"&gt;kubernetes&lt;/a&gt;, making JupyterHub more scalable, stable, and flexible.&lt;/p&gt;
&lt;p&gt;We, the JupyterHub team, are proud to announce the next version of the JupyterHub Helm Chart: version 0.5. This post describes a bit of what’s new in this release. We’ve nicknamed the releases of the JupyterHub Helm Chart after famous cricketers, in this case world-class bowler &lt;a href="http://www.espncricinfo.com/afghanistan/content/player/311427.html"&gt;Hamid Hassan*&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;tl;dr: The release bumps JupyterHub to 0.8, adds better HTTPS support, and improves scalability to ~4,000 simultaneous users. See the &lt;a href="https://github.com/jupyterhub/zero-to-jupyterhub-k8s/blob/master/CHANGELOG.md"&gt;Helm Chart Changelog&lt;/a&gt; for more information.&lt;/p&gt;
&lt;h2 id="new-features"&gt;New Features&lt;/h2&gt;
&lt;p&gt;The following major features have been added to v0.5:&lt;/p&gt;
&lt;h3 id="jupyterhub-08"&gt;JupyterHub 0.8&lt;/h3&gt;
&lt;p&gt;Version 0.8 of JupyterHub was &lt;a href="/posts/2017/jupyterhub-0-8/"&gt;released earlier this year&lt;/a&gt;. It is full of new features, many of which directly benefit the Kubernetes deployment of JupyterHub. Below is a list of relevant points along with the relevant sections of the &lt;a href="https://zero-to-jupyterhub-with-kubernetes.readthedocs.io/en/latest/reference.html"&gt;Helm Chart configuration&lt;/a&gt;:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Lots of performance improvements. &lt;strong&gt;We now know we can handle up to 4k active users.&lt;/strong&gt;&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Limit the number of users who can try to launch the hub at once.&lt;/strong&gt; This can be tuned to avoid crashes when hundreds of users try to launch at the same time. It gives them a friendly error message and asks them to try later. See &lt;code&gt;hub.concurrentSpawnLimit&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Limit the number of simultaneous active users .&lt;/strong&gt; The Active Server limit can be used to limit the total number of active users that can use the hub at any given time. This allows admins to control the size of their clusters more effectively. See &lt;code&gt;hub.activeServerLimit&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Memory limits &amp;amp; guarantees&lt;/strong&gt; can now contain fractional units. So you can say &lt;code&gt;0.5G&lt;/code&gt; instead of having to use &lt;code&gt;512M&lt;/code&gt;.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;No more ‘too many redirects’ errors at scale.&lt;/strong&gt; This fixes an annoying race condition causing users to get stuck in a redirect loop when starting their servers.&lt;/li&gt;
&lt;/ol&gt;
&lt;h3 id="easier-https"&gt;Easier HTTPS&lt;/h3&gt;
&lt;p&gt;Version 0.5 of the helm chart makes it easier for admins to set up HTTPS for their users with &lt;a href="https://letsencrypt.org/"&gt;Let’s Encrypt&lt;/a&gt;. Users often access a JupyterHub instance from a public URL. To avoid nefarious behavior and increase security, using HTTPS is important. You can now choose to use Let’s Encrypt or a valid HTTPS certificate and key. You can also use your own HTTPS certificates &amp;amp; keys rather than using Let’s Encrypt. You can find &lt;a href="https://zero-to-jupyterhub.readthedocs.io/en/latest/security.html#setting-up-https"&gt;the new instructions here&lt;/a&gt;.&lt;/p&gt;
&lt;h3 id="more-authenticators"&gt;More authenticators&lt;/h3&gt;
&lt;p&gt;Authenticators allow you to control who has access to your JupyterHub. The following new authentication providers have been added in 0.5:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;a href="https://about.gitlab.com/"&gt;GitLab&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="http://www.cilogon.org/"&gt;CILogon&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;&lt;a href="https://www.globus.org/"&gt;Globus&lt;/a&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;You can now also set up a whitelist of usernames that have access to the hub (in addition to other authenticators in use). Do so by adding to the list in &lt;code&gt;auth.whitelist.users&lt;/code&gt;.&lt;/p&gt;
&lt;h3 id="hub-services-support"&gt;Hub Services support&lt;/h3&gt;
&lt;p&gt;Services let you connect your JupyerHub to other web services (for example, in &lt;a href="https://mybinder.org"&gt;mybinder.org&lt;/a&gt;). You can now add &lt;a href="https://jupyterhub.readthedocs.io/en/latest/reference/services.html"&gt;external JupyterHub Services&lt;/a&gt; by adding them to &lt;code&gt;hub.services&lt;/code&gt;. Note that you are still responsible for actually running the service somewhere (perhaps as a deployment object in Kubernetes).&lt;/p&gt;
&lt;h3 id="more-customization-with-jupyterhub_configpy"&gt;More customization with &lt;code&gt;jupyterhub_config.py&lt;/code&gt;&lt;/h3&gt;
&lt;p&gt;Sometimes it is useful to be able to run arbitrary extra code when setting up your deployment. You can put extra snippets of &lt;code&gt;jupyterhub_config.py&lt;/code&gt; configuration in &lt;code&gt;hub.extraConfig&lt;/code&gt;. Now you can also add &lt;a href="http://zero-to-jupyterhub.readthedocs.io/en/latest/advanced.html#hub-extraenv"&gt;extra environment variables&lt;/a&gt; to the hub in &lt;code&gt;hub.extraEnv&lt;/code&gt; and &lt;a href="http://zero-to-jupyterhub.readthedocs.io/en/latest/advanced.html#hub-extraconfigmap"&gt;extra configmap items&lt;/a&gt; via &lt;code&gt;hub.extraConfigMap&lt;/code&gt;. This makes it cleaner to customize the hub’s configuration in ways that are not yet possible with &lt;code&gt;config.yaml&lt;/code&gt;. You can find more information in the &lt;a href="http://zero-to-jupyterhub.readthedocs.io/en/latest/advanced.html#arbitrary-code-in-jupyterhub-config-py"&gt;documentation&lt;/a&gt;.&lt;/p&gt;
&lt;h3 id="more-customization-options-for-user-server-environments"&gt;More customization options for user server environments&lt;/h3&gt;
&lt;p&gt;More options have been added under &lt;code&gt;singleuser&lt;/code&gt; to help you customize the environment that the user session is spawned in. You can…&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;Change the uid / gid of the user with &lt;code&gt;singleuser.uid&lt;/code&gt; and &lt;code&gt;singleuser.fsGid&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Mount extra volumes with &lt;code&gt;singleuser.storage.extraVolumes&lt;/code&gt; &amp;amp; &lt;code&gt;singleuser.storage.extraVolumeMounts&lt;/code&gt;&lt;/li&gt;
&lt;li&gt;Provide extra environment variables with &lt;code&gt;singleuser.extraEnv&lt;/code&gt;.&lt;/li&gt;
&lt;/ul&gt;
&lt;h2 id="more-information"&gt;More information&lt;/h2&gt;
&lt;p&gt;For more information about the JupyterHub Helm Chart, and the JupyterHub ecosystem more broadly, see the following links:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;The &lt;a href="https://jupyterhub.readthedocs.io/en/latest/"&gt;JupyterHub documentation&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;The &lt;a href="https://github.com/jupyterhub/jupyterhub"&gt;JupyterHub development repository&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;The &lt;a href="https://github.com/jupyterhub/zero-to-jupyterhub-k8s"&gt;JupyterHub Helm Chart development repository&lt;/a&gt;&lt;/li&gt;
&lt;li&gt;The &lt;a href="http://zero-to-jupyterhub.readthedocs.io"&gt;Zero to JupyterHub guide&lt;/a&gt; to deploying JupyterHub on Kubernetes&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;*&lt;em&gt;&lt;strong&gt;Hamid Hassan&lt;/strong&gt;&lt;/em&gt; &lt;em&gt;is a fast bowler who currently plays for the Afghanistan National Cricket Team. With nicknames ranging from&lt;/em&gt; &lt;a href="https://www.rferl.org/a/interview-afghan-cricketer-living-the-dream/24752618.html"&gt;&lt;em&gt;“Afghanistan’s David Beckham”&lt;/em&gt;&lt;/a&gt; &lt;em&gt;to&lt;/em&gt; &lt;a href="http://www.nzherald.co.nz/nz/news/article.cfm?c_id=1&amp;amp;objectid=11413633"&gt;&lt;em&gt;“Rambo”&lt;/em&gt;&lt;/a&gt;&lt;em&gt;, he is considered by many to be Afghanistan’s first Cricket Superhero. Currently known for fast (145km/h+) deliveries, cartwheeling celebrations, war painted face and having had to flee Afghanistan as a child to escape from war. He&lt;/em&gt; &lt;a href="http://www.nzherald.co.nz/nz/news/article.cfm?c_id=1&amp;amp;objectid=11413633"&gt;&lt;em&gt;says&lt;/em&gt;&lt;/a&gt; &lt;em&gt;he plays because “We are ambassadors for our country and we want to show the world that Afghanistan is not like people recognize it by terrorists and these things. We want them to know that we have a lot of talent as well.”&lt;/em&gt;&lt;/p&gt;
</content><category term="JupyterHub"/><category term="Kubernetes"/></entry><entry><title>Binder 2.0, a Tech Guide</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2017/binder-2-0-a-tech-guide-2017/" rel="alternate"/><published>2017-11-30T16:01:00+00:00</published><updated>2017-11-30T19:25:00+00:00</updated><author><name>Chris Holdgraf</name></author><id>tag:jasongrout.github.io,2017-11-30:/medium-archive/pelican/posts/2017/binder-2-0-a-tech-guide-2017/</id><summary type="html">&lt;p&gt;Authors: The Binder project is comprised of many individuals within and outside of the core Jupyter team. A list of members that…&lt;/p&gt;
</summary><content type="html">&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2017/binder-2-0-a-tech-guide-2017/images/001-1_cWQj_YdmY_p14eh628N_Kg.webp" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Authors: The Binder project is comprised of many individuals within and outside of the core Jupyter team. A list of members that contributed to this post is at the end of this article.&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;em&gt;Note: this post focuses more on technical changes in the Binder 2.0 reboot. For a post about user-facing features and future plans, see&lt;/em&gt; &lt;a href="https://elifesciences.org/labs/8653a61d"&gt;&lt;em&gt;this eLife blog post&lt;/em&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;We are undergoing a dramatic increase in the complexity of techniques for analyzing data, doing scientific research, and sharing our work with others. In early 2016, the &lt;a href="https://mybinder.org"&gt;Binder project&lt;/a&gt; was announced, attempting to connect these three components. A &lt;a href="https://elifesciences.org/labs/a7d53a88/toward-publishing-reproducible-computation-with-binder"&gt;blogpost in eLife&lt;/a&gt; described a vision where scientists could specify dependencies along with a collection of Jupyter notebooks. Binder builds a Docker image from these dependencies, and provides a URL where any user in the world can instantly recreate this environment.&lt;/p&gt;
&lt;p&gt;Want to see it in action? Click the button below.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://mybinder.org/v2/gh/wildtreetech/explore-open-data/binder20-elife?filepath=bikes-per-week.ipynb"&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2017/binder-2-0-a-tech-guide-2017/images/002-1_IA00K8fa8FvXedoBBDh2fg.webp" alt="" loading="lazy" data-body-image=""&gt;&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;With this post we are proud to announce the next version of Binder. It aims to be more modular, more flexible, more stable, faster, and more extensible than its predecessor. Powering this version of Binder is a collection of tools in the Jupyter ecosystem. Since being released, the Binder project has learned many things about implementing fast, flexible online deployments. In addition, its vision has expanded to include not only Jupyter notebooks, but many other computational workflows. You can access an open beta version of this deployment here:&lt;/p&gt;
&lt;p&gt;&lt;a href="https://mybinder.org"&gt;mybinder.org&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You can find a list of sample repositories to learn how to create “Binder”-ready repositories here:&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/binder-examples"&gt;github.com/binder-examples&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;You can also see what the Binder community has been up to in creating their own repositories by checking the GitHub Binder topic:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;https://github.com/topics/binder
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;Give it a shot, build some repositories, and importantly, tell us what could be improved on our &lt;a href="https://github.com/jupyterhub/binderhub"&gt;GitHub repo&lt;/a&gt;. Below we’ll describe a bit about what’s new.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="The new Binder UI. Users input a URL to a git repository (or specify a specific branch/tag/commit). Upon clicking “launch”, you will be directed to a live environment where you can interact with the code." src="https://jasongrout.github.io/medium-archive/pelican/posts/2017/binder-2-0-a-tech-guide-2017/images/003-1_lWcoBaRvNzXxzGPqV_3vew.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;The new Binder UI. Users input a URL to a git repository (or specify a specific branch/tag/commit). Upon clicking “launch”, you will be directed to a live environment where you can interact with the code.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="whats-new"&gt;What’s new?&lt;/h2&gt;
&lt;p&gt;First off we’ll describe how the experience will change for users. The short answer is: not much. The goal of Binder is still enabling you to intantly create interactive and shareable repositories. We’ve completely rebuilt the backend of Binder, but we’ve made minimal changes to the front-end user experience.&lt;/p&gt;
&lt;p&gt;The biggest difference you should notice is that Binder is both faster and more stable. You’ll still be able to generate Binder links from a git URL from a single web-page. However, there are a few key differences:&lt;/p&gt;
&lt;h3 id="new-default-environment"&gt;New Default Environment&lt;/h3&gt;
&lt;p&gt;Old versions of Binder were based off of a Docker image that contained a fairly heavy computational environment. The new Binder deployment makes minimal assumptions about what environment you want installed, by default the only thing that will be installed is the Jupyter Notebook and Python 3. This means you’ll need to be more expressive in the dependencies you include in your dependency files. For example, if you want &lt;code&gt;numpy&lt;/code&gt; or &lt;code&gt;matplotlib&lt;/code&gt;, you should specify them in a &lt;code&gt;requirements.txt&lt;/code&gt; or &lt;code&gt;environment.yml&lt;/code&gt; file. Since you’re explicitly listing your requirements it makes reproducing your work more reliable, and allows the Binder infrastructure to change more freely without breaking your repository code.&lt;/p&gt;
&lt;h3 id="new-url-structure"&gt;New URL structure&lt;/h3&gt;
&lt;p&gt;The new URL structure for Binder follows the following convention:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;https://mybinder.org/v2/gh/&amp;lt;org-name&amp;gt;/&amp;lt;repo-name&amp;gt;/&amp;lt;branch|tag|hash-name&amp;gt;?filepath=&amp;lt;path-to-file&amp;gt;
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;For example, below is the URL for a basic Binder-ready Python 3 repository, it includes basic information about the repository:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;https://mybinder.org/v2/gh/binder-examples/requirements/master
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;You can also specify parameters that do things like point users to a particular file or initialize a user-interface. For example, the following URL starts JupyterLab once users click the link:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;https://mybinder.org/v2/gh/binder-examples/jupyterlab/master?urlpath=lab
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;In each case, note the &lt;code&gt;gh&lt;/code&gt; at the very beginning — this specifies that the git URL exists on github.com. It is possible to build new URL parsers for other online repositories such as BitBucket, osf.io, or any other provider. We are currently focusing on Git and GitHub, but nothing prevents Binder from being compatible with other kinds of content providers. Shortly before this blog post we added support for arbitrary git URLs to binderhub, so watch this space.&lt;/p&gt;
&lt;h3 id="specify-a-specific-branch-tag-commit"&gt;Specify a specific branch / tag / commit&lt;/h3&gt;
&lt;p&gt;Also notice that in the URL above you can specify a branch, tag, or commit hash for the Binder image. This allows you to ensure that a Binder image will &lt;strong&gt;always&lt;/strong&gt; remain the same (if you specify package versions properly). This is a crucial step for reproducibility and maintaining consistency in how users experience the files in your repository.&lt;/p&gt;
&lt;p&gt;Notice that in the URL above you can specify a commit hash or git tag for the Binder image. This hash is unique to the state of the code at the moment that the commit was made, ensuring that Binder can rebuild the exact same environment any time. Note that if authors don’t want to guarantee the same reproducible environment, they can specify a branch and BinderHub will resolve it to the latest commit hash before building the environment.&lt;/p&gt;
&lt;h3 id="binder-auto-building"&gt;Binder auto-building&lt;/h3&gt;
&lt;p&gt;When a git repository is launched, Binder will now check whether an image has already been built for that repository at the same commit hash. If it has, then Binder will skip the building process and take you straight to a JupyterHub instance that serves this image.&lt;/p&gt;
&lt;p&gt;If the image hasn’t been built, then it will automatically be generated before sending the user to JupyterHub. The only difference will be the amount of time it takes before entering the JupyterHub environment. This means that authors no longer need to explicitly build their Binder images when they update a branch. The next time someone clicks a Binder link, it will happen automatically. If you don’t want this behavior, be sure to point Binder to a specific tag or commit hash, rather than a branch name or tag. For example, here’s a Binder URL that will always point to the same commit hash:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;https://mybinder.org/v2/gh/wildtreetech/explore-open-data/binder20-elife
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;While this URL points to a branch, and will thus be re-built each time a new commit is made to that branch:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;https://mybinder.org/v2/gh/wildtreetech/explore-open-data/master
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;h3 id="more-options-for-dependency-files"&gt;More options for dependency files&lt;/h3&gt;
&lt;p&gt;Users often want to specify a computational environment that is more complex than a simple list of Python requirements. While this is possible by specifying a Dockerfile, it’s often an overly-complicated solution to this problem. Binder now uses &lt;a href="https://github.com/jupyter/repo2docker"&gt;repo2docker&lt;/a&gt; to build a Docker image from your repository. This makes it possible to specify a more complex environment with text files. For example, you can use an &lt;code&gt;apt.txt&lt;/code&gt; file to install packages with &lt;code&gt;apt-get&lt;/code&gt;, or use a file called &lt;code&gt;postBuild&lt;/code&gt; to define shell commands that are run before generating the Docker image (e.g. for downloading some data or running scripts). See the &lt;a href="https://repo2docker.readthedocs.io/en/latest/samples.html"&gt;repo2docker documentation&lt;/a&gt; for a list of files that are supported with Binder.&lt;/p&gt;
&lt;p&gt;For a selection of examples that show off how to specify dependencies take a look at the example gallery: &lt;a href="https://github.com/binder-examples"&gt;https://github.com/binder-examples&lt;/a&gt;&lt;/p&gt;
&lt;h3 id="more-user-interfaces"&gt;More user interfaces&lt;/h3&gt;
&lt;figure&gt;
&lt;img alt="The JupyterLab interface running on Binder. You can access the JupyterLab demo repository at mybinder.org/v2/gh/binder-examples/jupyterlab/master?urlpath=lab" src="https://jasongrout.github.io/medium-archive/pelican/posts/2017/binder-2-0-a-tech-guide-2017/images/004-1_TW7Gnwl-02cejzgs1nnX2Q.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;The JupyterLab interface running on Binder. You can access the JupyterLab demo repository at ``mybinder&lt;code&gt;.org/v2/gh/binder-examples/jupyterlab/master?urlpath=lab&lt;/code&gt;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The previous iteration of Binder only supported the classic Jupyter Notebook user interface, while the new deployment will additionally support &lt;a href="https://github.com/binder-examples/dockerfile-rstudio"&gt;RStudio&lt;/a&gt; and &lt;a href="https://github.com/binder-examples/jupyterlab"&gt;JupyterLab&lt;/a&gt;. Because of the extra build configuration files specified above, you can also utilize more tools in the Jupyter widgets ecosystem, such as the &lt;a href="https://github.com/binder-examples/jupyter-rise"&gt;RISE plugin for interactive presentations&lt;/a&gt; or the &lt;a href="https://github.com/oschuett/appmode"&gt;appmode plugin&lt;/a&gt; to generate interactive apps from your repository. We also welcome contributions to add support for other user interfaces.&lt;/p&gt;
&lt;h3 id="more-online-repository-providers"&gt;More online repository providers&lt;/h3&gt;
&lt;p&gt;While GitHub is a fantastic repository of open-source code, it’s not the only repository. The new Binder deployment makes it straightforward to adding support for new sources of code (for example, GitLab, the OSF, or even non-git codebases). Currently GitHub is the only supported source for code, but we welcome contributions enabling support for new sources.&lt;/p&gt;
&lt;p&gt;We’re excited about this next step in Binder’s development, and hopeful that we can build a community around this powerful set of tools. Don’t hesitate to open an issue or pull request on our &lt;a href="https://github.com/jupyterhub/binderhub"&gt;GitHub repository&lt;/a&gt;, or to reach out via &lt;a href="https://gitter.im/jupyterhub/binder"&gt;our Gitter channel&lt;/a&gt;. We look forward to seeing what comes next, and to continue enabling reproducible and open workflows in data science and research.&lt;/p&gt;
&lt;h2 id="for-developers"&gt;For developers&lt;/h2&gt;
&lt;p&gt;The next few sections are meant for developers interested in deploying their own Binder, or for those interested in the technical details behind the new deployment.&lt;/p&gt;
&lt;h3 id="tech-components"&gt;Tech components&lt;/h3&gt;
&lt;p&gt;The three main technical components behind the new Binder backend are:&lt;/p&gt;
&lt;ul&gt;
&lt;li&gt;&lt;a href="https://binderhub.readthedocs.io/"&gt;BinderHub&lt;/a&gt;, currently on display at &lt;a href="https://mybinder.org"&gt;mybinder.org&lt;/a&gt; and contained in the &lt;a href="https://github.com/jupyterhub/binderhub"&gt;binderhub repository&lt;/a&gt;.&lt;/li&gt;
&lt;li&gt;&lt;a href="https://repo2docker.readthedocs.io/"&gt;repo2docker&lt;/a&gt;, a tool that converts a code repository into a Docker image with an environment specified via dependency files (e.g., &lt;code&gt;requirements.txt&lt;/code&gt;).&lt;/li&gt;
&lt;li&gt;&lt;a href="https://jupyterhub.readthedocs.io/en/latest/"&gt;JupyterHub&lt;/a&gt;, which hosts user instances with a server in the cloud. We use a distribution of JupyterHub that runs on top of Kubernetes.&lt;/li&gt;
&lt;/ul&gt;
&lt;p&gt;For more information on the infrastructure behind Binder, &lt;a href="https://binderhub.readthedocs.io/en/latest/"&gt;see the documentation&lt;/a&gt;.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="A prototype of RStudio running in a Binder. This is currently support with a Dockerfile, and we are working on supporting R build files natively. You can access this repository at:" src="https://jasongrout.github.io/medium-archive/pelican/posts/2017/binder-2-0-a-tech-guide-2017/images/005-1_EgMk1PYMl6ouIP_XQ5G5Fg.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;A prototype of RStudio running in a Binder. This is currently support with a Dockerfile, and we are working on supporting R build files natively. You can access this repository at: &lt;code&gt;&amp;lt;https://mybinder.org/v2/gh/binder-examples/dockerfile-rstudio/master&amp;gt;&lt;/code&gt;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h3 id="kubernetes"&gt;Kubernetes&lt;/h3&gt;
&lt;p&gt;Binder now also heavily relies on &lt;a href="https://kubernetes.io/"&gt;Kubernetes&lt;/a&gt; for scaling our image building service and the JupyterHub. Kubernetes is massively scalable and has a strong community of developers behind it. Moreover, Kubernetes is cloud-agnostic. It can be run on Google Cloud, Microsoft Azure, and AWS among others, as well as on your own bare metal hardware if needed. Because BinderHub is built to run on top of Kubernetes, you can deploy Binder off of any of these resources as well (see below).&lt;/p&gt;
&lt;h3 id="deploying-your-own-binder-server"&gt;Deploying your own Binder server&lt;/h3&gt;
&lt;p&gt;While mybinder.org will continue to exist as a public service, we hope to see new Binder deployments for many different use cases in the wild. One of our primary goals is to make it easier for users to deploy their own Binder servers. This is relatively straightforward by following the instructions on the &lt;a href="https://binderhub.readthedocs.io/en/latest/"&gt;BinderHub documentation&lt;/a&gt;, which are currently under active development to make ongoing improvements as the Kubernetes technology evolves. We’re continuously updating these steps to make them as clear as possible, so please don’t hesitate to open an issue or a pull request on our &lt;a href="https://github.com/jupyterhub/binderhub"&gt;github repository&lt;/a&gt; and make suggestions.&lt;/p&gt;
&lt;p&gt;We would love to see others deploy their own BinderHub servers, either for their own communities, or as part of a federated public service of BinderHubs.&lt;/p&gt;
&lt;h2 id="future-development"&gt;Future development&lt;/h2&gt;
&lt;p&gt;This is the just the beginning of new features and improvements to Binder. We’re working hard to grow an open-source community around these tools, and we encourage issues, comments, and PRs on the &lt;a href="https://github.com/jupyterhub/binderhub"&gt;BinderHub&lt;/a&gt;, &lt;a href="https://github.com/jupyter/repo2docker"&gt;repo2docker&lt;/a&gt;, and &lt;a href="https://github.com/jupyterhub/jupyterhub"&gt;JupyterHub&lt;/a&gt; repositories. We look forward to growing the Binder ecosystem, and we’re excited to see all of the Binders that people design.&lt;/p&gt;
&lt;h2 id="acknowledgements-alphabetical-order"&gt;Acknowledgements (alphabetical order)&lt;/h2&gt;
&lt;p&gt;&lt;em&gt;C. Titus Brown (UC Davis), Matthias Bussonnier (UC Berkeley), Jessica Forde (UC Berkeley), Brian Granger (Cal Poly), Tim Head (Wild Tree Tech), Chris Holdgraf (UC Berkeley), Andrew Osheroff (Google), Naomi Penfold (eLife Sciences), M Pacer (UC Berkeley), Yuvi Panda (UC Berkeley), Fernando Perez (UC Berkeley), Min Ragan-Kelley (Simula Research Laboratory), and Carol Willing (Cal Poly). The Binder project is currently being funded from a grant from the Moore Foundation.&lt;/em&gt;&lt;/p&gt;
</content><category term="Binder"/><category term="Docker"/><category term="Kubernetes"/><category term="open science"/></entry><entry><title>Introducing the Helm Chart for JupyterHub Deployment with Kubernetes</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2017/introducing-the-helm-chart-for-jupyterhub-deployment/" rel="alternate"/><published>2017-06-29T23:21:00+00:00</published><updated>2017-08-28T18:20:00+00:00</updated><author><name>Project Jupyter</name></author><id>tag:jasongrout.github.io,2017-06-29:/medium-archive/pelican/posts/2017/introducing-the-helm-chart-for-jupyterhub-deployment/</id><summary type="html">&lt;p&gt;The JupyterHub team proudly announces the release of a helm chart for deploying JupyterHub on Kubernetes clusters. We’ve …&lt;/p&gt;
</summary><content type="html">&lt;p&gt;The JupyterHub team proudly announces the release of a &lt;a href="https://github.com/jupyterhub/helm-chart"&gt;helm chart&lt;/a&gt; for deploying JupyterHub on Kubernetes clusters. We’ve designed the JupyterHub helm chart to save you time in creating JupyterHub deployments.&lt;/p&gt;
&lt;p&gt;You can find a repository with the helm chart &lt;a href="https://github.com/jupyterhub/helm-chart"&gt;here&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;&lt;strong&gt;This is a pre-release version of the helm-chart.&lt;/strong&gt; It will likely change in a breaking fashion sometime in the future, though we will make every effort to minimize this as much as possible as modifications are made for new features and increased stability. If you have any questions or comments, reach out to us on &lt;a href="https://gitter.im/jupyterhub/jupyterhub"&gt;Gitter&lt;/a&gt; or &lt;a href="https://github.com/jupyterhub/helm-chart/issues"&gt;open an issue&lt;/a&gt;. For help with deploying your own JupyterHub instance, see our &lt;a href="https://zero-to-jupyterhub.readthedocs.io/en/latest/"&gt;guide for setting up JupyterHub&lt;/a&gt; (which uses this helm chart). Or, come find us at the &lt;a href="https://conferences.oreilly.com/jupyter/jup-ny/public/schedule/detail/60074"&gt;JupyterHub talk&lt;/a&gt; at JupyterCon.&lt;/p&gt;
&lt;p&gt;The following principles guided development:&lt;/p&gt;
&lt;h3 id="easy-to-administer"&gt;Easy to administer&lt;/h3&gt;
&lt;p&gt;The nitty gritty of JupyterHub setup and management can be cumbersome, complicated, and time-consuming. With the JupyterHub helm chart, you will spend less time debugging your setup, and more time deploying, customizing to your needs, and successfully running your JupyterHub. Within a cloud computing infrastructure, using the helm chart typically requires only one or two commands to get started.&lt;/p&gt;
&lt;h3 id="open-source"&gt;Open source&lt;/h3&gt;
&lt;p&gt;The JupyterHub helm chart uses applications and codebases that are open and thriving. We prefer, prioritize, and select tools that have a history of stability and development, and which adhere to open-source principles when it comes to project and community growth. As a result, you’ll be able to easily connect with the many tools available in the open-source community, and you’ll have flexibility in where you deploy JupyterHub.&lt;/p&gt;
&lt;h3 id="cloud-agnostic"&gt;Cloud agnostic&lt;/h3&gt;
&lt;p&gt;We’ve made an effort to keep our helm chart as cloud-agnostic as possible. The only requirement is that your computing provider supports the Kubernetes infrastructure, an open-source platform that is widely available across many different online providers. You can run JupyterHub on cloud services such as Google Cloud, Microsoft Azure, and Amazon EC2, and even on your own hardware or institution-specific setup.&lt;/p&gt;
&lt;h3 id="scalable"&gt;Scalable&lt;/h3&gt;
&lt;p&gt;We’ve taken great care to develop the JupyterHub helm chart with the ability to be used in a variety of work, research, and education settings. Some JupyterHub deployments have a dynamic userbase that works in spurts of activity. Others have users that have long periods of inactivity. Rather than capping the amount of resources available for users, the JupyterHub helm chart utilizes Kubernetes to scale computational resources up (or down) as needed. This means that large changes in user behavior don’t result in system-wide instability or slowdown issues. It also means that you can quickly update user hardware, push new files to user disks, and alter the environment in which users are operating.&lt;/p&gt;
&lt;p&gt;JupyterHub has been deployed in a variety of places, including at least one class with nearly 1500 students. We’ve made sure that it can handle large groups of users, and we’re excited to see people push the limit even further.&lt;/p&gt;
&lt;h3 id="support"&gt;Support&lt;/h3&gt;
&lt;p&gt;If you have any questions or comments, reach out to us on &lt;a href="https://gitter.im/jupyterhub/jupyterhub"&gt;Gitter&lt;/a&gt; or &lt;a href="https://github.com/jupyterhub/helm-chart/issues"&gt;open an issue&lt;/a&gt;. For help with deploying your own JupyterHub instance, see our guide, &lt;a href="https://zero-to-jupyterhub.readthedocs.io/en/latest/"&gt;Zero to JupyterHub&lt;/a&gt;. Or, come find us at the &lt;a href="https://conferences.oreilly.com/jupyter/jup-ny/public/schedule/detail/60074"&gt;JupyterHub talk&lt;/a&gt; at JupyterCon.&lt;/p&gt;
&lt;h3 id="acknowledgements"&gt;Acknowledgements&lt;/h3&gt;
&lt;p&gt;JupyterHub and this helm chart wouldn’t have been possible without the goodwill, time, and funding from a lot of different people. In particular, we want to thank the Gordon and Betty Moore Foundation, the Sloan Foundation, the Helmsley Charitable Trust, the Berkeley Data Science Education Program, and the Wikimedia Foundation for supporting various members of our team. We also want to thank the individuals of the JupyterHub team (listed below), the Project Jupyter community, and our more than 100 contributors for continuing to grow and improve this technology.&lt;/p&gt;
&lt;p&gt;Sincerely,&lt;br&gt;
&lt;em&gt;The JupyterHub Team (in alphabetical order)&lt;/em&gt;&lt;/p&gt;
&lt;p&gt;&lt;a href="https://www.google.com/url?q=http://data.berkeley.edu/&amp;amp;sa=D&amp;amp;ust=1498755372168000&amp;amp;usg=AFQjCNFrmMhaSh1FrvBDYK6EYUZsWrK4vQ"&gt;Berkeley Data Science Education Program&lt;/a&gt; (Gunjan Baid, Sam Lau, Ryan Lovett, Yuvi Panda, Vinitra Swamy)&lt;/p&gt;
&lt;p&gt;&lt;a href="https://www.google.com/url?q=https://www.calblueprint.org/&amp;amp;sa=D&amp;amp;ust=1498755372168000&amp;amp;usg=AFQjCNFqs3qDHccsrEcZoO1MnXMkdcmGHw"&gt;Cal Blueprint&lt;/a&gt; Team (Jiefu Gong, Sam Lau, Derrick Mar, Peter Veerman, Tony Yang)&lt;/p&gt;
&lt;p&gt;&lt;a href="http://jupyter.org/"&gt;Project Jupyter&lt;/a&gt; (Chris Holdgraf, Yuvi Panda, Min Ragan-Kelley, Carol Willing)&lt;/p&gt;
</content><category term="JupyterHub"/><category term="Kubernetes"/></entry></feed>