<?xml version="1.0" encoding="utf-8"?>
<feed xmlns="http://www.w3.org/2005/Atom"><title>Jupyter Blog - yuvipanda</title><link href="https://jasongrout.github.io/medium-archive/pelican/" rel="alternate"/><link href="https://jasongrout.github.io/medium-archive/pelican/feeds/author-yuvipanda.atom.xml" rel="self"/><id>https://jasongrout.github.io/medium-archive/pelican/</id><updated>2026-04-08T23:34:00+00:00</updated><subtitle>The Project Jupyter blog: news, releases, and community stories, archived from blog.jupyter.org.</subtitle><entry><title>Berkeley Institute for Data Science (BIDS) joins the mybinder.org</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2026/berkeley-institute-for-data-science-bids-joins-the/" rel="alternate"/><published>2026-04-08T23:34:00+00:00</published><updated>2026-04-08T23:34:00+00:00</updated><author><name>yuvipanda</name></author><id>tag:jasongrout.github.io,2026-04-08:/medium-archive/pelican/posts/2026/berkeley-institute-for-data-science-bids-joins-the/</id><summary type="html">&lt;p&gt;The Berkeley Institute for Data Science (BIDS) is now a part of the mybinder.org federation!&lt;/p&gt;
</summary><content type="html">&lt;h2 id="berkeley-institute-for-data-science-bids-joins-the-mybinderorg-federation-in-partnership-with-2i2c"&gt;Berkeley Institute for Data Science (BIDS) joins the mybinder.org federation in partnership with 2i2c&lt;/h2&gt;
&lt;p&gt;The &lt;a href="https://bids.berkeley.edu/"&gt;Berkeley Institute for Data Science (BIDS)&lt;/a&gt; is the birthplace of the current iteration of &lt;a href="https://mybinder.org"&gt;mybinder.org&lt;/a&gt;, all the way back in 2017. In 2026, they are back as a member of the &lt;a href="https://mybinder.readthedocs.io/en/latest/about/federation.html"&gt;federation&lt;/a&gt;, joining &lt;a href="https://2i2c.org"&gt;2i2c&lt;/a&gt; and &lt;a href="https://www.gesis.org/en/home"&gt;GESIS&lt;/a&gt; in contributing to the cloud costs of keeping mybinder.org running!&lt;/p&gt;
&lt;p&gt;The BIDS node is running on the &lt;a href="https://us.ovhcloud.com/"&gt;OVH Cloud Provider&lt;/a&gt;, joining the two existing nodes (from 2i2c and GESIS) running on &lt;a href="https://www.hetzner.com/"&gt;Hetzner&lt;/a&gt;. By running on a different cloud provider than our existing nodes, we reduce the risk that a cloud provider outage or policy change will take down the entire mybinder.org service. This immediately came into play, as Hetzner is having &lt;a href="https://github.com/jupyterhub/mybinder.org-deploy/issues/3686"&gt;a lot of trouble with its object storage service&lt;/a&gt;, increasing failure rates on mybinder.org. We were able to &lt;a href="https://github.com/jupyterhub/mybinder.org-deploy/pull/3704"&gt;shift more of our users&lt;/a&gt; to the BIDS node seamlessly, reducing disruptions for our end users.&lt;/p&gt;
&lt;p&gt;The BIDS node thus temporarily served 50% of our traffic, taking a higher than usual share of the load as we wait for error rates on Hetzner object store to get better. It has served more than 64,000 users so far in the short time it’s been up, and expected to do more.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Graph with three lines, representing the three current mybinder.org federation members: 2i2c, BIDS and GESIS. Shows data from Dec 2025 — end of March 2026. BIDS starts roughly in mid 2025, and ramps up to roughly over 1000 launches per day. GESIS handles roughly 3000, 2i2c handles roughly 2000. The graph is periodic weekly, as usage drops during the weekends" src="https://jasongrout.github.io/medium-archive/pelican/posts/2026/berkeley-institute-for-data-science-bids-joins-the/images/001-1_7pRDcpGYyl1R6pBWsIC5UQ.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Graph of mybinder.org sessions across our federation members over the last 4 months, showing the BIDS member taking on more of the load over time as it is brought online. The graph is jagged due to dips in demand over the weekend&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This node was brought up in partnership with &lt;a href="https://2i2c.org"&gt;2i2c&lt;/a&gt;. 2i2c contributed expertise in extending the mybinder.org deployment infrastructure to run on OVH, as well as working with UC Berkeley procurement to ensure that the cloud costs can be paid for on a continuous basis. Many thanks to the ever awesome &lt;a href="http://github.com/minrk/"&gt;MinRK&lt;/a&gt; (who is now &lt;a href="https://bids.berkeley.edu/news/min-ragan-kelley-and-his-journey-back-bids"&gt;back to working at BIDS&lt;/a&gt;!) for a lot of the technical work in taking this through!&lt;/p&gt;
&lt;p&gt;Want your organization to join the mybinder.org federation, materially making a difference in resilience and availability of the service, even at just a few hundred dollars a month? Come talk to us &lt;a href="https://jupyter.zulipchat.com"&gt;on the Jupyter Zulip&lt;/a&gt;!&lt;/p&gt;
</content><category term="Binder"/></entry><entry><title>Scaling “Maintainer Intuition” with Pull Request Triage Boards</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2025/scaling-maintainer-intuition-with-pull-request-triage/" rel="alternate"/><published>2025-11-16T21:22:00+00:00</published><updated>2025-11-16T21:22:00+00:00</updated><author><name>yuvipanda</name></author><id>tag:jasongrout.github.io,2025-11-16:/medium-archive/pelican/posts/2025/scaling-maintainer-intuition-with-pull-request-triage/</id><summary type="html">&lt;p&gt;When I helped start 2i2c.org, one of my goals for the non-profit was to experiment with new and different ways of supporting the Jupyter…&lt;/p&gt;
</summary><content type="html">&lt;p&gt;When I helped start &lt;a href="https://2i2c.org"&gt;2i2c.org&lt;/a&gt;, one of my goals for the non-profit was to experiment with new and different ways of supporting the Jupyter ecosystem’s long term health. As part of that, we have identified how &lt;a href="https://2i2c.org/blog/2025/good-citizen/"&gt;foundational contributions are very important (and distinct from directed contributions)&lt;/a&gt;, and have been experimenting with &lt;a href="https://2i2c.org/blog/2025/foundational-contributions/"&gt;new ways to make foundational contributions&lt;/a&gt; in sustainable and structured ways.&lt;/p&gt;
&lt;p&gt;One such way was via systematically doing code reviews for Pull Requests from non-maintainers in the JupyterHub ecosystem. Reviewing PRs is a critical way that maintainers keep an open source project moving forward, but identifying PRs that can productively be merged is hard. This is a post describing our system for scaling this in our team. Each 2 week sprint, we:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Looked at all open PRs in the JupyterHub org who were not maintainers&lt;/li&gt;
&lt;li&gt;Picked a PR that I deemed was reviewable and ideally mergeable within this time window&lt;/li&gt;
&lt;li&gt;Have an engineer on the team pick up that PR, and try to get it to close&lt;/li&gt;
&lt;li&gt;&lt;a href="https://github.com/2i2c-org/infrastructure/issues/5058"&gt;Report back&lt;/a&gt; on what we have learnt, so we can iterate on our process&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;We managed to do this for a majority of sprints over the last roughly 12 months! One key bottleneck we identified in the process was Step 2. In particular, I was relying on my &lt;em&gt;maintainer intuition&lt;/em&gt; to pick a single PR that I &lt;em&gt;believe&lt;/em&gt; can be merged, so others in the team can do review work. I started exploring &lt;em&gt;what&lt;/em&gt; this intuition is, and if it can be scaled.&lt;/p&gt;
&lt;h2 id="what-is-this-maintainer-intuition"&gt;What is this maintainer intuition?&lt;/h2&gt;
&lt;p&gt;How did I pick a PR from a long list of open PRs? Observing my own behavior a few times, I noticed I was looking for:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;PRs that aren’t too &lt;em&gt;big&lt;/em&gt;, and are &lt;strong&gt;a reasonable size&lt;/strong&gt; that can be merged within a 2 week window&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;CI tests passing&lt;/strong&gt;, so at least our automated checks haven’t caught any issues with it&lt;/li&gt;
&lt;li&gt;Features or bug fixes that I believe &lt;strong&gt;add value to the project&lt;/strong&gt; and move us in the right direction towards being able to support our users as they need (this is the hardest!)&lt;/li&gt;
&lt;li&gt;If the author of the PR is a &lt;strong&gt;newish contributor&lt;/strong&gt;, as I want to encourage them to stick around by being responsive to their gift. All PRs are gifts that we may or may not choose to accept, but should do so with grace.&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;How long ago the PR was opened&lt;/strong&gt;. There is such a big difference between a response to your PR 2 days after you make it vs 2 months vs 2 years. I prioritized newer PRs.&lt;/li&gt;
&lt;li&gt;What &lt;strong&gt;kind of contribution&lt;/strong&gt; is it primarily? Different engineers on our team have different skillsets (JS, Python, etc) and I wanted to match the PR to what the engineer preferred code reviewing.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;While (3) is hard to scale, everything else seemed like something we could build systems for that let others follow a process, thus removing myself as a bottleneck. I experimented with some GitHub issue filters and project automations, and after finding them lacking, built out a brand new open source project to do this: &lt;a href="https://github.com/jupyter/pr-triage-board-bot"&gt;pr-triage-board-bot&lt;/a&gt;!&lt;/p&gt;
&lt;h2 id="pr-triage-github-boards"&gt;PR Triage GitHub Boards&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://github.com/jupyter/pr-triage-board-bot"&gt;pr-triage-board-bot&lt;/a&gt; automatically maintains PR Triage Boards for a Github organization. You can check out the current boards to get a sense of how it looks: &lt;a href="https://github.com/orgs/jupyterhub/projects/4"&gt;JupyterHub&lt;/a&gt;, &lt;a href="https://github.com/orgs/jupyterlab/projects/11"&gt;JupyterLab&lt;/a&gt;, and &lt;a href="https://github.com/orgs/geojupyter/projects/3"&gt;GeoJupyter&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;For each PR Triage Board, the bot will:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Add all open, ready for review PRs&lt;/li&gt;
&lt;li&gt;Remove all closed or draft PRs&lt;/li&gt;
&lt;li&gt;Annotate each PR with additional deterministic project fields that allow for sorting and filtering in the project board. To start with, it populates the following fields:
&lt;ol&gt;
&lt;li&gt;Author kind (Maintainer, Seasoned Contributor, First Time Contributor, Bot)&lt;/li&gt;
&lt;li&gt;Size (Number of lines touched)&lt;/li&gt;
&lt;li&gt;Date it was opened&lt;/li&gt;
&lt;li&gt;If a maintainer has already interacted with the PR (One, Many, None)&lt;/li&gt;
&lt;li&gt;Are there merge conflicts?&lt;/li&gt;
&lt;li&gt;What kind of files were mostly changed? (Python, JS, Docs)&lt;/li&gt;
&lt;/ol&gt;
&lt;/li&gt;
&lt;li&gt;Keep this up to date by automatically running every hour via GitHub actions&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Once these project fields are populated automatically, maintainers can create different Project Views for themselves to help with different workflows via filters and sorting. For example, &lt;a href="https://github.com/orgs/jupyterhub/projects/4"&gt;the JupyterHub board&lt;/a&gt; contains the following views:&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;PRs split by author kind, sorted by newness and size so we can try to respond to PRs from new users as early as possible&lt;/li&gt;
&lt;li&gt;PRs that have not had a single maintainer interaction on them, so we can acknowledge people’s contributions to us even if we can’t fully review it at the moment&lt;/li&gt;
&lt;li&gt;Bot PRs with all tests passing and no merge conflicts, so we can more easily stay on top of automated updates&lt;/li&gt;
&lt;li&gt;PRs that have been approved by a maintainer but not merged yet, as sometimes a maintainer wants to give others time to object but needs to go back and hit merge.&lt;/li&gt;
&lt;li&gt;PRs that are primarily updating documentation, as the review process for this can be different&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;For our original purpose of getting more people to do code review, this board has served well — we can roughly say ‘Pick a PR that looks good to you from the top of the “First Time Contributor” or “Seasoned Contributor” list’, and that relieves me from being the bottleneck quite a bit. Other maintainers are also finding a lot of value in this board. For example, if you only have 15 min, you can probably get a clean bot PR reviewed and merged. Or look for PRs that haven’t been acknowledged by a maintainer and engage with them. Our next step is to help run structured social experiments, where we try to establish specific ceremonies in the open source ecosystem to get specific lists of open PRs (such as “PRs with no maintainer engagement, or PRs older than 2y”) down to zero. Stay tuned to hear more :)&lt;/p&gt;
&lt;p&gt;We are essentially using GitHub Project Fields as a database, adding additional fields that are very valuable to maintainers but not available in GitHub by default. While ideally these would be contributed by us directly into GitHub, only Microsoft has control over what gets added to the product. So we find interesting workarounds like this to accomplish our goals :)&lt;/p&gt;
&lt;h2 id="successful-adoption"&gt;Successful Adoption&lt;/h2&gt;
&lt;p&gt;I was just playing with this, and Raniere from the JupyterHub team spotted it and &lt;a href="https://jupyter.zulipchat.com/#narrow/channel/469744-jupyterhub/topic/.22PR.20triage.20.28experimental.29.22.20project.20on.20GitHub/with/536668935"&gt;asked about it&lt;/a&gt; on Zulip. &lt;a href="https://jasongrout.org/"&gt;Jason Grout&lt;/a&gt; from the JupyterLab team was also super interested, and with contributions from him, the bot quickly got adopted by the JupyterLab GitHub org as well. We cleaned this up a bit more, and found adoption within the GeoJupyter project too with help from &lt;a href="https://mfisher87.github.io/"&gt;Matt Fisher&lt;/a&gt;. As with everything we do at 2i2c, I had tried to design this to be widely useful to many orgs and maintainers rather than just us, and looks like I have wildly succeeded :)&lt;/p&gt;
&lt;p&gt;I am happy to announce today that 2i2c is officially donating pr-triage-board-bot to Project Jupyter! I will still continue to contribute to maintaining the project, and welcome contributions from everyone else too!&lt;/p&gt;
&lt;p&gt;If this looks useful to your open source project, consider adopting it by &lt;a href="https://github.com/jupyter/pr-triage-board-bot?tab=readme-ov-file#set-up"&gt;following these instructions&lt;/a&gt;. The project is still fairly new, so if you run into issues please let us know on the &lt;a href="https://jupyter.zulipchat.com/"&gt;Project Jupyter Zulip chat&lt;/a&gt; or by &lt;a href="https://github.com/jupyter/pr-triage-board-bot/issues"&gt;opening an issue&lt;/a&gt;.&lt;/p&gt;
</content><category term="community"/><category term="JupyterHub"/></entry><entry><title>Accurately counting Daily, Weekly &amp; Monthly active users on JupyterHub</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2023/accurately-counting-daily-weekly-monthly-active-users/" rel="alternate"/><published>2023-03-27T09:16:00+00:00</published><updated>2023-03-27T09:16:00+00:00</updated><author><name>yuvipanda</name></author><id>tag:jasongrout.github.io,2023-03-27:/medium-archive/pelican/posts/2023/accurately-counting-daily-weekly-monthly-active-users/</id><summary type="html">&lt;p&gt;Being able to say ‘we served X unique users over the last month’ (the Monthly Active User metric) is very helpful when advocating for…&lt;/p&gt;
</summary><content type="html">&lt;p&gt;Being able to say ‘we served X unique users over the last month’ (the &lt;a href="https://en.wikipedia.org/wiki/Active_users"&gt;Monthly Active User metric&lt;/a&gt;) is very helpful when advocating for resources for a JupyterHub your organization is running. However, until now, figuring out that number &lt;em&gt;accurately&lt;/em&gt; has been difficult, requiring keeping and analysing JupyterHub logs.&lt;/p&gt;
&lt;p&gt;That changes with JupyterHub 3.1! &lt;a href="https://github.com/jupyterhub/jupyterhub/pull/4214"&gt;This Pull Request&lt;/a&gt; adds daily, weekly, and monthly active user metrics to JupyterHub, accessible via the prometheus interface by hitting the &lt;code&gt;/metrics&lt;/code&gt; URL on your JupyterHub. These metrics are calculated by JupyterHub itself, and are pretty accurate as it already keeps track of when a user was last active. This relies metric on each JupyterHub user matching an actual user, so if you are using your JupyterHub deployment purely as an API with ephemeral users (as &lt;a href="https://github.com/jupyterhub/binderhub/"&gt;binderhub&lt;/a&gt; does, for example) or delete inactive users, these will not be useful numbers.&lt;/p&gt;
&lt;p&gt;Ideally, you should have a &lt;a href="https://prometheus.io/"&gt;prometheus&lt;/a&gt; instance to scrape and store metrics from your JupyterHub over time, so you can track this metric over time. However, in a pinch, you can also go directly to &lt;code&gt;https://&amp;lt;your-hub-url&amp;gt;/hub/api/metrics&lt;/code&gt; and see the current value of these metrics. They are in the &lt;a href="https://github.com/prometheus/docs/blob/main/content/docs/instrumenting/exposition_formats.md"&gt;prometheus exposition format&lt;/a&gt;, but you can look for something like:&lt;/p&gt;
&lt;div class="highlight"&gt;&lt;pre&gt;&lt;span&gt;&lt;/span&gt;&lt;code&gt;&lt;span class="c1"&gt;# HELP jupyterhub_active_users number of users who were active in the given time period&lt;/span&gt;
&lt;span class="c1"&gt;# TYPE jupyterhub_active_users gauge&lt;/span&gt;
jupyterhub_active_users&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;period&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;24h&amp;quot;&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;610&lt;/span&gt;.0
jupyterhub_active_users&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;period&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;7d&amp;quot;&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;2800&lt;/span&gt;.0
jupyterhub_active_users&lt;span class="o"&gt;{&lt;/span&gt;&lt;span class="nv"&gt;period&lt;/span&gt;&lt;span class="o"&gt;=&lt;/span&gt;&lt;span class="s2"&gt;&amp;quot;30d&amp;quot;&lt;/span&gt;&lt;span class="o"&gt;}&lt;/span&gt;&lt;span class="w"&gt; &lt;/span&gt;&lt;span class="m"&gt;4526&lt;/span&gt;.0
&lt;/code&gt;&lt;/pre&gt;&lt;/div&gt;
&lt;p&gt;This denotes that this particular hub had 610 daily active users (active over the last 24 hours), 2800 weekly users (over last 7d) and 4526 monthly ones (over the last 30 days). Very helpful if you want to quickly add numbers to a report :)&lt;/p&gt;
&lt;p&gt;Note that depending on your hub’s configuration, access to the &lt;code&gt;/metrics&lt;/code&gt; endpoint might &lt;a href="https://jupyterhub.readthedocs.io/en/stable/api/app.html#jupyterhub.app.JupyterHub.authenticate_prometheus"&gt;require authentication&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;If you &lt;em&gt;do&lt;/em&gt; have a prometheus installation for your JupyterHub, you may benefit from deploying &lt;a href="https://github.com/jupyterhub/grafana-dashboards"&gt;the JupyterHub Grafana Dashboards&lt;/a&gt;. These are targeted at installations of &lt;a href="https://z2jh.jupyter.org"&gt;zero-to-jupyterhub on kubernetes&lt;/a&gt;, and provide a lot of useful usage &amp;amp; diagnostic information.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Grafana Dashboard showing “Hub Usage Stats” for a JupyterHub deployment. Four panels, clockwise: Current Active Users, Daily Active Users, Weekly Active Users, Monthly Active Users." src="https://jasongrout.github.io/medium-archive/pelican/posts/2023/accurately-counting-daily-weekly-monthly-active-users/images/001-0_VYHSKwQ5Jp-L3ArS.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Grafana Dashboard showing “Hub Usage Stats” for a JupyterHub deployment. Four panels, clockwise: Current Active Users, Daily Active Users, Weekly Active Users, Monthly Active Users.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Happy report writing!&lt;/p&gt;
</content><category term="DevOps"/><category term="JupyterHub"/></entry><entry><title>Securely pushing to GitHub from a JupyterHub with gh-scoped-creds</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2022/securely-pushing-to-github-from-a-jupyterhub/" rel="alternate"/><published>2022-04-21T16:54:00+00:00</published><updated>2022-04-21T16:57:00+00:00</updated><author><name>yuvipanda</name></author><id>tag:jasongrout.github.io,2022-04-21:/medium-archive/pelican/posts/2022/securely-pushing-to-github-from-a-jupyterhub/</id><summary type="html">&lt;p&gt;Many JupyterHub users want to push and pull their content from GitHub in order to collaborate and share their work. However, working on a…&lt;/p&gt;
</summary><content type="html">&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2022/securely-pushing-to-github-from-a-jupyterhub/images/001-1_9E0cif7g07xWAOfsFFmijw.webp" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;p&gt;Many JupyterHub users want to push and pull their content from GitHub in order to collaborate and share their work. However, working on a JupyterHub means working on &lt;em&gt;shared infrastructure&lt;/em&gt;, not your own laptop, and this poses some extra security risks that have made two-way sync with GitHub more difficult. This post describes &lt;code&gt;gh-scoped-creds&lt;/code&gt;, a new tool to make it quick and easy to authorize a JupyterHub session with push access to GitHub in a secure and simple manner.&lt;/p&gt;
&lt;p&gt;GitHub user credentials are &lt;a href="https://github.blog/2022-04-15-security-alert-stolen-oauth-user-tokens/"&gt;high value targets&lt;/a&gt; for cybercriminals in today’s security environment, and any system that stores these credentials long term paints an unwanted target on itself. Current solutions — putting an ssh key on the JupyterHub, using a &lt;a href="https://docs.github.com/en/authentication/keeping-your-account-and-data-secure/creating-a-personal-access-token"&gt;personal access token&lt;/a&gt; or deploy keys — involve storing long term valid GitHub credentials in the filesystem. As users can do this by themselves without admin intervention, admins often are not aware these (often unencrypted) credentials are on their filesystems. If an attacker compromises an ssh key or a personal access token, they have unlimited access to all GitHub repos the compromised user had access to, including repos in high-impact GitHub organizations. In the recent credential theft incident, Travis-CI and Heroku were ‘lucky’ in that the &lt;a href="https://github.blog/2022-04-15-security-alert-stolen-oauth-user-tokens/"&gt;attackers accessed npm infrastructure&lt;/a&gt; — and since npm is owned by GitHub, GitHub was able to detect that Travis CI and Heroku had compromised credentials. You and the users of repositories you have rights to might not be so lucky. It’s 2022, and &lt;a href="https://github.com/cncf/tag-security/blob/main/supply-chain-security/compromises/README.md"&gt;supply chain attacks are everywhere&lt;/a&gt; — you aren’t special, you’re just one link in a long chain attackers use to get to someone else.&lt;/p&gt;
&lt;p&gt;There is a clear need for a simple solution that lets users push to GitHub from JupyterHub in a secure manner without admins having to worry about securing high-value GitHub credentials long term. It is not acceptable to “Just Say No” to users wanting this functionality either — if you try to ‘sacrifice’ usability for security, you end up getting neither.&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/yuvipanda/gh-scoped-creds/"&gt;&lt;code&gt;gh-scoped-creds&lt;/code&gt;&lt;/a&gt; attempts to solve this problem by allowing users to grant &lt;em&gt;time-limited&lt;/em&gt; push access to &lt;em&gt;specific repositories&lt;/em&gt; to &lt;em&gt;specific JupyterHub installations&lt;/em&gt; in a user friendly way.&lt;/p&gt;
&lt;p&gt;Here’s a quick GIF running through the user workflow.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2022/securely-pushing-to-github-from-a-jupyterhub/images/002-1_B3qjACXLBG9pBOlzY8WNxA.mp4" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;p&gt;Push access is scoped both by time (credentials expire after 8 hours) as well as repository (access is granted per-repository, per-hub). While you need to refresh credentials every 8 hours, the list of repositories is remembered until you explicitly revoke access. You can always grant access to your own personal repositories, but repositories belonging to organisations might require admins to approve push access to them.&lt;/p&gt;
&lt;p&gt;You can also run the command from the terminal as &lt;code&gt;gh-scoped-creds&lt;/code&gt; instead of using the IPython magic &lt;code&gt;%ghscopedcreds&lt;/code&gt; as shown in the demo. This way, you can also use this from a HPC system, not just a JupyterHub!&lt;/p&gt;
&lt;p&gt;Setting this up for your JupyterHub requires a tiny bit of work from the admin — see &lt;a href="https://github.com/yuvipanda/gh-scoped-creds/"&gt;the project README&lt;/a&gt; for more details. Shouldn’t take long, and it’s a one-time task. Once that’s set up, your users can securely push to GitHub from the comfort of their JupyterHubs!&lt;/p&gt;
&lt;p&gt;Thanks to &lt;a href="https://twitter.com/fperez_org"&gt;Fernando Perez&lt;/a&gt; for using his &lt;a href="https://classes.berkeley.edu/content/2021-spring-stat-159-001-lec-001"&gt;stat159 class&lt;/a&gt; at &lt;a href="https://www.berkeley.edu/"&gt;UC Berkeley&lt;/a&gt; to test this project out.&lt;/p&gt;
</content><category term="GitHub"/><category term="JupyterHub"/></entry><entry><title>Setting up a “Production Ready” TLJH deployment</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2021/setting-up-a-production-ready-tljh-deployment/" rel="alternate"/><published>2021-06-04T17:05:00+00:00</published><updated>2021-06-04T17:05:00+00:00</updated><author><name>yuvipanda</name></author><id>tag:jasongrout.github.io,2021-06-04:/medium-archive/pelican/posts/2021/setting-up-a-production-ready-tljh-deployment/</id><summary type="html">&lt;p&gt;The Littlest JupyterHub is an extremely capable hub distribution that I’d recommend for situations where you expect, on average, under 100…&lt;/p&gt;
</summary><content type="html">&lt;p&gt;&lt;a href="http://tljh.jupyter.org/"&gt;The Littlest JupyterHub&lt;/a&gt; is an extremely capable hub distribution that I’d recommend for situations where you expect, on average, under 100 active users.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="The Littlest JupyterHub is a distribution of JupyterHub for single VM instances, best-used with 1–100 users." src="https://jasongrout.github.io/medium-archive/pelican/posts/2021/setting-up-a-production-ready-tljh-deployment/images/001-1_4zZBoheuyfPsemCzc_O7jg.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;The Littlest JupyterHub is a distribution of JupyterHub for single VM instances, best-used with 1–100 users.&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="why-not-kubernetes"&gt;Why not Kubernetes?&lt;/h2&gt;
&lt;p&gt;The primary reason to use &lt;a href="https://z2jh.jupyter.org"&gt;Zero to JupyterHub on k8s&lt;/a&gt; over TLJH in cases with a smaller number of users is to reduce costs — Kubernetes can spin down nodes when not in use. However, you’ll always have at least one node running (for the hub / proxy pods) and the extra complexity that comes with it — particularly around needing to build your own docker images — may not be worth it. TLJH works perfectly well for these cases!&lt;/p&gt;
&lt;h2 id="what-is-production"&gt;What is ‘production’?&lt;/h2&gt;
&lt;p&gt;A JupyterHub that you can run securely without lots of intervention from the person who created it is what I’ll call a &lt;em&gt;production-ready&lt;/em&gt; JupyterHub. It’s a pretty arbitrary standard. In this blog post, I’ll lay out what &lt;strong&gt;I&lt;/strong&gt; want in the TLJH hubs I run before I let users on them.&lt;/p&gt;
&lt;h2 id="authentication"&gt;Authentication&lt;/h2&gt;
&lt;p&gt;Use a &lt;em&gt;real&lt;/em&gt; &lt;a href="https://tljh.jupyter.org/en/latest/howto/index.html#authentication"&gt;Authenticator&lt;/a&gt;, not the default &lt;a href="https://github.com/jupyterhub/firstuseauthenticator"&gt;&lt;code&gt;FirstUseAuthenticator&lt;/code&gt;&lt;/a&gt;. The default authenticator is pretty insecure, and should really not be used in production. If you don’t know what to use, I’ll suggest the &lt;a href="https://tljh.jupyter.org/en/latest/howto/auth/google.html"&gt;Google&lt;/a&gt; or &lt;a href="https://tljh.jupyter.org/en/latest/howto/auth/github.html"&gt;GitHub&lt;/a&gt; authenticators.&lt;/p&gt;
&lt;h2 id="enable-https"&gt;Enable HTTPS&lt;/h2&gt;
&lt;p&gt;Enable &lt;a href="https://tljh.jupyter.org/en/latest/howto/admin/https.html"&gt;HTTPS&lt;/a&gt;. An absolute security requirement now, and TLJH makes it quite easy. You &lt;em&gt;do&lt;/em&gt; need to get a domain for this to work, which can be a source of friction. Totally worth it, though.&lt;/p&gt;
&lt;h2 id="resource-limits"&gt;Resource Limits&lt;/h2&gt;
&lt;p&gt;In many systems, a single user can often write code that accidentally crashes the whole system. By default, TLJH doesn’t have any memory limits enforced per-user, but it is very easy to configure it to enforce &lt;a href="https://tljh.jupyter.org/en/latest/topic/tljh-config.html#user-server-limits"&gt;memory limits&lt;/a&gt;. Tuning these to match your needs will help prevent a single student from accidentally taking down your whole hub. I’d highly recommend &lt;a href="https://tljh.jupyter.org/en/latest/howto/admin/nbresuse.html"&gt;checking&lt;/a&gt; how much memory your typical notebook uses, and making sure you have user limits set to above that.&lt;/p&gt;
&lt;h2 id="sizing-your-vm-correctly"&gt;Sizing your VM correctly&lt;/h2&gt;
&lt;p&gt;If you choose a VM that’s too big, you’ll end up spending a lot of cash for unused resources. If it’s too small, your users will not have the resources they need to do their work. TLJH provides &lt;a href="https://tljh.jupyter.org/en/latest/howto/admin/resource-estimation.html"&gt;some helpful docs&lt;/a&gt; estimating your VM size, and you can always &lt;a href="https://tljh.jupyter.org/en/latest/howto/admin/resize.html"&gt;resize&lt;/a&gt; your VM afterwards if you get it wrong.&lt;/p&gt;
&lt;h2 id="disk-backups"&gt;Disk backups&lt;/h2&gt;
&lt;p&gt;TLJH contains everything on the VM’s disk — your user environment, users’ home directories, current hub configuration, etc. It is very important you back this up, to recover in case of disasters. Automated disk snapshots from your cloud provider are an easy way to do this. Most major cloud providers offer a way to do this — &lt;a href="https://cloud.google.com/compute/docs/disks/create-snapshots"&gt;Google Cloud&lt;/a&gt;, &lt;a href="https://www.digitalocean.com/docs/images/snapshots/"&gt;Digital Ocean&lt;/a&gt;, &lt;a href="https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/EBSSnapshots.html"&gt;AWS&lt;/a&gt;, etc. Some let you automate it as well — Google &amp;amp; AWS certainly do, I’m not sure about other cloud providers. This isn’t the &lt;em&gt;best&lt;/em&gt; way to do backup — there’s approximately 1 billion ways to do so. However, this is an absolute minimum, and it might just be enough.&lt;/p&gt;
&lt;p&gt;If you want to be more fancy, I’d suggest using a separate disk / volume for your user home directories, possibly on &lt;a href="https://wiki.ubuntu.com/ZFS"&gt;ZFS&lt;/a&gt;, and snapshot much more aggressively. Talk to your nearest google search bar for your options.&lt;/p&gt;
&lt;h2 id="pin-your-public-ip"&gt;Pin your public IP&lt;/h2&gt;
&lt;p&gt;Some cloud providers change your VM’s public IP address if you start / stop them. This can be pretty bad — you’ll have to change your domain’s DNS entry, and re-acquire HTTPS. A hassle! You can tell your cloud provider to hang on to your IP even if your VM goes down / changes. And you should! DigitalOcean doesn’t require this, but &lt;a href="https://cloud.google.com/compute/docs/ip-addresses/reserve-static-external-ip-address"&gt;Google Cloud does&lt;/a&gt;. I think AWS does too, but I’m not sure how you can reserve the public IP for it — since it’s usually a domain name itself.&lt;/p&gt;
&lt;h2 id="base-environment-setup-snapshot"&gt;Base environment setup + snapshot&lt;/h2&gt;
&lt;p&gt;TLJH has a shared &lt;a href="https://conda.io"&gt;conda&lt;/a&gt; environment that is used by &lt;em&gt;all&lt;/em&gt; users. Everyone can read from it, but only users who are &lt;code&gt;admin&lt;/code&gt; can write to it (via &lt;code&gt;sudo&lt;/code&gt;). This is one of TLJH’s core design trade-offs - admins can install packages the way they are used to, without requiring a separate image-build step. But it also means the admin can mess it up - conda environments can be sometimes fickle! So it’s not a bad idea to spend some time in the beginning setting everything up - python packages, JupyterLab extensions, etc. Then make a disk snapshot, so you can revert to it if things go bad. This is where having a separate disk for your user home directories comes in handy, so you can reset your hub environment without losing your user home directories.&lt;/p&gt;
&lt;h2 id="ssh-admin-access"&gt;SSH admin access&lt;/h2&gt;
&lt;p&gt;The TLJH documentation strives hard to make sure SSH isn’t &lt;em&gt;required&lt;/em&gt; for setup and most common usage. However, if your TLJH breaks in certain ways, you can no longer access the machine — since all access is via TLJH! For this, I recommend making sure someone who is admin has SSH access to the VM. Most cloud providers offer a way to set the root ssh key on creation. If not, you can follow the many guides on the internet to making it happen.&lt;/p&gt;
&lt;p&gt;You can also just put your ssh keys in &lt;code&gt;$HOME/.ssh/authorized_keys&lt;/code&gt;, and ssh in as &lt;code&gt;jupyter-&amp;lt;username&amp;gt;@&amp;lt;hub-ip&amp;gt;&lt;/code&gt;. This works for any / all users!&lt;/p&gt;
&lt;h2 id="others"&gt;Others?&lt;/h2&gt;
&lt;p&gt;I’m sure this isn’t the end — probably need something about firewalls, monitoring and automated system package upgrades. But hey, great start!&lt;/p&gt;
</content><category term="JupyterHub"/><category term="TLJH"/></entry><entry><title>Connect to a JupyterHub from Visual Studio Code</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2019/connect-to-a-jupyterhub-from-visual-studio-code/" rel="alternate"/><published>2019-12-09T17:23:00+00:00</published><updated>2019-12-09T17:23:00+00:00</updated><author><name>yuvipanda</name></author><id>tag:jasongrout.github.io,2019-12-09:/medium-archive/pelican/posts/2019/connect-to-a-jupyterhub-from-visual-studio-code/</id><summary type="html">&lt;p&gt;Visual Studio Code has pretty good support for running Jupyter Notebooks. But what if your organization has a JupyterHub running remotely…&lt;/p&gt;
</summary><content type="html">&lt;p&gt;&lt;a href="https://code.visualstudio.com"&gt;Visual Studio Code&lt;/a&gt; has pretty good support&lt;br&gt;
for &lt;a href="https://code.visualstudio.com/docs/python/jupyter-support"&gt;running Jupyter Notebooks&lt;/a&gt;. But what if your organization has a &lt;a href="https://jupyter.org/hub"&gt;JupyterHub&lt;/a&gt; running remotely, with more compute resources &amp;amp; access to large amounts of data? How can you access that from Visual Studio Code running on your local machine?&lt;/p&gt;
&lt;p&gt;It’s pretty easy to do, and this blog post will guide you through it.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/connect-to-a-jupyterhub-from-visual-studio-code/images/001-1_FGIOFXmphgFGab3WOJwSZw_2x.webp" alt="jupyterhub and vscode logos" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;h2 id="step-1-get-a-jupyterhub-access-token"&gt;Step 1: Get a JupyterHub access token&lt;/h2&gt;
&lt;p&gt;JupyterHub lets you create tokens for yourself for use by third party applications. These tokens can be used anywhere a &lt;a href="https://jupyter-notebook.readthedocs.io/en/stable/security.html"&gt;Jupyter Notebook access token&lt;/a&gt; is needed. Since this is what Visual Studio Code needs, let’s acquire one.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;Log-in to your JupyterHub&lt;/li&gt;
&lt;li&gt;Access your &lt;em&gt;Control Panel&lt;/em&gt;. In classic notebook, there is a ‘Control Panel’ button on the top right. In JupyterLab, you can access it under ‘File -&amp;gt; Hub Control Panel’&lt;/li&gt;
&lt;/ol&gt;
&lt;figure&gt;
&lt;img alt="Top Right ‘Control Panel’ button in classic Jupyter Notebook" src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/connect-to-a-jupyterhub-from-visual-studio-code/images/002-1_ZNZ-jbJu8TbhWPZ596difg.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Top Right ‘Control Panel’ button in classic Jupyter Notebook&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;figure&gt;
&lt;img alt="File -&amp;gt; Hub Control Panel in JupyterLab" src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/connect-to-a-jupyterhub-from-visual-studio-code/images/003-1_Jl-Ug0_n0Pr22AW7Le6CWg.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;File -&amp;gt; Hub Control Panel in JupyterLab&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;ol start="3"&gt;
&lt;li&gt;Go to the ‘Token’ page by clicking ‘Token’ in the top bar&lt;/li&gt;
&lt;/ol&gt;
&lt;figure&gt;
&lt;img alt="Token link in the top bar to go to the token page" src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/connect-to-a-jupyterhub-from-visual-studio-code/images/004-1_BSKDpkMhSp_l0oAuMVrMRA.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Token link in the top bar to go to the token page&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;ol start="4"&gt;
&lt;li&gt;Type in a description for the new token you want, and click ‘Request new API Token’&lt;/li&gt;
&lt;/ol&gt;
&lt;figure&gt;
&lt;img alt="Type in a description for what this token will be used for" src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/connect-to-a-jupyterhub-from-visual-studio-code/images/005-1_6JY55qIAgIaYQlviKY1xYw.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Type in a description for what this token will be used for&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;ol start="5"&gt;
&lt;li&gt;Copy your token and keep it somewhere safe. You should treat this like a password to your JupyterHub. You can (and should!) revoke it (as I have done) from the same page when you are no longer using it.&lt;/li&gt;
&lt;/ol&gt;
&lt;figure&gt;
&lt;img alt="Copy your token" src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/connect-to-a-jupyterhub-from-visual-studio-code/images/006-1_1iCbErLsY3L8p5YFG-ch0Q.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Copy your token&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;This is all the information you need from JupyterHub! Now let’s go to vscode.&lt;/p&gt;
&lt;h2 id="step-2-connect-vs-code-to-your-jupyterhub"&gt;Step 2: Connect VS Code to your JupyterHub&lt;/h2&gt;
&lt;p&gt;Visual Studio Code supports connecting to a &lt;a href="https://code.visualstudio.com/docs/python/jupyter-support#_connect-to-a-remote-jupyter-server"&gt;remote notebook server&lt;/a&gt;, and we can use that to connect to our JupyterHub. You must perform these steps &lt;em&gt;before&lt;/em&gt; opening your notebook.&lt;/p&gt;
&lt;ol&gt;
&lt;li&gt;&lt;strong&gt;Open the command palette&lt;/strong&gt; in Visual Studio Code (‘Cmd+Shift+P’ on MacOS, ‘Ctrl+Shift+P’ elsewhere)&lt;/li&gt;
&lt;li&gt;&lt;strong&gt;Select ‘Python: Specify local or remote Jupyter server for connections’&lt;/strong&gt;&lt;/li&gt;
&lt;/ol&gt;
&lt;figure&gt;
&lt;img alt="vscode command palette" src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/connect-to-a-jupyterhub-from-visual-studio-code/images/007-1_hm9yZinnwlF3EqeXuossBQ.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;vscode command palette&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;ol start="3"&gt;
&lt;li&gt;&lt;strong&gt;Construct your &lt;em&gt;notebook server URL&lt;/em&gt;&lt;/strong&gt; with the following template: &lt;code&gt;https://&amp;lt;your-hub-url&amp;gt;/user/&amp;lt;your-hub-user-name&amp;gt;/?token=&amp;lt;your-token&amp;gt;&lt;/code&gt;. Note that your hub user name might sometimes be escaped from whatever you used to actually log in, if it has special characters. You can verify this by looking at the URL you get once you log in to your JupyterHub — it should have the right one after ‘user’.&lt;/li&gt;
&lt;/ol&gt;
&lt;figure&gt;
&lt;img alt="Enter your notebook server URL" src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/connect-to-a-jupyterhub-from-visual-studio-code/images/008-1__FILTBgRJ76bq5R7LYFGww.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Enter your notebook server URL&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;ol start="4"&gt;
&lt;li&gt;&lt;strong&gt;Create or open a new notebook.&lt;/strong&gt; The kernel for this should now live on your JupyterHub! You can verify this by running &lt;code&gt;!hostname&lt;/code&gt;, which should return the hostname of your remote JupyterHub server instead of your local hostname. You can also try importing libraries that are in the remote JupyterHub server, but not your local file system.&lt;/li&gt;
&lt;/ol&gt;
&lt;p&gt;Tada! That wasn’t so hard, was it?&lt;/p&gt;
&lt;h2 id="limitations"&gt;Limitations&lt;/h2&gt;
&lt;ol&gt;
&lt;li&gt;Your JupyterHub notebook server &lt;em&gt;must&lt;/em&gt; be running already when you try to open your notebook — Visual Studio Code will not automatically start it. If it has stopped, you need to log in to your JupyterHub &amp;amp; start it again. You do not need a new token though.&lt;/li&gt;
&lt;li&gt;Watch out for filesystem access. When you are calling &lt;code&gt;open()&lt;/code&gt;(or a helper function that eventually access a file) to read a file, it is going to be read from your &lt;em&gt;remote JupyterHub server’s home directory&lt;/em&gt;, not your local system’s current directory. So if you add a new file next to your Jupyter Notebook locally, that is &lt;em&gt;not&lt;/em&gt; automatically going to be available to the Jupyter Notebook to read.&lt;/li&gt;
&lt;li&gt;Installing pip or conda packages locally will have no effect, since your Python kernel is running on your JupyterHub. Use the &lt;a href="https://ipython.readthedocs.io/en/stable/interactive/magics.html#magic-pip"&gt;%pip&lt;/a&gt; or &lt;a href="https://ipython.readthedocs.io/en/stable/interactive/magics.html#magic-conda"&gt;%conda&lt;/a&gt; magics to install packages in the correct environment.&lt;/li&gt;
&lt;/ol&gt;
</content><category term="JupyterHub"/></entry><entry><title>99 ways to extend the Jupyter ecosystem</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2019/99-ways-to-extend-the-jupyter-ecosystem/" rel="alternate"/><published>2019-06-18T15:55:00+00:00</published><updated>2019-06-18T15:55:00+00:00</updated><author><name>yuvipanda</name></author><id>tag:jasongrout.github.io,2019-06-18:/medium-archive/pelican/posts/2019/99-ways-to-extend-the-jupyter-ecosystem/</id><summary type="html">&lt;p&gt;Whenever someone says ‘You can do that with an extension’ in the Jupyter ecosystem, it is often not clear what kind of extension they are…&lt;/p&gt;
</summary><content type="html">&lt;p&gt;Whenever someone says ‘&lt;em&gt;You can do that with an extension&lt;/em&gt;’ in the Jupyter ecosystem, it is often not clear what &lt;em&gt;kind&lt;/em&gt; of extension they are talking about. The Jupyter ecosystem is very modular and extensible, so there are lots of ways to extend it. This blog post aims to provide a quick summary of the most common ways to extend Jupyter, and links to help you explore the extension ecosystem.&lt;/p&gt;
&lt;h2 id="jupyterlab-extensions-labextension"&gt;JupyterLab extensions (labextension)&lt;/h2&gt;
&lt;figure&gt;
&lt;img alt="Draw vector graphics in JupyterLab with the jupyterlab-drawio extension" src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/99-ways-to-extend-the-jupyter-ecosystem/images/001-1_gCej3VJVI_8k27KR33wDLw.mp4" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Draw vector graphics in JupyterLab with the &lt;a href="https://github.com/QuantStack/jupyterlab-drawio"&gt;jupyterlab-drawio&lt;/a&gt; extension&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;&lt;a href="https://github.com/jupyterlab/jupyterlab"&gt;JupyterLab&lt;/a&gt; is a popular ‘new’ interface for working with Jupyter Notebooks. It is an interactive development environment for working with notebooks, code and data — and hence extremely extensible. Using JupyterLab extensions, you can add entirely new functionality or change almost any aspect of how the interface behaves. These are written in &lt;a href="https://www.typescriptlang.org/"&gt;TypeScript&lt;/a&gt; or JavaScript, and run in the browser.&lt;/p&gt;
&lt;p&gt;The JupyterLab documentation has information on how to &lt;a href="https://jupyterlab.readthedocs.io/en/stable/user/extensions.html"&gt;install &amp;amp; use extensions&lt;/a&gt;, as well as how to &lt;a href="https://jupyterlab.readthedocs.io/en/stable/developer/extension_dev.html"&gt;author &amp;amp; distribute them&lt;/a&gt;. You can also discover extensions by searching &lt;a href="https://github.com/topics/jupyterlab-extension"&gt;on GitHub&lt;/a&gt; or &lt;a href="https://www.npmjs.com/search?q=keywords:jupyterlab-extension"&gt;npmjs.com&lt;/a&gt;&lt;/p&gt;
&lt;p&gt;My favorite JupyterLab extension is &lt;a href="https://github.com/jwkvam/jupyterlab-vim"&gt;jupyterlab-vim&lt;/a&gt; — it lets you fully use Vim keybindings inside JupyterLab!&lt;/p&gt;
&lt;h2 id="classic-notebook-extensions-nbextension"&gt;Classic Notebook extensions (nbextension)&lt;/h2&gt;
&lt;figure&gt;
&lt;img alt="Table of Contents nbextension" src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/99-ways-to-extend-the-jupyter-ecosystem/images/002-1_PNV_xAcapA3ONJZwuCKnfg.mp4" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Table of Contents nbextension&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;When people think of ‘the notebook interface’, they are probably thinking of the &lt;a href="https://github.com/jupyter/notebook"&gt;classic Jupyter Notebook&lt;/a&gt;. You can extend any aspect of the notebook user experience with &lt;em&gt;nbextension&lt;/em&gt;s. These are little bits of client-side JavaScript that allow you to add / change functionality as you wish. They are the Classic Notebook equivalent to JupyterLab extensions.&lt;/p&gt;
&lt;p&gt;The Jupyter Notebook documentation has information on how to &lt;a href="https://jupyter-notebook.readthedocs.io/en/stable/examples/Notebook/Distributing%20Jupyter%20Extensions%20as%20Python%20Packages.html#Installation-of-Jupyter-Extensions"&gt;install&lt;/a&gt; or &lt;a href="https://jupyter-notebook.readthedocs.io/en/stable/extending/frontend_extensions.html"&gt;develop&lt;/a&gt; extensions. The &lt;a href="https://jupyter-contrib-nbextensions.readthedocs.io/en/latest/"&gt;Unofficial Jupyter Notebook extensions&lt;/a&gt; repository has a lot of popular extensions and a GUI extension manager you can use to install nbextensions.&lt;/p&gt;
&lt;p&gt;My favorite nbextension provides a collapsible &lt;a href="https://jupyter-contrib-nbextensions.readthedocs.io/en/latest/nbextensions/toc2/README.html"&gt;Table of Contents&lt;/a&gt; for your notebooks.&lt;/p&gt;
&lt;h2 id="notebook-server-extensions-serverextension"&gt;Notebook Server Extensions (serverextension)&lt;/h2&gt;
&lt;p&gt;Unlike JupyterLab or nbextensions, Jupyter Notebook &lt;a href="https://jupyter-notebook.readthedocs.io/en/stable/extending/handlers.html"&gt;Server extensions&lt;/a&gt; are written in Python to add some serverside functionality. There are two primary use cases for server extensions.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="LaTeX previews in JupyterLab" src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/99-ways-to-extend-the-jupyter-ecosystem/images/003-1_c5wlcF6SXXlblgiPrQzA0A.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;LaTeX previews in JupyterLab&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;The first use case is to provide a backend for a particular JupyterLab or classic notebook extension. An example is the &lt;a href="https://github.com/jupyterlab/jupyterlab-latex"&gt;jupyterlab-latex&lt;/a&gt; JupyterLab extension, which provides live previews of LaTeX files in JupyterLab. It has a frontend JupyterLab extension to integrate with the JupyterLab text editor, and a backend serverextension component that actually runs the LaTeX commands to produce the output displayed to you.&lt;/p&gt;
&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/99-ways-to-extend-the-jupyter-ecosystem/images/004-1_1EEekwYT-JCARQoQgUJdJg.mp4" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;p&gt;The second use case is to provide any user interface backed by any kind of server side processing. Server extensions can function as arbitrary &lt;a href="https://tornadoweb.org"&gt;Tornado&lt;/a&gt; HTTP handlers — so any web application you can think of, you can write as a Jupyter serverextension. An example is &lt;a href="https://jupyterhub.github.io/nbgitpuller/"&gt;nbgitpuller&lt;/a&gt;, which provides UI and mechanisms to distribute notebooks from git repositories to your users in a transparent way.&lt;/p&gt;
&lt;p&gt;My favorite here is &lt;a href="https://github.com/jupyterhub/jupyter-server-proxy/tree/master/contrib/rstudio"&gt;jupyter-rsession-proxy&lt;/a&gt;, which lets you run RStudio in JupyterHub environments!&lt;/p&gt;
&lt;h2 id="jupyter-kernels"&gt;Jupyter Kernels&lt;/h2&gt;
&lt;p&gt;You might be most familiar with using Jupyter notebooks with Python, but you can use &lt;a href="https://github.com/jupyter/jupyter/wiki/Jupyter-kernels"&gt;a ton of other languages&lt;/a&gt; when writing your notebook: &lt;a href="https://irkernel.github.io/"&gt;R&lt;/a&gt;, &lt;a href="https://github.com/JuliaLang/IJulia.jl"&gt;Julia&lt;/a&gt;, &lt;a href="https://github.com/n-riesco/ijavascript"&gt;JavaScript&lt;/a&gt;, &lt;a href="https://github.com/calysto/octave_kernel"&gt;Octave&lt;/a&gt;, &lt;a href="https://github.com/apache/incubator-toree"&gt;Scala/Spark&lt;/a&gt;, &lt;a href="https://github.com/QuantStack/xeus-cling"&gt;interactive C++&lt;/a&gt;, &lt;a href="https://github.com/takluyver/bash_kernel"&gt;bash&lt;/a&gt;, or even &lt;a href="https://github.com/calysto/matlab_kernel"&gt;Matlab&lt;/a&gt;! These are called &lt;em&gt;kernels&lt;/em&gt;, and they speak the language agnostic &lt;a href="https://jupyter-client.readthedocs.io/en/stable/messaging.html"&gt;Jupyter protocol&lt;/a&gt; over &lt;a href="http://zeromq.org/"&gt;zeromq&lt;/a&gt;. You can write a new kernel for your language &lt;a href="https://jupyter-client.readthedocs.io/en/stable/kernels.html"&gt;by directly implementing&lt;/a&gt; the Jupyter protocol, by wrapping it with the &lt;a href="https://github.com/Calysto/metakernel"&gt;metakernel&lt;/a&gt; project, or using C++ bindings via &lt;a href="https://github.com/QuantStack/xeus"&gt;Xeus&lt;/a&gt;. Once a kernel exists, it seamlessly works with any Jupyter frontend — classic notebook, JupyterLab, &lt;a href="http://nteract.io"&gt;nteract&lt;/a&gt;, the &lt;a href="https://github.com/jupyter/jupyter_console"&gt;terminal jupyter console&lt;/a&gt;, the &lt;a href="https://qtconsole.readthedocs.io/en/stable/"&gt;graphical Qt Console&lt;/a&gt; , etc.&lt;/p&gt;
&lt;p&gt;My favorite kernel is &lt;a href="https://www.kernel.org/"&gt;the linux kernel&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="ipython-magics"&gt;IPython Magics&lt;/h2&gt;
&lt;p&gt;If you’ve written &lt;code&gt;%matplotlib inline&lt;/code&gt;in a notebook, you have used an &lt;a href="https://ipython.readthedocs.io/en/stable/interactive/magics.html"&gt;IPython magic&lt;/a&gt;. These are &lt;em&gt;almost&lt;/em&gt; like macros for Python — you can write custom code that parses the rest of the line (or cell), and do whatever it is that you want.&lt;/p&gt;
&lt;p&gt;Line magics start with one &lt;code&gt;%&lt;/code&gt; symbol and take some action based on the rest of the line. For example, &lt;code&gt;%cd somedirectory&lt;/code&gt; changes the current directory of the python process. Cell magics start with &lt;code&gt;%%&lt;/code&gt; and operate on the entire cell contents after it. &lt;a href="https://ipython.readthedocs.io/en/stable/interactive/magics.html#magic-timeit"&gt;&lt;code&gt;%%timeit&lt;/code&gt;&lt;/a&gt; is probably the most famous – it’ll run the code a number of times and report stats on how long it takes to run.&lt;/p&gt;
&lt;p&gt;You can also &lt;a href="https://ipython.readthedocs.io/en/stable/config/custommagics.html"&gt;build your own magic command&lt;/a&gt; that integrates with IPython. For example, the &lt;a href="https://github.com/catherinedevlin/ipython-sql"&gt;ipython-sql&lt;/a&gt; package provides the&lt;code&gt;%%sql&lt;/code&gt; magic command for working seamlessly with databases. However, remember that in contrast to the extensions listed so far, IPython magics only work with the IPython kernel.&lt;/p&gt;
&lt;p&gt;My favorite use of IPython magics is this &lt;a href="/posts/2018/i-python-you-r-we-julia/"&gt;blog post&lt;/a&gt; by &lt;a href="https://matthiasbussonnier.com/"&gt;Matthias Bussonnier&lt;/a&gt;, which makes great use of custom magics to seamlessly integrate Python, R, C and Julia in the same notebook.&lt;/p&gt;
&lt;h2 id="ipython-widgets-ipywidgets"&gt;IPython Widgets (ipywidgets)&lt;/h2&gt;
&lt;figure&gt;
&lt;img alt="Play with plot options with dropdown. Courtesy Towards Data Science by Will Koehrsen" src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/99-ways-to-extend-the-jupyter-ecosystem/images/005-1_gZRZp4X1SM1tiC3FHpDUTw.mp4" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Play with plot options with dropdown. Courtesy &lt;a href="https://towardsdatascience.com/interactive-controls-for-jupyter-notebooks-f5c94829aee6"&gt;Towards Data Science&lt;/a&gt; by &lt;a href="http://twitter.com/@koehrsen_will"&gt;Will Koehrsen&lt;/a&gt;&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;IPython Widgets (ipywidgets) provide interactive GUI widgets for Jupyter notebooks and the IPython kernel. They let you and the people you share your notebooks with explore various options in your code with GUI elements rather than having to modify code. Coupled with something like &lt;a href="https://github.com/QuantStack/voila"&gt;voila&lt;/a&gt;, you can make dashboard-like applications for other people to consume without realizing it was created completely with a Jupyter Notebook!&lt;/p&gt;
&lt;p&gt;You can &lt;a href="https://ipywidgets.readthedocs.io/en/stable/examples/Widget%20Custom.html"&gt;build your own custom widgets&lt;/a&gt; to provide domain-specific interactive visualizations. For example, you can interactively visualize maps with &lt;a href="https://github.com/jupyter-widgets/ipyleaflet"&gt;ipyleaflet&lt;/a&gt;, use &lt;a href="https://github.com/InsightSoftwareConsortium/itk-jupyter-widgets"&gt;itk-jupyter-widget&lt;/a&gt; to explore image segmentation/registration problems interactively, or model 3D objects with &lt;a href="https://github.com/jupyter-widgets/pythreejs"&gt;pythreejs&lt;/a&gt;.&lt;/p&gt;
&lt;p&gt;Check out the &lt;a href="https://github.com/nteract/vdom"&gt;vdom project&lt;/a&gt; for a more &lt;em&gt;react&lt;/em&gt;ive take on the same problem space, and &lt;a href="https://github.com/QuantStack/xwidgets"&gt;xwidgets&lt;/a&gt; for a C++ implementation&lt;/p&gt;
&lt;h2 id="contents-manager"&gt;Contents Manager&lt;/h2&gt;
&lt;p&gt;Whenever you open or save a notebook or file through the web interface, a &lt;a href="https://jupyter-notebook.readthedocs.io/en/stable/extending/contents.html"&gt;ContentsManager&lt;/a&gt; decides what actually happens. By default it loads and saves files from the local filesystem, but a custom contents manager could do whatever it wants. A popular use case it to load/save contents from somewhere other than the local filesystem — &lt;a href="https://github.com/danielfrg/s3contents"&gt;Amazon S3 / Google Cloud Storage&lt;/a&gt;, &lt;a href="https://github.com/quantopian/pgcontents"&gt;PostgreSQL&lt;/a&gt;, &lt;a href="https://jcrist.github.io/hdfscm/"&gt;HDFS&lt;/a&gt;, etc. When using one of these, you can load / save notebooks &amp;amp; files via the web interface as if they are on your local filesystem! This is extremely useful if you are already using any of these to store your data.&lt;/p&gt;
&lt;p&gt;My favorite contents manager is &lt;a href="https://github.com/mwouts/jupytext"&gt;Jupytext&lt;/a&gt;. It does some magic during save/load to give you a &lt;code&gt;.py&lt;/code&gt; equivalent of your &lt;code&gt;.ipynb&lt;/code&gt;, and keeps them in sync. You can explore code interactively in your notebook, then open the &lt;code&gt;.py&lt;/code&gt; file in an IDE to do some heavy text editing, and automatically get all your changes back in your notebook when you open it again. It’s quite magical.&lt;/p&gt;
&lt;figure&gt;
&lt;img alt="Jupytext: .ipynb or .py? why not both!" src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/99-ways-to-extend-the-jupyter-ecosystem/images/006-1_DQA3TQcVSJBtWxSMD-sq-g.mp4" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Jupytext: .ipynb or .py? why not both!&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;h2 id="extending-jupyterhub"&gt;Extending JupyterHub&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://jupyter.org/hub"&gt;JupyterHub&lt;/a&gt; is a multi-user application for spawning notebooks &amp;amp; other interactive web applications, designed for use in classrooms, research labs and companies. These organizations probably have other systems they are using, and JupyterHub needs to integrate strongly with them. Here is a non-exhaustive list of ways JupyterHub can be extended.&lt;/p&gt;
&lt;h3 id="authenticators"&gt;Authenticators&lt;/h3&gt;
&lt;p&gt;JupyterHub is a &lt;em&gt;multi-user&lt;/em&gt; application, so users need to log in somehow — ideally the same way they log in to every other application in their organization. The &lt;a href="https://jupyterhub.readthedocs.io/en/stable/reference/authenticators.html"&gt;authenticator&lt;/a&gt; is responsible for this. Authenticators already exist for &lt;a href="https://github.com/jupyterhub/jupyterhub/wiki/Authenticators"&gt;many popular authentication services&lt;/a&gt; — &lt;a href="https://github.com/jupyterhub/ldapauthenticator"&gt;LDAP&lt;/a&gt;, &lt;a href="https://github.com/jupyterhub/oauthenticator"&gt;OAuth&lt;/a&gt; (Google, GitHub, CILogon, Globus, Okta, Canvas, etc), most &lt;a href="https://github.com/jupyterhub/ltiauthenticator"&gt;LMS with LTI&lt;/a&gt; , &lt;a href="https://github.com/bluedatainc/jupyterhub-samlauthenticator"&gt;SAML&lt;/a&gt;, &lt;a href="https://github.com/mogthesprog/jwtauthenticator"&gt;JWT&lt;/a&gt;, plain &lt;a href="http://github.com/jupyterhub/nativeauthenticator"&gt;usernames &amp;amp; passwords&lt;/a&gt;, &lt;a href="https://github.com/jupyterhub/jupyterhub/blob/5e60582ef319c17591f51ba78cd78719dd8fb179/jupyterhub/auth.py#L774"&gt;linux users&lt;/a&gt;, etc. You can &lt;a href="https://jupyterhub.readthedocs.io/en/stable/reference/authenticators.html#how-to-write-a-custom-authenticator"&gt;write your own&lt;/a&gt; or customize one that exists very easily, so whatever your authentication needs — JupyterHub has you covered.&lt;/p&gt;
&lt;h3 id="spawners"&gt;Spawners&lt;/h3&gt;
&lt;p&gt;Using pluggable spawners, you can start a Jupyter Notebook Server for each user in many different ways. You might want them to spawn on a node with &lt;a href="https://github.com/jupyterhub/dockerspawner"&gt;docker containers&lt;/a&gt;, scale them out with &lt;a href="https://github.com/jupyterhub/kubespawner"&gt;Kubernetes&lt;/a&gt;, use it on your &lt;a href="https://github.com/jupyterhub/batchspawner"&gt;HPC cluster&lt;/a&gt;, have them run along &lt;a href="https://github.com/jcrist/yarnspawner"&gt;your Hadoop / Spark cluster&lt;/a&gt;, contain them with &lt;a href="https://github.com/jupyterhub/systemdspawner"&gt;systemd&lt;/a&gt;, simply run them as &lt;a href="https://github.com/jupyterhub/jupyterhub/blob/5e60582ef319c17591f51ba78cd78719dd8fb179/jupyterhub/spawner.py#L1164"&gt;different linux users&lt;/a&gt; or in many other possible ways. The spawners themselves are usually extremely configurable, and of course you can write your own.&lt;/p&gt;
&lt;h3 id="services"&gt;Services&lt;/h3&gt;
&lt;p&gt;Often you want to provide additional services to your JupyterHub users — &lt;a href="https://github.com/jupyterhub/jupyterhub/tree/master/examples/cull-idle"&gt;cull their servers&lt;/a&gt; when idle, or allow them to &lt;a href="https://github.com/OpenHumans/jupyter-gallery"&gt;publish shareable notebooks&lt;/a&gt;. You can run a &lt;a href="https://jupyterhub.readthedocs.io/en/stable/reference/services.html"&gt;JupyterHub service&lt;/a&gt; to provide these — or similar — services. Users can make requests to them with their JupyterHub identities, and the services can make &lt;a href="https://jupyterhub.readthedocs.io/en/stable/reference/rest.html"&gt;API calls&lt;/a&gt; to JupyterHub too. These can be arbitrary processes or web services — &lt;a href="https://github.com/jupyterhub/binderhub"&gt;BinderHub&lt;/a&gt; is implemented as a JupyterHub service, for example.&lt;/p&gt;
&lt;h2 id="nbconvert-exporter"&gt;NBConvert Exporter&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://github.com/jupyter/nbconvert"&gt;nbconvert&lt;/a&gt; converts between the notebook format and various other formats — if you’ve exported your notebook to PDF, LaTeX, HTML, or used &lt;a href="https://nbviewer.jupyter.org/"&gt;nbviewer&lt;/a&gt;, you have used nbconvert. It has an exporter for each format it exports to, and you can &lt;a href="https://nbconvert.readthedocs.io/en/latest/external_exporters.html"&gt;write your own&lt;/a&gt; to export to a new format — or to just massively customize an existing export format. If you’re performing complex conversion operations involving notebooks, you might find writing an exporter to be the cleanest way to accomplish your goals.&lt;/p&gt;
&lt;p&gt;My happiest moment when researching for this blog post is finding out that a &lt;a href="https://github.com/m-rossi/jupyter-docx-bundler"&gt;docx exporter&lt;/a&gt; exists.&lt;/p&gt;
&lt;h2 id="bundler-extensions"&gt;Bundler Extensions&lt;/h2&gt;
&lt;figure&gt;
&lt;img alt="Discoverable way to enable nbconvert exporters" src="https://jasongrout.github.io/medium-archive/pelican/posts/2019/99-ways-to-extend-the-jupyter-ecosystem/images/007-1_1AzOXjN6O9SrqptsxVLeqQ.webp" loading="lazy" data-body-image=""&gt;
&lt;figcaption&gt;Discoverable way to enable nbconvert exporters&lt;/figcaption&gt;
&lt;/figure&gt;
&lt;p&gt;Bundler extensions let you add entries to the &lt;em&gt;Download as&lt;/em&gt; item in the menu bar. They are often paired with an nbconvert exporter to make the exporter more discoverable, though you can also write a custom bundler extension to do any kind of custom processing of a notebook before downloading. For example, &lt;a href="https://github.com/choldgraf/nbreport"&gt;nbreport&lt;/a&gt; provides a bundler extension that cleans up the notebook in a way suitable for viewing as a report &amp;amp; exports it as HTML.&lt;/p&gt;
&lt;h2 id="repo2docker"&gt;Repo2Docker&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://github.com/jupyter/repo2docker"&gt;repo2docker&lt;/a&gt; turns git (and other) repositories into reproducible, data science focused docker images. &lt;a href="https://mybinder.org"&gt;mybinder.org&lt;/a&gt; (and other &lt;a href="https://github.com/jupyterhub/binderhub"&gt;binderhub&lt;/a&gt; installations) rely on it to build and launch interactive Jupyter/RStudio sessions from git repositories. There are currently two ways to extend repo2docker.&lt;/p&gt;
&lt;h3 id="buildpacks"&gt;BuildPacks&lt;/h3&gt;
&lt;p&gt;repo2docker looks at the contents of the repository to decide how to build it. For example, if there is a &lt;code&gt;requirements.txt&lt;/code&gt; it sets up a miniconda environment to install python packages into, while if there is an &lt;code&gt;install.R&lt;/code&gt; file it makes sure R/RStudio is installed. &lt;a href="https://repo2docker.readthedocs.io/en/latest/architecture.html#buildpacks"&gt;Writing a new BuildPack&lt;/a&gt; lets you extend this behavior to add support for your favorite language, or customize how an existing language is built.&lt;/p&gt;
&lt;h3 id="contentproviders"&gt;ContentProviders&lt;/h3&gt;
&lt;p&gt;The &lt;em&gt;repo&lt;/em&gt; part of repo2docker is a misnomer — you can turn anything into a docker image. Currently, it supports &lt;code&gt;git&lt;/code&gt;, &lt;code&gt;local folder&lt;/code&gt; and &lt;a href="https://zenodo.org/"&gt;zenodo&lt;/a&gt; repositories — but you can &lt;a href="https://repo2docker.readthedocs.io/en/latest/architecture.html#contentproviders"&gt;add support&lt;/a&gt; for your favorite source of reproducible code by making a new ContentProvider!&lt;/p&gt;
&lt;h2 id="is-that-all"&gt;Is that all?&lt;/h2&gt;
&lt;p&gt;Of course not? The Jupyter ecosystem is vast, and no one blog post can cover them all. This blog post is already missing a few — &lt;a href="https://jupyter-enterprise-gateway.readthedocs.io/en/latest/system-architecture.html#enterprise-gateway-process-proxy-extensions"&gt;enterprise gateway&lt;/a&gt;, &lt;a href="http://tljh.jupyter.org/en/latest/contributing/plugins.html"&gt;TLJH Plugins&lt;/a&gt;, etc. As time marches on, there will be newer components and newer ways of extending things that have not even been imagined yet. Leave a comment about what else is missing here.&lt;/p&gt;
&lt;p&gt;Look forward to seeing what kinda beautiful extensions y’all create!&lt;/p&gt;
</content><category term="JupyterLab"/></entry><entry><title>Outreachy &amp; Jupyter: Supporting diversity in open communities</title><link href="https://jasongrout.github.io/medium-archive/pelican/posts/2018/outreachy-jupyter-supporting-diversity-in-open/" rel="alternate"/><published>2018-11-28T17:26:00+00:00</published><updated>2018-11-28T18:35:00+00:00</updated><author><name>yuvipanda</name></author><id>tag:jasongrout.github.io,2018-11-28:/medium-archive/pelican/posts/2018/outreachy-jupyter-supporting-diversity-in-open/</id><summary type="html">&lt;p&gt;Project Jupyter has accepted 2 interns through the Outreachy program, which supports open community members from under-represented…&lt;/p&gt;
</summary><content type="html">&lt;p&gt;&lt;img src="https://jasongrout.github.io/medium-archive/pelican/posts/2018/outreachy-jupyter-supporting-diversity-in-open/images/001-1_OsCmvuJ-lLeC7UtWK8CkNA.webp" alt="" loading="lazy" data-body-image=""&gt;&lt;/p&gt;
&lt;p&gt;Project Jupyter has accepted 2 interns through the &lt;a href="https://www.outreachy.org/"&gt;Outreachy program&lt;/a&gt;, which supports open community members from under-represented backgrounds. Our Outreachy interns will work on important problems in the JupyterHub community with help from 2 mentors from December 2018 through March 2019. This is a short post describing why Jupyter is committing to this program, as well as what we’ll be working on with our Outreachy participants.&lt;/p&gt;
&lt;p&gt;&lt;em&gt;We are grateful to the Berkeley Institute for Data Science &amp;amp; NumFocus for jointly sponsoring our interns. This article is cross-posted with the&lt;/em&gt; &lt;a href="https://bids.berkeley.edu/news/outreachy-jupyter-supporting-diversity-open-communities"&gt;&lt;em&gt;Berkeley Institute for Data Science blog&lt;/em&gt;&lt;/a&gt;.&lt;/p&gt;
&lt;h2 id="why-focus-on-diversity-and-inclusion"&gt;Why focus on diversity and inclusion?&lt;/h2&gt;
&lt;p&gt;The tech industry isn’t doing great on diversity, and the Open Source community &lt;a href="https://www.wired.com/2017/06/diversity-open-source-even-worse-tech-overall/"&gt;is doing worse&lt;/a&gt; . This limits the community’s ability to tackle challenging, diverse projects. The amount of responsibility and power held by open source communities is increasing, so we have a responsibility to be more representative of the diversity present in our world. Many projects grow their community by encouraging people to start making small patches on their own. However, &lt;a href="https://www.ashedryden.com/blog/the-ethics-of-unpaid-labor-and-the-oss-community"&gt;requiring uncompensated work&lt;/a&gt; is a big barrier to getting more diverse representation in our open source communities. Paid internships with dedicated mentors are a great way to help people break through this particular barrier and make the open source community more diverse.&lt;/p&gt;
&lt;h2 id="what-is-outreachy"&gt;What is Outreachy?&lt;/h2&gt;
&lt;p&gt;&lt;a href="http://outreachy.org"&gt;Outreachy&lt;/a&gt; is an internship program coordinated by the &lt;a href="https://sfconservancy.org/"&gt;Software Freedom Conservancy&lt;/a&gt;. It is run twice a year with a goal to bring people from underrepresented backgrounds in tech into open source projects. 21 open source organizations are participating in this round, and will be working with a total of 46–47 interns. Importantly, these are &lt;em&gt;paid&lt;/em&gt; internships, which make them more viable for a much broader slice of the population. Jupyter and BIDS will both contribute funding and mentorship for two Outreachy interns.&lt;/p&gt;
&lt;h2 id="our-interns-and-mentors"&gt;Our Interns and Mentors&lt;/h2&gt;
&lt;p&gt;We have two amazing interns for this Outreachy round, working on important projects in the JupyterHub ecosystem. Below we’ll describe the projects that they’ll work on.&lt;/p&gt;
&lt;h2 id="a-highly-available-proxy-for-jupyterhub-link"&gt;A highly Available Proxy for JupyterHub (&lt;a href="https://github.com/jupyterhub/outreachy/blob/master/ideas/traefik-jupyterhub-proxy.rst"&gt;link&lt;/a&gt;)&lt;/h2&gt;
&lt;p&gt;Georgiana Dolocan will be mentored by Min RK and Yuvi Panda in building &lt;a href="https://github.com/jupyterhub/outreachy/blob/master/ideas/traefik-jupyterhub-proxy.rst"&gt;a highly available &amp;amp; scalable proxy for JupyterHub&lt;/a&gt; using the &lt;a href="https://traefik.io/"&gt;Traefik&lt;/a&gt; project. This will help JupyterHub deployments scale more easily for thousands of active users with minimal service disruptions. Georgiana will be working from Bucharest, Romania where she likes to paint and explore the city and the villages nearby alongside her dog and camera.&lt;/p&gt;
&lt;h2 id="improvements-to-user-management-link"&gt;Improvements to user management (&lt;a href="https://github.com/jupyterhub/outreachy/blob/master/ideas/native-jupyterhub-user-management.rst"&gt;link&lt;/a&gt;)&lt;/h2&gt;
&lt;p&gt;&lt;a href="https://leportella.com/"&gt;Leticia Portella&lt;/a&gt; will be mentored by Yuvi Panda and Min RK in building &lt;a href="https://github.com/jupyterhub/outreachy/blob/master/ideas/native-jupyterhub-user-management.rst"&gt;better native user management features&lt;/a&gt; into JupyterHub. Small to medium installations of JupyterHub that do not want to depend entirely on an external authentication provider will benefit greatly from this. Leticia will be working from Dublin, Ireland (although she used to live in Florianópolis, Brazil), where she likes to read a lot (especially The Chronicles of Ice and Fire), to swim, and to work on her podcast (the first Brazilian podcast specialized in Data Science topics), &lt;a href="https://pizzadedados.com/"&gt;Pizza de Dados&lt;/a&gt; (Data Pizza, free translation).&lt;/p&gt;
&lt;h2 id="about-our-mentors"&gt;About our mentors&lt;/h2&gt;
&lt;p&gt;Outreachy requires more than just funding, but also mentorship. Two members of the JupyterHub community have offered their time to help mentor our Outreachy contributors, some information about them is below!&lt;/p&gt;
&lt;p&gt;&lt;a href="https://github.com/minrk"&gt;Min RK&lt;/a&gt; is a research engineer at Simula Research Laboratory in Norway and has been working on the IPython and Jupyter open source projects since joining in 2006 as an undergraduate in Engineering Physics at Santa Clara University with Brian Granger, one of the founders of the Jupyter project. Min currently focuses on JupyterHub, and is excited to welcome new folks to the Jupyter team.&lt;/p&gt;
&lt;p&gt;&lt;a href="http://yuvi.in"&gt;Yuvi Panda&lt;/a&gt; is an operations engineer at University of California, Berkeley. He works at the &lt;a href="https://bids.berkeley.edu/"&gt;Berkeley Institute for Data Science&lt;/a&gt;, where he maintains the JupyterHub infrastructure for the &lt;a href="https://data.berkeley.edu/"&gt;Division of Data Sciences&lt;/a&gt;. He has been involved in various Open Source communities in the last ten years, spending time in the GNOME, Wikimedia &amp;amp; Jupyter communities. His life was changed drastically by participating as a student in Google Summer of Code 2010, and he’s excited to give back to the Open Source community.&lt;/p&gt;
&lt;p&gt;We would like to thank everyone who applied to JupyterHub for this round of Outreachy. We received a number of robust proposals, out of which we were only able to accept two. JupyterHub mentors spent quite a lot of time mentoring candidates during the application period, in reviewing their pull requests, and giving them feedback on their proposals. We look forward to bringing these new open source contributors into our community, and hope that we can set a path that other open source projects may follow in the future. Open source works best when it is diverse and inclusive, we think this is a small step in that direction.&lt;/p&gt;
&lt;p&gt;Thanks to &lt;a href="https://people.eecs.berkeley.edu/~odemasi/"&gt;Orianna DeMasi&lt;/a&gt;, &lt;a href="https://sastoudt.github.io/"&gt;Sara Stoudt&lt;/a&gt;, &lt;a href="https://predictablynoisy.com/"&gt;Chris Holdgraf&lt;/a&gt; &amp;amp; others for contributing heavily to this blog post.&lt;/p&gt;
</content><category term="community"/><category term="diversity"/><category term="Outreachy"/></entry></feed>