Day 2 — Working on the Yens Cluster with JupyterHub

Learning Goals

By the end of today you will be able to

  • understand what a “path” is and why it matters;

  • create and activate Python virtual environments on Yens;

  • understand how a virtual environment can assist with reproducible code;

  • open and run Jupyter notebooks in JupyterHub;

  • understand how to manage passwords and other “secrets” in your code;

  • run API calls to OpenAI from JupyterHub.


Exercise 0: Log in using ssh

Open a local terminal and use ssh to connect to the Yens.

  • What server are you on?
  • What directory are you in?
  • What files are in your current directory?
  • Do you see the files from yesterday?

Exercise 1: Accessing JupyterHub

1: Open the Hub

Choose any of the following links to access JupyterHub on the Yens cluster:

You should see the folders you created on the previous day, including your exercise notebooks.

2: Start a New Notebook and Terminal
  • Click the blue “+” to open a Python 3 notebook.

  • Click the Terminal icon to launch a shell (we’ll use it for environment commands).

Exercise 1.1: First Notebook Cells

Copy each of the following into separate cells in your notebook, then run them using Shift + Enter.

# Print "Hello, World!"
print("Hello, World!")  
# Import the math module and print the value of pi
# Use the math module
import math
print(math.pi)
numbers = [1, 2, 3, 4, 5]
print(sum(numbers))

Exercise 1.2: Saving + Running Python Scripts

Open the terminal tab you started above and run:

touch interactive.py

Now find the script in your finder, double click on it, and paste the code you wrote above into interactive.py (save the file once you’re done).

To verify that it works, run

python interactive.py

and check that you get the following output:

Hello, World!
3.141592653589793
15

If you ever need to run a Python script in the future (for instance, later in this course), you can now use the script you created as a starting point.

Exercise 1.3: Jupyter Terminal Basics

Open the terminal tab you started above and try:

# List files in your home directory
ls

# navigate into your pokemon_images directory
cd technical_data_important

# List files in the pokemon_images directory
ls

# Find out which python version you are using
which python3

Exercise 1.4: Display a Pokémon Image

  1. Locate any PNG in your images folder (use the file browser or ls).

You can double click on it to view it natively in JupyterHub.

  1. Replace /path/to/your/pokemon_image.png below with that full path.
from PIL import Image
from IPython.display import display      
# Load and display a Pokémon image
img_path = '/path/to/your/pokemon_image.png'  # Replace with your image
img = Image.open(img_path)
display(img)

Exercise 1.5: Manipulate an image on the terminal

  1. Let’s manipulate an image on the terminal using a tool called imagemagick.

Go to your terminal and type:

module load imagemagick

Pick the Pokemon you displayed above and flip it upside down like this:

magick /path/to/your/pokemon_image.png -flip /path/to/your/output_image.png
  1. Did it work? Go check in your notebook.

  2. Type the following command in your terminal:

which magick

Does it look the same as when you did which python3?

Whiteboarding

I’ll do my best to build some intuition for paths for you on the whiteboard!

Exercise 2: Understand Paths

Let’s explore your own path, and see how it can change. Earlier, you ran, module load imagemagick, and typed which magick.

Exercise 2.1

In your (Jupyter) terminal, type:

which python
which magick
echo $PATH

The $PATH (anything with a $ in front, actually) is a variable. Find the python and magick programs in your $PATH. The command echo is just like print.

Exercise 2.2

We used module load to load the imagemagick module. Let’s explore the module command more. Try the following:

module list

What is listed? Does it make sense?

Now try this:

module unload imagemagick

Is magick there for you to use? Verify by trying the following:

  • running the command to flip your Pokemon image
  • using which to see if the command is available

Take a look at your $PATH – what changed?

Run module avail – what do you see? Try and load a specific version of R, and verify that it works as you expect. Why would you care about versions?

Exercise 2.3

Let’s think a little bit more about Python in particular.

  • Go to the terminal within Jupyter. Which python3 do you see? How do you know?
  • Go to the terminal you get from logging in with ssh. Which python3 do you see? How do you know?

Whiteboarding

All this path and version stuff is important for reproducibility. Let’s take a beat to think through what reproducibility means in research.

Exercise 3: Creating a Python Virtual Environment

A Python virtual environment is a self-contained directory that includes its own Python installation and packages. It allows you to manage dependencies separately for different projects.

More detailed directions can be found on our website.

First, we will create a dedicated directory for our work and set up our environment inside it.

  1. Open a new Terminal from the JupyterHub Launcher.

  2. Run the following commands in your terminal to create a new directory and navigate into it:

mkdir day2
cd day2
  1. Next, create the virtual environment. We’ll name it venv:
    /usr/bin/python3  -m venv venv
    # This should make a new folder called venv in your day2 directory
    

Find the path to python3, and echo the entire $PATH. Where is python3?

  1. Activate the environment:
source venv/bin/activate

This runs a script that’s located in the ./venv/bin directory called activate. The bin directory doesn’t mean like, a literal bin. It’s short for binary, things that can be executed as programs, as opposed to data or configuration files.

Tip: You will know the activation was successful when you see (venv) at the beginning of your terminal prompt. This indicates that the virtual environment is active.

  1. Check which Python version is being used in your virtual environment:
which python3

The output should point to the Python executable inside your day2/venv directory.

What is in your path now? What changed?

Run deactivate to exit the virtual environment, and then check which python3 and your path again.

Now, reactivate the environment.

Step 2: Installing Packages

Your environment is activated, so now you can install packages using pip. Let’s try it.

pip install dotenv

This library is now installed in this environment. You can load it while the environment is activated, but it’s not installed for anyone else. Test it out! Try importing numpy and dotenv in the Jupyter terminal with your virtual environment activated and deactivated.

Step 3: Integrating Jupyter Notebooks

Now, let’s install the ipykernel package, which provides the tools to connect your environment to Jupyter:

pip install ipykernel

Now, create a new Jupyter kernel linked to your virtual environment. Replace <kernel_name> with day2-venv. Make sure you’re in your active venv when you run this command!

python -m ipykernel install --user --name=<kernel_name>

In the Jupyter interface, go to your day2 folder, and start a new notebook. Name it Interactive.ipynb. Change the kernel to day2-env.

You should be able to run:

import dotenv

If you can’t, let’s get help!

Whiteboarding

Let’s whiteboard out a real task, where we use an external large language model to process a SEC filing.


Exercise 4: Calling OpenAI API

Go ahead and install the openai package in your day2 venv. Once you’re done, we’re ready to call the OpenAI API.

Here’s how we initialize calls to the OpenAI API:

from openai import OpenAI
client = OpenAI(api_key=<your api key>)

DO NOT put your API key in here. Instead, we’re going to use an API key stored in a ‘hidden’ file, and load it in as an environment variable.

4.1 Try out the dotenv library

We can use os.getenv to retrieve an environment variable. This a variable (like $PATH) that exists on the shell, and you can also read it from Python.

import os
os.getenv("PATH")

We can use the dotenv library to put things into our environment variables, which is a good practice for storing things like API keys.

We prepared a hidden file for you – take a look at it in the terminal.

cat /scratch/shared/rf_bootcamp_2025/.env

Now we’re going to use it in Jupyter.

from dotenv import load_dotenv
load_dotenv('/scratch/shared/rf_bootcamp_2025/.env')

This should load the environment variable – test by seeing if OPENAI_API_KEY is there in your environment.

Look at the code below – we can publish this on the internet no problem, because our API key isn’t in there!

client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))

4.2 OpenAI: “Hello, World!”

We can try this simple example to confirm it works!

completion = client.chat.completions.create(
    model="gpt-4.1-nano",
    messages=[
        {"role": "system", "content": "You are a helpful assistant."},
        {"role": "user", "content": "Say hello world!"}
    ]
)

# Print the model's response
print(completion.choices[0].message.content)

4.3 OpenAI with a SEC Document

We have a mirror of the SEC filings on the Yen servers. Here’s one example:

/zfs/data/NODR/EDGAR_HTTPS/edgar/data/1656998/0000950103-24-000077.txt

Go ahead and load in your notebook at take a look.

One neat thing about Jupyter is that if the document is HTML, you can render the HTML in the document itself:

from IPython.display import display, HTML

sec_doc = open(filing_path).read()
display(HTML(sec_doc))

Now, you pass the document and use the LLM to extract a useful piece of information. Try changing the system prompt to describe what you want from the document, and then just pass sec_doc as the user prompt. Here’s a simple example:

    {"role": "system", "content": "You are a terse assistant that reads SEC Form 3 documents and extracts a list of the names of all of the attorneys-in-fact."},
    {"role": "user", "content": sec_doc}

4.4 From Jupyter to Command Line

A notebook is a great place to explore, investigating data, testing things out, and developing skills. Let’s make a cell in our notebook that concisely does the following:

  • Load the relevant packages
  • Set up your OpenAI client (using dotenv)
  • Set up the path to your sample SEC document
  • Call the OpenAI model to extract information from the sample document
  • Print the results from OpenAI

Once you’ve prepared that cell, copy its contents to a new file called form3_test.py. Run it in the terminal and verify that it works.

4.5 OpenAI with Structured Outputs

We’re not out of time yet? Amazing!

We’re going to extract key information from a Form 3 filing, which is a filing company directors (“insiders”) have to submit to the SEC to disclose their financial interests (to prevent things like insider trading). Each Form 3 filing includes the insider’s name, their role(s), the company name, CIK (an index for companies and individuals filing with the SEC), and the filing date.

We want to find and return this information in a structured, standardized format (e.g., JSON or a dictionary). This will make the data easy to validate, analyze, and store for downstream use (like building a dataset or running queries).

Note that since large language models are sometimes a little unreliable, we will use structured outputs to ensure that OpenAI returns a consistent format. We can use a library called pydantic to define what that format is.

So, to recap, your tasks are:

  • Write a system prompt that’s going to extract the information listed above (i.e., the insider’s name, role, etc.).
  • Try running it.

To do this, you will want to:

  • Install pydantic in your virtual environment;
  • Build a pydantic model (as in the example below) that contains the information in your filing;
  • Write a prompt to extract the information we care about in the SEC filing;
  • Send that model to the OpenAI API and confirm you get a structured output back.

Here’s an example for a setting in which we want to query the name and price of an item on a lunch menu:

import os
from dotenv import load_dotenv
from openai import OpenAI
from pydantic import BaseModel, Field

# Load environment variables
load_dotenv('/scratch/shared/rf_bootcamp_2025/.env')
client = OpenAI(api_key=os.getenv("OPENAI_API_KEY"))

# Define pydantic model specifying desired LLM outputs
class MenuItem(BaseModel):
    name: str = Field(..., description="Name of the menu item")
    price: float = Field(..., description="Price of the menu item in dollars")

# Specify the menu to be parsed (this is a subset of yesterday's Arbuckle menu)
user_prompt = """
    cheese pizza
    with house-made dough and tomato sauce with fontina-mozzarella-provolone cheese blend — $5.20

    chef blend mushroom pizza
    with house-made pizza dough and four-cheese garlic sauce, roasted global chef blend mushroom, caramelized shallot, truffle oil, frisée, chives — $5.20

    blue cheese buffalo chicken pizza
    with house-made pizza dough and pizza sauce, Point Reyes blue cheese, buffalo chicken, green bell pepper, lemon yogurt sauce, green onion — $5.20

    pepperoni pizza
    with thinly sliced pepperoni with house-made dough and sauce with fontina-mozzarella-provolone cheese blend — $5.20

    classic cheeseburger
    with Niman Ranch 100% Angus beef, onion, lettuce, heirloom tomato, cheddar cheese, pickle mustard sauce, brioche bun, choice of fries or onion rings
    regular — $12.95

    grilled sesame chicken cabbage crunch salad
    with mixed cabbage slaw, jalapeño, red onion, green onion, carrot, cucumber, celery, crispy wonton, sesame ginger tamari soy dressing — $12.95

    Croque Monsieur Sandwich
    with Panorama Baking Co sourdough bread, béchamel sauce, parmesan cheese, Swiss cheese, provolone cheese, sliced Black Forest ham, Dijon mustard, choice of fries or onion rings — $12.95
"""

# Tell the LLM what to do
system_prompt = """
    You are at a cafeteria and you want to extract the name and price of the cheeseburger on the menu.
    Please extract the following fields:
    
    - name: The name of the cheeseburger item.
    - price: The price, in dollars, of the cheeseburger item.

    Return valid JSON matching the provided Pydantic model.
"""

# Query the API
response = client.responses.parse(
    model="gpt-4.1-nano",
    input=[
        {"role": "system", "content": system_prompt},
        {"role": "user", "content": user_prompt}
    ],
    text_format=MenuItem,
)

# Print the LLM result
print(response.output_parsed.model_dump())