Picture this. You’ve just been granted access to a shiny HPC cluster. Four GPUs waiting for you. You open your terminal, type ssh cluster:

Permission denied (publickey).

I spent an embarrassing amount of time staring at that line. My GPU hours were ticking. I had a deadline. And SSH felt like some arcane ritual that everyone else just… understood.

Turns out, it’s not that complicated. But nobody ever explains the gotchas. So here’s everything I learned the hard way.

The Padlock and the Key

SSH keys work like a padlock system. You have two files:

  • id_rsa.pub is the padlock. You hand out copies to every server you want access to. It’s public. Share it freely.
  • id_rsa is the key. This one stays on your machine. Never, ever share it.

When you connect to a server, there’s a four-step handshake. The server picks a random challenge, encrypts it with your public key, and sends it over. Only your private key can decrypt it. You send the proof back. If it checks out, you’re in.

sequenceDiagram participant L as Your Laptop (private key) participant S as Server (public key) L->>S: 1. Here is my key ID S->>L: 2. Encrypted challenge L->>S: 3. Decrypted proof S->>L: 4. Access granted! Note over L,S: Your private key never leaves your machine

The beautiful thing? Your private key never crosses the network. Not once.

Generating Keys

One command:

ssh-keygen -t rsa -b 4096 -C "you@example.com"

There’s also Ed25519, which is newer and faster. But some older clusters don’t accept it, so RSA is the safe bet. Hit enter for the default path, optionally set a passphrase, and you’re done.

Your ~/.ssh/ folder now has:

~/.ssh/
├── id_rsa          # Private key (NEVER share)
├── id_rsa.pub      # Public key (copy to servers)
├── config          # Your best friend (more on this)
├── known_hosts     # Fingerprints of servers you've visited
└── authorized_keys # On the server side: who's allowed in

Getting your public key onto the server is one more command:

ssh-copy-id -i ~/.ssh/id_rsa.pub you@cluster.example.com

The Config File

Nobody wants to type ssh -p 2222 alice@login.hpc.uni.edu every time. Life is short. So SSH has a config file.

Host mylab
  HostName login.hpc.uni.edu
  User alice
  Port 2222
  IdentityFile ~/.ssh/id_rsa

Now ssh mylab does the same thing. Six words saved, sanity preserved.

Host mylab HostName login.hpc.uni.edu User alice Port 2222 IdentityFile ~/.ssh/id_rsa ForwardAgent yes Your alias: ssh mylab Server address Username Custom port Which key to use Forward SSH agent
Anatomy of an SSH config block.

You can stack as many of these as you want. One per cluster, one per lab server, one for your Raspberry Pi at home. Whatever.

The IdentityFile Trap

This one cost me two hours.

I had two keys on my laptop: id_rsa and id_ed25519. My config pointed at the Ed25519 key. Seemed fine. I ran ssh mylab. Rejected.

But here’s the weird part: the full command worked perfectly.

ssh -p 2222 alice@login.hpc.uni.edu  # works fine
ssh mylab                             # Permission denied

Same server. Same laptop. Same keys sitting in ~/.ssh/. What gives?

The answer is subtle. When you set IdentityFile in your config, SSH only offers that one key. It doesn’t try the others. So if the server has your RSA key in its authorized_keys but your config says “use Ed25519”: you’re locked out. SSH won’t even try the key that would work.

Without the config? SSH tries all your keys, one by one. RSA gets accepted. You never notice the problem.

BROKEN IdentityFile ~/.ssh/id_ed25519 Only offers ed25519 "I don't know this key" REJECTED FIXED IdentityFile ~/.ssh/id_rsa Offers the right key "Welcome in!" ACCEPTED Wrong IdentityFile = instant rejection, even if the right key exists on your machine.
The IdentityFile trap. Wrong key in your config = locked out.

Finding the culprit

The way I figured this out was ssh -vvv. Triple-v verbose mode. It dumps everything SSH is doing: which keys it’s trying, what the server says back. Pipe it through grep and the answer jumps out:

ssh -vvv mylab 2>&1 | grep -E "Offering|Authentication"
debug1: Offering public key: /Users/you/.ssh/id_ed25519 ED25519
debug1: Authentications that can continue: publickey    # REJECTED

There it is. SSH offered id_ed25519. The server said no. Compare that to the raw command:

ssh -vvv -p 2222 alice@login.hpc.uni.edu 2>&1 | grep -E "Offering|Authentication"
debug1: Offering public key: /Users/you/.ssh/id_rsa RSA
debug1: Authentication succeeded (publickey).           # ACCEPTED

Mystery solved. Change IdentityFile in the config, done.

Hopping Through Login Nodes

Most HPC clusters won’t let you SSH directly into a compute node. Those machines sit on an internal network. You have to go through a login node first: a two-hop connection.

graph LR A["🖥️ Your Laptop
SSH agent + keys"] -->|Hop 1| B["🖧 Login Node
Public-facing"] B -->|Hop 2| C["⚡ Compute Node
GPUs here"] C -.->|Agent forwarding| A

SSH has ProxyJump for exactly this:

Host alpha
  HostName cluster.example.com
  User alice
  IdentityFile ~/.ssh/id_rsa
  ForwardAgent yes

Host alpha-gpu
  HostName gpu-node-042
  User alice
  ProxyJump alpha

ssh alpha-gpu and SSH handles both hops automatically. But there’s a catch. The second hop needs your private key to authenticate: but your private key is on your laptop, not on the login node.

That’s what ForwardAgent yes does. It lets the login node ask your laptop’s SSH agent to sign the authentication challenge. The key itself never touches the login node. Elegant.

Just make sure your agent actually has keys loaded:

ssh-add -l              # list loaded keys
ssh-add ~/.ssh/id_rsa   # add one if it's empty

A Multi-Cluster Config

Here’s roughly what a config looks like when you’re juggling multiple clusters:

Host alpha
  HostName alpha.example.com
  User alice
  IdentityFile ~/.ssh/id_rsa
  ForwardAgent yes

# This block gets updated automatically (see below)
# BEGIN alpha-gpu
Host alpha-gpu
  HostName gpu-node-042
  User alice
  ProxyJump alpha
# END alpha-gpu

Host beta
  HostName beta.example.com
  User alice
  IdentityFile ~/.ssh/id_rsa
  ForwardAgent yes

Host lab
  HostName login.lab.uni.edu
  User alice
  Port 2222
  IdentityFile ~/.ssh/id_rsa
  ForwardAgent yes

Automating salloc with Dynamic Config Updates

Here’s the part that made my life dramatically better.

Every time you run salloc, the scheduler assigns you a different compute node. gpu-node-042 today, gpu-node-117 tomorrow. You’d have to open ~/.ssh/config, find the right block, change the hostname, save. Every. Single. Time.

I got tired of this after day two.

The fix is a shell function that runs salloc, watches its output for the node name, and patches your SSH config in real-time: while the allocation is still running. By the time you see the prompt, your config is already updated.

_salloc_helper() {
  local cluster="$1" node_pattern="$2" account="$3"
  shift 3

  local defaults="--nodes=1 --ntasks=1 --cpus-per-task=4 \
    --mem=16G --time=3:00:00 --account=$account"

  # Let user args override defaults
  local arg key
  for arg in "$@"; do
    key="${arg%%=*}"
    defaults=$(echo "$defaults" | sed "s|${key}=[^ ]*||g")
  done

  local config=~/.ssh/config
  local marker_start="# BEGIN ${cluster}-gpu"
  local marker_end="# END ${cluster}-gpu"
  local node_detected=false

  ssh "$cluster" "salloc $defaults $*" 2>&1 \
    | while IFS= read -r line; do
    echo "$line"

    if ! $node_detected; then
      local node
      node=$(echo "$line" \
        | perl -ne "print \"\$1\" if /Nodes? ($node_pattern)/")
      if [[ -n "$node" ]]; then
        node_detected=true

        local new_block="$marker_start
Host ${cluster}-gpu
  HostName $node
  User your_username
  ProxyJump $cluster
  StrictHostKeyChecking no
$marker_end"

        cp "$config" "$config.bak"

        if grep -q "$marker_start" "$config"; then
          local tmp=$(mktemp)
          perl -0pe \
            "s|\Q$marker_start\E.*?\Q$marker_end\E|$new_block|s" \
            "$config" > "$tmp"
          mv "$tmp" "$config"
          echo "Updated ${cluster}-gpu → $node"
        else
          echo -e "\n$new_block" >> "$config"
          echo "Added ${cluster}-gpu → $node"
        fi

        echo "Connect: ssh ${cluster}-gpu"
        echo "VSCode:  Remote-SSH → ${cluster}-gpu"
      fi
    fi
  done
}

Then you wrap it in tiny one-liners per cluster:

alpha-gpu()  { _salloc_helper alpha 'gpu-\w+' my-account "$@"; }
beta-gpu()   { _salloc_helper beta  'node\d+' other-account "$@"; }

And the workflow becomes:

$ alpha-gpu --gres=gpu:1
salloc: Granted job allocation 12345
salloc: Nodes gpu-node-042
Updated alpha-gpu → gpu-node-042
Connect: ssh alpha-gpu
VSCode:  Remote-SSH → alpha-gpu

$ ssh alpha-gpu   # you're in

The trick is the while IFS= read -r line loop. It processes salloc output as it streams, line by line. No waiting for the session to end. The moment the scheduler prints the node name, your config is patched. Open VSCode, pick alpha-gpu from Remote-SSH, and you’re coding on a GPU.

When Things Break

They will. Here’s how to think through it.

flowchart TD A["❌ Permission denied publickey"] --> B["Run ssh -vvv your-alias"] B --> C{"Which key is SSH offering?"} C -->|Wrong key| D["Fix IdentityFile in config"] C -->|Right key, still rejected| E["Key not on server yet"] D --> F["✅ Change IdentityFile → retry"] E --> G["✅ ssh-copy-id → retry"]

Quick reference

Error What’s probably happening What to do
Permission denied (publickey) Wrong key, or key not on server ssh -vvv, check IdentityFile, ssh-copy-id
Second hop fails No agent forwarding, or agent empty ForwardAgent yes, ssh-add
Connection refused Wrong port, or server is down Double-check port and hostname
Host key verification failed Server was reinstalled Delete old entry in ~/.ssh/known_hosts

Commands worth memorizing

ssh-keygen -t rsa -b 4096    # generate keys
ssh-copy-id user@host         # put your key on a server
ssh-add -l                    # what keys does the agent have?
ssh-add ~/.ssh/id_rsa         # load a key
ssh -vvv host                 # debug everything
ssh -J jumphost target        # one-off ProxyJump

That’s It

Public keys go on servers. Private keys stay home. The config file saves you from typing. IdentityFile can trick you if you’re not careful. Agent forwarding makes multi-hop work. And the salloc helper means you never manually edit a hostname again.

None of this is hard. But nobody tells you about the gotchas until you’ve already lost an afternoon to them. Hopefully this saves you that afternoon.

Happy SSH-ing.