Picture this. You’ve just been granted access to a shiny HPC cluster. Four GPUs waiting for you. You open your terminal, type ssh cluster:
Permission denied (publickey).
I spent an embarrassing amount of time staring at that line. My GPU hours were ticking. I had a deadline. And SSH felt like some arcane ritual that everyone else just… understood.
Turns out, it’s not that complicated. But nobody ever explains the gotchas. So here’s everything I learned the hard way.
The Padlock and the Key
SSH keys work like a padlock system. You have two files:
id_rsa.pubis the padlock. You hand out copies to every server you want access to. It’s public. Share it freely.id_rsais the key. This one stays on your machine. Never, ever share it.
When you connect to a server, there’s a four-step handshake. The server picks a random challenge, encrypts it with your public key, and sends it over. Only your private key can decrypt it. You send the proof back. If it checks out, you’re in.
The beautiful thing? Your private key never crosses the network. Not once.
Generating Keys
One command:
ssh-keygen -t rsa -b 4096 -C "you@example.com"
There’s also Ed25519, which is newer and faster. But some older clusters don’t accept it, so RSA is the safe bet. Hit enter for the default path, optionally set a passphrase, and you’re done.
Your ~/.ssh/ folder now has:
~/.ssh/
├── id_rsa # Private key (NEVER share)
├── id_rsa.pub # Public key (copy to servers)
├── config # Your best friend (more on this)
├── known_hosts # Fingerprints of servers you've visited
└── authorized_keys # On the server side: who's allowed in
Getting your public key onto the server is one more command:
ssh-copy-id -i ~/.ssh/id_rsa.pub you@cluster.example.com
The Config File
Nobody wants to type ssh -p 2222 alice@login.hpc.uni.edu every time. Life is short. So SSH has a config file.
Host mylab
HostName login.hpc.uni.edu
User alice
Port 2222
IdentityFile ~/.ssh/id_rsa
Now ssh mylab does the same thing. Six words saved, sanity preserved.
You can stack as many of these as you want. One per cluster, one per lab server, one for your Raspberry Pi at home. Whatever.
The IdentityFile Trap
This one cost me two hours.
I had two keys on my laptop: id_rsa and id_ed25519. My config pointed at the Ed25519 key. Seemed fine. I ran ssh mylab. Rejected.
But here’s the weird part: the full command worked perfectly.
ssh -p 2222 alice@login.hpc.uni.edu # works fine
ssh mylab # Permission denied
Same server. Same laptop. Same keys sitting in ~/.ssh/. What gives?
The answer is subtle. When you set IdentityFile in your config, SSH only offers that one key. It doesn’t try the others. So if the server has your RSA key in its authorized_keys but your config says “use Ed25519”: you’re locked out. SSH won’t even try the key that would work.
Without the config? SSH tries all your keys, one by one. RSA gets accepted. You never notice the problem.
Finding the culprit
The way I figured this out was ssh -vvv. Triple-v verbose mode. It dumps everything SSH is doing: which keys it’s trying, what the server says back. Pipe it through grep and the answer jumps out:
ssh -vvv mylab 2>&1 | grep -E "Offering|Authentication"
debug1: Offering public key: /Users/you/.ssh/id_ed25519 ED25519
debug1: Authentications that can continue: publickey # REJECTED
There it is. SSH offered id_ed25519. The server said no. Compare that to the raw command:
ssh -vvv -p 2222 alice@login.hpc.uni.edu 2>&1 | grep -E "Offering|Authentication"
debug1: Offering public key: /Users/you/.ssh/id_rsa RSA
debug1: Authentication succeeded (publickey). # ACCEPTED
Mystery solved. Change IdentityFile in the config, done.
Hopping Through Login Nodes
Most HPC clusters won’t let you SSH directly into a compute node. Those machines sit on an internal network. You have to go through a login node first: a two-hop connection.
SSH agent + keys"] -->|Hop 1| B["🖧 Login Node
Public-facing"] B -->|Hop 2| C["⚡ Compute Node
GPUs here"] C -.->|Agent forwarding| A
SSH has ProxyJump for exactly this:
Host alpha
HostName cluster.example.com
User alice
IdentityFile ~/.ssh/id_rsa
ForwardAgent yes
Host alpha-gpu
HostName gpu-node-042
User alice
ProxyJump alpha
ssh alpha-gpu and SSH handles both hops automatically. But there’s a catch. The second hop needs your private key to authenticate: but your private key is on your laptop, not on the login node.
That’s what ForwardAgent yes does. It lets the login node ask your laptop’s SSH agent to sign the authentication challenge. The key itself never touches the login node. Elegant.
Just make sure your agent actually has keys loaded:
ssh-add -l # list loaded keys
ssh-add ~/.ssh/id_rsa # add one if it's empty
A Multi-Cluster Config
Here’s roughly what a config looks like when you’re juggling multiple clusters:
Host alpha
HostName alpha.example.com
User alice
IdentityFile ~/.ssh/id_rsa
ForwardAgent yes
# This block gets updated automatically (see below)
# BEGIN alpha-gpu
Host alpha-gpu
HostName gpu-node-042
User alice
ProxyJump alpha
# END alpha-gpu
Host beta
HostName beta.example.com
User alice
IdentityFile ~/.ssh/id_rsa
ForwardAgent yes
Host lab
HostName login.lab.uni.edu
User alice
Port 2222
IdentityFile ~/.ssh/id_rsa
ForwardAgent yes
Automating salloc with Dynamic Config Updates
Here’s the part that made my life dramatically better.
Every time you run salloc, the scheduler assigns you a different compute node. gpu-node-042 today, gpu-node-117 tomorrow. You’d have to open ~/.ssh/config, find the right block, change the hostname, save. Every. Single. Time.
I got tired of this after day two.
The fix is a shell function that runs salloc, watches its output for the node name, and patches your SSH config in real-time: while the allocation is still running. By the time you see the prompt, your config is already updated.
_salloc_helper() {
local cluster="$1" node_pattern="$2" account="$3"
shift 3
local defaults="--nodes=1 --ntasks=1 --cpus-per-task=4 \
--mem=16G --time=3:00:00 --account=$account"
# Let user args override defaults
local arg key
for arg in "$@"; do
key="${arg%%=*}"
defaults=$(echo "$defaults" | sed "s|${key}=[^ ]*||g")
done
local config=~/.ssh/config
local marker_start="# BEGIN ${cluster}-gpu"
local marker_end="# END ${cluster}-gpu"
local node_detected=false
ssh "$cluster" "salloc $defaults $*" 2>&1 \
| while IFS= read -r line; do
echo "$line"
if ! $node_detected; then
local node
node=$(echo "$line" \
| perl -ne "print \"\$1\" if /Nodes? ($node_pattern)/")
if [[ -n "$node" ]]; then
node_detected=true
local new_block="$marker_start
Host ${cluster}-gpu
HostName $node
User your_username
ProxyJump $cluster
StrictHostKeyChecking no
$marker_end"
cp "$config" "$config.bak"
if grep -q "$marker_start" "$config"; then
local tmp=$(mktemp)
perl -0pe \
"s|\Q$marker_start\E.*?\Q$marker_end\E|$new_block|s" \
"$config" > "$tmp"
mv "$tmp" "$config"
echo "Updated ${cluster}-gpu → $node"
else
echo -e "\n$new_block" >> "$config"
echo "Added ${cluster}-gpu → $node"
fi
echo "Connect: ssh ${cluster}-gpu"
echo "VSCode: Remote-SSH → ${cluster}-gpu"
fi
fi
done
}
Then you wrap it in tiny one-liners per cluster:
alpha-gpu() { _salloc_helper alpha 'gpu-\w+' my-account "$@"; }
beta-gpu() { _salloc_helper beta 'node\d+' other-account "$@"; }
And the workflow becomes:
$ alpha-gpu --gres=gpu:1
salloc: Granted job allocation 12345
salloc: Nodes gpu-node-042
Updated alpha-gpu → gpu-node-042
Connect: ssh alpha-gpu
VSCode: Remote-SSH → alpha-gpu
$ ssh alpha-gpu # you're in
The trick is the while IFS= read -r line loop. It processes salloc output as it streams, line by line. No waiting for the session to end. The moment the scheduler prints the node name, your config is patched. Open VSCode, pick alpha-gpu from Remote-SSH, and you’re coding on a GPU.
When Things Break
They will. Here’s how to think through it.
Quick reference
| Error | What’s probably happening | What to do |
|---|---|---|
Permission denied (publickey) |
Wrong key, or key not on server | ssh -vvv, check IdentityFile, ssh-copy-id |
| Second hop fails | No agent forwarding, or agent empty | ForwardAgent yes, ssh-add |
Connection refused |
Wrong port, or server is down | Double-check port and hostname |
Host key verification failed |
Server was reinstalled | Delete old entry in ~/.ssh/known_hosts |
Commands worth memorizing
ssh-keygen -t rsa -b 4096 # generate keys
ssh-copy-id user@host # put your key on a server
ssh-add -l # what keys does the agent have?
ssh-add ~/.ssh/id_rsa # load a key
ssh -vvv host # debug everything
ssh -J jumphost target # one-off ProxyJump
That’s It
Public keys go on servers. Private keys stay home. The config file saves you from typing. IdentityFile can trick you if you’re not careful. Agent forwarding makes multi-hop work. And the salloc helper means you never manually edit a hostname again.
None of this is hard. But nobody tells you about the gotchas until you’ve already lost an afternoon to them. Hopefully this saves you that afternoon.
Happy SSH-ing.