 Command

Sam Foreman's personal site. Vim-style keybinds for navigation; theme + font pickers below.

Theme
 Font Body Code
Reader
Keybinds
Navigation
j / ↓ Next item k / ↑ Previous item g First item in region G Last item in region zz Center focused item h / l Sidebar / main content ] / [ Next/previous heading } / { Next/previous block d / u Half-page down/up
Layout
<zh> Toggle sidebar <zr> Toggle reader view <zj> / <zk> Focus main / actions ⇧C / ⇧E  ·  <zM> / <zR> Collapse / expand all sections
Dialogs
⌃P / : Command palette ⌃X Theme picker / Search ? Show keybinds ⌃N / ⌃P Next/prev search result Esc Close dialog / exit reader
History
n Next document b Previous document ⌃O History back ⌃I History forward
Sections
a about p posts t talks m more s style
 Search
about: Sam Foreman about/more: 🪪 More ideas: 💡 Ideas more: ➕ More now: Now posts: 📬 Posts posts/2023/12/05: 🔳 l2hmc-qcd Example: 4D SU(3) posts/2025: 📆 2025 posts/2025/04/28: 🔥 Building PyTorch 2.6 from Source on Aurora posts/2025/05/03: 🚧 Frameworks Issue with numpy \› 2 posts/2025/06: 06 posts/2025/06/01: 📰 Nice Headings posts/2025/06/02: 🧜‍♀️ Mermaid posts/2025/06/14: 🏗️ Building PyTorch 2.8 from Source on Aurora posts/2025/09/12: 🍹 BlendCorpus + TorchTitan @ ALCF posts/2025/09/17: 📊 pbs-tui: TUI for PBS Job Scheduler Monitoring posts/2025/10/06: 🎨 Mixing Between Distributions While Training posts/2025/11/12: 🧊 Cooling Down Checkpoints: Best Practices for Model Evaluation posts/2026/01/07: 🎉 Happy New Year! posts/2026/01/10: 🍋 ezpz: distributed PyTorch across any hardware posts/2026/02/28: ⏱️ Comparing Launchers on Aurora posts/2026/02/28: ## torchrun posts/2026/02/28: ## ezpz posts/2026/04/27: Pre-Training AuroraGPT with TorchTitan posts/2026/04/27: ## Two-Week Summary (Apr 12–27, 2026) posts/2026/04/27: ## Detailed Breakdown posts/2026/04/27: ### Week 1: Apr 12–18 — Benchmarking, LR Finder, XPU Fixes posts/2026/04/27: #### Benchmarking (Apr 12–15) posts/2026/04/27: #### LR Finder (Apr 12–14) posts/2026/04/27: #### Scaling Study (Apr 12) posts/2026/04/27: #### Upstream Syncs (Apr 12–18, syncs 6–14) posts/2026/04/27: #### XPU Bug Fixes (Apr 18) posts/2026/04/27: #### RL Experiment (Apr 18) posts/2026/04/27: ### Week 1.5: Apr 18–25 — Production Readiness posts/2026/04/27: #### Torch 2.12 Benchmarks (Apr 18) posts/2026/04/27: #### LR Finder Extensions (Apr 20–21) posts/2026/04/27: #### XPU Fixes (Apr 23) posts/2026/04/27: #### Torch 2.13 Environment (Apr 25) posts/2026/04/27: #### 2B Scaling Study on Torch 2.13 (Apr 25) posts/2026/04/27: #### Production Training (Apr 25) posts/2026/04/27: ### Week 2: Apr 26–27 — Optimizer Competition posts/2026/04/27: #### RL Multi-Task Refactor (Apr 26) posts/2026/04/27: #### Docs Reorganization (Apr 26) posts/2026/04/27: #### Generic HF Dataset Streaming (Apr 26) posts/2026/04/27: #### New Optimizers (Apr 26) posts/2026/04/27: #### Architecture Tweaks (Apr 26–27) posts/2026/04/27: ## Competition Results posts/2026/04/27: ### Round 1–3: Speedrun — 2N, GBS=48, 1000 steps posts/2026/04/27: ### 10B Full Training — 8N, GBS=384, ~3,178 steps posts/2026/04/27: ### Round 4: Reproducible Speedrun — 2N, GAS=8, GBS=384, 1000 steps posts/2026/04/27: ## Key Discoveries posts/2026/04/27: ## Infrastructure Built posts/2026/04/27: ## High-Level posts/2026/04/27: ## Detailed Breakdown posts/2026/04/27: ### Week 1: Apr 12–18 — Benchmarking, LR Finder, XPU Fixes posts/2026/04/27: #### Benchmarking (Apr 12–15) posts/2026/04/27: #### LR Finder (Apr 12–14) posts/2026/04/27: #### Scaling Study (Apr 12) posts/2026/04/27: #### Upstream Syncs (Apr 12–18, syncs 6–14) posts/2026/04/27: #### XPU Bug Fixes (Apr 18) posts/2026/04/27: #### RL Experiment (Apr 18) posts/2026/04/27: ### Week 1.5: Apr 18–25 — Production Readiness posts/2026/04/27: #### Torch 2.12 Benchmarks (Apr 18) posts/2026/04/27: #### LR Finder Extensions (Apr 20–21) posts/2026/04/27: #### XPU Fixes (Apr 23) posts/2026/04/27: #### Torch 2.13 Environment (Apr 25) posts/2026/04/27: #### 2B Scaling Study on Torch 2.13 (Apr 25) posts/2026/04/27: #### Production Training (Apr 25) posts/2026/04/27: ### Week 2: Apr 26–27 — Optimizer Competition posts/2026/04/27: #### RL Multi-Task Refactor (Apr 26) posts/2026/04/27: #### Docs Reorganization (Apr 26) posts/2026/04/27: #### Generic HF Dataset Streaming (Apr 26) posts/2026/04/27: #### New Optimizers (Apr 26) posts/2026/04/27: #### Architecture Tweaks (Apr 26–27) posts/2026/04/27: ## Competition Results posts/2026/04/27: ### Round 1–3: 1000-step speedruns, 2 nodes, GBS=48 (17 configs) posts/2026/04/27: ### Round 4 (10B full training, 8 nodes, GBS=384, 5 configs) posts/2026/04/27: ### Round 5 (2 nodes, GAS=8, GBS=384, local dataset, 8 configs — in progress) posts/2026/04/27: ## Key Discoveries posts/2026/04/27: ## Infrastructure Built posts/2026/05/01: Running 50k Python Processes on Aurora with ezpz yeet posts/2026/06/27: Local AI Apps on ALCF: Argo, Inference Endpoints, and One Gateway posts/2026/06/28: Migrating from Quarto to Astro: samforeman.me → samf.sh posts/2026/08/08: Pre-Training LLMs on a Supercomputer posts/2026/09/22: Working From Anywhere: Persistent Access to Compute and Context posts/2026/09/25: A Small Service Mesh for My Macs and Supercomputers posts/ai-for-physics: ⚛️ AI for Physics posts/ai-for-physics/diffusion: 🎲 MCMC + Diffusion Sampling posts/ai-for-physics/l2hmc-qcd: 🎢 L2HMC for LQCD posts/ai-for-physics/l2hmc-qcd/2du1: 🎢 l2hmc-qcd Example: 2D U(1) posts/auroragpt: 🤖 AuroraGPT posts/auroragpt/aurora-gpt: 🏎️ Megatron-DeepSpeed on Intel XPU posts/auroragpt/checkpoints: 💾 Converting Checkpoints posts/auroragpt/determinstic-flash-attn/deterministic-flash-attn: 🎰 Deterministic flash-attn posts/auroragpt/flash-attn-sunspot: 📸 flash-attn on Sunspot posts/auroragpt/long-sequences: 🚂 Loooooooong Sequence Lengths posts/auroragpt/mpi4py-reproducer: 🐛 mpi4py bug on Sunspot posts/auroragpt/spike-skipper: 🏔️ Spike Skipper posts/auroragpt/startup-times: 🐢 Starting Up Distributed Training on Aurora posts/auroragpt/startup-times: ## Response posts/auroragpt/startup-times: ### Measuring / Calculating Startup Time posts/auroragpt/startup-times: ## Minimal Working Example posts/croon: croon: Synced Lyrics in the Terminal, for Whatever Is Playing posts/dope-slides: 💅 How to Make Dope Slides posts/drafts/2025/09/22: 📝 2025 Annual Report posts/ezpz-at-alcf: 🍋 ezpz @ ALCF posts/ezpz-v1: 📝 ezpz-v1 posts/globusfs: globusfs: an fsspec Filesystem for Globus Collections posts/jupyter: 📗 Jupyter posts/jupyter/test: 🏁 l2hmc Example: 2D $U(1)$ posts/resume: 🧑🏻‍💻 Sam Foreman’s Résumé posts/svgbob: 🫥 svgbob posts/torchtune-aurora: 🪛 Torchtune on Aurora posts/torchtune-patch-aurora: 🚑 Torchtune Patch on Aurora posts/wandb-tui: wandb-tui: Comparing W&B Runs Without Leaving the Terminal projects: 📚 Projects talks: 🎙️ Talks talks/2025/09/24: Training Foundation Models on Supercomputers talks/2025/10/08: AERIS: Argonne's Earth Systems Model talks/2025/10/15: Training Foundation Models on Supercomputers talks/2025/10/24: Training Foundation Models on Supercomputers talks/2025/12/16: AuroraGPT: Training Foundation Models on Supercomputers talks/2026/06/03: Production Pre-Training at Scale: The Good, the Bad, and the Restarts talks/2026/07/14: Pre-Training AuroraGPT at Scale on Aurora talks/2026/08/03: Pre-Training LLMs on a Supercomputer talks/ai-for-science-2024: Parallel Training Methods talks/alcf-hpc-workshop-2024/alcf-hpc-workshop-2024: Deep Learning and Foundation Models at Scale talks/aurora-gpt-fm-for-electric-grid/auroragpt-fm-for-electric-grid: AuroraGPT: Foundation Models for Science talks/auroragpt-siam25: AuroraGPT talks/auroragpt/alcf-hpc-workshop-2024/auroragpt-alcf-hands-on-hpc-workshop-2024: AuroraGPT: ANL's General Purpose Scientific LLM talks/demo-slides: AuroraGPT: Training Foundation Models on Supercomputers talks/hpc-user-forum/auroragpt: AuroraGPT talks/incite-hackathon-2025: ALCF Incite Hackathon 2025 talks/incite-hackathon-2025/auroragpt: LLMs on Aurora: Overview talks/incite-hackathon-2025/ezpz: LLMs on Aurora: Hands-On talks/llms-at-scale: Training LLMs at Scale talks/llms-on-polaris: Training LLMs on Polaris talks/openskai25: Open SkAI2025 talks/openskai25/ai4science: Scientific AI at Scale: AuroraGPT talks/openskai25/training: Scientific AI at Scale: Distributed Training webtui: Style webtui/components/accordion: Accordion webtui/components/badge: Badge webtui/components/button: Button webtui/components/checkbox: Checkbox webtui/components/dialog: Dialog webtui/components/input: Input webtui/components/popover: Popover webtui/components/pre: Pre webtui/components/progress: Progress webtui/components/radio: Radio webtui/components/range: Range webtui/components/separator: Separator webtui/components/spinner: Spinner webtui/components/switch: Switch webtui/components/table: Table webtui/components/textarea: Textarea webtui/components/tooltip: Popover webtui/components/typography: Typography webtui/components/view: View webtui/contributing/contributing: Contributing webtui/contributing/contributing: ## Local Development webtui/contributing/contributing: ## Issues webtui/contributing/contributing: ## Pull Requests webtui/contributing/style-guide: Style Guide webtui/contributing/style-guide: ## CSS Units webtui/contributing/style-guide: ## Selectors webtui/contributing/style-guide: ## Documentation webtui/installation/astro: Astro webtui/installation/astro: ## Scoping webtui/installation/astro: ### Frontmatter Imports webtui/installation/astro: ### ‹style› tag webtui/installation/astro: ### Full Library Import webtui/installation/nextjs: Next.js webtui/installation/vite: Vite webtui/plugins/plugin-dev: Developing Plugins webtui/plugins/plugin-dev: ### Style Layers webtui/plugins/plugin-nf: Nerd Font Plugin webtui/plugins/theme-catppuccin: Catppuccin Theme webtui/plugins/theme-custom: Custom Theme webtui/plugins/theme-everforest: Everforest Theme webtui/plugins/theme-gruvbox: Gruvbox Theme webtui/plugins/theme-nord: Nord Theme webtui/plugins/theme-vitesse: Vitesse Theme webtui/start/ascii-boxes: ASCII Boxes webtui/start/changelog: Changelog webtui/start/installation: Installation webtui/start/installation: ## Installation webtui/start/installation: ## Using CSS webtui/start/installation: ## Using ESM webtui/start/installation: ## Using a CDN webtui/start/installation: ## Full Library Import webtui/start/installation: ### CSS webtui/start/installation: ### ESM webtui/start/installation: ### CDN webtui/start/intro: Introduction webtui/start/intro: ## Features webtui/start/plugins: Plugins webtui/start/plugins: ## Official Plugins webtui/start/plugins: ### Themes webtui/start/plugins: ## Community Plugins webtui/start/theming: Theming webtui/start/theming: ## CSS Variables webtui/start/theming: ### Font Styles webtui/start/theming: ### Colors webtui/start/theming: ### Light & Dark webtui/start/theming: ## Theme Plugins webtui/start/theming: ### Using Multiple Theme Accents webtui/start/tuis-vs-guis: TUIs vs GUIs webtui/start/tuis-vs-guis: ## Monospace Fonts webtui/start/tuis-vs-guis: ## Character Cells
 Theme Current: Light j/k or ↑/↓ + Enter

Working From Anywhere: Persistent Access to Compute and Context

A portable working setup that follows me across networks: attaching to always-on agent sessions on my home Mac and on ALCF/CELS machines with herdr, keeping a single cross-agent memory store reachable from every client, staying model-flexible across Claude, GPT, and ALCF-hosted open models, and what happens when a restrictive network breaks the transport underneath all of it.

Most of my day is spent talking to long-running agents — on my home machine, on ALCF login and compute nodes, and on CELS hosts — from whatever laptop I happen to be sitting at. Over time that has turned into a small system with a single design goal: the work, the machines, and the accumulated context should follow me, not the other way around.

This post walks through four pieces of that system and how they fit together:

  1. Attaching to remote sessions with herdr --remote, so an agent session on a distant machine feels local and stays alive when I disconnect.
  2. One transport underneath everything — and what happened when a restrictive campus network broke it.
  3. A single, persistent cross-agent memory store, reachable from every client so context is not siloed per-machine or per-tool.
  4. Model flexibility — moving fluidly between Claude, GPT, and ALCF-hosted open models, and between clients (Claude Code, Hermes, OpenCode, …), without re-plumbing anything.

The model-routing internals have their own post — Local AI Apps on ALCF — so here I’ll link there rather than repeat it, and focus on how these layers compose.

TL;DR — the shape of the system
  • herdr --remote <host> attaches to a persistent agent session on a remote machine over SSH; the session survives disconnects, so long jobs and long conversations both outlive my laptop’s network.
  • Everything rides one point-to-point transport between my machines. When a campus network blocked the usual mesh, I swapped the transport layer without changing anything above it.
  • A single memory service holds cross-agent session history and a wiki-style knowledge store; every client points at the same endpoint, so context is shared rather than fragmented.
  • Model choice is a routing decision, not a client decision: one local gateway fans out to Claude (via Argo), GPT, and ALCF-hosted open models, and any OpenAI/Anthropic-compatible client can use any of them.

1. Attaching to remote sessions with herdr --remote

The foundation is being able to attach to a session running somewhere else, rather than starting a fresh one each time I connect. I use herdr for this:

herdr --remote mbph          # attach to a session on my home Mac
herdr --remote polaris       # ... or an ALCF login node
herdr --remote logins.cels   # ... or a CELS host

Under the hood this is “just” SSH plus a terminal session on the far end — but the ergonomics matter:

  • Persistence. The session lives on the remote host, not in my terminal. If my laptop sleeps, my wifi drops, or I close the lid and walk to a meeting, the agent on the other end keeps going. I re-attach and the conversation (and any running job) is exactly where I left it.
  • Locality of compute. When I’m working on Polaris or Aurora, I want the agent on the machine with the data, the module environment, and the scheduler — not shuttling files back to my laptop. herdr --remote puts the session where the work is.
  • One muscle-memory for many machines. The same command attaches to home, ALCF, and CELS. I don’t context-switch between different tools per host.
  • It’s deliberately minimal. Because it’s SSH + a terminal, it needs none of the heavier “let an agent drive a whole spare Mac” machinery (Screen Recording, Accessibility, synthetic input). That keeps it robust and portable across very different hosts.

The important consequence: a remote session is only as reachable as the network path to it. Which is exactly where the next piece comes in.


2. One transport underneath everything

All of this — remote attach, the memory store below, file sync — depends on my machines being able to find and reach each other regardless of which network I’m on. For a long time that was a mesh VPN: every machine gets a stable address, and point-to-point encrypted links form automatically.

That works beautifully until you’re on a network that doesn’t want it to.

When a restrictive network breaks the transport

On one campus network (Argonne-Auth), my mesh VPN simply would not connect. The failure was specific and worth understanding, because the diagnosis is reusable even if your particulars differ.

The symptom looked like a login failure, but the real story was at the TLS layer. I could resolve the control server’s DNS fine, and I could open a raw TCP connection to it on 443 — but the moment the TLS handshake began, the connection was reset by peer. Not a timeout, not a certificate error: an immediate reset.

The decisive test was to vary only the hostname advertised in the TLS ClientHello (the SNI field) while holding the destination IP constant:

TestResult
DNS resolution of the control host✅ correct answers
Raw TCP connect to :443✅ succeeds
TLS handshake with the VPN’s SNI❌ reset every time
TLS handshake with an unrelated SNI, same IP✅ reaches the server
TLS handshake with the VPN’s SNI, pointed at an unrelated IP❌ reset

That last row is the giveaway: send the VPN’s hostname in the handshake and the connection dies even when the IP belongs to something else entirely. So the filter isn’t blocking an IP range or doing TLS interception (no substitute certificate is ever presented — it’s a clean reset). It’s reading the plaintext hostname out of the TLS ClientHello and resetting connections whose SNI matches a category — here, VPN/anonymizer services. A well-known info site in the same product family loaded fine; a commercial VPN’s endpoint got the same reset my mesh did. Classic SNI-based category filtering.

Note

This is a legitimate network security control on infrastructure I don’t administer. The right long-term fix is to request an allowlist exception from IT for the endpoints I need — not to treat the filter as an adversary. I raised it through that channel.

Swapping the transport, keeping everything above it

The useful architectural point is this: because every layer above sat on a generic point-to-point tunnel, I only had to replace the transport — not the remote-attach workflow, not the memory store, not the model gateway.

The replacement I settled on is a peer-to-peer data-plane tool that establishes an encrypted tunnel between two machines I control without routing through the filtered control endpoint — the two ends exchange what they need out-of-band and then connect directly (falling back to a relay that isn’t on the filter’s list). I’m intentionally not publishing a step-by-step recipe for getting around a specific institution’s network policy here; the reusable lesson is the one worth taking away:

  • Diagnose at the right layer. “VPN won’t connect” was really “SNI category filter resets the control channel.” You cannot fix what you’ve mis-located.
  • Keep a clean seam between transport and everything else. Remote attach, memory, and model routing all spoke “connect to this host:port.” Swapping the thing that provides that host:port was a localized change, not a rebuild.
  • Distrust stale status. At one point the VPN UI still displayed “connected” with leftover byte counters from an earlier network — but a live reachability probe timed out. On a network like this, verify the data path, don’t trust the indicator light.

Once the tunnel was back — by a different mechanism — herdr --remote, the memory store, and everything else came straight back to life, unchanged.


3. A single, persistent cross-agent memory store

The second thing that has to follow me is context. I don’t want each agent, on each machine, starting amnesiac — and I especially don’t want my home Mac’s agent and an ALCF agent to have separate, divergent memories of the same project.

So the memory store is centralized and persistent, not per-client:

  • One service holds session history (what each agent did, across every project and harness) and a wiki-style knowledge store that agents read from and write to.
  • It lives on an always-on host, and every client — on every machine — points at the same endpoint. Capture hooks stream session events to it; retrieval reads from it. The store is the source of truth; its index is rebuildable from flat markdown, so “the files are the database.”
  • Because it’s centralized, cross-agent handoffs and cross-project messages work: one agent can leave a note or an open handoff that another agent — in a different project or on a different machine — picks up.

The payoff is that “what was I doing on this?” has a single answer regardless of which client or machine asks. But centralization has a sharp edge that ties directly back to §2.

The dependency you inherit

A centralized store reachable “from everywhere” is only reachable if the transport to it is up. When the campus network broke my mesh VPN (§2), it didn’t just break remote attach — it broke the memory store too, for the same SNI reason, because the client reached it over the same kind of tunnel.

Two lessons crystallized here:

  • Point every surface at one indirection, then move the indirection. Once the transport was restored, I pointed the memory clients at a single local address that the tunnel forwards to the always-on host. Now the store is reachable the same way on every network, with one code path instead of per-network special cases.
  • Beware config that ignores your override. Some capture hooks hardcode the server URL on their command line, where it beats any environment variable. If you “fix” things by exporting an env var, those hooks keep dialing the dead address and your capture silently breaks while everything else looks fine. The fix has to land where each surface actually reads its configuration — which means finding all of them.

That second point is a recurring theme in this whole setup: the failure modes are rarely “it’s totally down.” They’re “it’s down here, silently, while the dashboard says green.”


4. Model flexibility, and moving between clients

The last piece is not being locked to one model or one client. On any given day I might want Claude for one task, GPT for another, and an ALCF-hosted open model for a third — and I might want to reach them from Claude Code, from Hermes (CLI or desktop), or from OpenCode.

I’ve written the full stack up separately, so here’s just the shape and why it matters for a portable setup:

  • One local gateway (llm-rosetta) presents a single localhost API and routes by model name to:
    • Claude via Argo (ALCF’s internal frontier-model gateway), reached through a small SSH-tunneling shim (argo-shim);
    • GPT-family models, likewise;
    • ALCF Inference Endpoints — open models on Sophia/Metis/Minerva behind ALCF’s OpenAI-compatible API.
  • Every client points at that one port. Switching from Claude Code to Hermes to OpenCode is a client preference, not a re-plumbing job — they all speak to the same gateway, so they all get the same model menu.
  • Routing, fallback, and billing are centralized. Model choice is a routing decision; a request can prefer Argo Claude and fall back through GPT to an open ALCF model automatically. Everything bills through ALCF and stays on 127.0.0.1.

Why this belongs in a post about portability: because model access is behind the same “one local endpoint” indirection as everything else, it moves with me for the same reason the memory store does. When the transport changed in §2, the model gateway didn’t care — it was already talking to a local port.


How the four layers compose

Read top to bottom, the stack is:

  1. Transport — a point-to-point encrypted tunnel between my machines, chosen so it survives hostile networks. Everything else assumes only “I can reach host:port.”
  2. Remote attach — herdr --remote puts a persistent agent session on the right machine (home, ALCF, CELS) and lets me come and go.
  3. Memory — one centralized, persistent store gives every agent on every machine a shared, durable context and a channel to hand work off.
  4. Models — one local gateway makes Claude, GPT, and ALCF open models interchangeable across every client.

The recurring design principle across all four: put a clean indirection between “what I use” and “where it physically lives,” so that when the network or the host changes, the change is localized. The campus-network incident in §2 was the stress test — and the reason it was a swap rather than a rebuild is that the seams were already there.

The uncomfortable corollary is that centralization concentrates failure: one broken transport took out remote attach and memory at once. Worth it, in my experience — but only because each layer fails loudly enough to diagnose, and because the seams make each layer independently replaceable.

 samf.sh / posts / 2026 / 09 / 22 · Top 1:1