 Command

Sam Foreman's personal site. Vim-style keybinds for navigation; theme + font pickers below.

Theme
 Font Body Code
Reader
Keybinds
Navigation
j / ↓ Next item k / ↑ Previous item g First item in region G Last item in region zz Center focused item h / l Sidebar / main content ] / [ Next/previous heading } / { Next/previous block d / u Half-page down/up
Layout
<zh> Toggle sidebar <zr> Toggle reader view <zj> / <zk> Focus main / actions ⇧C / ⇧E  ·  <zM> / <zR> Collapse / expand all sections
Dialogs
⌃P / : Command palette ⌃X Theme picker / Search ? Show keybinds ⌃N / ⌃P Next/prev search result Esc Close dialog / exit reader
History
n Next document b Previous document ⌃O History back ⌃I History forward
Sections
a about p posts t talks m more s style
 Search
about: Sam Foreman about/more: 🪪 More ideas: 💡 Ideas more: ➕ More now: Now posts: 📬 Posts posts/2023/12/05: 🔳 l2hmc-qcd Example: 4D SU(3) posts/2025: 📆 2025 posts/2025/04/28: 🔥 Building PyTorch 2.6 from Source on Aurora posts/2025/05/03: 🚧 Frameworks Issue with numpy \› 2 posts/2025/06: 06 posts/2025/06/01: 📰 Nice Headings posts/2025/06/02: 🧜‍♀️ Mermaid posts/2025/06/14: 🏗️ Building PyTorch 2.8 from Source on Aurora posts/2025/09/12: 🍹 BlendCorpus + TorchTitan @ ALCF posts/2025/09/17: 📊 pbs-tui: TUI for PBS Job Scheduler Monitoring posts/2025/10/06: 🎨 Mixing Between Distributions While Training posts/2025/11/12: 🧊 Cooling Down Checkpoints: Best Practices for Model Evaluation posts/2026/01/07: 🎉 Happy New Year! posts/2026/01/10: 🍋 ezpz: distributed PyTorch across any hardware posts/2026/02/28: ⏱️ Comparing Launchers on Aurora posts/2026/02/28: ## torchrun posts/2026/02/28: ## ezpz posts/2026/04/27: Pre-Training AuroraGPT with TorchTitan posts/2026/04/27: ## Two-Week Summary (Apr 12–27, 2026) posts/2026/04/27: ## Detailed Breakdown posts/2026/04/27: ### Week 1: Apr 12–18 — Benchmarking, LR Finder, XPU Fixes posts/2026/04/27: #### Benchmarking (Apr 12–15) posts/2026/04/27: #### LR Finder (Apr 12–14) posts/2026/04/27: #### Scaling Study (Apr 12) posts/2026/04/27: #### Upstream Syncs (Apr 12–18, syncs 6–14) posts/2026/04/27: #### XPU Bug Fixes (Apr 18) posts/2026/04/27: #### RL Experiment (Apr 18) posts/2026/04/27: ### Week 1.5: Apr 18–25 — Production Readiness posts/2026/04/27: #### Torch 2.12 Benchmarks (Apr 18) posts/2026/04/27: #### LR Finder Extensions (Apr 20–21) posts/2026/04/27: #### XPU Fixes (Apr 23) posts/2026/04/27: #### Torch 2.13 Environment (Apr 25) posts/2026/04/27: #### 2B Scaling Study on Torch 2.13 (Apr 25) posts/2026/04/27: #### Production Training (Apr 25) posts/2026/04/27: ### Week 2: Apr 26–27 — Optimizer Competition posts/2026/04/27: #### RL Multi-Task Refactor (Apr 26) posts/2026/04/27: #### Docs Reorganization (Apr 26) posts/2026/04/27: #### Generic HF Dataset Streaming (Apr 26) posts/2026/04/27: #### New Optimizers (Apr 26) posts/2026/04/27: #### Architecture Tweaks (Apr 26–27) posts/2026/04/27: ## Competition Results posts/2026/04/27: ### Round 1–3: Speedrun — 2N, GBS=48, 1000 steps posts/2026/04/27: ### 10B Full Training — 8N, GBS=384, ~3,178 steps posts/2026/04/27: ### Round 4: Reproducible Speedrun — 2N, GAS=8, GBS=384, 1000 steps posts/2026/04/27: ## Key Discoveries posts/2026/04/27: ## Infrastructure Built posts/2026/04/27: ## High-Level posts/2026/04/27: ## Detailed Breakdown posts/2026/04/27: ### Week 1: Apr 12–18 — Benchmarking, LR Finder, XPU Fixes posts/2026/04/27: #### Benchmarking (Apr 12–15) posts/2026/04/27: #### LR Finder (Apr 12–14) posts/2026/04/27: #### Scaling Study (Apr 12) posts/2026/04/27: #### Upstream Syncs (Apr 12–18, syncs 6–14) posts/2026/04/27: #### XPU Bug Fixes (Apr 18) posts/2026/04/27: #### RL Experiment (Apr 18) posts/2026/04/27: ### Week 1.5: Apr 18–25 — Production Readiness posts/2026/04/27: #### Torch 2.12 Benchmarks (Apr 18) posts/2026/04/27: #### LR Finder Extensions (Apr 20–21) posts/2026/04/27: #### XPU Fixes (Apr 23) posts/2026/04/27: #### Torch 2.13 Environment (Apr 25) posts/2026/04/27: #### 2B Scaling Study on Torch 2.13 (Apr 25) posts/2026/04/27: #### Production Training (Apr 25) posts/2026/04/27: ### Week 2: Apr 26–27 — Optimizer Competition posts/2026/04/27: #### RL Multi-Task Refactor (Apr 26) posts/2026/04/27: #### Docs Reorganization (Apr 26) posts/2026/04/27: #### Generic HF Dataset Streaming (Apr 26) posts/2026/04/27: #### New Optimizers (Apr 26) posts/2026/04/27: #### Architecture Tweaks (Apr 26–27) posts/2026/04/27: ## Competition Results posts/2026/04/27: ### Round 1–3: 1000-step speedruns, 2 nodes, GBS=48 (17 configs) posts/2026/04/27: ### Round 4 (10B full training, 8 nodes, GBS=384, 5 configs) posts/2026/04/27: ### Round 5 (2 nodes, GAS=8, GBS=384, local dataset, 8 configs — in progress) posts/2026/04/27: ## Key Discoveries posts/2026/04/27: ## Infrastructure Built posts/2026/05/01: Running 50k Python Processes on Aurora with ezpz yeet posts/2026/06/27: Local AI Apps on ALCF: Argo, Inference Endpoints, and One Gateway posts/2026/06/28: Migrating from Quarto to Astro: samforeman.me → samf.sh posts/2026/08/08: Pre-Training LLMs on a Supercomputer posts/2026/09/22: Working From Anywhere: Persistent Access to Compute and Context posts/2026/09/25: A Small Service Mesh for My Macs and Supercomputers posts/ai-for-physics: ⚛️ AI for Physics posts/ai-for-physics/diffusion: 🎲 MCMC + Diffusion Sampling posts/ai-for-physics/l2hmc-qcd: 🎢 L2HMC for LQCD posts/ai-for-physics/l2hmc-qcd/2du1: 🎢 l2hmc-qcd Example: 2D U(1) posts/auroragpt: 🤖 AuroraGPT posts/auroragpt/aurora-gpt: 🏎️ Megatron-DeepSpeed on Intel XPU posts/auroragpt/checkpoints: 💾 Converting Checkpoints posts/auroragpt/determinstic-flash-attn/deterministic-flash-attn: 🎰 Deterministic flash-attn posts/auroragpt/flash-attn-sunspot: 📸 flash-attn on Sunspot posts/auroragpt/long-sequences: 🚂 Loooooooong Sequence Lengths posts/auroragpt/mpi4py-reproducer: 🐛 mpi4py bug on Sunspot posts/auroragpt/spike-skipper: 🏔️ Spike Skipper posts/auroragpt/startup-times: 🐢 Starting Up Distributed Training on Aurora posts/auroragpt/startup-times: ## Response posts/auroragpt/startup-times: ### Measuring / Calculating Startup Time posts/auroragpt/startup-times: ## Minimal Working Example posts/croon: croon: Synced Lyrics in the Terminal, for Whatever Is Playing posts/dope-slides: 💅 How to Make Dope Slides posts/drafts/2025/09/22: 📝 2025 Annual Report posts/ezpz-at-alcf: 🍋 ezpz @ ALCF posts/ezpz-v1: 📝 ezpz-v1 posts/globusfs: globusfs: an fsspec Filesystem for Globus Collections posts/jupyter: 📗 Jupyter posts/jupyter/test: 🏁 l2hmc Example: 2D $U(1)$ posts/resume: 🧑🏻‍💻 Sam Foreman’s Résumé posts/svgbob: 🫥 svgbob posts/torchtune-aurora: 🪛 Torchtune on Aurora posts/torchtune-patch-aurora: 🚑 Torchtune Patch on Aurora posts/wandb-tui: wandb-tui: Comparing W&B Runs Without Leaving the Terminal projects: 📚 Projects talks: 🎙️ Talks talks/2025/09/24: Training Foundation Models on Supercomputers talks/2025/10/08: AERIS: Argonne's Earth Systems Model talks/2025/10/15: Training Foundation Models on Supercomputers talks/2025/10/24: Training Foundation Models on Supercomputers talks/2025/12/16: AuroraGPT: Training Foundation Models on Supercomputers talks/2026/06/03: Production Pre-Training at Scale: The Good, the Bad, and the Restarts talks/2026/07/14: Pre-Training AuroraGPT at Scale on Aurora talks/2026/08/03: Pre-Training LLMs on a Supercomputer talks/ai-for-science-2024: Parallel Training Methods talks/alcf-hpc-workshop-2024/alcf-hpc-workshop-2024: Deep Learning and Foundation Models at Scale talks/aurora-gpt-fm-for-electric-grid/auroragpt-fm-for-electric-grid: AuroraGPT: Foundation Models for Science talks/auroragpt-siam25: AuroraGPT talks/auroragpt/alcf-hpc-workshop-2024/auroragpt-alcf-hands-on-hpc-workshop-2024: AuroraGPT: ANL's General Purpose Scientific LLM talks/demo-slides: AuroraGPT: Training Foundation Models on Supercomputers talks/hpc-user-forum/auroragpt: AuroraGPT talks/incite-hackathon-2025: ALCF Incite Hackathon 2025 talks/incite-hackathon-2025/auroragpt: LLMs on Aurora: Overview talks/incite-hackathon-2025/ezpz: LLMs on Aurora: Hands-On talks/llms-at-scale: Training LLMs at Scale talks/llms-on-polaris: Training LLMs on Polaris talks/openskai25: Open SkAI2025 talks/openskai25/ai4science: Scientific AI at Scale: AuroraGPT talks/openskai25/training: Scientific AI at Scale: Distributed Training webtui: Style webtui/components/accordion: Accordion webtui/components/badge: Badge webtui/components/button: Button webtui/components/checkbox: Checkbox webtui/components/dialog: Dialog webtui/components/input: Input webtui/components/popover: Popover webtui/components/pre: Pre webtui/components/progress: Progress webtui/components/radio: Radio webtui/components/range: Range webtui/components/separator: Separator webtui/components/spinner: Spinner webtui/components/switch: Switch webtui/components/table: Table webtui/components/textarea: Textarea webtui/components/tooltip: Popover webtui/components/typography: Typography webtui/components/view: View webtui/contributing/contributing: Contributing webtui/contributing/contributing: ## Local Development webtui/contributing/contributing: ## Issues webtui/contributing/contributing: ## Pull Requests webtui/contributing/style-guide: Style Guide webtui/contributing/style-guide: ## CSS Units webtui/contributing/style-guide: ## Selectors webtui/contributing/style-guide: ## Documentation webtui/installation/astro: Astro webtui/installation/astro: ## Scoping webtui/installation/astro: ### Frontmatter Imports webtui/installation/astro: ### ‹style› tag webtui/installation/astro: ### Full Library Import webtui/installation/nextjs: Next.js webtui/installation/vite: Vite webtui/plugins/plugin-dev: Developing Plugins webtui/plugins/plugin-dev: ### Style Layers webtui/plugins/plugin-nf: Nerd Font Plugin webtui/plugins/theme-catppuccin: Catppuccin Theme webtui/plugins/theme-custom: Custom Theme webtui/plugins/theme-everforest: Everforest Theme webtui/plugins/theme-gruvbox: Gruvbox Theme webtui/plugins/theme-nord: Nord Theme webtui/plugins/theme-vitesse: Vitesse Theme webtui/start/ascii-boxes: ASCII Boxes webtui/start/changelog: Changelog webtui/start/installation: Installation webtui/start/installation: ## Installation webtui/start/installation: ## Using CSS webtui/start/installation: ## Using ESM webtui/start/installation: ## Using a CDN webtui/start/installation: ## Full Library Import webtui/start/installation: ### CSS webtui/start/installation: ### ESM webtui/start/installation: ### CDN webtui/start/intro: Introduction webtui/start/intro: ## Features webtui/start/plugins: Plugins webtui/start/plugins: ## Official Plugins webtui/start/plugins: ### Themes webtui/start/plugins: ## Community Plugins webtui/start/theming: Theming webtui/start/theming: ## CSS Variables webtui/start/theming: ### Font Styles webtui/start/theming: ### Colors webtui/start/theming: ### Light & Dark webtui/start/theming: ## Theme Plugins webtui/start/theming: ### Using Multiple Theme Accents webtui/start/tuis-vs-guis: TUIs vs GUIs webtui/start/tuis-vs-guis: ## Monospace Fonts webtui/start/tuis-vs-guis: ## Character Cells
 Theme Current: Light j/k or ↑/↓ + Enter

A Small Service Mesh for My Macs and Supercomputers

How I run local dashboards, separate personal and agent knowledge vaults, shared agent memory, email, terminal history, model gateways, and persistent agent sessions across two Macs and ALCF systems -- with one health checker for the whole stack.

I have accumulated a surprising number of small web services around my daily work: a training dashboard, terminal history, separate personal and agent knowledge vaults, an agent memory API, an email API, a model gateway, and persistent agent sessions. None is a large application. The interesting part is making each one available in the right places without making everything public.

This post is an inventory of that system: what each service does, how the network paths fit together, and the commands and launchd patterns I use to bring them up again. It is the implementation-level companion to Working From Anywhere, which explains the larger remote-work design.

TL;DR — the pattern
  • Applications listen on 127.0.0.1, not every network interface.
  • launchd keeps long-lived processes alive on macOS.
  • Tailscale Serve gives selected HTTP services private HTTPS URLs inside my tailnet.
  • SSH forwards move a port to the machine that needs it.
  • Tailcat is a second point-to-point path when a network blocks Tailscale’s control plane.
  • Cloudflare Tunnel is reserved for one browser-facing service that must be reachable without joining my tailnet, and Cloudflare Access authenticates it.
  • A listening process is not a health check. I verify the final URL from the machine that will consume it.
  • One deterministic checker now verifies 14 contracts across the local services, private routes, MCP connections, schedulers, and agent panes.

The topology

The always-on MacBook, mbph, is the hub and the public edge for the one service published through Cloudflare. A second MacBook, mbpr, is often on campus networks and acts as a client. Aurora and CELS are remote compute environments behind SSH bastions.

                      mbpr (mobile Mac)
                        │           │
          SSH forwards  │           │ Tailcat forwards
                        ▼           ▼
                       mbph (home Mac)
      ┌────────────┬──────────┴─────────┬─────────────────┐
      │            │                    │                 │
  Tailscale    localhost            SSH jump          herdr
    Serve      services               hosts           server
      │            │                    │                 │
      ▼            ▼                    ▼                 ▼
  browsers      agents            Aurora / CELS    Heeler / relay

Here is the concrete inventory. Ports are included because they make debugging far easier than descriptions like “the notes service.”

ServiceLocal backendReachabilityWhat it is for
Scrollbackmbph:8766Tailscale HTTPS :10443Search and read terminal history away from the terminal
SilverBulletmbph:3000Tailscale HTTPS :9443Human-friendly editing of the shared agent-memory Markdown
ai-memorymbph:49375Tailscale HTTPS :8443 and TailcatOne persistent memory/MCP service for agents on every machine
AGPT dashboardmbph:8720Tailscale HTTPS :8720 and an SSH forward to mbpr:8720Monitor AuroraGPT runs from a browser
Personal vaultmbph:8891Tailscale HTTPS :443Read-only browsing and search for my Obsidian vault
Agent vaultloopbackSeparate private Tailscale routeGit-backed history, evidence, and generated agent notes
llm-rosettambph:8765Loopback; clients reach it through the hostOne OpenAI-compatible endpoint for local and ALCF models
Hermes gatewaylocal IPCLocal clients and configured messaging transportsRoutes prompts, tools, MCP servers, cron, and notifications
AgentMail MCPsubprocess/APIHermes MCP onlyDedicated agent inbox, drafts, send, receive, and replies
Herdrlocal socketLocal clients and the separately authenticated relayPersistent panes, process detection, and agent status
Aurora data APIAurora :8712SSH local forward to mbph:8712Feed live training data into the AGPT dashboard
Argo / CELSCELS HTTPS :443SSH jump forward to mbph:25939Reach the internal model gateway from local clients
Tailcat SSHmbph:22mbpr:2222A backup path to SSH when Tailscale is unavailable
Tailcat memorymbph:49375mbpr:8443The same backup path for agent memory
AGPT reverse viewmbph:8720SSH local forward to mbpr:8720Keep the dashboard at a stable localhost URL on the mobile Mac
Herdr relaymbph:8375Cloudflare at relay.sf.onlAttach a browser to persistent terminal/agent sessions

The table contains both applications and multiple routes to some applications. That distinction is useful: an application owns its local port; a transport merely decides who can reach it. I can replace Tailscale with an SSH forward without changing the application.

The common service recipe

Every local application starts life on loopback:

my-service --host 127.0.0.1 --port 9000
curl --fail http://127.0.0.1:9000/healthz

Binding to 127.0.0.1 means the process is not accidentally exposed on Wi-Fi, Ethernet, or a VPN interface. I then make persistence explicit with a macOS LaunchAgent:

<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN"
  "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0">
<dict>
  <key>Label</key>
  <string>dev.example.my-service</string>
  <key>ProgramArguments</key>
  <array>
    <string>/Users/me/.local/bin/my-service</string>
  </array>
  <key>RunAtLoad</key>
  <true/>
  <key>KeepAlive</key>
  <true/>
  <key>StandardOutPath</key>
  <string>/Users/me/Library/Logs/my-service.log</string>
  <key>StandardErrorPath</key>
  <string>/Users/me/Library/Logs/my-service.err.log</string>
</dict>
</plist>

Install it once, or restart it after a change:

label=dev.example.my-service
plist="$HOME/Library/LaunchAgents/$label.plist"

plutil -lint "$plist"
launchctl bootstrap "gui/$(id -u)" "$plist"     # first install
launchctl kickstart -k "gui/$(id -u)/$label"    # restart
launchctl print "gui/$(id -u)/$label"           # inspect

RunAtLoad starts the job at login. KeepAlive restarts it if it exits. Logs go to stable files instead of disappearing with a terminal window.

Warning

KeepAlive is not proof that a service works. It can keep restarting a crashing process, or keep a wedged process alive forever. Always test the backend and the externally consumed route separately.

Private web apps with Tailscale Serve

Tailscale Serve terminates HTTPS on the tailnet hostname and proxies to a loopback HTTP server. These are the public-safe route definitions from mbph; the agent-vault route stays in private configuration:

tailscale serve --bg --https=10443 http://127.0.0.1:8766   # Scrollback
tailscale serve --bg --https=9443  http://127.0.0.1:3000   # SilverBullet
tailscale serve --bg --https=8443  http://127.0.0.1:49375  # ai-memory
tailscale serve --bg --https=8720  http://127.0.0.1:8720   # AGPT dashboard
tailscale serve --bg --https=443   http://127.0.0.1:8891   # VaultServe
tailscale serve status

Serve configuration persists in Tailscale, so these commands define routes; they do not need to remain running in a shell. The backend processes still need their own supervision.

This is Tailscale Serve, not Funnel. Serve is tailnet-only. Funnel would put a service on the public internet, which is not what I want for terminal history, notes, dashboards, or agent memory.

Scrollback: searchable terminal history

Scrollback turns captured terminal output into a small searchable web application. This is useful when I remember seeing an error or command but not which terminal, host, or session contained it. It is also much easier to read long output on a phone than through a terminal multiplexer.

My wrapper allows the private Tailscale hostname while keeping the socket local:

#!/usr/bin/env python3
import uvicorn
from scrollback.web.app import create_app

uvicorn.run(
    create_app(allowed_hosts=["my-mac.my-tailnet.ts.net"]),
    host="127.0.0.1",
    port=8766,
)

Run the wrapper from a LaunchAgent, then add the :10443 Serve route shown above. The hostname allowlist matters: the proxy preserves an HTTP host that a default local-only configuration may reject.

SilverBullet: a human view of agent memory

SilverBullet is a Markdown knowledge base with wiki links, search, and a pleasant browser editor. I point it at a mounted/synchronized view of my agent-memory wiki:

mount-agent-memory
silverbullet --single \
  --hostname 127.0.0.1 \
  --port 3000 \
  "$HOME/AgentMemory"

This gives humans a useful complement to programmatic retrieval: I can browse a project’s decisions, edit a durable page, or follow links across related notes.

SilverBullet in this configuration does not add application-level authentication. Tailscale makes it private to the tailnet, not private to one person inside that tailnet. On a shared tailnet I would add a restrictive Tailscale grant/ACL for port 9443, and preferably application authentication as defense in depth.

ai-memory: one memory service for every agent

ai-memory stores captured agent sessions plus a Git-backed Markdown wiki. Every supported coding harness points at the same MCP endpoint, so Claude Code, Codex, OpenCode, and other clients can retrieve the same project history and hand work to one another.

A minimal local launch looks like:

export AI_MEMORY_AUTH_TOKEN="$(openssl rand -hex 32)"
ai-memory serve \
  --transport http \
  --bind 127.0.0.1:49375 \
  --enable-web

I put the token in the LaunchAgent environment or a protected environment file, never in a public repository. The service is exposed through Tailscale on :8443; clients use the /mcp path and send the bearer token.

/mcp is a machine endpoint, not a web page. Opening it in a normal browser correctly returns 401 because the browser did not send the bearer token. That response proves only that the route reached ai-memory and that authentication ran; it does not prove an MCP client can log in. A complete health check sends an authenticated initialize request and verifies the JSON-RPC response:

curl --fail-with-body \
  -H "Authorization: Bearer $AI_MEMORY_AUTH_TOKEN" \
  -H 'Content-Type: application/json' \
  -H 'Accept: application/json, text/event-stream' \
  --data '{
    "jsonrpc": "2.0",
    "id": 1,
    "method": "initialize",
    "params": {
      "protocolVersion": "2025-06-18",
      "capabilities": {},
      "clientInfo": {"name": "healthcheck", "version": "1.0"}
    }
  }' \
  https://my-mac.my-tailnet.ts.net:8443/mcp

A healthy response is HTTP 200 with result.serverInfo.name equal to ai-memory. The optional human web interface is a separate /web route and needs browser/session authentication configured before it is useful as a page.

For a new client, the project README’s installers are preferable to hand-editing every harness:

ai-memory install-mcp --client claude-code --apply
ai-memory install-hooks --agent claude-code --apply

AGPT dashboard: training state without an SSH terminal

The AGPT dashboard is a small local web app for current and historical AuroraGPT training runs. Its backend merges cached run metadata with a live data API forwarded from Aurora. I start the dashboard itself on mbph:

cd ~/agpt-dash
uv run python3 server.py --port 8720
curl --fail http://127.0.0.1:8720/

Tailscale Serve makes it available to my tailnet at HTTPS port 8720. On mbpr, I additionally keep the same dashboard at a predictable localhost URL:

ssh -N \
  -o BatchMode=yes \
  -o ControlMaster=no \
  -o ControlPath=none \
  -o ExitOnForwardFailure=yes \
  -o ServerAliveInterval=30 \
  -o ServerAliveCountMax=3 \
  -L 127.0.0.1:8720:127.0.0.1:8720 \
  mbph

That SSH command lives in a script supervised by launchd on mbpr. Disabling SSH connection sharing is intentional: a multiplexed master can accept the forward and detach from the process that launchd is supervising, leaving job state and actual port ownership out of sync.

VaultServe: read-only notes in a browser

VaultServe is my intentionally small, read-only Obsidian viewer. It renders Markdown, resolves wiki links and embedded media, and searches the vault without giving the browser a write API.

VAULT_ROOT="$HOME/Obsidian/Notes" \
VAULT_BIND=127.0.0.1 \
VAULT_PORT=8891 \
  ~/.hermes/vaultserve/.venv/bin/python \
  ~/.hermes/vaultserve/server.py

This is useful on devices where I want to consult notes without installing or synchronizing the full Obsidian vault. The default Tailscale HTTPS route proxies to it, but it remains tailnet-only.

Two vaults, two trust boundaries

The personal vault and the agent vault use the same small Python server, but they are separate services with separate roots, repositories, ports, and sync rules.

The personal service on 8891 reads my Obsidian tree. Automation treats that tree as read-only. The agent service on 8892 serves a private Git repository containing curated memory, compiled session evidence, project histories, and a generated index of recently modified notes. Its daily sync job stages only the paths it owns; it cannot sweep an unrelated hand-edited history page into an automated commit.

This split is simpler than teaching one application which pages are personal, generated, editable, or publishable. A route now implies one source tree and one policy.

Hermes, AgentMail, and Herdr

The browser services are only half of the stack. Hermes runs the agent control plane: model routing, MCP servers, scheduled jobs, and notifications. A dedicated AgentMail inbox is connected through an MCP subprocess with a narrow tool allowlist. I verified the full loop with one outbound message, one inbound reply, and one threaded reply. Credentials stay in a protected environment file; the mailbox address and provider identifiers do not belong in public configuration.

Herdr owns persistent terminal panes and reports process state to Heeler. Current Hermes processes start through a Python bootstrap, so process detection has to recognize the bootstrap signature rather than only a binary named hermes. The local detector now does that without treating arbitrary Python source containing hermes_cli as an agent. The three HPC panes remain attached while Herdr reports each as hermes instead of unknown.

Bringing HPC services home with SSH

Some services cannot originate on either Mac. They run behind institutional login nodes, so SSH is the transport.

Aurora data for the dashboard

The training-data service listens on Aurora’s loopback port 8712. A local forward makes it look like an mbph service:

ssh -N \
  -o ExitOnForwardFailure=yes \
  -o ServerAliveInterval=30 \
  -L 127.0.0.1:8712:127.0.0.1:8712 \
  aurora

curl --fail http://127.0.0.1:8712/api/backbone

The dashboard only knows about 127.0.0.1:8712; it does not need to know about Aurora’s bastions or network topology. This is the same indirection principle as Tailscale Serve, in the opposite direction.

Argo through CELS

ALCF’s Argo service is HTTPS inside the CELS network. An SSH jump host carries that endpoint back to a local TLS port:

ssh -N -f \
  -o BatchMode=yes \
  -o ExitOnForwardFailure=yes \
  -o ServerAliveInterval=15 \
  -J "$USER@logins.cels.anl.gov" \
  -L 127.0.0.1:25939:apps.inside.anl.gov:443 \
  "$USER@compute-01.cels.anl.gov"

My actual entry point is argo-shim, which wraps the tunnel and the authentication details:

uvx --no-cache argo-shim --host compute-01.cels.anl.gov

Local model gateways can now speak to 127.0.0.1:25939 while preserving the correct upstream TLS server name. A useful acceptance test is a complete TLS handshake, not merely seeing the SSH process:

openssl s_client \
  -connect 127.0.0.1:25939 \
  -servername apps.inside.anl.gov \
  -brief </dev/null

A fallback data plane with Tailcat

Tailscale is normally the cleanest path between my Macs, but one campus network has blocked its control connection. I wanted the services above to depend on “a port exists here,” not on one particular VPN implementation, so I added a second point-to-point transport using Tailcat.

On mbph, the server is restricted to the peer’s public node key:

tailcat serve --allow='nodekey:<MBPR_PUBLIC_NODE_KEY>' all

On mbpr, one client process creates two loopback forwards:

tailcat forward '<MBPH_TAILCAT_ADDRESS>' \
  2222:22 \
  8443:49375

The first mapping makes mbph SSH available at mbpr:2222; the second makes the ai-memory API available at mbpr:8443. Both commands run under LaunchAgents with KeepAlive and a short throttle interval.

The node key and Tailcat address are capabilities. I keep the real values in private configuration and use placeholders here. The key design point is that the services above do not change: SSH still speaks SSH and ai-memory still speaks HTTP/MCP. Only the bytes’ route between the Macs changes.

One carefully public route with Cloudflare Tunnel

Everything so far assumes the client can join my tailnet or reach an SSH path. The exception is Herdr’s browser relay, which lets me attach to a persistent terminal/agent session from an ordinary browser. That needs a public hostname, so it gets a separate and more explicit boundary.

The relay binds to loopback on mbph:

HERDR_RELAY_PORT=8375 uv run ./relay/herdr_relay.py
curl -I http://127.0.0.1:8375/

This one detail matters more than it looks: the relay shells out to the local herdr binary, so it can only ever report the sessions on the machine it runs on. The relay has to live wherever the agents live. It originally ran on mbpr, and once most of my agents had moved to mbph the public URL was faithfully serving the wrong machine’s terminals. Moving it was the fix.

A named Cloudflare Tunnel publishes only that backend:

tunnel: herdr-relay
credentials-file: /Users/me/.cloudflared/<TUNNEL-ID>.json

ingress:
    - hostname: relay.example.com
      service: http://localhost:8375
    - service: http_status:404

Create and run it with:

cloudflared tunnel login
cloudflared tunnel create herdr-relay
cloudflared tunnel route dns herdr-relay relay.example.com
cloudflared tunnel \
  --config "$HOME/.cloudflared/config-herdr.yml" \
  run herdr-relay

The final catch-all 404 prevents the tunnel from becoming an accidental general-purpose proxy. In production I also put the hostname behind Cloudflare Access; the relay retains its own token as a fallback. The tunnel credential, access audience, and relay token live in private files and never in the LaunchAgent plist or repository.

Two LaunchAgents supervise this path independently:

  1. the Herdr relay on 127.0.0.1:8375;
  2. cloudflared, which maintains outbound connections to Cloudflare.

Splitting them makes failures legible. I can test the local relay first, then the authenticated public URL, and know which half is broken.

Moving the relay to another machine

Because the relay reports on whatever machine it runs on, “which host serves the public URL” is a decision I expect to revisit. The useful property of a named tunnel is that moving it needs no DNS change at all: relay.example.com is a CNAME pointing at the tunnel’s UUID, not at a host. Moving the connector is invisible from the outside.

The move is a cutover, not a parallel run. Two copies of the relay on one LAN collide on mDNS (both register the same hardcoded herdr-remote._herdr-remote._tcp.local. name), and two connectors for one tunnel is not a state worth reasoning about. So: stop the old host first.

Before touching anything, get any uncommitted work off the old machine. Mine had an unpushed feature sitting in the checkout, which a migration is an excellent way to lose:

git -C ~/projects/herdr-remote status --short

Install the prerequisites on the new host — cloudflared, uv, and herdr itself — then clone the relay at the same commit the old host was running.

Copy the credentials and configuration. The tunnel credential and the cert.pem are what let the new host claim the same tunnel; without them cloudflared has no identity:

scp ~/.cloudflared/cert.pem newhost:~/.cloudflared/cert.pem
scp ~/.cloudflared/<TUNNEL-ID>.json newhost:~/.cloudflared/
ssh newhost 'chmod 600 ~/.cloudflared/cert.pem; chmod 400 ~/.cloudflared/*.json'

If the two machines have different usernames — mine do — every absolute path in the config, the env files, and the LaunchAgent plists has to be rewritten, and then verified, because a plist with a bad path fails quietly:

sed 's#/Users/olduser#/Users/newuser#g' config.env > /tmp/config.env
ssh newhost 'for p in $(grep -oE "/Users/[^ ]*" ~/.config/herdr-remote/config.env); do
  [ -e "$p" ] && echo "OK   $p" || echo "MISS $p"
done'

Then cut over. Unload on the old host, confirm the port is actually released, and only then load on the new one:

# Old host
launchctl unload ~/Library/LaunchAgents/com.herdr-remote.{relay,tunnel}.plist
lsof -nP -iTCP:8375 -sTCP:LISTEN   # must be empty

# New host
launchctl load ~/Library/LaunchAgents/com.herdr-remote.{relay,tunnel}.plist

Verifying this is where it gets interesting, because the obvious check is useless. curl https://relay.example.com/ returns the same 302 before and after the move — that is Cloudflare Access redirecting to SSO, and it would keep returning 302 even if I had migrated nothing. The public URL cannot tell me which machine is behind it.

Two checks actually can. The tunnel’s own connector list reports the connector’s cloudflared version and creation time, so if the new host runs a different build the version alone identifies it:

cloudflared tunnel info herdr-relay

I expect exactly one connector, created at cutover time. And then the only check that really matters — ask the relay what it is actually serving, and confirm the session paths belong to the new machine:

ssh newhost 'herdr pane list' | head

Finally, keep the old host’s LaunchAgents on disk but renamed, so they do not silently reload at next login while remaining available for rollback:

mv com.herdr-remote.relay.plist com.herdr-remote.relay.plist.migrated-YYYYMMDD

One log line will look alarming and is not: the relay throws a zeroconf NonUniqueNameException if anything else on the LAN still advertises that mDNS name. It runs in a daemon thread, and the relay logs relay on :8375 and Polling: local immediately afterwards. Read the lines after the traceback before concluding anything is broken.

Starting and checking the whole stack

I used to keep this as a list of launchctl kickstart lines per machine, and copy-pasted the relevant block. That was a bad habit for the reason this whole post keeps circling: kickstarting an agent only asks launchd to run something. It says nothing about whether the service came back. The list also drifted — I added the spool drainer and the export job and never updated the snippet.

So restart commands live in one start-services script, while a separate check-services.py verifies the steady state without changing it:

start-services            # restart everything, then verify
start-services --check    # verify only, change nothing
start-services --serve    # also re-declare the Tailscale Serve routes
~/.hermes/scripts/check-services.py
~/.hermes/scripts/check-services.py --json

It picks its service list from hostname -s, so the same script is correct on both machines. Each entry is a LaunchAgent label, an optional probe URL, and a description:

read -r -d '' SERVICES_MBPH <<'EOF'
dev.saforem2.scrollback|http://127.0.0.1:8766/|Scrollback terminal history
dev.saforem2.silverbullet-agent-memory|http://127.0.0.1:3000/|SilverBullet
com.github.akitaonrails.ai-memory|http://127.0.0.1:49375/mcp|ai-memory server
dev.saforem2.vaultserve|http://127.0.0.1:8891/|VaultServe
sh.samf.tailcat-serve||Tailcat listener
com.herdr-remote.relay|http://127.0.0.1:8375/|Herdr relay
com.herdr-remote.tunnel||Cloudflare tunnel
dev.saforem2.ai-memory-hook-drain||ai-memory spool drainer
dev.saforem2.ai-memory-silverbullet-sync||ai-memory -> SilverBullet export
EOF

The verdict logic is the part worth stealing. 401 and 403 count as healthy: they prove TLS, routing, and the authorization boundary all ran, and the service is correctly refusing an unauthenticated probe. Only 000 — nothing answered at all — is an unambiguous failure:

case "$code" in
  2*|3*|401|403) printf '  %-34s ok (HTTP %s)\n' "$desc" "$code" ;;
  000)           printf '  %-34s NO RESPONSE\n'  "$desc"; fail=$((fail+1)) ;;
  *)             printf '  %-34s HTTP %s\n'      "$desc" "$code"; fail=$((fail+1)) ;;
esac

Two smaller details matter. A periodic agent such as the spool drainer shows - instead of a PID between runs, which is healthy rather than missing. And the probe URL has to be a path the service actually routes: ai-memory answers 404 on / and 401 on /mcp, so probing / reports a false failure.

The checker has explicit contracts for the Hermes gateway, AgentMail MCP, ai-memory MCP, llm-rosetta, both vaults, SilverBullet, the required Tailscale Serve routes, Herdr, the three HPC agent panes, the Hermes cron scheduler, the AGPT dashboard, and Scrollback. The local Hermes dashboard is optional. A required failure exits 1; an optional failure does not change the exit status.

The important distinction is protocol health versus port health. The two MCP checks perform real MCP connection tests. The Tailscale check parses the route table and confirms each backend is listening. The Herdr check requires a compatible live server, then verifies that the Aurora, Sunspot, and Polaris panes are still classified as hermes rather than unknown.

The current result is 14 OK, 0 warnings, 0 failures:

OK   Hermes gateway
OK   AgentMail MCP
OK   ai-memory MCP
OK   llm-rosetta
OK   personal vault
OK   agent vault
OK   SilverBullet
OK   Tailscale
OK   Herdr
OK   Hermes HPC panes
OK   Hermes cron
OK   AGPT dashboard
OK   Scrollback viewer
OK   Hermes dashboard [optional]
SUMMARY ok=14 warn=0 fail=0

The same checker runs twice daily. Its wrapper prints nothing when every contract is healthy, so the notification path stays quiet unless there is a warning or failure. The JSON form carries the same verdict for other automation.

Warning

Restarting a service restarts everything downstream of it, and the downstream side may not notice. When I restarted ai-memory on the hub, the Tailcat forward on the other Mac kept listening on its local port while its far end was gone — launchctl list reported the forward as -9, and a request through it hung instead of failing. Run start-services --check on both machines after restarting anything shared.

Verify paths, not processes

This stack recently produced five instructive failures:

  • Tailscale Serve was healthy but proxied the AGPT URL to stale port 8726 instead of the live dashboard on 8720.
  • VaultServe still accepted TCP connections on 8890 but closed every HTTP request without sending a response.
  • An SSH control master owned mbpr:8720, while the LaunchAgent that was supposed to own the forward was not actually running.
  • The Herdr relay served relay.sf.onl correctly from the wrong machine for weeks. Every probe passed, because the public URL returns an identical Cloudflare Access 302 no matter which host is behind the tunnel.
  • llm-rosetta answered correctly on IPv4 while a stale review server owned the same port on IPv6. 127.0.0.1 returned the model list; localhost resolved to the other process and returned 404.

All five could look “up” in a process list. The acceptance checks need to cross the same boundary as the real consumer:

# Backend first
curl --fail http://127.0.0.1:8891/healthz

# Then the tailnet route
curl --fail https://my-mac.my-tailnet.ts.net/

# Reachability/auth-boundary check only; this is not a full health check
curl -sS -o /dev/null -w '%{http_code}\n' \
  https://my-mac.my-tailnet.ts.net:8443/mcp

# Test the forwarded dashboard on the consuming Mac, not the server
ssh mbpr 'curl --fail http://127.0.0.1:8720/'

# Inspect ownership as well as presence
lsof -nP -iTCP:8720 -sTCP:LISTEN
launchctl print gui/$(id -u)/sh.samf.agpt-forward-mbph

# Check both address families when localhost and 127.0.0.1 disagree
curl --fail http://127.0.0.1:8765/v1/models
curl --fail http://[::1]:8765/v1/models

For an authenticated API, 401 proves that TLS, routing, proxying, and the authorization boundary ran. It does not prove valid credentials or protocol initialization; use the authenticated MCP request above for that. Conversely, a Tailscale-generated 502 means the route exists but its backend is unavailable or failed to produce HTTP.

Security boundaries

The useful question is not “is it encrypted?” Every path here is encrypted. The question is who is allowed to arrive at the application after decryption?

BoundaryWho can reach itAdditional control
Loopback backendProcesses on that hostFilesystem/user permissions
Tailscale ServeDevices/users admitted to the tailnetTailnet grants/ACLs; app auth where available
SSH forwardA user with an accepted SSH credentialSSH config and remote account permissions
Tailcat forwardA peer holding the allowed key/addressLoopback binding at the client
Cloudflare TunnelThe public internet can reach the edgeCloudflare Access plus application auth

There are a few rules I now apply consistently:

  1. Bind local unless public is deliberate. No application here needs 0.0.0.0.
  2. Treat tailnet-only as private, not secret. Every authorized tailnet member can potentially reach an unauthenticated Serve route unless grants say otherwise.
  3. Keep credentials out of examples, repositories, and process arguments. Use protected environment files or the platform’s credential store.
  4. Publish one hostname, not a network. The Cloudflare tunnel has one ingress rule and a terminating 404 rule.
  5. Test from the consumer. A browser route is only healthy if a browser-side HTTP request works; a forwarded API is only healthy if the client host can use it.

What this buys me

The obvious payoff is convenience: terminal history on my phone, training state without an Aurora login shell, notes from a browser, and agent context that does not depend on which laptop or coding harness I opened.

The larger payoff is replaceability. Each application sees a local port. Each consumer sees a stable URL or local port. Between them I can choose Tailscale, SSH, Tailcat, or Cloudflare according to the trust boundary and the network I am currently on. When one transport fails, I do not have to redesign the service.

That is enough “service mesh” for two Macs and a few supercomputers: not a cluster orchestrator, just small processes with explicit ownership, narrow network paths, persistent supervision, and health checks that exercise the whole route.

 samf.sh / posts / 2026 / 09 / 25 · Top 1:1