 Command

Sam Foreman's personal site. Vim-style keybinds for navigation; theme + font pickers below.

Theme
 Font Body Code
Reader
Keybinds
Navigation
j / ↓ Next item k / ↑ Previous item g First item in region G Last item in region zz Center focused item h / l Sidebar / main content ] / [ Next/previous heading } / { Next/previous block d / u Half-page down/up
Layout
<zh> Toggle sidebar <zr> Toggle reader view <zj> / <zk> Focus main / actions ⇧C / ⇧E  ·  <zM> / <zR> Collapse / expand all sections
Dialogs
⌃P / : Command palette ⌃X Theme picker / Search ? Show keybinds ⌃N / ⌃P Next/prev search result Esc Close dialog / exit reader
History
n Next document b Previous document ⌃O History back ⌃I History forward
Sections
a about p posts t talks m more s style
 Search
about: Sam Foreman about/more: 🪪 More ideas: 💡 Ideas more: ➕ More now: Now posts: 📬 Posts posts/2023/12/05: 🔳 l2hmc-qcd Example: 4D SU(3) posts/2025: 📆 2025 posts/2025/04/28: 🔥 Building PyTorch 2.6 from Source on Aurora posts/2025/05/03: 🚧 Frameworks Issue with numpy \› 2 posts/2025/06: 06 posts/2025/06/01: 📰 Nice Headings posts/2025/06/02: 🧜‍♀️ Mermaid posts/2025/06/14: 🏗️ Building PyTorch 2.8 from Source on Aurora posts/2025/09/12: 🍹 BlendCorpus + TorchTitan @ ALCF posts/2025/09/17: 📊 pbs-tui: TUI for PBS Job Scheduler Monitoring posts/2025/10/06: 🎨 Mixing Between Distributions While Training posts/2025/11/12: 🧊 Cooling Down Checkpoints: Best Practices for Model Evaluation posts/2026/01/07: 🎉 Happy New Year! posts/2026/01/10: 🍋 ezpz: distributed PyTorch across any hardware posts/2026/02/28: ⏱️ Comparing Launchers on Aurora posts/2026/02/28: ## torchrun posts/2026/02/28: ## ezpz posts/2026/04/27: Pre-Training AuroraGPT with TorchTitan posts/2026/04/27: ## Two-Week Summary (Apr 12–27, 2026) posts/2026/04/27: ## Detailed Breakdown posts/2026/04/27: ### Week 1: Apr 12–18 — Benchmarking, LR Finder, XPU Fixes posts/2026/04/27: #### Benchmarking (Apr 12–15) posts/2026/04/27: #### LR Finder (Apr 12–14) posts/2026/04/27: #### Scaling Study (Apr 12) posts/2026/04/27: #### Upstream Syncs (Apr 12–18, syncs 6–14) posts/2026/04/27: #### XPU Bug Fixes (Apr 18) posts/2026/04/27: #### RL Experiment (Apr 18) posts/2026/04/27: ### Week 1.5: Apr 18–25 — Production Readiness posts/2026/04/27: #### Torch 2.12 Benchmarks (Apr 18) posts/2026/04/27: #### LR Finder Extensions (Apr 20–21) posts/2026/04/27: #### XPU Fixes (Apr 23) posts/2026/04/27: #### Torch 2.13 Environment (Apr 25) posts/2026/04/27: #### 2B Scaling Study on Torch 2.13 (Apr 25) posts/2026/04/27: #### Production Training (Apr 25) posts/2026/04/27: ### Week 2: Apr 26–27 — Optimizer Competition posts/2026/04/27: #### RL Multi-Task Refactor (Apr 26) posts/2026/04/27: #### Docs Reorganization (Apr 26) posts/2026/04/27: #### Generic HF Dataset Streaming (Apr 26) posts/2026/04/27: #### New Optimizers (Apr 26) posts/2026/04/27: #### Architecture Tweaks (Apr 26–27) posts/2026/04/27: ## Competition Results posts/2026/04/27: ### Round 1–3: Speedrun — 2N, GBS=48, 1000 steps posts/2026/04/27: ### 10B Full Training — 8N, GBS=384, ~3,178 steps posts/2026/04/27: ### Round 4: Reproducible Speedrun — 2N, GAS=8, GBS=384, 1000 steps posts/2026/04/27: ## Key Discoveries posts/2026/04/27: ## Infrastructure Built posts/2026/04/27: ## High-Level posts/2026/04/27: ## Detailed Breakdown posts/2026/04/27: ### Week 1: Apr 12–18 — Benchmarking, LR Finder, XPU Fixes posts/2026/04/27: #### Benchmarking (Apr 12–15) posts/2026/04/27: #### LR Finder (Apr 12–14) posts/2026/04/27: #### Scaling Study (Apr 12) posts/2026/04/27: #### Upstream Syncs (Apr 12–18, syncs 6–14) posts/2026/04/27: #### XPU Bug Fixes (Apr 18) posts/2026/04/27: #### RL Experiment (Apr 18) posts/2026/04/27: ### Week 1.5: Apr 18–25 — Production Readiness posts/2026/04/27: #### Torch 2.12 Benchmarks (Apr 18) posts/2026/04/27: #### LR Finder Extensions (Apr 20–21) posts/2026/04/27: #### XPU Fixes (Apr 23) posts/2026/04/27: #### Torch 2.13 Environment (Apr 25) posts/2026/04/27: #### 2B Scaling Study on Torch 2.13 (Apr 25) posts/2026/04/27: #### Production Training (Apr 25) posts/2026/04/27: ### Week 2: Apr 26–27 — Optimizer Competition posts/2026/04/27: #### RL Multi-Task Refactor (Apr 26) posts/2026/04/27: #### Docs Reorganization (Apr 26) posts/2026/04/27: #### Generic HF Dataset Streaming (Apr 26) posts/2026/04/27: #### New Optimizers (Apr 26) posts/2026/04/27: #### Architecture Tweaks (Apr 26–27) posts/2026/04/27: ## Competition Results posts/2026/04/27: ### Round 1–3: 1000-step speedruns, 2 nodes, GBS=48 (17 configs) posts/2026/04/27: ### Round 4 (10B full training, 8 nodes, GBS=384, 5 configs) posts/2026/04/27: ### Round 5 (2 nodes, GAS=8, GBS=384, local dataset, 8 configs — in progress) posts/2026/04/27: ## Key Discoveries posts/2026/04/27: ## Infrastructure Built posts/2026/05/01: Running 50k Python Processes on Aurora with ezpz yeet posts/2026/06/27: Local AI Apps on ALCF: Argo, Inference Endpoints, and One Gateway posts/2026/06/28: Migrating from Quarto to Astro: samforeman.me → samf.sh posts/2026/08/08: Pre-Training LLMs on a Supercomputer posts/2026/09/22: Working From Anywhere: Persistent Access to Compute and Context posts/2026/09/25: A Small Service Mesh for My Macs and Supercomputers posts/ai-for-physics: ⚛️ AI for Physics posts/ai-for-physics/diffusion: 🎲 MCMC + Diffusion Sampling posts/ai-for-physics/l2hmc-qcd: 🎢 L2HMC for LQCD posts/ai-for-physics/l2hmc-qcd/2du1: 🎢 l2hmc-qcd Example: 2D U(1) posts/auroragpt: 🤖 AuroraGPT posts/auroragpt/aurora-gpt: 🏎️ Megatron-DeepSpeed on Intel XPU posts/auroragpt/checkpoints: 💾 Converting Checkpoints posts/auroragpt/determinstic-flash-attn/deterministic-flash-attn: 🎰 Deterministic flash-attn posts/auroragpt/flash-attn-sunspot: 📸 flash-attn on Sunspot posts/auroragpt/long-sequences: 🚂 Loooooooong Sequence Lengths posts/auroragpt/mpi4py-reproducer: 🐛 mpi4py bug on Sunspot posts/auroragpt/spike-skipper: 🏔️ Spike Skipper posts/auroragpt/startup-times: 🐢 Starting Up Distributed Training on Aurora posts/auroragpt/startup-times: ## Response posts/auroragpt/startup-times: ### Measuring / Calculating Startup Time posts/auroragpt/startup-times: ## Minimal Working Example posts/croon: croon: Synced Lyrics in the Terminal, for Whatever Is Playing posts/dope-slides: 💅 How to Make Dope Slides posts/drafts/2025/09/22: 📝 2025 Annual Report posts/ezpz-at-alcf: 🍋 ezpz @ ALCF posts/ezpz-v1: 📝 ezpz-v1 posts/globusfs: globusfs: an fsspec Filesystem for Globus Collections posts/jupyter: 📗 Jupyter posts/jupyter/test: 🏁 l2hmc Example: 2D $U(1)$ posts/resume: 🧑🏻‍💻 Sam Foreman’s Résumé posts/svgbob: 🫥 svgbob posts/torchtune-aurora: 🪛 Torchtune on Aurora posts/torchtune-patch-aurora: 🚑 Torchtune Patch on Aurora posts/wandb-tui: wandb-tui: Comparing W&B Runs Without Leaving the Terminal projects: 📚 Projects talks: 🎙️ Talks talks/2025/09/24: Training Foundation Models on Supercomputers talks/2025/10/08: AERIS: Argonne's Earth Systems Model talks/2025/10/15: Training Foundation Models on Supercomputers talks/2025/10/24: Training Foundation Models on Supercomputers talks/2025/12/16: AuroraGPT: Training Foundation Models on Supercomputers talks/2026/06/03: Production Pre-Training at Scale: The Good, the Bad, and the Restarts talks/2026/07/14: Pre-Training AuroraGPT at Scale on Aurora talks/2026/08/03: Pre-Training LLMs on a Supercomputer talks/ai-for-science-2024: Parallel Training Methods talks/alcf-hpc-workshop-2024/alcf-hpc-workshop-2024: Deep Learning and Foundation Models at Scale talks/aurora-gpt-fm-for-electric-grid/auroragpt-fm-for-electric-grid: AuroraGPT: Foundation Models for Science talks/auroragpt-siam25: AuroraGPT talks/auroragpt/alcf-hpc-workshop-2024/auroragpt-alcf-hands-on-hpc-workshop-2024: AuroraGPT: ANL's General Purpose Scientific LLM talks/demo-slides: AuroraGPT: Training Foundation Models on Supercomputers talks/hpc-user-forum/auroragpt: AuroraGPT talks/incite-hackathon-2025: ALCF Incite Hackathon 2025 talks/incite-hackathon-2025/auroragpt: LLMs on Aurora: Overview talks/incite-hackathon-2025/ezpz: LLMs on Aurora: Hands-On talks/llms-at-scale: Training LLMs at Scale talks/llms-on-polaris: Training LLMs on Polaris talks/openskai25: Open SkAI2025 talks/openskai25/ai4science: Scientific AI at Scale: AuroraGPT talks/openskai25/training: Scientific AI at Scale: Distributed Training webtui: Style webtui/components/accordion: Accordion webtui/components/badge: Badge webtui/components/button: Button webtui/components/checkbox: Checkbox webtui/components/dialog: Dialog webtui/components/input: Input webtui/components/popover: Popover webtui/components/pre: Pre webtui/components/progress: Progress webtui/components/radio: Radio webtui/components/range: Range webtui/components/separator: Separator webtui/components/spinner: Spinner webtui/components/switch: Switch webtui/components/table: Table webtui/components/textarea: Textarea webtui/components/tooltip: Popover webtui/components/typography: Typography webtui/components/view: View webtui/contributing/contributing: Contributing webtui/contributing/contributing: ## Local Development webtui/contributing/contributing: ## Issues webtui/contributing/contributing: ## Pull Requests webtui/contributing/style-guide: Style Guide webtui/contributing/style-guide: ## CSS Units webtui/contributing/style-guide: ## Selectors webtui/contributing/style-guide: ## Documentation webtui/installation/astro: Astro webtui/installation/astro: ## Scoping webtui/installation/astro: ### Frontmatter Imports webtui/installation/astro: ### ‹style› tag webtui/installation/astro: ### Full Library Import webtui/installation/nextjs: Next.js webtui/installation/vite: Vite webtui/plugins/plugin-dev: Developing Plugins webtui/plugins/plugin-dev: ### Style Layers webtui/plugins/plugin-nf: Nerd Font Plugin webtui/plugins/theme-catppuccin: Catppuccin Theme webtui/plugins/theme-custom: Custom Theme webtui/plugins/theme-everforest: Everforest Theme webtui/plugins/theme-gruvbox: Gruvbox Theme webtui/plugins/theme-nord: Nord Theme webtui/plugins/theme-vitesse: Vitesse Theme webtui/start/ascii-boxes: ASCII Boxes webtui/start/changelog: Changelog webtui/start/installation: Installation webtui/start/installation: ## Installation webtui/start/installation: ## Using CSS webtui/start/installation: ## Using ESM webtui/start/installation: ## Using a CDN webtui/start/installation: ## Full Library Import webtui/start/installation: ### CSS webtui/start/installation: ### ESM webtui/start/installation: ### CDN webtui/start/intro: Introduction webtui/start/intro: ## Features webtui/start/plugins: Plugins webtui/start/plugins: ## Official Plugins webtui/start/plugins: ### Themes webtui/start/plugins: ## Community Plugins webtui/start/theming: Theming webtui/start/theming: ## CSS Variables webtui/start/theming: ### Font Styles webtui/start/theming: ### Colors webtui/start/theming: ### Light & Dark webtui/start/theming: ## Theme Plugins webtui/start/theming: ### Using Multiple Theme Accents webtui/start/tuis-vs-guis: TUIs vs GUIs webtui/start/tuis-vs-guis: ## Monospace Fonts webtui/start/tuis-vs-guis: ## Character Cells
 Theme Current: Light j/k or ↑/↓ + Enter

Local AI Apps on ALCF: Argo, Inference Endpoints, and One Gateway

Claude Code, OpenCode, Codex, and Hermes (CLI + desktop) through one local llm-rosetta gateway that reaches ALCF Argo (via argo-shim) and ALCF Inference Endpoints (Sophia/Metis/Minerva). Claude, GPT, and open models from any client, kept running on an always-on host, with automatic model fallback and an iMessage front end.

ALCF has two model gateways. Argo serves Claude and GPT-family frontier models from an internal endpoint reachable only from inside the lab network. The newer ALCF Inference Endpoints service serves open models on Sophia and Metis through a public OpenAI-compatible API behind Globus auth.

I wanted both from local clients on my Mac (Claude Code, OpenCode, Codex, and Hermes, both its CLI and its hermes desktop app), without exposing ports, juggling client-specific API keys, or paying a third-party provider.

The stack I ended up with is one local llm-rosetta gateway that fans out to argo-shim (for Argo) or ALCF’s native inference API (for Sophia/Metis), and presents one localhost API to every client. Everything runs on 127.0.0.1, bills through ALCF, and survives token rotation. The minimal version is two steps (Quick start below); everything after that is optional.

TL;DR — what this builds, and where each piece lives

argo-shim turns the SSH-gated Argo service into a localhost API. llm-rosetta sits in front of it, translates OpenAI ⇆ Anthropic, and exposes ALCF’s native Sophia/Metis/Minerva inference endpoints. Point Claude Code, OpenCode, Codex, Hermes, or anything OpenAI-compatible at one local port and route by model name.

The two-step minimum

  • Start argo-shim — one uvx command; it opens the SSH tunnel and writes the token.
  • Claude Code — no config needed, the shim writes ~/.claude/settings.json for you.

Everything else is optional

When it breaks

Quick start

This assumes macOS with uv installed and working SSH access to the ALCF jump host.

Two steps: start argo-shim, and Claude Code picks it up automatically. If Claude Code with Argo is all you want, stop after this section.

Start argo-shim

One line to try it, one to keep it. argo-shim creates the SSH tunnel and writes a token to ~/.claude/settings.json. A recent build already includes the bearer-auth + user-injection support (see Contributing back).

# Kick the tires without installing anything.
# Drop CELS_USERNAME if your ALCF username matches your local login name.
CELS_USERNAME=<your-alcf-username> uvx argo-shim

# Keep it: this is a long-running daemon you restart often, and the shell
# helpers further down invoke it as a bare `argo-shim`, so put it on PATH.
uv tool install argo-shim

argo-shim listens on 127.0.0.1:25940, derives a per-user port for the SSH tunnel, and prints the token. Each restart rotates the token (important later).

Sanity check. The token lives in ~/.claude/settings.json, so the helper digs it out itself. It builds the request body with jq (brew install jq) rather than string interpolation, so a prompt containing quotes or an apostrophe still produces valid JSON:

# argo-ask <model> <prompt...>  — Claude via Anthropic Messages.
argo-ask() {
  local model="$1"; shift
  local token
  token=$(python3 -c "import json;print(json.load(open('$HOME/.claude/settings.json'))['apiKeyHelper'].split()[-1])") || return 1
  jq -nc --arg m "$model" --arg p "$*" \
    '{model:$m, max_tokens:64, messages:[{role:"user",content:$p}]}' \
  | curl -sS http://127.0.0.1:25940/argoapi/v1/messages \
      -H "x-api-key: $token" -H "anthropic-version: 2023-06-01" \
      -H "content-type: application/json" --data @-
}

argo-ask "Claude Opus 5" say hi          # should return JSON
argo-ask "Claude Opus 5" "what's 2+2?"   # quotes are safe

Claude Code (Anthropic, native)

Claude Code needs no extra config; argo-shim writes the base URL and token into ~/.claude/settings.json for you:

{
    "apiKeyHelper": "echo <ROTATING_TOKEN>",
    "env": {
        "ANTHROPIC_BASE_URL": "http://127.0.0.1:25940/argoapi"
    }
}

claude now routes through Argo.

Everything below is optional: a shared gateway so OpenAI-format tools (OpenCode, Codex, Hermes) reach the same models, plus the ALCF Inference Endpoints. Add only the pieces you need.

The full picture

Argo speaks the Anthropic Messages API for Claude models (/v1/messages) and the OpenAI Chat Completions API for GPT/Gemini (/v1/chat/completions), authenticated with an x-api-key header. The ALCF Inference Endpoints service speaks OpenAI-compatible chat/completions directly, but needs a Globus access token.

Four problems follow from that:

  1. Argo (apps.inside.anl.gov) is only reachable through an SSH jump host behind MFA.
  2. Different clients want different things: Claude Code speaks Anthropic; OpenCode and Hermes speak OpenAI; some send Authorization: Bearer, some send x-api-key.
  3. Argo’s OpenAI endpoint has two undocumented quirks (covered in Gotchas) that make it return HTTP 500 unless you massage the request.
  4. The inference service has its own auth lifecycle: Globus access tokens expire and must be present when the rosetta gateway starts.

One local front door (llm-rosetta) and one shim for the SSH-gated Argo path cover all four:

      ┌─────────────┐ ┌─────────────┐         ┌─────────────┐ ┌─────────────┐
      │ Claude Code │ │    vtcode   │         │   OpenCode  │ │    Hermes   │
      └─────────────┘ └─────────────┘         └─────────────┘ └─────────────┘
            │ Anthropic     │ Anthropic             │ OpenAI        │ OpenAI
            └───────┬───────┘                       └───────┬───────┘
                    │                                       │
 ┌╴╴╴╴╴╴╴╴╴╴╴╴╴╴╴╴╴╴├  Anthropic-native clients             │
 ┊                  │  can skip rosetta entirely            │
 ┊                  └───────────────────┬───────────────────┘
 ┊                                      │
 ┊                                      ▼
 ┊                       ┌─────────────────────────────┐
 ┊                       │  llm-rosetta gateway :8765  │
 ┊                       │      OpenAI ⇆ Anthropic     │
 ┊                       └───────┬─────────────┬───────┘
 ┊                   Argo        │             │ ALCF Inference
 ┊               Claude, GPT/o   │             │ Sophia, Metis
 ┊                               ▼             ▼
 ┊        ┌──────────────────────────┐  ┌────────────────────────────────┐
 └╴╴╴╴╴╴╴▶│     argo-shim :25940     │  │   inference-api.alcf.anl.gov   │
          │     (auth + fixups)      │  │  /resource_server/{cluster}/…  │
          └────────────┴─────────────┘  └───────────────┴────────────────┘
                       │ SSH tunnel :25939              │
                       ▼                                ▼
         ALCF Argo (apps.inside.anl.gov)      Sophia / Metis endpoints
  • Claude Code, OpenCode, Codex, Hermes, and vtcode can all point at rosetta if you want one dashboard/request log. Codex is another OpenAI-format client, so it joins the OpenCode/Hermes group above (talking OpenAI to rosetta).
  • The Anthropic-native clients (Claude Code, vtcode) can skip rosetta and talk to argo-shim directly, since argo-shim already speaks Anthropic /v1/messages (the dotted path above). Routing them through rosetta buys only the unified request log.
  • Argo Claude models route through rosetta → argo-shim → Argo /messages.
  • Argo GPT/o-series route through rosetta → argo-shim → Argo /chat/completions.
  • ALCF inference models route through rosetta directly to inference-api.alcf.anl.gov (no SSH tunnel or argo-shim needed).

The pieces

argo-shim: tunnel + auth + fixups

argo-shim (by n-getty) is a single-file Python proxy that:

  1. Manages an SSH tunnel to apps.inside.anl.gov:443.
  2. Listens on 127.0.0.1:25940 and rewrites any path to /argoapi/... before forwarding upstream.
  3. Authenticates local clients with a random token it writes into ~/.claude/settings.json (so Claude Code picks it up automatically).

I made two additive patches to get the OpenAI path working and let OpenAI-format clients authenticate (covered in Gotchas), both since contributed upstream. Neither touches the Anthropic /v1/messages path that Claude Code uses.

llm-rosetta + the gateway

llm-rosetta (by Oaklight) is an LLM API translation layer with an optional HTTP gateway. It converts between OpenAI Chat Completions, Anthropic Messages, and Google GenAI formats through a central intermediate representation. The gateway routes by model name:

  • A request for Argo claude-* → translated OpenAI → Anthropic → posted to argo-shim’s /v1/messages.
  • A request for Argo gpt-* / o* → forwarded as OpenAI Chat Completions to argo-shim’s /argoapi/v1/chat/completions.
  • A request for alcf-sophia/* or alcf-metis/* → forwarded directly to the ALCF Inference Endpoints service with a Globus access token.

Hermes (CLI + desktop)

Hermes is a tool-calling agent. One install ships both a CLI (hermes) and a native Electron desktop app you launch with hermes desktop (alias hermes gui); both read the same ~/.hermes/config.yaml. Point its model at the rosetta gateway with provider: custom and every Argo model shows up in either the terminal or a native chat UI.

Adding the other clients

Optional add-ons to the Quick start above. Each is one config block pointed at the same local stack; add only the clients you use. OpenCode talks to argo-shim directly. Codex and Hermes are OpenAI-format, so they go through the llm-rosetta gateway (set up in step 2 below).

1. OpenCode (OpenAI-format, via a custom provider)

OpenCode reads ~/.config/opencode/opencode.json. Add an Anthropic-compatible custom provider pointed at the shim (OpenCode’s @ai-sdk/anthropic sends x-api-key, which the shim wants):

{
    "provider": {
        "argo": {
            "npm": "@ai-sdk/anthropic",
            "name": "Argo (via argo-shim)",
            "options": {
                "baseURL": "http://127.0.0.1:25940/argoapi/v1",
                "apiKey": "{env:ARGO_SHIM_TOKEN}",
                "headers": { "anthropic-version": "2023-06-01" }
            },
            "models": {
                "claudeopus5": { "name": "claude-opus-5" }
            }
        }
    }
}

Export the token so {env:ARGO_SHIM_TOKEN} resolves (see Token rotation for a helper that does this automatically):

export ARGO_SHIM_TOKEN=$(python3 -c "import json;print(json.load(open('$HOME/.claude/settings.json'))['apiKeyHelper'].split()[-1])")

Then prove it end-to-end rather than trusting the config. --standalone keeps this out of the background service, so it exercises the provider block as written:

$ opencode run --standalone -m argo/claude-opus-5 "Reply with exactly: rosetta ok"
rosetta ok

Once the gateway in step 2 is up, the same command reaches it through the llm-rosetta provider, which is the OpenAI-format path rather than the Anthropic one:

$ opencode run --standalone -m llm-rosetta/anthropic/claude-opus-5 "Reply with exactly: gateway ok"
gateway ok

Both were run against OpenCode v2.0.18. The model key in models is the id sent upstream, so claudeopus5 and claude-opus-5 both work here only because the shim normalizes the name — don’t count on that with other providers.

2. llm-rosetta gateway (for OpenAI-only clients)

Install the gateway. Same reasoning as argo-shim: it is a daemon you restart often, and later sections call llm-rosetta-gateway by bare name.

uv tool install "llm-rosetta[gateway]"

Don’t start it yet. It needs the config below first.

Create ~/.config/llm-rosetta-gateway/config.jsonc. The Argo providers point at the local shim; the ALCF inference providers point directly at the public OpenAI-compatible endpoint:

{
    "providers": {
        // Claude models: Anthropic Messages format.
        // base_url has NO /v1: the anthropic template appends /v1/messages.
        "argo": {
            "type": "anthropic",
            "api_key": "${ARGO_SHIM_TOKEN}",
            "base_url": "http://127.0.0.1:25940",
        },
        // GPT/o models: OpenAI Chat Completions.
        // base_url includes /argoapi/v1: the template appends /chat/completions.
        "argo-openai": {
            "type": "openai_chat",
            "api_key": "${ARGO_SHIM_TOKEN}",
            "base_url": "http://127.0.0.1:25940/argoapi/v1",
        },
        // ALCF Inference Endpoints: already OpenAI-compatible.
        // ${ALCF_INFERENCE_TOKEN} is substituted when the gateway starts.
        "alcf-sophia": {
            "type": "openai_chat",
            "api_key": "${ALCF_INFERENCE_TOKEN}",
            "base_url": "https://inference-api.alcf.anl.gov/resource_server/sophia/vllm/v1",
        },
        "alcf-metis": {
            "type": "openai_chat",
            "api_key": "${ALCF_INFERENCE_TOKEN}",
            "base_url": "https://inference-api.alcf.anl.gov/resource_server/metis/api/v1",
        },
    },
    "models": {
        // Register each model in BOTH forms clients send: bare ("claude-opus-5")
        // and vendor-prefixed ("anthropic/claude-opus-5"). Declare `capabilities`
        // explicitly: the default is ["text"], which makes rosetta strip images.
        "claude-opus-5": {
            "provider": "argo",
            "upstream_model": "Claude Opus 5",
            "capabilities": ["text", "vision", "tools", "reasoning"],
        },
        "anthropic/claude-opus-5": {
            "provider": "argo",
            "upstream_model": "Claude Opus 5",
            "capabilities": ["text", "vision", "tools", "reasoning"],
        },
        "gpt-5.6-sol": {
            "provider": "argo-openai",
            "upstream_model": "GPT-5.6 Sol",
            "capabilities": ["text", "vision", "tools", "reasoning"],
        },
        "openai/gpt-5.6-sol": {
            "provider": "argo-openai",
            "upstream_model": "GPT-5.6 Sol",
            "capabilities": ["text", "vision", "tools", "reasoning"],
        },

        // ALCF inference models can be registered by exact upstream ID and/or a
        // cluster-prefixed alias to make filtering easier in dashboards.
        "alcf-sophia/openai/gpt-oss-120b": {
            "provider": "alcf-sophia",
            "upstream_model": "openai/gpt-oss-120b",
            "capabilities": ["text", "tools", "reasoning"],
        },
        "alcf-metis/gpt-oss-120b": {
            "provider": "alcf-metis",
            "upstream_model": "gpt-oss-120b",
            "capabilities": ["text"],
        },
        // … repeat for every model you want …
    },
    // No server.api_key: the gateway binds to 127.0.0.1 only, so no auth needed.
    "server": { "host": "127.0.0.1", "port": 8765 },
}

If you only use Argo models, start it directly:

llm-rosetta-gateway --no-banner   # listens on 127.0.0.1:8765

If you also want ALCF Inference Endpoints, authenticate once with Globus and export a short-lived access token before starting/restarting the gateway:

# First-time / monthly-ish auth; opens a Globus browser flow.
uvx --from alcf-ai alcf-ai auth login

# Refresh/export the access token for this shell, then restart rosetta so
# ${ALCF_INFERENCE_TOKEN} gets substituted into the config.
alcf-inference-token
rosetta-gateway restart

My day-to-day post-shim-restart ritual is one command:

argo-rosetta-sync --with-inference

Test the translation: OpenAI request in, Claude answer out. Same shape as argo-ask above, but pointed at the gateway and speaking OpenAI’s format, so one helper covers every model rosetta knows about:

# rosetta-ask <model> <prompt...>
rosetta-ask() {
  local model="$1"; shift
  jq -nc --arg m "$model" --arg p "$*" \
    '{model:$m, max_tokens:64, messages:[{role:"user",content:$p}]}' \
  | curl -sS http://127.0.0.1:8765/v1/chat/completions \
      -H "Content-Type: application/json" --data @-
}

# An Argo model...
rosetta-ask claude-opus-5 say hi

# ...and an ALCF inference model, if ALCF_INFERENCE_TOKEN was exported before
# the gateway started. A stale token shows up here as
# {"error":{"code":"unauthorized"}} — rerun alcf-inference-token, restart.
rosetta-ask alcf-sophia/openai/gpt-oss-120b say hi

The gateway ships a web admin panel at http://127.0.0.1:8765/admin/ with live metrics and request logs.

Heads up: the admin panel’s Fetch from Provider button re-discovers Argo’s full model list and writes entries that don’t work (some Gemini variants still 500 upstream; spaced display-names get mapped to the wrong provider). Manage the model list from the config file instead. I keep a regen script for exactly this.

3. Codex (OpenAI-format, via a custom provider)

Codex reads ~/.codex/config.toml. It’s OpenAI-format, so like OpenCode and Hermes it points at the rosetta gateway rather than the shim. Define a custom provider and select it as the default:

model = "gpt-5.6-sol"
model_provider = "rosetta"

[model_providers.rosetta]
name = "ALCF via llm-rosetta"
base_url = "http://127.0.0.1:8765/v1"
# Either wire API works: the gateway serves both /v1/chat/completions and
# /v1/responses. "responses" is Codex's default and what I actually run.
wire_api = "responses"
# No env_key: the gateway is loopback-only and runs no-auth, so there's no key
# to supply (same reason Codex's built-in Ollama provider omits it). Codex is
# happy without one against a local endpoint.

Correction (2026-09-28). An earlier version of this post set wire_api = "chat" and claimed the gateway doesn’t serve /v1/responses. That was wrong — I never tested it. It serves both:

$ curl -s -X POST http://127.0.0.1:8765/v1/responses \
    -H 'Content-Type: application/json' \
    -d '{"model":"gpt-5.6-sol","input":"Reply with exactly: responses ok"}' \
  | python3 -c 'import json,sys; d=json.load(sys.stdin); print([c["text"] for o in d["output"] for c in o.get("content",[]) if c.get("type")=="output_text"][0])'
responses ok

The 400 I originally read as “no such route” was upstream rejecting my request body: /v1/responses wants input, not messages. A 400 means the route exists and disliked what you sent; a missing route gives 404.

A few Codex-specific notes:

  • The provider id (rosetta here) can be anything except the reserved built-ins openai, ollama, and lmstudio.

  • Any rosetta model works as the top-level model: swap gpt-5.6-sol for claude-opus-5, alcf-sophia/openai/gpt-oss-120b, etc. (use the exact ids you registered in the gateway config above).

  • To switch models per session without editing the config, pass them on the command line:

    codex --model claude-opus-5 --model-provider rosetta

Check that Codex reaches the gateway. This hits the same /v1/chat/completions the earlier curl did, just through Codex’s config:

$ codex exec --skip-git-repo-check --model gpt-5.6-sol "Reply with exactly: codex responses ok" < /dev/null
codex
codex responses ok

tokens used
19,817

4. Hermes (CLI + hermes desktop)

Hermes installs a CLI (hermes) and a native Electron app together, and both read the same ~/.hermes/config.yaml. Register the gateway under providers: and point the main model at it by name:

model:
    base_url: http://127.0.0.1:8765/v1
    default: gpt-5.6-sol
    provider: llm-rosetta
    api_mode: chat_completions
providers:
    llm-rosetta:
        name: llm-rosetta
        base_url: http://127.0.0.1:8765/v1
        model: gpt-5.6-sol
        discover_models: true

That is what I run, minus a 130-entry models: map the hermes model wizard wrote for itself. You don’t need it: discover_models: true pulls the catalog from the gateway’s /v1/models at runtime, and a models: dict is treated as per-model metadata rather than an allowlist, so it doesn’t narrow anything either way. Pin a catalog with discover_models: false if you want the opposite.

llm-rosetta isn’t a name Hermes ships with — asking its resolver directly gives Unknown provider 'llm-rosetta'. Any key under providers: becomes a usable provider id that resolves internally to the custom backend, tagged with where it came from:

>>> from hermes_cli.runtime_provider import resolve_runtime_provider
>>> resolve_runtime_provider(requested="llm-rosetta")
{'api_key': 'no-key-required',
 'api_mode': 'chat_completions',
 'base_url': 'http://127.0.0.1:8765/v1',
 'provider': 'custom',
 'requested_provider': 'llm-rosetta',
 'source': 'custom_provider:llm-rosetta'}

api_key answers the other question: the gateway runs no-auth (loopback only), and Hermes supplies the no-key-required placeholder itself rather than making you invent one. Use any id you registered in the gateway config for default (claude-opus-5, alcf-sophia/openai/gpt-oss-120b, …).

hermes status confirms the name landed:

$ hermes status
◆ Environment
  Project:      /Users/sam/.hermes/hermes-agent
  Python:       3.14.7
  .env file:    ✗ not found
  Model:        gpt-5.6-sol
  Provider:     llm-rosetta

The shorter anonymous form also works, if you’d rather not name a provider:

model:
    base_url: http://127.0.0.1:8765/v1
    default: gpt-5.6-sol
    provider: custom
    api_mode: chat_completions

Five lines, no providers: block. The tradeoff is that custom is a single anonymous slot, and the fallback chain below refers to llm-rosetta and argo-shim by name. Those names only exist because they are registered — a fallback_providers entry naming an unregistered provider dies with Unknown provider 'llm-rosetta', and giving the tier an inline base_url does not rescue it.

For an interactive setup, hermes model walks you through adding a custom endpoint and writes the same block. Then confirm the CLI reaches the gateway:

hermes chat "say hi"

That run used a throwaway HOME whose config.yaml was the eleven lines above and nothing else, with no credentials anywhere in it. That is the point: if it needed something else from my home directory, it would have failed rather than answering in four seconds. The anonymous custom form answers the same way.

Launch the desktop app from the same install (it picks up the same config):

hermes desktop     # alias: hermes gui

Start a New Chat and you’re talking to Argo (or any registered model) through the local gateway. 🎉

Always-on: the stationary host

Everything above assumes the stack runs on the machine you’re typing at. That breaks the moment you travel: argo-shim’s SSH tunnel dies with the network, and re-authenticating through MFA on hotel wifi is the friction this setup was supposed to remove.

The fix is to stop treating the laptop as the host. I run the whole stack on a stationary Mac that never sleeps and attach to it from the laptop with herdr, a terminal workspace manager with a remote-attach mode:

        laptop (travels)                stationary host (always on)
   ┌──────────────────────────┐      ┌────────────────────────────────┐
   │  herdr --remote  ────────┼─ssh─▶│  herdr server                  │
   │  (terminal session)      │      │    ├─ hermes gateway           │
   │                          │      │    ├─ llm-rosetta      :8765   │
   │  ...or local clients     │      │    ├─ argo-shim        :25940  │
   │    :18765 ───────────────┼─ssh─▶│    └─ SSH tunnel       :25939  │
   │    :25941 ───────────────┼─ssh─▶│                                │
   └──────────────────────────┘      └────────────────────────────────┘
                                                    │
                                          ALCF Argo + Inference

The laptop holds no tokens and runs no daemons. It opens a terminal session on the host and detaches when the lid closes; the agent keeps working. Reconnecting from a different network is one command, no re-auth:

herdr --remote <host> --remote-keybindings server

Two things make this stick. First, the host must actually never sleep. Verify rather than assume:

Or: keep working locally and forward the ports

Attaching to a remote terminal session means the agent runs on the host. That is the right answer when the work has to survive a closed lid, and the wrong answer when you want the laptop’s own editor, clipboard and windows. The alternative is to leave the clients on the laptop and move only the services: forward the host’s two loopback ports and point every client at the forwarded copies.

ssh -N -L 127.0.0.1:18765:127.0.0.1:8765  myhost   # llm-rosetta
ssh -N -L 127.0.0.1:25941:127.0.0.1:25940 myhost   # argo-shim

The local ports are deliberately different. 18765 rather than 8765 lets the forward coexist with a still-running local gateway, so the cutover is reversible one client at a time instead of a flag day.

The laptop then runs no gateway, no shim, and no second MFA-gated tunnel of its own. Two machines each holding their own tunnel to the same login node is twice the exposure for no benefit, and login nodes are not always gracious about it.

To be precise, since this is the sort of claim that quietly stops being true: what the laptop sheds is the shim’s tunnel, the long-lived :25939 hop that argo-shim opens and keeps open. Ordinary SSH to the same hosts carries on: an interactive login, a ControlMaster socket, an ssh -fN for some other forward. Those are short-lived or cheap and are not what this avoids. My own laptop while writing this held exactly that mix: no argo-shim and no :25939, but a control master to the CELS login node and a backgrounded forward to Aurora, both days old.

Wrap each forward in a script that retries and falls back to a second path:

for host in myhost myhost-alt; do
  ssh -N -o ExitOnForwardFailure=yes -o ServerAliveInterval=30 \
      -L "127.0.0.1:$LOCAL:127.0.0.1:$REMOTE" "$host"
done

ExitOnForwardFailure matters. Without it ssh connects while silently failing to bind the local port, and you get a forward that looks alive and answers nothing. The second host is whatever path survives when the first does not; in my case a tunnel that works on networks which block my VPN’s control plane.

Every argo-shim mints its own key

argo-shim generates a fresh gateway key per process, so the shim on the host and the shim on the laptop never share one. The failure is invisible until you hit it.

Point a client at the forwarded shim while it still presents the laptop’s key and every request returns 401 Invalid API Key. Change only the key and you get the mirror image. The URL and the key have to move in the same edit.

The fix is to stop copying the key at all. Fetch the host’s live key at use time:

ssh myhost 'cat ~/path/to/shim-key'   # cache it briefly; it rotates on restart

One more wrinkle for Claude Code specifically: its apiKeyHelper runs on every request, so an SSH round trip there puts the network in the auth path of your own session. I did exactly that, with the base URL still pointing at the local shim, and locked myself out of the session I was using to make the change. Edit that file from a different session, change the URL and the key together, and keep a one-line rollback script to hand.

pmset -g | grep -E 'SleepDisabled|^ sleep'
# SleepDisabled  1
# sleep          0

Second, give the SSH client room to ride out a stall. Restarting any service on the host briefly blocks the session, and an aggressive keepalive budget tears down a working connection:

Host myhost
  ServerAliveInterval 30
  ServerAliveCountMax 8   # 8 x interval before giving up
  TCPKeepAlive yes
The Host * precedence trap

SSH takes the first value it sees for each keyword, not the most specific one. If Host * appears at the top of your config and sets ServerAliveInterval, a per-host block further down cannot override it. The per-host line is silently dead.

Don’t verify by reading the file. Ask ssh what it resolved:

ssh -G myhost | grep -E 'serveralive|controlpersist'

The same applies to ControlPersist: a Host * value of 600 will beat a per-host 4h that appears below it.

Model fallback: Argo first, ALCF as backstop

Once both upstreams work, you can prefer one and let the other cover for it. Argo serves the frontier models (Claude, the GPT/o-series); the ALCF Inference Endpoints serve open models on lab hardware. Argo is what I want for day-to-day work, so it goes first, and ALCF becomes the backstop that keeps a session alive when Argo is unreachable.

Argo lives behind an SSH tunnel and an MFA-gated jump host, so it has strictly more ways to fail than the public inference endpoints: the tunnel drops, the token rotates, the shim restarts. ALCF Inference needs only a Globus token that refreshes non-interactively. The chain is ordered by reliability, not preference: most-wanted first, most-available last.

In ~/.hermes/config.yaml:

model:
    base_url: 'http://127.0.0.1:8765/v1'
    default: 'gpt-5.6-sol'
    provider: 'llm-rosetta'
    api_mode: chat_completions

providers:
    llm-rosetta: # as above
        name: llm-rosetta
        base_url: http://127.0.0.1:8765/v1
        model: gpt-5.6-sol
        discover_models: true
    argo-shim:
        name: argo-shim
        base_url: http://127.0.0.1:25940
        key_env: ARGO_SHIM_TOKEN
        api_mode: anthropic_messages
        context_length: 1000000
        model: claude-opus-5

fallback_providers:
    - provider: llm-rosetta
      model: anthropic/claude-opus-5
    - provider: argo-shim
      model: GPT-5.6 Sol
      base_url: http://127.0.0.1:25940/v1
      key_env: ARGO_SHIM_TOKEN
      api_mode: chat_completions
    - provider: argo-shim
      model: claude-opus-5
      base_url: http://127.0.0.1:25940
      key_env: ARGO_SHIM_TOKEN
      api_mode: anthropic_messages
    - provider: llm-rosetta
      model: alcf-minerva/inkling-bf16

argo-shim is registered here for the same reason llm-rosetta was: the tiers below name it, and a name that isn’t under providers: doesn’t resolve. Note that providers: appears once — this block extends the one from the previous section rather than repeating it. Two top-level providers: keys in one file is the quiet failure: yaml.safe_load keeps only the last and drops the other’s providers without a word, and the ruamel round-trip Hermes writes config with raises DuplicateKeyError the next time anything edits the file.

Hermes walks that list in order, so a turn degrades GPT → Claude → open model rather than failing.

Tiers 2 and 3 are the ones worth copying. They name argo-shim directly on :25940 instead of going through rosetta, which means they survive a failure of the gateway itself — not just of an upstream behind it. A chain that routes every tier through one process only covers upstream outages; the process is still a single point of failure. Skipping it costs four extra lines per tier.

Note that the two shim tiers have different base URLs, and that is not a typo. argo-shim serves /v1/chat/completions and /v1/messages, but the Anthropic SDK appends /v1/messages to whatever base it is handed. So the chat_completions tier needs the /v1 written out and the anthropic_messages tier must not have it. Hermes normalizes this automatically for its OpenCode-family providers and for nothing else, so a custom provider is on its own here.

The api_mode on tier 2 is load-bearing for the same reason. A tier that doesn’t declare one inherits the wire from its providers: block — and argo-shim is registered as anthropic_messages, so dropping that one line sends the chat_completions tier down the Messages path instead:

>>> from agent.chat_completion_helpers import _fallback_api_mode_hint
>>> tier = {'provider': 'argo-shim', 'model': 'GPT-5.6 Sol',
...         'base_url': 'http://127.0.0.1:25940/v1',
...         'api_mode': 'chat_completions'}
>>> _fallback_api_mode_hint(tier, tier['provider'], tier['base_url'])
(True, 'chat_completions')
>>> del tier['api_mode']
>>> _fallback_api_mode_hint(tier, tier['provider'], tier['base_url'])
(True, 'anthropic_messages')

Get it wrong in the chat_completions direction and the tier 404s — which is exactly what mine did, silently, for as long as it had been configured. The turn still completed, because the chain just kept walking, and that is the whole problem with a quiet fallback: a dead tier and a tier that never gets reached look identical from the outside. Worth an actual curl per tier:

curl -s -o /dev/null -w '%{http_code}\n' -X POST \
  http://127.0.0.1:25940/v1/chat/completions \
  -H "Authorization: Bearer $ARGO_SHIM_TOKEN" \
  -H 'Content-Type: application/json' \
  -d '{"model":"GPT-5.6 Sol","messages":[{"role":"user","content":"hi"}],"max_tokens":5}'

You can see the chain fire in ~/.hermes/logs/agent.log:

Fallback activated: gpt-5.6-sol → anthropic/claude-opus-5 (llm-rosetta)
Fallback activated: anthropic/claude-opus-5 → GPT-5.6 Sol (argo-shim)
Fallback activated: GPT-5.6 Sol → claude-opus-5 (argo-shim)

Those three are one turn, nine seconds apart, and they are a better advertisement for the design than a clean run would be. GPT-5.6 Sol hit a 429 token rate limit upstream, Claude came back with an empty stream three times, and the chain walked off rosetta and onto the shim on its own.

That line is worth knowing by sight. Fallback is quiet by design — the session continues normally — so a model that’s silently unavailable looks like a model that’s working, just with different output. If you see that line constantly, your preferred upstream is down, not merely slow.

Check the direction. This is easy to get backwards: set default to whichever model you were last testing, and the chain now prefers the backstop. Confirm what served a turn by watching the rosetta log rather than trusting the config:

tail -f ~/.hermes/logs/rosetta-gateway.log | grep -o 'model=[^ ]*'

Note: Minerva is a third ALCF inference cluster alongside Sophia and Metis, registered the same way:

"alcf-minerva": {
  "type": "openai_chat",
  "api_key": "${ALCF_INFERENCE_TOKEN}",
  "base_url": "https://inference-api.alcf.anl.gov/resource_server/minerva/api/v1",
}

iMessage via Photon

Hermes ships a Photon plugin that bridges iMessage, so the agent running on the stationary host is reachable from a phone.

Enable the platform in ~/.hermes/config.yaml:

platforms:
    photon:
        enabled: true

The plugin runs a Node sidecar that the gateway talks to over gRPC on loopback. A healthy start says so in the gateway log:

[photon] connected — sidecar on 127.0.0.1:8789, streaming inbound over gRPC
✓ photon connected
Gateway running with 2 platform(s)

From there a text message is a prompt. Inbound and outbound both show up in the gateway log:

inbound message: platform=photon user=+1######### msg='...'
response ready: platform=photon time=228.8s api_calls=16 response=1536 chars

Note the time=228.8s with api_calls=16. A texted request runs a real multi-step tool-calling turn on the host, not a thin chat relay.

Checking that Photon is actually connected

The sidecar being alive is not the same as the platform being connected. Check all three:

# 1. platform registered
hermes status | grep -i photon
#   iMessage via Photon  ✓ configured (plugin)

# 2. sidecar listening
lsof -nP -iTCP:8789 -sTCP:LISTEN

# 3. gateway actually bridged to it
grep -E '\[photon\] connected|photon connected' ~/.hermes/logs/gateway.log | tail -2

The sidecar logs noisy ZodError lines for message shapes it doesn’t recognize (polls, some attachments). Those are non-fatal and don’t indicate a broken connection. Look for an explicit disconnect instead.

Gotchas

The non-obvious problems I hit. Two needed additive patches to argo-shim, leaving the Anthropic /v1/messages path Claude Code uses untouched; both are now upstream. The rest were configuration or upstream-app quirks.

1. OpenAI models need a user field

Argo’s /chat/completions returns a bare HTTP 500 for GPT/Gemini requests unless the body includes a user field set to a valid ALCF username. With it, the 500 turns into a real completion. The fix is to have the shim auto-inject it.

# In argo-shim, alongside the existing /messages handling:
if method == "POST" and body and "/chat/completions" in self.path:
    req = json.loads(body)
    if isinstance(req, dict) and not (req.get("user") or "").strip():
        req["user"] = ARGO_USER          # $ARGO_USER / $CELS_USERNAME / login
        body = json.dumps(req).encode()

This took a while to find: curl worked but the real clients all returned 500. The difference was that none of my curl tests happened to send a user field Argo accepted. Once I tried a valid ALCF username, GPT-4o, GPT-5, and the o-series all came alive.

2. OpenAI clients send Authorization: Bearer, not x-api-key

The shim originally accepted only x-api-key. rosetta’s openai_chat provider authenticates with Authorization: Bearer <key>. The shim accepts both now:

client_key = self.headers.get("x-api-key", "")
if not client_key:
    auth = self.headers.get("Authorization", "")
    if auth.lower().startswith("bearer "):
        client_key = auth[7:].strip()

3. Hermes’ custom provider and model-name prefixes

Two Hermes-specific quirks:

  • The custom provider doesn’t read CUSTOM_API_KEY for the actual request: only an inline api_key in config.yaml (which the desktop UI strips on save). Running the gateway no-auth sidesteps this entirely.
  • Hermes prepends a vendor prefix to model names (anthropic/claude-opus-5), then matches that exact string against the endpoint’s /v1/models list. So the gateway must register both the bare and vendor-prefixed forms of every model, hence the duplicated entries in the config above.

4. The capabilities default silently drops images

The rosetta admin panel showed every model with a single text capability badge, even though Claude and GPT-4o do vision. I assumed it was cosmetic. It isn’t.

A model entry with no capabilities list defaults to ["text"]. A model without "vision" has images stripped from the request before it’s forwarded upstream (enforce_vision() replaces them with text placeholders). Pasting an image into Hermes would have silently dropped it: no error, just a model that “couldn’t see” the image.

The fix is to declare real capabilities per model in the gateway config:

"claude-opus-5": {
  "provider": "argo",
  "upstream_model": "Claude Opus 5",
  "capabilities": ["text", "vision", "tools", "reasoning"]
}

My regen script sets these per model family (Claude 4.x/5 and the GPT-5/o-series get text + vision + tools + reasoning; GPT-4o/4.1 get text + vision + tools). After regenerating, images flow through to Argo instead of being quietly discarded.

5. Duplicated responses in the Hermes desktop app

After everything worked, hermes desktop started showing every reply twice: once as a slightly-reworded partial, then again in full. My first instinct was that the proxy chain was double-emitting.

It wasn’t. I tested the rosetta gateway directly — streaming, non-streaming, and a fresh non-stream call — and every response came back clean and singular. The shim and Argo were innocent. The duplication appeared only inside the Hermes desktop renderer, and only for chatty tool-calling turns (the offending turn had tool_turns=10).

The culprit was Hermes’ display.interim_assistant_messages: true, which renders the model’s interim commentary between tool calls and the final answer. With a verbose reasoning model (GPT-5.6 Sol) those two are near-identical, so the reply reads twice with slightly different wording. The fix is one line in ~/.hermes/config.yaml:

display:
    interim_assistant_messages: false

Update (2026-09-27): treat that as a workaround, not the setting you should be running. true is the upstream default, and Hermes has since reworked how interim commentary is projected across history and tool boundaries. My own config has been back on the default for months without me noticing doubled replies. Leave it alone unless you actually see the duplication, then flip it.

The meta-lesson, again: when something looks broken, bisect the layers. A two-minute curl against the gateway saved me from “fixing” a proxy that was working perfectly.

6. ALCF inference models appear even when auth is missing

The ALCF inference providers use "api_key": "${ALCF_INFERENCE_TOKEN}" in the rosetta config. That substitution happens when the gateway starts. With the env var missing, the models still appear in /v1/models and the admin UI, but calls 401 because the literal placeholder (or no useful bearer token) gets sent upstream.

The fix is to refresh the Globus token before restarting rosetta:

alcf-inference-token
rosetta-gateway restart

Or, after an argo-shim restart, do the whole thing:

argo-rosetta-sync --with-inference

7. Gemini is partially available

I first thought Gemini was completely broken, because I kept hitting 'NoneType' object is not iterable after fixing the user field. Careful re-testing showed three of four models work.

These are Argo Gemini models, not Google GenAI API models. Argo exposes them through its OpenAI-compatible /chat/completions endpoint, so in rosetta they use the openai_chat provider type, not the google provider type.

  • Gemini 2.5 Pro works through Argo’s OpenAI /chat/completions path.
  • Gemini 2.5 Flash works too.
  • Gemini 3.5 Flash also works.
  • Gemini 3.1 Flash Lite returned the upstream NoneType 500 whenever the request carried the usual OpenAI-style max_tokens field, and only worked if you omitted it or sent max_completion_tokens instead.

Update (2026-09-27): Argo fixed the last one. gemini-3.1-flash-lite now answers with max_tokens set and without it, so all four are registered as bare and google/... aliases. The request transform I was going to write turned out to be unnecessary — worth re-testing an upstream quirk before building around it.

Token rotation

Every argo-shim restart mints a new token. Shell helpers in ~/.config/zsh/functions.zsh keep everything in sync. These are the actual pieces, not magic commands hidden elsewhere.

# Re-export the live shim token into the env vars various clients read.
refresh-argo-token() {
  local settings="$HOME/.claude/settings.json"
  local token
  token=$(python3 -c "import json; print(json.load(open('$settings'))['apiKeyHelper'].split()[-1])") || return 1
  export ARGO_SHIM_TOKEN="$token"          # OpenCode, vtcode, direct shim clients
  export ANTHROPIC_API_KEY="$token"        # Anthropic-native clients
  export ANTHROPIC_BASE_URL="http://127.0.0.1:25940/v1"
  print "ARGO_SHIM_TOKEN refreshed (${token:0:8}…)"
}

# Point Claude Code at rosetta, not directly at argo-shim, so Claude Code
# traffic appears in the rosetta dashboard. argo-shim rewrites Claude settings
# back to :25940/argoapi on startup, so run this after starting argo-shim.
route-claude-through-rosetta() {
  python3 - <<'PY'
import json, pathlib
p = pathlib.Path.home() / ".claude" / "settings.json"
d = json.load(open(p))
env = d.setdefault("env", {})
env["ANTHROPIC_BASE_URL"] = "http://127.0.0.1:8765"
for key in ("NO_PROXY", "no_proxy"):
    vals = [x.strip() for x in env.get(key, "").split(",") if x.strip()]
    for val in ("localhost", "127.0.0.1"):
        if val not in vals:
            vals.append(val)
    env[key] = ",".join(vals)
p.write_text(json.dumps(d, indent=2) + "\n")
PY
}

# Export a fresh Globus access token for ALCF Inference Endpoints. First run may
# require an interactive browser auth flow (see below).
alcf-inference-token() {
  local token
  token=$(uvx --from alcf-ai alcf-ai auth get-access-token) || return 1
  export ALCF_INFERENCE_TOKEN="$token"
  print "ALCF_INFERENCE_TOKEN refreshed (${token:0:8}…)"
}

# Minimal rosetta restart helper. My full version also has start/stop/status,
# but this is the essential part: regenerate config if needed, then restart.
rosetta-gateway-restart() {
  pkill -f llm-rosetta-gateway 2>/dev/null || true
  sleep 1
  nohup llm-rosetta-gateway --no-banner >> "$HOME/.hermes/logs/rosetta-gateway.log" 2>&1 &
}

# One-shot recovery after starting/restarting argo-shim.
argo-rosetta-sync() {
  refresh-argo-token || return 1
  [[ "$1" == "--with-inference" ]] && alcf-inference-token
  # If you keep a regen script, run it here. Otherwise ensure config.jsonc uses
  # the current $ARGO_SHIM_TOKEN / $ALCF_INFERENCE_TOKEN before restarting.
  rosetta-gateway-restart || return 1
  route-claude-through-rosetta || return 1
}

The startup ritual after a reboot:

argo-shim                    # in a terminal: creates tunnel, rotates token
refresh-argo-token           # sync Argo token into env
rosetta-gateway-restart      # restart rosetta after config/token changes
route-claude-through-rosetta # optional: send Claude Code through rosetta too
# then run `hermes desktop` (or `hermes chat`)

For the ALCF Inference Endpoints in that same gateway:

alcf-inference-token      # export Globus access token
rosetta-gateway restart   # reload config with ${ALCF_INFERENCE_TOKEN}

Or one bundle after an argo-shim restart:

argo-rosetta-sync --with-inference

Letting launchd do it

Running that by hand gets old, and it fails in a specific way: the gateway substitutes ${ARGO_SHIM_TOKEN} and ${ALCF_INFERENCE_TOKEN} into config.jsonc once, at startup. A gateway that is up and listening can still be holding two dead tokens. Nothing crashes. Every model just 401s.

So there are two launchd jobs now. ai.llm-rosetta.gateway runs the server under KeepAlive and resolves both credentials on the way up, which is the whole point — the restart is the refresh mechanism:

# ~/.config/llm-rosetta-gateway/launch-gateway.sh (excerpt)
cd "$HOME" || exit 1

ARGO_SHIM_TOKEN=$("$PY" -c \
  "import json;print(json.load(open('$HOME/.claude/settings.json'))['apiKeyHelper'].split()[-1])")
export ARGO_SHIM_TOKEN

ALCF_INFERENCE_TOKEN=$(timeout 120 uv run --with globus-sdk --with openai \
  "$HELPER" get_access_token 2>/dev/null | tail -1)
export ALCF_INFERENCE_TOKEN

exec "$HOME/.local/bin/llm-rosetta-gateway" --no-banner --config "$CFG"

Two things in there are load-bearing. exec in the foreground, because launchd supervises the real process and nohup-ing it means KeepAlive watches a wrapper that already exited. And cd "$HOME" before the ALCF helper, because uv run resolves the enclosing project first — from a repo whose uv.lock your uv is too old to parse, you get a TOML error instead of a token.

The second job, ai.llm-rosetta.gateway.refresh, is the interesting one. It runs every six hours (StartInterval 21600) and probes both upstreams rather than restarting on a schedule:

argo=$(probe "anthropic/claude-haiku-4-5")
alcf=$(probe "alcf-metis/gpt-oss-120b")

stale=0
[[ "$argo" == 401 || "$argo" == 400 ]] && stale=1
[[ "$alcf" == 401 ]] && stale=1

(( stale )) && launchctl kickstart -k "gui/$(id -u)/ai.llm-rosetta.gateway"

A real inference call is the only honest test. /v1/models answers fine with expired credentials — the catalog is local config, not an upstream query. That is the same trap as gotcha #6, just automated.

Note 400 counts as stale for Argo. That one cost me a while: the shim 401s without draining the POST body, so the desynced keep-alive stream reports the next request as 400 Bad request syntax rather than an auth error. The auth failure shows up one request late, wearing a different status code.

The log is boring, which is the goal:

2026-09-27T05:09:49-0500 token-refresh: probe argo=200 alcf=200
2026-09-27T05:09:49-0500 token-refresh: both upstreams OK
2026-09-27T11:09:51-0500 token-refresh: probe argo=200 alcf=401
2026-09-27T11:09:51-0500 token-refresh: ALCF token stale
2026-09-27T11:09:51-0500 token-refresh: kickstarting ai.llm-rosetta.gateway to re-resolve tokens
2026-09-27T11:10:18-0500 token-refresh: post-restart argo=200 alcf=200

The Globus token expired, the probe caught it, and the gateway came back with a fresh one 27 seconds later. I found out by reading the log afterward.

argo-shim itself is still a manual foreground process, since the SSH tunnel it owns can need MFA and there is no point supervising something that may stop to ask a human a question. The shell functions above all still work, and are still what I reach for when I’m actively changing things — launchd handles the unattended case, not every case.

A regen script

Clicking the admin panel’s Fetch from Provider button re-pollutes the model list: it re-discovers Argo’s full set and writes broken entries (see the heads-up above). So I keep a script that regenerates a known-good config.jsonc from small Python maps of { alias: (upstream_model_id, capabilities) }. It registers bare + vendor-prefixed forms for Argo, adds alcf-sophia/* and alcf-metis/* aliases, attaches the right capabilities (so images aren’t stripped), pulls the live Argo token, and restarts the gateway:

rosetta-gateway regen-models   # rebuild model list + restart

Editing the model maps at the top of that script is the one place to add or remove models.

Contributing back

Two of the argo-shim fixes were general improvements rather than local workarounds, so they went back upstream as a PR:

  1. Accept Authorization: Bearer <token> in addition to x-api-key, so any OpenAI-format client can authenticate to the shim.
  2. Auto-inject the user field on /chat/completions, so OpenAI/Gemini models stop returning HTTP 500. The value resolves from $ARGO_USER, then $CELS_USERNAME, then the login user: no hardcoded usernames.

An automated review on the PR flagged a real edge case: a non-dict JSON body (a top-level array, say) would crash the handler, since the injection code assumed req.get(...) always worked. Guarding on isinstance(req, dict) and treating a blank user as missing fixed it. That’s the isinstance check in the snippet above.

 samf.sh / posts / 2026 / 06 / 27 · Top 1:1