Local AI Apps on ALCF: Argo, Inference Endpoints, and One Gateway
Claude Code, OpenCode, Codex, and Hermes (CLI + desktop) through one local llm-rosetta gateway that reaches ALCF Argo (via argo-shim) and ALCF Inference Endpoints (Sophia/Metis/Minerva). Claude, GPT, and open models from any client, kept running on an always-on host, with automatic model fallback and an iMessage front end.
ALCF has two model gateways. Argo serves Claude and GPT-family frontier models from an internal endpoint reachable only from inside the lab network. The newer ALCF Inference Endpoints service serves open models on Sophia and Metis through a public OpenAI-compatible API behind Globus auth.
I wanted both from local clients on my Mac (Claude Code,
OpenCode, Codex, and Hermes, both its CLI and its
hermes desktop app), without exposing ports, juggling client-specific API keys,
or paying a third-party provider.
The stack I ended up with is one local llm-rosetta gateway that fans out to
argo-shim (for Argo) or ALCF’s native inference API (for Sophia/Metis), and
presents one localhost API to every client. Everything runs on 127.0.0.1, bills
through ALCF, and survives token rotation. The minimal version is two steps
(Quick start below); everything after that is optional.
TL;DR — what this builds, and where each piece lives
argo-shim turns the SSH-gated Argo service into a localhost API.
llm-rosetta sits in front of it, translates OpenAI ⇆ Anthropic, and exposes
ALCF’s native Sophia/Metis/Minerva inference endpoints. Point Claude Code,
OpenCode, Codex, Hermes, or anything OpenAI-compatible at one local port and
route by model name.
The two-step minimum
- Start
argo-shim— oneuvxcommand; it opens the SSH tunnel and writes the token. - Claude Code — no config needed, the shim
writes
~/.claude/settings.jsonfor you.
Everything else is optional
- The full picture — architecture diagram of the whole stack.
llm-rosettagateway — theconfig.jsoncthat registers providers and models.- Clients: OpenCode · Codex · Hermes
- Always-on host — run the stack on one machine, attach to it from a laptop.
- Model fallback — prefer Argo Claude, fall back through GPT to an open ALCF model automatically.
- iMessage — talk to the agent from your phone.
When it breaks
- Gotchas — seven real ones, including
the
userfield 500, silently stripped images, and missing ALCF auth. - Token rotation — shell helpers for the post-restart ritual.
Quick start
This assumes macOS with uv installed and working SSH access to the
ALCF jump host.
Two steps: start argo-shim, and Claude Code picks it up automatically. If
Claude Code with Argo is all you want, stop after this section.
Start argo-shim
One line to try it, one to keep it. argo-shim creates the SSH tunnel and
writes a token to ~/.claude/settings.json. A recent build already includes
the bearer-auth + user-injection support (see Contributing
back).
# Kick the tires without installing anything.
# Drop CELS_USERNAME if your ALCF username matches your local login name.
CELS_USERNAME=<your-alcf-username> uvx argo-shim
# Keep it: this is a long-running daemon you restart often, and the shell
# helpers further down invoke it as a bare `argo-shim`, so put it on PATH.
uv tool install argo-shim
argo-shim listens on 127.0.0.1:25940, derives a per-user port for the SSH
tunnel, and prints the token. Each restart rotates the token (important
later).
Sanity check. The token lives in ~/.claude/settings.json, so the helper digs
it out itself. It builds the request body with jq (brew install jq)
rather than string interpolation, so a prompt containing quotes or an apostrophe
still produces valid JSON:
# argo-ask <model> <prompt...> — Claude via Anthropic Messages.
argo-ask() {
local model="$1"; shift
local token
token=$(python3 -c "import json;print(json.load(open('$HOME/.claude/settings.json'))['apiKeyHelper'].split()[-1])") || return 1
jq -nc --arg m "$model" --arg p "$*" \
'{model:$m, max_tokens:64, messages:[{role:"user",content:$p}]}' \
| curl -sS http://127.0.0.1:25940/argoapi/v1/messages \
-H "x-api-key: $token" -H "anthropic-version: 2023-06-01" \
-H "content-type: application/json" --data @-
}
argo-ask "Claude Opus 5" say hi # should return JSON
argo-ask "Claude Opus 5" "what's 2+2?" # quotes are safe
Claude Code (Anthropic, native)
Claude Code needs no extra config; argo-shim writes the base URL and token
into ~/.claude/settings.json for you:
{
"apiKeyHelper": "echo <ROTATING_TOKEN>",
"env": {
"ANTHROPIC_BASE_URL": "http://127.0.0.1:25940/argoapi"
}
}
claude now routes through Argo.
Everything below is optional: a shared gateway so OpenAI-format tools (OpenCode, Codex, Hermes) reach the same models, plus the ALCF Inference Endpoints. Add only the pieces you need.
The full picture
Argo speaks the Anthropic Messages API for Claude models
(/v1/messages) and the OpenAI Chat Completions API for GPT/Gemini
(/v1/chat/completions), authenticated with an x-api-key header. The ALCF
Inference Endpoints service speaks OpenAI-compatible chat/completions directly,
but needs a Globus access token.
Four problems follow from that:
- Argo (
apps.inside.anl.gov) is only reachable through an SSH jump host behind MFA. - Different clients want different things: Claude Code speaks Anthropic;
OpenCode and Hermes speak OpenAI; some send
Authorization: Bearer, some sendx-api-key. - Argo’s OpenAI endpoint has two undocumented quirks (covered in
Gotchas) that make it return
HTTP 500unless you massage the request. - The inference service has its own auth lifecycle: Globus access tokens expire and must be present when the rosetta gateway starts.
One local front door (llm-rosetta) and one shim for the SSH-gated Argo path
cover all four:
┌─────────────┐ ┌─────────────┐ ┌─────────────┐ ┌─────────────┐
│ Claude Code │ │ vtcode │ │ OpenCode │ │ Hermes │
└─────────────┘ └─────────────┘ └─────────────┘ └─────────────┘
│ Anthropic │ Anthropic │ OpenAI │ OpenAI
└───────┬───────┘ └───────┬───────┘
│ │
┌╴╴╴╴╴╴╴╴╴╴╴╴╴╴╴╴╴╴├ Anthropic-native clients │
┊ │ can skip rosetta entirely │
┊ └───────────────────┬───────────────────┘
┊ │
┊ ▼
┊ ┌─────────────────────────────┐
┊ │ llm-rosetta gateway :8765 │
┊ │ OpenAI ⇆ Anthropic │
┊ └───────┬─────────────┬───────┘
┊ Argo │ │ ALCF Inference
┊ Claude, GPT/o │ │ Sophia, Metis
┊ ▼ ▼
┊ ┌──────────────────────────┐ ┌────────────────────────────────┐
└╴╴╴╴╴╴╴▶│ argo-shim :25940 │ │ inference-api.alcf.anl.gov │
│ (auth + fixups) │ │ /resource_server/{cluster}/… │
└────────────┴─────────────┘ └───────────────┴────────────────┘
│ SSH tunnel :25939 │
▼ ▼
ALCF Argo (apps.inside.anl.gov) Sophia / Metis endpoints
- Claude Code, OpenCode, Codex, Hermes, and vtcode can all point at rosetta if you want one dashboard/request log. Codex is another OpenAI-format client, so it joins the OpenCode/Hermes group above (talking OpenAI to rosetta).
- The Anthropic-native clients (Claude Code, vtcode) can skip rosetta
and talk to
argo-shimdirectly, sinceargo-shimalready speaks Anthropic/v1/messages(the dotted path above). Routing them through rosetta buys only the unified request log. - Argo Claude models route through rosetta →
argo-shim→ Argo/messages. - Argo GPT/o-series route through rosetta →
argo-shim→ Argo/chat/completions. - ALCF inference models route through rosetta directly to
inference-api.alcf.anl.gov(no SSH tunnel orargo-shimneeded).
The pieces
argo-shim: tunnel + auth + fixups
argo-shim (by n-getty) is a single-file Python proxy
that:
- Manages an SSH tunnel to
apps.inside.anl.gov:443. - Listens on
127.0.0.1:25940and rewrites any path to/argoapi/...before forwarding upstream. - Authenticates local clients with a random token it writes into
~/.claude/settings.json(so Claude Code picks it up automatically).
I made two additive patches to get the OpenAI path working and let
OpenAI-format clients authenticate (covered in Gotchas), both since
contributed upstream. Neither touches the Anthropic
/v1/messages path that Claude Code uses.
llm-rosetta + the gateway
llm-rosetta (by Oaklight) is an LLM API translation
layer with an optional HTTP gateway. It converts between OpenAI Chat
Completions, Anthropic Messages, and Google GenAI formats through a central
intermediate representation. The gateway routes by model name:
- A request for Argo
claude-*→ translated OpenAI → Anthropic → posted toargo-shim’s/v1/messages. - A request for Argo
gpt-*/o*→ forwarded as OpenAI Chat Completions toargo-shim’s/argoapi/v1/chat/completions. - A request for
alcf-sophia/*oralcf-metis/*→ forwarded directly to the ALCF Inference Endpoints service with a Globus access token.
Hermes (CLI + desktop)
Hermes is a tool-calling agent. One install ships both a CLI
(hermes) and a native Electron desktop app you launch with hermes desktop
(alias hermes gui); both read the same ~/.hermes/config.yaml. Point its
model at the rosetta gateway with provider: custom and every Argo model shows
up in either the terminal or a native chat UI.
Adding the other clients
Optional add-ons to the Quick start above. Each is one config
block pointed at the same local stack; add only the clients you use. OpenCode
talks to argo-shim directly. Codex and Hermes are OpenAI-format, so they go
through the llm-rosetta gateway (set up in step 2 below).
1. OpenCode (OpenAI-format, via a custom provider)
OpenCode reads ~/.config/opencode/opencode.json. Add an Anthropic-compatible
custom provider pointed at the shim (OpenCode’s @ai-sdk/anthropic sends
x-api-key, which the shim wants):
{
"provider": {
"argo": {
"npm": "@ai-sdk/anthropic",
"name": "Argo (via argo-shim)",
"options": {
"baseURL": "http://127.0.0.1:25940/argoapi/v1",
"apiKey": "{env:ARGO_SHIM_TOKEN}",
"headers": { "anthropic-version": "2023-06-01" }
},
"models": {
"claudeopus5": { "name": "claude-opus-5" }
}
}
}
}
Export the token so {env:ARGO_SHIM_TOKEN} resolves (see
Token rotation for a helper that does this automatically):
export ARGO_SHIM_TOKEN=$(python3 -c "import json;print(json.load(open('$HOME/.claude/settings.json'))['apiKeyHelper'].split()[-1])")
Then prove it end-to-end rather than trusting the config. --standalone keeps
this out of the background service, so it exercises the provider block as
written:
$ opencode run --standalone -m argo/claude-opus-5 "Reply with exactly: rosetta ok"
rosetta ok
Once the gateway in step 2 is up, the same command reaches it through the
llm-rosetta provider, which is the OpenAI-format path rather than the
Anthropic one:
$ opencode run --standalone -m llm-rosetta/anthropic/claude-opus-5 "Reply with exactly: gateway ok"
gateway ok
Both were run against OpenCode v2.0.18. The model key in models is the id
sent upstream, so claudeopus5 and claude-opus-5 both work here only because
the shim normalizes the name — don’t count on that with other providers.
2. llm-rosetta gateway (for OpenAI-only clients)
Install the gateway. Same reasoning as argo-shim: it is a daemon you restart
often, and later sections call llm-rosetta-gateway by bare name.
uv tool install "llm-rosetta[gateway]"
Don’t start it yet. It needs the config below first.
Create ~/.config/llm-rosetta-gateway/config.jsonc. The Argo providers point
at the local shim; the ALCF inference providers point directly at the public
OpenAI-compatible endpoint:
{
"providers": {
// Claude models: Anthropic Messages format.
// base_url has NO /v1: the anthropic template appends /v1/messages.
"argo": {
"type": "anthropic",
"api_key": "${ARGO_SHIM_TOKEN}",
"base_url": "http://127.0.0.1:25940",
},
// GPT/o models: OpenAI Chat Completions.
// base_url includes /argoapi/v1: the template appends /chat/completions.
"argo-openai": {
"type": "openai_chat",
"api_key": "${ARGO_SHIM_TOKEN}",
"base_url": "http://127.0.0.1:25940/argoapi/v1",
},
// ALCF Inference Endpoints: already OpenAI-compatible.
// ${ALCF_INFERENCE_TOKEN} is substituted when the gateway starts.
"alcf-sophia": {
"type": "openai_chat",
"api_key": "${ALCF_INFERENCE_TOKEN}",
"base_url": "https://inference-api.alcf.anl.gov/resource_server/sophia/vllm/v1",
},
"alcf-metis": {
"type": "openai_chat",
"api_key": "${ALCF_INFERENCE_TOKEN}",
"base_url": "https://inference-api.alcf.anl.gov/resource_server/metis/api/v1",
},
},
"models": {
// Register each model in BOTH forms clients send: bare ("claude-opus-5")
// and vendor-prefixed ("anthropic/claude-opus-5"). Declare `capabilities`
// explicitly: the default is ["text"], which makes rosetta strip images.
"claude-opus-5": {
"provider": "argo",
"upstream_model": "Claude Opus 5",
"capabilities": ["text", "vision", "tools", "reasoning"],
},
"anthropic/claude-opus-5": {
"provider": "argo",
"upstream_model": "Claude Opus 5",
"capabilities": ["text", "vision", "tools", "reasoning"],
},
"gpt-5.6-sol": {
"provider": "argo-openai",
"upstream_model": "GPT-5.6 Sol",
"capabilities": ["text", "vision", "tools", "reasoning"],
},
"openai/gpt-5.6-sol": {
"provider": "argo-openai",
"upstream_model": "GPT-5.6 Sol",
"capabilities": ["text", "vision", "tools", "reasoning"],
},
// ALCF inference models can be registered by exact upstream ID and/or a
// cluster-prefixed alias to make filtering easier in dashboards.
"alcf-sophia/openai/gpt-oss-120b": {
"provider": "alcf-sophia",
"upstream_model": "openai/gpt-oss-120b",
"capabilities": ["text", "tools", "reasoning"],
},
"alcf-metis/gpt-oss-120b": {
"provider": "alcf-metis",
"upstream_model": "gpt-oss-120b",
"capabilities": ["text"],
},
// … repeat for every model you want …
},
// No server.api_key: the gateway binds to 127.0.0.1 only, so no auth needed.
"server": { "host": "127.0.0.1", "port": 8765 },
}
If you only use Argo models, start it directly:
llm-rosetta-gateway --no-banner # listens on 127.0.0.1:8765
If you also want ALCF Inference Endpoints, authenticate once with Globus and export a short-lived access token before starting/restarting the gateway:
# First-time / monthly-ish auth; opens a Globus browser flow.
uvx --from alcf-ai alcf-ai auth login
# Refresh/export the access token for this shell, then restart rosetta so
# ${ALCF_INFERENCE_TOKEN} gets substituted into the config.
alcf-inference-token
rosetta-gateway restart
My day-to-day post-shim-restart ritual is one command:
argo-rosetta-sync --with-inference
Test the translation: OpenAI request in, Claude answer out. Same shape as
argo-ask above, but pointed at the gateway and speaking OpenAI’s format, so
one helper covers every model rosetta knows about:
# rosetta-ask <model> <prompt...>
rosetta-ask() {
local model="$1"; shift
jq -nc --arg m "$model" --arg p "$*" \
'{model:$m, max_tokens:64, messages:[{role:"user",content:$p}]}' \
| curl -sS http://127.0.0.1:8765/v1/chat/completions \
-H "Content-Type: application/json" --data @-
}
# An Argo model...
rosetta-ask claude-opus-5 say hi
# ...and an ALCF inference model, if ALCF_INFERENCE_TOKEN was exported before
# the gateway started. A stale token shows up here as
# {"error":{"code":"unauthorized"}} — rerun alcf-inference-token, restart.
rosetta-ask alcf-sophia/openai/gpt-oss-120b say hi
The gateway ships a web admin panel at http://127.0.0.1:8765/admin/ with live
metrics and request logs.
Heads up: the admin panel’s Fetch from Provider button re-discovers Argo’s full model list and writes entries that don’t work (some Gemini variants still 500 upstream; spaced display-names get mapped to the wrong provider). Manage the model list from the config file instead. I keep a regen script for exactly this.
3. Codex (OpenAI-format, via a custom provider)
Codex reads ~/.codex/config.toml. It’s OpenAI-format, so like
OpenCode and Hermes it points at the rosetta gateway rather than the shim.
Define a custom provider and select it as the default:
model = "gpt-5.6-sol"
model_provider = "rosetta"
[model_providers.rosetta]
name = "ALCF via llm-rosetta"
base_url = "http://127.0.0.1:8765/v1"
# Either wire API works: the gateway serves both /v1/chat/completions and
# /v1/responses. "responses" is Codex's default and what I actually run.
wire_api = "responses"
# No env_key: the gateway is loopback-only and runs no-auth, so there's no key
# to supply (same reason Codex's built-in Ollama provider omits it). Codex is
# happy without one against a local endpoint.
Correction (2026-09-28). An earlier version of this post set
wire_api = "chat"and claimed the gateway doesn’t serve/v1/responses. That was wrong — I never tested it. It serves both:$ curl -s -X POST http://127.0.0.1:8765/v1/responses \ -H 'Content-Type: application/json' \ -d '{"model":"gpt-5.6-sol","input":"Reply with exactly: responses ok"}' \ | python3 -c 'import json,sys; d=json.load(sys.stdin); print([c["text"] for o in d["output"] for c in o.get("content",[]) if c.get("type")=="output_text"][0])' responses okThe
400I originally read as “no such route” was upstream rejecting my request body:/v1/responseswantsinput, notmessages. A400means the route exists and disliked what you sent; a missing route gives404.
A few Codex-specific notes:
-
The provider id (
rosettahere) can be anything except the reserved built-insopenai,ollama, andlmstudio. -
Any rosetta model works as the top-level
model: swapgpt-5.6-solforclaude-opus-5,alcf-sophia/openai/gpt-oss-120b, etc. (use the exact ids you registered in the gateway config above). -
To switch models per session without editing the config, pass them on the command line:
codex --model claude-opus-5 --model-provider rosetta
Check that Codex reaches the gateway. This hits the same
/v1/chat/completions the earlier curl did, just through Codex’s config:
$ codex exec --skip-git-repo-check --model gpt-5.6-sol "Reply with exactly: codex responses ok" < /dev/null
codex
codex responses ok
tokens used
19,817
4. Hermes (CLI + hermes desktop)
Hermes installs a CLI (hermes) and a native Electron app together, and both
read the same ~/.hermes/config.yaml. Register the gateway under providers:
and point the main model at it by name:
model:
base_url: http://127.0.0.1:8765/v1
default: gpt-5.6-sol
provider: llm-rosetta
api_mode: chat_completions
providers:
llm-rosetta:
name: llm-rosetta
base_url: http://127.0.0.1:8765/v1
model: gpt-5.6-sol
discover_models: true
That is what I run, minus a 130-entry models: map the hermes model wizard
wrote for itself. You don’t need it: discover_models: true pulls the catalog
from the gateway’s /v1/models at runtime, and a models: dict is treated
as per-model metadata rather than an allowlist, so it doesn’t narrow anything
either way. Pin a catalog with discover_models: false if you want the
opposite.
llm-rosetta isn’t a name Hermes ships with — asking its resolver directly
gives Unknown provider 'llm-rosetta'. Any key under providers: becomes a
usable provider id that resolves internally to the custom backend, tagged
with where it came from:
>>> from hermes_cli.runtime_provider import resolve_runtime_provider
>>> resolve_runtime_provider(requested="llm-rosetta")
{'api_key': 'no-key-required',
'api_mode': 'chat_completions',
'base_url': 'http://127.0.0.1:8765/v1',
'provider': 'custom',
'requested_provider': 'llm-rosetta',
'source': 'custom_provider:llm-rosetta'}
api_key answers the other question: the gateway runs no-auth (loopback
only), and Hermes supplies the no-key-required placeholder itself rather than
making you invent one. Use any id you registered in the gateway config for
default (claude-opus-5, alcf-sophia/openai/gpt-oss-120b, …).
hermes status confirms the name landed:
$ hermes status
◆ Environment
Project: /Users/sam/.hermes/hermes-agent
Python: 3.14.7
.env file: ✗ not found
Model: gpt-5.6-sol
Provider: llm-rosetta
The shorter anonymous form also works, if you’d rather not name a provider:
model:
base_url: http://127.0.0.1:8765/v1
default: gpt-5.6-sol
provider: custom
api_mode: chat_completions
Five lines, no providers: block. The tradeoff is that custom is a single
anonymous slot, and the fallback chain below refers to llm-rosetta and
argo-shim by name. Those names only exist because they are registered — a
fallback_providers entry naming an unregistered provider dies with
Unknown provider 'llm-rosetta', and giving the tier an inline base_url
does not rescue it.
For an interactive setup, hermes model walks you through adding a custom
endpoint and writes the same block. Then confirm the CLI reaches the gateway:
hermes chat "say hi"
That run used a throwaway HOME whose config.yaml was the eleven lines above
and nothing else, with no credentials anywhere in it. That is the point: if it
needed something else from my home directory, it would have failed rather than
answering in four seconds. The anonymous custom form answers the same way.
Launch the desktop app from the same install (it picks up the same config):
hermes desktop # alias: hermes gui
Start a New Chat and you’re talking to Argo (or any registered model) through the local gateway. 🎉
Always-on: the stationary host
Everything above assumes the stack runs on the machine you’re typing at. That
breaks the moment you travel: argo-shim’s SSH tunnel dies with the network,
and re-authenticating through MFA on hotel wifi is the friction this setup was
supposed to remove.
The fix is to stop treating the laptop as the host. I run the whole stack on a
stationary Mac that never sleeps and attach to it from the laptop with
herdr, a terminal workspace manager with a remote-attach mode:
laptop (travels) stationary host (always on)
┌──────────────────────────┐ ┌────────────────────────────────┐
│ herdr --remote ────────┼─ssh─▶│ herdr server │
│ (terminal session) │ │ ├─ hermes gateway │
│ │ │ ├─ llm-rosetta :8765 │
│ ...or local clients │ │ ├─ argo-shim :25940 │
│ :18765 ───────────────┼─ssh─▶│ └─ SSH tunnel :25939 │
│ :25941 ───────────────┼─ssh─▶│ │
└──────────────────────────┘ └────────────────────────────────┘
│
ALCF Argo + Inference
The laptop holds no tokens and runs no daemons. It opens a terminal session on the host and detaches when the lid closes; the agent keeps working. Reconnecting from a different network is one command, no re-auth:
herdr --remote <host> --remote-keybindings server
Two things make this stick. First, the host must actually never sleep. Verify rather than assume:
Or: keep working locally and forward the ports
Attaching to a remote terminal session means the agent runs on the host. That is the right answer when the work has to survive a closed lid, and the wrong answer when you want the laptop’s own editor, clipboard and windows. The alternative is to leave the clients on the laptop and move only the services: forward the host’s two loopback ports and point every client at the forwarded copies.
ssh -N -L 127.0.0.1:18765:127.0.0.1:8765 myhost # llm-rosetta
ssh -N -L 127.0.0.1:25941:127.0.0.1:25940 myhost # argo-shim
The local ports are deliberately different. 18765 rather than 8765 lets the
forward coexist with a still-running local gateway, so the cutover is reversible
one client at a time instead of a flag day.
The laptop then runs no gateway, no shim, and no second MFA-gated tunnel of its own. Two machines each holding their own tunnel to the same login node is twice the exposure for no benefit, and login nodes are not always gracious about it.
To be precise, since this is the sort of claim that quietly stops being true:
what the laptop sheds is the shim’s tunnel, the long-lived :25939 hop that
argo-shim opens and keeps open. Ordinary SSH to the same hosts carries on: an
interactive login, a ControlMaster socket, an ssh -fN for some other
forward. Those are short-lived or cheap and are not what this avoids. My own
laptop while writing this held exactly that mix: no argo-shim and no :25939,
but a control master to the CELS login node and a backgrounded forward to
Aurora, both days old.
Wrap each forward in a script that retries and falls back to a second path:
for host in myhost myhost-alt; do
ssh -N -o ExitOnForwardFailure=yes -o ServerAliveInterval=30 \
-L "127.0.0.1:$LOCAL:127.0.0.1:$REMOTE" "$host"
done
ExitOnForwardFailure matters. Without it ssh connects while silently failing
to bind the local port, and you get a forward that looks alive and answers
nothing. The second host is whatever path survives when the first does not; in
my case a tunnel that works on networks which block my VPN’s control plane.
Every argo-shim mints its own key
argo-shim generates a fresh gateway key per process, so the shim on the
host and the shim on the laptop never share one. The failure is invisible
until you hit it.
Point a client at the forwarded shim while it still presents the laptop’s key
and every request returns 401 Invalid API Key. Change only the key and you
get the mirror image. The URL and the key have to move in the same edit.
The fix is to stop copying the key at all. Fetch the host’s live key at use time:
ssh myhost 'cat ~/path/to/shim-key' # cache it briefly; it rotates on restart
One more wrinkle for Claude Code specifically: its apiKeyHelper runs on
every request, so an SSH round trip there puts the network in the auth path of
your own session. I did exactly that, with the base URL still pointing at the
local shim, and locked myself out of the session I was using to make the
change. Edit that file from a different session, change the URL and the key
together, and keep a one-line rollback script to hand.
pmset -g | grep -E 'SleepDisabled|^ sleep'
# SleepDisabled 1
# sleep 0
Second, give the SSH client room to ride out a stall. Restarting any service on the host briefly blocks the session, and an aggressive keepalive budget tears down a working connection:
Host myhost
ServerAliveInterval 30
ServerAliveCountMax 8 # 8 x interval before giving up
TCPKeepAlive yes
The Host * precedence trap
SSH takes the first value it sees for each keyword, not the most specific
one. If Host * appears at the top of your config and sets
ServerAliveInterval, a per-host block further down cannot override it.
The per-host line is silently dead.
Don’t verify by reading the file. Ask ssh what it resolved:
ssh -G myhost | grep -E 'serveralive|controlpersist'
The same applies to ControlPersist: a Host * value of 600 will beat a
per-host 4h that appears below it.
Model fallback: Argo first, ALCF as backstop
Once both upstreams work, you can prefer one and let the other cover for it. Argo serves the frontier models (Claude, the GPT/o-series); the ALCF Inference Endpoints serve open models on lab hardware. Argo is what I want for day-to-day work, so it goes first, and ALCF becomes the backstop that keeps a session alive when Argo is unreachable.
Argo lives behind an SSH tunnel and an MFA-gated jump host, so it has strictly more ways to fail than the public inference endpoints: the tunnel drops, the token rotates, the shim restarts. ALCF Inference needs only a Globus token that refreshes non-interactively. The chain is ordered by reliability, not preference: most-wanted first, most-available last.
In ~/.hermes/config.yaml:
model:
base_url: 'http://127.0.0.1:8765/v1'
default: 'gpt-5.6-sol'
provider: 'llm-rosetta'
api_mode: chat_completions
providers:
llm-rosetta: # as above
name: llm-rosetta
base_url: http://127.0.0.1:8765/v1
model: gpt-5.6-sol
discover_models: true
argo-shim:
name: argo-shim
base_url: http://127.0.0.1:25940
key_env: ARGO_SHIM_TOKEN
api_mode: anthropic_messages
context_length: 1000000
model: claude-opus-5
fallback_providers:
- provider: llm-rosetta
model: anthropic/claude-opus-5
- provider: argo-shim
model: GPT-5.6 Sol
base_url: http://127.0.0.1:25940/v1
key_env: ARGO_SHIM_TOKEN
api_mode: chat_completions
- provider: argo-shim
model: claude-opus-5
base_url: http://127.0.0.1:25940
key_env: ARGO_SHIM_TOKEN
api_mode: anthropic_messages
- provider: llm-rosetta
model: alcf-minerva/inkling-bf16
argo-shim is registered here for the same reason llm-rosetta was: the tiers
below name it, and a name that isn’t under providers: doesn’t resolve. Note
that providers: appears once — this block extends the one from the previous
section rather than repeating it. Two top-level providers: keys in one file
is the quiet failure: yaml.safe_load keeps only the last and drops the
other’s providers without a word, and the ruamel round-trip Hermes writes
config with raises DuplicateKeyError the next time anything edits the file.
Hermes walks that list in order, so a turn degrades GPT → Claude → open model rather than failing.
Tiers 2 and 3 are the ones worth copying. They name argo-shim directly on
:25940 instead of going through rosetta, which means they survive a failure
of the gateway itself — not just of an upstream behind it. A chain that
routes every tier through one process only covers upstream outages; the
process is still a single point of failure. Skipping it costs four extra lines
per tier.
Note that the two shim tiers have different base URLs, and that is not a
typo. argo-shim serves /v1/chat/completions and /v1/messages, but the
Anthropic SDK appends /v1/messages to whatever base it is handed. So the
chat_completions tier needs the /v1 written out and the
anthropic_messages tier must not have it. Hermes normalizes this
automatically for its OpenCode-family providers and for nothing else, so a
custom provider is on its own here.
The api_mode on tier 2 is load-bearing for the same reason. A tier that
doesn’t declare one inherits the wire from its providers: block — and
argo-shim is registered as anthropic_messages, so dropping that one line
sends the chat_completions tier down the Messages path instead:
>>> from agent.chat_completion_helpers import _fallback_api_mode_hint
>>> tier = {'provider': 'argo-shim', 'model': 'GPT-5.6 Sol',
... 'base_url': 'http://127.0.0.1:25940/v1',
... 'api_mode': 'chat_completions'}
>>> _fallback_api_mode_hint(tier, tier['provider'], tier['base_url'])
(True, 'chat_completions')
>>> del tier['api_mode']
>>> _fallback_api_mode_hint(tier, tier['provider'], tier['base_url'])
(True, 'anthropic_messages')
Get it wrong in the chat_completions direction and the tier 404s — which is
exactly what mine did, silently, for as long as it had been configured. The
turn still completed, because the chain just kept walking, and that is the
whole problem with a quiet fallback: a dead tier and a tier that never gets
reached look identical from the outside. Worth an actual curl per tier:
curl -s -o /dev/null -w '%{http_code}\n' -X POST \
http://127.0.0.1:25940/v1/chat/completions \
-H "Authorization: Bearer $ARGO_SHIM_TOKEN" \
-H 'Content-Type: application/json' \
-d '{"model":"GPT-5.6 Sol","messages":[{"role":"user","content":"hi"}],"max_tokens":5}'
You can see the chain fire in ~/.hermes/logs/agent.log:
Fallback activated: gpt-5.6-sol → anthropic/claude-opus-5 (llm-rosetta)
Fallback activated: anthropic/claude-opus-5 → GPT-5.6 Sol (argo-shim)
Fallback activated: GPT-5.6 Sol → claude-opus-5 (argo-shim)
Those three are one turn, nine seconds apart, and they are a better
advertisement for the design than a clean run would be. GPT-5.6 Sol hit a
429 token rate limit upstream, Claude came back with an empty stream three
times, and the chain walked off rosetta and onto the shim on its own.
That line is worth knowing by sight. Fallback is quiet by design — the session continues normally — so a model that’s silently unavailable looks like a model that’s working, just with different output. If you see that line constantly, your preferred upstream is down, not merely slow.
Check the direction. This is easy to get backwards: set
defaultto whichever model you were last testing, and the chain now prefers the backstop. Confirm what served a turn by watching the rosetta log rather than trusting the config:tail -f ~/.hermes/logs/rosetta-gateway.log | grep -o 'model=[^ ]*'
Note: Minerva is a third ALCF inference cluster alongside Sophia and Metis, registered the same way:
"alcf-minerva": { "type": "openai_chat", "api_key": "${ALCF_INFERENCE_TOKEN}", "base_url": "https://inference-api.alcf.anl.gov/resource_server/minerva/api/v1", }
iMessage via Photon
Hermes ships a Photon plugin that bridges iMessage, so the agent running on the stationary host is reachable from a phone.
Enable the platform in ~/.hermes/config.yaml:
platforms:
photon:
enabled: true
The plugin runs a Node sidecar that the gateway talks to over gRPC on loopback. A healthy start says so in the gateway log:
[photon] connected — sidecar on 127.0.0.1:8789, streaming inbound over gRPC
✓ photon connected
Gateway running with 2 platform(s)
From there a text message is a prompt. Inbound and outbound both show up in the gateway log:
inbound message: platform=photon user=+1######### msg='...'
response ready: platform=photon time=228.8s api_calls=16 response=1536 chars
Note the time=228.8s with api_calls=16. A texted request runs a real
multi-step tool-calling turn on the host, not a thin chat relay.
Checking that Photon is actually connected
The sidecar being alive is not the same as the platform being connected. Check all three:
# 1. platform registered
hermes status | grep -i photon
# iMessage via Photon ✓ configured (plugin)
# 2. sidecar listening
lsof -nP -iTCP:8789 -sTCP:LISTEN
# 3. gateway actually bridged to it
grep -E '\[photon\] connected|photon connected' ~/.hermes/logs/gateway.log | tail -2
The sidecar logs noisy ZodError lines for message shapes it doesn’t
recognize (polls, some attachments). Those are non-fatal and don’t indicate a
broken connection. Look for an explicit disconnect instead.
Gotchas
The non-obvious problems I hit. Two needed additive patches to argo-shim,
leaving the Anthropic /v1/messages path Claude Code uses untouched; both are
now upstream. The rest were configuration or upstream-app
quirks.
1. OpenAI models need a user field
Argo’s /chat/completions returns a bare HTTP 500 for GPT/Gemini requests
unless the body includes a user field set to a valid ALCF username. With it,
the 500 turns into a real completion. The fix is to have the shim auto-inject
it.
# In argo-shim, alongside the existing /messages handling:
if method == "POST" and body and "/chat/completions" in self.path:
req = json.loads(body)
if isinstance(req, dict) and not (req.get("user") or "").strip():
req["user"] = ARGO_USER # $ARGO_USER / $CELS_USERNAME / login
body = json.dumps(req).encode()
This took a while to find: curl worked but the real clients all returned 500.
The difference was that none of my curl tests happened to send a user field
Argo accepted. Once I tried a valid ALCF username, GPT-4o, GPT-5, and the
o-series all came alive.
2. OpenAI clients send Authorization: Bearer, not x-api-key
The shim originally accepted only x-api-key. rosetta’s openai_chat provider
authenticates with Authorization: Bearer <key>. The shim accepts both now:
client_key = self.headers.get("x-api-key", "")
if not client_key:
auth = self.headers.get("Authorization", "")
if auth.lower().startswith("bearer "):
client_key = auth[7:].strip()
3. Hermes’ custom provider and model-name prefixes
Two Hermes-specific quirks:
- The
customprovider doesn’t readCUSTOM_API_KEYfor the actual request: only an inlineapi_keyinconfig.yaml(which the desktop UI strips on save). Running the gateway no-auth sidesteps this entirely. - Hermes prepends a vendor prefix to model names
(
anthropic/claude-opus-5), then matches that exact string against the endpoint’s/v1/modelslist. So the gateway must register both the bare and vendor-prefixed forms of every model, hence the duplicated entries in the config above.
4. The capabilities default silently drops images
The rosetta admin panel showed every model with a single text capability
badge, even though Claude and GPT-4o do vision. I assumed it was cosmetic. It
isn’t.
A model entry with no capabilities list defaults to ["text"]. A model
without "vision" has images stripped from the request before it’s
forwarded upstream (enforce_vision() replaces them with text placeholders).
Pasting an image into Hermes would have silently dropped it: no error, just a
model that “couldn’t see” the image.
The fix is to declare real capabilities per model in the gateway config:
"claude-opus-5": {
"provider": "argo",
"upstream_model": "Claude Opus 5",
"capabilities": ["text", "vision", "tools", "reasoning"]
}
My regen script sets these per model family (Claude 4.x/5 and
the GPT-5/o-series get text + vision + tools + reasoning; GPT-4o/4.1 get
text + vision + tools). After regenerating, images flow through to Argo
instead of being quietly discarded.
5. Duplicated responses in the Hermes desktop app
After everything worked, hermes desktop started showing every reply twice:
once as a slightly-reworded partial, then again in full. My first instinct was
that the proxy chain was double-emitting.
It wasn’t. I tested the rosetta gateway directly — streaming, non-streaming, and
a fresh non-stream call — and every response came back clean and singular.
The shim and Argo were innocent. The duplication appeared only inside the
Hermes desktop renderer, and only for chatty tool-calling turns (the offending
turn had tool_turns=10).
The culprit was Hermes’ display.interim_assistant_messages: true, which
renders the model’s interim commentary between tool calls and the final
answer. With a verbose reasoning model (GPT-5.6 Sol) those two are
near-identical, so the reply reads twice with slightly different wording. The
fix is one line in ~/.hermes/config.yaml:
display:
interim_assistant_messages: false
Update (2026-09-27): treat that as a workaround, not the setting you should
be running. true is the upstream default, and Hermes has since reworked how
interim commentary is projected across history and tool boundaries. My own
config has been back on the default for months without me noticing doubled
replies. Leave it alone unless you actually see the duplication, then flip it.
The meta-lesson, again: when something looks broken, bisect the layers.
A two-minute curl against the gateway saved me from “fixing” a proxy that was
working perfectly.
6. ALCF inference models appear even when auth is missing
The ALCF inference providers use "api_key": "${ALCF_INFERENCE_TOKEN}" in the
rosetta config. That substitution happens when the gateway starts. With the
env var missing, the models still appear in /v1/models and the admin UI, but
calls 401 because the literal placeholder (or no useful bearer token) gets sent
upstream.
The fix is to refresh the Globus token before restarting rosetta:
alcf-inference-token
rosetta-gateway restart
Or, after an argo-shim restart, do the whole thing:
argo-rosetta-sync --with-inference
7. Gemini is partially available
I first thought Gemini was completely broken, because I kept hitting
'NoneType' object is not iterable after fixing the user field. Careful
re-testing showed three of four models work.
These are Argo Gemini models, not Google GenAI API models. Argo exposes them
through its OpenAI-compatible /chat/completions endpoint, so in rosetta they
use the openai_chat provider type, not the google provider type.
Gemini 2.5 Proworks through Argo’s OpenAI/chat/completionspath.Gemini 2.5 Flashworks too.Gemini 3.5 Flashalso works.Gemini 3.1 Flash Litereturned the upstreamNoneType500 whenever the request carried the usual OpenAI-stylemax_tokensfield, and only worked if you omitted it or sentmax_completion_tokensinstead.
Update (2026-09-27): Argo fixed the last one. gemini-3.1-flash-lite now
answers with max_tokens set and without it, so all four are registered as
bare and google/... aliases. The request transform I was going to write
turned out to be unnecessary — worth re-testing an upstream quirk before
building around it.
Token rotation
Every argo-shim restart mints a new token. Shell helpers in
~/.config/zsh/functions.zsh keep everything in sync. These are the actual
pieces, not magic commands hidden elsewhere.
# Re-export the live shim token into the env vars various clients read.
refresh-argo-token() {
local settings="$HOME/.claude/settings.json"
local token
token=$(python3 -c "import json; print(json.load(open('$settings'))['apiKeyHelper'].split()[-1])") || return 1
export ARGO_SHIM_TOKEN="$token" # OpenCode, vtcode, direct shim clients
export ANTHROPIC_API_KEY="$token" # Anthropic-native clients
export ANTHROPIC_BASE_URL="http://127.0.0.1:25940/v1"
print "ARGO_SHIM_TOKEN refreshed (${token:0:8}…)"
}
# Point Claude Code at rosetta, not directly at argo-shim, so Claude Code
# traffic appears in the rosetta dashboard. argo-shim rewrites Claude settings
# back to :25940/argoapi on startup, so run this after starting argo-shim.
route-claude-through-rosetta() {
python3 - <<'PY'
import json, pathlib
p = pathlib.Path.home() / ".claude" / "settings.json"
d = json.load(open(p))
env = d.setdefault("env", {})
env["ANTHROPIC_BASE_URL"] = "http://127.0.0.1:8765"
for key in ("NO_PROXY", "no_proxy"):
vals = [x.strip() for x in env.get(key, "").split(",") if x.strip()]
for val in ("localhost", "127.0.0.1"):
if val not in vals:
vals.append(val)
env[key] = ",".join(vals)
p.write_text(json.dumps(d, indent=2) + "\n")
PY
}
# Export a fresh Globus access token for ALCF Inference Endpoints. First run may
# require an interactive browser auth flow (see below).
alcf-inference-token() {
local token
token=$(uvx --from alcf-ai alcf-ai auth get-access-token) || return 1
export ALCF_INFERENCE_TOKEN="$token"
print "ALCF_INFERENCE_TOKEN refreshed (${token:0:8}…)"
}
# Minimal rosetta restart helper. My full version also has start/stop/status,
# but this is the essential part: regenerate config if needed, then restart.
rosetta-gateway-restart() {
pkill -f llm-rosetta-gateway 2>/dev/null || true
sleep 1
nohup llm-rosetta-gateway --no-banner >> "$HOME/.hermes/logs/rosetta-gateway.log" 2>&1 &
}
# One-shot recovery after starting/restarting argo-shim.
argo-rosetta-sync() {
refresh-argo-token || return 1
[[ "$1" == "--with-inference" ]] && alcf-inference-token
# If you keep a regen script, run it here. Otherwise ensure config.jsonc uses
# the current $ARGO_SHIM_TOKEN / $ALCF_INFERENCE_TOKEN before restarting.
rosetta-gateway-restart || return 1
route-claude-through-rosetta || return 1
}
The startup ritual after a reboot:
argo-shim # in a terminal: creates tunnel, rotates token
refresh-argo-token # sync Argo token into env
rosetta-gateway-restart # restart rosetta after config/token changes
route-claude-through-rosetta # optional: send Claude Code through rosetta too
# then run `hermes desktop` (or `hermes chat`)
For the ALCF Inference Endpoints in that same gateway:
alcf-inference-token # export Globus access token
rosetta-gateway restart # reload config with ${ALCF_INFERENCE_TOKEN}
Or one bundle after an argo-shim restart:
argo-rosetta-sync --with-inference
Letting launchd do it
Running that by hand gets old, and it fails in a specific way: the gateway
substitutes ${ARGO_SHIM_TOKEN} and ${ALCF_INFERENCE_TOKEN} into
config.jsonc once, at startup. A gateway that is up and listening can
still be holding two dead tokens. Nothing crashes. Every model just 401s.
So there are two launchd jobs now. ai.llm-rosetta.gateway runs the server
under KeepAlive and resolves both credentials on the way up, which is the
whole point — the restart is the refresh mechanism:
# ~/.config/llm-rosetta-gateway/launch-gateway.sh (excerpt)
cd "$HOME" || exit 1
ARGO_SHIM_TOKEN=$("$PY" -c \
"import json;print(json.load(open('$HOME/.claude/settings.json'))['apiKeyHelper'].split()[-1])")
export ARGO_SHIM_TOKEN
ALCF_INFERENCE_TOKEN=$(timeout 120 uv run --with globus-sdk --with openai \
"$HELPER" get_access_token 2>/dev/null | tail -1)
export ALCF_INFERENCE_TOKEN
exec "$HOME/.local/bin/llm-rosetta-gateway" --no-banner --config "$CFG"
Two things in there are load-bearing. exec in the foreground, because
launchd supervises the real process and nohup-ing it means KeepAlive
watches a wrapper that already exited. And cd "$HOME" before the ALCF
helper, because uv run resolves the enclosing project first — from a repo
whose uv.lock your uv is too old to parse, you get a TOML error instead of
a token.
The second job, ai.llm-rosetta.gateway.refresh, is the interesting one. It
runs every six hours (StartInterval 21600) and probes both upstreams
rather than restarting on a schedule:
argo=$(probe "anthropic/claude-haiku-4-5")
alcf=$(probe "alcf-metis/gpt-oss-120b")
stale=0
[[ "$argo" == 401 || "$argo" == 400 ]] && stale=1
[[ "$alcf" == 401 ]] && stale=1
(( stale )) && launchctl kickstart -k "gui/$(id -u)/ai.llm-rosetta.gateway"
A real inference call is the only honest test. /v1/models answers fine with
expired credentials — the catalog is local config, not an upstream query. That
is the same trap as gotcha #6,
just automated.
Note 400 counts as stale for Argo. That one cost me a while: the shim 401s
without draining the POST body, so the desynced keep-alive stream reports the
next request as 400 Bad request syntax rather than an auth error. The
auth failure shows up one request late, wearing a different status code.
The log is boring, which is the goal:
2026-09-27T05:09:49-0500 token-refresh: probe argo=200 alcf=200
2026-09-27T05:09:49-0500 token-refresh: both upstreams OK
2026-09-27T11:09:51-0500 token-refresh: probe argo=200 alcf=401
2026-09-27T11:09:51-0500 token-refresh: ALCF token stale
2026-09-27T11:09:51-0500 token-refresh: kickstarting ai.llm-rosetta.gateway to re-resolve tokens
2026-09-27T11:10:18-0500 token-refresh: post-restart argo=200 alcf=200
The Globus token expired, the probe caught it, and the gateway came back with a fresh one 27 seconds later. I found out by reading the log afterward.
argo-shim itself is still a manual foreground process, since the SSH tunnel
it owns can need MFA and there is no point supervising something that may stop
to ask a human a question. The shell functions above all still work, and are
still what I reach for when I’m actively changing things — launchd handles the
unattended case, not every case.
A regen script
Clicking the admin panel’s Fetch from Provider button re-pollutes the model
list: it re-discovers Argo’s full set and writes broken entries (see the
heads-up above). So I keep a script that regenerates a known-good
config.jsonc from small Python maps of
{ alias: (upstream_model_id, capabilities) }. It registers bare +
vendor-prefixed forms for Argo, adds alcf-sophia/* and alcf-metis/*
aliases, attaches the right capabilities (so images aren’t stripped), pulls the
live Argo token, and restarts the gateway:
rosetta-gateway regen-models # rebuild model list + restart
Editing the model maps at the top of that script is the one place to add or remove models.
Contributing back
Two of the argo-shim fixes were general improvements rather than local
workarounds, so they went back upstream as a PR:
- Accept
Authorization: Bearer <token>in addition tox-api-key, so any OpenAI-format client can authenticate to the shim. - Auto-inject the
userfield on/chat/completions, so OpenAI/Gemini models stop returningHTTP 500. The value resolves from$ARGO_USER, then$CELS_USERNAME, then the login user: no hardcoded usernames.
An automated review on the PR flagged a real edge case: a non-dict JSON body
(a top-level array, say) would crash the handler, since the injection code
assumed req.get(...) always worked. Guarding on isinstance(req, dict) and
treating a blank user as missing fixed it. That’s the isinstance check in
the snippet above.