| name | multi-host-live-mount |
| description | Use when one Relayfile workspace must be mounted on several machines and an agent placed on a remote host inside the live tree without cloning. Covers joining an existing workspace with write opt-in, per-node mount scopes and credentials, fleet enrollment and placement, and byte-level proof that remote mirrors are current. Treats daemon liveness, lag/status lines, and fleet-node presence as non-proof. |
One Workspace, Many Hosts, Nothing Cloned
Take a fresh machine to a state where it mounts an existing Relayfile
workspace, joins the fleet, and hosts a placed agent that does its work inside
the live-mounted tree — no git clone, no copy of the data.
This skill is the composition of three primitives that are documented
separately and lie to you when combined:
| Primitive | Owning skill | What it does NOT tell you |
|---|
relayfile setup / mount | setting-up-relayfile | how a second host joins a workspace that already exists |
agent-relay fleet enroll / spawn | orchestrating-agent-relay | whether the node's mount is current |
mount layout / LAYOUT.md | workspace-layout | that the file you just read may be days stale |
The load-bearing section is Proving the mirror is current.
Everything before it is setup. If you read one section, read that one.
The model
- One workspace, many mirrors. A workspace (
rw_<8hex>) lives in the
cloud. Each host runs its own relayfile-mount daemon projecting some
scope of that workspace into a local directory. The cloud is the single
writer of record; every local tree is a cache.
- A mirror is a cache, and caches go stale silently. There is no
self-healing guarantee and no loud failure. A mount daemon can be alive,
connected, and reporting
lag: 0s while serving two-day-old bytes. Verified
live — see the observed failure.
- Hosts are not symmetric. Two hosts mounting the same workspace routinely
hold different scopes, different credentials, and different write
permissions. "Host B sees what host A sees" is a claim to test, never assume.
- Placement and mounting are independent.
agent-relay places an agent
onto a node. relayfile mounts data on that node. Neither checks the other.
A perfectly placed agent on a node with a dead mount fails in the most
confusing way available: it reads plausible, wrong, old files and never errors.
cloud workspace rw_7ccfea89 ← single source of truth
/ | \
host A (--write) host B (ro) host C (ro, /linear only)
/linear /github /digests scoped mirror
creds: node-A creds: node-B creds: node-C
↑ ↑ ↑
placed agent placed agent placed agent
works in-tree works in-tree works in-tree
Verified against
| Component | Version |
|---|
relayfile | 0.10.39 |
agent-relay | 11.4.2 |
Re-verify with relayfile --version / agent-relay --version; probe the
installed binary rather than trusting this table, since flags have moved before.
Repo surface vs provider surface
The most common wrong turn is mounting the wrong kind of thing. Decide first:
| Repo surface (git) | Provider surface (relayfile mount) |
|---|
| Holds | source code, tests, build config | Linear issues, GitHub PR metadata, Notion pages, Slack messages, digests |
| Get it by | git clone / worktree | relayfile workspace join + relayfile mount |
| Write by | commit + push + PR | writing a file in a writable resource dir → writeback |
| Consistency | explicit fetch; staleness is visible in git status | implicit poll; staleness is invisible |
| Multi-host | every host has a full independent copy | every host has a partial, scoped, possibly-stale cache |
Do not mount source code. The GitHub adapter surfaces PR/review metadata,
not a working tree. An agent that must compile, test, or edit code needs a real
checkout. An agent that must read a Linear issue, comment on a PR, or summarize
a Notion page needs the mount, and cloning nothing is the correct outcome.
A "nothing cloned" host is therefore only coherent when the agent's work is
entirely provider-surface work. State that in the task prompt, because an
agent that discovers it needs source will otherwise clone one silently and
diverge from the design.
Step 1 — Join the existing workspace (the second-host primitive)
relayfile setup creates a workspace. On host two through host N you must
not run it — you will end up with two workspaces and a mystery about why the
hosts disagree. The primitive is join:
relayfile workspace join WORKSPACE_ID [--name NAME] [--write]
# On the fresh machine, after `relayfile login`:
relayfile login --server https://agentrelay.com/cloud # or --api-key for self-hosted
relayfile workspace join rw_7ccfea89 --name shared-ws # read-only by default
relayfile workspace list # '*' marks the active one
relayfile workspace current --verbose
--write is opt-in and it is the whole permission model for this host.
Omit it and the host mirrors read-only; writes never become provider
mutations. Grant it only to hosts whose agents are meant to mutate providers.
WORKSPACE_ID is the rw_<8hex> form. Do not substitute a UUID — most
internal surfaces use UUIDs for the same workspace and they are not
interchangeable (see setting-up-relayfile G5).
- Joining registers the workspace locally in
~/.relayfile/workspaces.json. It
does not start a mount.
Step 2 — Mount a scope, with this node's own credentials
relayfile mount rw_7ccfea89 /path/to/mirror --background
relayfile status rw_7ccfea89
That is the simple form. Real multi-host deployments run the scoped form. This
is the actual argv of a live production mount (secrets elided) — it is the
shape to copy:
relayfile-mount
--base-url https://file.agentrelay.com
--workspace rw_7ccfea89
--local-dir /Users/you/Projects/thing/senses
--local-layout scoped
--creds-file /Users/you/Projects/thing/.agentworkforce/relayfile/<node>-mount.json
--state-dir /Users/you/Projects/thing/.agentworkforce/relayfile/state
--mode poll
--interval 30s
--websocket=true
--remote-path /linear
--remote-path /github
--remote-path /notion
--remote-path /digests
The four flags that make multi-host work:
--remote-path (repeatable) — the node's scope. Each occurrence adds one
subtree to this host's projection. A host that only needs Linear mounts
/linear and never materializes the rest. This is the main lever for keeping
a mirror small enough to actually stay current.
--creds-file — this node's credential, not the human's. Per-node
credentials live under ~/.relayfile/delegated/<shard>/<id>.json (with
.lock siblings). Scoping credentials per node means revoking one host does
not disturb the others.
--state-dir — per-mount sync state. Two mounts sharing a state dir
corrupt each other's revision tracking. One state dir per mount, always.
--local-layout scoped — lay the tree out to match the scope rather than
the full workspace root.
Credential hazard, observed live. Broker-spawned agents receive
rk_live_… workspace keys and at_live_… agent tokens as command-line
arguments, which makes them readable by any local user via ps auxww. On a
shared or multi-tenant host, treat every token handed to a placed agent as
disclosed. Prefer per-node delegated credentials with the narrowest scope, and
rotate anything that has appeared in a process listing or a transcript.
One repo can hold several mounts, from different workspaces
Do not assume a machine has one mount. A single working directory routinely
carries more than one mounted surface, and they may project different
workspaces. Observed live in one repo:
chief/senses/github/… ← workspace rw_7ccfea89 (full projection)
chief/.integrations/github/… ← workspace 50587328-… (UUID, different workspace)
and .integrations is a SYMLINK to
~/.agentworkforce/pear/relayfile/workspaces/<uuid>
Both expose a github/ subtree. An agent that reads github/... without
knowing which root it came from is reading an unidentified workspace. Three
consequences:
- Certify per mount, not per machine. A currency assertion naming
rw_7ccfea89 says nothing about the other mount. Pair every assertion with
the (workspace, mirror root) it actually ran against, and name the pairing in
the result.
- Mounts are often symlinks.
find without -L will not traverse one, and
a naive directory walk reports 0 files for a symlinked provider dir — a
coverage counter that silently measures nothing. The assertion traverses
symlinks while refusing cycles and repeated resolved directories.
- Workspace ids come in both shapes.
rw_<8hex> and bare UUIDs are both
live, sometimes on the same host. They are not interchangeable; pass the exact
id the mount was created with.
Not every mount is a full projection — absence can be by design
Some surfaces deliberately do not download history: records are fetched on
demand or arrive through webhook events, and only writeback command roots plus
discovery/<provider>/ schemas are materialized up front. On such a mount a
cloud file with no local counterpart is expected, not stale.
This breaks the naive reading of Assertion A: MISSING would fire in bulk and
mean nothing. Before treating MISSING as a defect, establish which kind of
mount you have — a full projection, or an on-demand/event-scoped one. For the
latter, make only missing non-fatal; STALE, an incomplete cloud listing,
and every read/listing error still fail the assertion. An on-demand assertion
also needs at least one local/cloud byte comparison — an all-missing listing is
unknown, not current. Say in the result which mode you asserted under. STALE
(present locally but byte-divergent from cloud) remains a real defect in both
modes.
Never write provider records into discovery/ — it holds schemas and examples.
Writeback files belong under the provider's command root.
Scope and write-permission are per host — check, don't assume
relayfile integration list --workspace rw_7ccfea89 --json # which providers this host sees
pgrep -fl relayfile-mount # which --remote-path scopes are live
A file being absent on host B may mean it does not exist, or may mean host B
simply does not mount that scope. Those are different bugs. Distinguish them
with a cloud-side read (relayfile tree), which is scope-independent.
Proving the mirror is current
This is the section that matters.
lag: 0, pending: 0, a live daemon process, and a .relay/state.json
whose mtime is ticking right now are all simultaneously compatible with
content that is days stale. Every one of those signals measures the daemon's
own activity, not the freshness of the bytes it serves. Only a
content-level assertion — comparing mounted content against a fact you know
to be true at this moment — proves currency.
The daemon can poll forever, rewrite its state file every 30 seconds, report
zero lag and zero pending writebacks, and never reconcile a single content file.
That is not a hypothetical; it is the observed
failure.
What is NOT proof
Every one of these was true on a host serving two-day-old content:
| Non-proof | Why it fails |
|---|
pgrep -fl relayfile-mount returns a pid | Proves a process exists. Says nothing about whether its last sync cycle succeeded. |
relayfile supervisor status | Independent of mount health, and often not installed at all — a healthy mount here reported Could not find service "com.relayfile.listen". |
relayfile status shows lag: 0s | The single most dangerous false positive. Observed printing lag: 0s for every provider simultaneously — including one flagged lagging reason: no sync cursor or watermark and one whose last event was 1486 hours earlier. lag: 0s is not a measurement of mirror freshness. |
pending writebacks: 0 | That is the outbound queue. It says nothing about inbound freshness. |
.relay/state.json mtime is seconds old | The most seductive non-proof, because it looks like liveness with a timestamp. Measured on the stale host: all four per-provider state.json files rewritten within ~4 minutes of the check, while the newest actual content file was 1h31m old and digests/today.md was two days old. The daemon writes its state file every cycle whether or not the cycle reconciled anything. |
| Some content files updated recently | Partial reconciliation is the norm in this failure. On the stale host, 9 content files had changed within 3h — and today.md, the file that changes most, had not moved in two days. Freshness is per-path, never global. |
Node appears in agent-relay fleet nodes | Fleet registration is about the relay node, not the mount. Disjoint subsystems. |
Node shows status: online, live: true | Registration fields, not liveness of the mount. |
| The file you read had plausible content | Stale files are perfectly well-formed. That is the entire problem. |
What IS proof
Three assertions. Run all three.
Assertion A — mirror-matches-cloud (read-side currency)
relayfile tree is a live cloud-side listing carrying authoritative
revision, size, and updatedAt per file. Compare it against the bytes on
local disk: the cloud is the oracle, the local daemon is the thing under test.
First, the trap that makes a naive version of this assertion pass by
omission. relayfile tree is paginated, and the pagination is not usable:
| Call | File rows returned |
|---|
tree /linear --depth 20 | 100 |
tree /linear --depth 3 | 325 |
tree /linear --depth 3 --json | 596 entries, plus nextCursor |
files actually present locally under /linear | 3072 |
- Deeper
--depth returned fewer rows — a per-response row cap interacts with
traversal, so a high --depth is actively worse.
- The human-readable form truncates silently. Only
--json reveals the
nextCursor field that tells you the listing was incomplete.
--cursor is not implemented (error: flag provided but not defined: -cursor). The CLI hands you a cursor it cannot consume, so a single tree
call is a page, not a tree.
The workaround is to narrow the path instead of deepening it: walk directory by
directory at --depth 1, and report coverage so a partial verification can
never read as a clean pass.
#!/usr/bin/env bash
# assert-mirror-current.sh — read-side currency proof, coverage-explicit.
# usage: assert-mirror-current.sh <workspace> <local-mirror-dir> <remote-scope> <full|on-demand>
set -uo pipefail
WS="${1:?workspace}"; MIRROR="${2:?local mirror dir}"
SCOPE="${3:?remote scope}"; MODE="${4:?projection mode: full or on-demand}"
python3 - "$WS" "$MIRROR" "$SCOPE" "$MODE" <<'PY'
import json, os, subprocess, sys
ws, mirror, scope, mode = (sys.argv[1], sys.argv[2].rstrip('/'),
sys.argv[3], sys.argv[4])
if mode not in {"full", "on-demand"}:
raise SystemExit("projection mode must be 'full' or 'on-demand'")
def page(path):
"""One cloud-side page. `tree` returns nextCursor but the CLI has no
--cursor flag, so a page is all you get for this path."""
try:
r = subprocess.run(["relayfile", "tree", ws, path, "--depth", "1", "--json"],
capture_output=True, text=True, timeout=180)
except subprocess.TimeoutExpired:
return [], None, "tree timed out after 180s"
if r.returncode != 0:
return [], None, r.stderr.strip()[:120]
try:
d = json.loads(r.stdout[r.stdout.index('{'):])
except (ValueError, json.JSONDecodeError) as e:
return [], None, f"unparseable: {e}"
if not isinstance(d, dict):
return [], None, "tree response is not a JSON object"
entries = d.get("entries")
if not isinstance(entries, list):
return [], None, "tree response has no entries array"
cursor = d.get("nextCursor")
if cursor is not None and not isinstance(cursor, str):
return [], None, "tree response has an invalid nextCursor"
return entries, cursor, None
def cloud_bytes(path):
"""Read the actual cloud artifact. Revision/size are diagnostics, not proof."""
try:
r = subprocess.run(["relayfile", "read", ws, path],
capture_output=True, timeout=180)
except subprocess.TimeoutExpired:
return None, "read timed out after 180s"
if r.returncode != 0:
return None, r.stderr.decode(errors="replace").strip()[:120]
return r.stdout, None
def local_files(root):
"""Return logical paths below root; fail closed on cycles or walk errors."""
paths, walk_errors, visited = set(), [], set()
def visit(local_dir, logical_dir):
resolved = os.path.realpath(local_dir)
if resolved in visited:
walk_errors.append(f"repeated resolved directory (cycle/alias): {logical_dir}")
return
visited.add(resolved)
try:
entries = list(os.scandir(local_dir))
except OSError as e:
walk_errors.append(f"cannot read {logical_dir}: {e}")
return
for entry in entries:
logical = logical_dir.rstrip("/") + "/" + entry.name
if ".relay" in logical.lstrip("/").split("/"):
continue # reserved daemon state is not cloud content
try:
if entry.is_dir(follow_symlinks=True):
visit(entry.path, logical)
elif entry.is_file(follow_symlinks=True):
paths.add(logical)
else:
walk_errors.append(f"non-regular local entry: {logical}")
except OSError as e:
walk_errors.append(f"cannot inspect {logical}: {e}")
visit(root, scope.rstrip("/") or "/")
return paths, walk_errors
match = stale = missing = 0
bad, truncated, errors, cloud_paths = [], [], [], set()
queue, seen = [scope], set()
while queue:
path = queue.pop(0)
if path in seen:
continue
seen.add(path)
entries, cursor, err = page(path)
if err:
errors.append((path, err)); continue
if cursor:
truncated.append(path) # this directory was NOT fully listed
for e in entries:
if not isinstance(e, dict):
errors.append((path, f"invalid cloud entry: {e!r}")); continue
p, typ = e.get("path"), e.get("type")
if not isinstance(p, str) or not p.startswith("/"):
errors.append((path, f"invalid cloud entry path: {p!r}")); continue
parent = path.rstrip("/")
if path != "/" and not p.startswith(parent + "/"):
errors.append((path, f"cloud entry escaped listed directory: {p!r}")); continue
if any(part in {"", ".", ".."} for part in p.split("/")[1:]):
errors.append((path, f"cloud entry contains empty/dot component: {p!r}")); continue
if ".relay" in p.lstrip("/").split("/"):
continue # reserved daemon state is not cloud content
if typ == "dir":
queue.append(p); continue
if typ != "file":
errors.append((p, f"invalid cloud entry type: {typ!r}")); continue
size, rev = e.get("size"), e.get("revision")
cloud_paths.add(p)
lp = mirror + p
if not os.path.lexists(lp):
missing += 1; bad.append(("MISSING", p, size, None, rev))
elif not os.path.isfile(lp):
errors.append((p, "local path is not a regular file")); continue
else:
try:
with open(lp, "rb") as f:
local = f.read()
except OSError as e:
errors.append((p, f"cannot read local file: {e}")); continue
cloud, err = cloud_bytes(p)
if err:
errors.append((p, f"cloud read failed: {err}")); continue
if local == cloud: # compare the artifact, never only its size
match += 1
else:
stale += 1
bad.append(("STALE", p, size, len(local), rev))
checked = match + stale + missing
local_paths, walk_errors = local_files(mirror + scope)
errors.extend((scope, e) for e in walk_errors)
extra = sorted(local_paths - cloud_paths)
if not cloud_paths:
errors.append((scope, "cloud listing contained no file paths; currency not proven"))
if mode == "on-demand" and match + stale == 0:
errors.append((scope, "no local/cloud bytes were compared; currency not proven"))
fatal = stale or truncated or errors or extra
if mode == "full":
fatal = fatal or missing
ok = not fatal
print(f"ASSERT mirror-matches-cloud ({mode}): {'PASS' if ok else 'FAIL'}")
print(f" mount: workspace={ws} mirror={mirror} scope={scope}")
print(f" checked={checked} match={match} stale={stale} missing={missing}")
print(f" coverage: {len(cloud_paths)} cloud paths listed; {len(local_paths)} local paths under {scope}")
if truncated:
print(f" INCOMPLETE: {len(truncated)} dir(s) returned nextCursor and were only "
f"partially listed — currency NOT proven for them: {truncated[:3]}")
if errors:
print(f" ERRORS: {errors[:3]}")
if extra:
print(f" EXTRA LOCAL PATHS (not in cloud listing): {extra[:10]}")
for b in bad[:10]:
print(" ", b)
sys.exit(0 if ok else 1)
PY
The package ships this exact executable as
scripts/assert-mirror-current.sh. Use that shipped file in a gate; do not
retype the listing into a different script. The composition workflow below
checks that each packaged assertion executable is present before it invokes one.
Non-zero exit ⇒ the mirror is not current; do not place an agent on this
host. Read the coverage: line every time — checked far below the local
file count means you verified a corner of the tree, not the tree.
The walk is serial and makes one cloud tree call per directory plus one full
relayfile read per listed file, so a deep scope (/github, /linear) can
take substantially longer than minutes. Run it against the scope the agent
will actually read, not the whole workspace. For a byte-exact check on one
file that matters:
diff <(relayfile read "$WS" /digests/today.md) "$MIRROR/digests/today.md" && echo CURRENT
Assertion B — cross-host-write-visible (the real end-to-end proof)
Read-side currency on one host does not prove two hosts agree. The only proof
that host A and host B share one live workspace is a write on A observed on
B, through the cloud, within a bounded time.
Direction matters: run it both ways, because --write is per host and an
asymmetric grant is invisible until you test the direction that lacks it.
# Preconditions: `/linear` is in BOTH hosts' --remote-path sets, HOST A joined
# with --write, and this is a dedicated throwaway Linear issue/comment surface.
PROBE_SCOPE=/linear
MARK="xhost-$(date -u +%Y%m%dT%H%M%SZ)-$$"
# This concrete example is for a Linear comments directory whose .schema.json
# requires a non-empty `body`; construct a different payload only from its schema.
PROBE_DIR="$PROBE_SCOPE/issues/<issue-id>__<uuid>/comments"
PROBE_FILE="$PROBE_DIR/wb-$MARK.json" # remote path, not a basename search
case "$PROBE_FILE" in "$PROBE_SCOPE"/*) ;; *) exit 2;; esac
# ---- on HOST A (the writer) ----
jq -n --arg body "cross-host visibility probe: $MARK" '{body: $body}' \
> "$MIRROR_A$PROBE_FILE" || { echo 'Could not create source probe' >&2; exit 1; }
SOURCE_PROBE="$MIRROR_A$PROBE_FILE"
if [[ ! -f "$SOURCE_PROBE" ]] || ! grep -qF -- "$MARK" "$SOURCE_PROBE"; then
echo 'ASSERT cross-host-source-written: FAIL — marker absent from exact source file' >&2
exit 1
fi
echo "ASSERT cross-host-source-written: PASS file=$PROBE_FILE"
relayfile writeback status "$WS" --json | jq '{pending, deadLettered: (.deadLettered|length)}'
echo "marker: $MARK"
# ---- on HOST B (the reader) ----
# This polls the cloud artifact and the exact expected target file for the marker.
/path/to/installed-skill/scripts/assert-cross-host-write-visible.sh \
"$WS" "$MIRROR_B" "$PROBE_SCOPE" "$PROBE_FILE" "$MARK"
Rules that make this assertion honest:
- Bound the wait. An unbounded poll turns a failed assertion into a hang,
which reads as "still working" instead of "broken."
- Use a unique marker per run. A previous run's marker still on disk turns
the assertion into a tautology that always passes.
- Bind the file to a shared declared scope. The source and target must both
mount
PROBE_SCOPE, and the source must have --write; a broad grep across a
different mount or scope is not evidence.
- Write into a resource the adapter actually accepts, discovered from
.adapter.md / .schema.json (see writeback-as-files). Never write under
<local-dir>/.relay/ — it is reserved daemon state.
- Prefer a throwaway record. This probe emits a real provider mutation.
Do not aim it at a live issue, page, or channel that people read.
- A pass proves the pair, at that moment, for that scope. It is not
transitive: A↔B passing says nothing about host C, which may mount a
different scope entirely.
- Test only permitted directions as visibility. If B is intentionally
read-only, A→B must pass and B→A must be a schema-valid probe rejected at the
exact resource path; do not call a forbidden reverse write a visibility pass.
In harnessed environments a bare foreground sleep in a wait loop is often
blocked. Run the poll backgrounded, or with the harness's Monitor/until-loop.
Assertion C — known-true-now (the content-level proof)
Assertions A and B compare the mount against the cloud. If the cloud
projection itself is behind the provider, both can pass while the agent still
reads stale reality. The only defense is to anchor on a fact you know is true
right now, independently of the mount.
Pick something whose existence you can confirm out-of-band — an issue you just
filed, a PR number you can see in the provider UI, a message you just posted —
then assert the mount contains it:
# Anchor: a GitHub issue you KNOW exists right now (confirm out-of-band first).
KNOWN=2949
BY_ID_DIR="/github/repos/<owner>__<repo>/issues/by-id"
/path/to/installed-skill/scripts/assert-known-true-now.sh "$WS" "$MIRROR" "$BY_ID_DIR" "$KNOWN"
A gap between newest projected and known-true-now is a projection
failure, upstream of your mirror. Assertion A cannot see it, because local and
cloud agree — on stale data.
Uncertified is not a verdict. A scope you did not assert against is neither
current nor stale; it is unknown. Say so. Reporting "the mount is fine" on the
strength of one certified scope is the same error as lag: 0s, one level up.
The observed failure: every health signal green, content days stale
Recorded on a live production host, and the reason this skill exists:
$ ps -o pid,etime -p 2429 → daemon up 3h06m ← "healthy"
$ relayfile status rw_7ccfea89 → mode: poll lag: 0s ← "healthy"
pending writebacks: 0 ← "healthy"
$ find . -name state.json -mmin -5 → all 4 provider state files ← "healthy"
rewritten minutes ago
$ relayfile supervisor status → service not found (never installed)
$ /path/to/installed-skill/scripts/assert-mirror-current.sh rw_7ccfea89 ./senses /digests full
ASSERT mirror-matches-cloud (full): FAIL
mount: workspace=rw_7ccfea89 mirror=./senses scope=/digests
checked=80 match=78 stale=2 missing=0
coverage: 80 cloud paths listed; 80 local paths under /digests
('STALE', '/digests/this-week.md', 169047, 70301, 'rev_1553124')
('STALE', '/digests/today.md', 38416, 10112, 'rev_1553123')
$ stat -f%Sm senses/digests/today.md → Aug 5 16:07 (wall clock: Aug 7 12:58)
Cloud says today.md is ~38KB and climbing; disk holds 10,112 bytes last
written two days earlier. Repeated sampling showed the cloud revision
advancing (rev_1553031 → rev_1553123 → rev_1553124) while the local size
never moved — so this is a wedged mirror, not a sampling race.
The mount was not idle, which is what makes this failure so hard to see:
| Signal | Measured |
|---|
all 4 per-provider .relay/state.json | rewritten within ~4 min of the check |
| content files changed in last 3h | 9 — but newest mtime 1h31m old |
digests/today.md | 2 days old, cloud revision still advancing |
github projection | newest projected cloud issue #2935; issue #2949 known to exist and absent — the cloud projection is behind the provider, not just the mirror |
linear, notion | uncertified — not asserted, therefore neither current nor stale |
Every status surface reported healthy throughout. An agent placed on this host
would have read a two-day-old digest, produced confident and wrong output, and
nothing would have errored.
The github row is the one to internalize: it is a projection gap, so
Assertion A would have passed — local and cloud agreed with each other, on
stale data. Only Assertion C catches that class. (GitHub projection gap
contributed by Chief; state-tick and content-mtime figures measured directly.)
today.md is also the worst possible file to have stale — the failure
concentrates in exactly the files that change most.
Composing enrollment, placement, and the mount
Order matters. Mount before placing, verify between.
# 1. Fleet must be enabled for the workspace before a node is brought up.
agent-relay fleet enable
# 2. Enroll the fresh machine as a node (mint on control plane, redeem on node).
# Script + authoritative README: dev-stack/fleet-node-bootstrap/ in the
# AgentWorkforce/cloud repo — not shipped with this skill.
read -r -s -p 'Enrollment token: ' RELAY_ENROLLMENT_TOKEN; printf '\n'
export RELAY_ENROLLMENT_TOKEN
export RELAY_ENROLLMENT_URL='https://<app>/api/v1/fleet/register'
export RELAY_NODE_NAME='<node>'
trap 'unset RELAY_ENROLLMENT_TOKEN RELAY_ENROLLMENT_URL RELAY_NODE_NAME RELAY_AGENT_TOKEN' EXIT
sandbox-node-bootstrap.sh preflight && sandbox-node-bootstrap.sh enroll
unset RELAY_ENROLLMENT_TOKEN RELAY_ENROLLMENT_URL RELAY_NODE_NAME
# 3. On the TARGET host, declare and mount the target-local manifest.
# `/github` supplies the Assertion C anchor and `/linear` supplies the
# Assertion B probe in this example. Substitute compatible in-scope paths;
# do not drop an assertion without a real replacement.
TARGET_WS=rw_7ccfea89
TARGET_MIRROR=/path/to/mirror
TARGET_MOUNT_SCOPES=(/digests /github /linear)
TARGET_PROJECTION_MODES=(full on-demand on-demand) # one verified mode per scope
(( ${#TARGET_MOUNT_SCOPES[@]} )) || { echo 'Refusing placement: no mount scopes declared' >&2; exit 2; }
(( ${#TARGET_MOUNT_SCOPES[@]} == ${#TARGET_PROJECTION_MODES[@]} )) || {
echo 'Each mount scope needs one projection mode' >&2; exit 2
}
case "$TARGET_MIRROR" in /*) ;; *) echo 'Target mirror path must be absolute' >&2; exit 2;; esac
for i in "${!TARGET_MOUNT_SCOPES[@]}"; do
scope=${TARGET_MOUNT_SCOPES[$i]}
mode=${TARGET_PROJECTION_MODES[$i]}
case "$scope" in /*) ;; *) echo "Invalid remote scope: $scope" >&2; exit 2;; esac
case "$mode" in full|on-demand) ;; *) echo "Invalid projection mode for $scope: $mode" >&2; exit 2;; esac
done
MOUNT_ARGS=()
for scope in "${TARGET_MOUNT_SCOPES[@]}"; do MOUNT_ARGS+=(--remote-path "$scope"); done
mount_fingerprint() { printf '%s\n' "$TARGET_WS" "$TARGET_MIRROR" "${TARGET_MOUNT_SCOPES[@]}"; }
if command -v shasum >/dev/null 2>&1; then
MOUNT_ID=$(mount_fingerprint | shasum -a 256 | awk '{print $1}')
elif command -v sha256sum >/dev/null 2>&1; then
MOUNT_ID=$(mount_fingerprint | sha256sum | awk '{print $1}')
else
echo 'Need shasum or sha256sum to derive a per-mount state directory' >&2; exit 2
fi
TARGET_STATE_DIR="/path/to/relayfile-state/$MOUNT_ID" # unique to workspace + mirror + exact scope set
relayfile workspace join "$TARGET_WS" --name shared-ws # add --write only if it must mutate
relayfile mount "$TARGET_WS" "$TARGET_MIRROR" --background --local-layout scoped \
--creds-file '/path/to/<node>-mount.json' --state-dir "$TARGET_STATE_DIR" \
"${MOUNT_ARGS[@]}"
# 4. TARGET-HOST GATE: all three assertions must pass before placement.
ASSERT_DIR=/path/to/installed-skill/scripts
ASSERT_MIRROR_CURRENT="$ASSERT_DIR/assert-mirror-current.sh"
ASSERT_KNOWN_TRUE_NOW="$ASSERT_DIR/assert-known-true-now.sh"
ASSERT_CROSS_HOST="$ASSERT_DIR/assert-cross-host-write-visible.sh"
for assertion in "$ASSERT_MIRROR_CURRENT" "$ASSERT_KNOWN_TRUE_NOW" "$ASSERT_CROSS_HOST"; do
[[ -x "$assertion" ]] || { echo "Missing packaged assertion: $assertion" >&2; exit 2; }
done
for i in "${!TARGET_MOUNT_SCOPES[@]}"; do
scope=${TARGET_MOUNT_SCOPES[$i]}
mode=${TARGET_PROJECTION_MODES[$i]}
"$ASSERT_MIRROR_CURRENT" "$TARGET_WS" "$TARGET_MIRROR" "$scope" "$mode" || exit 1
done
# Assertion C: choose an anchor that is independently true now and in a mounted scope.
KNOWN=2949 # replace after confirming in the provider UI
BY_ID_DIR="/github/repos/<owner>__<repo>/issues/by-id"
"$ASSERT_KNOWN_TRUE_NOW" "$TARGET_WS" "$TARGET_MIRROR" \
"$BY_ID_DIR" "$KNOWN" || exit 1
# Assertion B: HOST A must mount /linear too and have --write. It writes the
# exact schema-valid probe shown above and verifies its source file on HOST A;
# run this cloud-and-target poll on the TARGET host (B).
# For an intentionally read-only target, assert its attempted reverse probe is
# rejected at that exact resource path instead of expecting B→A visibility.
MARK='<paste marker from host A>'
PROBE_FILE="/linear/issues/<issue-id>__<uuid>/comments/wb-$MARK.json"
"$ASSERT_CROSS_HOST" "$TARGET_WS" "$TARGET_MIRROR" /linear \
"$PROBE_FILE" "$MARK" || exit 1
# This target joined read-only. Its reverse direction must reject the exact
# schema-valid write. A local permission rejection can be immediate; if the
# filesystem accepts it, wait for the exact asynchronous result. Use a dedicated
# throwaway resource because a FAIL below means the write may have mutated cloud.
REJECT_MARK="xhost-reject-$(date -u +%Y%m%dT%H%M%SZ)-$$"
REJECT_FILE="/linear/issues/<issue-id>__<uuid>/comments/wb-$REJECT_MARK.json"
REJECT_LOCAL="$TARGET_MIRROR$REJECT_FILE"
if WRITE_ERR=$( { jq -n --arg body "must be rejected: $REJECT_MARK" \
'{body: $body}' > "$REJECT_LOCAL"; } 2>&1 ); then
LOCAL_REJECTED=0
elif grep -qF "$REJECT_LOCAL" <<<"$WRITE_ERR" &&
grep -Eiq 'permission denied|read-only file system|operation not permitted' <<<"$WRITE_ERR"; then
LOCAL_REJECTED=1
else
echo "ASSERT cross-host-write-rejected: FAIL — unexpected local write error: $WRITE_ERR" >&2
exit 1
fi
# A filesystem-enforced rejection is valid only when the exact cloud path is
# also absent. Other read errors are not absence evidence.
if (( LOCAL_REJECTED )); then
if READ_RESULT=$(relayfile read "$TARGET_WS" "$REJECT_FILE" 2>&1); then
echo 'ASSERT cross-host-write-rejected: FAIL — exact cloud path exists' >&2
exit 1
fi
if grep -qF "$REJECT_FILE" <<<"$READ_RESULT" &&
grep -Eiq '404|not found|does not exist' <<<"$READ_RESULT"; then
echo "ASSERT cross-host-write-rejected: PASS — local permission rejection plus exact cloud absence"
echo " mount: workspace=$TARGET_WS mirror=$TARGET_MIRROR scope=/linear file=$REJECT_FILE"
else
echo 'ASSERT cross-host-write-rejected: FAIL — cloud absence was not proven' >&2
exit 1
fi
else
REJECTED=0
for i in $(seq 1 12); do
READONLY_SEEN=0
while IFS= read -r -d '' state_file; do
if jq -e --arg path "$REJECT_FILE" \
'.files[$path].readonly == true and ((.files[$path].dirty // false) == false)' \
"$state_file" >/dev/null; then
READONLY_SEEN=1
break
fi
done < <(find "$TARGET_STATE_DIR" -type f -name state.json -print0)
if READ_RESULT=$(relayfile read "$TARGET_WS" "$REJECT_FILE" 2>&1); then
echo 'ASSERT cross-host-write-rejected: FAIL — exact cloud path exists' >&2
exit 1
fi
if (( READONLY_SEEN )) && grep -qF "$REJECT_FILE" <<<"$READ_RESULT" &&
grep -Eiq '404|not found|does not exist' <<<"$READ_RESULT"; then
REJECTED=1
break
fi
[ "$i" -lt 12 ] && sleep 10
done
if (( ! REJECTED )); then
echo 'ASSERT cross-host-write-rejected: FAIL — no exact read-only result plus cloud absence within 120s' >&2
exit 1
fi
echo "ASSERT cross-host-write-rejected: PASS after a bounded result check"
echo " mount: workspace=$TARGET_WS mirror=$TARGET_MIRROR scope=/linear file=$REJECT_FILE"
fi
# If the target instead joined with --write, run the same schema-valid probe in
# the reverse direction on HOST A and require assert-cross-host-write-visible.sh
# there to pass before returning to this placement step.
# 5. On the CONTROL host, declare the target manifest again. These are target
# paths and scopes, not variables inherited from the target shell.
TARGET_NODE='<node>'
TARGET_CWD=/path/to/mirror
TARGET_SCOPE_LIST=(/digests /github /linear)
(( ${#TARGET_SCOPE_LIST[@]} )) || { echo 'Refusing placement: target scopes missing' >&2; exit 2; }
case "$TARGET_CWD" in /*) ;; *) echo 'Target cwd must be absolute' >&2; exit 2;; esac
for scope in "${TARGET_SCOPE_LIST[@]}"; do
case "$scope" in /*) ;; *) echo "Invalid target scope: $scope" >&2; exit 2;; esac
done
# Read the agent token without printing it. A targeted fleet spawn requires this
# credential; enrollment tokens cannot substitute.
read -r -s -p 'Agent token: ' RELAY_AGENT_TOKEN; printf '\n'
export RELAY_AGENT_TOKEN
# Only now place the agent, with --cwd inside the live mount.
agent-relay fleet spawn claude --name worker-1 --node "$TARGET_NODE" --channel general \
--cwd "$TARGET_CWD" \
--task "Work inside the mounted tree at $TARGET_CWD. Read only these mounted scopes: ${TARGET_SCOPE_LIST[*]}. Do not clone any repository."
Never skip preflight on a machine that already runs brokers.
agent-relay node up kills every broker whose CWD resolves to the same
project root; a $HOME-rooted workdir reaps every $HOME-rooted broker
(relay#1328, a real production incident). Pin AGENT_RELAY_PROJECT to a
unique per-instance dir and drop a physical .agentworkforce/relay marker.
Step 4 is the step everyone skips. Placement succeeds against a stale mount and
reports success.
Fleet node lists are partial — absence is not evidence
Measured on a live workspace, same binary, minutes apart:
| Query | Records | Live |
|---|
agent-relay fleet nodes (default) | 3 | 3 |
agent-relay fleet nodes --all | 400 (exactly, on 3/3 consecutive runs) | 20, 20, 21 |
agent-relay fleet nodes --all --capability spawn:claude | 33 | 4 |
What this means in practice:
- The default view omitted a live, spawn-capable node.
sf-mini was
status: online, live: true, advertising spawn:claude|codex|gemini|opencode
— and absent from the default 3. Never conclude a node is missing from the
default view.
--all returns exactly 400 and there is no --limit flag. 400 on every
run is a server-side cap, not a coincidence. Beyond it, records are silently
dropped. A capability-filtered query returned 33, well under the cap — which
is why filtering is the reliable form.
- The live subset changes between consecutive calls (20 → 20 → 21) as
sessions come and go, while the id set held stable across those three runs.
Do not build logic on a node count.
- Most of the
--all bulk is node_direct_* session records, not placement
targets. spawn:* capability is what makes a node spawnable.
The reliable placement query, and the parse that survives the output format:
# Redirect, never pipe: output truncates at 64KB through a pipe.
agent-relay fleet nodes --all --capability spawn:claude > /tmp/nodes.raw
python3 - <<'PY'
import json, re
raw = open('/tmp/nodes.raw').read()
m = re.search(r"^\{", raw, re.M) # skip any human-readable preamble
if not m:
raise SystemExit("No JSON. Raw:\n" + raw[:500])
for n in json.loads(raw[m.start():]).get("nodes", []):
if n.get("live"):
print(n["name"], n["id"], n.get("status"),
[c["name"] for c in n.get("capabilities", [])])
PY
Compare dispatchedNodeId from a spawn against the id (node_…), not the
name. And per orchestrating-agent-relay, dispatch recorded by the control
plane still is not execution — confirm with pgrep on the target host.
What the placed agent sees, and what to tell it
The agent gets an ordinary directory. Read, Write, Edit, Glob, Grep
all behave normally, which is precisely why staleness is invisible to it.
Put this in the task prompt:
You are working inside a live Relayfile mount at <MIRROR>. It is a mounted
projection of a shared cloud workspace, not a checkout.
- Do NOT clone any repository. Everything you need is in the mounted tree.
- Start at <MIRROR>/LAYOUT.md and each <provider>/LAYOUT.md. Use the by-*
alias indexes rather than find/grep -r across the tree.
- This host mounts only these scopes: <list>. A path outside them is not
missing — it is unmounted. Confirm with `relayfile tree <ws> <path>` before
concluding anything does not exist.
- Before you rely on any file whose freshness matters, verify it against the
cloud: `diff <(relayfile read <ws> <path>) <MIRROR><path>`. A local file can
be days stale while every status command reports healthy.
- To write back, read the resource's .adapter.md and .schema.json first, then
write the JSON to a non-canonical filename in the resource dir. Never write
under <MIRROR>/.relay/ — it is reserved daemon state.
- Writes here become REAL provider mutations. No dry run exists.
Troubleshooting
| Symptom | Cause / fix |
|---|
| Second host created its own workspace | Ran relayfile setup instead of relayfile workspace join <rw_id>. Delete the stray local registration and join the real id. |
| Agent's writes never reach the provider | Host joined without --write. Re-join with --write; confirm with a cross-host-write-visible run. |
| File exists on host A, missing on host B | Almost always scope, not sync. Compare --remote-path sets; confirm existence cloud-side with relayfile tree. |
lag: 0s but content is old | Expected — lag is not a freshness measurement. Run mirror-matches-cloud. |
Provider row reads lagging reason: no sync cursor or watermark | That provider never established a sync cursor; its subtree may never have populated. Do not treat its absence as data. |
| Two mounts corrupting each other's state | Shared --state-dir. Give every mount its own. |
delegated credential workspace "<id>" was not uniquely resolved warning | Multiple/alias shards in ~/.relayfile/workspaces.json; the CLI falls back to probing. Harmless to reads, but pin the exact workspace id in scripts. |
| Writebacks stuck / dead-lettered | relayfile writeback list --state dead --json, inspect <mirror>/.relay/dead-letter/<opId>.json, fix cause, relayfile writeback retry --opId <op> <ws>. relayfile writeback skip-stuck as a last resort. |
| Mirror never converges; sync cycle repeats "forcing full reconcile" | Adapter emitted one path as both file and directory — POSIX cannot hold both, so bootstrap never completes. Adapter-side fix; see setting-up-relayfile. |
relayfile tree shows fewer files than exist | It is paginated and capped per response; higher --depth returns fewer rows. Use --json (only it exposes nextCursor), walk --depth 1 per directory, and never treat one call as a complete listing. is not implemented. |
Named assertions
Report these by name. A run that did not execute an assertion must say so
rather than implying it passed.
| Assertion | Proves | Fails when |
|---|
workspace-joined-not-created | host registers the existing rw_* id | relayfile workspace list shows a new id |
scope-declared | this host's --remote-path set is known and intentional | scope inferred rather than read from the live daemon argv |
mirror-matches-cloud | local bytes equal the artifact returned by cloud relayfile read for the complete scoped path set | a content mismatch, extra local path, non-regular local entry, cloud-list/read error, or incomplete listing; MISSING also fails in full mode, and on-demand needs at least one byte comparison |
listing-coverage-reported | the currency check states how much of the tree it actually verified | any directory returned nextCursor, or coverage went unreported |
known-true-now | mounted content contains a fact confirmed true out-of-band right now | the anchor record is absent — the projection is behind the provider |
uncertified-scopes-named | every scope the agent will read was asserted, or is explicitly listed as unknown | an unasserted scope is reported as current (or as stale) |
mount-identified | each result names the exact (workspace id, mirror root) pair it ran against | a repo holds several mounts and the result says only "the mount" |
projection-mode-known | the mount is known to be a full projection vs on-demand/event-scoped before MISSING is called a defect | bulk MISSING on an on-demand mount reported as staleness, or stale/incomplete/error results ignored in either mode |
cross-host-write-visible | visibility through the cloud in every permitted direction, bounded; exact write rejection in forbidden directions | marker not observed at its expected path within the ceiling, or a read-only target accepts the probe |
write-permission-matches-intent |
Related skills
setting-up-relayfile — first-time setup, OAuth, integrations, writeback recovery
orchestrating-agent-relay — broker lifecycle, spawning, fleet enrollment, placement proof
workspace-layout — navigating a mount via LAYOUT.md and by-* indexes
writeback-as-files — the file-creation writeback contract
using-agent-relay — participant-side messaging for the placed agent