用 Codex 或 Claude 帮你安装 复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它检查 Skill 页面并帮你完成安装。
直接命令不会经过审查 Prompt;运行前请先检查来源。
npx skills add https://github.com/ZeusLN/zeus --skill zeus-node-lifecycle-campaign命令会保持在同一行。复制前请横向滚动并检查完整内容。
想先保存到本地?可下载 SkillsMP 当前能够提供的文件。
Load when setting up the Zeus dev environment from scratch, running yarn install / postinstall, or debugging build/toolchain failures. Triggers - "yarn install fails", "postinstall error", missing Lndmobile.aar / Lndmobile.xcframework / LDKNodeFFI / cdkFFI / zeusRestoreFFI, "checksum failed", fetch-libraries.sh, rn-nodeify, shim.js, pod install / CocoaPods errors, Gradle / Kotlin / NDK / minSdk questions, Xcode framework-not-found, autolinking / native module not found, MainApplication.kt package registration, zeus_modules / lnc-rn vendoring, uniffi binding version mismatch, gen-proto, Hermes or New Architecture build crashes.
Load when you need to understand WHY Zeus is structured the way it is before changing it: app boot/connection sequence (App.tsx vs views/Wallet/Wallet.tsx fetchData, fetchLock, connecting), the stores/Stores.ts singleton DI graph, MobX 5 decorators and @inject/@observer, utils/BackendUtils.ts dispatch and the call()-returns-false invariant, supports*() capability gating vs implementation branching, EmbeddedLND/LndHub inheritance fallthrough, NavigationService/popTo navigation contract, New Architecture (Fabric) constraints, and the catalog of known weak points. Symptoms/triggers: 'false.then is not a function' crashes, a backend method that silently does nothing, a feature appearing on a backend that cannot support it, infinite 'connecting' spinner, reconnect re-entrancy, embedded node hitting a REST URL, navigation going forward instead of back, 'where does the app actually connect to the node', 'where do I add a new store/backend method'.
Load when working across Zeus's 7 Lightning backends (lnd, embedded-lnd, ldk-node, lightning-node-connect, cln-rest, lndhub, nostr-wallet-connect) - adding a new RPC or feature to backends/ or utils/BackendUtils.ts, adding or checking a supports* capability flag, gating a view on backend support, debugging "works on LND but crashes on CLN/NWC/LndHub", ".then is not a function" / "false.then" crashes, features silently doing nothing on one backend, payment success reported despite timeout, LNC showing send UI on read-only pairings, CLN invoice/payment lists missing entries, LDK Node channel-open or payment result handling, stale Tor request errors, or filling in the PR-template backend testing matrix.
基于 SOC 职业分类
正在显示 SKILL.md
| name | zeus-node-lifecycle-campaign |
| description | Load when working on Zeus's node lifecycle races - the |
This is an executable, decision-gated investigation campaign for Zeus's hardest live problem (maintainer-confirmed, 2026-07-06): races in starting, stopping, deleting, and switching embedded nodes (embedded LND and LDK Node), including Tokio-thread crashes and Wallet.tsx fetchData re-entrancy. Everything below was verified by reading the code at master c5fd094fb (v13.1.3-alpha) unless explicitly labeled HYPOTHESIS.
Use this skill when you are reproducing, diagnosing, or fixing node start/stop/delete/switch races, or reviewing any PR that touches views/Wallet/Wallet.tsx connection flow, SettingsStore lifecycle flags, utils/LndMobileUtils.ts, utils/LdkNodeUtils.ts, or the native node modules.
Do NOT use this skill for:
| Need | Go to sibling skill |
|---|---|
| General symptom→triage for other bugs | zeus-debugging-playbook |
| How to capture logs / measure (general tooling) | zeus-diagnostics-and-tooling |
| Past incident narratives beyond lifecycle | zeus-failure-archaeology |
| Backend capability matrix, adding an RPC | zeus-backends-and-capabilities |
| Formal race-analysis method (generic recipes) | zeus-proof-and-analysis-toolkit |
| PR/commit/review rules, the non-negotiables | zeus-change-control |
| Running the app / connecting nodes at all | zeus-run-and-operate |
| Settings blob, migrations, keychain | zeus-storage-and-migrations |
| What counts as test evidence, jest traps | zeus-validation-and-qa |
'embedded-lnd'.ldk-node crate, ZeusLN fork) embedded via uniffi FFI bindings; implementation string 'ldk-node'.ldk_node.kt / LDKNode.swift are its output.LndMobileService.java, a bound Android Service the RN module talks to via Messenger message passing (Blixt-derived code, TODO(hsjoberg) debt).'focus' every time a screen becomes active — including returning from another screen. Wallet.tsx re-runs its connect logic on focus.active/background/inactive); Wallet also re-runs connect logic on active.vss.zeusln.com).views/Wallet/Wallet.tsx that resets stores, starts/attaches to the node for the active backend, and loads wallet data. It IS the boot sequence (not App.tsx).There is no single owner. Lifecycle state is smeared across four layers:
| Layer | File | Owns |
|---|---|---|
| Orchestration | views/Wallet/Wallet.tsx | handleFocus / handleAppStateChange / handleOpenURL → getSettingsAndNavigate() → fetchData(); the _navigating instance flag; a 60s slow-startup alert timer |
| Flags | stores/SettingsStore.ts | connecting (init true), fetchLock, embeddedLndStarted, walletJustCreated, triggerSettingsRefresh, posWasEnabled, ldkNodeSyncing; setConnectingStatus(); updateNodeProperties() (swaps implementation, dirs, creds on every settings write) |
| JS node control | utils/LndMobileUtils.ts (initializeLnd, startLnd, stopLnd, waitForRpcReady, createLndWallet, deleteLndWallet) and utils/LdkNodeUtils.ts (startLdkNodeWallet, stopLdkNode, waitForLdkNodeReady, createLdkNodeWallet, deleteLdkNodeWallet) | retry loops, sleeps, error classification (utils/LndMobileErrors.ts) |
| Native | Android android/app/src/main/java/com/zeus/LndMobileService.java (+ LndMobile.java, LndMobileScheduledSyncWorker.java), android/app/src/main/java/app/zeusln/zeus/LdkNodeModule.kt; iOS ios/LndMobile/, ios/LdkNodeMobile/LdkNodeModule.swift | the actual daemon/Node object, lndStarted boolean, nodeLock/takeNode() |
| Mutation entry points | views/Settings/Wallets.tsx (switch), views/Settings/WalletConfiguration.tsx (create/edit/delete), (Quick Start), |
_navigating (Wallet.tsx instance field) — guards getSettingsAndNavigate re-entry from the focus event and AppState active. Added in c71a04ce9 (2026-04-07) because the focus event fires twice on startup and both raced into fetchData → duplicate buildNode, second build times out holding VSS resources. Bypassed by three call sites that invoke getSettingsAndNavigate() directly: pull-to-refresh (LayerBalances onRefresh), the error-screen Restart button (iOS branch), and handleOpenURL for zeusln://share.fetchLock (SettingsStore observable) — checked/set at fetchData entry (if (fetchLock) return; SettingsStore.fetchLock = true;). Cleared in exactly two places: the happy-path end of fetchData, and setConnectingStatus(true). Every early-return error path leaves fetchLock === true — by design: the app then shows the error screen, and recovery requires setConnectingStatus(true) (Restart button, wallet switch, transient-retry loop), which clears the lock. Consequence: any code path that errors out and then expects a plain focus event to reconnect will hang silently.connecting — initialized true, so the first fetchData after cold start runs the full connect branch (stop→init→start + reset of ~12 stores). Set false at the end of a successful connect or on fatal errors; set true by every wallet-switch/create/restart path. When false, fetchData becomes a light data refresh (embedded-lnd still calls waitForRpcReady() + full data fetch on every focus).walletJustCreated — set by createLndWallet / LDK create paths; makes fetchData skip the stop→init→start cycle because the node is already running from wallet creation. Consumed (set false) on first use.embeddedLndStarted — whether stopLnd() is called before re-init (if (!recovery && SettingsStore.embeddedLndStarted) await stopLnd();). Set true in (only when a wallet password is present), in and several views.setConnectingStatus(true) side effects (SettingsStore.ts)Resets error/errorMsg/lndFolderMissing, calls BackendUtils.clearCachedCalls() (flushes the request dedup cache — see zeus-backends-and-capabilities), and clears fetchLock. This is the universal "reconnect" lever; treat any call to it as a lifecycle transition.
fetchData (only when connecting)stopLdkNode() (unless walletJustCreated) → startLdkNodeWallet (initializeNode FFI + node.start() retried up to 5×500ms on "Node not initialized" + syncWallets) → later branch waitForLdkNodeReady() (polls status() every 500ms up to 60s, tolerating "Node not initialized" and not-running — added by b54e8f138, 2026-05-17; the companion buildNode-race fix 8fc9698a0, same day, added the walletJustCreated skip).stopLnd() (if started) → initializeLnd (writes lnd.conf, binds service) → optional express graph sync → startLnd (retry ladder incl. force-stop on LND_ALREADY_RUNNING, wallet unlock on WALLET_LOCKED) → waitForRpcReady() (polls getInfo 500ms up to 30s). Transient RPC errors route through retryOnTransientError (max 5 retries, 2s base exponential backoff; it toggles connecting false→true, which is also what re-clears fetchLock).connect(), NWC connectNWC(), REST default) have no native lifecycle but share the same flags.From LdkNodeModule.kt (stop, buildNode) and LdkNodeModule.swift (same methods), established by d416052cd + 0e2b89164 (both 2026-03-10):
nodeLock/takeNode() first, so no other call can reach a stopping node.node.stop() on a dedicated non-Tokio thread (raw Thread on Android, GCD global queue on iOS) — never on a Tokio callback thread.stop() returns — callers may delete wallet files or build a new node the moment stop() resolves.Thread.sleep(2000) + hashCode() on Android, withExtendedLifetime on iOS) so internal Tokio tasks drop their Arc refs before ours — otherwise the "Runtime dropped on worker thread" panic.buildNode applies the same stop-and-hold to any existing node before building, and clears the JS-visible node reference up front — which is why any FFI call during a rebuild window rejects with "Node not initialized" (JS must tolerate it: LDK_NODE_NOT_INITIALIZED handling in utils/LdkNodeUtils.ts).Embedded LND has no equivalent contract; stopLnd() in utils/LndMobileUtils.ts approximates one (status check → stopDaemon → killLnd → poll status up to 10×500ms).
Each window is verified by code read at c5fd094fb. Runtime confirmation status noted.
| # | Window | Status |
|---|---|---|
| W1 | Double focus → duplicate getSettingsAndNavigate | FIXED by _navigating (c71a04ce9); residual: the three bypass call sites rely only on fetchLock, which does NOT stop a second getSettingsAndNavigate from re-running the pre-fetchData navigation logic concurrently |
| W2 | JS call into LDK during buildNode rebuild → "Node not initialized" rejection | Tolerated in waitForLdkNodeReady (b54e8f138) and the pre-existing start retries; unguarded for every other store that calls the backend during the window (e.g. polling stores) |
| W3 | Cross-implementation switch leaves the previous node running: fetchData's ldk-node branch calls only stopLdkNode(), the embedded-lnd branch only stopLnd(). Switching embedded-lnd → ldk-node (or reverse) stops NEITHER the old node; views/Settings/Wallets.tsx onWalletPress only calls setPersistentMode(false) (dismisses the Android notification, does not stop the node) and BackendUtils.disconnect() for LNC | Confirmed by code read; runtime impact (two chain-syncing nodes, memory/battery, port/db contention) needs Phase 3 measurement |
| W4 | Delete-active-wallet flow: WalletConfiguration.deleteNodeConfig stops node → updateSettings({ justDeletedWallet: active }) → deletes files → navigates to Wallets; selecting the next wallet runs handleJustDeletedWallet which does setConnectingStatus(true) then sleep(2000) — a bare time-based wait standing in for "old node fully gone" | Open; the sleep is a fenced symptom-patch pattern (below) |
| W5 | JS reload with live daemon: LND runs in the app process (android/app/src/main/AndroidManifest.xml declares LndMobileService with no android:process attribute), so a JS-only restart (RNRestart, dev-menu reload, JS crash) leaves the Go daemon running → next start hits → force-stop ladder in with 2–4s cleanup sleeps |
Goal: a numeric crash/hang rate per scenario BEFORE any fix, so any fix can be judged. Do not skip; the maintainer's revert-first culture means an unmeasured "fix" is unlandable.
yarn android / yarn ios, Metro via yarn start).embedded-lnd (Mainnet or Testnet) and one ldk-node. Create via IntroSplash Quick Start + Settings → Wallets → Add.# Android — native lifecycle tags + crashes (works on debug and release):
adb logcat -v time | grep -E "ReactNativeJS|LdkNodeModule|LndMobileService|LndMobile|AndroidRuntime|FATAL|Fatal signal"
# Android — crash buffer only (count crashes across trials):
adb logcat -b crash -v time
# Android JS console in DEBUG builds goes to the Metro terminal, not logcat.
# In RELEASE builds console.log appears under the ReactNativeJS logcat tag.
# LDK Node's own Rust log (debug builds; substitute the wallet's uuid dir):
adb shell run-as app.zeusln.zeus sh -c 'ls files/ldk-node/'
adb shell run-as app.zeusln.zeus sh -c 'tail -f files/ldk-node/<uuid>/ldk_node.log'
# Embedded LND's log: in-app Settings → Embedded node → LND Logs, or (debug):
adb shell run-as app.zeusln.zeus sh -c 'tail -f files/<lndDir>/logs/bitcoin/mainnet/lnd.log'
# iOS — native NSLog (simulator; process name is "zeus"):
xcrun simctl spawn booted log stream --predicate 'process == "zeus"' --style compact
# iOS JS console: Metro terminal.
All verified present in source:
| Milestone | Exact string (grep-able) | Source |
|---|---|---|
| Focus fired, connect starting | [Wallet] handleFocus: triggering getSettingsAndNavigate | Wallet.tsx |
| Re-entry suppressed | skipping — getSettingsAndNavigate already in flight | Wallet.tsx (focus + AppState variants) |
| LDK teardown | [LDK startup] stopping previous instance / [LDK startup] stopped | Wallet.tsx |
| LDK build+start | [LDK startup] calling startLdkNodeWallet / LDK Node: Started successfully / LDK Node: Sync complete | Wallet.tsx / LdkNodeUtils.ts |
| LDK ready-wait | [LDK startup] waiting for node to be ready | Wallet.tsx |
| LDK native timing | LdkNodeModule: [timing] FFI buildWithDualStore starting/completed | LdkNodeModule.kt/.swift |
| LND ready | RPC ready - Node pubkey: | LndMobileUtils.ts waitForRpcReady |
| LND states | Current LND state: (NON_EXISTING/LOCKED/UNLOCKED/RPC_ACTIVE/SERVER_ACTIVE) | LndMobileUtils.ts |
| LND double-start recovery | LND already started - force stop (Go thinks running) | LndMobileUtils.ts |
| Connect finished | [Wallet] connect time: | Wallet.tsx |
| LDK race symptom | Node not initialized | native modules / LdkNodeUtils.ts |
| Transient-retry loop | Transient ... - retry N/5 in Xms | LndMobileErrors.ts |
Kill the app (swipe from recents), relaunch, wait for balance screen.
triggering line, typically followed by one skipping line (the startup focus event fires twice — see c71a04ce9), the backend's start milestones once each, exactly one connect time: line.triggering lines with no skipping → the _navigating guard regressed or a bypass path fired → branch to Phase 2.connect time: never appears and the spinner persists → note which milestone was last; fetchLock is likely stranded (Phase 2's flag table tells you which return path).Settings → Wallets → tap the other wallet; as soon as the Wallet screen shows anything (don't wait for full sync), switch back. Record: crashes (logcat -b crash), stuck-spinner incidents, and whether the OLD node keeps running.
tail -f on the old wallet's lnd.log — if new log lines keep appearing (neutrino banner/filter activity), the LND daemon is still running alongside LDK. Expected per code read: it is. Symmetrically, ldk_node.log continues after switching ldk→embedded.triggering → store resets → old-backend stop milestone only if switching within the same implementation → new-backend start milestones → connect time:.Fatal signal 6 (SIGABRT) mentioning tokio/runtime → LDK ref released on a Tokio thread; capture the full tombstone; this is the d416052cd class — branch to Phase 5.Node not initialized surfaces in the UI (error screen) rather than being retried → an unguarded call site hit W2 — branch to Phase 2 call-site census.triggering line at all → focus fired while _navigating was stuck true (a getSettingsAndNavigate that never resolved) — capture the last milestone.While the node is mid-sync (do this within ~30s of connect), Settings → Wallets → gear icon → Delete wallet → confirm; then immediately select the remaining wallet.
justDeletedWallet flow (2s pause in Wallets.handleJustDeletedWallet before reconnect), new wallet connects.d416052cd regression — branch to Phase 5.LND_FOLDER_MISSING alert on the NEXT connect → stale lndDir/ldkNodeDir read from a not-yet-propagated settings write — record selectedNode at each step (Phase 1 instrumentation).Rapidly background/foreground the app (Android: home + reopen; iOS: app switcher). Then once more with ≥30s in background.
skipping lines for suppressed re-entries; at most one connect flow per foregrounding; NWC store reset lines on iOS inactive are normal.connecting got stuck true, making every fetchData run the full reset branch.With embedded-lnd running: dev-menu Reload (debug) or trigger the in-app restart. The process survives, so the Go daemon does too.
LND already started - force stop (Go thinks running) → stop ladder → successful restart within ~10s (includes the 4s ANDROID_PROCESS_CLEANUP_DELAY_MS).LND still running (attempt N) - force stop until LND_START_FAILED → the stop ladder can't actually kill the daemon (consistent with W6: killLnd is a no-op on this manifest) → Phase 4.Baseline record: for each experiment, crashes / trials, hangs / trials, and the last milestone before each failure. This table is the campaign's ground truth.
Add TEMPORARY logging so every flag transition is attributable. Instrumentation must not ship: strip it before any PR, or land it only behind an explicit maintainer-approved debug gate — unconditional console noise in the connect path will be rejected in review (zeus-change-control; there is no developer-mode toggle to hide behind, see zeus-config-and-flags).
Instrumentation points (grep anchors, all verified):
| Where | Anchor | Log |
|---|---|---|
stores/SettingsStore.ts setConnectingStatus | public setConnectingStatus = (status = false) | old→new value + new Error().stack (caller attribution — MobX gives you none) |
views/Wallet/Wallet.tsx fetchData entry | if (fetchLock) return; | log BOTH the rejected entries (currently silent) and the acquisitions, with implementation, connecting, walletJustCreated |
views/Wallet/Wallet.tsx fetchData exit | SettingsStore.fetchLock = false; | matched release log; any acquisition without a release = a stranded lock |
utils/LndMobileUtils.ts stopLnd/startLnd | function heads | entry/exit + embeddedLndStarted value |
utils/LdkNodeUtils.ts stopLdkNode/startLdkNodeWallet | function heads | entry/exit timestamps (measures overlap with builds) |
views/Settings/Wallets.tsx onWalletPress | const onWalletPress | old/new selectedNode, old/new implementation |
Re-run E1–E5 with instrumentation. Decision gate: produce a timeline for every Phase 0 failure showing which flag was in the wrong state. If any failure shows two node-start sequences interleaved (start A begins, start B begins before A's ready milestone), the serialization defect is confirmed and Phase 2 is mandatory.
Enumerate every path into getSettingsAndNavigate/fetchData and every exit's flag outcome. Starting set (verified at c5fd094fb — re-derive, don't trust):
grep -n "getSettingsAndNavigate\|_navigating" views/Wallet/Wallet.tsx
grep -rn "setConnectingStatus(true)" --include="*.ts" --include="*.tsx" . | grep -v node_modules
Entry paths: focus (guarded), AppState active (guarded), componentDidMount→handleFocus, pull-to-refresh (UNguarded), error-screen Restart (UNguarded, iOS), zeusln://share URL (UNguarded), transient-retry recursion (retryOnTransientError → getSettingsAndNavigate(undefined, count+1)).
For each return/throw inside fetchData, record: fetchLock left true? connecting left true? Is there a UI affordance that will call setConnectingStatus(true) to recover? Build the full table (there are ~12 early returns).
Decision gate:
Experiment: E2's old-node liveness probe, plus resource measurement:
# Memory of the app process while both nodes run vs one:
adb shell dumpsys meminfo app.zeusln.zeus | head -30
Also test: switch embedded→ldk, wait 5 min, switch back — does the returning stopLnd()+startLnd() find a healthy daemon or a wedged one?
Decision gate: if the old node demonstrably keeps syncing after a switch (expected), classify severity: (i) resource-only → fold the fix into solution (a)/(c); (ii) causes E2 crashes or db contention → promote solution (c) to first.
killLnd's return value during E5. Expected per code read: always false (no :zeusLndMobile process exists to kill). If confirmed, the "force stop" ladder is stopDaemon-only, and the manifest/no-process finding should be written into the fix design.LndMobileScheduledSyncWorker interaction: enable persistent mode / scheduled sync, background the app, watch for a worker-initiated MSG_START_LND colliding with a foreground start.android:process), or keep in-process and delete the dead killLnd scan + fix the thread.run() FIXME.Decision gate: only proceed to native changes if Phase 0/1 shows failures that JS-level serialization cannot prevent (e.g. daemon survives all stop attempts). Native service changes require full re-test of: background receive, persistent mode notification, scheduled sync, graceful-stop notification action.
buildNode/stop is in flight (grep -rn "LdkNode\." stores/ backends/ utils/ --include="*.ts"); classify each as tolerant (retries on Node not initialized) or crash/error-surfacing.ldk_node.log and the crash buffer. If unlink-during-cleanup errors appear, the fix is to move the delete into native (delete after the Arc-hold) or extend the JS wait — with a real completion signal, NOT a longer sleep.Decision gate: any crash here re-opens the d416052cd/0e2b89164 contract — fix in the native module (both platforms symmetrically; the Kotlin and Swift implementations are deliberate mirrors), never by adding JS-side sleeps.
Ranking reflects maintainer culture: minimal diffs first, big designs only with prior sign-off. Every option below is funds-adjacent (a node torn down mid-payment can strand an HTLC or force-close exposure), so all inherit: no drive-by refactors in payment paths, manual 2-platform testing by the author, revert-first if release testing regresses (zeus-change-control).
fetchData lock-hygiene + serialization audit (smallest, do first)Make lock acquire/release provably paired: single exit point (try/finally) for fetchLock; extend the _navigating guard (or fold it into one SettingsStore-owned in-flight promise) to cover the three bypass call sites.
setConnectingStatus(true) clearing the lock — don't break that contract silently).Before starting backend B, stop whichever embedded node type is actually running — e.g. track "last started native node" (type + dir) in SettingsStore and stop it in fetchData's connect branch regardless of the NEW implementation, or stop-old in onWalletPress before updateSettings.
stopLnd() is safe when LND isn't running (it is — status-check path returns early) and stopLdkNode() likewise (native stop resolves ok on null node); measure added switch latency; prove the delete flow (which stops nodes itself) doesn't double-stop into an error state.Give delete a real completion signal (native-side delete after Arc-hold, or a stopAndAwaitCleanup FFI method); make non-tolerant FFI call sites tolerant or gate them on a ready flag.
LdkNodeModule.kt + LdkNodeModule.swift (must change in lockstep) + utils/LdkNodeUtils.ts. Native binary/bindings untouched (pure module-layer change) — if the ldk-node fork itself must change, that's a fetch-libraries-versions.json bump with its own process (zeus-build-and-env).One owner (e.g. a NodeLifecycleStore / manager) with explicit states — STOPPED, STARTING, RUNNING, STOPPING, DELETING, SWITCHING — through which ALL start/stop/delete/switch flows pass; Wallet.tsx becomes a consumer, not an owner.
start has exactly one matching stop on every path including error paths and app-kill (invariant: at most one native node per implementation, and after the W3 fix, at most one total); (iii) mapping of ALL existing flags (connecting, fetchLock, embeddedLndStarted, walletJustCreated, _navigating) onto machine states, with a staged migration so each PR stays reviewable; (iv) re-entrancy story: transitions are serialized on one async queue, concurrent requests coalesce or queue, never interleave.fdad118ed — land it as a series of minimal PRs, each independently revertible.Fix/delete dead killLnd, resolve the thread.run() FIXME, decide the process model, fix or remove the broken unbindLndMobileService.
fdad118ed (2025-10-24) refactored Wallet's React lifecycle methods and caused infinite loading, fixed only three weeks later by 93227029e. Fence: no mechanical lifecycle refactors of Wallet.tsx; behavior-preserving proof + full Phase 0 suite or don't touch it.d416052cd + 0e2b89164 (2026-03-10). Fence: the four-step native contract above is load-bearing; any "simplification" that drops the dedicated thread, the blocking resolve, or the 2s hold re-introduces a hard native crash.b54e8f138 (2026-05-17), startup had no ready-wait, so "Node not initialized" from post-start FFI calls could surface as a fatal error (the node.start() retry itself predates that commit). Fence: during rebuild windows, retry-with-classification (specific error strings, bounded retries) — but do NOT blanket-retry all FFI errors; only the two sentinels in utils/LdkNodeUtils.ts.sleep(2000) in Wallets.handleJustDeletedWallet, ANDROID_PROCESS_CLEANUP_DELAY_MS = 4000, the 2s Arc-holds, STATE_SUBSCRIPTION_SETTLE_MS). Fence: adding another timeout to make a race "go away" is symptom-patching; every new wait must be a real completion signal (status poll, event, promise) or it will be NACKed. Existing sleeps may only be removed by a change that replaces them with an owned signal.c71a04ce9 shows the right shape — the commit body names the exact double-fire mechanism (focus fires twice with initialLoad=true) before adding the guard. Fence: no guard/lock may be added without a written statement of which concrete interleaving it prevents.A lifecycle fix is promotable when ALL of the following hold:
triggering per user action (plus skipping lines); one connect time: per connect; zero UI-surfaced Node not initialized; zero Fatal signal in adb logcat -b crash; after a cross-implementation switch, the old node's log file goes quiet (post W3 fix).fetchLock acquisition has a matching release in every experiment (then strip the instrumentation).yarn verify green (Test, Lint, Prettier, Typescript Check — see zeus-validation-and-qa). Note stores/views have zero jest coverage; your manual evidence IS the test.fetch-libraries-versions.json (zeus-build-and-env).Facts verified 2026-07-06 against master c5fd094fb (v13.1.3-alpha) by direct code read; commit hashes verified with git log. The maintainer priority statement dates from 2026-07-06. Runtime numbers (crash rates) are deliberately absent — Phase 0 produces them.
One-line re-verification commands (run from repo root):
git log -1 --format="%h %s" # still c5fd094fb?
grep -n "_navigating" views/Wallet/Wallet.tsx # re-entry guard still present
grep -n "if (fetchLock) return" views/Wallet/Wallet.tsx # lock acquire site
grep -n "remove fetchLock on reconnect" stores/SettingsStore.ts # lock cleared in setConnectingStatus(true)
grep -n "connecting = true" stores/SettingsStore.ts # connecting initialized true
grep -n "walletJustCreated" views/Wallet/Wallet.tsx | head -3 # skip-start flag still consumed
grep -n "stopLdkNode()" views/Wallet/Wallet.tsx # ldk branch stops only ldk (W3)
grep -rn "stopLnd" views/Wallet/Wallet.tsx | head -3 # lnd branch stops only lnd (W3)
grep -n "thread.run()" android/app/src/main/java/com/zeus/LndMobileScheduledSyncWorker.java # W7 FIXME
grep -n "zeusLndMobile" android/app/src/main/java/com/zeus/LndMobileTools.java # W6 dead scan
grep -n "android:process" android/app/src/main/AndroidManifest.xml # expect NO output (in-process LND)
grep -n "looks broken" lndmobile/LndMobile.d.ts # W8 unbind still flagged broken
grep -n "Thread.sleep(2000)" android/app/src/main/java/app/zeusln/zeus/LdkNodeModule.kt # 2s Arc-hold
grep -n "withExtendedLifetime" ios/LdkNodeMobile/LdkNodeModule.swift # iOS mirror of the hold
grep -n "LDK_NODE_NOT_INITIALIZED" utils/LdkNodeUtils.ts # tolerated sentinel
grep -n "MAX_TRANSIENT_RPC_RETRIES\|TRANSIENT_RPC_RETRY_BASE_MS" utils/LndMobileErrors.ts
grep -n views/Settings/Wallets.tsx
git -1 --format= --=short c71a04ce9 d416052cd 0e2b89164 b54e8f138 8fc9698a0 fdad118ed 93227029e
If any of these greps come back empty or changed, the corresponding claim above is stale — re-read the file before acting on it.
utils/WalletCreationUtils.tsviews/Settings/SeedRecovery.tsxwrite settings.nodes/selectedNode via updateSettings, then setConnectingStatus(true) + popTo('Wallet') |
startLndfalsestopLndnodeLock (Kotlin synchronized) / takeNode() (Swift) makes stop/build swap the node reference atomically; LND Android: a single lndStarted boolean inside LndMobileService; the Go daemon itself rejects double starts with "lnd already started".LND_ALREADY_RUNNINGstartLndWithRetry| Handled but fragile; see W6 |
| W6 | LndMobileTools.killLnd() scans running processes for packageName + ":zeusLndMobile", but no such process is ever created (no android:process in the manifest) — so killLnd can never find/kill anything and resolves false. Recovery therefore rests entirely on the Go stopDaemon call. Also killLndProcess() uses getCurrentActivity() which is null when no Activity is attached (background) → NPE risk | Code-read finding; runtime confirmation = log killLnd result in Phase 4. Blixt heritage: Blixt runs the service in a separate process, Zeus does not |
| W7 | LndMobileScheduledSyncWorker.java (background sync) binds the same LndMobileService and — per its own FIXME at ~line 310 — calls thread.run() instead of thread.start(), executing the "thread" body synchronously on the worker's calling thread with handlers on the main looper. A scheduled sync racing a foreground app start can send MSG_START_LND against an already-started daemon (the worker's own comment: "Make sure we don't attempt to start lnd twice... better fix would be MSG_CHECKSTATUS") | Open, acknowledged debt (TODO(hsjoberg)) |
| W8 | LndMobileService.onUnbind: when the last client unbinds it calls stopLnd(null, -1) + stopSelf() — but the RN module (LndMobile.java) binds once in initialize() and its unbindLndMobileService is flagged in lndmobile/LndMobile.d.ts as "function looks broken" and is never called from TS (verified: zero TS call sites). So service lifetime ≈ process lifetime; unbind-triggered stop is effectively dead code in the app (only the sync worker unbinds) | Confirmed by code read |
| W9 | Background/foreground storm: AppState active triggers getSettingsAndNavigate (guarded by _navigating), inactive resets NostrWalletConnectStore; iOS fires inactive on mere app-switcher/notification-shade interactions | Guarded but untested under storm; Phase 0 E4 |
| W10 | LDK delete-then-create same dir: stop() resolves before the 2s native Arc-hold expires; deleteLdkNodeWallet then RNFS.unlinks the storage dir while background Tokio cleanup may still touch it | HYPOTHESIS — the 2s hold was added for the ref-drop panic, not for file-handle safety; needs Phase 5 experiment |