OpenClaw merged a P1 Gateway reliability fix tonight with PR #165107, "fix(gateway): small agents stay unavailable until larger agents finish startup preparation". The patch targets a painful restart behavior: small agents could remain unavailable after Gateway restart because earlier, larger agents were still working through startup preparation.
The user-facing symptom was blunt. Sessions, cron jobs, and chat history for a small agent could report UNAVAILABLE even after that agent's own database had opened cleanly.
The Startup Bottleneck
Before this change, deferred startup recovery put each agent's full preparation into one FIFO path. That included database opening, session migration, credential refresh, model refresh, and final admission. If an early agent had a large store or slow integrity work, later small agents waited silently.
The PR describes a production incident on OpenClaw 2026.9.8 with three agents under high host load. All three writable database opens finished within about six minutes, including a 1.69 GB store that took 44 seconds for full integrity checking. But the recoveries completed one after another at roughly 22, 41, and 42 minutes. The final agent had a 12 MB store and only needed 64 seconds of its own preparation.
Most of that wait was queueing, not actual work for the small agent.
What Changed
OpenClaw now separates agent-local startup work from global publication work. Database opening and session migration can run per agent, while credential and model publication still keep the serialized turn because they depend on shared secrets and admitted auth state.
That means each deferred agent can become available as soon as its own database and sessions are ready. A large agent can still take time, but it no longer keeps unrelated smaller agents hostage.
The operator experience also improves:
- Per-agent phase durations are logged on recovery and failure
- A progress warning appears every 60 seconds while recovery is pending
- Existing authority checks remain in place for config identity, secrets revision, auth store, prepared model, and deletion journal
Why It Matters
OpenClaw is increasingly used with multiple agents, persistent stores, scheduled jobs, and background sessions. Restart behavior is therefore not just a boot-time detail; it determines whether automations and chat histories come back predictably after upgrades, host pressure, or maintenance.
This fix is especially useful for mixed fleets where one agent has a very large history and another is lightweight but time-sensitive.
Verification
The PR reports regression coverage for the slow-FIFO case and preserves the existing concurrency limits: two concurrent opens, up to four concurrent session migrations, and serialized final admission. The important outcome is practical: smaller agents can resume service independently instead of waiting for the slowest database in the line.
For operators running multi-agent Gateway setups, PR #165107 is the nightly fix to watch.