feat(server): write back and flush the world on a graceful stop #8

Merged
Serkyo merged 2 commits from feat/graceful-stop into dev 2026-09-25 23:55:05 +00:00
Owner

What does this PR do?

I gave the server a stop path that saves the world before exiting.

  • ServerWorld::shutdown(self) in crates/server/src/world_server.rs tears the world down in channel order. It evicts every resident chunk through reconcile(&HashSet::new()) so each one queues its Write or Remove, drops job_tx, joins every generation worker, sends a final Flush and waits for the reply, then drops its save sender and joins the save actor. Workers are joined while the actor is still alive because a worker mid-load is blocked on a Read reply.
  • SaveActor::shutdown(self) in crates/server/src/save/region_actor.rs drops the actor's own sender and joins its thread. A panicked worker or actor thread is logged with warn! and the teardown continues.
  • I pulled the flush round trip into a private request_flush helper so flush and shutdown share it.
  • main installs a ctrlc handler (with the termination feature, so SIGTERM as well as SIGINT) before the loading phase. The handler only swaps an Arc<AtomicBool>; a second signal while the flag is already set exits with status 130.
  • run_simulation now checks the flag at the top of each tick and returns instead of looping forever. A new shut_down function in main.rs then drops the NetworkServer, removes ServerWorld from the ECS world, calls shutdown, and logs the elapsed time. A failed final flush makes the process exit with an error.
  • I removed the #[expect(dead_code)] placeholders on SaveActor.handle, ServerWorld.save_actor and ServerWorld.workers, along with their TODO.
  • docs/save_format.md now states that a graceful stop writes back and flushes, and that a kill without a signal still loses up to one autosave interval.

Why is this change necessary?

The server had no way to stop. run_simulation returned ! and a signal killed the process wherever it landed. A chunk evicted since the last autosave had its diff only in the save actor's in-memory region image, and resident chunks that were never evicted had never been diffed at all, so both were lost on every exit. The 45 second autosave narrowed that window without closing it.

Scope of Changes

  • server
  • workspace

Testing

Automated:

  • cargo check --workspace --all-targets: passes.
  • cargo clippy --all-targets --all-features -- -D warnings: passes.
  • cargo fmt --all -- --check: passes.
  • cargo test -p server: 56 passed, 0 failed.
  • cargo deny check: advisories, bans, licenses and sources all ok with ctrlc added.

Two new tests in crates/server/src/tests/world_server.rs:

  • shutdown_writes_back_resident_edits edits a resident chunk, calls shutdown with no manual reconcile or flush, and checks that a fresh world over the same directory streams the edit back.
  • shutdown_with_loads_in_flight_returns dispatches a radius 4 cylinder and shuts down without draining, which proves the join order cannot deadlock while workers are mid-load.

Manual:

  • Started the server and sent SIGINT once it was listening: logged "Stop requested; saving the world" then "Server stopped cleanly" after about 6.5 ms, exit status 0.
  • Same with SIGTERM: same log lines, exit status 0.
  • Two SIGINTs 10 ms apart: exited at once with status 130 after the stop request line.
  • With one client connected and 13549 chunks resident after a minute of streaming, Ctrl-C logged the stop request, the network accept loop shutting down, and "Server stopped cleanly" after 690 ms.

Additional Context

This adds ctrlc 3.5.2 to the server, which brings in nix on Unix and two macOS-only crates (dispatch2, block2 0.6). The workspace already had block2 0.5.1 through the client, so the lockfile now carries two versions of it; cargo deny accepts that.

Connected clients still see a server stop as an idle timeout. NetworkServer's drop path does not send Disconnect to open sessions yet; that is a separate change in net.

None.

Checklist

  • I have branched from dev (or a feature branch off dev) and my PR targets dev.
  • I have kept my changes focused to a single concept.
  • I have added or updated documentation (/// doc comments for Rust) where necessary.
  • I have tested my changes and described any relevant automated or manual testing above.
### What does this PR do? I gave the server a stop path that saves the world before exiting. - `ServerWorld::shutdown(self)` in `crates/server/src/world_server.rs` tears the world down in channel order. It evicts every resident chunk through `reconcile(&HashSet::new())` so each one queues its `Write` or `Remove`, drops `job_tx`, joins every generation worker, sends a final `Flush` and waits for the reply, then drops its save sender and joins the save actor. Workers are joined while the actor is still alive because a worker mid-load is blocked on a `Read` reply. - `SaveActor::shutdown(self)` in `crates/server/src/save/region_actor.rs` drops the actor's own sender and joins its thread. A panicked worker or actor thread is logged with `warn!` and the teardown continues. - I pulled the flush round trip into a private `request_flush` helper so `flush` and `shutdown` share it. - `main` installs a `ctrlc` handler (with the `termination` feature, so SIGTERM as well as SIGINT) before the loading phase. The handler only swaps an `Arc<AtomicBool>`; a second signal while the flag is already set exits with status 130. - `run_simulation` now checks the flag at the top of each tick and returns instead of looping forever. A new `shut_down` function in `main.rs` then drops the `NetworkServer`, removes `ServerWorld` from the ECS world, calls `shutdown`, and logs the elapsed time. A failed final flush makes the process exit with an error. - I removed the `#[expect(dead_code)]` placeholders on `SaveActor.handle`, `ServerWorld.save_actor` and `ServerWorld.workers`, along with their TODO. - `docs/save_format.md` now states that a graceful stop writes back and flushes, and that a kill without a signal still loses up to one autosave interval. ### Why is this change necessary? The server had no way to stop. `run_simulation` returned `!` and a signal killed the process wherever it landed. A chunk evicted since the last autosave had its diff only in the save actor's in-memory region image, and resident chunks that were never evicted had never been diffed at all, so both were lost on every exit. The 45 second autosave narrowed that window without closing it. ### Scope of Changes - server - workspace ### Testing Automated: - `cargo check --workspace --all-targets`: passes. - `cargo clippy --all-targets --all-features -- -D warnings`: passes. - `cargo fmt --all -- --check`: passes. - `cargo test -p server`: 56 passed, 0 failed. - `cargo deny check`: advisories, bans, licenses and sources all ok with `ctrlc` added. Two new tests in `crates/server/src/tests/world_server.rs`: - `shutdown_writes_back_resident_edits` edits a resident chunk, calls `shutdown` with no manual reconcile or flush, and checks that a fresh world over the same directory streams the edit back. - `shutdown_with_loads_in_flight_returns` dispatches a radius 4 cylinder and shuts down without draining, which proves the join order cannot deadlock while workers are mid-load. Manual: - Started the server and sent SIGINT once it was listening: logged "Stop requested; saving the world" then "Server stopped cleanly" after about 6.5 ms, exit status 0. - Same with SIGTERM: same log lines, exit status 0. - Two SIGINTs 10 ms apart: exited at once with status 130 after the stop request line. - With one client connected and 13549 chunks resident after a minute of streaming, Ctrl-C logged the stop request, the network accept loop shutting down, and "Server stopped cleanly" after 690 ms. ### Additional Context This adds `ctrlc` 3.5.2 to the server, which brings in `nix` on Unix and two macOS-only crates (`dispatch2`, `block2` 0.6). The workspace already had `block2` 0.5.1 through the client, so the lockfile now carries two versions of it; `cargo deny` accepts that. Connected clients still see a server stop as an idle timeout. `NetworkServer`'s drop path does not send `Disconnect` to open sessions yet; that is a separate change in `net`. ### Related Issues None. ### Checklist - [x] I have branched from `dev` (or a feature branch off `dev`) and my PR targets `dev`. - [x] I have kept my changes focused to a single concept. - [x] I have added or updated documentation (`///` doc comments for Rust) where necessary. - [x] I have tested my changes and described any relevant automated or manual testing above.
feat(server): stop the simulation loop on SIGINT and SIGTERM
Some checks failed
CI / Rust Check, Lint & Test (pull_request) Has been cancelled
CI / Dependency Licenses & Advisories (pull_request) Has been cancelled
CI / Lua Lint & Format (pull_request) Has been cancelled
CI / Commit Message Lint (pull_request) Has been cancelled
CI / LFS Pointer Guard (pull_request) Has been cancelled
Auto Labeler / label-scope (pull_request_target) Successful in 2s
CLA Signed All authors have signed the CLA.
CLA Check / cla-check (pull_request_target) Successful in 2s
fcdfcb9ffd
Serkyo deleted branch feat/graceful-stop 2026-09-25 23:55:05 +00:00
Sign in to join this conversation.
No reviewers
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set

Reference
Synvael/synvael!8
No description provided.