Skip to content

Discard dirty guest memory at exit - #415

Open
henrybear327 wants to merge 3 commits into
sysprog21:mainfrom
henrybear327:slab/exit-truncate
Open

henrybear327 wants to merge 3 commits into
sysprog21:mainfrom
henrybear327:slab/exit-truncate

Conversation

@henrybear327

@henrybear327 henrybear327 commented Oct 6, 2026 •

Copy link
Copy Markdown
Collaborator

Currently, the disk IO is very heavy when gvisor / LTP conformance test is run. On local development machine, it's easy to see around 1TB of accumulated write to SSD.

The main issue lies in the flush mechanism of our file-backed memory file - when a process terminates, the file is flushed, though no one would be reading it anymore.

Experiments

For workloads that isn't fork-heavy, we can see improvements.

Workload main branch Wall time
gVisor readv_test 2056.3 MiB 12.2 MiB 1.15 s -> 0.41 s
gVisor close_range_test 832.9 MiB 0.4 MiB 2.01 s -> 0.10 s

Summary by cubic

Discards a guest's dirty RAM at process exit instead of writing it back to disk, eliminating a major source of disk I/O (roughly 1 TB of writes on the worst local workloads; e.g. gVisor readv_test dropped from 2056 MiB to 12 MiB written).

guest_destroy now truncates the unlinked slab file before its last close. The truncate is skipped when fclonefileat failed and a fork child received the live slab fd (fallback path, marked via shm_exported), because truncating under that child's private mapping would stall the host.

  • Adds test-slab-exit-writes, which reads the process's block-write count from its zombie to confirm dirtying 256 MiB writes nothing at exit, with a 32 MiB fsync control.
  • Adds test-slab-exit-fork, which runs both the APFS clone path and the forced live-fd fallback, verifying a fork child still reads its memory after the parent exits and that the truncate is skipped only in the fallback; probe-disk-writes records the counter.
  • Documents the truncate and fallback in docs/internals.md, including ELFUSE_DISABLE_FORK_CLONEFILE=1 to force the fallback on APFS.

Written for commit 5a3ab6f. Summary will update on new commits.

Review in cubic

@henrybear327
henrybear327 requested a review from jserv October 6, 2026 14:49
@henrybear327 henrybear327 self-assigned this Oct 6, 2026
cubic-dev-ai[bot]

This comment was marked as resolved.

Guest RAM uses an unlinked file mapped MAP_SHARED. Closing that file
writes dirty pages to disk even when no process needs them. Truncate
it during normal guest_destroy teardown to discard those pages.

A native fork can hand the live slab fd to its child when fclonefileat
fails. Truncating beneath that child's MAP_PRIVATE mapping can stall
the host, so skip the truncate once the slab has been exported. Set
shm_exported only after admission succeeds: a failed fork must leave
the previous export state intact. Live worker vCPUs defer teardown to
process exit.

Check exit writes against an fsync control and verify that a clone
child retains its memory after the parent exits. Wait for the child's
result before skipping the fallback lane, and let removal of the
scratch directory cancel its wait. Prebuilt bundles without the
fixture skip both lanes.

Validation: make check-format, make check, and the AArch64 matrix
(297 passed, 0 failed, 10 skipped), plus failed-admission fault
injection and the slab exit lanes.
Comment thread mk/tests.mk Outdated
Comment thread mk/tests.mk
Require the orphan's memory check to succeed before inspecting the fork
path. A live-fd fallback must also confirm that teardown skipped slab
truncation; treating that path as a skip hides corrupt child memory.

The clone lane passes, and six simulated verdict cases reject missing
results, corrupt bytes, and inconsistent truncation logs.
Add ELFUSE_DISABLE_FORK_CLONEFILE=1 so the orphan lane can force the
live-fd fallback on APFS. Require intact child memory and the teardown
skip log in that run, alongside the normal clone attempt.

The switch follows the existing clone-failure branch, so Rosetta keeps
its region-copy fallback. Document both paths and the diagnostic switch.

The slab lanes, make check, native aarch64 matrix, and format checks
pass. A forced Rosetta fork confirms region-copy selection. The Rosetta
signal audit hits the same xstate assertion on the original PR head.
@henrybear327
henrybear327 requested a review from jserv October 6, 2026 21:45

@cubic-dev-ai cubic-dev-ai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

1 issue found across 3 files (changes from recent commits).

Prompt for AI agents (unresolved issues)

Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.


<file name="src/runtime/forkipc.c">

<violation number="1" location="src/runtime/forkipc.c:1842">
P3: This new branch routes a deliberately disabled clone through the error path, so on Rosetta the pre-existing handler logs "clone: rosetta CoW snapshot via fclonefileat failed (Operation not supported)" even though fclonefileat was never attempted, and native guests log a fabricated ENOTSUP. Emit the diagnostic before faking the failure, or use a distinct message for the disabled case, so the logs report the real cause.</violation>
</file>

Tip: Review your code locally with the cubic CLI to iterate faster.

Turn on auto-fix | Re-trigger cubic

Comment thread src/runtime/forkipc.c
const char *disable_clone = getenv("ELFUSE_DISABLE_FORK_CLONEFILE");
if (disable_clone && strcmp(disable_clone, "1") == 0) {
snapshot_shm_fd = -1;
errno = ENOTSUP;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P3: This new branch routes a deliberately disabled clone through the error path, so on Rosetta the pre-existing handler logs "clone: rosetta CoW snapshot via fclonefileat failed (Operation not supported)" even though fclonefileat was never attempted, and native guests log a fabricated ENOTSUP. Emit the diagnostic before faking the failure, or use a distinct message for the disabled case, so the logs report the real cause.

Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/runtime/forkipc.c, line 1842:

<comment>This new branch routes a deliberately disabled clone through the error path, so on Rosetta the pre-existing handler logs "clone: rosetta CoW snapshot via fclonefileat failed (Operation not supported)" even though fclonefileat was never attempted, and native guests log a fabricated ENOTSUP. Emit the diagnostic before faking the failure, or use a distinct message for the disabled case, so the logs report the real cause.</comment>

<file context>
@@ -1834,8 +1834,15 @@ int64_t sys_clone(hv_vcpu_t vcpu,
+        const char *disable_clone = getenv("ELFUSE_DISABLE_FORK_CLONEFILE");
+        if (disable_clone && strcmp(disable_clone, "1") == 0) {
+            snapshot_shm_fd = -1;
+            errno = ENOTSUP;
+        } else {
+            snapshot_shm_fd = fork_snapshot_shm_via_clonefile(g->shm_fd);
</file context>

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants