Repository navigation
Discard dirty guest memory at exit - #415
Open
henrybear327 wants to merge 3 commits into
Open
henrybear327 wants to merge 3 commits into
henrybear327 wants to merge 3 commits into
Conversation
henrybear327
force-pushed
the
slab/exit-truncate
branch
from
October 6, 2026 14:49
d37720f to
9c74257
Compare
henrybear327
force-pushed
the
slab/exit-truncate
branch
from
October 6, 2026 15:10
9c74257 to
4ce4562
Compare
Guest RAM uses an unlinked file mapped MAP_SHARED. Closing that file writes dirty pages to disk even when no process needs them. Truncate it during normal guest_destroy teardown to discard those pages. A native fork can hand the live slab fd to its child when fclonefileat fails. Truncating beneath that child's MAP_PRIVATE mapping can stall the host, so skip the truncate once the slab has been exported. Set shm_exported only after admission succeeds: a failed fork must leave the previous export state intact. Live worker vCPUs defer teardown to process exit. Check exit writes against an fsync control and verify that a clone child retains its memory after the parent exits. Wait for the child's result before skipping the fallback lane, and let removal of the scratch directory cancel its wait. Prebuilt bundles without the fixture skip both lanes. Validation: make check-format, make check, and the AArch64 matrix (297 passed, 0 failed, 10 skipped), plus failed-admission fault injection and the slab exit lanes.
henrybear327
force-pushed
the
slab/exit-truncate
branch
from
October 6, 2026 15:50
4ce4562 to
b23b46e
Compare
jserv
reviewed
Oct 6, 2026
jserv
reviewed
Oct 6, 2026
Require the orphan's memory check to succeed before inspecting the fork path. A live-fd fallback must also confirm that teardown skipped slab truncation; treating that path as a skip hides corrupt child memory. The clone lane passes, and six simulated verdict cases reject missing results, corrupt bytes, and inconsistent truncation logs.
Add ELFUSE_DISABLE_FORK_CLONEFILE=1 so the orphan lane can force the live-fd fallback on APFS. Require intact child memory and the teardown skip log in that run, alongside the normal clone attempt. The switch follows the existing clone-failure branch, so Rosetta keeps its region-copy fallback. Document both paths and the diagnostic switch. The slab lanes, make check, native aarch64 matrix, and format checks pass. A forced Rosetta fork confirms region-copy selection. The Rosetta signal audit hits the same xstate assertion on the original PR head.
There was a problem hiding this comment.
1 issue found across 3 files (changes from recent commits).
Prompt for AI agents (unresolved issues)
Check if these issues are valid — if so, understand the root cause of each and fix them. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. If appropriate, use sub-agents to investigate and fix each issue separately.
<file name="src/runtime/forkipc.c">
<violation number="1" location="src/runtime/forkipc.c:1842">
P3: This new branch routes a deliberately disabled clone through the error path, so on Rosetta the pre-existing handler logs "clone: rosetta CoW snapshot via fclonefileat failed (Operation not supported)" even though fclonefileat was never attempted, and native guests log a fabricated ENOTSUP. Emit the diagnostic before faking the failure, or use a distinct message for the disabled case, so the logs report the real cause.</violation>
</file>
Tip: Review your code locally with the cubic CLI to iterate faster.
Turn on auto-fix | Re-trigger cubic
| const char *disable_clone = getenv("ELFUSE_DISABLE_FORK_CLONEFILE"); | ||
| if (disable_clone && strcmp(disable_clone, "1") == 0) { | ||
| snapshot_shm_fd = -1; | ||
| errno = ENOTSUP; |
There was a problem hiding this comment.
P3: This new branch routes a deliberately disabled clone through the error path, so on Rosetta the pre-existing handler logs "clone: rosetta CoW snapshot via fclonefileat failed (Operation not supported)" even though fclonefileat was never attempted, and native guests log a fabricated ENOTSUP. Emit the diagnostic before faking the failure, or use a distinct message for the disabled case, so the logs report the real cause.
Prompt for AI agents
Check if this issue is valid — if so, understand the root cause and fix it. When an issue isn't valid or won't be fixed in this PR, reply in its thread with the reason and then resolve the thread. At src/runtime/forkipc.c, line 1842:
<comment>This new branch routes a deliberately disabled clone through the error path, so on Rosetta the pre-existing handler logs "clone: rosetta CoW snapshot via fclonefileat failed (Operation not supported)" even though fclonefileat was never attempted, and native guests log a fabricated ENOTSUP. Emit the diagnostic before faking the failure, or use a distinct message for the disabled case, so the logs report the real cause.</comment>
<file context>
@@ -1834,8 +1834,15 @@ int64_t sys_clone(hv_vcpu_t vcpu,
+ const char *disable_clone = getenv("ELFUSE_DISABLE_FORK_CLONEFILE");
+ if (disable_clone && strcmp(disable_clone, "1") == 0) {
+ snapshot_shm_fd = -1;
+ errno = ENOTSUP;
+ } else {
+ snapshot_shm_fd = fork_snapshot_shm_via_clonefile(g->shm_fd);
</file context>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Currently, the disk IO is very heavy when gvisor / LTP conformance test is run. On local development machine, it's easy to see around 1TB of accumulated write to SSD.
The main issue lies in the flush mechanism of our file-backed memory file - when a process terminates, the file is flushed, though no one would be reading it anymore.
Experiments
For workloads that isn't fork-heavy, we can see improvements.
readv_testclose_range_testSummary by cubic
Discards a guest's dirty RAM at process exit instead of writing it back to disk, eliminating a major source of disk I/O (roughly 1 TB of writes on the worst local workloads; e.g. gVisor
readv_testdropped from 2056 MiB to 12 MiB written).guest_destroynow truncates the unlinked slab file before its last close. The truncate is skipped whenfclonefileatfailed and a fork child received the live slab fd (fallback path, marked viashm_exported), because truncating under that child's private mapping would stall the host.test-slab-exit-writes, which reads the process's block-write count from its zombie to confirm dirtying 256 MiB writes nothing at exit, with a 32 MiB fsync control.test-slab-exit-fork, which runs both the APFS clone path and the forced live-fd fallback, verifying a fork child still reads its memory after the parent exits and that the truncate is skipped only in the fallback;probe-disk-writesrecords the counter.docs/internals.md, includingELFUSE_DISABLE_FORK_CLONEFILE=1to force the fallback on APFS.Written for commit 5a3ab6f. Summary will update on new commits.