-
Notifications
You must be signed in to change notification settings - Fork 5
Document that encrypted VM disks already work, no code needed #7
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Open
doccaz
wants to merge
4
commits into
cloudbase:main
Choose a base branch
from
doccaz:encryption-investigation
base: main
Could not load branches
Branch not found: {{ refName }}
Loading
Could not load tags
Nothing to show
Loading
Are you sure you want to change the base?
Some commits from the old base branch may be removed from the timeline,
and old review comments may become outdated.
Open
Changes from all commits
Commits
Show all changes
4 commits
Select commit
Hold shift + click to select a range
10d16f5
Add direct-ESXi support, GetInfo, QueryAllocatedBlocks, DDB_GET, and CBT
doccaz 7b61452
Document that encrypted VM disks already work, no code needed
doccaz 61dceb3
Close the cross-host key provisioning gap in the encryption investiga…
doccaz e7367d3
Address review feedback on the encryption investigation
doccaz File filter
Filter by extension
Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
There are no files selected for viewing
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,110 @@ | ||
| # Changed Block Tracking (CBT) | ||
|
|
||
| This is not a reverse-engineered NFC feature. VixDiskLib does not expose | ||
| CBT itself: `VixDiskLib_QueryAllocatedBlocks` (implemented separately; | ||
| see `docs/nfc_read.md`) reports which blocks are *allocated* | ||
| (non-sparse) within a single NFC-opened disk, not which byte ranges | ||
| *changed* between two points in time. Real backup tools | ||
| get changed-range information from vSphere's public | ||
| `VirtualMachine.QueryChangedDiskAreas` VIM call instead, used alongside | ||
| VDDK/NFC reads for the actual bytes. `openvixdisklib.nfc_auth` wraps | ||
| that public pyVmomi call directly — no capture, no wire format to | ||
| document, per the project rule to reuse pyVmomi for anything it already | ||
| exposes. | ||
|
|
||
| ## Workflow | ||
|
|
||
| 1. `nfc_auth.enable_change_tracking(vm)` — sets | ||
| `VirtualMachineConfigSpec.changeTrackingEnabled = True` via | ||
| `ReconfigVM_Task`. Takes effect for writes from that point forward; | ||
| it does not retroactively track earlier changes. | ||
| 2. Take a snapshot (or power-cycle the VM). A disk's `changeId` is | ||
| empty until this happens. | ||
| 3. `nfc_auth.disk_change_id(vm, device_key)` — reads the current | ||
| `changeId` off `VirtualDisk.backing.changeId` (for example | ||
| `"52 f3 b6 37 30 8d ea 3e-58 70 c0 fd 61 44 26 62/2"`). | ||
| 4. Do backup work (VDDK/NFC reads of the disk at that point). | ||
| 5. Later, take another snapshot. | ||
| 6. `nfc_auth.query_changed_disk_areas(vm, new_snapshot, device_key, | ||
| change_id_from_step_3)` — returns the byte ranges written between | ||
| the two snapshots. | ||
| 7. Read only those ranges via VDDK/NFC on the new snapshot's disk | ||
| chain for an incremental backup. | ||
|
|
||
| For an initial full backup, pass `change_id="*"` in step 6 without a | ||
| prior snapshot. **Correction from an earlier draft of this doc:** | ||
| this does *not* report the entire disk as one changed extent — see | ||
| "Wildcard `changeId='*'` reports allocated regions, not the whole | ||
| disk" below. | ||
|
|
||
| ## Validated in this lab | ||
|
|
||
| Confirmed end-to-end against a temporary VM on the standalone ESXi | ||
| 8.0.3 lab host (no vCenter): enabled CBT, snapshotted, wrote one | ||
| sector via `openvixdisklib.openvixdisklib` at a known offset, | ||
| snapshotted again, and called `query_changed_disk_areas` with the | ||
| first snapshot's `changeId`. The single reported extent | ||
| (`start=3932160, length=65536`, i.e. sectors 7680–7807) correctly | ||
| covered the written sector (7777). Extents were 64 KiB-aligned in this | ||
| lab's observations; that granularity is server-defined, not part of | ||
| the function's contract. | ||
|
|
||
| ### Wildcard `changeId="*"` reports allocated regions, not the whole disk | ||
|
|
||
| Tested `query_changed_disk_areas(vm, snapshot, device_key, "*")` (the | ||
| initial-full-backup path, no prior snapshot needed) against a fresh | ||
| 10 GiB thin-provisioned temp-VM disk. `result.length` correctly | ||
| reports the full declared virtual capacity (10737418240 bytes), but | ||
| `result.changed_areas` only covered **1 MiB** total — not the whole | ||
| disk. For a thin-provisioned disk, `"*"` reports the regions that are | ||
| actually *allocated* (backed by real data on the datastore), not the | ||
| full sparse virtual capacity; unwritten/unallocated regions have | ||
| nothing to back up regardless. A backup tool doing an initial full | ||
| backup with `"*"` should read exactly the reported extents, not assume | ||
| it needs to read `result.length` bytes. | ||
|
|
||
| ### One large contiguous write is one extent; scattered writes are not | ||
|
|
||
| Wrote a single 4 MiB contiguous region plus three separate one-sector | ||
| writes at scattered offsets (same disk, one CBT interval), then | ||
| queried changed areas: | ||
|
|
||
| ``` | ||
| 4 extents reported: | ||
| start= 196608 length= 65536 (64 KiB) | ||
| start= 51183616 length= 4259840 (4160 KiB) <-- covers the whole 4 MiB write as ONE extent | ||
| start= 460783616 length= 65536 (64 KiB) | ||
| start= 921567232 length= 65536 (64 KiB) | ||
| ``` | ||
|
|
||
| The 4 MiB write came back as a single extent (padded slightly beyond | ||
| 4 MiB — 4259840 bytes vs. the exact 4194304 written — to the 64 KiB | ||
| tail-end granularity). Each scattered single-sector write produced its | ||
| own separate 64 KiB extent. **The extent list scales with the number | ||
| of discontiguous changed regions, not with the total volume of changed | ||
| data.** A multi-hundred-GB sequential write is still one small extent | ||
| record; thousands of scattered small writes (e.g. a busy database VM | ||
| doing random I/O across a large disk) produce thousands of extent | ||
| records in one `QueryChangedDiskAreas` response, since the API has no | ||
| pagination. Real backup tools facing that scenario typically chunk the | ||
| query with `start_offset` over fixed-size windows rather than querying | ||
| the whole disk in one call — `query_changed_disk_areas`'s | ||
| `start_offset` parameter exists for this, but nothing in this module | ||
| does the chunking loop itself; that is caller responsibility. | ||
|
|
||
| Disk-size scaling itself (e.g., whether extent granularity increases | ||
| for very large disks) was not tested — only reasoned about above as an | ||
| open question, not verified against a large ESXi 8 disk. | ||
|
|
||
| ## What this does not cover | ||
|
|
||
| - `VixDiskLib_QueryAllocatedBlocks` (NFC-level allocated-block bitmap | ||
| within a single disk, useful for skipping sparse regions inside a | ||
| delta disk) — implemented separately, see `docs/nfc_read.md`. Pairs | ||
| naturally with CBT: `query_changed_disk_areas` says which byte | ||
| ranges changed, `query_allocated_blocks` says which parts of a | ||
| snapshot's delta disk are actually worth reading. Note its "same | ||
| still-open write handle" staleness gotcha in `docs/nfc_read.md` if | ||
| chaining a CBT-driven write with an allocation check. | ||
| - `DDB_GET` fields (`biosGeo`, `adapterType`, `uuid`) — also | ||
| implemented, see `docs/nfc_open.md`. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,141 @@ | ||
| # Encrypted VM disks | ||
|
|
||
| This is not a reverse-engineered NFC feature. When the ESXi host | ||
| serving an NFC session already holds the target disk's encryption key | ||
| — the common case, and the only one testable without a second ESXi | ||
| host (see "What this does not cover" below) — VM/VMDK encryption is | ||
| handled entirely by ESXi's storage stack, **below and invisible to the | ||
| NFC protocol**. `openvixdisklib` already reads and writes encrypted | ||
| disks correctly with no encryption-specific code at all; this was | ||
| verified, not implemented. | ||
|
|
||
| ## Investigation | ||
|
|
||
| The backlog item was "encrypted disks" (VDDK supports VM/VMDK | ||
| encryption; OpenVixDiskLib had no code path for it and no lab evidence | ||
| either way). `strings` on `libvddkVimAccess.so` found suggestive | ||
| tokens — `"cannot cast the crypto manager to CryptoManagerHostKMS"`, | ||
| `"Cannot push crypto key to host"` — implying VDDK does some kind of | ||
| key-provisioning VIM call before NFC I/O on an encrypted disk. Testing | ||
| this needed an actual encrypted VM, which needed vCenter (VM | ||
| encryption / Key Management Server configuration is a vCenter-only | ||
| concept — a bare standalone ESXi host's `cryptoManager` is the base | ||
| `vim.encryption.CryptoManagerHost`, not `CryptoManagerHostKMS`, and | ||
| `AddKey` fails with `InvalidState`/"crypto state incapable": no KMS | ||
| trust relationship is possible without vCenter). Building that lab | ||
| (vCenter Server Appliance + a vSphere Native Key Provider + an | ||
| encrypted test VM) is documented separately in | ||
| `docs/encryption_lab_setup.md`. | ||
|
|
||
| With a real encrypted disk in hand, the same SSL-hook technique as | ||
| every other NFC feature (`docs/ssl_hook.md`) captured native VDDK | ||
| 8.0.2 opening it, compared against an identical capture against an | ||
| unencrypted disk on the same host. Findings: | ||
|
|
||
| - VDDK's own log makes the branch explicit. Unencrypted: | ||
| `VddkVimAccess_HandleDiskCryptoKey: Handle the key of disk ...` | ||
| immediately followed by | ||
| `HandleDiskCryptoKey: Disk '...' is not encrypted. No need to handle | ||
| disk crypto key.` Encrypted: the same `HandleDiskCryptoKey` line | ||
| fires, but that second message is absent — it goes straight to | ||
| `VddkVimAccess_FreeNfcTicket` with no further narration. | ||
| - The encrypted-disk capture (844 SSL/socket records covering both the | ||
| vCenter SOAP session and the ESXi NFC session) contains **no** | ||
| `AddKey`, `ConfigureCryptoKey`, `EnableCrypto`, `PrepareCrypto`, or | ||
| `QueryCryptoKeyStatus` calls — the VIM methods that looked like | ||
| candidates for "push key to host" from the binary strings. The only | ||
| crypto-related SOAP content anywhere in the capture is routine | ||
| `RetrieveServiceContent` boilerplate (`cryptoManager | ||
| type="CryptoManagerKmip"` at vCenter, `type="CryptoManagerHostKMS"` | ||
| at the host) that appears on every connection, encrypted disk or not. | ||
|
|
||
| That's negative evidence (no extra call *seen*), which only proves the | ||
| happy path. The decisive test was positive: used `openvixdisklib` | ||
| itself — no crypto-specific code — to write a known byte pattern to | ||
| an encrypted disk's sector 0 through normal `VixDiskLib_Write`-equivalent | ||
| NFC I/O, read it back through a **separate** `open()` (avoiding the | ||
| same-handle staleness class of bug noted in `docs/nfc_read.md`), and | ||
| got an exact match: a completely ordinary, transparent NFC round-trip. | ||
| Then fetched the same byte range directly from the disk's | ||
| `-flat.vmdk` file over ESXi's datastore-browser HTTP API | ||
| (`Range: bytes=0-511`) and confirmed the raw bytes on disk are genuine | ||
| ciphertext — the plaintext pattern does not appear. Real | ||
| encryption-at-rest is happening (confirmed independently at the wire | ||
| level too: the `.vmdk` descriptor carries `keyID=`, an `encryptionKeys=` | ||
| wrapped-key blob, and `ddb.iofilters = "vmwarevmcrypt"`), and it is | ||
| completely invisible to any NFC client, VDDK or otherwise. | ||
|
|
||
| ## Why this is the expected result | ||
|
|
||
| VDDK's "push crypto key to host" logic exists for *provisioning*: the | ||
| case where the ESXi host about to serve NFC I/O does **not** already | ||
| have the disk's key cached locally — for example, right after a | ||
| cross-host vMotion, clone, or restore places the encrypted VM on a | ||
| host that has never served it before. In that case VDDK (running with | ||
| appropriate vCenter privileges) fetches the key and pushes it to the | ||
| target host via `CryptoManagerHostKMS.AddKey` before NFC can succeed. | ||
| Once the host has the key — which it always will for a VM that has | ||
| simply been encrypted and left in place, the scenario this lab's | ||
| single ESXi host can produce — ESXi's storage stack decrypts/encrypts | ||
| transparently at the IOFilter layer for every client, with nothing | ||
| encryption-specific in the NFC wire protocol at all. | ||
|
|
||
| ### Transports not covered | ||
|
|
||
| Only the NBD/NFC path was tested. The `hotadd` transport most likely | ||
| behaves the same way, since the disk is still read through the ESXi | ||
| storage stack. The `san` transport is an open question: it reads the | ||
| underlying LUNs (iSCSI, Fibre Channel) directly, mounts VMFS and parses | ||
| the VMDK itself, bypassing the IOFilter layer where decryption happens, | ||
| so it would presumably see ciphertext. Not verified. | ||
|
|
||
| ## Validated in this lab | ||
|
|
||
| Single-ESXi-host lab (`docs/encryption_lab_setup.md`), VM | ||
| `ovdl-crypto-test` with both its home directory and its virtual disk | ||
| encrypted via a vSphere Native Key Provider: | ||
|
|
||
| - Read/write round-trip via `openvixdisklib.openvixdisklib` over the | ||
| vCenter-managed path (`server_name=<vcenter>`, `vmx_spec="moref=vm-N"`) | ||
| — matches the pattern written, confirmed as genuine ciphertext at | ||
| rest as described above. | ||
| - Native VDDK 8.0.2 (`ConnectEx` + `Open` + `Read`) succeeds against | ||
| the same disk with no special handling required, decrypting | ||
| correctly (a freshly-created 1 GiB disk reads back as zeros, as | ||
| expected for an unwritten region). | ||
|
|
||
| ## Cross-host key provisioning (resolved 2026-09-20) | ||
|
|
||
| The gap flagged below — a host that doesn't already have the disk's | ||
| key cached — was resolved once a second ESXi host became available | ||
| (built for the host-switch investigation, `docs/host_switch_lab_setup.md`). | ||
| Relocated the encrypted test VM (compute **and** its disk, cold — | ||
| powered off) to a host that had never held its key at all, then opened | ||
| the disk on that new host: it just worked, both via `openvixdisklib` | ||
| and via cross-checked native VDDK, decrypting correctly with zero key | ||
| management on the client's part. | ||
|
|
||
| vCenter pushes the key to the destination host automatically as part | ||
| of any relocation of an encrypted VM (`RelocateVM_Task`, cold or live) | ||
| — entirely transparent to any NFC client. The `CryptoManagerHostKMS.AddKey` | ||
| ("push crypto key to host") mechanism the binary strings pointed to | ||
| does exist and does get exercised, but it's driven by vCenter's own | ||
| migration workflow, not by VDDK or any NFC client. There is nothing | ||
| for OpenVixDiskLib to implement here either: the "host doesn't have | ||
| the key yet" case simply cannot arise for a client reading a disk | ||
| through a normal, already-completed migration, because vCenter | ||
| resolves it before the migration finishes. | ||
|
|
||
| ## What this does not cover | ||
|
|
||
| - Standard/KMIP external Key Management Servers were not tested — only | ||
| vSphere Native Key Provider. The NFC-transparency conclusion should | ||
| hold regardless of which KMS variant provisioned the key (the | ||
| key-caching-on-host mechanism is the same either way), but this | ||
| wasn't independently verified. | ||
| - Direct-ESXi (no vCenter) access to an encrypted disk was attempted | ||
| but not completed — hit an unrelated, reproducible pyVmomi issue | ||
| (a bare `vim.VirtualMachine("vm-N", stub)` construction raising | ||
| `ManagedObjectNotFound` against this specific host, independent of | ||
| encryption) that wasn't investigated further. Worth retrying if this | ||
| recurs elsewhere. | ||
Oops, something went wrong.
Add this suggestion to a batch that can be applied as a single commit.
This suggestion is invalid because no changes were made to the code.
Suggestions cannot be applied while the pull request is closed.
Suggestions cannot be applied while viewing a subset of changes.
Only one suggestion per line can be applied in a batch.
Add this suggestion to a batch that can be applied as a single commit.
Applying suggestions on deleted lines is not supported.
You must change the existing code in this line in order to create a valid suggestion.
Outdated suggestions cannot be applied.
This suggestion has been applied or marked resolved.
Suggestions cannot be applied from pending reviews.
Suggestions cannot be applied on multi-line comments.
Suggestions cannot be applied while the pull request is queued to merge.
Suggestion cannot be applied right now. Please check back later.
Uh oh!
There was an error while loading. Please reload this page.