Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
5 changes: 4 additions & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -24,8 +24,11 @@ Implemented against vCenter 8 / ESXi 8. Default transport is `nbdssl`
- `VixDiskLib_Open` (datastore path, read-only or read-write)
- `VixDiskLib_Read` (optional ``skip_decompression`` packs FastLZ extras)
- `VixDiskLib_Write`
- Changed Block Tracking: `openvixdisklib.nfc_auth.enable_change_tracking`
/ `disk_change_id` / `query_changed_disk_areas` (public VIM API, not
part of VixDiskLib itself; see `docs/cbt.md`)

Not implemented: compression open flags other than FastLZ, CBT /
Not implemented: compression open flags other than FastLZ,
allocated-block queries, disk geometry (`DDB_GET`), encrypted disks,
and direct ESXi `ha-nfc` without vCenter `vpxa-nfc`.

Expand Down
110 changes: 110 additions & 0 deletions docs/cbt.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,110 @@
# Changed Block Tracking (CBT)

This is not a reverse-engineered NFC feature. VixDiskLib does not expose
CBT itself: `VixDiskLib_QueryAllocatedBlocks` (implemented separately;
see `docs/nfc_read.md`) reports which blocks are *allocated*
(non-sparse) within a single NFC-opened disk, not which byte ranges
*changed* between two points in time. Real backup tools
get changed-range information from vSphere's public
`VirtualMachine.QueryChangedDiskAreas` VIM call instead, used alongside
VDDK/NFC reads for the actual bytes. `openvixdisklib.nfc_auth` wraps
that public pyVmomi call directly — no capture, no wire format to
document, per the project rule to reuse pyVmomi for anything it already
exposes.

## Workflow

1. `nfc_auth.enable_change_tracking(vm)` — sets
`VirtualMachineConfigSpec.changeTrackingEnabled = True` via
`ReconfigVM_Task`. Takes effect for writes from that point forward;
it does not retroactively track earlier changes.
2. Take a snapshot (or power-cycle the VM). A disk's `changeId` is
empty until this happens.
3. `nfc_auth.disk_change_id(vm, device_key)` — reads the current
`changeId` off `VirtualDisk.backing.changeId` (for example
`"52 f3 b6 37 30 8d ea 3e-58 70 c0 fd 61 44 26 62/2"`).
4. Do backup work (VDDK/NFC reads of the disk at that point).
5. Later, take another snapshot.
6. `nfc_auth.query_changed_disk_areas(vm, new_snapshot, device_key,
change_id_from_step_3)` — returns the byte ranges written between
the two snapshots.
7. Read only those ranges via VDDK/NFC on the new snapshot's disk
chain for an incremental backup.

For an initial full backup, pass `change_id="*"` in step 6 without a
prior snapshot. **Correction from an earlier draft of this doc:**
this does *not* report the entire disk as one changed extent — see
"Wildcard `changeId='*'` reports allocated regions, not the whole
disk" below.

## Validated in this lab

Confirmed end-to-end against a temporary VM on the standalone ESXi
8.0.3 lab host (no vCenter): enabled CBT, snapshotted, wrote one
sector via `openvixdisklib.openvixdisklib` at a known offset,
snapshotted again, and called `query_changed_disk_areas` with the
first snapshot's `changeId`. The single reported extent
(`start=3932160, length=65536`, i.e. sectors 7680–7807) correctly
covered the written sector (7777). Extents were 64 KiB-aligned in this
lab's observations; that granularity is server-defined, not part of
the function's contract.

### Wildcard `changeId="*"` reports allocated regions, not the whole disk

Tested `query_changed_disk_areas(vm, snapshot, device_key, "*")` (the
initial-full-backup path, no prior snapshot needed) against a fresh
10 GiB thin-provisioned temp-VM disk. `result.length` correctly
reports the full declared virtual capacity (10737418240 bytes), but
`result.changed_areas` only covered **1 MiB** total — not the whole
disk. For a thin-provisioned disk, `"*"` reports the regions that are
actually *allocated* (backed by real data on the datastore), not the
full sparse virtual capacity; unwritten/unallocated regions have
nothing to back up regardless. A backup tool doing an initial full
backup with `"*"` should read exactly the reported extents, not assume
it needs to read `result.length` bytes.

### One large contiguous write is one extent; scattered writes are not

Wrote a single 4 MiB contiguous region plus three separate one-sector
writes at scattered offsets (same disk, one CBT interval), then
queried changed areas:

```
4 extents reported:
start= 196608 length= 65536 (64 KiB)
start= 51183616 length= 4259840 (4160 KiB) <-- covers the whole 4 MiB write as ONE extent
start= 460783616 length= 65536 (64 KiB)
start= 921567232 length= 65536 (64 KiB)
```

The 4 MiB write came back as a single extent (padded slightly beyond
4 MiB — 4259840 bytes vs. the exact 4194304 written — to the 64 KiB
tail-end granularity). Each scattered single-sector write produced its
own separate 64 KiB extent. **The extent list scales with the number
of discontiguous changed regions, not with the total volume of changed
data.** A multi-hundred-GB sequential write is still one small extent
record; thousands of scattered small writes (e.g. a busy database VM
doing random I/O across a large disk) produce thousands of extent
records in one `QueryChangedDiskAreas` response, since the API has no
pagination. Real backup tools facing that scenario typically chunk the
query with `start_offset` over fixed-size windows rather than querying
the whole disk in one call — `query_changed_disk_areas`'s
`start_offset` parameter exists for this, but nothing in this module
does the chunking loop itself; that is caller responsibility.

Disk-size scaling itself (e.g., whether extent granularity increases
for very large disks) was not tested — only reasoned about above as an
open question, not verified against a large ESXi 8 disk.

## What this does not cover

- `VixDiskLib_QueryAllocatedBlocks` (NFC-level allocated-block bitmap
within a single disk, useful for skipping sparse regions inside a
delta disk) — implemented separately, see `docs/nfc_read.md`. Pairs
naturally with CBT: `query_changed_disk_areas` says which byte
ranges changed, `query_allocated_blocks` says which parts of a
snapshot's delta disk are actually worth reading. Note its "same
still-open write handle" staleness gotcha in `docs/nfc_read.md` if
chaining a CBT-driven write with an allocation check.
- `DDB_GET` fields (`biosGeo`, `adapterType`, `uuid`) — also
implemented, see `docs/nfc_open.md`.
19 changes: 16 additions & 3 deletions docs/reverse_engineering_procedure.md
Original file line number Diff line number Diff line change
Expand Up @@ -381,8 +381,21 @@ not an OPEN_FILE bit. Capture VDDK with that flag (NBD + the port-902
Replay: pip `pyfastlz` via `openvixdisklib/fastlz.py` (NFC extra is
raw FastLZ, without the wrapper's 4-byte length prefix) plus `NfcDisk`
compression on each IO. Proof:
`tests/integration/test_nfc_read_write.py` (`fastlz`) and
`tests/perf/test_compare.py`.
## Note — CBT needed no reverse engineering

Investigated change-block tracking (backlog item "CBT /
`QueryAllocatedBlocks`") expecting an NFC capture like the steps above.
It turned out `VirtualMachine.QueryChangedDiskAreas` — the actual
changed-byte-range query backup tools use — is public pyVmomi API with
no VDDK/NFC involvement at all; only `VixDiskLib_QueryAllocatedBlocks`
(disk-internal allocated-block bitmap, a different and lesser feature)
needed NFC work (done separately, see Step 15 below). Implemented as
`openvixdisklib.nfc_auth.enable_change_tracking` /
`disk_change_id` / `query_changed_disk_areas`; validated end-to-end
against a temp VM on the lab (enable CBT, snapshot, write a known
sector, snapshot, query — the written sector fell inside the reported
extent). Full workflow and lab evidence: `docs/cbt.md`.


## What to write down

Expand All @@ -405,7 +418,7 @@ OpenVixDiskLib.
Not yet reversed, same loop as above:

- `DDB_GET` / disk geometry, zlib/skipz compression, encrypted disks
- `NFC_DELTA_DISK`, CBT / `QueryAllocatedBlocks`
- `NFC_DELTA_DISK`, `QueryAllocatedBlocks`
- `VixDiskLib_GetInfo` capacity
- Host-switch AIO messages
- Direct ESXi `ha-nfc` without vCenter `vpxa-nfc`
107 changes: 107 additions & 0 deletions openvixdisklib/nfc_auth.py
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,8 @@
import hashlib
import socket
import ssl
import time
from dataclasses import dataclass

from pyVim.connect import Disconnect, SmartConnect
from pyVmomi import vim
Expand All @@ -29,6 +31,8 @@
NFC_SERVICE_MOID = "nfcService"
AUTHD_DEFAULT_PORT = 902
_NFC_TYPES_REGISTERED = False
_TASK_POLL_S = 0.5
_TASK_TIMEOUT_S = 300


def _ssl_client_context(verify: bool = True) -> ssl.SSLContext:
Expand Down Expand Up @@ -469,3 +473,106 @@ def authenticate(
Disconnect(si)
raise
return NfcAuthSession(si, ticket, authd_sock, nfc_ssl=nfc_ssl)


# --- Changed Block Tracking (CBT) ---
#
# VixDiskLib does not expose CBT itself: ``VixDiskLib_QueryAllocatedBlocks``
# reports which blocks are allocated (non-sparse) within a single
# NFC-opened disk, not which byte ranges changed between two points in
# time. Real backup tools get that from vSphere's public
# ``VirtualMachine.QueryChangedDiskAreas`` VIM call instead, used
# alongside VDDK/NFC reads. The helpers below are thin wrappers around
# that public pyVmomi call — no NFC reverse engineering was needed for
# them. See ``docs/cbt.md``.


@dataclass(frozen=True, slots=True)
class ChangedExtent:
"""One changed byte range, as returned by ``QueryChangedDiskAreas``."""

start: int
length: int


@dataclass(frozen=True, slots=True)
class ChangedDiskAreas:
"""Result of ``QueryChangedDiskAreas``, converted to plain dataclasses."""

start_offset: int
length: int
changed_areas: tuple[ChangedExtent, ...]


def _wait_for_cbt_task(task: vim.Task):
deadline = time.monotonic() + _TASK_TIMEOUT_S
while task.info.state in (vim.TaskInfo.State.running, vim.TaskInfo.State.queued):
if time.monotonic() > deadline:
raise TimeoutError(f"timed out waiting for vSphere task {task}")
time.sleep(_TASK_POLL_S)
if task.info.state != vim.TaskInfo.State.success:
raise RuntimeError(f"vSphere task failed: {task.info.error}")
return task.info.result


def enable_change_tracking(vm: vim.VirtualMachine) -> None:
"""Enable CBT on ``vm``.

Takes effect for writes from this point forward; it does not
retroactively track earlier changes. A ``changeId`` for a disk only
becomes available after the next snapshot or power cycle once this
is set.
"""
spec = vim.vm.ConfigSpec(changeTrackingEnabled=True)
_wait_for_cbt_task(vm.ReconfigVM_Task(spec=spec))


def disk_change_id(vm: vim.VirtualMachine, device_key: int) -> str:
"""Return the current ``changeId`` for the disk with ``device_key``.

Requires CBT to be enabled and at least one snapshot (or power
cycle) to have happened since. Raises ``ValueError`` if the disk
isn't found or has no ``changeId`` yet (CBT not active for it).
"""
for device in vm.config.hardware.device:
if isinstance(device, vim.vm.device.VirtualDisk) and device.key == device_key:
change_id = getattr(device.backing, "changeId", None)
if not change_id:
raise ValueError(
f"disk {device_key} on {vm._moId} has no changeId yet "
"(enable CBT and take a snapshot first)"
)
return change_id
raise ValueError(f"no VirtualDisk with device key {device_key} on {vm._moId}")


def query_changed_disk_areas(
vm: vim.VirtualMachine,
snapshot: vim.vm.Snapshot,
device_key: int,
change_id: str,
start_offset: int = 0,
) -> ChangedDiskAreas:
"""Return byte ranges changed since ``change_id``, up to ``snapshot``.

Thin wrapper around the public
``VirtualMachine.QueryChangedDiskAreas`` VIM call. ``change_id`` is
the value from an earlier ``disk_change_id()`` call (or ``"*"`` for
the entire disk, e.g. for an initial full backup). Extents are
64 KiB-aligned in this lab's observations, but that granularity is
server-defined and not part of this function's contract.
"""
result = vm.QueryChangedDiskAreas(
snapshot=snapshot,
deviceKey=device_key,
startOffset=start_offset,
changeId=change_id,
)
return ChangedDiskAreas(
start_offset=result.startOffset,
length=result.length,
changed_areas=tuple(
ChangedExtent(start=extent.start, length=extent.length)
for extent in result.changedArea
),
)
Loading
Loading