Skip to content

backup writer: accept already compressed chunks - #525

Open
petrutlucian94 wants to merge 3 commits into
cloudbase:mainfrom
petrutlucian94:encoding
Open

petrutlucian94 wants to merge 3 commits into
cloudbase:mainfrom
petrutlucian94:encoding

Conversation

@petrutlucian94

Copy link
Copy Markdown
Member

Unlike VDDK, OpenVixDiskLib can skip decompressing chunks originating
from ESXi.

We can take advantage of this and forward the already compressed
FastLZ chunks to coriolis-writer.

To do so, we'll add the "encoding" and "uncompressed_length"
parameters to the "write" method of the backup writer interface.
The HTTP backup writer will skip the "compressor" queue when
receiving already chunks, submitting them directly to the "sender"
queue.

The caller may also use the "incompressible" encoding
as a hint that a given chunk cannot be compressed and that it should
be sent right away.

Other backup writers (ssh, file) will error out if an encoding
is specified.

Note that with FastLZ we also need to declare the uncompressed chunk
length. We went with a request header as opposed to a buffer prefix
in order to reduce the number of buffer copy operations.

While at it, we're including a coriolis-writer build that includes
FastLZ encoding support.

We'll add a minimal AGENTS.md file in order to guide AI tools.
@petrutlucian94
petrutlucian94 marked this pull request as ready for review September 30, 2026 08:33

@Dany9966 Dany9966 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Apart from a couple of nits, LGTM

Comment thread AGENTS.md Outdated
- **Deployment**: creates the destination VM from a completed transfer.
- **Replica** vs **migration** (`transfer.scenario`: `replica` /
`live_migration`) is a licensing split, not two engines. Replicas can
be re-executed and re-deployed.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

NIT:

Suggested change
be re-executed and re-deployed.
be re-executed and re-deployed even after license fulfillment.

self._comp_q.put(payload)
if encoding is None:
self._comp_q.put(payload)
elif encoding in ("fastlz", "gzip", "zlib", "deflate"):

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can we document what encoding is used for under the method definition please?
What I'd like to avoid is confussion between already encoded data (which is what the arg is used for) and encoding used by the compressor (which is also gzip by default: https://github.com/cloudbase/coriolis/blob/main/coriolis/providers/backup_writers.py#L758)

)
payload["uncompressed_size"] = uncompressed_size
self._sender_q.put(payload)
elif encoding == "incompressible":

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

What are the usual reasons for a chunk to be incompressible?

Copy link
Copy Markdown
Member Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Repeated sequences (e.g. zero blocks, text, etc) can easily be compressed. On the other hand, encrypted data usually cannot be compressed. Same applies to already compressed chunks (e.g. rotated log files).

High entropy leads to low compression rates.

The AGENTS.md file can be condensed, preserving only the information
that is actually useful to AI agents.

The previous commit is kept intentionally as a reference.
Unlike VDDK, OpenVixDiskLib can skip decompressing chunks originating
from ESXi.

We can take advantage of this and forward the already compressed
FastLZ chunks to coriolis-writer.

To do so, we'll add the "encoding" and "uncompressed_length"
parameters to the "write" method of the backup writer interface.
The HTTP backup writer will skip the "compressor" queue when
receiving already chunks, submitting them directly to the "sender"
queue.

The caller may also use the "incompressible" encoding
as a hint that a given chunk cannot be compressed and that it should
be sent right away.

Other backup writers (ssh, file) will error out if an encoding
is specified.

Note that with FastLZ we also need to declare the uncompressed chunk
length. We went with a request header as opposed to a buffer prefix
in order to reduce the number of buffer copy operations.

While at it, we're including a coriolis-writer build that includes
FastLZ encoding support.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants