backup writer: accept already compressed chunks - #525
petrutlucian94 wants to merge 3 commits into
Conversation
c16edf0 to
b8f3496
Compare
b8f3496 to
6799df0
Compare
We'll add a minimal AGENTS.md file in order to guide AI tools.
6799df0 to
4dada6b
Compare
Dany9966
left a comment
There was a problem hiding this comment.
Apart from a couple of nits, LGTM
| - **Deployment**: creates the destination VM from a completed transfer. | ||
| - **Replica** vs **migration** (`transfer.scenario`: `replica` / | ||
| `live_migration`) is a licensing split, not two engines. Replicas can | ||
| be re-executed and re-deployed. |
There was a problem hiding this comment.
NIT:
| be re-executed and re-deployed. | |
| be re-executed and re-deployed even after license fulfillment. |
| self._comp_q.put(payload) | ||
| if encoding is None: | ||
| self._comp_q.put(payload) | ||
| elif encoding in ("fastlz", "gzip", "zlib", "deflate"): |
There was a problem hiding this comment.
Can we document what encoding is used for under the method definition please?
What I'd like to avoid is confussion between already encoded data (which is what the arg is used for) and encoding used by the compressor (which is also gzip by default: https://github.com/cloudbase/coriolis/blob/main/coriolis/providers/backup_writers.py#L758)
| ) | ||
| payload["uncompressed_size"] = uncompressed_size | ||
| self._sender_q.put(payload) | ||
| elif encoding == "incompressible": |
There was a problem hiding this comment.
What are the usual reasons for a chunk to be incompressible?
There was a problem hiding this comment.
Repeated sequences (e.g. zero blocks, text, etc) can easily be compressed. On the other hand, encrypted data usually cannot be compressed. Same applies to already compressed chunks (e.g. rotated log files).
High entropy leads to low compression rates.
The AGENTS.md file can be condensed, preserving only the information that is actually useful to AI agents. The previous commit is kept intentionally as a reference.
Unlike VDDK, OpenVixDiskLib can skip decompressing chunks originating from ESXi. We can take advantage of this and forward the already compressed FastLZ chunks to coriolis-writer. To do so, we'll add the "encoding" and "uncompressed_length" parameters to the "write" method of the backup writer interface. The HTTP backup writer will skip the "compressor" queue when receiving already chunks, submitting them directly to the "sender" queue. The caller may also use the "incompressible" encoding as a hint that a given chunk cannot be compressed and that it should be sent right away. Other backup writers (ssh, file) will error out if an encoding is specified. Note that with FastLZ we also need to declare the uncompressed chunk length. We went with a request header as opposed to a buffer prefix in order to reduce the number of buffer copy operations. While at it, we're including a coriolis-writer build that includes FastLZ encoding support.
4dada6b to
c0254b1
Compare
Unlike VDDK, OpenVixDiskLib can skip decompressing chunks originating
from ESXi.
We can take advantage of this and forward the already compressed
FastLZ chunks to coriolis-writer.
To do so, we'll add the "encoding" and "uncompressed_length"
parameters to the "write" method of the backup writer interface.
The HTTP backup writer will skip the "compressor" queue when
receiving already chunks, submitting them directly to the "sender"
queue.
The caller may also use the "incompressible" encoding
as a hint that a given chunk cannot be compressed and that it should
be sent right away.
Other backup writers (ssh, file) will error out if an encoding
is specified.
Note that with FastLZ we also need to declare the uncompressed chunk
length. We went with a request header as opposed to a buffer prefix
in order to reduce the number of buffer copy operations.
While at it, we're including a coriolis-writer build that includes
FastLZ encoding support.