-
Notifications
You must be signed in to change notification settings - Fork 68
backup writer: accept already compressed chunks #525
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,41 @@ | ||
| # Agent guidelines | ||
|
|
||
| Coriolis migrates VMs between clouds. Cloud support lives in separate | ||
| provider plugins (often private). Target Python 3.10 and 3.12. | ||
|
|
||
| ## Coriolis terminology | ||
|
|
||
| - **Minion**: temporary worker VM for disk transfer and os-morphing | ||
| (guest prep: network, packages, drivers). **Minion pools** reuse them. | ||
| Some source providers skip minions and read disks directly. | ||
| - **Transfer**: creates destination volumes and copies disk data. | ||
| Re-run a transfer as a new **execution** to pick up later changes. | ||
| Most providers can do that incrementally. | ||
| - **Deployment**: creates the destination VM from a completed transfer. | ||
| - **Replica** vs **migration** (`transfer.scenario`: `replica` / | ||
| `live_migration`) is a licensing split, not two engines. Replicas can | ||
| be re-executed and re-deployed even after license fulfillment. | ||
|
|
||
| Users configure source/destination **endpoints** (credentials) and | ||
| **environment options** (transfer and resulting VM settings). | ||
|
|
||
| ## Working in this repo | ||
|
|
||
| - Use `.tox/py3/bin/` (`stestr`, `ruff`); it has project deps. Ignore | ||
| `.mypy_cache`, `.ruff_cache`, `.tox`. | ||
| - Follow ruff (`tox.ini`, `ruff.toml`). | ||
| - Public methods need docstrings (subclasses may inherit). Use type | ||
| hints when the type is known. | ||
| - Do not add helpers for trivial checks such as | ||
| `server.power_status == "RUNNING"`; keep those inline. | ||
| - Do not strip still-relevant inline comments. | ||
| - If regenerating a file, replace its contents; do not append duplicates. | ||
| - Empty `__init__.py` files must not have license headers. | ||
|
|
||
| ## Tests | ||
|
|
||
| - Prefer `@mock.patch` / `@mock.patch.object` over `with mock.patch`. | ||
| - For multiple mock calls, use `assert_has_calls`. For a single call, | ||
| `assert_called_once_with` is fine. Match nearby tests. | ||
| - Integration tests use Docker providers; external cloud providers are | ||
| optional. |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -257,7 +257,16 @@ def truncate(self, size): | |
| pass | ||
|
|
||
| @abc.abstractmethod | ||
| def write(self, data): | ||
| def write(self, data, encoding=None, uncompressed_size=None): | ||
| """Write the given data chunk. | ||
|
|
||
| Use the `seek` method to specify the write offset. | ||
|
|
||
| :param data: data chunk to write (raw bytes) | ||
| :param encoding: a data encoding format supported by the backup writer. | ||
| :param uncompressed_size: the uncompressed size in case of | ||
| compressed chunks. | ||
| """ | ||
| pass | ||
|
|
||
| @abc.abstractmethod | ||
|
|
@@ -319,7 +328,19 @@ def seek(self, pos): | |
| def truncate(self, size): | ||
| self._file.truncate(size) | ||
|
|
||
| def write(self, data): | ||
| def write(self, data, encoding=None, uncompressed_size=None): | ||
| """Write the given data chunk. | ||
|
|
||
| Use the `seek` method to specify the write offset. | ||
|
|
||
| :param data: data chunk to write (raw bytes) | ||
| :param encoding: unsupported, this backup writer only accepts raw chunks. | ||
| :param uncompressed_size: unsupported. | ||
| """ | ||
| if encoding: | ||
| raise exception.InvalidInput( | ||
| f"The file backup writer does not support {encoding} encoded chunks." | ||
| ) | ||
| self._file.write(data) | ||
|
|
||
| def close(self): | ||
|
|
@@ -463,7 +484,15 @@ def _encoder(self): | |
| self._enc_q.task_done() | ||
| LOG.debug("Backup encoder stopped.") | ||
|
|
||
| def write(self, data): | ||
| def write(self, data, encoding=None, uncompressed_size=None): | ||
| """Write the given data chunk. | ||
|
|
||
| Use the `seek` method to specify the write offset. | ||
|
|
||
| :param data: data chunk to write (raw bytes) | ||
| :param encoding: unsupported, this backup writer only accepts raw chunks. | ||
| :param uncompressed_size: unsupported. | ||
| """ | ||
| if self._closing: | ||
| raise exception.CoriolisException("Attempted to write to a closed writer.") | ||
|
|
||
|
|
@@ -472,6 +501,11 @@ def write(self, data): | |
| "Failed to write data. See log for details." | ||
| ) from self._exception | ||
|
|
||
| if encoding: | ||
| raise exception.InvalidInput( | ||
| f"The file backup writer does not support {encoding} encoded chunks." | ||
| ) | ||
|
|
||
| payload = { | ||
| "offset": self._offset, | ||
| "data": data, | ||
|
|
@@ -785,6 +819,10 @@ def _sender(self): | |
| if payload.get("encoding", None): | ||
| enc = copy.copy(payload["encoding"]) | ||
| headers["content-encoding"] = enc | ||
| if payload.get("uncompressed_size") is not None: | ||
| headers["X-Uncompressed-Content-Length"] = str( | ||
| payload["uncompressed_size"] | ||
| ) | ||
|
|
||
| @utils.retry_on_error() | ||
| def send(): | ||
|
|
@@ -831,7 +869,18 @@ def send(): | |
| LOG.debug("Backup sender stopped.") | ||
|
|
||
| @utils.retry_on_error() | ||
| def write(self, data): | ||
| def write(self, data, encoding=None, uncompressed_size=None): | ||
| """Write the given data chunk. | ||
|
|
||
| Use the `seek` method to specify the write offset. | ||
|
|
||
| :param data: data chunk to write (raw bytes) | ||
| :param encoding: the encoding of the data chunk, one of the following: | ||
| * compression algorithms: fastlz, deflate, gzip, zlib | ||
| * None: raw data, no encoding (default) | ||
| :param uncompressed_size: the uncompressed size in case of | ||
| compressed chunks. Required for fastlz. | ||
| """ | ||
| if self._closing: | ||
| raise exception.CoriolisException("Attempted to write to a closed writer.") | ||
| if self._exception: | ||
|
|
@@ -841,7 +890,27 @@ def write(self, data): | |
| "offset": self._offset, | ||
| "data": data, | ||
| } | ||
| self._comp_q.put(payload) | ||
| if encoding is None: | ||
| self._comp_q.put(payload) | ||
| elif encoding in ("fastlz", "gzip", "zlib", "deflate"): | ||
| # The payload is already compressed, skip the compressor | ||
| # queue, use the sender queue directly. | ||
| payload["encoding"] = encoding | ||
| payload["chunk"] = data | ||
| if encoding == "fastlz": | ||
| if uncompressed_size is None: | ||
| raise exception.InvalidInput( | ||
| "fastlz without explicit uncompressed size." | ||
| ) | ||
| payload["uncompressed_size"] = uncompressed_size | ||
| self._sender_q.put(payload) | ||
| elif encoding == "incompressible": | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. What are the usual reasons for a chunk to be incompressible?
Member
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Repeated sequences (e.g. zero blocks, text, etc) can easily be compressed. On the other hand, encrypted data usually cannot be compressed. Same applies to already compressed chunks (e.g. rotated log files). High entropy leads to low compression rates. |
||
| # The caller determined that the chunk is incompressible, | ||
| # skip the compression queue. | ||
| payload["chunk"] = data | ||
| self._sender_q.put(payload) | ||
| else: | ||
| raise exception.InvalidInput("Unsupported write encoding: %s" % encoding) | ||
| self._offset += len(data) | ||
|
|
||
| def _wait_for_queues(self): | ||
|
|
||
Uh oh!
There was an error while loading. Please reload this page.