Skip to content
Merged
72 changes: 72 additions & 0 deletions .grype.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -13,6 +13,78 @@
# stable Python 3.14 or Debian 13 packages become available upstream.

ignore:
# Renewed by Adam Hernandez on 2026-09-13 for eight exact findings in the
# refreshed DHI runtime. These are accepted-risk exceptions, NOT fixes.
# Keep all other High findings blocking, including other package versions.
# Preserve the previous entries for cached images and the review dates below.
#
# Debian 13 still lists CVE-2026-5435 as vulnerable/no-DSA (TSIG printing).
# No direct Pullbox calls were found, but third-party reachability and DHI
# backport status remain unproven. Re-review by 2026-09-30 or the next base
# refresh, whichever comes first; remove when a fixed stable package ships.
# https://security-tracker.debian.org/tracker/CVE-2026-5435
- vulnerability: CVE-2026-5435
package:
name: libc6
version: 2.41-12+deb13u4
type: deb

# Debian 13 still has no fixed stable Expat package for these findings.
# The libexpat1 copy supports Fontconfig/Poppler. The separate Python parser
# reports Expat 2.8.2: its >=2.8.1 runtime guard does NOT establish protection
# against these newer advisories. Do not interpret this renewal as proving
# either parser safe. No Python-binary or libexpat1-dev exception is added.
# Re-review by 2026-10-07 or the next base refresh, whichever comes first;
# remove as soon as a fixed stable package is available.
# https://security-tracker.debian.org/tracker/CVE-2026-66046
# https://security-tracker.debian.org/tracker/CVE-2026-76956
# https://security-tracker.debian.org/tracker/CVE-2026-76957
- vulnerability: CVE-2026-66046
package:
name: libexpat1
version: 2.8.3-1~deb13u1+dhi3
type: deb
- vulnerability: CVE-2026-76956
package:
name: libexpat1
version: 2.8.3-1~deb13u1+dhi3
type: deb
- vulnerability: CVE-2026-76957
package:
name: libexpat1
version: 2.8.3-1~deb13u1+dhi3
type: deb

# These util-linux source-package findings attach to libuuid1, but concern
# privileged mount/nsenter helpers. The inspected ARM64 production image
# contains libuuid1 and neither executable. Other architectures still need
# their normal image validation. Re-review by 2026-10-04 or the next base
# refresh, whichever comes first; remove when fixed stable packages ship.
# https://security-tracker.debian.org/tracker/CVE-2026-76642
# https://security-tracker.debian.org/tracker/CVE-2026-78408
# https://security-tracker.debian.org/tracker/CVE-2026-78409
# https://security-tracker.debian.org/tracker/CVE-2026-78410
- vulnerability: CVE-2026-76642
package:
name: libuuid1
version: 2.41.5-0+deb13u1+dhi3
type: deb
- vulnerability: CVE-2026-78408
package:
name: libuuid1
version: 2.41.5-0+deb13u1+dhi3
type: deb
- vulnerability: CVE-2026-78409
package:
name: libuuid1
version: 2.41.5-0+deb13u1+dhi3
type: deb
- vulnerability: CVE-2026-78410
package:
name: libuuid1
version: 2.41.5-0+deb13u1+dhi3
type: deb

# Debian Trixie has postponed the OpenSSL 3.5 QUIC listener connection-limit
# fix for CVE-2026-14456. Pullbox does not expose an OpenSSL QUIC listener;
# retain this exact DHI base-image exception only until Debian ships a fix.
Expand Down
11 changes: 10 additions & 1 deletion docs/development/ARCHITECTURE_OVERVIEW.md
Original file line number Diff line number Diff line change
Expand Up @@ -316,7 +316,8 @@ session and WebSocket lifecycle are supervised per exact configured client.
**Current repo nuances**

- Prowlarr-synced Torznab indexers are aggregated through a single Prowlarr
search path.
search path. Persisted indexer IDs resolve to that aggregate for acquisition,
without duplicating searches or adding manager-owned per-indexer health checks.
- Prowlarr-synced Newznab indexers are kept as individual Newznab proxy
endpoints because the direct proxy behavior can produce better category and
result fidelity.
Expand All @@ -326,6 +327,14 @@ session and WebSocket lifecycle are supervised per exact configured client.
policy and retires missing manager rows instead of deleting their history.
- Download clients may be cached across task cycles when config values have not
changed.
- HTTP torrent metadata is fetched and validated inside Pullbox, then uploaded
to the selected qBittorrent, Transmission, or Deluge client. This is independent
of browser-resolver opt-in and applies to manual grabs, automatic acquisition,
intervention approval, and retries. Magnets remain URL submissions. Descriptor
fetching is bounded by the configured indexer origin, redirect count, timeout,
response size, and shared bencode validation. A failed fetch never falls back
to forwarding the private indexer URL to a remote client. Legacy HTTP downloads
without an available originating indexer require a new search.
- Pullbox Data is not a general metadata proxy. Installed clients default to
the public `https://api.pullbox.app` release API, while deployments may use
`PULLBOX_DATA_API_BASE_URL` for an intentional private-network override.
Expand Down
23 changes: 23 additions & 0 deletions docs/development/IMPORT_REVIEW_RECOVERY.md
Original file line number Diff line number Diff line change
Expand Up @@ -76,6 +76,29 @@ Available recovery actions are intentionally narrow:
issue. Ambiguous titles, filename-only guesses, conflicting target files,
stale references, and managed files remain review-only. The source path and
source artifact are never changed.
- **Recover known series** re-evaluates legacy series-level identity rejections
using saved evidence, without a new scan or provider requests. A unique Mylar
series ID, or a retained trusted folder/ComicInfo series match, can restore
only files whose saved series and issue identities agree. Existing catalog
ownership, issue numbers, per-file conflicts, manual decisions, skips, and
safety blocks are checked before an actor-bound preview is issued. Ambiguous
duplicate candidates and a series containing an unpreviewed ready file are
excluded. The confirmed scope runs through normal background Step 4 source
validation and import rules. Successful files and source paths are untouched;
unresolved files remain in Follow-up. This is not a blanket repair of stale
IDs or a replacement for manual review when trusted evidence disagrees.

For older jobs, **Retry failed** also corrects the misleading series-level
`No eligible files available for import` outcome when the series already has
successful imported files and no failed or ready files remain. It retains the
prior error in diagnostics, rebuilds the counters, and leaves unresolved file
decisions visible. A status-only correction does not launch another import.

Recovery queries must not expand an entire library into SQL bind parameters.
Mixed-folder lookups join existing references and discard exact same-title
rows before loading archive diagnostics; the final shared identity rules still
decide eligibility. Known-series recovery batches catalog lookups and fails
closed if saved or current ownership evidence changes after preview.

During a Mylar scan, Pullbox can also reconcile one stale recorded path with a
file found in another series folder when the embedded ComicInfo issue ID is an
Expand Down
5 changes: 3 additions & 2 deletions src/pullbox/api/v1/downloads.py
Original file line number Diff line number Diff line change
Expand Up @@ -7,14 +7,15 @@

from pullbox.api.deps import AuthenticatedUser, DbSession
from pullbox.core.acquisition import AcquisitionProtocol
from pullbox.core.exceptions import NotFoundError
from pullbox.core.exceptions import NotFoundError, ProviderError
from pullbox.models.blocklist import BlocklistReason
from pullbox.models.direct_acquisition import DirectAcquisitionAttempt
from pullbox.models.download import DownloadClientType, DownloadHistory, DownloadState
from pullbox.models.issue import Issue, IssueStatus
from pullbox.providers.base import DownloadClient, ProviderRegistry
from pullbox.providers.download.qbittorrent import QBittorrentError
from pullbox.providers.indexer.newznab import NewznabError
from pullbox.providers.indexer.prowlarr import ProwlarrError
from pullbox.schemas.blocklist import BlocklistEntryResponse
from pullbox.schemas.download import (
DirectSourceAlternative,
Expand Down Expand Up @@ -853,7 +854,7 @@ async def retry_download(
status_code=409,
detail=f"Unsupported acquisition protocol: {protocol.value}",
)
except (NewznabError, QBittorrentError) as exc:
except (NewznabError, ProwlarrError, QBittorrentError, ProviderError) as exc:
logger.warning(
"download_retry_client_rejected",
download_id=download.id,
Expand Down
16 changes: 14 additions & 2 deletions src/pullbox/api/v1/issues.py
Original file line number Diff line number Diff line change
Expand Up @@ -24,6 +24,7 @@
from pullbox.models.client import DownloadClientConfig
from pullbox.models.direct_acquisition import DirectAcquisitionAttempt
from pullbox.models.download import DownloadClientType, DownloadState
from pullbox.models.indexer import IndexerConfig
from pullbox.models.issue import Issue, IssueStatus, IssueType
from pullbox.models.library import LibraryFile, MatchConfidence
from pullbox.models.search_log import SearchLog, SearchType
Expand Down Expand Up @@ -765,8 +766,19 @@ async def grab_release(

raise ProviderError("download", "No download clients configured")

download_svc, indexer_configs = built
if body.indexer_id is not None and body.indexer_id not in indexer_configs:
download_svc, _health_configs = built
# Aggregated Prowlarr sources intentionally do not have per-indexer health entries.
if (
body.indexer_id is not None
and await session.scalar(
select(IndexerConfig.id).where(
IndexerConfig.id == body.indexer_id,
IndexerConfig.enabled.is_(True),
IndexerConfig.manager_available.is_(True),
)
)
is None
):
raise ValidationError(
"The originating indexer is no longer available. Run the search again."
)
Expand Down
2 changes: 2 additions & 0 deletions src/pullbox/composition/providers.py
Original file line number Diff line number Diff line change
Expand Up @@ -143,6 +143,8 @@ async def register_indexers(
indexer_rankings=prowlarr_indexer_rankings,
)
registry.register_indexer(_PROWLARR_AGGREGATE_CONFIG_ID, aggregate)
for config_id, _priority in prowlarr_indexer_rankings.values():
registry.register_indexer_alias(config_id, _PROWLARR_AGGREGATE_CONFIG_ID)
logger.debug(
"prowlarr_aggregate_registered",
indexer_count=len(prowlarr_torznab_ids),
Expand Down
89 changes: 89 additions & 0 deletions src/pullbox/core/torrent_metadata.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,89 @@
"""Bounded bencode validation and exact torrent metadata hashes."""

from __future__ import annotations

import hashlib


def torrent_info_hashes(content: bytes) -> frozenset[str]:
"""Validate one descriptor and hash its original info bytes without re-encoding."""
info_start, info_end = _top_level_info_span(content)
info = content[info_start:info_end]
return frozenset(
(
hashlib.sha1(info, usedforsecurity=False).hexdigest(),
hashlib.sha256(info).hexdigest(),
)
)


def _top_level_info_span(content: bytes) -> tuple[int, int]:
if not content or content[0] != ord("d"):
raise ValueError("torrent descriptor must be a dictionary")
index = 1
info_span: tuple[int, int] | None = None
while index < len(content) and content[index] != ord("e"):
key, index = _parse_bencoded_bytes(content, index)
value_start = index
index = _skip_bencoded_value(content, index, depth=1)
if key == b"info":
if info_span is not None:
raise ValueError("torrent descriptor has duplicate info dictionaries")
info_span = (value_start, index)
if index >= len(content) or content[index] != ord("e") or index + 1 != len(content):
raise ValueError("torrent descriptor is truncated or has trailing data")
if info_span is None or content[info_span[0]] != ord("d"):
raise ValueError("torrent descriptor has no info dictionary")
return info_span


def _skip_bencoded_value(content: bytes, index: int, *, depth: int) -> int:
if depth > 100 or index >= len(content):
raise ValueError("invalid bencode nesting")
marker = content[index]
if 48 <= marker <= 57:
_, end = _parse_bencoded_bytes(content, index)
return end
if marker == ord("i"):
end = content.find(b"e", index + 1)
if end < 0:
raise ValueError("unterminated bencoded integer")
value = content[index + 1 : end]
digits = value[1:] if value.startswith(b"-") else value
if (
not digits
or not digits.isdigit()
or (len(digits) > 1 and digits.startswith(b"0"))
or value == b"-0"
):
raise ValueError("invalid bencoded integer")
return end + 1
if marker not in {ord("l"), ord("d")}:
raise ValueError("invalid bencoded value")
cursor = index + 1
while cursor < len(content) and content[cursor] != ord("e"):
if marker == ord("d"):
_, cursor = _parse_bencoded_bytes(content, cursor)
cursor = _skip_bencoded_value(content, cursor, depth=depth + 1)
if cursor >= len(content):
raise ValueError("unterminated bencoded collection")
return cursor + 1


def _parse_bencoded_bytes(content: bytes, index: int) -> tuple[bytes, int]:
colon = content.find(b":", index)
if colon < 0:
raise ValueError("invalid bencoded byte string")
raw_length = content[index:colon]
if (
not raw_length
or not raw_length.isdigit()
or (len(raw_length) > 1 and raw_length.startswith(b"0"))
):
raise ValueError("invalid bencoded byte-string length")
length = int(raw_length)
start = colon + 1
end = start + length
if end > len(content):
raise ValueError("truncated bencoded byte string")
return content[start:end], end
8 changes: 7 additions & 1 deletion src/pullbox/providers/base.py
Original file line number Diff line number Diff line change
Expand Up @@ -360,6 +360,7 @@ def __init__(self) -> None:
self._download_clients: dict[int, DownloadClient] = {}
self._download_priorities: dict[int, int] = {}
self._indexers: dict[int, Indexer] = {}
self._indexer_aliases: dict[int, int] = {}

def register_metadata_provider(self, name: str, provider: MetadataProvider) -> None:
self._metadata_providers[name] = provider
Expand All @@ -374,8 +375,13 @@ def register_download_client(
self._download_priorities[config_id] = priority

def register_indexer(self, config_id: int, indexer: Indexer) -> None:
self._indexer_aliases.pop(config_id, None)
self._indexers[config_id] = indexer

def register_indexer_alias(self, config_id: int, aggregate_id: int) -> None:
"""Resolve a persisted source ID without adding another search request."""
self._indexer_aliases[config_id] = aggregate_id

def get_metadata_provider(self, name: str = "comicvine") -> MetadataProvider:
"""Get a metadata provider by name."""
return self._metadata_providers[name]
Expand All @@ -394,7 +400,7 @@ def get_indexers(self) -> list[Indexer]:

def get_indexer(self, config_id: int) -> Indexer | None:
"""Get one registered indexer by its persisted config ID."""
return self._indexers.get(config_id)
return self._indexers.get(self._indexer_aliases.get(config_id, config_id))

def get_indexer_items(self) -> list[tuple[int, Indexer]]:
"""Get (config_id, indexer) pairs for all registered indexers."""
Expand Down
Loading
Loading