Environment
CloudNativePG 1.30.0, plugin-barman-cloud v0.14.0, PostgreSQL 17/18 clusters, VersityGW (posix) S3 endpoint.
What happens
The instance sidecar collects plugin metrics over the plugin gRPC with no deadline on the collect path. Consequence: if any plugin operation gets stuck (in my case a retention/catalog-maintenance run that never completes against a posix-backend S3 gateway — filed separately), the next scrape blocks on the collector and the entire /metrics endpoint of that instance hangs indefinitely. Prometheus marks the target down and every metric of that instance disappears, not just the plugin's own collectors.
Observed in production: the primary instance's exporter dark for 6.7 hours while the database itself was healthy — WAL-archiver and backup-staleness alerting was blind precisely on the instance where it matters.
Expected
- A per-collect deadline (a few seconds) on the plugin metrics gRPC call.
- On timeout/error: return the rest of the metrics plus an error counter (e.g.
..._collector_errors_total), instead of hanging the whole endpoint. A misbehaving plugin should degrade its own metrics, never the instance's.
Reproduction sketch
- ObjectStore against VersityGW (posix backend) with any
retentionPolicy set.
- Wait for a retention run to start (~30 min cadence) — it never completes against this backend.
- Scrape the instance sidecar's metrics port: the request hangs until the client gives up;
up goes to 0 for the instance.
Companion issue describes the retention hang itself; this one is about the missing deadline that turns any such hang into a full metrics outage.
Environment
CloudNativePG 1.30.0, plugin-barman-cloud v0.14.0, PostgreSQL 17/18 clusters, VersityGW (posix) S3 endpoint.
What happens
The instance sidecar collects plugin metrics over the plugin gRPC with no deadline on the collect path. Consequence: if any plugin operation gets stuck (in my case a retention/catalog-maintenance run that never completes against a posix-backend S3 gateway — filed separately), the next scrape blocks on the collector and the entire
/metricsendpoint of that instance hangs indefinitely. Prometheus marks the target down and every metric of that instance disappears, not just the plugin's own collectors.Observed in production: the primary instance's exporter dark for 6.7 hours while the database itself was healthy — WAL-archiver and backup-staleness alerting was blind precisely on the instance where it matters.
Expected
..._collector_errors_total), instead of hanging the whole endpoint. A misbehaving plugin should degrade its own metrics, never the instance's.Reproduction sketch
retentionPolicyset.upgoes to 0 for the instance.Companion issue describes the retention hang itself; this one is about the missing deadline that turns any such hang into a full metrics outage.