Skip to content

fix(collector): keep collecting temperatures without CPU power status - #3767

Open
SaiPisey2 wants to merge 2 commits into
prometheus:masterfrom
SaiPisey2:fix/thermal-darwin-no-cpu-power-status
Open

fix(collector): keep collecting temperatures without CPU power status#3767
SaiPisey2 wants to merge 2 commits into
prometheus:masterfrom
SaiPisey2:fix/thermal-darwin-no-cpu-power-status

Conversation

@SaiPisey2

Copy link
Copy Markdown

Addresses #2906.

Apple Silicon does not implement IOPMCopyCPUPowerStatus, so fetchCPUPowerStatus gets back kIOReturnNotFound. Update returned that error straight away, which meant it never reached updateTemperatures, so no temperature metrics were collected at all. The error also isn't ErrNoData, so it was logged at error level on every scrape.

Measured on an M5 Pro (darwin/arm64) against master:

fetchCPUPowerStatus()      -> status=map[] err=no CPU power status has been recorded
Update()                   -> err=no CPU power status has been recorded ; metrics emitted=0
IsNoDataError(err)         -> false
updateTemperatures() alone -> err=<nil> ; temperature metrics available=52

52 usable temperature sensors on that machine, none of them exported, because an expected and unrelated condition aborted the collector.

Change

Treat kIOReturnNotFound as "this system does not report CPU power status" rather than a failure: skip the three CPU power metrics, log at debug level, and continue on to the temperature sensors. Any other non-success return code is still returned as an error, exactly as before, so systems that do report CPU power status are unaffected.

After the change, on the same machine, Update() returns no error and emits all 52 temperature metrics.

Scope

This does not make the CPU power metrics appear on Apple Silicon and does not resolve #2218 — the underlying API provides no data there, as @rexagod already established on #2906. It only stops their absence from suppressing the temperature metrics. I've left Closes #2906 out of the commit deliberately, since whether that issue is fully answered by this is a maintainer call.

Testing

Added collector/thermal_darwin_test.go, which fails on master:

--- FAIL: TestThermalUpdateWithoutCPUPowerStatus
    thermal_darwin_test.go:44: Update returned errNoCPUPowerStatus; a system
    without CPU power status must still collect temperatures

and passes with the change. Also verified locally:

  • go build ./...
  • go vet ./collector/
  • gofmt -l clean on both touched files
  • go test ./collector/ full package
  • go build -tags notherm ./collector/

@nicolastakashi

Copy link
Copy Markdown
Contributor

/workflow-approve

@nicolastakashi

Copy link
Copy Markdown
Contributor

The Darwin/macOS e2e job is failing on a fixture mismatch, node_scrape_collector_success{collector="thermal"} changed from 0 to 1. Can you regenerate the golden output on a Darwin/arm64 box and push it?

./end-to-end-test.sh -u

Then commit the updated collector/fixtures/e2e-output-darwin.txt.

Apple Silicon does not implement IOPMCopyCPUPowerStatus, so
fetchCPUPowerStatus returns kIOReturnNotFound there. Update returned that
error straight away, which aborted the collector before updateTemperatures
ran, so no temperature metrics were collected at all. The error was also
not ErrNoData, so it was logged at error level on every scrape.

On an M5 Pro, Update emitted 0 metrics and failed, while updateTemperatures
on its own returned 52 temperature metrics.

Treat kIOReturnNotFound as a system that does not report CPU power status:
skip the three CPU power metrics, log at debug level and carry on to the
temperature sensors. Any other non-success return code is still returned as
an error. Systems that do report CPU power status are unaffected.

This does not make the CPU power metrics available on Apple Silicon, since
the underlying API provides no data. It only stops their absence from
suppressing the temperature metrics.

Adds a regression test covering the case.

Signed-off-by: SaiPisey2 <piseysai0202@gmail.com>
@SaiPisey2
SaiPisey2 force-pushed the fix/thermal-darwin-no-cpu-power-status branch from 8d1e633 to 907bf2f Compare August 6, 2026 10:22
@SaiPisey2

Copy link
Copy Markdown
Author

Thanks for the review. Pushed, but regenerating the fixture turned up two things worth flagging rather than just committing the output.

The fixture change itself

I did not commit a regenerated file. Running ./end-to-end-test.sh -u on real hardware produced a 469-line diff, because this machine reports actual sensors while the CI runner reports none, and the values are live temperatures. So the fixture now carries only the single line the CI diff asked for:

-node_scrape_collector_success{collector="thermal"} 0
+node_scrape_collector_success{collector="thermal"} 1

./end-to-end-test.sh -u overwrites the Linux fixture on macOS

Worth knowing before anyone else follows the same instruction. The Darwin fixture path is built with ${fixture_metrics::-4}, and a negative substring length needs bash 4.2. macOS still ships bash 3.2, where that expansion fails, fixture_metrics keeps pointing at e2e-output.txt, and the update copies Darwin output over the Linux fixture. My first run rewrote 5330 lines of collector/fixtures/e2e-output.txt before I noticed. ${fixture_metrics%.txt} is equivalent and works on both, so I've changed it.

I also added node_thermal_temperature_celsius to non_deterministic_metrics. It is per-machine in both its sensor names and its values, so without it anyone running the suite on real hardware sees the whole sensor list as a diff. It makes no difference in CI, which reports no sensors.

(Unrelated and left alone: the sed -i /pattern/d calls in that loop need an extension argument on BSD sed, so the non-deterministic stripping aborts on stock macOS. It behaves the same on master, so it is not something this PR introduces.)

A real bug the fixture regeneration exposed

This is the part I would not have caught otherwise. Once temperatures are actually collected, the collector emits duplicate label sets:

emitted=52  uniqueSensorNames=17  duplicateSeries=35

and the scrape logs 35 error(s) occurred: ... was collected before with the same name and label values every time, with those samples dropped.

A sensor is identified only by its IOHID Product name, and several services report the same one. Probing the services directly, 76 temperature-capable services carry just 25 distinct Product values, and there is nothing to tell them apart: RegistryID, UniqueID, SerialNumber and Manufacturer are all unset, VendorID/ProductID/PrimaryUsage are constant across every service, and two services can share both Product and LocationID (PMU tdie2 at 1414541922 appears twice). There are also genuinely distinct sensors sharing a name — two gas gauge battery services with different LocationIDs.

Since there is no property that makes them addressable, I report the first reading per name and count the rest in a debug line. The endpoint now returns 200 with 17 unique series and no gather errors. TestThermalTemperaturesAreUnique covers it and fails without the change.

CI never sees this, since the runner has no sensors, but every Apple Silicon machine would have.

Verified locally: go build ./..., go vet ./collector/, gofmt clean, go test ./collector/ including 20 repeats of the thermal tests, and a -tags notherm build.

Happy to split the two end-to-end-test.sh changes into their own PR if you would rather keep this one to the collector.

@nicolastakashi

Copy link
Copy Markdown
Contributor

Thanks for the fixes. Let's split the end-to-end-test.sh changes into a separate PR.

On the duplicate sensors: dropping readings loses data. #3646 hit the same colliding-label problem for hwmon and fixed it by disambiguating the label instead (suffix it, falling back to something always unique when the first suffix still collides). Take a look there for the approach.

Collecting temperatures on Apple Silicon surfaced a second problem that
was previously unreachable, because Update returned before
updateTemperatures ever ran.

A sensor is labelled with its IOHID "Product" property, which is not
unique. On an M5 Pro, 52 temperature-reporting services carry only 17
distinct product names, so the registry rejects the colliding samples and
logs an error on every scrape.

The collisions are of two kinds. Six "gas gauge battery" services are
distinct sensors that merely share a name; they have distinct
LocationIDs. The "PMU tdie*" services come in threes that share both the
name and the LocationID, and report readings that differ by a few tenths
of a degree, so they are separate sensing elements rather than the same
one listed repeatedly. Dropping either kind loses readings.

Following the approach taken for hwmon in prometheus#3646, read all sensors first,
then qualify a name shared by several services with its location, falling
back to the service registry ID when the location does not separate them
either. Registry IDs are unique per service, which the OS guarantees and
which held for every service observed here. Names that do not collide are
left untouched.

The endpoint now reports all 52 readings with unique label sets and no
gather errors, where before the change 35 of them were rejected.

Also update collector/fixtures/e2e-output-darwin.txt for
node_scrape_collector_success{collector="thermal"}, which is now 1
because the collector no longer fails.

resolveSensorNames is kept free of cgo so that the labelling rules are
covered by table-driven tests, which matters because the CI runner
reports no thermal sensors at all.

Signed-off-by: SaiPisey2 <piseysai0202@gmail.com>
@SaiPisey2

Copy link
Copy Markdown
Author

Both done. The end-to-end-test.sh changes are now #3798, and this PR is back to the collector plus the one fixture line.

You were right that dropping loses data — more than I realised before looking properly. Probing the services directly, the collisions turn out to be two different things:

  • Six gas gauge battery services with six distinct LocationIDs. Genuinely different sensors that only share a product name.
  • Fifteen PMU * names, each reported by three services that share the product name and the LocationID. Those three are not the same reading twice: they have distinct registry IDs and report values a few tenths of a degree apart, so they are separate sensing elements.

So both kinds carry real readings, and the earlier "first one wins" would have thrown away 35 of 52.

Following #3646: read all sensors first, then qualify a colliding name with its location, falling back to the registry ID when the location does not separate them either. Names that do not collide are untouched.

IOHIDServiceClientGetRegistryID is the always-unique part — 52 distinct IDs for 52 services here, and 76 for 76 before the sub-absolute-zero readings are filtered. I had missed it earlier because I was looking for a RegistryID property, which is unset; it is a function on the service client.

Result on the same machine, where before the change 35 of 52 samples were rejected:

http=200  thermalSeries=52  uniqueLabelSets=52  gatherErrors=0
node_scrape_collector_success{collector="thermal"} 1
sensor="NAND CH0 temp"                        # unique, untouched
sensor="gas gauge battery_1413951554"         # ... and five more, by location
sensor="PMU tdie2_12724010287800361880"       # ... and two more, by registry ID

One deliberate choice worth flagging: the label resolution lives in resolveSensorNames, which takes a plain slice and has no cgo in it, so the rules are covered by table-driven tests — the collision-by-location case, the fallback, a missing location, a mix of both within one name, and the empty case. That matters here because the CI runner reports no thermal sensors, so anything that depends on real hardware does not actually get exercised there. The hardware-level check that no two label sets collide is still present as a separate test.

Registry IDs are not stable across reboots, so those labels will change if the machine restarts. That is the same trade-off as the hwmonX fallback in #3646, and it only affects sensors that cannot be told apart any other way.

Verified: go build ./..., go vet ./collector/, gofmt clean, go test ./collector/, -race, 30 repeats against live hardware, and a -tags notherm build.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

SIGTRAP: trace trap on M1

2 participants