Conversation
|
I did this too, actually. But an easier way to get the benchmarks running is to push your branch, then run ASV benchmarks directly in the Actions tab. That way you don't have to mess with a PR to get them running. |
ASV BenchmarkingBenchmark Comparison ResultsBenchmarks that have improved:
Benchmarks that have stayed the same:
Benchmarks that have got worse:
|
Ah, that makes sense! Does that lead to a nice comment with the benchmark results too? I suppose even if there isn't a comment it's definitely less noisy than making a draft PR |
|
Nah, you'd have to dig. I've been mostly looking at total time though, so for me it doesn't matter. |
|
@cmdupuis3 I've been noticing that the "import uxarray" benchmark is always showing an improvement, and the EDIT: Ah, I agree now I was wrong about the latter… I felt like I've been seeing NeighborhoodDask benchmarks in that category in a few different PRs, but now I can see that it isn't actually always showing up as "got worse", only sometimes! |
|
@Sevans711 If it's around 300M on a memory benchmark, that usually indicates caching behavior. I think the pure import is pretty indicative. It can be hard to measure consistently under caching conditions, so for the other benchmarks, maybe they aren't rigorous enough. The other one seems like a one-off, you can just compare it with previous benchmark runs and if there's no pattern, I wouldn't worry about it much. |

Not intended to be merged to main, just trying to check how ASV benchmarks look currently when the code is unchanged.
There's probably a better solution than having a draft PR for this, but this is the simplest/fastest way I could think of right now. I already attempted but struggled to generate ASV benchmark runs on my local machine with the same output format as what the run-benchmark label produces via github CI.
I originally created this with the intention of determining whether the reported benchmark performance degradation in #1705 is actually caused by the code changes there or not. Maybe this should be closed soon after that check, or maybe it should be open for a longer-term basis to repeat checks like this in the future? For example, it might have been useful to have something like this available for initially discovering #1605.