Skip to content

CI: unit and regression tests on a simulated GPU - #110

Merged
MartinPulec merged 1 commit into
CESNET:masterfrom
saqibkh:ci-simulated-gpu
Oct 6, 2026
Merged

MartinPulec merged 1 commit into
CESNET:masterfrom
saqibkh:ci-simulated-gpu

Conversation

@saqibkh

@saqibkh saqibkh commented Sep 30, 2026

Copy link
Copy Markdown
Contributor

ccpp.yml builds GPUJPEG on GitHub's runners, which have no GPU, so nothing that encodes or decodes runs in CI. This adds a workflow that runs the unit tests and the regression suite on a simulated NVIDIA T4, on pushes to devel and master and on pull requests.

How: PantheonSim, via pantheongpu/setup-pantheonsim, simulates the GPU on the runner's CPU. The kernels run from their PTX (CMAKE_CUDA_ARCHITECTURES=75 embeds it), so the results are real. There is no timing model, so this checks correctness, not performance.

Changes: one new workflow, .github/workflows/simulated-gpu.yml, on ubuntu-24.04:

  • set up Ubuntu's CUDA toolkit and a simulated T4;
  • install ImageMagick and ffmpeg, which the regression suite compares images with;
  • build with CMAKE_CUDA_ARCHITECTURES=75 and the shared CUDA runtime;
  • run ctest -R 'unittests|regression' under vgpu run, which puts the simulated GPU's libraries in place for the tests only. The build links against the toolkit's own libcudart, because libgpujpeg references the OpenGL interop calls even when OpenGL is off.

ccpp.yml is unchanged.

What it can touch: the action is pinned by commit (v0.1.1), not by a tag, and it pins the simulator it builds by commit too, so nothing in the job changes until you move the pin. The workflow's token is read-only (permissions: contents: read), and the action uses no token or secrets; it installs packages only from Ubuntu's and NVIDIA's repositories. The job isn't a required check, so a red run never blocks a merge unless you decide it should.

What it gives, from a run on my fork:

  • unittests and regression both pass; the tests take under a minute.
  • The first run in the repository builds the simulator; the action caches it, per pinned version, so later runs skip that.

Disclosure: I maintain PantheonSim. Running GPUJPEG's tests found three gaps in its CUDA runtime that are now fixed: copy directions, symbol copies from device memory, and the cache-preference calls. If the job is ever flaky or wrong, please open an issue at pantheongpu/pantheonsim and I'll fix it on our side.

🤖 Generated with Claude Code

ccpp.yml builds GPUJPEG on runners with no GPU, so nothing that encodes
or decodes runs in CI. A new workflow runs the unit tests and the
regression suite on a simulated NVIDIA T4 (PantheonSim, via
setup-pantheonsim), on pushes and pull requests.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@MartinPulec MartinPulec self-assigned this Oct 2, 2026
@MartinPulec
MartinPulec merged commit 393ae20 into CESNET:master Oct 6, 2026
@MartinPulec

Copy link
Copy Markdown
Collaborator

thanks, I think that PantheonSim is superb

because libgpujpeg references the OpenGL interop calls even when OpenGL is off.

I've already moved cudaGraphicsUnmapResources() inside OpenGL interop include guard so that it can be linked also with simulator libcudart if the interop is disabled.

As a consequence, vgpu run should IMO no longer be needed and now all tests, including test-colors can be run. test-colors is actually important for us because the coverage with unit tests and regression tests is a bit weak and the functional/CLI tests ("test-colors") have the slightly better coverage.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants