-
Notifications
You must be signed in to change notification settings - Fork 159
Enable native code on AARCH64 #723
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
base: main
Are you sure you want to change the base?
Changes from all commits
05e2efe
ed9c97d
c002a11
d0e4e89
a54bb5f
7962fb7
5815aa0
9678d8b
96fc0a6
d6ad7ff
d8178f6
9e9da2b
e50a33e
200f726
5225643
a9a8504
1d4fd1a
589d257
068f286
7b8b6fb
e572c77
4966d4a
a4fb28a
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,137 @@ | ||
| name: Unit Test CI — ARM64 (NEON / SVE) | ||
|
|
||
| on: | ||
| workflow_dispatch: | ||
| pull_request: | ||
| push: | ||
| branches: | ||
| - main | ||
| paths: | ||
| - .github/workflows/unit-tests-arm64.yaml | ||
| - '**.java' | ||
| - '**/pom.xml' | ||
|
|
||
| jobs: | ||
| build-arm64: | ||
| concurrency: | ||
| group: arm64-${{ matrix.max_isa }}-${{ matrix.jdk }} | ||
| cancel-in-progress: false | ||
| strategy: | ||
| matrix: | ||
| jdk: [ 24 ] | ||
| # Three ISA tiers in ascending capability order, mirroring avx512f/avx2/sse42. | ||
| # GitHub-hosted ubuntu-24.04-arm is a Neoverse-N1 (Graviton 2): NEON only, no SVE. | ||
| # The sve/sve2 matrix entries still exercise the JVECTOR_MAX_ISA cap path and | ||
| # compile all three ISA variants; the native kernel tests that require actual SVE | ||
| # hardware are gated on the runtime feature check below. | ||
| max_isa: [ neon, sve, sve2 ] | ||
| runs-on: ubuntu-24.04-arm | ||
| steps: | ||
| - name: Report ARM64 ISA capabilities | ||
| id: cpu-features | ||
| run: | | ||
| # Parse the "Features" line from /proc/cpuinfo — the kernel only exposes a | ||
| # token here when the OS has set up context-switch support for it, so this is | ||
| # the same authority as getauxval(AT_HWCAP / AT_HWCAP2). | ||
| # "asimd" is the NEON token; "sve"/"sve2"/"sveaes" appear on Graviton 3/4. | ||
| flags="$(grep '^Features' /proc/cpuinfo | head -1 | cut -d: -f2)" | ||
| has_neon=false; has_sve=false; has_sve2=false | ||
| [[ " $flags " == *" asimd "* ]] && has_neon=true | ||
| [[ " $flags " == *" sve "* ]] && has_sve=true | ||
| [[ " $flags " == *" sve2 "* ]] && has_sve2=true | ||
| printf "NEON=%s SVE=%s SVE2=%s\n" "$has_neon" "$has_sve" "$has_sve2" | ||
| if [[ "$has_neon" != "true" ]]; then | ||
| echo "ERROR: NEON (asimd) not found in /proc/cpuinfo — not a valid AArch64 runner" | ||
| exit 2 | ||
| fi | ||
| # Expose as step outputs for conditional steps below. | ||
| echo "has_neon=$has_neon" >> "$GITHUB_OUTPUT" | ||
| echo "has_sve=$has_sve" >> "$GITHUB_OUTPUT" | ||
| echo "has_sve2=$has_sve2" >> "$GITHUB_OUTPUT" | ||
|
|
||
| - name: Set up GCC | ||
| run: | | ||
| sudo apt install -y gcc g++ | ||
|
|
||
| - name: Install Meson, Ninja, and GTest | ||
| run: | | ||
| sudo apt update && sudo apt install -y meson ninja-build pkg-config libgtest-dev | ||
|
|
||
| - uses: actions/checkout@v4 | ||
|
|
||
| - name: Initialize Git Submodules | ||
| run: git submodule update --init | ||
|
|
||
| - name: Build test_simd_kernels (native C++) | ||
| # Meson detects aarch64 and compiles all three ISA variants (neon/sve/sve2) | ||
| # regardless of what the host CPU supports at runtime. | ||
| working-directory: jvector-native/src/main/native | ||
| run: | | ||
| meson setup build --wipe | ||
| ninja -C build test_simd_kernels | ||
|
|
||
| - name: Run test_simd_kernels — no ISA cap (auto-detect, neon job) | ||
| if: matrix.max_isa == 'neon' | ||
| working-directory: jvector-native/src/main/native | ||
| run: ./build/test_simd_kernels | ||
|
|
||
| - name: Run test_simd_kernels — capped at neon (sve job, host may lack SVE) | ||
| if: matrix.max_isa == 'sve' | ||
| working-directory: jvector-native/src/main/native | ||
| env: | ||
| JVECTOR_MAX_ISA: neon | ||
| run: ./build/test_simd_kernels | ||
|
|
||
| - name: Run test_simd_kernels — no ISA cap on SVE hardware (sve job) | ||
| if: matrix.max_isa == 'sve' && steps.cpu-features.outputs.has_sve == 'true' | ||
| working-directory: jvector-native/src/main/native | ||
| run: ./build/test_simd_kernels | ||
|
|
||
| - name: Run test_simd_kernels — capped at neon (sve2 job baseline check) | ||
| if: matrix.max_isa == 'sve2' | ||
| working-directory: jvector-native/src/main/native | ||
| env: | ||
| JVECTOR_MAX_ISA: neon | ||
| run: ./build/test_simd_kernels | ||
|
|
||
| - name: Run test_simd_kernels — no ISA cap on SVE2 hardware (sve2 job) | ||
| if: matrix.max_isa == 'sve2' && steps.cpu-features.outputs.has_sve2 == 'true' | ||
| working-directory: jvector-native/src/main/native | ||
| run: ./build/test_simd_kernels | ||
|
|
||
| - name: Set up JDK ${{ matrix.jdk }} | ||
| uses: actions/setup-java@v3 | ||
| with: | ||
| java-version: ${{ matrix.jdk }} | ||
| distribution: temurin | ||
| cache: maven | ||
|
|
||
| - name: Verify native-access vector support (JDK ${{ matrix.jdk }}) | ||
| env: | ||
| JVECTOR_MAX_ISA: ${{ matrix.max_isa }} | ||
| run: >- | ||
| mvn -B -Punix-amd64-profile -pl jvector-tests -am test | ||
| -DTest_RequireSpecificVectorizationProvider=NativeVectorizationProvider | ||
| -Dsurefire.failIfNoSpecifiedTests=false | ||
| -Dtest=TestVectorizationProvider | ||
|
|
||
| - name: Test full suite with native vectorization (JDK ${{ matrix.jdk }}) | ||
| env: | ||
| JVECTOR_MAX_ISA: ${{ matrix.max_isa }} | ||
| run: >- | ||
| mvn -B -Punix-amd64-profile test | ||
| -DTest_RequireSpecificVectorizationProvider=NativeVectorizationProvider | ||
|
|
||
| - name: Test Summary for (ARM64/max:${{ matrix.max_isa }},JDK${{ matrix.jdk }}) | ||
| if: always() | ||
| uses: test-summary/action@v2 | ||
| with: | ||
| paths: | | ||
| **/target/surefire-reports/TEST-*.xml | ||
|
|
||
| - name: Upload Surefire Test Results | ||
| uses: actions/upload-artifact@v4 | ||
| if: always() | ||
| with: | ||
| name: surefire-results--arm64-${{ matrix.max_isa }}-${{ matrix.jdk }} | ||
| path: "**/target/surefire-reports/**" |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,91 @@ | ||
| ### Native SIMD Acceleration on AArch64 (NEON, SVE, SVE2) | ||
|
|
||
| **Description** | ||
|
|
||
| This PR extends JVector's native Highway SIMD backend to **AArch64 (64-bit ARM)**, bringing | ||
| the same native acceleration that x86-64 users already have to ARM-based servers and | ||
| development machines (e.g. AWS Graviton, Apple Silicon under Linux, Ampere Altra). | ||
|
|
||
| AArch64 support was previously absent from the native layer — ARM hosts silently fell back to | ||
| the Panama Vector API path. The native binary now ships three AArch64 ISA tiers, all compiled | ||
| as vector-length-agnostic (scalable) code: | ||
|
|
||
| | Tier | Highway target | Notes | | ||
| |---|---|---| | ||
| | NEON | `HWY_NEON` | Baseline; all AArch64 CPUs | | ||
| | SVE | `HWY_SVE` | Scalable Vector Extension | | ||
| | SVE2 | `HWY_SVE2` | SVE2 + `i8mm` + `bf16` | | ||
|
|
||
| SVE and SVE2 are compiled as VL-agnostic (scalable) code, so the same binary runs correctly | ||
| on any SVE/SVE2-capable CPU regardless of its physical vector length. NEON operates on fixed | ||
| 128-bit vectors. At JVM startup, the best available tier is selected automatically via CPU | ||
| feature detection — identical to the existing CPUID-based dispatch on x86-64. | ||
|
|
||
| The following changes were made: | ||
|
|
||
| - **`NativeVectorizationProvider`** — the hard-coded `x86_64` architecture check is extended | ||
| to include `aarch64`, so the provider is offered on ARM Linux hosts. | ||
| - **`jvector_arch.h`** — new header introducing `JV_ARCH_X86_64` / `JV_ARCH_AARCH64` | ||
| preprocessor macros; all arch-specific code is guarded with `#if`/`#elif`/`#endif`. | ||
| - **`jvector_cpu_features.h`** — AArch64 feature probing (NEON, SVE, SVE2 presence and | ||
| vector-length queries) added alongside the existing x86 CPUID path. | ||
| - **Meson build** — three AArch64 ISA variant targets added; SVE/SVE2 targets skip gracefully | ||
| on Clang < 22 or non-Linux hosts where SVE is unavailable. | ||
| - **Cross-compilation** — Meson cross-files and updated READMEs provided for building the | ||
| AArch64 native library from an x86-64 host. | ||
| - **CI** — a dedicated AArch64 GitHub Actions runner added to the matrix so native builds and | ||
| kernel correctness are verified on real ARM hardware on every PR. | ||
|
|
||
| **Purpose / Impact** | ||
|
|
||
| Benchmarked on 1M-scale datasets on both AWS Graviton 3 and Graviton 4, Native SIMD (Highway) | ||
| consistently outperforms Panama SIMD on AArch64: | ||
|
|
||
| - Up to **40% lower search latency** at the same recall level | ||
| - Up to **29% faster index construction** | ||
|
|
||
| - All kernels already ported to Highway for x86-64 (FP32 similarity, PQ, NVQ, element-wise | ||
| arithmetic) benefit immediately on AArch64 — no separate ARM implementation was required. | ||
| - On AArch64 Linux with native vectorization enabled, the Highway backend is used exclusively | ||
| for all similarity and quantization operations; the Panama Vector API path is bypassed. | ||
| - On platforms where the native library cannot be loaded, JVector falls back to the Panama or | ||
| pure-Java provider transparently, as before. | ||
|
|
||
| **How to Enable** | ||
|
|
||
| Both `libjvector.so` variants — x86-64 and AArch64 — are built and bundled into the release | ||
| JAR. At JVM startup, `NativeVectorizationProvider` detects the current architecture and loads | ||
| the matching native library automatically. The same flags work on both architectures: | ||
|
|
||
| ```bash | ||
| java --enable-native-access=ALL-UNNAMED \ | ||
| -Djvector.experimental.enable_native_vectorization=true \ | ||
| -jar your-app.jar | ||
| ``` | ||
|
|
||
| To cap the ISA tier for debugging or benchmarking: | ||
|
|
||
| ```bash | ||
| JVECTOR_MAX_ISA=neon java --enable-native-access=ALL-UNNAMED \ | ||
| -Djvector.experimental.enable_native_vectorization=true \ | ||
| -jar your-app.jar | ||
| ``` | ||
|
|
||
| **Building releases:** | ||
|
|
||
| To produce a release JAR with native libraries for all supported | ||
| architectures (x86-64 and AArch64), the build must be run with the `-Dnative.crossarch` | ||
| Maven property. Without this flag, only the native library for the current host architecture | ||
| is built and bundled. | ||
|
|
||
| ```bash | ||
| mvn package -Dnative.crossarch | ||
|
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Note that this requires installing additional packages, at least in some cases. On my test system I had to Is there any way we can simply add this as part of the standard
Contributor
Author
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. We could pass
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. sorry, I was unclear and mixed 2 things in my comment
Contributor
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. Clarification: |
||
| ``` | ||
|
|
||
|
|
||
| **Notes** | ||
|
|
||
| - SVE/SVE2 compilation requires Clang ≥ 22; older compilers silently skip those targets and | ||
| fall back to NEON. | ||
| - There is no change to the on-disk index format; AArch64 and x86-64 indexes are fully | ||
| interchangeable. | ||
Uh oh!
There was an error while loading. Please reload this page.