Repository navigation
Conversation
Account for the Hermitian overlap derivative for nspin=4 as well as nspin=2. Add a converged Fe2 component regression with an independently differentiated frozen-state reference and document remaining total-force finite-difference limits.
aboys-cb
added a commit
to MagTheoryLab/abacus-develop
that referenced
this pull request
Oct 10, 2026
Apply the source fix from upstream PR deepmodeling#8117 to c516. Include the Hermitian overlap-derivative partner for nspin=4 as well as nspin=2; the factor two follows from the product rule, not spin degeneracy. Validation on Sai V100 with the rebuilt c516 executable: - 14 independent Fe2 SOC+DeltaSpin SCFs passed strict convergence gates. - 100 Ry Fe2 x force error: 2.52e-5 eV/Angstrom at h=0.0005 Angstrom. - 200 Ry force errors: 5.12e-6 and 8.51e-6 at h=0.00025 and 0.000125. - Retain the coarse 200 Ry result: 1.28e-4 at h=0.0005, above tolerance. - Fe2 z magnetic-force errors: 1.40e-5 and 1.49e-6 eV/muB at 100/200 Ry. - git diff --check and agent governance checks passed. The finite-difference results cover Fe2 x and Mz, not all components or physical convergence with respect to cutoff and basis. No INPUT changes.
7 tasks done
mohanchen
reviewed
Oct 10, 2026
hujieting
approved these changes
Oct 10, 2026
mohanchen
reviewed
Oct 10, 2026
mohanchen
reviewed
Oct 10, 2026
| scf_thr 1.0e-5 | ||
| scf_nmax 100 | ||
| out_chg 0 | ||
| ecutwfc 100 |
Collaborator
There was a problem hiding this comment.
maybe you can set up a smaller value to accelerate the calculations?
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Reminder
AGENTS.mdanddocs/developers_guide/agent_governance.md.source/changes.Linked Issue
No separate issue: this PR documents and fixes the force-normalization regression introduced by #7513 (
698e1876211b9fdf4f334fa8ac3efa40b9e4801f, 2026-06-30). Base: upstreamdevelopat97697c26b4562faf7fbcfffd75ba970c635d945a; no downstream optimization commits are included.Unit Tests and/or Case Tests for my changes
Update the existing
tests/03_NAO_multik/scf_deltaspin4case. The regular03_NAO_multikCI job already selects this case throughCASES_CPU.txt. There is no new case directory, orbital asset, Python test runner, shared comparison logic, or CTest entry.The case uses the existing
Fe.upfandFe_gga_6au_100Ry_4s2p2d1f.orbfromtests/PP_ORB. It is periodic Fe2 in a 2.86 Angstrom cubic cell, LCAO + SOC (nspin 4), 12 Ry, a 2x2x2 k mesh, no Hubbard U, and target moments (0,0,3.0) and (2.59807621135332,0,1.5) Bohr magnetons: both have magnitude 3.0, with a 60-degree relative tilt. Fe1 is at the origin; Fe2 is at (1.573,1.287,1.430) Angstrom, displaced by (+0.143,-0.143,0) from the body center. The geometric displacement produces nonzero atomic forces, while the magnetic magnitude and direction perturbation exercise the DeltaSpin constraint-force contribution.The old case deliberately accepted an unconverged 100-step SCF with tolerances of 1 eV, 10 eV/Angstrom and 500 kbar. Those thresholds could hide the force error. The lightweight CI case uses
ecutwfc 12andscf_nmax 10to bound its cost. SCF convergence and basis/grid convergence are not acceptance criteria for this output regression. Its regeneratedresult.refchecks reproducibility at those exact settings, with the force threshold kept at 1e-6 eV/Angstrom. Energy/stress use the standard 1e-7 eV / 1e-3 kbar defaults. The short, extensionlessREADMEhas two lines.The unchanged
Autotest.shreads reference values from the case'sresult.ref. Itstotalforcerefis the sum of absolute values of the final printed atomic-force components, not the net vector force. This is a standard output regression, not a finite-difference calculation inside CI. Independent derivative evidence is reported separately below.Current lightweight regression: 12 Ry, at most 10 SCF steps
Per review, restore the original benchmark cutoff (
ecutwfc 100 -> 12) and reducescf_nmax 240 -> 10to bound runtime. All other computational inputs, assets, comparison tolerances and shared test scripts are unchanged in this follow-up. The README still has two lines. The initial 12 Ry / 240-step probe was stopped after observing charge oscillation; retaining 240 steps would defeat the low-cost regression purpose.Job
1730113, Sai16v100n08,16V100/flood-1o2gpu, completed successfully. The CPUscalapack_gvxone-rank reference and independent four-rank corrected run both completed exactly 10 steps without charge convergence, as allowed for this test. OMP/OpenBLAS=1. The checked fixed/original executable SHA256 values are the same as those recorded below. The sharedAutotest.shcommands below were rerun unchanged at the new input settings.etotref(eV)etotperatomref(eV)totalforceref(eV/Angstrom)totalstressref(kbar)These reference changes follow the deliberately different cutoff and finite SCF trajectory; they are not converged physical predictions or a force-only code change in SCF energy.
totaltimerefis refreshed from 142.51 to 27.24 seconds and is not a numerical acceptance assertion. Observed runtimes were 27.24 seconds (one rank) and 29.70 seconds (four ranks), single observations on this node, not a controlled performance speedup claim.The four-rank corrected run passed every standard assertion: energy differed from the one-rank reference by 1.34e-9 eV, aggregate force was identical at six decimals, and aggregate stress differed by 2e-6 kbar. The original buggy executable at four ranks reproduced the corrected energy and stress exactly but gave
totalforceref=94.929118; its force error against the new reference is 1.243250 eV/Angstrom, and the standard runner returned 1 (expected force-only failure). Thus lowering the cutoff and limiting the iteration count preserve detection of the missing factor of two, without relaxing tolerances.The earlier converged 100 Ry controls and independent derivative evidence below remain separate. No new 12 Ry GPU or finite-difference claim is made.
git diff --checkand the staged governance checker pass for this follow-up.Earlier 100 Ry converged negative control (separate physical validation)
Slurm job
1726202, node16v100n01: all SCFs passed the convergence audit. The one-rank reference converged in 37 steps; both four-rank runs converged in 35 steps.totalforceref(eV/Angstrom)totalstressref(kbar)The independently cold-started one-rank reference gives
totalforceref=7.902208, identical to the four-rank fixed result at the collector's six-decimal precision. Its total energy differs from the four-rank result by 5.40e-10 eV. The original-condition run fails on force only, with an aggregate difference of 0.115626 eV/Angstrom (115,626 times the force threshold); energy and stress pass. Thus restoring the bug is demonstrably detected by the existing shared comparison flow.The stronger constraint was chosen using a controlled magnetic perturbation: at the same geometry and 60-degree tilt, increasing both target magnitudes from 2.2 to 3.0 Bohr magnetons increases Fe1's DeltaSpin x force from 0.0084497251 to 0.0579262041 eV/Angstrom (6.86x). The force contribution itself is measured here, rather than assuming a larger total atomic force implies a stronger DeltaSpin signal.
Executed from the staged
tests/03_NAO_multikdirectory, with the repository's unmodified integration scripts and the final case name:Reference generation uses the fixed executable at one rank; validation uses a fresh cold start at four ranks. The original executable is a matched negative control with upstream
dspin_fs.cpp. Both use the unmodified upstream AO implementation and the same base objects and link settings. The check also audits the SCF-converged marker, DRHO < 1e-10, constrained-moment RMS < 1.01e-10 (printed rounding), and |DeltaE_womix| < 1e-9 eV. Those extra convergence audits are external validation, not new logic in the shared CI script.Environment: Sai
16V100/flood-1o2gpu, CPUscalapack_gvx, one or four MPI ranks, OMP/OpenBLAS=1; GCC 13.3.0, CUDA 12.9, OpenMPI 5.0.10 and Libxc 7.0.0. Reused the previously built/verifiedv3.11.0-beta10executables because this follow-up changes only case data/layout. No performance comparison is asserted.Fixed executable SHA256:
97f9ec738013190ff33e0d4012b56737fcce8f9099469704f5e349c52da6dbee.Original executable SHA256:
194b6e94a96df0ee9fe51b3b42964bdfe51136ddac30c6bc5f073bb714c942a3.Independent derivative at 100 Ry (separate physical validation)
Slurm job
1726205, node16v100n01, one V100 GPU / one MPI rank, OMP/OpenBLAS=1,cusolver. This diagnostic uses the same geometry, orbital, k mesh and target moments, but retains the earlier 100 Ry cutoff and converged SCF. It is independent of the current 12 Ry / 10-step CI case; the device/solver and asset paths also differ. A cold-start SCF converged in 35 steps (DRHO=6.1242e-11, moment RMS=9.5650e-11, DeltaE_womix=-2.5522e-11 eV).For Fe1 x, the DeltaSpin contribution is:
The finite-difference errors relative to the fixed expression are 1.88e-9 and 6.82e-9 eV/Angstrom. The two explicit analytical overlap terms and the fixed production expression agree within 1.4e-14 eV/Angstrom. This independently supports the product-rule normalization rather than merely observing that the new output is twice the old output.
The diagnostic executable was checked against SHA256
2ee90c3678159f9d676870b8e425bbd467990f37518efd5519ca100640f19ab2before execution. It was built from the same fixed source with only a separately instrumented force translation unit.The separate diagnostic holds the converged density matrix D and constraint field lambda fixed, recomputes the two projector overlaps at displaced coordinates, and central-differences their contracted trace. It also evaluates both analytical overlap derivatives explicitly. It is not part of the generic test runner and introduces no diagnostic code into this PR. The CI result.ref is an output regression reference at 12 Ry / 10 steps; its values are not finite-difference reference values. The converged 100 Ry derivative evidence remains separate.
Local layout and governance checks
Registration-only CMake/CTest inspection passed: the existing
03_NAO_multiktest invokesAutotest.sh, and its unchanged CPU case list containsscf_deltaspin4. Byte comparisons confirm the shared integration scripts, case lists and category CMake files are unchanged from the PR base. The README has exactly two lines. For that earlier 100 Ry validation, INPUT/STRU/KPT/threshold matched the tested files byte-for-byte. The current 12 Ry / 10-step validation is recorded above. Whitespace, staged governance, and full base-to-head governance checks passed with no findings. This registration harness does not claim a full repository build or full category runtime test.Verification limits and negative results
In earlier validation on the smaller-displacement geometry (not the updated regression above), full-SCF total-energy differences at h=0.0005 Angstrom did not pass a 1e-4 eV/Angstrom force tolerance on the fixed upstream-based executable. The mean net-force subtraction was undone when comparing an individual atom to its energy derivative:
For the relative displacement,
R1x += q/2,R2x -= q/2, and the analytical derivative is(F1x-F2x)/2. For Fe2-only displacement the raw force isF2x_printed + net_force_x/2.All SCFs met the electronic/magnetic convergence gates. The CPU/GPU agreement rules out calling this a GPU-only residual. These full-force discrepancies are not claimed fixed by this PR. The total-force regression above detects changes in the final printed force, while the frozen-state derivative isolates the diagnosed DeltaSpin error. Neither establishes full total-energy/force consistency for production relaxation or MD. The existing loose thresholds are tightened, not relaxed.
Full repository CI, PW finite differences, a broader scan of noncollinear directions, stress finite differences and ROCm execution were not run. This is a correctness fix, with no performance claim.
What's changed?
For
basis_type lcao,sc_mag_switch 1andnspin 4, the DeltaSpin atomic-force contribution was half of its complete value. Total forces are affected whenever this contribution is nonzero; the entire total force is not halved.At fixed D and lambda, the relevant trace contains products of orbital-projector overlaps:
cal_force_IJRaccumulates only the first derivative. To derive the missing factor, write the contracted trace as E = Re sum C_mu,nu B_mu* B_nu, with C_nu,mu = C_mu,nu* (the spin/constraint indices are implicit). Its derivative is Re(S1 + S2), where S1 = sum C_mu,nu (dB_mu*) B_nu and S2 = sum C_mu,nu B_mu* (dB_nu). Exchanging mu and nu in S2 and using Hermiticity gives S2 = S1*. Hence dE/dR = 2 Re(S1), and F = -2 Re(S1). Both terms survive in the Pauli representation as well. The final factor of two accounts for the two overlap derivatives, not spin degeneracy. Transforming D to Pauli components does not remove either derivative.#7513 restricted the final multiplication to
nspin != 4, reasoning that the Pauli basis already covers the spin channels. This conflated spin counting with the product rule. This PR restores unconditional multiplication by two and explains why in a code comment. The other Pauli-to-spinor corrections from #7513 are retained.Scope:
nspin=2retains its existing multiplication.Governance Notes
docs/parameters.yamlorinput-main.mdis required. Existing parameters are used for a bounded-cost force regression; the current case uses 12 Ry and at most 10 SCF steps without requiring convergence. Case references changed because the geometry, SOC, cutoff and iteration settings changed; the before/after numerical evidence is above.nspin=2and stress retain their existing formulas.