Skip to content

Simulation state is process-wide, not per-simulation #175

Description

@HugoFara

Question for @filippi / @antonio-leblanc: should one process be able to run several simulations (for ensembles, #172, or free-threaded Python)? If not, we merge the small fixes below and stop there.

ForeFire keeps simulation state in statics and a singleton: SimulationParameters::instance, FireDomain::propModelsTable/fluxModelsTable, the FireNode numerical settings, ForeFireAtom::instanceNRCount and StringRepresentation::outputstr. The GIL hides this today. Without it, 8 concurrent simulations segfault in 4 out of 5 runs, and the fifth gives wrong node counts (27–209 instead of 62).

Plan:

  1. Stress test (Add the concurrency stress test that reproduces #175 #176, merged)
  2. Atomic id counter and thread-safe singleton (Make the atom id counter atomic and the parameters singleton safe #177, merged)
  3. One output buffer per StringRepresentation (Give each string representation its own buffer #178, merged)
  4. A lock around NetCDF/HDF5

Steps 1–4 are correctness fixes. They are not enough for thread safety on their own.

  1. Move state into each domain. This changes FireDomain, DataBroker, Command and CLibForeFire, so the coupling ABI that Meso-NH uses must be agreed first.
  2. Ship cp314t wheels again (reverses part of Stop shipping free-threaded (cp314t) wheels #155), only after step 5.

Drafted with Claude Opus 5, reviewed by a maintainer.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions