Skip to content

Open OnDemand integration blueprint - #951

Open
pablodelarco wants to merge 44 commits into
masterfrom
feature/open-ondemand-docs
Open

pablodelarco wants to merge 44 commits into
masterfrom
feature/open-ondemand-docs

Conversation

@pablodelarco

Copy link
Copy Markdown

Documentation of the Open OnDemand Service appliance (marketplace-community #132), as an integration blueprint under Solutions, in the folder of the page created for it.

  • Six pages: overview, quick start, architecture, configuration, operations, monitoring and troubleshooting.
  • The overview opens with the usage model in a multi-tenant cloud and its three actors, Infrastructure Admin, Tenant Operator and OnDemand End User.

Preview: http://51.159.138.137:8082/devel/solutions/integration_blueprints/open_ondemand/

Replaces #942, which targeted one-7.4.

…eenshots

The quick start now explains the two Virtual Networks before the wizard and
shows how to create the compute one in Sunstone. The Service Inputs step is
described tab by tab with the redesigned inputs of appliance 1.0.0-20260916,
where every input is optional and each optional feature sits behind a switch.
Configuration regroups the inputs by tab and adds the advanced attributes.
Screenshots retaken on the rebuilt appliance.
…s, two-column input tables, access from outside the management network
…nt network, uid optional in the initial users
…he service

The portal runs the controller and the accounting, every worker joins as a dynamic node,
the forms ask for cores, memory and GPUs, the pool grows on pending jobs and the oldest
worker drains before OneFlow removes it. The wizard loses the Slurm tab and an external
cluster becomes an advanced attribute.
…apter pattern

Shorter sentences and common words on the six pages, every value, command and
heading kept. Dates removed from the measurements. The chapter index carries one
paragraph in the body, as the Elastic Slurm chapter does, and says that the
service runs its own Slurm cluster.
…U model, faster scale up

The overview explains Slurm, munge and MPI for a reader who does not know
them, the configuration page shows an MPI job on two workers, the scale up
cooldown is 180 seconds and the requirements name the host CPU passthrough.
Two troubleshooting cases added: a service stuck in DEPLOYING_NETS and a
worker started before its portal.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants