End-to-end spatial-genomics workflow for 10x Xenium slides: proseg cell-segmentation preprocessing → celltype-marker reference building → RCTD deconvolution + SPLIT processing.
This pipeline standardises the workflow that start from xenium cells to cleaned-up proseg segmented cells via a scRNA reference-based purification method called SPLIT. xenium-preprocess-pipeline bundles the chain — proseg → xenium-ranger → RCTD
reference → RCTD + SPLIT processing — into three publishable packages plus one
shell driver, so a new sample runs end-to-end from a single command, with
every intermediate product traceable to the exact config that built it.
git clone https://github.com/settylab/xenium-preprocess-pipeline
cd xenium-preprocess-pipeline
scripts/create-env.sh -n xenium -f environments/xenium.yml
micromamba activate xenium # needs `micromamba shell hook` sourced first — see docs/installation.md § Prerequisites
scripts/uv-pip-install.sh -e packages/xenium-preprocess
scripts/uv-pip-install.sh -e packages/ref-build
scripts/uv-pip-install.sh -e packages/rctd-split
./scripts/submit_workflow.sh \
--sample-id SAMPLE1 \
--output-root /data/workflow_runs \
--run-id demo_v1 \
--flex-h5ad /data/SAMPLE1/scRNA/SAMPLE1_flex.h5ad \
--celltype-marker-json /data/markers/markers.json \
--proseg-dir /data/SAMPLE1/proseg \
--xenium-ranger-dir /data/SAMPLE1/xenium_rangerOutputs land under /data/workflow_runs/SAMPLE1/SAMPLE1_demo_v1/.
See Outputs for the tree.
Three publishable packages + one shell driver, chained by Slurm afterok:
dependencies:
Xenium raw xenium-preprocess spatial_adata/
┌──────────┐ ┌──────────────────┐ ┌──────────────────┐
│ proseg │───────▶│ xenium-preprocess│─────────▶│ *_proseg_raw.h5ad│
│ xranger │ │ (5 stages) │ │ *_xenium_ranger │
└──────────┘ └──────────────────┘ │ .h5ad │
│ └──────────────────┘
│ │
▼ │
rctd/*_test_object.rds │
│
Flex scRNA ref-build │
┌──────────┐ ┌──────────────────┐ │
│ flex.h5ad│───────▶│ ref-build │ │
│ donors │ │ (celltype-marker │ │
│ markers │ │ driven ref) │ │
└──────────┘ └──────────────────┘ │
│ │
▼ │
rctd/*_reference.rds │
│
rctd-split │
┌──────────────────┐ │
│ rctd-split │◀──────────────────┘
│ (RCTD + SPLIT │
│ typing) │
└──────────────────┘
│
▼
rctd/*_rctd_results.rds
spatial_adata/ (typed)
Each package is independently installable, testable, and reusable outside
this pipeline. The shell driver (scripts/submit_workflow.sh) is a thin
Slurm layer on top.
- OS: Linux (developed and tested on a Slurm cluster).
- Scheduler: Slurm (
sbatch,--dependency=afterok:support). - RAM: ~64 GB per step for typical samples; rctd-split benefits from 16 CPUs.
- Disk: ~50 GB per sample per run for intermediate + final outputs.
- Micromamba (or
conda),uv, and R 4.4+ withspacexr,SPLIT, andSeurat. - Outbound network access to
github.com/api.github.com—spacexrandSPLITinstall from GitHub, not CRAN. No GitHub credentials required (both repos are public), but seedocs/installation.md§ GitHub access — GitHub's unauthenticated API rate limit is easy to hit on shared infrastructure.
# Once per user / per machine
# (scripts/create-env.sh wraps `micromamba create`; if MAMBA_ROOT_PREFIX is
# set for an isolated install, it also keeps the package cache isolated —
# see docs/installation.md § 2)
scripts/create-env.sh -n xenium -f environments/xenium.yml
micromamba activate xenium # needs `micromamba shell hook` sourced first — see docs/installation.md § Prerequisites
# Editable installs of the three packages
# (scripts/uv-pip-install.sh wraps `uv pip install`; if MAMBA_ROOT_PREFIX
# is set, it also keeps uv's cache isolated — see docs/installation.md § 3)
scripts/uv-pip-install.sh -e packages/xenium-preprocess
scripts/uv-pip-install.sh -e packages/ref-build
scripts/uv-pip-install.sh -e packages/rctd-split
# Verify
xenium-preprocess --help
ref-build --help
rctd-split --helpref-build and rctd-split shell out to Rscript. Both stages need
Seurat, Matrix, spacexr (RCTD), and SPLIT.
On the cluster's fhR/4.4.1-foss-2023b module provides
Seurat + Matrix + SpatialExperiment; spacexr and SPLIT install into
a user library — see docs/installation.md for
the recipe. That install goes over GitHub's API, not CRAN — no
credentials required (both repos are public), but see § GitHub access
there for the anonymous rate-limit caveat.
All invocations below assume you are at the repo root
(the folder that contains scripts/, packages/, and
environments/). If you haven't already, cd xenium-preprocess-pipeline first — the ./scripts/…
paths in the examples are relative.
scripts/submit_workflow.sh is the primary entry point — it chains
xenium-preprocess → ref-build → rctd-split via Slurm
--dependency=afterok: and exposes step + sub-stage selection flags.
For running a single package's CLI in isolation (no Slurm; quick
sanity checks on a laptop; debugging one step in isolation), see
Advanced: single-package usage without the driver
below.
# Replace the YOUR_PARTITION with your own cluster's partition. At here, campus-new is the example.
grep -rn YOUR_PARTITION scripts/ # inventory hits
sed -i 's/YOUR_PARTITION/campus-new/g' scripts/*.sbatch scripts/*.sh
grep -rn YOUR_PARTITION scripts/ # should print nothing now
./scripts/submit_workflow.sh \
--sample-id SAMPLE1 \
--output-root /data/workflow_runs \
--run-id demo_v1 \
--flex-h5ad /data/SAMPLE1/scRNA/SAMPLE1_flex.h5ad \
--celltype-marker-json /data/markers/markers.json \
--proseg-dir /data/SAMPLE1/proseg \
--xenium-ranger-dir /data/SAMPLE1/xenium_ranger \
--donor-h5ad /data/SAMPLE2/scRNA/SAMPLE2_flex.h5ad \
--donor-h5ad /data/SAMPLE3/scRNA/SAMPLE3_flex.h5adSubmits three chained Slurm jobs via --dependency=afterok: and prints
their job ids + dependency chain. All jobs share the --run-id so
outputs colocate under one folder.
# xenium-preprocess already ran; resume the chain at ref-build
./scripts/submit_workflow.sh \
--sample-id SAMPLE1 \
--output-root /data/workflow_runs \
--run-id demo_v1 \
--flex-h5ad /data/SAMPLE1/scRNA/SAMPLE1_flex.h5ad \
--celltype-marker-json /data/markers/markers.json \
--start-step ref-build
# --start-step past xenium-preprocess implies --reuse-run-dirResume behaviour depends on which flag you pass. Three levels of "force", coarsest first:
-
--force— DESTRUCTIVE:rm -rfthe entire<output-root>/<sample>/<sample>_<run-id>/folder, then re-run the whole pipeline fromxenium-preprocess. Use ONLY for a clean-slate re-run under an already-used--run-id. Rejected together with--start-step ref-build/--start-step rctd-split(which would wipe the very outputs those flags need to resume from), and with--reuse-run-dir. -
--reuse-run-dir(implied by--start-step ref-buildor--start-step rctd-split) — proceed against the existing run folder; each step overwrites only the files it writes, everything else preserved. This is the everyday "resume the chain" flag. -
Per-step force-rerun — narrower than
--force; wipes only ONE step's sentinels + outputs and re-runs that step, leaving other steps' outputs intact:--xenium-preprocess-force-rerun— re-run xenium-preprocess only.--ref-build-force-rerun— re-run ref-build only.--rctd-split-force-rerun— re-run rctd-split only.
Combines with
--start-stepand--<pkg>-stages(see next section) for even narrower re-runs — e.g. re-render only rctd-split'sqc_reportsub-stage against an existing run:./scripts/submit_workflow.sh \ --sample-id SAMPLE1 \ --output-root /data/workflow_runs \ --run-id demo_v1 \ --flex-h5ad /data/SAMPLE1/scRNA/SAMPLE1_flex.h5ad \ --celltype-marker-json /data/markers/markers.json \ --start-step rctd-split \ --rctd-split-stages qc_report \ --rctd-split-force-rerun
The driver forwards a per-step sub-stage subset via three semantic
flags — --xenium-preprocess-stages, --ref-build-stages,
--rctd-split-stages — each taking a comma-separated stage list.
Composes with --start-step and the per-step force-rerun flags:
# ref-build was interrupted mid-pipeline; resume from `census`
# (assumes load_primary_and_donors already wrote intermediate/loaded/concat.h5ad
# on the previous run under the same --run-id).
./scripts/submit_workflow.sh \
--sample-id SAMPLE1 \
--output-root /data/workflow_runs \
--run-id demo_v1 \
--flex-h5ad /data/SAMPLE1/scRNA/SAMPLE1_flex.h5ad \
--celltype-marker-json /data/markers/markers.json \
--start-step ref-build \
--ref-build-stages census,assemble,export_mtx,rctd_reference_buildOmit --<pkg>-stages to fall back to that package's
DEFAULT_STAGES (the full list). Sub-stages are sentinel-gated —
re-runs skip work whose sentinel already exists — so if you want
to force a specific sub-stage to redo, combine --<pkg>-stages <name> with the matching --<pkg>-force-rerun (see above).
Use these paths when you don't have Slurm, or when you want to run ONE
package in isolation (debugging one step, quick sanity check on a
laptop). Each package installs its own CLI via
pip install -e packages/<name> and accepts the same run scope
(--sample-id, --run-id, --output-root) as the driver.
xenium-preprocess run \
--sample-id SAMPLE1 --run-id demo_v1 \
--output-root /data/workflow_runs \
--proseg-dir /data/SAMPLE1/proseg --xenium-ranger-dir /data/SAMPLE1/xenium_ranger
ref-build run \
--sample-id SAMPLE1 --run-id demo_v1 \
--output-root /data/workflow_runs \
--flex-h5ad /data/SAMPLE1/scRNA/SAMPLE1_flex.h5ad \
--celltype-marker-json /data/markers/markers.json \
--donor-h5ad /data/SAMPLE2/scRNA/SAMPLE2_flex.h5ad
rctd-split run \
--sample-id SAMPLE1 --run-id demo_v1 \
--output-root /data/workflow_runs \
--max-cores 12Each run subcommand takes a --stages flag; pass a subset to run
only part of a step:
xenium-preprocess run \
--sample-id SAMPLE1 --run-id demo_v1 \
--output-root /data/workflow_runs \
--stages proseg_to_anndata qc_filter
ref-build run \
--sample-id SAMPLE1 --run-id demo_v1 \
--output-root /data/workflow_runs \
--flex-h5ad /data/SAMPLE1/scRNA/SAMPLE1_flex.h5ad --celltype-marker-json /data/markers/markers.json \
--stages census assembleEach package ships a config/default.yaml with sensible defaults.
Overrides in precedence order (highest first):
- CLI flag (e.g.
--x-source expected_counts). --config <user.yaml>— user YAML overriding default keys.config/default.yaml— package-shipped defaults.
The full effective config for a run is recorded at
<run-dir>/config.yaml.
Per-step docs live under each package's docs/ directory.
scripts/submit_workflow.sh accepts --config <path>. The file gathers
the values you'd otherwise pass as CLI flags to the driver — no shared
folder needed; inputs point at any absolute path. CLI flags on the
same invocation win over the YAML, so the file is a defaults source you
can layer overrides on top of.
Recommended location: colocated with the run's outputs at
<output_root>/<sample>/<sample>_<run_id>/config.yaml.
Per-step keys use the semantic step names (xenium_preprocess:,
ref_build:, rctd_split:).
A ready-to-copy template listing every supported key with comments
lives at configs/example.yaml. Minimal shape:
# /data/workflow_runs/SAMPLE1/SAMPLE1_demo_v1/config.yaml
sample_id: SAMPLE1
run_id: demo_v1
output_root: /data/workflow_runs
flex_h5ad: /data/SAMPLE1/scRNA/SAMPLE1_flex.h5ad
celltype_marker_json: /data/markers/markers.json
# Optional rctd-split explicit inputs — bypasses the run-folder layout
# auto-discovery. Use to mix a test_object from one sample with a
# reference from another.
test_object: /data/SAMPLE_A/rctd/SAMPLE_A_test_object.rds
reference_rds: /data/SAMPLE_B/rctd/SAMPLE_B_reference.rds
xenium_preprocess:
x_source: maxpost_counts
qc_min_counts_cell: 10
ref_build:
donor_borrow_cap: 100
rctd_split:
umi_min: 10
doublet_mode: doublet./scripts/submit_workflow.sh --config /data/workflow_runs/SAMPLE1/SAMPLE1_demo_v1/config.yaml
# CLI still wins — same YAML with a knob overridden:
./scripts/submit_workflow.sh --config /data/workflow_runs/SAMPLE1/SAMPLE1_demo_v1/config.yaml \
--ref-build-donor-borrow-cap 200ref-build needs a JSON file declaring the celltypes it will build a
reference for, with one list of marker genes per celltype. The
pipeline doesn't ship one — it's data-dependent (which celltypes are
in your Flex scRNA and which genes read as markers in your panel).
Pass the path via --celltype-marker-json (driver flag), or under
celltype_marker_json: in a YAML config.
Schema. One top-level JSON object. Keys are <CellType>_marker or
<CellType>_markers (both accepted; trailing suffix stripped when
the census matches celltype names in .obs). Values are lists of
HGNC gene symbols. Example:
{
"B_marker": ["MS4A1", "CD79A", "CD79B", "CD19", "BANK1"],
"Plasma_marker": ["MZB1", "XBP1", "DERL3", "FKBP11", "TENT5C"],
"Myeloid_marker": ["CD14", "CD68", "MRC1", "CSF1R", "CD163", "SPI1"],
"Fibroblast_marker": ["POSTN", "THY1", "PDGFRA", "CXCL12", "FAP", "COL5A1"],
"T/NK_marker": ["CD8A", "CD3D", "CD3E", "NKG7", "GZMB", "PRF1"],
"endothelial_marker":["PECAM1", "VWF", "CDH5", "KDR", "ENG"],
"RBC_markers": ["HBB", "HBA1", "HBA2", "ALAS2"],
"epithelial_markers":["EPCAM", "KRT8", "KRT18", "KRT19", "CDH1"]
}The keys must cover every celltype present in your Flex data's
.obs[celltype_col] (default column name celltypes; override
with --celltype-col-for-ref-build). Missing celltypes get a
WARN in
census.csv and are dropped from the reference. Extra celltypes
(present in the JSON but with zero cells across primary + donors +
fallback) get decision=missing_no_donor in the census with a
warning.
A ready-to-copy template lives at
configs/example_markers.json
(5 generic celltypes with 2-4 markers each). Copy + edit for
your own celltype set.
All three steps write into a single run folder:
<output-root>/<sample>/<sample>_<run-id>/
├── spatial_adata/ # xenium-preprocess (rctd-split augments in place)
│ ├── <sample>_proseg_raw.h5ad # proseg cell×gene
│ ├── <sample>_xenium_ranger.h5ad # xenium-ranger cell×gene
│ └── <sample>_proseg_purified.h5ad # rctd-split — SPLIT-purified cell×gene
├── rctd/
│ ├── <sample>_test_object.rds # xenium-preprocess — RCTD test object (spatial query)
│ ├── <sample>_reference.rds # ref-build — RCTD reference (celltype pool)
│ └── <sample>_rctd_results.rds # rctd-split — RCTD + SPLIT typing result
├── config.yaml # merged effective config across the three steps
└── logs/
├── slurm-<jobid>-xenium-preprocess.log # slurm-captured stdout+stderr of the xenium-preprocess sbatch job
├── slurm-<jobid>-ref-build.log # same, for ref-build
├── slurm-<jobid>-rctd-split.log # same, for rctd-split
└── workflow-submit.log # authoritative record of jobids + dep chain
Each slurm-<jobid>-<stage>.log is the FULL per-stage log — the stage's
Python + R shell-outs print progress lines to stdout, slurm captures both
stdout and stderr into that single file. There is no separate
<stage>.log sidecar; the sbatch stubs don't write one.
pytest itself is gated behind the [test] extra in each package's
pyproject.toml — not part of environments/xenium.yml — so install
it before running the suites (see docs/installation.md § 5):
scripts/uv-pip-install.sh -e "packages/xenium-preprocess[test]"
scripts/uv-pip-install.sh -e "packages/ref-build[test]"
scripts/uv-pip-install.sh -e "packages/rctd-split[test]"
# Sanity check before running pytest — catches ml/PYTHONPATH env
# shadowing loudly instead of a misleading ModuleNotFoundError (see
# docs/installation.md § 4-5).
scripts/check-python-env.sh --env-name xenium
# Per-package pytest suites
pytest packages/xenium-preprocess/tests
pytest packages/ref-build/tests
pytest packages/rctd-split/tests
# Shell driver smoke test
pytest scripts/testsContributions welcome:
- Open an issue on this repo before starting non-trivial work.
- Branch from
main; open a PR intomain. - Keep per-package changes in per-package PRs when possible.
pytestmust pass locally before pushing.
If you use xenium-preprocess-pipeline in a publication, please cite:
TBD — placeholder pending the associated manuscript.
Per-package citations live in each packages/*/CITATION.cff.
MIT — see LICENSE.