AREE is an open, reproducible research resource that converts fragmented public omics datasets into harmonized, comparable resilience-biomarker evidence for Pacific oyster (Crassostrea gigas), and is deliberately extensible to other shellfish and aquaculture organisms.
It operationalizes Objective 1 of the AREE proposal:
Develop standardized open-access, user-friendly, reproducible bioinformatics pipelines for resilience biomarker discovery through systematic reanalysis, data integration, and meta-analysis.
AREE is an evidence engine, not a claims engine. It identifies associations and evidence convergence across studies and molecular layers — never confirmed mechanistic causation — and it never presents a statistically significant hit in a single study as a validated biomarker.
Every demo dataset is clearly labeled SIMULATED and is kept separate from real evidence. The repository also contains registered real public studies; only harmonized real-study outputs may be interpreted as real evidence (see docs/adding_a_study.md).
Progress and findings dashboard: https://sr320.github.io/AREE/
— what has been registered, harmonized, pooled, and ranked so far, regenerated
from the pipeline's own outputs on every push to main
(how it is built).
- Study registry & intake (
registry/,src/intake/) — machine-readable dataset registration with controlled vocabularies for phenotypes and stressors, plus schema validation. - Standardized reanalysis workflows (
workflows/,modules/) — Nextflow DSL2 scaffolds for RNA-seq, methylation, proteomics, and metabolomics, in raw-reanalysis or processed-results modes. - Cross-study harmonization (
src/harmonize/) — a shared, assay-agnostic evidence schema plus identifier mapping with explicit confidence levels. - Meta-analysis & candidate prioritization (
src/meta_analysis/,src/prioritize/) — random-effects pooling, heterogeneity statistics, and a transparent (non-black-box) candidate score with hard tier gates. - User-facing outputs (
app/,docs/,src/reporting/) — a Streamlit interface, a Quarto documentation site, and per-candidate evidence cards.
Requires Python ≥ 3.9.
python -m venv .venv && source .venv/bin/activate
pip install -e ".[dev,app]"
# Add `intake` if you will convert published .xls/.xlsx supplementary tables:
# pip install -e ".[dev,app,intake]"See docs/installation.md for details.
# 1. Validate a study registration file
aree validate-study registry/studies/GIGAS_HEAT01.yaml
# 2. Register every demo study
for f in registry/studies/GIGAS_*.yaml; do aree register-study "$f"; done
aree list-studies
# 3. Harmonize each study into the shared evidence table
for sid in GIGAS_HEAT01 GIGAS_OA02 GIGAS_PATH03 GIGAS_SAL04 GIGAS_LARV05 GIGAS_GROW06; do
aree harmonize --study "$sid"
done
# (or harmonize a single processed results table)
aree harmonize --study GIGAS_SAL04 \
--input data/demo/proteomics/GIGAS_SAL04_low_salinity_vs_control_protein_abundance_demo.tsv
# 4. Run a meta-analysis
aree meta-analyze --phenotype thermal_tolerance --feature-type gene
aree meta-analyze --feature-type gene
# 5. Generate biomarker evidence cards, then summarize the top candidates per phenotype
aree build-evidence-cards --phenotype thermal_tolerance
aree top-candidates --n 10
# 6. Build the docs site (with the progress dashboard) / launch the interface
aree build-dashboard
quarto render docs
streamlit run app/main.pyOr run the whole demo in one step:
make demoAREE/
├── README.md, LICENSE, CITATION.cff, CONTRIBUTING.md, CODE_OF_CONDUCT.md
├── pyproject.toml, Makefile
├── docs/ # Quarto site + narrative + how-to documentation
├── schemas/ # JSON Schema: study.schema.json, evidence.schema.json
├── registry/
│ ├── studies/ # per-study YAML registrations (+ templates)
│ ├── controlled_vocabularies/ # phenotype, stressor, tissue, life-stage, etc.
│ └── study_registry.csv # flat index (generated by `aree register-study`)
├── workflows/ # Nextflow DSL2 scaffolds (rnaseq/methylation/proteomics/metabolomics)
├── modules/ # reusable Nextflow process modules per assay
├── containers/ # container image strategy (documentation)
├── config/ # shared Nextflow config (base + demo)
├── data/
│ ├── demo/ # SIMULATED demo result tables per assay
│ ├── reference/ # genome/annotation metadata
│ │ └── crosswalk/ # REAL identifier crosswalk + provenance sidecar
│ └── mappings/ # SYNTHETIC demo crosswalk + ambiguous-symbol map
├── src/
│ ├── common/ # shared paths, IO, vocabulary loaders
│ ├── intake/ # schema validation + registry ingestion
│ ├── harmonize/ # per-assay harmonizers -> shared evidence schema
│ ├── meta_analysis/ # random-effects pooling + heterogeneity
│ ├── prioritize/ # transparent scoring + tier gating
│ ├── reporting/ # evidence cards + provenance manifests
│ ├── mappings/ # builds real crosswalks from NCBI Gene + UniProtKB
│ ├── validation/ # reusable validation checks
│ └── aree/ # the `aree` CLI
├── app/ # Streamlit interface
├── reports/ # generated outputs (gitignored)
├── tests/ # pytest suite
└── .github/workflows/ # CI
| Command | Purpose |
|---|---|
aree validate-study <file> |
Validate a study YAML against schema + controlled vocabularies |
aree register-study <file> [--update] |
Add (or update) a study in the registry |
aree fetch-samplesheet --bioproject <acc> --study <id> |
Build a sample sheet + checksummed FASTQ manifest from a BioProject's deposited metadata |
aree intake-supplementary <config> [--check] |
Convert a published supplementary table into AREE result files; --check verifies committed files still reproduce |
aree list-studies |
List registered studies and their pipeline status |
aree harmonize --study <id> [--input <file>] |
Harmonize a study (or one processed table) into the evidence table |
aree meta-analyze [--phenotype <p>] [--feature-type <t>] |
Random-effects meta-analysis over the evidence table, with BH-adjusted p-values per phenotype / feature-type family |
aree build-evidence-cards [--phenotype <p>] [--feature-type <t>] [--max-adjusted-p <q>] [--all-cards] |
Rank every candidate into reports/evidence_cards/candidates.tsv and write an evidence card for each one with a BH-adjusted p ≤ q (default 0.05, pooled or in any contributing study) |
aree top-candidates [--n <N>] [--candidates <path>] [--out <path>] |
Regroup an existing candidates.tsv by phenotype and write the top N (default 10) per phenotype, ranked by tier then score, to reports/top_candidates_summary.md |
aree build-crosswalk [--taxid <n>] |
Build a real identifier crosswalk from NCBI Gene + UniProtKB |
aree build-dashboard [--out <path>] |
Summarize registry, pipeline, and findings state into docs/dashboard/data.json (plus the top candidates' evidence cards) for the Quarto dashboard page |
The quick start above runs on simulated studies and a synthetic crosswalk.
AREE also ships one real registered study — HESSER2024_VCOR, curated from the
open-access supplementary tables of
Hesser et al. 2024 (Vibrio
coralliilyticus challenge of C. gigas larvae). Harmonizing it requires the real
crosswalk:
export AREE_CROSSWALK=data/reference/crosswalk/mgigas_gene_id_crosswalk.tsv
aree harmonize --study HESSER2024_VCORThe result tables it harmonizes are themselves derived from the published supplementary spreadsheet by a committed, re-runnable intake config — no manual copy-paste step sits between the publication and the evidence:
aree intake-supplementary data/studies/HESSER2024_VCOR/intake.yaml --check--check regenerates the tables into a temporary directory and compares
checksums against both the committed files and the committed provenance, so a
hand-edited result table or a swapped source artifact fails loudly. CI runs this
on every push. Drop --check to actually rewrite the files.
87.2% of its published identifiers resolve (274 exact, 32 inferred via NCBI's
retired-GeneID remapping); every unresolved identifier is a gene NCBI discontinued
without a replacement. Read
docs/first_real_study.md before curating your own — it
records what broke the first time real data hit the pipeline.
Real and simulated evidence are kept strictly apart: simulated is a column in the
evidence schema and part of the meta-analysis grouping key, and each study selects
its crosswalk from its own simulated flag, so a real study will refuse to run
against the demo crosswalk.
Rebuild it only when the NCBI annotation changes (streams ~230 MB, 15-20 min):
aree build-crosswalkThe demo and real crosswalks are deliberately never merged — the demo's LOC
numbers collide with real NCBI GeneIDs that denote different genes. See
docs/identifier_mapping.md for the coverage
figures and their consequences, in particular that UniProt links only 8.4% of
M. gigas genes, so proteomics evidence carries a much higher unresolved
rate than transcriptomics evidence. Identifiers retired by NCBI re-annotation
(9,057 of them) resolve to their current replacement as inferred; the 13,601
discontinued with no replacement stay unresolved by design.
Start with docs/about.qmd or render the site with
quarto render docs. Key pages:
- Why this resource matters
- Design document · Technical architecture
- Adding a study · Identifier mapping
- Interpreting meta-analysis · Interpreting candidate scores
- Governance & provenance · Roadmap
The site's home page, https://sr320.github.io/AREE/, is a one-screen dashboard of where the effort stands: how many real studies, comparisons, evidence records, and high-priority candidates exist, what the pooled evidence currently says and what it does not, the top-ranked candidates with links to their evidence cards, and the status of each real study. Simulated demo evidence is excluded from its numbers.
Nothing on the page is typed in. .github/workflows/pages.yml registers every
study, harmonizes the demo studies against the demo crosswalk and the real
studies against the real crosswalk, pools and ranks, then runs
aree build-dashboard, which writes docs/dashboard/data.json from those
outputs, and quarto render docs renders docs/index.qmd from it. To build
the same page locally:
make demo # simulated studies
make real-pool # the three real studies, against the real crosswalk
make dashboard # aree build-dashboard + quarto render docsThe site is served from docs/_site/. For the first deployment, set the
repository's Pages source to GitHub Actions (Settings → Pages), or run:
gh api -X POST repos/sr320/AREE/pages -f build_type=workflowThis is a functioning MVP. See docs/roadmap.md for a candid breakdown of what is complete and runnable (schemas, registry, harmonization for all four assay types, meta-analysis, prioritization, evidence cards, CLI, Streamlit app, Quarto docs, tests, CI) versus scaffolded but not yet production-ready (the Nextflow raw-data workflows, which are structurally complete but have not been executed against real sequencing data in this build) versus planned (real ortholog mapping, additional species, hosted deployment).
MIT — see LICENSE.
See CITATION.cff.