Skip to content

Latest commit

 

History

48 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

xenium-panel-convert

A command-line utility for converting files to the formats accepted by the 10x Genomics Xenium Panel Designer.

In order to create a custom Xenium Panel, the 10x Genomics Xenium Panel Designer requires a list of targets (typically genes) and one or more reference datasets. This command-line tool takes as input the list of targets as a CSV-file and one or more scanpy-generated H5AD files and converts them to the formats accepted by the Xenium Panel Designer.

Installation

Install the latest version from the releases page.

Usage

xp-convert has 3 subcommands:

xp-convert targets

Inputs

The main input for this command is a CSV-formatted target-list. The target-list must contain the following fields:

Field Description Allowed Values Example Required
ensembl_id The Ensembl ID of the target. If the row has custom == true, this will not be validated against the reference genome. Any Ensembl ID from the allowed list of genes, unless custom == true, in which case any string ENSG00000141510 yes
gene_symbol The gene-symbol of the target. If the row has custom == true, this will not be validated against the reference genome. Any gene-symbol from the allowed list of genes, unless custom == true, in which case any string TP53 yes
group A key to create various gene "groups". If the Xenium Panel Designer cannot include a target, you can replace it with a target with the same group. This string can be whatever you want, though it may be useful to use biologically-meaningful terms. Note that all group-names will be lowercased in the output. Any string senescence yes
priority A priority to assign to the target, used to sort the output and prevent the Xenium Panel Designer from dropping must-have genes from the panel. backup | desired | must_have must_have yes
custom Whether this is a "custom" target, in which case it will not be validated against the reference genome. true | false true no, defaults to false

Any additional fields will be propagated to the output untouched. This is useful for notes or any information you want to attach to targets.

Example target-list
ensembl_id,gene_symbol,group,priority,notes
ENSG00000141510,TP53,tumor,must_have,some interesting fact

If the target-list has different fieldnames and you don't want to edit it, you can provide a set of field-aliases that map the fieldname in the file to one of the canonical fieldnames above. For example, if the field containing Ensembl IDs is called "gene ID", you can do:

xp-convert targets --targets-path targets.csv --field-alias 'gene ID=ensembl_id'

If you have a set of aliases you use frequently, you can store them in a TOML file mapping the alias to the canonical fieldname:

"gene ID" = "ensembl_id"

and then pass that file on the command-line:

xp-convert targets --targets-path targets.csv --field-alias-file aliases.toml

If you use both flags, aliases passed on the command-line take precedence over aliases in the file.

Outputs

Outputs will be written to the directory provided as the --output-dir option, which defaults to xp-convert. If there are no errors, it will look like:

xp-convert
├── validated-targets.csv
└── xenium-panel-designer-targets.csv

validated-targets.csv is a copy of the input file with renamed fields and lowercased group-names, while xenium-panel-designer-targets.csv is a sorted list of targets suitable for copy-pasting into the panel designer. For example, the input file shown above would result in:

Gene,Ensembl ID,Probe sets,Force
TP53,ENSG00000141510,,forced

All encountered errors are collected and written to <OUTPUT_DIR>/target-list-errors.json. For example, a row without the field priority and with an incorrect gene-symbol would result in:

[
  {
    "line_number": 2,
    "submitted_target": {
      "ensembl_id": "ENSG00000141510",
      "gene_symbol": "SOME GENE SYMBOL",
      "group": "group0",
      "priority": null,
      "custom": null
    },
    "errors": [
      {
        "type": "missing_field",
        "fieldname": "priority",
        "hint": "the field 'priority' is missing - add it to the CSV"
      },
      {
        "type": "ensembl_id_gene_symbol_mismatch",
        "ensembl_id": "ENSG00000141510",
        "correct_gene_symbol": "TP53",
        "hint": "the gene symbol corresponding to the Ensembl ID ENSG00000141510 is TP53 - change either the Ensembl ID or the gene symbol so they match"
      }
    ]
  }
]

xp-convert references

Inputs

The main input for this command is one or more H5AD files generated using scanpy. The AnnData object must contain cell-barcodes, cell-annotations, Ensembl IDs, and gene-symbols. All of these are automatically populated for you by scanpy besides cell-annotations. You can generate such a file from the outputs of cellranger like so:

import scanpy as sc

adata = sc.read_10x_h5("/path/to/data.h5")

# Whatever logic you want to annotate each cell
adata.obs["annotation"] = ...

sc.write("/path/to/data.h5ad", adata)

Outputs

If no errors are found, the following invocation:

xp-convert references 1k_mouse_kidney_CNIK_3pv3_filtered_feature_bc_matrix.h5ad,annotation-col=annotation,transcriptome=mm10-2020-A

would produce the following output:

xp-convert
└── 1k_mouse_kidney_CNIK_3pv3_filtered_feature_bc_matrix
    ├── annotations.csv
    └── matrix.h5

This directory is ready for upload to the Xenium Panel Designer in one of the accepted formats. See the full list of options by running xp-convert references --help.

All encountered errors are collected and written to <OUTPUT_DIR>/<DATASET_NAME>-errors.json. For example, supplying the wrong column for cell-annotations would result in:

{
  "path": "crates/core/test-data/1k_mouse_kidney_CNIK_3pv3_filtered_feature_bc_matrix.h5ad",
  "errors": [
    {
      "component": "cell_annotations",
      "type": "invalid_h5_object_path",
      "hdf5_error": "H5Dopen2(): unable to synchronously open dataset: object 'foo' doesn't exist",
      "object_path": "obs/foo",
      "field_type": "container",
      "available_objects": [
        "/X/data",
        "/X/indices",
        "/X/indptr",
        "/obs/_index/mask",
        "/obs/_index/values",
        "/obs/annotation/categories",
        "/obs/annotation/codes",
        "/obs/barcode/mask",
        "/obs/barcode/values",
        "/var/_index/mask",
        "/var/_index/values",
        "/var/feature_types/categories",
        "/var/feature_types/codes",
        "/var/gene_ids/mask",
        "/var/gene_ids/values",
        "/var/gene_symbol/categories",
        "/var/gene_symbol/codes",
        "/var/genome/categories",
        "/var/genome/codes"
      ],
      "hint": "obs/foo could not be read as a container (H5Dopen2(): unable to synchronously open dataset: object 'foo' doesn't exist) - ensure the correct column name was provided"
    }
  ]
}

xp-convert all

This command is the same as combining the above commands with the added checks that:

  • the species of the target-list and reference datasets match
  • the gene-symbols in the target-list and reference datasets match
  • all genes in the target-list are in the reference datasets

If any of these conditions are not met, warnings are written to <OUTPUT_DIR>/<DATASET_NAME>-warnings.json.

About

A command-line utility for converting files to the formats accepted by the 10x Genomics Xenium Panel Designer.

Resources

Stars

2 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages