diff --git a/source/_static/opmodels/annotation/annotation-1.png b/source/_static/opmodels/annotation/annotation-1.png new file mode 100644 index 0000000..696971d Binary files /dev/null and b/source/_static/opmodels/annotation/annotation-1.png differ diff --git a/source/_static/opmodels/annotation/annotation-2.png b/source/_static/opmodels/annotation/annotation-2.png new file mode 100644 index 0000000..8e08783 Binary files /dev/null and b/source/_static/opmodels/annotation/annotation-2.png differ diff --git a/source/_static/opmodels/annotation/annotation-3.png b/source/_static/opmodels/annotation/annotation-3.png new file mode 100644 index 0000000..e028c0d Binary files /dev/null and b/source/_static/opmodels/annotation/annotation-3.png differ diff --git a/source/_static/opmodels/cluster/cluster-1.png b/source/_static/opmodels/cluster/cluster-1.png new file mode 100644 index 0000000..8aaf939 Binary files /dev/null and b/source/_static/opmodels/cluster/cluster-1.png differ diff --git a/source/_static/opmodels/cluster/cluster-2.png b/source/_static/opmodels/cluster/cluster-2.png new file mode 100644 index 0000000..d83bbfa Binary files /dev/null and b/source/_static/opmodels/cluster/cluster-2.png differ diff --git a/source/_static/opmodels/cluster/cluster-3.png b/source/_static/opmodels/cluster/cluster-3.png new file mode 100644 index 0000000..aa4289e Binary files /dev/null and b/source/_static/opmodels/cluster/cluster-3.png differ diff --git a/source/_static/opmodels/cluster/cluster-4.png b/source/_static/opmodels/cluster/cluster-4.png new file mode 100644 index 0000000..e67c9e0 Binary files /dev/null and b/source/_static/opmodels/cluster/cluster-4.png differ diff --git a/source/_static/opmodels/cluster/cluster-5.png b/source/_static/opmodels/cluster/cluster-5.png new file mode 100644 index 0000000..37b0e59 Binary files /dev/null and b/source/_static/opmodels/cluster/cluster-5.png differ diff --git a/source/_static/opmodels/cluster/cluster-6.gif b/source/_static/opmodels/cluster/cluster-6.gif new file mode 100644 index 0000000..02215b7 Binary files /dev/null and b/source/_static/opmodels/cluster/cluster-6.gif differ diff --git a/source/_static/tools/poet/prediction-1.png b/source/_static/tools/poet/prediction-1.png new file mode 100644 index 0000000..f94432a Binary files /dev/null and b/source/_static/tools/poet/prediction-1.png differ diff --git a/source/_static/tools/poet/prediction-2.png b/source/_static/tools/poet/prediction-2.png new file mode 100644 index 0000000..cd03182 Binary files /dev/null and b/source/_static/tools/poet/prediction-2.png differ diff --git a/source/_static/tools/poet/prediction-3.png b/source/_static/tools/poet/prediction-3.png new file mode 100644 index 0000000..71cb2fc Binary files /dev/null and b/source/_static/tools/poet/prediction-3.png differ diff --git a/source/_static/tools/poet/prediction-4.png b/source/_static/tools/poet/prediction-4.png new file mode 100644 index 0000000..16cd334 Binary files /dev/null and b/source/_static/tools/poet/prediction-4.png differ diff --git a/source/_static/tools/poet/system-prompt-1.png b/source/_static/tools/poet/system-prompt-1.png new file mode 100644 index 0000000..01fe7c8 Binary files /dev/null and b/source/_static/tools/poet/system-prompt-1.png differ diff --git a/source/_static/walkthroughs/antibody-hit-selection-ngs/ngs-advanced-filters.png b/source/_static/walkthroughs/antibody-hit-selection-ngs/ngs-advanced-filters.png new file mode 100644 index 0000000..b5fdc87 Binary files /dev/null and b/source/_static/walkthroughs/antibody-hit-selection-ngs/ngs-advanced-filters.png differ diff --git a/source/_static/walkthroughs/antibody-hit-selection-ngs/ngs-antibody-view.png b/source/_static/walkthroughs/antibody-hit-selection-ngs/ngs-antibody-view.png new file mode 100644 index 0000000..ef7e140 Binary files /dev/null and b/source/_static/walkthroughs/antibody-hit-selection-ngs/ngs-antibody-view.png differ diff --git a/source/_static/walkthroughs/antibody-hit-selection-ngs/ngs-cluster.png b/source/_static/walkthroughs/antibody-hit-selection-ngs/ngs-cluster.png new file mode 100644 index 0000000..a0c2893 Binary files /dev/null and b/source/_static/walkthroughs/antibody-hit-selection-ngs/ngs-cluster.png differ diff --git a/source/_static/walkthroughs/antibody-hit-selection-ngs/ngs-dataset-assay-overview.png b/source/_static/walkthroughs/antibody-hit-selection-ngs/ngs-dataset-assay-overview.png new file mode 100644 index 0000000..d6b4658 Binary files /dev/null and b/source/_static/walkthroughs/antibody-hit-selection-ngs/ngs-dataset-assay-overview.png differ diff --git a/source/_static/walkthroughs/antibody-hit-selection-ngs/ngs-predict.png b/source/_static/walkthroughs/antibody-hit-selection-ngs/ngs-predict.png new file mode 100644 index 0000000..75037c3 Binary files /dev/null and b/source/_static/walkthroughs/antibody-hit-selection-ngs/ngs-predict.png differ diff --git a/source/walkthroughs/antibody-hit-selection-ngs.rst b/source/walkthroughs/antibody-hit-selection-ngs.rst index 52b410f..76aac83 100644 --- a/source/walkthroughs/antibody-hit-selection-ngs.rst +++ b/source/walkthroughs/antibody-hit-selection-ngs.rst @@ -7,10 +7,9 @@ This recommended end-to-end workflow guides you through selecting antibody hits from NGS-derived libraries using the **Dataset Assay Details** page. Each step assumes the previous step's output is in place. -This walkthrough is task-oriented. For a detailed feature reference of the controls used below like Predict, Clustering, Advanced Filters, and the Antibody -settings panel, see comprehensive guide at:doc:`/web-app/opmodels/dataset-assay`. +This walkthrough is task-oriented. For a detailed feature reference of the controls used below, view the following pages: `predict withina a table `, 'Clustering `, and the `Antibody settings panel`. -.. figure:: /_static/walkthroughs/antibody-hit-selection-ngs/dataset-assay-overview.png +.. figure:: /_static/walkthroughs/antibody-hit-selection-ngs/ngs-dataset-assay-overview.png :alt: Dataset Assay Details page overview, showing tabs, header chips, and action bar @@ -44,6 +43,9 @@ On the **Dataset** tab, open the **Antibody** panel, then configure the followin You now have a fully annotated table view of the library. +.. figure:: /_static/walkthroughs/antibody-hit-selection-ngs/ngs-antibody-view.png + :alt: open the antibody panel + Reduce redundancy with Clustering ================================= @@ -61,6 +63,8 @@ downstream steps operate on diverse families. You now have a ``Cluster Number`` column. +.. figure:: /_static/walkthroughs/antibody-hit-selection-ngs/ngs-cluster.png + :alt: view cluster column Pre-filter using NGS / antibody metadata ======================================== @@ -86,6 +90,8 @@ Open **Advanced Filters** from the Dataset tab and apply the following filters i Toggle **Show select column** if you want to see what got rejected instead of hiding it. +.. figure:: /_static/walkthroughs/antibody-hit-selection-ngs/ngs-advanced-filters.png + :alt: view cluster column Score with Predict ====================== @@ -103,6 +109,9 @@ With the candidate set narrowed, run a model to rank within it. **Scale with parallel predictions**: You can run multiple predictions in parallel — for example, one for binding and one for developability. Each gets its own chip and its own column. +.. figure:: /_static/walkthroughs/antibody-hit-selection-ngs/ngs-predict.png + :alt: view cluster column + Combine signals ================ diff --git a/source/web-app/opmodels/antibody-annotation.rst b/source/web-app/opmodels/antibody-annotation.rst new file mode 100644 index 0000000..4dbf0e2 --- /dev/null +++ b/source/web-app/opmodels/antibody-annotation.rst @@ -0,0 +1,116 @@ +Antibody annotation +=================== + +This tutorial shows you how the platform automatically annotates antibody sequences on upload: identifying CDR regions, flagging known liabilities, and calling germline V-genes, alleles, and mutation load, all without a separate annotation step. + +Use this as a starting point for screening a dataset for developability risk or germline diversity before moving on to embedding, clustering, or scoring. + +If you run into any challenges or have questions while getting started, please contact `OpenProtein.AI support `_. + + +What you need before starting +------------------------------ + +You need a sequence-only CSV file of antibody sequences. No header row or extra metadata columns are required, the platform detects VH and VL chains on its own and needs no manual chain labeling or numbering. + + +Upload your dataset +^^^^^^^^^^^^^^^^^^^^^ + +Upload your CSV the same way you would a `dataset`. You do not need to upload a csv with properties. If the sequences are recognized as antibodies, the table automatically gains a set of **CDR1** / **CDR2** / **CDR3** / **Liability** chips above the grid, and an **Antibody** entry appears in the table toolbar alongside **Dataset Info**, **Kabat**, **Consensus**, **Settings**, **Collapse**, **Filters**, and **Export**. + +No separate annotation job is needed. Non-antibody protein datasets will not show the Antibody control, since there's no CDR or germline structure to annotate. + + + +Viewing the antibody settings +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +Click **Antibody** in the toolbar to open the annotation panel. + +- **Highlight CDRs** lets you toggle **Show CDR1** / **CDR2** / **CDR3** independently. Each region is color-coded directly inside the VH and VL sequence text in the table. +- **Sequence view** offers **Aligned** (pad sequences to a common length for side-by-side comparison) and **Trim non-standard positions**. +- The numbering scheme used to define the CDR boundaries (Kabat, by default) is set from the separate **Kabat** dropdown next to Antibody in the toolbar. + +.. image:: /_static/opmodels/annotation/annotation-1.png + :alt: Antibody panel open showing Highlight CDRs, Sequence view, Liabilities, and Show antibody columns controls + + +Review flagged liabilities +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +The **Liabilities** section flags residues or motifs known to affect antibody developability. + +- Choose **Highlight** to mark liabilities directly in the sequence text while still seeing every row. +- Switch to **Filter** to narrow the table down to rows that contain a flagged liability. +- **Show column** adds a dedicated Liability column to the grid, matching the red **Liability** chip shown above the table alongside the CDR chips. + +Use Highlight while you're still exploring the dataset broadly, and switch to Filter once you're ready to narrow in on sequences that need to be deprioritized or redesigned around a specific liability. + + +Choose which antibody columns to show +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +**Show antibody columns** controls which germline and mutation metrics appear in the grid. It's split into two groups: + +- **Gene**: Germline pair, Heavy V-Gene, and Light V-Gene, each with an **Allele** toggle that switches the calls between gene-level (for example ``IGHV1-69``) and allele-level (for example ``IGHV1-69*01``) precision. +- **Metrics**: Germline pair frequency, Total Mutations, CDR3 length, and Germline distance, numeric summaries computed from each sequence's alignment back to its called germline. + + +Read the annotated table +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +Every selected column appears directly in the grid. VH and VL show the full sequence with CDR1, CDR2, and CDR3 shaded in distinct colors inline, followed by the germline and mutation columns you selected: Germline Pair, Germline Pair Frequency, Heavy V-Gene, Light V-Gene, Total Mutations, and more. + +.. image:: /_static/opmodels/annotation/annotation-2.png + :alt: Dataset table with VH and VL columns showing color-coded CDR highlighting, plus Germline Pair, Germline Pair Frequency, Heavy V-Gene, Light V-Gene, and Total Mutations columns + +From here, sort or filter on any of these columns the same way you would elsewhere in the table, and combine them with Embedding, Cluster, or Predictions to bring germline and liability context into hit selection. + + +Column reference +----------------- + +.. list-table:: + :header-rows: 1 + :widths: 20 30 + :align: left + + * - Column + - What it tells you + * - Germline Pair + - The closest matching heavy and light germline gene (or allele, if the Allele toggle is on) called together, for example ``IGHV1-69*01_IGLV1-44*01``. + * - Germline Pair Frequency + - How often this exact germline pairing occurs across the dataset, a quick signal of whether a sequence sits in a common or rare germline background. + * - Heavy V-Gene / Light V-Gene + - The called germline V-gene for each chain individually, with the Allele toggle switching between gene-level and allele-level precision. + * - Total Mutations + - Count of amino acid differences between the sequence and its called germline, a proxy for how far a sequence has diverged through affinity maturation or engineering. + * - CDR3 length + - Length of the CDR3 loop, useful for spotting unusually long or short CDR3s that may affect developability or expression. + * - Germline distance + - Overall sequence distance from the called germline, a broader divergence measure than Total Mutations alone. + * - Liability + - Flags residues or motifs associated with known developability risks (for example deamidation, oxidation, glycosylation sites), shown inline via Highlight or as its own column via Show column. + + +Tips and troubleshooting +-------------------------- + +.. list-table:: + :header-rows: 1 + :widths: 20 20 + :align: left + + * - Question + - Answer + * - Do I need to tell the platform which columns are heavy chain versus light chain? + - No. A sequence-only upload is enough, the platform detects VH and VL chains and numbers them automatically. There's no separate setup step before the Antibody panel becomes available. + * - Why would I turn on Allele instead of leaving germline calls at the gene level? + - Allele-level calls (for example ``IGHV1-69*01``) are more specific and useful when you need to track fine-grained germline differences, such as comparing sequences that share a V-gene but differ by allele. Gene-level calls are easier to scan when you just want a broad family view. + * - My sequences aren't getting annotated as antibodies. + - Confirm the upload is a plain sequence-only CSV or FASTA (no unexpected extra columns before the sequence data) and that the sequences resemble recognizable antibody variable domains. Non-antibody protein datasets won't show the Antibody control since there's no CDR or germline structure to annotate. + * - How does this relate to Embedding, Cluster, and Predictions? + - Antibody annotation is descriptive context computed directly from sequence, it doesn't require an embedding job to run first. You can use it on its own to screen for liabilities and germline diversity, or alongside Embedding, Cluster, and Predictions for a fuller view when selecting hits. + +Please contact `OpenProtein.AI support `_ if the suggested solutions don't resolve the issue. diff --git a/source/web-app/opmodels/cluster-sequences.rst b/source/web-app/opmodels/cluster-sequences.rst new file mode 100644 index 0000000..61402ae --- /dev/null +++ b/source/web-app/opmodels/cluster-sequences.rst @@ -0,0 +1,159 @@ +Cluster sequences +================== + +Overview +-------- + +Clustering groups the sequences in a table by similarity in embedding space, using hierarchical clustering on top of a protein language model embedding (for example, PoET-2). Once a clustering job finishes, every sequence gets a cluster label you can view as a color-coded UMAP, browse as a column in the table, and use to filter or select groups of related sequences. + + +Where this applies +--------------------- + +The Cluster control lives in the same toolbar (alongside Embedding and Predictions) above every sequence table in the product, so the steps below work the same way across **Generate**, **Score**, and **Design**. + + +Before you start +------------------- + +- Have a sequence table open. +- Clustering runs on top of an **embedding**. You will be able to pick an embedding model as part of the setup, so you don't need to precompute one separately. +- How to build the **prompt** for the embedding model, see Embedding Model and Prompt below. + + +1. Open the cluster panel +---------------------------- + +In the toolbar above the table, click the **Cluster** dropdown (it reads ``None`` if nothing is clustered yet). This opens the **Cluster** panel, which lists any existing clustering jobs that had already been run against the table and their settings (embedding, method, linkage, distance metric). + +To start a new one, click **New clustering** in the top-right of the panel. + +.. figure:: /_static/opmodels/cluster/cluster-1.png + :alt: new cluster job + +**Tip:** If a clustering job has already been run on this table, you can just select it from this list instead of creating a new one, jump to Step 4 below. + + +2. Configure the clustering job +----------------------------------- + +**New clustering** opens the **Cluster Sequences** dialog. This clusters every sequence in the table using the embedding model and method you choose here. + +.. figure:: /_static/opmodels/cluster/cluster-2.png + :alt: configuring settings for embedding model + +*The Cluster sequences dialog: pick an embedding model and prompt on top, then a reduction type and hierarchical clustering method below.* + +Embedding model and prompt +~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +- **Embedding model**: choose which protein language model generates the embeddings sequences are clustered on. The current recommended default is **PoET-2**. +- **Prompt**: PoET family models are conditional, so they need a prompt for context. Reuse an existing saved prompt from the list, or click **+ Create new prompt** to build a new one. See `prompt and prompt sampling methods <./prompts.rst>`_ on how to build a prompt. + +.. figure:: /_static/opmodels/cluster/cluster-3.png + :alt: building a prompt for PoET-2 as the selected embedding model + +**Note:** Models without a conditional prompt requirement (e.g. ESM) will skip the prompt step. If you see the error *"Prompt Query: Please enter a sequence or upload a file..."*, either finish building/selecting a prompt or switch to a model that doesn't require one. + +Reduction type and clustering method +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +- **Reduction type**: how per-residue embeddings are collapsed into a single vector per sequence (Recommended: ``Mean``). +- **Linkage method**: the hierarchical clustering linkage criterion (Recommended: ``Ward``). +- **Distance metric**: the distance used between embedding vectors (Recommended: ``Euclidean``). Some linkage methods, like Ward, require Euclidean distance and will lock this field automatically. + +.. figure:: /_static/opmodels/cluster/cluster-4.png + :alt: configuring cluster method and reduction types + +When everything is set, click **Run**. + + +3. Run the job and wait for it to finish +-------------------------------------------- + +Clustering runs asynchronously. After you click Run, a job status bar appears above the table, and the **Jobs** counter in the top-right increments. You can keep working, open the Jobs panel any time to check progress, and the new cluster becomes selectable in the Cluster dropdown once it completes. Refresh your browser to view the completed jobs. + + +4. Select the cluster and tune its resolution +------------------------------------------------- + +Open the Cluster dropdown and click a clustering run to select it. Two additional controls appear for hierarchical clusterings: + +- **Number of clusters**: cuts the hierarchical dendrogram to produce exactly this many clusters. +- **Cluster distance**: alternatively, cut the dendrogram at a given distance threshold, adjusting one updates the other. + +.. figure:: /_static/opmodels/cluster/cluster-5.png + :alt: configuring cluster distance and number of clusters + +Both update instantly against the already-computed job, so you can explore coarser or finer groupings without re-running the clustering. Click **Deselect cluster** to go back to ``None``. + +*With a cluster selected, use Number of clusters or Cluster distance to change resolution on the fly.* + + +5. Use the results +---------------------- + +- **UMAP tab**: in the right-hand Dataset panel, set **Discrete** to **Cluster** to color every point by its cluster assignment, using a distinct-colors legend numbered 1, 2, 3, etc. +- **Dataset / results table**: each row shows its assigned cluster once a clustering is selected, so you can sort or filter the table by cluster. +- **Downstream actions**: select a cluster's points on the UMAP (click, or Shift-drag to multi-select) to view them in the table. + +.. figure:: /_static/opmodels/cluster/cluster-6.gif + :alt: selecting sequences in a cluster to view in the table + +*UMAP colored by cluster (Discrete to Cluster), with each of the 10 clusters shown in a distinct color.* + +**Tip:** Switch **Discrete** back to a continuous property at any time to compare cluster structure against an experimental readout side by side. + + +Settings reference +---------------------- + +.. list-table:: + :header-rows: 1 + :widths: 15 12 30 + :align: left + + * - Setting + - Required? + - What it controls + * - Embedding model + - Required + - Which protein language model produces the per-sequence embeddings clustering runs on (PoET-2, ESM variants, AbLang, etc.). + * - Prompt + - Model-dependent + - Context sequences used by conditional models (PoET family). Reuse a saved prompt or build one via Homology Search, MSA upload, Property Based Sample, or direct Upload. + * - Reduction type + - Required + - How per-residue embeddings are pooled into one vector per sequence (e.g. Mean). + * - Linkage method + - Required + - Hierarchical clustering linkage criterion, e.g. Ward, complete, average. + * - Distance metric + - Required + - Distance function between embeddings, e.g. Euclidean. Some linkage methods force this to Euclidean. + * - Number of clusters + - Post-run + - Cuts the dendrogram to a target cluster count. Adjustable after the job completes, no re-run needed. + * - Cluster distance + - Post-run + - Cuts the dendrogram at a distance threshold instead of a fixed count. Linked to Number of clusters. + + +Tips and troubleshooting +---------------------------- + +.. list-table:: + :header-rows: 1 + :widths: 20 20 + :align: left + + * - Question + - Answer + * - The Run button gives a "Prompt Query" error. + - The selected embedding model needs a prompt but none is attached yet. Select an existing prompt from the list, finish building a new one and submit it, or pick a model that doesn't require a prompt. + * - Can I re-cluster with different settings without losing my current one? + - Yes. Click New clustering again to start another run with different embedding/method settings. Every run is saved and listed in the Cluster dropdown, so you can switch between them freely. + * - Do I need to rerun the job to see more or fewer clusters? + - No. Number of clusters and Cluster distance are applied on top of the already-computed dendrogram, so changing them is instant. + * - Does this work the same in Design and Predict results tables? + - Yes. The Cluster control sits in the same toolbar position in Dataset, Design, and Predict result views, and the setup dialog and UMAP coloring behave identically. diff --git a/source/web-app/opmodels/index.rst b/source/web-app/opmodels/index.rst index 1439409..bf73f8c 100644 --- a/source/web-app/opmodels/index.rst +++ b/source/web-app/opmodels/index.rst @@ -27,6 +27,8 @@ Learn more and get started with our tutorials - `Model training and evaluation <./model-train-evaluate.rst>`_ - `Substitution analysis with OP Models <./sub-analysis.rst>`_ - `Designing sequences <./design.rst>`_ +- `Antibody annotation <./antibody-annotation.rst>`_ +- `Cluster sequences <./cluster-sequences.rst>`_ .. toctree:: :maxdepth: 0 @@ -44,4 +46,5 @@ Learn more and get started with our tutorials Substitution analysis with OP Models Design Aligning sequences - + Automatic Antibody Annotation + Cluster sequences in a table diff --git a/source/web-app/poet/prompts.rst b/source/web-app/poet/prompts.rst index 5647286..24dc3f1 100644 --- a/source/web-app/poet/prompts.rst +++ b/source/web-app/poet/prompts.rst @@ -82,6 +82,8 @@ You can create a prompt context in three ways: If you've previously uploaded prompts, you can reuse them. In the **Choose from project**, select an existing prompt. The sequences from that prompt will automatically load. +This list also includes any **System** prompts available to you, see System Prompts below for the full catalog. + .. image:: /_static/tools/poet/prompt-context-use-existing-1.png :alt: Use existing prompt @@ -275,3 +277,53 @@ The **homology level** field allows you to generate more or less diverse prompt - If you need more focused generation, use a higher homology level and set a minimum similarity threshold to ensure the prompt focuses on the local sequence landscape around your seed. The default **maximum** and **minimum similarity parameters** are set to values which perform well across a wide range of protein families. These can be tuned to adjust the diversity of sequences that will be modeled by PoET. + +Antibody Prompts +----------------- + +In addition to prompts you build yourself, OpenProtein.AI provides **system prompts**: platform-curated prompts available out-of-the-box. + +Antibody prompts appear in the same prompt list used across the PoET tools (Score Sequences, Create Embedding, Cluster, Predictions, and Train a Model), marked with a **System** badge. Prompts recommended for your current model and property selections are additionally marked **Recommended**, and prompts built for a specific chain configuration carry a short tag, for example ``Antibody_vh_vl``. A prompt name ending in **(Virtual)** indicates a pre-computed model memory rather than a sampled context, see below. + +**Note:** Antibody prompts can't be edited or deleted the way a prompt of your own can. If you need a variation, for example a different sample size or clustering threshold, build a new prompt following the same reference database as a starting point, see Creating a Context below. + +.. image:: /_static/tools/poet/system-prompt-1.png + :alt: system level prompts + +Virtual vs. Sampled Prompts +~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~~ + +Most antibody prompts are built the same way a user-defined prompt is: as an ensemble of replicates sampled from a reference database. For the antibody prompts below, each replicate is a random sample of 200 naive sequences from the OAS paired antibody database, clustered at 70% sequence identity, with 10 replicates per ensemble. + +A prompt labeled **(Virtual)** works differently. It's a pre-computed PoET-2 memory, trained ahead of time using cluster representatives from the reference database (via mmseqs2 linclust) with a frozen PoET-2 backbone and a fixed number of virtual sequences, using the ``poet-2-vcontext`` prompt type. Because the context is pre-computed rather than sampled per job, it has a single replicate and only supports PoET-2. + +Available Prompts +~~~~~~~~~~~~~~~~~~~~ + +Antibody Human VH-VL (Virtual) +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +Pre-computed PoET-2 memory trained using mmseqs2 linclust cluster reps from OAS human paired antibody sequences clustered at 70% sequence identity. Trained with a frozen PoET-2 backbone and ``nvs=40`` virtual sequences using ``poet-2-vcontext``. + +Antibody Human VH-VL +^^^^^^^^^^^^^^^^^^^^^ + +An ensemble of 10 replicates, each a random sample of 200 naive paired human antibody sequences from the OAS paired database, clustered at 70% sequence identity. Each complex is the heavy chain (VH) followed by the light chain (VL). + + +Antibody Human VL-VH +^^^^^^^^^^^^^^^^^^^^^ + +An ensemble of 10 replicates, each a random sample of 200 naive paired human antibody sequences from the OAS paired database, clustered at 70% sequence identity. Each complex is the light chain (VL) followed by the heavy chain (VH). + +Antibody Human VH +^^^^^^^^^^^^^^^^^^ + +An ensemble of 10 replicates, each a random sample of 200 naive human antibody heavy chain (VH) sequences from the OAS paired database, clustered at 70% sequence identity. + + +Antibody Human VL +^^^^^^^^^^^^^^^^^^ + +An ensemble of 10 replicates, each a random sample of 200 naive human antibody light chain (VL) sequences from the OAS paired database, clustered at 70% sequence identity. + diff --git a/source/web-app/poet/score-sequences.rst b/source/web-app/poet/score-sequences.rst index 10c76c9..00dec15 100644 --- a/source/web-app/poet/score-sequences.rst +++ b/source/web-app/poet/score-sequences.rst @@ -94,6 +94,74 @@ Improve your results by adding more sequences with your desired properties to yo To improve scores, increase the number of the **ensemble** setting. This will result in higher scoring sequences, but will take longer to complete. +Running predictions within a dataset +------------------------------------ + +If the sequences you want to score already live in a dataset, design results, or predict results table, you can score them in place using the **Predictions** panel instead, without leaving the table. + +Predictions supports two kinds of models: + +- A **user model** you've already trained on your own assay data, so its held-out accuracy is known before you trust its ranking. +- A **foundation model**, such as PoET-2, for zero-shot scoring when you don't yet have labeled data for the property you care about. + +Step 1: Create prediction +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +In the toolbar above the table, click the **Predictions** dropdown (it shows **None** if the table hasn't been scored yet). This lists any prediction jobs already run against the table. Click **New prediction** to open the **Create Prediction** dialog. + +.. image:: /_static/tools/poet/prediction-1.png + :alt: open new prediction window + +Step 2: Choose a model +^^^^^^^^^^^^^^^^^^^^^^^ + +Create prediction offers two tabs: + +- **User models** lists trained models available in your project, along with the property each one predicts, what dataset it was trained on, and its held-out Spearman's rho and Pearson's r against measured values. Select one or more models and click **Run** to score the whole table. + +.. image:: /_static/tools/poet/prediction-2.png + :alt: choosing models + +- **Foundation models** lets you score without a trained model of your own. Models are grouped by family, with **PoET-2** recommended. PoET-family models are conditional and require a **prompt**, the same prompt mechanism used by Score Sequences (see `prompts and prompt sampling methods <./prompts.rst>`_). Reuse a saved prompt or build a new one before running. + +.. image:: /_static/tools/poet/prediction-3.png + :alt: choosing plm + +Step 3: Run the scoring job +^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +Scoring runs as a background job. After clicking Run, a job status bar appears above the table and the **Jobs** counter increments. The Predictions dropdown shows the run as in progress until it finishes, at which point its predicted column becomes available in the table. + +Step 4: Read the predicted column and select hits +^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ + +Once a run finishes, select it from the Predictions dropdown. Each selected run adds its own column to the table, named after the source model, sitting right alongside any measured column it was trained to predict. + +.. image:: /_static/tools/poet/prediction-4.png + :alt: reviewing predict results + + +With a predicted column in the table, hit selection is a matter of working the table: + +.. list-table:: + :header-rows: 1 + :widths: 20 20 + :align: left + + * - Action + - Why it helps + * - Sort by the predicted column + - Brings your top-scoring candidates to the top. + * - Filter above a score threshold + - Cuts the table down to only the rows worth reviewing. + * - Cross-check against cluster assignment, if available + - Keeps your shortlist diverse instead of pulling near-duplicates from one region of sequence space. + * - Select more than one completed run + - Lets you compare predictions across multiple properties at once, for example activity and stability. + +Once you've selected your rows, carry the shortlist forward into Create Design, Substitution Analysis, or Train Model. + + Next Steps ----------