This is a demo instance. You can upload any data you like, including test data to try things out. You can find test data to upload in the example data repository. This instance should not be used for analysis.

Preprocessing

Pathoplexus processes submitted data to validate, harmonize, and standardize it. We ensure users have maximum flexibility in accessing the most useful data for their needs by rejecting only submissions that lack essential metadata values, whose sequences are not identifiable as the specified pathogen, or whose raw reads contain an excessive number of likely-human reads.

We use Nextclade for alignment, mutation calling, quality checks and clade assignment.

The data preprocessing steps encompass:

  • Sequences:
    • Verification that the sequence corresponds to the virus specified by the user.
    • Alignment with the reference genome.
    • Removal of terminal Ns.
    • Translation of genes/coding regions.
    • Quantification of mutations by number and type (including nucleotide and amino acid variations).
    • Identification and labeling of deletions and insertions.
    • Assignment of specific clades/lineages.
  • Metadata:
    • Standardization of collection date formats.
    • Standardization of location information using INSDC-standards.
    • Ensure required values are set
    • Ensure metadata fields are of correct type (e.g. string, int, date)
  • Raw reads:
    • File format and content validation

Submissions that fail to meet these requirements are rejected by the preprocessing pipeline, which provides a detailed error message explaining the reason for rejection.