Pathoplexus processes submitted data to validate, harmonize, and standardize it. We ensure users have maximum flexibility in accessing the most useful data for their needs by rejecting only submissions that lack essential metadata values, whose sequences are not identifiable as the specified pathogen, or whose raw reads contain an excessive number of likely-human reads.
We use Nextclade for alignment, mutation calling, quality checks and clade assignment.
The data preprocessing steps encompass:
Submissions that fail to meet these requirements are rejected by the preprocessing pipeline, which provides a detailed error message explaining the reason for rejection.