This is a preview deployment banner

Uploading raw reads

Contents:

Pathoplexus allows users to include raw reads when submitting sequences (see here for our general submission documentation). To include your raw reads files, you will need to upload them and link them to their respective sequences via the files.rawReads field in your metadata file. The files.rawReads field should be a space-separated list containing the names of all files associated with the sequence.

When on the submission page, you can either click on ‘Upload folder’ to upload an entire folder containing raw reads files, or ‘Upload files’ to select individual files to upload. Once you’ve selected files to upload, the webpage will let you know if there are files listed in the files.rawReads metadata column that are missing or if you’ve uploaded files that are not listed in files.rawReads. Once all the files included in files.rawReads have been uploaded and linked and no unexpected files are present, you can proceed with submission.

If the raw reads for a sequence are already available in the INSDC, do not upload them again. Instead, supply their run, BioSample, and BioProject accessions in the INSDC-linking metadata fields insdcRawReadsAccession, biosampleAccession, and bioprojectAccession, and we will link to them rather than submitting a second copy.

Raw reads specific metadata

When including raw reads files in your submission you will also need to provide additional metadata fields in your metadata file. Please refer to the table below to see which fields are required and desired when submitting raw reads and include them in your metadata file. For further details on the metadata required in your submission see our metadata docs.

Required Fields

Field nameTypeDescriptionExample
sequencingInstrument
Sequencing instrument
enumThe model of the sequencing instrument used. This field is required when raw reads are supplied. Select the sequencing instrument from the pick list.
  • 454 GS
  • 454 GS 20
  • 454 GS FLX
  • 454 GS FLX Titanium
  • 454 GS FLX+
  • 454 GS Junior
  • AB 310 Genetic Analyzer
  • AB 3130 Genetic Analyzer
  • AB 3130xL Genetic Analyzer
  • AB 3500 Genetic Analyzer
  • AB 3500xL Genetic Analyzer
  • AB 3730 Genetic Analyzer
  • AB 3730xL Genetic Analyzer
  • AVITI 24
  • BGISEQ-50
  • BGISEQ-500
  • DNBSEQ-G400
  • DNBSEQ-G400 FAST
  • DNBSEQ-G50
  • DNBSEQ-T7
  • DNBSEQ-T10x4RS
  • Element AVITI
  • FASTASeq 300
  • GENIUS
  • GS111
  • Genapsys Sequencer
  • GenoCare 1600
  • GenoLab M
  • GridION
  • HiSeq X Five
  • HiSeq X Ten
  • Illumina Genome Analyzer
  • Illumina Genome Analyzer II
  • Illumina Genome Analyzer IIx
  • Illumina HiScanSQ
  • Illumina HiSeq 1000
  • Illumina HiSeq 1500
  • Illumina HiSeq 2000
  • Illumina HiSeq 2500
  • Illumina HiSeq 3000
  • Illumina HiSeq 4000
  • Illumina HiSeq X
  • Illumina MiSeq
  • Illumina MiniSeq
  • Illumina NovaSeq 6000
  • Illumina NovaSeq X
  • Illumina NovaSeq X Plus
  • Illumina iSeq 100
  • Ion GeneStudio S5
  • Ion GeneStudio S5 Plus
  • Ion GeneStudio S5 Prime
  • Ion Torrent Genexus
  • Ion Torrent PGM
  • Ion Torrent Proton
  • Ion Torrent S5
  • Ion Torrent S5 XL
  • MGISEQ-2000RS
  • MinION
  • NextSeq 1000
  • NextSeq 2000
  • NextSeq 500
  • NextSeq 550
  • Onso
  • PacBio RS
  • PacBio RS II
  • PromethION
  • Revio
  • Sentosa SQ301
  • Sequel
  • Sequel II
  • Sequel IIe
  • UG 100
  • ILLUMINA
  • PACBIO_SMRT
  • OXFORD_NANOPORE
  • BGISEQ
  • LS454
  • ION_TORRENT
  • CAPILLARY
  • DNBSEQ
  • ELEMENT
  • ULTIMA
  • VELA_DIAGNOSTICS
  • GENAPSYS
  • GENEMIND
  • TAPESTRI
  • AVITI
MinION

Desired Fields

Field nameTypeDescriptionExample
depthOfCoverage
Depth of coverage
intThe average number of reads representing a given nucleotide in the reconstructed sequence. Provide value as a fold of coverage (as a number).400
sequencingAssayType
Sequencing assay type
enumThe overarching sequencing technique that was used to determine the sequence of a biomaterial, also called library strategy in the INSDC. Select the sequencing technology from the pick list.
  • WGS
  • WGA
  • WXS
  • RNA-Seq
  • ssRNA-seq
  • snRNA-seq
  • miRNA-Seq
  • ncRNA-Seq
  • FL-cDNA
  • EST
  • Hi-C
  • ATAC-seq
  • WCS
  • RAD-Seq
  • CLONE
  • POOLCLONE
  • AMPLICON
  • CLONEEND
  • FINISHING
  • ChIP-Seq
  • MNase-Seq
  • DNase-Hypersensitivity
  • Bisulfite-Seq
  • CTS
  • MRE-Seq
  • MeDIP-Seq
  • MBD-Seq
  • Tn-Seq
  • VALIDATION
  • FAIRE-seq
  • SELEX
  • RIP-Seq
  • ChIA-PET
  • Synthetic-Long-Read
  • Targeted-Capture
  • Tethered Chromatin Conformation Capture
  • NOMe-seq
  • ChM-Seq
  • GBS
  • Ribo-seq
  • OTHER
WGS
sequencingDate
Sequencing date
dateThe date the sample was sequenced. Please format the date as YYYY-MM-DD. If day is unknown, use YYYY-MM (e.g., 2020-03). If both day and month are unknown, use YYYY (e.g., 2020). Please provide at least a year.2021-04-26
sequencingLibrarySelection
Sequencing library selection
enumThe method used to select or enrich the source material. Provide the name of the DNA or RNA sequencing library selection technology used in your study, please select an option from the pick list.
  • RANDOM
  • PCR
  • RANDOM PCR
  • RT-PCR
  • HMPR
  • MF
  • repeat fractionation
  • size fractionation
  • MSLL
  • cDNA
  • cDNA_randomPriming
  • cDNA_oligo_dT
  • PolyA
  • Oligo-dT
  • Inverse rRNA
  • Inverse rRNA selection
  • ChIP
  • ChIP-Seq
  • MNase
  • DNase
  • Hybrid Selection
  • Reduced Representation
  • Restriction Digest
  • 5-methylcytidine antibody
  • MBD2 protein methyl-CpG binding domain
  • CAGE
  • RACE
  • MDA
  • padlock probes capture method
  • other
  • unspecified
PCR
sequencingLibrarySource
Sequencing library source
enumThe type of biomaterial that was sequenced, also called library source in the INSDC. Provide the name of the DNA or RNA sequencing library source used in your study, please select an option from the pick list.
  • GENOMIC
  • GENOMIC SINGLE CELL
  • TRANSCRIPTOMIC
  • TRANSCRIPTOMIC SINGLE CELL
  • METAGENOMIC
  • METATRANSCRIPTOMIC
  • SYNTHETIC
  • VIRAL RNA
  • OTHER
VIRAL RNA
sequencingProtocol
Sequencing protocol
stringThe protocol used to generate the sequence. Provide methods and materials used to generate the sequence (e.g. amplicon strategy, primer scheme, sequencing instrument, library kit).Genomes were generated through amplicon sequencing of 1200 bp amplicons with Freed schema primers. Libraries were created using Illumina DNA Prep kits, and sequence data was produced using Miseq Micro v2 (500 cycles) sequencing kits.

Optional Fields

Field nameTypeDescriptionExample
ampliconPcrPrimerScheme
Amplicon PCR primer scheme
stringThe specifications of the primers (primer sequences, binding positions, fragment size generated etc) used to generate the amplicons to be sequenced.https://github.com/joshquick/artic-ncov2019/blob/master/primer_schemes/nCoV-2019/V3/nCoV-2019.tsv
ampliconSize
Amplicon size
stringThe length of the amplicon generated by PCR amplification including units.300bp
assemblyReferenceGenomeAccession
Reference genome accession
stringA persistent, unique identifier of a genome database entry. Provide the INSDC accession number of the reference genome used for mapping/assembly.NC_045512.2
breadthOfCoverage
Breadth of coverage
intThe percentage of the reference genome covered by the sequenced data, at a prescribed depth (depthOfCoverage). Provide the percentage as a number.95
consensusSequenceSoftwareName
Consensus sequence software name
stringThe name of software used to generate the consensus sequence. Provide the name of the software used to generate the consensus sequence.Ivar
consensusSequenceSoftwareVersion
Consensus sequence software version
stringThe version of the software used to generate the consensus sequence. Provide the version of the software used to generate the consensus sequence.1.3
dehostingMethod
Dehosting method
stringThe method used to remove host reads from the pathogen sequence. Provide the name and version number of the software used to remove host reads.Nanostripper 1.2.3
pairedEndInsertSize
Paired end insert size
intThe insert size for paired-end sequencing libraries. Provide the insert size in base pairs (as an integer).500
purposeOfSequencing
Purpose of sequencing
stringThe reason that the sample was sequenced. The reason why a sample was originally collected may differ from the reason why it was selected for sequencing. The reason a sample was sequenced may provide information about potential biases in sequencing strategy.Baseline surveillance (random sampling) [GENEPIO:0100005]
qualityControlDetails
Quality control details
stringThe details surrounding a low quality determination in a quality control assessment.CT value of 39. Low viral load. Low DNA concentration after amplification.
qualityControlDetermination
Quality control determination
stringThe determination (result) of a quality control assessment.sequence failed quality control
qualityControlIssues
Quality control issues
stringThe reason contributing to, or causing, a low quality determination in a quality control assessment.low average genome coverage
qualityControlMethodName
Quality control method name
stringThe name of the method used to assess whether a sequence passed a predetermined quality control threshold. Method names can be provided as the name of a pipeline or a link to a GitHub repository. Multiple methods should be listed and separated by a semi-colon.ncov-tools
qualityControlMethodVersion
Quality control method version
stringThe version number of the method used to assess whether a sequence passed a predetermined quality control threshold. If multiple methods were used, record the version numbers in the same order as the method names. Separate the version numbers using a semi-colon.1.2.3
rawSequenceDataProcessingMethod
Raw sequence data processing method
stringThe method used for raw data processing such as removing barcodes, adapter trimming, filtering etc. Provide the name and version numbers of the software used to process the raw data.Porechop 0.2.3
sequencedByContactEmail
Sequenced by - contact email
stringThe email address of the contact responsible for follow-up regarding the sequence. This will be displayed publicly. As personnel turnover may render an individual's email obsolete, it is more preferable to provide an address for a position or lab, to ensure accuracy of information and institutional memory.enterics@lab.ca
sequencedByContactName
Sequenced by - contact name
stringThe name or title of the contact responsible for follow-up regarding the sequence. As personnel turnover may render the contact's name obsolete, it is more preferable to provide a job title for ensuring accuracy of information and institutional memory.Enterics Lab Manager
sequencedByOrganization
Sequenced by - organization
stringThe name of the agency, organization or institution responsible for sequencing the isolate's genome.Public Health Agency of Canada (PHAC) [GENEPIO:0100551]

Compressing your files

Raw reads files must be compressed with gzip before you upload them (extension .gz). If your files are not yet compressed, you can use the instructions below.

macOS and Linux

Open a terminal, navigate to the folder containing your files, and run:

gzip raw_reads_1.fastq

This replaces raw_reads_1.fastq with raw_reads_1.fastq.gz. To compress all fastq files in the current folder at once:

gzip *.fastq

Windows

Windows does not include gzip by default, but you can use 7-Zip to compress your files (free and open source):

  1. Install 7-Zip.
  2. In File Explorer, right-click your fastq file.
  3. Choose 7-ZipAdd to archive…
  4. Set Archive format to gzip and click OK.

This creates a .gz file next to your original file. When you select several files, 7-Zip creates one .gz file per input file, which is what we need. Please do not put multiple raw reads files into a single archive.

Alternatively, if you use the Windows Subsystem for Linux (WSL), you can follow the macOS and Linux instructions above.

Dehosting (host read removal)

Overview

Users must only submit files free of any human genetic data. This is particularly relevant to raw-read sequencing data, and any files uploaded should be screened for human reads and have those reads removed. This is commonly called dehosting or host read removal. It is the responsibility of data providers to ensure that submitted data does not contain human reads that could compromise participant privacy or ethical standards. Pathoplexus performs additional screening and will reject files where it detects an excessive number of reads likely to be human. This screening should not be relied upon as a substitute for dehosting prior to submission. Even for samples without a human host, it is advised to remove reads from the host species from samples, to ensure downstream analyses are performed only on relevant viral sequences.

Available Tools

Several established tools are available for host read removal. Different projects may choose different implementations depending on their workflows, performance requirements, and supported reference databases. Examples include:

These examples are provided for reference only; any approach that effectively removes host reads while preserving non-host sequences is acceptable.

Validation

Regardless of the tool used, users should validate that:

  • Host-derived reads, particularly human reads, have been effectively removed.
  • Non-host reads are retained at an acceptable rate.
  • The filtering strategy is appropriate for the host species and sequencing protocol used.
  • The resulting dataset complies with applicable privacy, ethical, and institutional requirements before submission.

Preparing your metadata: examples

This section shows some examples of how to include raw reads in your submission to Pathoplexus. Each example shows the submission of three sequences — seq1, seq2, and seq3 — with associated raw reads. The raw reads for seq1 were generated using paired-end sequencing and are uploaded as a forward and reverse fastq file. Raw reads for seq2 and seq3 are single-ended.

Example 1: uploading individual files

To upload the raw reads files, you can click ‘Upload files’ and select the raw reads files from your local filesystem: raw_reads1_fw.fq.gz, raw_reads1_rv.fq.gz, raw_reads_2.fq.gz, and raw_reads_3.fq.gz

Metadata file

You must then configure your files.rawReads metadata column to link each of these files to its respective sequence:

id files.rawReads
seq1 raw_reads1_fw.fq.gz raw_reads1_rv.fq.gz
seq2 raw_reads_2.fq.gz
seq3 raw_reads_3.fq.gz

End result

After sequence processing and approval, the files will be shown as follows when viewing the sequences on Pathoplexus:

  • seq1: raw_reads1_fw.fq.gz raw_reads1_rv.fq.gz
  • seq2: raw_reads_2.fq.gz
  • seq3: raw_reads_3.fq.gz

Example 2: uploading a single folder

Instead of uploading files individually, you can click ‘Upload folder’ and upload a folder that contains all four raw reads files at once.

Folder structure on your machine

uploaded_folder
├── raw_reads_1_fw.fq.gz
├── raw_reads_1_rv.fq.gz
├── raw_reads_2.fq.gz
└── raw_reads_3.fq.gz

Metadata file

You must then configure your files.rawReads metadata column to link each of these files to its respective sequence:

id files.rawReads
seq1 raw_reads_1_fw.fq.gz raw_reads_1_rv.fq.gz
seq2 raw_reads_2.fq.gz
seq3 raw_reads_3.fq.gz

End result

After sequence processing and approval, the files will be shown as follows when viewing the sequences on Pathoplexus:

  • seq1: raw_reads_1_fw.fq.gz raw_reads_1_rv.fq.gz
  • seq2: raw_reads_2.fq.gz
  • seq3: raw_reads_3.fq.gz
Advanced usage: name::path syntax

So far, we’ve shown examples where the files.rawReads field contained a list of filenames, resulting in the file that’s attached to the sequence in Pathoplexus having the same name as it does on your local filesystem. However, it is also possible to use the optional name::path syntax, where name denotes how the file will be called in Pathoplexus, and path indicates the location of the file on your local machine, starting at the root of the folder you’ve uploaded.

The examples below show some of the usecases of the name::path syntax.

Advanced example 1: uploading a single folder and updating file names

If you wish to use different names for your uploaded raw reads files on Pathoplexus then how they’re called on your local file system, you can use the name::path syntax in the files.rawReads metadata field:

Folder structure on your machine

uploaded_folder
├── raw_reads_1_fw.fq.gz
├── raw_reads_1_rv.fq.gz
├── raw_reads_2.fq.gz
└── raw_reads_3.fq.gz

Metadata file

You must now configure your files.rawReads metadata column to link each of these files to its respective sequence and provide a new name (note the name::path syntax):

id files.rawReads
seq1 raw_reads_fw.fq.gz::raw_reads1_fw.fq.gz raw_reads_rv.fq.gz::raw_reads1_rv.fq.gz
seq2 raw_reads.fq.gz::raw_reads_2.fq.gz
seq3 raw_reads.fq.gz::raw_reads_3.fq.gz

End result

After sequence processing and approval, the files will be shown as follows when viewing the sequences on Pathoplexus:

  • seq1: raw_reads_fw.fq.gz raw_reads_rv.fq.gz
  • seq2: raw_reads.fq.gz
  • seq3: raw_reads.fq.gz

Advanced example 2: uploading a nested folder

If you already have your files organised into subfolders on your local file system (with a subfolder per sequence, for example) you can still upload the complete folder that contains these subfolders and use the name::path syntax to link raw reads files to sequences:

Folder structure on your machine

uploaded_folder
├── subfolder_1
│   └── raw_reads_fw.fq.gz
│   └── raw_reads_rv.fq.gz

├── subfolder_2
│   └── raw_reads.fq.gz

└── subfolder_3
    └── raw_reads.fq.gz

Metadata file

You can configure your files.rawReads metadata column to link each of these files to its respective sequence, note the name::path syntax:

id files.rawReads
seq1 raw_reads_fw.fq.gz::subfolder_1/raw_reads_fw.fq.gz raw_reads_rv.fq.gz::subfolder_1/raw_reads_rv.fq.gz
seq2 raw_reads.fq.gz::subfolder_2/raw_reads.fq.gz
seq3 raw_reads.fq.gz::subfolder_3/raw_reads.fq.gz

End result

After sequence processing and approval, the files will be shown as follows when viewing the sequences on Pathoplexus:

  • seq1: raw_reads_fw.fq.gz raw_reads_rv.fq.gz
  • seq2: raw_reads.fq.gz
  • seq3: raw_reads.fq.gz