Pathoplexus allows users to include raw reads when submitting sequences (see here for our general submission documentation). To include your raw reads files, you will need to upload them and link them to their respective sequences via the files.rawReads field in your metadata file. The files.rawReads field should be a space-separated list containing the names of all files associated with the sequence.
When on the submission page, you can either click on ‘Upload folder’ to upload an entire folder containing raw reads files, or ‘Upload files’ to select individual files to upload. Once you’ve selected files to upload, the webpage will let you know if there are files listed in the files.rawReads metadata column that are missing or if you’ve uploaded files that are not listed in files.rawReads. Once all the files included in files.rawReads have been uploaded and linked and no unexpected files are present, you can proceed with submission.
Raw reads files must be submitted in fastq.gz format, fastq files should be compressed using gzip (see below). You can provide either a single fastq.gz file or a pair of fastq.gz files for consensus sequences generated from single-end or paired-end reads, respectively. Note that if you are submitting paired-end reads, they must be provided as two separate files: one with the forward reads and the other with the reverse. Interleaved files are not supported.
If the raw reads for a sequence are already available in the INSDC, do not upload them again. Instead, supply their run, BioSample, and BioProject accessions in the INSDC-linking metadata fields insdcRawReadsAccession, biosampleAccession, and bioprojectAccession, and we will link to them rather than submitting a second copy.
When including raw reads files in your submission you will also need to provide additional metadata fields in your metadata file. Please refer to the table below to see which fields are required and desired when submitting raw reads and include them in your metadata file. For further details on the metadata required in your submission see our metadata docs.
| Field name | Type | Description | Example |
|---|---|---|---|
sequencingInstrument Sequencing instrument | enum | The model of the sequencing instrument used. This field is required when raw reads are supplied. Select the sequencing instrument from the pick list.
| MinION |
| Field name | Type | Description | Example |
|---|---|---|---|
depthOfCoverage Depth of coverage | int | The average number of reads representing a given nucleotide in the reconstructed sequence. Provide value as a fold of coverage (as a number). | 400 |
sequencingAssayType Sequencing assay type | enum | The overarching sequencing technique that was used to determine the sequence of a biomaterial, also called library strategy in the INSDC. Select the sequencing technology from the pick list.
| WGS |
sequencingDate Sequencing date | date | The date the sample was sequenced. Please format the date as YYYY-MM-DD. If day is unknown, use YYYY-MM (e.g., 2020-03). If both day and month are unknown, use YYYY (e.g., 2020). Please provide at least a year. | 2021-04-26 |
sequencingLibrarySelection Sequencing library selection | enum | The method used to select or enrich the source material. Provide the name of the DNA or RNA sequencing library selection technology used in your study, please select an option from the pick list.
| PCR |
sequencingLibrarySource Sequencing library source | enum | The type of biomaterial that was sequenced, also called library source in the INSDC. Provide the name of the DNA or RNA sequencing library source used in your study, please select an option from the pick list.
| VIRAL RNA |
sequencingProtocol Sequencing protocol | string | The protocol used to generate the sequence. Provide methods and materials used to generate the sequence (e.g. amplicon strategy, primer scheme, sequencing instrument, library kit). | Genomes were generated through amplicon sequencing of 1200 bp amplicons with Freed schema primers. Libraries were created using Illumina DNA Prep kits, and sequence data was produced using Miseq Micro v2 (500 cycles) sequencing kits. |
| Field name | Type | Description | Example |
|---|---|---|---|
ampliconPcrPrimerScheme Amplicon PCR primer scheme | string | The specifications of the primers (primer sequences, binding positions, fragment size generated etc) used to generate the amplicons to be sequenced. | https://github.com/joshquick/artic-ncov2019/blob/master/primer_schemes/nCoV-2019/V3/nCoV-2019.tsv |
ampliconSize Amplicon size | string | The length of the amplicon generated by PCR amplification including units. | 300bp |
assemblyReferenceGenomeAccession Reference genome accession | string | A persistent, unique identifier of a genome database entry. Provide the INSDC accession number of the reference genome used for mapping/assembly. | NC_045512.2 |
breadthOfCoverage Breadth of coverage | int | The percentage of the reference genome covered by the sequenced data, at a prescribed depth (depthOfCoverage). Provide the percentage as a number. | 95 |
consensusSequenceSoftwareName Consensus sequence software name | string | The name of software used to generate the consensus sequence. Provide the name of the software used to generate the consensus sequence. | Ivar |
consensusSequenceSoftwareVersion Consensus sequence software version | string | The version of the software used to generate the consensus sequence. Provide the version of the software used to generate the consensus sequence. | 1.3 |
dehostingMethod Dehosting method | string | The method used to remove host reads from the pathogen sequence. Provide the name and version number of the software used to remove host reads. | Nanostripper 1.2.3 |
pairedEndInsertSize Paired end insert size | int | The insert size for paired-end sequencing libraries. Provide the insert size in base pairs (as an integer). | 500 |
purposeOfSequencing Purpose of sequencing | string | The reason that the sample was sequenced. The reason why a sample was originally collected may differ from the reason why it was selected for sequencing. The reason a sample was sequenced may provide information about potential biases in sequencing strategy. | Baseline surveillance (random sampling) [GENEPIO:0100005] |
qualityControlDetails Quality control details | string | The details surrounding a low quality determination in a quality control assessment. | CT value of 39. Low viral load. Low DNA concentration after amplification. |
qualityControlDetermination Quality control determination | string | The determination (result) of a quality control assessment. | sequence failed quality control |
qualityControlIssues Quality control issues | string | The reason contributing to, or causing, a low quality determination in a quality control assessment. | low average genome coverage |
qualityControlMethodName Quality control method name | string | The name of the method used to assess whether a sequence passed a predetermined quality control threshold. Method names can be provided as the name of a pipeline or a link to a GitHub repository. Multiple methods should be listed and separated by a semi-colon. | ncov-tools |
qualityControlMethodVersion Quality control method version | string | The version number of the method used to assess whether a sequence passed a predetermined quality control threshold. If multiple methods were used, record the version numbers in the same order as the method names. Separate the version numbers using a semi-colon. | 1.2.3 |
rawSequenceDataProcessingMethod Raw sequence data processing method | string | The method used for raw data processing such as removing barcodes, adapter trimming, filtering etc. Provide the name and version numbers of the software used to process the raw data. | Porechop 0.2.3 |
sequencedByContactEmail Sequenced by - contact email | string | The email address of the contact responsible for follow-up regarding the sequence. This will be displayed publicly. As personnel turnover may render an individual's email obsolete, it is more preferable to provide an address for a position or lab, to ensure accuracy of information and institutional memory. | enterics@lab.ca |
sequencedByContactName Sequenced by - contact name | string | The name or title of the contact responsible for follow-up regarding the sequence. As personnel turnover may render the contact's name obsolete, it is more preferable to provide a job title for ensuring accuracy of information and institutional memory. | Enterics Lab Manager |
sequencedByOrganization Sequenced by - organization | string | The name of the agency, organization or institution responsible for sequencing the isolate's genome. | Public Health Agency of Canada (PHAC) [GENEPIO:0100551] |
Raw reads files must be compressed with gzip before you upload them (extension .gz). If your files are not yet compressed, you can use the instructions below.
macOS and Linux
Open a terminal, navigate to the folder containing your files, and run:
gzip raw_reads_1.fastq
This replaces raw_reads_1.fastq with raw_reads_1.fastq.gz. To compress all fastq files in the current folder at once:
gzip *.fastq
Windows
Windows does not include gzip by default, but you can use 7-Zip to compress your files (free and open source):
fastq file.gzip and click OK.This creates a .gz file next to your original file. When you select several files, 7-Zip creates one .gz file per input file, which is what we need. Please do not put multiple raw reads files into a single archive.
Alternatively, if you use the Windows Subsystem for Linux (WSL), you can follow the macOS and Linux instructions above.
Users must only submit files free of any human genetic data. This is particularly relevant to raw-read sequencing data, and any files uploaded should be screened for human reads and have those reads removed. This is commonly called dehosting or host read removal. It is the responsibility of data providers to ensure that submitted data does not contain human reads that could compromise participant privacy or ethical standards. Pathoplexus performs additional screening and will reject files where it detects an excessive number of reads likely to be human. This screening should not be relied upon as a substitute for dehosting prior to submission. Even for samples without a human host, it is advised to remove reads from the host species from samples, to ensure downstream analyses are performed only on relevant viral sequences.
Several established tools are available for host read removal. Different projects may choose different implementations depending on their workflows, performance requirements, and supported reference databases. Examples include:
These examples are provided for reference only; any approach that effectively removes host reads while preserving non-host sequences is acceptable.
Regardless of the tool used, users should validate that:
This section shows some examples of how to include raw reads in your submission to Pathoplexus. Each example shows the submission of three sequences — seq1, seq2, and seq3 — with associated raw reads. The raw reads for seq1 were generated using paired-end sequencing and are uploaded as a forward and reverse fastq file. Raw reads for seq2 and seq3 are single-ended.
To upload the raw reads files, you can click ‘Upload files’ and select the raw reads files from your local filesystem: raw_reads1_fw.fq.gz, raw_reads1_rv.fq.gz, raw_reads_2.fq.gz, and raw_reads_3.fq.gz
Metadata file
You must then configure your files.rawReads metadata column to link each of these files to its respective sequence:
| id | … | files.rawReads |
|---|---|---|
| seq1 | … | raw_reads1_fw.fq.gz raw_reads1_rv.fq.gz |
| seq2 | … | raw_reads_2.fq.gz |
| seq3 | … | raw_reads_3.fq.gz |
End result
After sequence processing and approval, the files will be shown as follows when viewing the sequences on Pathoplexus:
seq1: raw_reads1_fw.fq.gz raw_reads1_rv.fq.gzseq2: raw_reads_2.fq.gzseq3: raw_reads_3.fq.gzInstead of uploading files individually, you can click ‘Upload folder’ and upload a folder that contains all four raw reads files at once.
Folder structure on your machine
uploaded_folder
├── raw_reads_1_fw.fq.gz
├── raw_reads_1_rv.fq.gz
├── raw_reads_2.fq.gz
└── raw_reads_3.fq.gz
Metadata file
You must then configure your files.rawReads metadata column to link each of these files to its respective sequence:
| id | … | files.rawReads |
|---|---|---|
| seq1 | … | raw_reads_1_fw.fq.gz raw_reads_1_rv.fq.gz |
| seq2 | … | raw_reads_2.fq.gz |
| seq3 | … | raw_reads_3.fq.gz |
End result
After sequence processing and approval, the files will be shown as follows when viewing the sequences on Pathoplexus:
seq1: raw_reads_1_fw.fq.gz raw_reads_1_rv.fq.gzseq2: raw_reads_2.fq.gzseq3: raw_reads_3.fq.gzname::path syntaxSo far, we’ve shown examples where the files.rawReads field contained a list of filenames, resulting in the file that’s attached to the sequence in Pathoplexus having the same name as it does on your local filesystem. However, it is also possible to use the optional name::path syntax, where name denotes how the file will be called in Pathoplexus, and path indicates the location of the file on your local machine, starting at the root of the folder you’ve uploaded.
The examples below show some of the usecases of the name::path syntax.
If you wish to use different names for your uploaded raw reads files on Pathoplexus then how they’re called on your local file system, you can use the name::path syntax in the files.rawReads metadata field:
Folder structure on your machine
uploaded_folder
├── raw_reads_1_fw.fq.gz
├── raw_reads_1_rv.fq.gz
├── raw_reads_2.fq.gz
└── raw_reads_3.fq.gzMetadata file
You must now configure your files.rawReads metadata column to link each of these files to its respective sequence and provide a new name (note the name::path syntax):
| id | … | files.rawReads |
|---|---|---|
| seq1 | … | raw_reads_fw.fq.gz::raw_reads1_fw.fq.gz raw_reads_rv.fq.gz::raw_reads1_rv.fq.gz |
| seq2 | … | raw_reads.fq.gz::raw_reads_2.fq.gz |
| seq3 | … | raw_reads.fq.gz::raw_reads_3.fq.gz |
End result
After sequence processing and approval, the files will be shown as follows when viewing the sequences on Pathoplexus:
seq1: raw_reads_fw.fq.gz raw_reads_rv.fq.gzseq2: raw_reads.fq.gzseq3: raw_reads.fq.gzIf you already have your files organised into subfolders on your local file system (with a subfolder per sequence, for example) you can still upload the complete folder that contains these subfolders and use the name::path syntax to link raw reads files to sequences:
Folder structure on your machine
uploaded_folder
├── subfolder_1
│ └── raw_reads_fw.fq.gz
│ └── raw_reads_rv.fq.gz
│
├── subfolder_2
│ └── raw_reads.fq.gz
│
└── subfolder_3
└── raw_reads.fq.gzMetadata file
You can configure your files.rawReads metadata column to link each of these files to its respective sequence, note the name::path syntax:
| id | … | files.rawReads |
|---|---|---|
| seq1 | … | raw_reads_fw.fq.gz::subfolder_1/raw_reads_fw.fq.gz raw_reads_rv.fq.gz::subfolder_1/raw_reads_rv.fq.gz |
| seq2 | … | raw_reads.fq.gz::subfolder_2/raw_reads.fq.gz |
| seq3 | … | raw_reads.fq.gz::subfolder_3/raw_reads.fq.gz |
End result
After sequence processing and approval, the files will be shown as follows when viewing the sequences on Pathoplexus:
seq1: raw_reads_fw.fq.gz raw_reads_rv.fq.gzseq2: raw_reads.fq.gzseq3: raw_reads.fq.gz