Submitting Mitochondrial and Chloroplast Genomes to GenBank
Introduction
This document guides submitters through the process of submitting organelle genome(s) to GenBank. It also provides help for the submitters who were asked to resubmit their data because of lacking or inaccurate annotation. If you are already familiar with the submission process but you need to fix your annotation, go to the following sections of the document:
The Features tab/Annotation with the Five-Column Feature Table
Troubleshooting Annotation/Feature Table Format
Submit organelle genomes through the Genome Submission Portal ONLY if they accompany organism's nuclear genome submission.
Organelle Genome Submission Path/Working with Submission Portal-GenBank (SP-GenBank)
Use SP-GenBank to submit complete or incomplete organelle genome(s) to GenBank:
- Log in to the Submission Portal
- Select -> Eukaryote
-
Under What do your sequences contain?
Select -> any other feature (coding regions, organelle genomes, regulatory regions, etc.)
-
Complete the rest of the page and click Continue
- Complete the following tabs: Submitter and Sequencing Technology
- Complete the Sequences tab:
When should this submission be released to the public?
What type of molecule did you isolate and sequence? Select -> genomic DNA
Uploaded sequences are from?
Select -> mitohondrial or chloroplast. Use Other if it is a different organelle
Selection of the organelle results in the following question:Are these complete mitochondrial genomes? Select -> Yes or No
If Yes is selected the following question appearsAre these complete mitochondrial genomes circular or linear? Select -> circular or linear
-
Complete Source Info and Source Modifiers tabs
-
The Features tab/Annotation with the Five-Column Feature Table
The Features step is an essential step in submitting your organelle genome(s). You are required to provide genes, coding regions (CDS), and other feature annotations on your genomes. Submitting your sequences without annotation or inaccurate annotation will delay issuing your accession numbers. GenBank curators will ask you to resubmit your genomes with added/corrected annotation. It means that you will need to use the FIX button to correct the SP-GenBank submission.
SP-GenBank offers two ways of providing features: (1) Add features by web forms and (2) Add features by file upload. Input forms are impractical for annotating organelle genomes with multiple features as you would need to work through the forms for each individual feature. Hence, select Add features by file upload and click 5-column tab-delimited feature table.
The format of the feature table is explained in SP-GenBank by clicking the + sign next to 5-column tab-delimited feature table file help. Our illustrated example shows a small section of the feature table as the submitter of the Chinese short-limbed skink mitochondrion (the MW327509.1 record) would have prepared:

The top panel (blue background) shows how to format the table using spreadsheet software. Just like with the source table, you need to save your final feature table as a plain-text tab-delimited file before you can import it in SP-GenBank. If you are working directly in a text editor, you need to move from one column to another with a single stroke on the tab key on your computer keyboard.
The total number of columns in the feature table is five. For each feature, locations are provided in columns 1 and 2. For a feature on the complement strand, invert the locations (the larger one of the two should be in column 1). The feature key is in column 3. Qualifiers and their values that follow are offset in columns 4 and 5.
The bottom panel (yellow background) in the image reflects the information in the FEATURES section of the processed GenBank record. The processed record has additional qualifiers that were not in the original feature table. SP-GenBank will use codon_start 1 (reading frame 1) as default if none is provided in the feature table.
You should not be providing protein_ids in a feature table as these will be uniquely assigned for each translated CDS on your genome at the time your record is processed.
The feature file that you import in SP-GenBank should contain the features of all the sequences that you imported at the Sequences step. For example, there will be three sets of feature annotations in the file for the three Physalis chloroplast genomes. Each starts with a feature definition that contains the ">" symbol and the word "Feature". The word "Feature" is followed by a space and then by the sequence ID that matches in the corresponding FASTA sequence:

You should always provide the feature locations as they are on the genome that you are submitting. You are required to determine the correct feature locations.
Your annotation software may offer the five-column feature table as one of the annotation outputs. However, as we detail in the troubleshooting section, you need to review the annotation and the table format to meet GenBank (INSDC) standards.
You can also use a feature table from a GenBank record that represents a similar genome to the one that you are submitting as your template. The troubleshooting section provides steps for obtaining a template from the NCBI web and a tip for removing unwanted information from the table.
Once you have prepared the table, upload it in SP-GenBank at the Features tab
- Select Add features by file upload
- At What Annotation File type to you have?, Select -> 5-column feature table
Click Upload file Click Choose file Click Accept (after the file is uploaded)
Upon uploading the table, SP-GenBank will translate the coding regions based on their annotation (locations). SP-GenBank will apply the proper translation table (genetic code) for translation.
SP-GenBank will also validate your submission. If there are errors, such as internal stop codons in protein translation, these will be reported at the top of the page.
Address any error reported. Use other approaches as listed in the troubleshooting section to assure accuracy of your annotation.
On the right will be Features Preview.
At this point all errors should be resolved and the features reviewed for accuracy. Click Continue
-
- Complete References tab
-
- Click Submit
Once you complete your submission, GenBank curators will communicate with you to address any remaining issues or assign accession numbers for your records. Your records will require manual processing before they are publicly released.
Do not send a new submission if you cannot find your records in GenBank. Please write gb-admin@ncbi.nlm.nih.gov and inquire about the status.
Troubleshooting Annotation/Feature Table Format
Annotation programs and copying features from a similar genome is a good way to begin annotating a new genome. However, you need to review the resulting annotations for accuracy and completeness. For example, annotation programs may put on a partial coding region because it cannot determine the complete coding region. Or an essential gene may be annotated as a misc_feature or pseudo because there is an internal stop codon in translation. This may also happen if RNA editing is required to produce a complete protein without stops or with a start methionine.
If using an annotation program:
-
Determine that the complete coding region has been annotated. Use other tools such as BLAST to evaluate the accuracy of the predicted starts and stops of the coding regions. If several features are partial this indicates that the annotation program was unable to determine the correct endpoints.
-
Review the product names. The following are examples of product names that you should modify/change:
a. beginning with 'TPA:'
b. containing the phrase 'No product string in profile'
c. containing the word 'partial' or a bracketed term
-
Review the gene features to assure that they have corresponding CDS/tRNA/rRNA features. A gene feature should not be listed or linked with a misc_feature. Lacking CDS features for protein-coding genes usually indicate the annotation software cannot predict a valid conceptual translation due to presence of internal stop codons or missing the start and/or stop codon.
-
You should check your sequence/assembly quality if the annotation tools are failing with CDS annotation of essential genes. You can find guidance on using BLAST for checking sequence quality in a series of knowledgebase articles.
-
Check the protein starts and ends against BLASTP.
-
Use alignment tools to check protein start and stop locations for groups of related proteins.
-
Annotation tools may introduce unwanted annotation lines/features that do not meet INSDC standards. Remove the lines/features that contain:
a. misc_features with the annotation program name.
b. the 'label' term.
c. systematic locus tags such as 'gene 1'.
d. the standard_name qualifier
e. BLAT output
If copying features from a similar genome:
-
Exclude (remove): locus_tags (these must be unique for each genome if included), gene_xref, GeneID, and protein_ID lines from the table.
-
Check that the annotation is reasonable. For example, a variation feature on one genome may not apply to your genome.
Downloading and adjusting feature table template:
You might have identified a GenBank record that can serve as an annotation template. Here we use the MT161478.1 record as an example. To download its feature table:
-
Click the Send to link top right above the record.
-
Select: Complete record File Feature table
-
Click the Create File button
The downloaded file will contain protein_id qualifiers (lines) that you need to remove.
You can use various text editors functions to remove repetitive unwanted lines. For example, open the downloaded file with the Notepad++ editor. To remove the lines:
-
Select: Search Mark...
-
Check the Bookmark line check box in the Mark dialogue.
-
Enter: protein_id as your search term in the Find what box of the Mark dialogue.
-
Click the Mark All button.
-
Close the Mark dialogue.
-
Select: Search Bookmark Remove Bookmarked lines
Annotation Checks for individual organism groups:
Here are some of the basic annotation checks that you need to meet for your organism groups:
I. Metazoan Mitochondria
- Each genome should have 13 CDS features, 22 tRNAs and 2 rRNAs.
There are known exceptions such as Cnidaria and Porifera. If your annotation is missing one of the expected features. Send a message to gb-admin@ncbi.nlm.nih.gov describing the difference and include the SUB number for your submission.
-
ND3 in some Testudines/Aves has an extra base that is skipped over (Mindell PMID 12572620). Annotate with a join in the nucleotide span and the exception /ribosomal_slippage
-
The CDS feature rarely overlaps the tRNA feature by more than a few nucleotides. These coding regions use the mitochondrial TAA stop.
-
Use protein BLAST to evaluate the length and similarity of the predicted proteins against those from related organisms.
-
With few exceptions (for example mollusks), the features are on both strands. Check the strandedness of the tRNAS and rRNAs.
II. Fungi
- Intronic ORFs should be located within an intron.
III. Embryophyta Plants
-
Rps12 is a trans-spliced gene consisting of 2-3 exons. This gene is not RNA edited.
-
Check the size of the tRNAs. tRNAs can consist of two exons in some cases.
-
Commonly RNA edited genes are ndhD, psbL, rpl2.
-
The ribosomal RNAs that should be annotated: 4.5S, 5S, 16S and 23S ribosomal RNA.
For more help, contact the NCBI helpdesk at: info@ncbi.nlm.nih.gov .
If you need to update already accessioned records, follow the instructions on the GenBank Update page.