Complete chloroplast genome sequence data of Cannabis sativa
College of Pharmacy, Heilongjiang University of Chinese Medicine, Harbin, Heilongjiang 150006, China
College of Jiamusi, Heilongjiang University of Chinese Medicine (TCM), Jiamusi, China
Key Laboratory of Basic and Applied Research of North Medicine, Heilongjiang University of Traditional Chinese Medicine, Ministry of Education, Harbin, Heilongjiang 150006, China
⁎Corresponding authors at: College of Pharmacy, Heilongjiang University of Chinese Medicine, Harbin, Heilongjiang 150006, China. renweichao@hljucm.edu.cnmawei@hljucm.edu.cnAbstract
Cannabis sativa, which is an annual erect herb belonging to the Cannabis family, possesses significant medicinal value. The entire genome of the cannabis chloroplast was sequenced using the Illumina NovaSeq 6000 on the Illumina high-throughput sequencing platform. The circular genome has a size of 153,871 base pairs (bp), consisting of two inverted repeats (IRs) that are 26,010 bp each, a large single copy region (LSC) measuring 84,038 bp, and a small single copy region (SSC) of 17,813 bp. A sum of 86 genes were identified and annotated. Among them, there were 48 protein-coding genes (PCGs), 30 tRNA genes, and 8 rRNA genes. Relying on the entire chloroplast genome and the PCGs sequence, a phylogenetic tree of six cannabis plants was established, with Arabidopsis serving as the outgroup.The phylogenetic results showed that Humulus scandens of Cannabaceae was closely related to the hemp in this study. These data enriched the genetic knowledge of Cannabidae plants and provided a basis for further research on molecular breeding and genetic breeding of Cannabidae.
Specifications TableSubject Biology Specific subject area Omics: Genomics. Type of data Tables, Figures
Raw, Analyzed.Data collection The chloroplast genomes of Cannabis sativa were sequenced using an Illumina Novaseq 6000 platform. Data source location City: Harbin City, Hei LongJiang Province.
Country: China.
Latitude and longitude: 126.65108,45.725346.
The voucher specimens were deposited at the Pharmacy of college, Heilongjiang University of Chinese Medicine, Harbin, China
(xujiao2007@sina.com) under the voucher numbers: HLJDA20240328001.Data accessibility Repository name:NGDC
Data identification number: https://ngdc.cncb.ac.cn/ under the accession
number C_AA060371.2
Direct URL to data: https://ngdc.cncb.ac.cn/genbase/search/gb/C_AA060371.2
The associated BioProject, Bio-Sample and GSA numbers are PRJCA024032, SAM118808 and CRA015623, respectively.Related research article ‘none’.
1Value of the Data
- •Cannabis has a protracted history in the realm of drug utilization. Moreover, its comprehensive chloroplast genome data can be employed for comparative analysis with the phylogenetic relationships of allied species.
- •These data will contribute to elucidating the chloroplast genome architecture of this genus and facilitate the conduction of systematic genomics investigations into cannabis.
- •These cp genome data can be harnessed to formulate potentially valuable molecular markers and to appraise the genetic diversity within cannabis populations and related species.
2Background
Cannabis sativa, which belongs to the genus Cannabis of the Cannabaceae family, is mainly distributed in Heilongjiang, Yunnan, Xinjiang and other regions in China, as well as in Sikkim, Bhutan, India and Central Asia. Currently, it exists in the wild or is cultivated in various countries, and the whole grass of cannabis is very commonly used in the medical field. It encompasses a sophisticated blend of secondary metabolites, incorporating both cannabinoids and non-cannabinoids. It has been reported that over 500 compounds have been isolated from cannabis. Of these, 125 cannabinoids have been isolated and/or verified as cannabinoids. Cannabinoid represents a distinctive C21 terpenoid phenolic compound specific to cannabis. In contrast, the non-cannabinoid components comprise non-cannabinoid phenols, flavonoids, terpenes, alkaloids and the like [1]. The medicinal constituents within it can be utilized for the treatment of constipation [2], emphysema, cholelithiasis, biliary ascariasis, hypertension, facial paralysis and other diseases triggered by diverse factors. With the aim of furnishing a basis for the resource management, protection, breeding as well as genetic research regarding this medicinal species. At present, the complete chloroplast genomes of many different varieties of hemp have been sequenced and published. For example, the chloroplast genome of Yunma 7 in Yunnan, China is 153899bp in length and has a GC content of 36.67 %. A single copy region (LSC) of 84046 bp, a small single copy region of 17831 bp, and a pair of inverted repeat regions (IR) of 26011 bp were identified in the cp genome [3]. Italian ‘‘Carmagnola’’ and Russian‘‘Dagestani’’, Both C. sativa genomes are 153 871 bp in length,. The genomes from the two C. sativa varieties differ in 16 single nucleotide polymorphisms (SNPs) [4]. Based on the published chloroplast genomes of different varieties of cannabis, important reference sequences and variation information can be provided for our research, aiding in phylogenetic analysis or variety identification. It also provides theoretical basis for subsequent genetic improvement or evolutionary research (Fig. 1).
3Data Description
Cannabis sativa is an annual erect herb that typically reaches a height ranging from 1 to 3 m. Its branches bear longitudinal grooves and are densely covered with gray-white appressed hairs. The leaves are palmately divided, with the lobes being lanceolate or linear-lanceolate in shape. The male flowers are yellow-green in color, and their pedicels are slender and drooping. The achenes are flat on one side and are encircled by persistent yellowish-brown bracts. The pericarp is brittle and features a fine reticulated pattern. The seeds are flat, while the endosperm is fleshy. The embryos are curved, and the cotyledons are thick and fleshy as well. The leaves of Heilongjiang hemp in this study were sourced from the Heilongjiang Academy of Agricultural Sciences (longitude: 126.621016, latitude: 45.685433). Extract total genomic DNA using an improved CTAB method [3]. The specimens are stored in the specimen room of Heilongjiang University of Traditional Chinese Medicine in Heilongjiang Province (HLJDA20240328001). The whole genome of the chloroplast of Cannabis sativa was sequenced using the Illumina NovaSeq 6000 platform [5]. The resultant original data amounted to approximately 4.7GB, with a GC content accounting for 37 %. The size of its chloroplast genome is 153,871 bp.The chloroplast genome exhibits a circular structure. This circular genome is composed of two copies of reverse repeats (designated as IRa and IRb, each with a length of 26,010 bp), which are separated by two distinct regions: the large single copy region (abbreviated as LSC, measuring 84,038 bp) and the small single copy region (referred to as SSC, with a length of 17,813 bp) (Fig. 2). Compared with previous studies on cannabis chloroplasts, our latest research shows that the genome of Yunnan Yunma 7 in China contains 74 protein coding genes, while the Korean non-drug variety, Cheungsam, the African variety, Yoruba Nigeria The genome contains 86 protein coding genes, suggesting that different protein coding genes may affect different genetic traits. However, Heilongjiang hemp is superior to Yunnan hemp Yunma 7 in China the Korean non-drug variety, Cheungsam, the African variety, Yoruba Nigeria [6], Italian, “Carmagnola” and Russian‘‘Dagestani’’ The total length of chloroplast genomes is almost the same. Especially with the same duplicated genes as Yunma 7 in Yunnan, China (ndhB, rps7, rps12, rps19, rpl2, rpl23, ycf1, ycf2, rrn4.5, rrn5, rrn16, and rrn23) [3]. From this, it can be inferred that the genetic background of cannabis from Heilongjiang and Yunnan provinces in China is similar, which will be beneficial for future research on chloroplasts in the cannabis genus (Table 1).Groups of Genes Genes Protein genes ATP synthase atpA, atpB, atpE, atpF, atpH, atpI Cytochrome b/f complex petA, petB, petD, petG, petL, petN RubisCO large subunit rbcL NADH dehydrogenase ndhA, ndhB ( × 2), ndhC, ndhD, ndhE, ndhF, ndhG, ndhH, ndhI, ndhJ, ndhK Photosystem I psaA, psaB, psaC, psaI, psaJ Photosystem II psbA, psbB, psbC, psbD, psbE, psbF, psbI, psbJ, psbK, psbM, psbN, psbT, psbZ, ycf3 Hypothetical chloroplast ycf1 ( × 2), ycf2 ( × 2), ycf4 reading frame Ribosomal proteins Large subunits rpl14, rpl16, rpl2 ( × 2), rpl20, rpl22, rpl23 ( × 2), rpl32, rpl33, rpl36 Small subunits rps11, rps12 ( × 2), rps14, rps15, rps16, rps18, rps19 ( × 2), rps2, rps3, rps4, rps7 ( × 2), rps8 RNA polymerase rpoA, rpoB, rpoC1, rpoC2 Acetyl-CoA carboxylase accD Inner envelope membrane cemA Cytochrome c biogenesis ccsA Protein Protease clpP Maturase matK RNAs Genes ribosomal RNAs rrn4.5S ( × 2), rrn5S ( × 2), rrn16S ( × 2), rrn23S ( × 2) transfer RNAs trnA-UGC( × 2), trnC-GCA, trnD-GUC, trnE-UUC, trnF-GAA, trnfM-CAU, trnG-GCC, trnG-UCC, trnH-GUG, trnI-CAU( × 2), trnI-GAU( × 2), trnK-UUU, trnL-CAA( × 2), trnL-UAA, trnL-UAG, trnM-CAU, trnN-GUU( × 2), trnP-UGG, trnQ-UUG, trnR-ACG( × 2), trnR-UCU, trnS-GCU( × 2), trnS-UGA, trnT-GGU, trnT-UGU, trnV-GAC( × 2), trnV-UAC, trnW-CCA, trnY-GUA
From the center going outward, the first circle indicates that Cannabis sativa. has 86 genes with a length of 153871 bp and a GC content of 38 %. The next circle displays the positions of the large single copy gene region (LSC), small single copy gene region (SSC), and reverse repeat regions (IRA and IRB). The third circle represents the content of GC. The genes located outside the fourth circle are transcribed counterclockwise. As shown in the figure, genes from different functional groups are color coded. The classification of gene functions is shown in the bottom left corner.
Among these 86 genes, eight genes, specifically trnA-UGC, trnI-CAU, trnI-GAU, trnL-CAA, trnN-GUU, trnR-ACG, trnS-GCU, and trnV-GAC, are duplicated. Moreover, 16 genes, namely trnK-UUU, rps16, trnG-UCC, atpF, rpoC1, trnL-UAA, trnV-UAC, petB, petD, rpl16, rpl2, ndhB, trnI-GAU, trnA-UGC, ndhA, and trnI-GAU, are found to contain an intron. In addition, two particular genes, namely ycf3 and clpP, have been found to incorporate two introns each.
A total of 241 simple sequence repeats (SSRs) were identified through the application of MicroSAtellite (MISA) technology.Six types of SSRs were identified using the relevant detection methods: 165 single nucleotides, 55 dinucleotides, 7 trinucleotides, 10 tetranucleotides, 2 pentanucleotides, and 2 hexanucleotides. Fig. 3 presents the types and quantities of the SSRs that have been determined (Fig. 3A). For the purpose of investigating the phylogenetic position of cannabis, the cp genome sequences of cannabis and five members of the Cannabaceae family, for which complete cp genome sequences can be obtained from the NCBI, were aligned, and a phylogenetic tree was constructed. The phylogenetic analysis demonstrated that Humulus scandens was closely affiliated with cannabis (Fig. 3B).
4Experimental Design, Materials and Methods
4.1Plant Materials and DNA Extraction
In this research endeavor, fresh cannabis leaves (with specific latitude and longitude details) were gathered from Heilongjiang Province. The experimental specimens have been deposited in the College of Pharmacy at Heilongjiang University of Chinese Medicine (accessible via the website https://yxy.hljucm.net/). The total DNA of the cannabis was isolated from approximately 110mg of dried leaves.
4.2Library Preparation, Sequencing and Sequence Analysis
The genomic DNA was fragmented, and the size was selected via agarose-gel electrophoresis. Subsequently, the selected DNA fragment was passivated a-nd ligated to the sequencing adapter. Employing the TruSeq Nano DNA Kit (I-llumina, USA), the DNA library was constructed following the standard Illuminaoperating protocol. Then, shallow sequencing (with approximately 20x coverag-e)was carried out on the Novaseq 6000 platform (Illumina, USA) with a runni-ngconfiguration of 2 × 150bp. High-quality reads were assembled using NOVOPlasty [10]. The clean data were stored in FASTA format and uploaded to the National Genome Science Data Center (NGDC; https://ngdc.cncb.ac.cn/). The subsequent analysis was based on this document. MAFFT [11] was utilized to align the DNA sequences of Morus alba, Morus nigra and five related species(onlineversion: https://mafft.cbrc.jp/alignment/se-erver/), and the phylogenetic tree was reconstructed. Phylogenetic tree based on maximum composite likelihood method, using MEGA-X [12] (https://www.megasoftware.net/) Build a neighbor connected phylogenetic tree with 1000 guided replicates.
4.3Gene Annotation
The web-based program geseq (https://chlorobox.mpimp-golm.mpg.de/geseq.html) [13] was employed to annotate the assembled genome with default parameters for the prediction of protein coding genes, tRNA genes, and rRNA genes. The Chloe program was invoked to enhance the annotation accuracy. Subseque-ntly, after the results were submitted to the National Genome Science Data Center (NGDC; https://ngdc.cncb.ac.cn/), GB2Sequin was utilized to generate a five-column tab-separated feature table. Eventually, the genome map of the roundchloroplast was drawn using OrganellarGenomeDRAW [14] (OGDRAW).
4.4Detection of SSR Markers in the Chloroplast Genome
The MIcroSAtellite (MISA) identification tool [15] (from the website that had an issue being fully parsed, but the partial URL suggests it's related to the mentioned resource) was used to perform a Simple Sequence Repetition (SSRs) scan on the chloroplast genome of hemp. The parameters set on this platform were: a minimum of 10 repetitions for single motif, 5 for two motifs, 4 for three motifs, 3 for four motifs, 3 % for five motifs, and 3′ for six nucleotide motifs. Using these parameters, the platform identified 204 simple sequence repeats (SSR) or microsatellites with lengths ranging from 3 to 27 units within the chloroplast genome of hemp.
Limitations
‘None’.
Ethics Statement
All authors have read and follow the ethical requirements for publication in Data in Brief and confirming that the current work does not involve human subjects, animal experiments, or any data collected from social media platforms.
Data Availability
Acknowledgments
This work was supported by Mechanism of WRKY Transcription Factor Regulating Platycodin Biosynthesis (CZKYF2022-F04). National Key Research and Development PROJECT, Research and Demonstration of Collection, Screening and Breeding Technology of ginseng and other genuine medicinal materials (2021YFD1600901); Talent training project supported by the Central Government for the Reform and Development of Local Colleges and Universities (ZYRCB2021008); General Program of 10.13039/100016072Postdoctoral Fund of Heilongjiang Province (LBH-Z21028).
Declaration of Competing Interest
The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper.