Overview

Natural Products

Natural products are organic compounds produced by living organisms not considered essential for normal growth or reproduction but instead having ecological functions related to communication, defense, competition, and reproduction.

Natural products from plants and microorganisms have historically been an important source of medicines including antibiotics and anticancer agents (1). Natural product drug discovery efforts traditionally used bioassays to guide the isolation of active compounds, which were then checked for structural novelty. Decreasing returns using this approach led the pharmaceutical industry to move away from natural products in favor of other drug discovery platforms.

In the genomic era, access to DNA sequence data has revealed a wealth of unrealized biosynthetic potential and generated considerable excitement in the development of sequence-based approaches to natural product discovery (2). NaPDoS provides a simple, straightforward approach to assess polyketide synthase (PKS) and nonribosomal peptide synthetase (NRPS) diversity based on the analysis of ketosynthase (KS) and condensation (C) domains derived from the considerably larger genes that encode these enzymes.

Polyketides and Nonribosomal Peptides

Two major classes of bacterial secondary metabolites are polyketides and nonribosomal peptides. These compounds are often biosynthesized by large, multifunctional enzymes that sequentially construct products in an assembly line process from small carboxylic acid and amino acid building blocks [3].

Non-ribosomal peptide synthetases (NRPS) are large multi-enzyme complexes. They have a modular structure, with each module being responsible for the activation, thiolation, modification and condensation of one specific amino acid. Each module consists of a number of domains, two of which are commonly referred to as A (adenylation) and C (condensation). The A domain activates a specific amino acid (analogous to a t-RNA) and transfers it to the PCP (peptidyl carrier protein) which holds on to the growing peptide as a thioester. The C domain forms a peptide bond between the next amino acyl and the peptidyl unit. Modifying domains for epimerisation, heterocyclisation or oxidation can be additionally integrated [4].

Polyketides can be assembled by modular polyketide synthases, which are functionally related to NRPS. Contrary to peptides and their amino acid building blocks, polyketides are assembled from acyl units. Selection of the monomers is performed by acyltransferase- (AT) domains. Domains responsible for the elongation step are termed ketosynthase- (KS) domains [5].

KS and C domains Domains of PKS and NRPS
KS ketosynthase
AT acyltransferase
ACP acyl carrier protein
  

Phylogeny of KS and C domains

Domain-specific phylogenetic analyses of the different modules has shown the elongation domains in NRPS and PKS assembly lines, called C and KS domains, respectively, to be the most informative in terms of predicting pathway associations. Protein sequences for majority of these domains are grouped in pathway-specific clades according to their chemical activity [6,7].

The phylogeny of KS domains delineates the major classes of polyketide synthases, according to the architecture of their biosynthesis enzymes. Type I PKS possess a multidomain architecture, that consist of either sets of modules corresponding to the number of acyl units in the product, or a single set of catalytic domains that act iteratively. Type II PKS carry each catalytic site on a separate protein [8].

The phylogeny of NRPS C domains clearly reflects the six functional categories that are known. An LCL domain catalyzes a peptide bond between two L-amino acids, a DCL domain links an L-amino acid to a growing peptide ending with a D-amino acid, and heterocyclization domains catalyze both peptide bond formation and subsequent cyclization of cysteine, serine or threonine residues. Three other subtypes of C domains are starter domains, epimerization domains and dual epimerization/condensation domains [6].

Domain Clades

Phylogenetic trees of trimmed amino acid sequences for PKS domains can be divided into three major clades, fatty acid synthases (FAS), Type II PKS, and Type I PKS. In NaPDoS version 2, KS domains have been further subdivided into a much wider diversity of classes and subclasses, while the 8 major C domains classes have remained unchanged.

Please see the CLASSIFICATION section below for detailed descriptions of KS and C domain classes and subclasses.

KS domains

KS_phylogeny_051420.png
Cdomainstree.png

Classification

Phylogenetic analyses of ketosynthase (KS) and condensation (C) domains derived from polyketide synthase (PKS) and non-ribosomal peptide synthetase (NRPS) genes, respectively, have proven highly informative of gene architecture and function (Rausch et al., 2007, Jenke-Kodama et al., 2009). Following this logic, and as per the original NaPDoS release, the sequences in the reference database were used to generate KS and C domain maximum likelihood phylogenies from which cladding patterns and established gene functions and architectures were used to develop classification schemes. This resulted in an expansion of the type I classification scheme that now includes class designations within type II PKSs. The type II aromatic class was further delineated into subclasses based on a separate phylogeny generated from concatenated KS alpha and beta sequences, which provided better resolution than either subunit alone.

The C domain classification scheme has not changed from the original NaPDoS release, but the database and reference tree have been modified to reflect the updated naming system. After assigning a phylogeny-based classification to each sequence in the database, a subset of the sequences (414 KS and 172 C domains), representing all classes and subclasses identified, was used to construct reference trees within which query sequences can be placed.

KS Domains

The primary bifurcation in the KS phylogeny delineates type I PKSs from type II PKSs and FASs. These lineages can be further resolved into classes and subclasses as described below. The abbreviations used in the database, reference tree, and NaPDoS output are given in parentheses..

Fatty Acid Synthases

KS domains from type I and II FASs have been included in the NaPDoS database and can be distinguished in the reference tree.

Class: Type I FAS (FASI)
Large multifunctional proteins with a modular organization similar to cis-AT type I PKSs (Schweizer and Hofmann 2004).
Subclass: Bacteria and fungi (bfFASI)
Found in bacteria and fungi with the associated KSs forming a monophyletic clade in the reference tree.
Subclass: Metazoa (mFASI)
Observed to date in the phyla Chordata and Nematoda. Included in the reference database but not included in the reference tree.
Subclass: Subclass: Protist (pFASI)
Observed in the phylum Apicomplexa (alveolata). Included in the reference database but not included in the reference tree.
Class: Type II FAS (FASII)
Discrete, monofunctional proteins (Campbell and Cronan 2001). Commonly observed in bacteria and archaea.

Type I PKS

Large multifunctional proteins with a domain and modular organization (Shen et al., 2003). Three broad classes can be recognized in the KS phylogeny based on function (modular assembly line or iterative) and position of the AT domain either within or outside of the module (cis-AT or trans-AT)

Class: Modular cis-AT (cisAT)
Canonical type I PKSs that function in an assembly-line fashion (Jenke-Kodama et al., 2005). Some members of this class can be further distinguished into four additional subclasses.
Subclass: Olefin synthases (cisOLS).
KSs in the two OLS clades are associated with the biosynthesis of a terminal olefin. These are best known from cyanobacteria (Coates et al., 2014).
Subclass: Loading module (cisloading).
KSs in the first module of cis-AT modular PKSs in which the catalytic cysteine has been replaced with glutamine (KSQ). This monophyletic clade also includes KSs that are responsible for both loading and the first elongation reaction (e.g. the salA KS in salinosporamide biosynthesis).
Subclass: Hybrid (cisHybridKS).
Located immediately downstream of a peptidyl carrier protein (PCP) domain and facilitate condensation reactions on a PCP-tethered intermediate. Thus, they and are found in genes that contain both PKS and NRPS components. Hybrid KSs are present in both cis- and trans-AT PKSs and form a monophyletic clade in the reference tree.
Subclass: Tandem ECH (cistandemECH).
Associated with gene cassettes involved in the introduction of a Β-branch to a Β-keto group (often referred to as Β-branching cassettes). KSs within this subclass are located in modules immediately downstream of Β-branching cassettes and are associated with enoyl CoA hydratase (ECH) and enoyl reductase (ER) domains (Nakamura et al., 2012).
Class: Iterative cis-AT (iPKS)
These cis-AT type I PKSs function iteratively. They have been documented in bacteria and fungi and can be further divided into seven subclasses (Chen & Du, 2016; Chooi & Tang, 2012).
Subclass: Polyunsaturated fatty acids (iPKSPUFA).
Produce long chain, unsaturated fatty acids that contain multiple cis double bonds. Most commonly observed in bacteria and are more closely related to PKSs than to FASs. PUFA PKSs contain three KS domains and repetitive (from five to nine), tandem ACP domains. The three PUFA KSs form three distinct clades in the reference tree.
Subclass: Enediynes (iPKSenediyne).
Produce molecules characterized by a nine- or ten-membered ring that contains a conjugated alkyne-alkene-alkyne moiety. Observed in bacteria and comprised of a single module that contains a PPTase at the C-terminal (Chen & Du, 2016). Members of this subclass form a monophyletic clade in the reference tree.
Subclass: Aromatic (iPKSaromatic).
Produce simple aromatic compounds that usually consist of mono- or bicyclic rings. The domain organization of the single module that is observed in bacteria is homologous to that observed in partially reducing fungal PKSs (Chen & Du, 2016). Members of this subclass form a monophyletic clade in the reference phylogeny.
Subclass: Polycyclic tetramate macrolactam-like (iPKSPTM).
Produce compounds usually consisting of a tetramic acid moiety and 2-3 rings fused to a macrolactam. The BGCs that produce these molecules are comprised of a single PKS-NRPS gene despite the presence of two distinct polyketide moieties (Chen & Du, 2016). The KS domains form a monophyletic clade and have been observed in bacteria.
Subclass: Non-reducing (iPKSNR).
Produce mono- or polycyclic aromatic polyketides from non-reduced and reactive poly-β-keto chains. Characterized by the absence of all three β-keto processing domains; ketoreductase (KR), dehydratase (DH), and enoyl reductase (ER). Distinguished by the presence of starter unit ACP transacylase (SAT), product template (PT), and thioesterase (TE) domains (Chooi & Tang, 2012). Members of this subclass are observed in fungi and form a monophyletic clade in the reference tree.
Subclass: Partially reducing (iPKSPR).
Produce simple mono- or bicyclic aromatic compounds similar to the products of bacterial aromatic iPKSs. These PKSs lack ER domains and have either DH and/or KR domain(s) (Chen & Due, 2016; Gallo et al., 2013). Members of this subclass form a monophyletic clade in the reference tree and are observed in fungi.
Subclass: Highly reducing (iPKSHR).
Produce linear and cyclic non-aromatic compounds. These PKS genes possess all three β-keto processing domains (KR, DH, and ER) and often contain a C-methylation (CMeT) domain responsible for α-carbon methylation (Chooi & Tang, 2012). Some highly reducing iPKSs can be fused with a C-terminal NRPS module. Members of this subclass form a monophyletic clade in the reference tree and are observed in fungi.
Class: trans-AT
Distinguished from cis-AT PKSs by the absence of a cognate AT domain, this activity is instead provided by a discrete protein that functions in trans with the type I PKS gene(s) (Piel, 2010). The polyphyletic trans-AT KS phylogeny reflects substrate-specificity (Nguyen et al., 2008). Here three additional subclasses that related to the function of the KS domain are defined. They are observed in bacteria.
Subclass: B domain (transBdomain).
Observed within modules that facilitate chain branching and are comprised of a KS domain, a cryptic “B” domain, and an ACP domain. The functionally divergent KS domain, together with the B domain, catalyzes the addition of a malonyl unit onto an α,β-unsaturated intermediate (Bretschneider et al., 2013). These modules can also facilitate ring formation, as observed in rhizoxin biosynthesis (Bretschneider et al., 2013).
Subclass: Hybrid (transHybridKS).
Similar to cisHybridKS (see above) except the AT domain occurs in trans. Members of this subclass clade with the cis-AT KSs in the reference tree.
Subclass: Hybrid non-elongating KS (transHybridKS0).
Non-elongating KS domains (KS0) that follow an NRPS module. Identified by the lack of histidine in the conserved HGTGT catalytical motif required for decarboxylative condensation. Proposed to chaperone the peptidyl-intermediate from the upstream PCP onto the downstream module, where the subsequent KS domain performs the condensation reaction (Tang et al., 2004). Currently characterized members of this subclass follow NRPS modules that generate oxazole or thioazole moieties form the incorporation of either serine or cysteine, respectively.
Class: Metazoa (MetazoaPKS)
Detected in Metazoans including birds (Melopsittacus undulates), fish (Oryzias latipes), sea urchins (Strongylocentrotus purpuratus), and nematodes (Caenorhabditis elegans). Representative PKS sequences are included in the NaPDoS database but not in the reference tree due to a lack of experimental characterization and a newly emerging understanding of their evolutionary relationships (Sabatini et al., 2018).

Class: Protist (ProtistPKS)
Reported in the amoeba Dictyostelium discoideum and the parasite Cryptosporidium parvum. Sequences are included in the NaPDoS database but not in the reference tree because few have been experimentally characterized.

Type II PKS

Discrete, monofunctional proteins (Shen, 2003). The most commonly studied type II PKSs function iteratively and produce aromatic polyketides. Five distinct classes can be recognized in the KS phylogeny with the aromatic type II PKSs further delineated into subclasses.

Class: Aromatic (aromaticKSa) or (aromatic KSb)
Produce polycyclic aromatic compounds through the iterative decarboxylative condensation of malonyl-CoA extender units. The nascent poly-β-keto intermediate is further cyclized, aromatized, and subjected to other post-processing modifications (Hertweck et al., 2007). The KS domain encodes a heterodimer consisting of an alpha subunit (KSα) that catalyzes the condensation reaction and a beta subunit (KSΒ) that lacks the catalytic cysteine residue and has been implicated in control of the chain length (Hertweck et al., 2007). A concatenated phylogeny of the alpha and beta subunits provide insight into the early biosynthetic steps in the formation of the polyketide intermediate and was used to establish the subclass designations below. Some sequences were unable to be assigned to a subclass and are annotated as "aromaticKSa" or "aromaticKSb". While concatenated sequences were used to establish the subclass classification, the alpha and beta sequences remain un-concatenated in the reference database and are phylogenetically distinct in the reference tree

Subclass: angucycline-derived I (angucyclineIKSa) or (angucyclineIKSb).
Derived from or contain an angular tetracyclic structure comprising a benz[a]anthracene moiety. Biosynthesis is most often initiated with an acetyl-CoA starting unit followed by nine extensions using malonyl-CoA to produce a decaketide intermediate (Kharel et al., 2012). The C-C bond formed during first ring cyclization occurs between C7 and C12. Some compounds within this class do not retain the angular tetracyclic core structure despite passing through a common benz[a]anthracene-containing intermediate (e.g. kinamycin).
Subclass: angucycline-derived II (angucyclineIIKSa) or (angucyclineIIKSb)
Distinguished from angucycline-derived I by the incorporation of a methylmalonyl-CoA-derived propionyl starting unit (Waldman & Balskus 2014). Otherwise, shares a similar benz[a]anthracene moiety produced by nine extensions (with malonyl-CoA) yielding a decaketide.
Subclass: anthracycline-derived I (anthracyclineIKSa) or (anthracyclineIKSb).
Possess a linear tetracyclic core structure derived from 7,8,9,10-tetrahydro-5,12-naphtacenoquinones (Metsä-Ketelä et al., 2008). Initiated with an acetyl-CoA starting unit followed by nine extensions to yield a decaketide and a first ring cyclization between C7 and C12. Members of this group can follow different folding patterns during the subsequent three ring cyclization reactions and are not a monophyletic group.
Subclass: anthracycline-dervied II (anthracyclineIIKSa) or (anthracyclineIIKSb).
Distinguished from anthracycline-derived I by the incorporation of a methylmalonyl-CoA-derived propionyl starting unit and additional enzymes for starter unit selection in the BGCs (Metsä-Ketelä et al., 2008). Otherwise shares a tetracyclic core formed by nine extensions.
Subclass: pentangular polyphenol-derived (pentangularpolyphenolKSa) or (pentangularpolyphenolKSb).
Long-chain polyphenols that form angular polycyclic core structures (Lackner et al., 2007). Use either an acetyl-CoA starting unit or longer chain acyl units (e.g. hexanoyl and butyryl starting units). Distinguished from other aromatic type II subclasses by longer pentacyclic and hexacyclic aromatic core structures formed from dodecaketide and tridecaketide chains, respectively. The C-C bond of the first ring cyclization occurs between C9 and C14.
Subclass: tetracenomycin-derived (tetracenomycinKSa) or (tetracenomycinKSb).
Linear tetracyclic decaketide core structures resulting from nine elongations of an acetyl-CoA starting unit (Hutchinson 1997). While the core structure is anthracycline-like, the first ring cyclization occurs between C9 and C14 rather than C7 and C12.
Subclass: spore pigment (sporepigmentKSa) or (sporepigmentKSb).
Not well characterized but attributed to the biosynthesis of streptomycete spore pigments (e.g. whiE in S. coelicolor).
Class: Β-branching cassettes (betabranch)
Clustered sets of free-standing genes responsible for the introduction of Β-keto branches (often collectively referred to as Β-branching cassettes). Usually comprised of KS, HMGS (3-hydroxy-3-methylglutaryl-CoA synthase), CR (crotonyl reductase), and/or ECH (enoyl-CoA dehydratases) genes. The free-standing KS is responsible for the decarboxylation of malonyl-ACP and lacks the catalytic residues required to catalyze the condensation reaction (Piel 2010). Members of this group form a monophyletic clade.

Class: Polyenes (polyeneKSa) or (polyeneKSb)
Produce reduced linear polyenes such as ishigamide (Du et al., 2016) rather than polycyclic aromatic compounds. Iteratively acting and include multiple KS domains that are similar to the alpha and beta ketosynthase subunits of type II aromatic PKSs. The two different KSs form separate clades in the reference phylogeney. No subclasses are identified.

Class: Aryl polyenes (arylpolyeneKSa) or (arylpolyeneKSb).
Produce a polyene chain with an aryl moiety that is often substituted (Cimermancic et al., 2014). They act iteratively and include two KS domains that share similarity with alpha and beta ketosynthase subunits and form separate clades in the reference phylogeny. No subclasses are identified.

Class: Non-iterative (noniterative).
Function as an assembly line to produce compounds such as pamamycin (Rebets et al., 2015) and nonactin. The associated BGCs may also contain type III KSs. These KS domains form a monophyletic clade in the reference tree.

C Domains

The phylogeny of NRPS condensation (C) domains reflects eight functional categories (Rausch et al., 2007). Condensation domains that could not be classified are annotation as "condensation" in the database. These are not included in the reference tree.

Class: Starter (starter).
Typically, the first module of a NRPS usually does not contain a C domain. But, when present, these starter C domains acylate the first amino acid with a fatty acid, polyketide, or other molecule.

Class: LCL (LCL)
Catalyze the formation of a peptide bond between two L-amino acids.

Class: DCL (DCL)
Catalyze the formation of a peptide bond between two L-amino acids.

Class: Cyclization (cyclization)
Catalyze both peptide bond formation and the subsequent cyclization of cysteine, serine or threonine residues.

Class: Epimerization (epimerization)
Change the chirality of the last amino acid in the chain from L to D. These domains can occur either before or after a C domain that performs a condensation reaction.

Class: Dual (dual)
Catalyzes both condensation and epimerization reactions.

Class: Modified amino acid (modifiedAA)
Modifies of the incorporated amino acid, for example the dehydration of serine to dehydroalanine.

Class: Hybrid (hybridC)
Located downstream of an aminotransferase domain. Involved in the condensation of an amino acid to a growing polyketide resulting in a hybrid PKS/NRPS metabolite.

References

  1. Li JW, Vederas JC (2009) Drug discovery and natural products: end of an era or an endless frontier? Science 325, 161-165.
  2. Zerikly M, Challis GL (2009) Strategies for the discovery of new natural products by genome mining. Chembiochem 10, 625-633.
  3. Fischbach MA, Walsh CT (2006) Assembly-line enzymology for polyketide and nonribo- somal Peptide antibiotics: logic, machinery, and mechanisms. Chem Rev 106, 3468-3496.
  4. Schwarzer D, Finking R, Marahiel MA (2003) Nonribosomal peptides: from genes to products. Nat Prod Rep 20, 275-287.
  5. Hertweck, C. (2009) The biosynthetic logic of polyketide diversity. Angew Chem Int Ed Engl 48, 4688-716.
  6. Rausch, C., I. Hoof, T. Weber, W. Wohlleben, and D. H. Huson (2007) Phylogenetic analysis of condensation domains in NRPS sheds light on their functional evolution. BMC Evol Biol 7:78.
  7. Jenke-Kodama H, Dittmann E (2009) Evolution of metabolic diversity: Insights from microbial polyketide synthases. Phytochemistry.
  8. Jenke-Kodama, H., A. Sandmann, R. Mueller, and E. Dittmann. 2005. Evolutionary implications of bacterial polyketide synthases. Mol. Biol. Evol. 22:2027-2039.
  9. R. Durbin, S. Eddy, A. Krogh, and G. Mitchison, Biological sequence analysis: probabilistic models of proteins and nucleic acids, Cambridge University Press (1998)
  10. Metsa-Ketela, M. et al. (2002). Molecular evolution of aromatic polyketides and comparative sequence analysis of polyketide ketosynthase and 16S ribosomal DNA genes from various streptomyces species. AEM 68, 4472-4479
  11. Shen, B. (2003) Polyketide Biosynthesis beyond the Type I, II, and III Polyketide Synthase Paradigms. Curr. Opinion Chem. Biol.,7:285-295.