Tutorial
Workflow
Preliminary Candidate Screening
Basic Procedures
|
|
BLAST search output
| A BLAST search is performed against curated reference database examples to identify matches to known PKS/NRPS pathways. Some suggested guidelines for interpreting BLAST scores are presented below. To proceed with further analysis, one or more candidate sequences must be selected using check boxes. Three different output options are available: | |
|
|
Tree Construction
|
Tools are provided for construction of phylogenetic trees from predicted KS and
C domain sequences, because these trees consistently out-perform simple BLAST searches
in predicting the type of compound produced and/or it's relative novelty.
Selected candidate sequences plus their BLAST matches are trimmed and inserted into a manually curated amino acid reference alignment, keeping the original reference alignment intact. This alignment is used to build a tree, which is often more useful than BLAST results alone in predicting whether pathway products for candidate domains are likely to be similar or different from previously known examples [2]. |
|
Tree output options
|
|
|
|
|
|
Newick format output
(hctox1_C2_dual:1.21247,(hctox5_C3_dual:1.58329, (hctox1_C3_dual:0.94480,hctox4_C3_dual:1 .08446) 0.842:0.19209)0.855:0.17790, (cyclo1_C12_dual:1.37115,((NC_013790. 1_3_5_1279_1556:0.76822,surfa4_C3_LCL:0. 60447)1.000:1.11061, (syrin1_C2_dual:0.42670,(syrin1_C8_dual: 0.45307,(syrin1_C4_dual:0.04019, syrin1_C3_dual:0.04876) 1.000:0.37152)0.909:0.19372)0.995:0. 59611)0.884:0.23364)0.761:0.11321); |
SVG format output
|
Interpreting Results
BLAST searches
|
BLAST matches to reference domain sequences provide a good starting point, but are less
reliable for predicting domain functionality than multiple sequence alignments
and the phylogentic trees built from them. This is because BLAST scores
represent average, overall similarities distributed over the entire domain, considering
only two sequences at a time, whereas multiple sequence alignments and tree topologies
combine information from many sequences together to reveal relationships based on
more specific, potentially localized features they may have in common.
Typical amino acid percent identities vary widely within different functional domain categories, due at least partly to variability in the number and diversity of reference examples available. The following chart of leave-one-out (jack-knife) cross validation scores, shows averages and expected ranges for members of different functional classes. These values were obtained by finding the closest, non-self database match for each NaPDoS2 reference domain, and comparing amino acid match statistics to other members of the same functional classification category. Brackets after each category on the x-axis indicate the number of database reference sequences analyzed for that category. |
|
|
|
Large declines in sensitivity and classification accuracy were observed when full-length, non-redundant positive control domains were split into smaller, overlapping subsequences of less than 100 amino acids (300 nucleotides), as shown below. Relatively novel domain classes, with fewer available reference examples, were the ones most likely to be mis-classified in leave-one-out cross validation tests. Sensitivity declines can be partially offset by decreasing BLAST stringency, but this adjustment may also increase false positives and mis-classifications, depending on the nature of the data being analyzed [5]. |
|
|
|
References
Edgar RC: MUSCLE: multiple sequence alignment with high accuracy and high throughput. Nucleic Acids Res 2004, 32(5):1792-1797.
Jenke-Kodama H, Sandmann A, Muller R, Dittmann E. Evolutionary implications of bacterial polyketide synthases. Mol Biol Evol. 2005 Oct;22(10):2027-39.
Guindon S, Gascuel O: A simple, fast, and accurate algorithm to estimate large phylogenies by maximum likelihood. Syst Biol 2003, 52(5):696-704.
Goksuluk D, Korkmaz S, Zararsiz G, Karaağaoğlu AE (2016). easyROC: An Interactive Web-tool for ROC Curve Analysis Using R Language Environment. The R Journal, 8(2):213-230.
Habener LA, Podell S, Creamer, CE, Demko, AM, Allen, E, Moore, BM, Ziemert N, Letzel AC, Jensen PR. The Natural Product Domain Seeker (NaPDoS) version 2: Relating ketosynthase phylogeny to biosynthetic function. (in preparation)
