| Period | Institution and degree | Location |
|---|---|---|
| Jan 2016 – Aug 2020 | The University of Queensland PhD, School of Chemistry & Molecular Biosciences | Brisbane, Australia |
| Sep 2006 – Jun 2009 | National Taiwan University MSc in Biomedical Engineering | Taipei, Taiwan |
| Sep 2001 – Jun 2006 | National Chung Cheng University BSc in Computer Science and Information Engineering | Chiayi, Taiwan |
| Period | Position | Institution | Location |
|---|---|---|---|
| Nov 2009 – May 2015 | Research Assistant | Academia Sinica, Institute of Information Science | Taipei, Taiwan |
Lai, Jhih-Siang. Protein structural phylogeny, a missing chapter in molecular evolutionary biology. PhD thesis, School of Chemistry and Molecular Biosciences, The University of Queensland, 2020. DOI: 10.14264/uql.2020.984
Abstract Protein evolution has been studied intensively based on protein amino acids, i.e. the protein sequence. Thanks to the next generation sequencing, a huge number of protein sequences can be collected and compared to extract evolutionary information. The evolutionary information contributes to protein bioinformatics and structural biology fields such as protein structure prediction, protein-protein interaction prediction and phylogenetic tree analysis. When proteins evolve over long evolutionary time, their homologous relationships become remote because of the low sequence identity, which makes it difficult to explore protein functions. Thus, what is required to explain functional mechanisms is determining the high-resolution protein structure and comparatively studying superposed structures. It is because protein function is connected to protein structure in that the latter is more conserved than sequence under the evolutionary pressure. However, research on evolutionary effect of structure (i.e., structural variations) in folded protein structures has been limited, although protein structure varies during the history of life because mutations in DNA influence protein amino acid sequence and affect protein structure. Thus, I hypothesised that protein structural evolutionary information helps us to explore protein function by capturing the remote homologous relationships at the folded structure level and complement protein evolutionary information.
In my PhD project, I built a structure-based evolutionary model, developed analytic tools to reconstruct structure-based phylogenetic trees and developed a novel structure-supported pairwise sequence alignment method. Specifically, I constructed a structure-based evolutionary model based on the transition probabilities of 7 secondary structure states, laying the foundation for phylogenetic analysis of superposed structures. Furthermore, based on the concept of structural evolution, a novel pairwise sequence alignment was developed to align the sequences with a 60-state model of all combinations of 3 predicted secondary structure states and the 20 amino acids; the new alignment method outperforms the amino acid-based substitution matrix with remote homologous pairwise alignments. These techniques were further applied to superfamilies of the Toll-like/interleukin-1 receptor (TIR) domains and armadillo-repeat (ARM) domains. Members of these two superfamilies are not only difficult to analyse because of the low sequence identities but also important in biological functions. TIR domains have been studied in innate immune responses across species. My structurebased phylogenetic analysis is the first to link plant TIR domains to domains with enzyme activity, explaining signalling in plant cell death pathways. Human SARM1 (sterile alpha and TIR motif containing 1 protein) plays an important role in the Wallerian degeneration pathway, and it has an armadillo-repeat (ARM) domain, which has been hypothesised to regulate SARM1 itself. The 60-state alignment method identified a possible regulatory region, and the results have been verified using the Drosophila SARM1 protein.
In conclusion, this thesis proposes the first structural evolutionary model for secondary structure states, based on all high-resolution protein structures deposited in the Protein Data Bank. This model makes phylogenetic analysis possible for any superposed structures, facilitating a number of applications, including exploring protein function. Using the phylogenetic tree inference with the novel evolutionary model and the novel pairwise alignment method developed in this PhD project, a new enzymatic function has been identified for plant TIR domains, and the binding sequence within the SARM1 protein for its ARM domain has been verified subsequently. Such critical knowledge of functional mechanisms analysed by techniques developed in this project supports the concept of the structural evolution in protein phylogenetic analysis. It is expected that structural evolutionary information for any protein structure can be studied without being limited to amino acid sequence information, revealing how protein structures evolve and bridging the gap between protein amino acid sequence and protein structure.
1. Lai J-S, Rost B, Kobe B, and Bodén M. Evolutionary model of protein secondary structure capable of revealing new biological relationships. Proteins (2020). DOI
Abstract
Ancestral sequence reconstruction has had recent success in decoding the origins and the determinants of complex protein functions. However, phylogenetic analyses of remote homologues must handle extreme amino-acid sequence diversity resulting from extended periods of evolutionary change. We exploited the wealth of protein structures to develop an evolutionary model based on protein secondary structure. The approach follows the differences between discrete secondary structure states observed in modern proteins and those hypothesised in their immediate ancestors. We implemented maximum likelihood-based phylogenetic inference to reconstruct ancestral secondary structure. The predictive accuracy from the use of the evolutionary model surpasses that of comparative modelling and sequence-based prediction; the reconstruction extracts information not available from modern structures or the ancestral sequences alone. Based on a phylogenetic analysis of a sequence-diverse protein family, we showed that the model can highlight relationships that are evolutionarily rooted in structure and not evident in amino acid-based analysis.
BibTeX
@article{Lai2020,
author = {Lai, Jhih-Siang and Rost, Burkhard and Kobe, Bostjan and Bodén, Mikael},
title = {Evolutionary model of protein secondary structure capable of revealing new biological relationships},
journal = {Proteins},
publisher = {John Wiley & Sons, Ltd},
year = {2020},
url = {https://doi.org/10.1002/prot.25898},
doi = {10.1002/prot.25898}
}
2. Horsefield S, Burdett H, Zhang X, Manik MK, Shi Y, Chen J, Qi T, Gilley J, Lai J-S, Rank MX, Casey LW, Gu W, Ericsson DJ, Foley G, Hughes RO, Bosanac T, von Itzstein M, Rathjen JP, Nanson JD, Boden M, Dry IB, Williams SJ, Staskawicz BJ, Coleman MP, Ve T, Dodds PN, and Kobe B. NAD+ cleavage activity by animal and plant TIR domains in cell death pathways. Science 365(6455):793 (2019). Article
Abstract
One way that plants respond to pathogen infection is by sacrificing the infected cells. The nucleotide-binding leucine-rich repeat immune receptors responsible for this hypersensitive response carry Toll/interleukin-1 receptor (TIR) domains. In two papers, Horsefield et al. and Wan et al. report that these TIR domains cleave the metabolic cofactor nicotinamide adenine dinucleotide (NAD+) as part of their cell-death signaling in response to pathogens. Similar signaling links mammalian TIR-containing proteins to NAD+ depletion during Wallerian degeneration of neurons.
SARM1 (sterile alpha and TIR motif containing 1) is responsible for depletion of nicotinamide adenine dinucleotide in its oxidized form (NAD+) during Wallerian degeneration associated with neuropathies. Plant nucleotide-binding leucine-rich repeat (NLR) immune receptors recognize pathogen effector proteins and trigger localized cell death to restrict pathogen infection. Both processes depend on closely related TIR domains in these proteins, which feature self-association-dependent NAD+ cleavage activity associated with cell death signaling. SARM1 SAM (sterile alpha motif) domains form an octamer essential for axon degeneration that contributes to TIR domain enzymatic activity. The crystal structures of ribose and NADP+ complexes of SARM1 and plant NLR RUN1 TIR domains, respectively, reveal a conserved substrate binding site. NAD+ cleavage by TIR domains is therefore a conserved feature of animal and plant cell death signaling pathways.
BibTeX
@article{Horsefield2019,
author = {Horsefield, Shane and Burdett, Hayden and Zhang, Xiaoxiao and Manik, Mohammad K. and Shi, Yun and Chen, Jian and Qi, Tiancong and Gilley, Jonathan and Lai, Jhih-Siang and Rank, Maxwell X. and Casey, Lachlan W. and Gu, Weixi and Ericsson, Daniel J. and Foley, Gabriel and Hughes, Robert O. and Bosanac, Todd and von Itzstein, Mark and Rathjen, John P. and Nanson, Jeffrey D. and Boden, Mikael and Dry, Ian B. and Williams, Simon J. and Staskawicz, Brian J. and Coleman, Michael P. and Ve, Thomas and Dodds, Peter N. and Kobe, Bostjan},
title = {NAD+ cleavage activity by animal and plant TIR domains in cell death pathways},
journal = {Science},
year = {2019},
volume = {365},
number = {6455},
pages = {793},
url = {http://science.sciencemag.org/content/365/6455/793.abstract}
}
3. Lai J-S, Cheng C-W, Lo A, Sung T-Y, and Hsu W-L. Lipid exposure prediction enhances the inference of rotational angles of transmembrane helices. BMC Bioinformatics 14(1):304 (2013). DOI
Abstract
Since membrane protein structures are challenging to crystallize, computational approaches are essential for elucidating the sequence-to-structure relationships. Structural modeling of membrane proteins requires a multidimensional approach, and one critical geometric parameter is the rotational angle of transmembrane helices. Rotational angles of transmembrane helices are characterized by their folded structures and could be inferred by the hydrophobic moment; however, the folding mechanism of membrane proteins is not yet fully understood. The rotational angle of a transmembrane helix is related to the exposed surface of a transmembrane helix, since lipid exposure gives the degree of accessibility of each residue in lipid environment. To the best of our knowledge, there have been few advances in investigating whether an environment descriptor of lipid exposure could infer a geometric parameter of rotational angle.
BibTeX
@article{Lai2013,
author = {Lai, Jhih-Siang and Cheng, Cheng-Wei and Lo, Allan and Sung, Ting-Yi and Hsu, Wen-Lian},
title = {Lipid exposure prediction enhances the inference of rotational angles of transmembrane helices},
journal = {BMC Bioinformatics},
year = {2013},
volume = {14},
number = {1},
pages = {304},
url = {https://doi.org/10.1186/1471-2105-14-304}
}
4. Lai J-S, Cheng C-W, Sung T-Y, and Hsu W-L. Computational comparative study of tuberculosis proteomes using a model learned from signal peptide structures. PLOS ONE 7(4):e35018 (2012). PubMed
Abstract
Secretome analysis is important in pathogen studies. A fundamental and convenient way to identify secreted proteins is to first predict signal peptides, which are essential for protein secretion. However, signal peptides are highly complex functional sequences that are easily confused with transmembrane domains. Such confusion would obviously affect the discovery of secreted proteins. Transmembrane proteins are important drug targets, but very few transmembrane protein structures have been determined experimentally; hence, prediction of the structures is essential. In the field of structure prediction, researchers do not make assumptions about organisms, so there is a need for a general signal peptide predictor.
To improve signal peptide prediction without prior knowledge of the associated organisms, we present a machine-learning method, called SVMSignal, which uses biochemical properties as features, as well as features acquired from a novel encoding, to capture biochemical profile patterns for learning the structures of signal peptides directly. We tested SVMSignal and five popular methods on two benchmark datasets from the SPdb and UniProt/Swiss-Prot databases, respectively. Although SVMSignal was trained on an old dataset, it performed well, and the results demonstrate that learning the structures of signal peptides directly is a promising approach. We also utilized SVMSignal to analyze proteomes in the entire HAMAP microbial database. Finally, we conducted a comparative study of secretome analysis on seven tuberculosis-related strains selected from the HAMAP database. We identified ten potential secreted proteins, two of which are drug resistant and four are potential transmembrane proteins.
SVMSignal was originally made available at http://bio-cluster.iis.sinica.edu.tw/SVMSignal.
BibTeX
@article{Lai2012,
author = {Lai, Jhih-Siang and Cheng, Cheng-Wei and Sung, Ting-Yi and Hsu, Wen-Lian},
title = {Computational comparative study of tuberculosis proteomes using a model learned from signal peptide structures},
journal = {PLOS ONE},
publisher = {Public Library of Science},
year = {2012},
volume = {7},
number = {4},
pages = {e35018},
url = {https://pubmed.ncbi.nlm.nih.gov/22496884}
}
5. Tsai K-N, Lin S-H, Shih S-R, Lai J-S, and Chen C-M. Genomic splice site prediction algorithm based on nucleotide sequence pattern for RNA viruses. Computational Biology and Chemistry 33(2):171–175 (2009). Article
Abstract
Splice site prediction on an RNA virus has two potential difficulties seriously degrading the performance of most conventional splice site predictors. One is a limited number of strains available for a virus species and the other is the diversified sequence patterns around the splice sites caused by the high mutation frequency. To overcome these two difficulties, a new algorithm called Genomic Splice Site Prediction (GSSP) algorithm was proposed for splice site prediction of RNA viruses. The key idea of the GSSP algorithm was to characterize the interdependency among the nucleotides and base positions based on the eigen-patterns. Identified by a sequence pattern mining technique, each eigen-pattern specified a unique composition of the base positions and the nucleotides occurring at the positions. To remedy the problem of insufficient training data due to the limited number of strains for an RNA virus, a cross-species strategy was employed in this study. The GSSP algorithm was shown to be effective and superior to two conventional methods in predicting the splice sites of five RNA species in the Orthomyxoviruses family. The sensitivity and specificity achieved by the GSSP algorithm was higher than 99% and 94%, respectively, for the donor sites, and was higher than 96% and 92%, respectively, for the acceptor sites.
BibTeX
@article{Tsai2009,
author = {Tsai, Kun-Nan and Lin, Shu-Hung and Shih, Shin-Ru and Lai, Jhih-Siang and Chen, Chung-Ming},
title = {Genomic splice site prediction algorithm based on nucleotide sequence pattern for RNA viruses},
journal = {Computational Biology and Chemistry},
year = {2009},
volume = {33},
number = {2},
pages = {171--175},
url = {http://www.sciencedirect.com/science/article/pii/S1476927108001278}
}
| Scholarship | Year | Description |
|---|---|---|
| Candidate Travel Award | 2019 | Scholarship support for a six-month visit to Technische Universität München, Germany (January–July 2019). |
| University of Queensland International Scholarship (UQI) | 2016 | Tuition fee award for three years. |
| Research Higher Degree Scholarship | 2016 | Living cost support for three years. |
| Year | Conference | Presentation |
|---|---|---|
| 2019 | ISMB — Intelligent Systems for Molecular Biology | Phylogenetic analysis in the predicted secondary structure space (poster) |
| 2017 | ISMB — Intelligent Systems for Molecular Biology | Modelling the evolution of protein secondary structure (poster) |
| 2014 | CASP — Critical Assessment of Structure Prediction | Protein residue-residue contact prediction by co-evolution analysis and machine learning (poster) |
Java, C/C++, MATLAB, R, Python
Created by sean (Jhih-Siang Lai) on 2020/03/20 13:21.
Contributing authors: