Annelida represents a large and morphologically diverse group of bilaterian organisms. The recently published polychaete and leech genome sequences revealed an equally dynamic range of diversity at the genomic level. The availability of more annelid genomes will allow for the identification of evolutionary genomic events that helped shape the annelid lineage and better understand the diversity within the group. We sequenced and assembled the genome of the common earthworm, Eisenia fetida. As a first pass at understanding the diversity within the group, we classified 440 earthworm homeoboxes and compared them to those of the leech Helobdella robusta and the polychaete Capitella teleta. We inferred many gene expansions occurring in the lineage connecting the most recent common ancestor (MRCA) of Capitella and Eisenia to the Eisenia/Helobdella MRCA. Likewise, the lineage leading from the Eisenia/Helobdella MRCA to the leech Helobdella robusta has experienced substantial gains and losses. However, the lineage leading from Eisenia/Helobdella MRCA to E. fetida is characterized by extraordinary levels of homeobox gain. The evolutionary dynamics observed in the homeoboxes of these lineages are very likely to be generalizable to all genes. These genome expansions and losses have likely contributed to the remarkable biology exhibited in this group. These results provide a new perspective from which to understand the diversity within these lineages, show the utility of sub-draft genome assemblies for understanding genomic evolution, and provide a critical resource from which the biology of these animals can be studied. The genome data can be accessed through the Eisenia fetida Genome Portal: http://ryanlab.whitney.ufl.edu/genomes/Efet/
Category Archives: Uncategorized
Genome-wide characterization of PRE-1 reveals a hidden evolutionary relationship between suidae and primates
Determination of Ubiquitin Fitness Landscapes Under Different Chemical Stresses in a Classroom Setting
Ubiquitination is an essential post-translational regulatory process that can control protein stability, localization, and activity. Ubiquitin is essential for eukaryotic life and is highly conserved, varying in only 3 amino acid positions between yeast and humans. However, recent deep sequencing studies in S. cerevisiae indicate that ubiquitin is highly tolerant to single amino acid mutations. To resolve this paradox, we hypothesized that the set of tolerated substitutions would be reduced when the cultures are not grown in rich media conditions and that chemically induced physiologic perturbations might unmask constraints on the ubiquitin sequence. To test this hypothesis, a class of first year UCSF graduate students employed a deep mutational scanning procedure to determine the fitness landscape of a library of all possible single amino acid mutations of ubiquitin in the presence of one of five small molecule perturbations: MG132, Dithiothreitol (DTT), Hydroxyurea (HU), Caffeine, and DMSO. Our data reveal that the number of tolerated substitutions is greatly reduced by DTT, HU, or Caffeine, and that these perturbations uncover “shared sensitized positions” localized to areas around the hydrophobic patch and to the C-terminus. We also show perturbation specific effects including the sensitization of His68 in HU and tolerance to mutation at Lys63 in DTT. Taken together, our data suggest that chemical stress reduces buffering effects in the ubiquitin proteasome system, revealing previously hidden fitness defects. By expanding the set of chemical perturbations assayed, potentially by other classroom-based experiences, we will be able to further address the apparent dichotomy between the extreme sequence conservation and the experimentally observed mutational tolerance of ubiquitin. Finally, this study demonstrates the realized potential of a project lab-based interdisciplinary graduate curriculum.
Genomic and gene-expression comparisons among phage-resistant type-IV pilus mutants of Pseudomonas syringae pathovar phaseolicola
Pseudomonas syringae pv. phaseolicola (Pph) is a significant bacterial pathogen of agricultural crops, and phage φ6 and other members of the dsRNA virus family Cystoviridae undergo lytic (virulent) infection of Pph, using the type IV pilus as the initial site of cellular attachment. Despite the popularity of Pph/phage φ6 as a model system in evolutionary biology, Pph resistance to phage φ6 remains poorly characterized. To investigate differences between phage φ6 resistant Pph strains, we examined genomic and gene expression variation among three bacterial genotypes that differ in the number of type IV pili expressed per cell: ordinary (wild-type), non-piliated, and super-piliated. Genome sequencing of non-piliated and super-piliated Pph identified few mutations that separate these genotypes from wild type Pph – and none present in genes known to be directly involved in type IV pilus expression. Expression analysis revealed that 81.1% of GO terms up-regulated in the non-piliated strain were down-regulated in the super-piliated strain. This differential expression is particularly prevalent in genes associated with respiration — specifically genes in the tricarboxylic acid cycle (TCA) cycle, aerobic respiration, and acetyl-CoA metabolism. The expression patterns of the TCA pathway appear to be generally up and down-regulated, in non-piliated and super-piliated Pph respectively. As pilus retraction is mediated by an ATP motor, loss of retraction ability might lead to a lower energy draw on the bacterial cell, leading to a different energy balance than wild type. The lower metabolic rate of the super-piliated strain is potentially a result of its loss of ability to retract.
Phylogenetic community structure metrics and null models: a review with new methods and software
Phylogenetic community structure metrics and null models: a review with new methods and software
Construction of relatedness matrices using genotyping-by-sequencing data
Construction of relatedness matrices using genotyping-by-sequencing data
Background Genotyping-by-sequencing (GBS) is becoming an attractive alternative to array-based methods for genotyping individuals for a large number of single nucleotide polymorphisms (SNPs). Costs can be lowered by reducing the mean sequencing depth, but this results in genotype calls of lower quality. A common analysis strategy is to filter SNPs to just those with sufficient depth, thereby greatly reducing the number of SNPs available. We investigate methods for estimating relatedness using GBS data, including results of low depth, using theoretical calculation, simulation and application to a real data set. Results We show that unbiased estimates of relatedness can be obtained by using only those SNPs with genotype calls in both individuals. The expected value of this estimator is independent of the SNP depth in each individual, under a model of genotype calling that includes the special case of the two alleles being read at random. In contrast, the estimator of self-relatedness does depend on the SNP depth, and we provide a modification to provide unbiased estimates of self-relatedness. We refer to these methods of estimation as kinship using GBS with depth adjustment (KGD). The estimators can be calculated using matrix methods, which allow efficient computation. Simulation results were consistent with the methods being unbiased, and suggest that the optimal sequencing depth is around 2-4 for relatedness between individuals and 5-10 for self-relatedness. Application to a real data set revealed that some SNP filtering may still be necessary, for the exclusion of SNPs which did not behave in a Mendelian fashion. A simple graphical method (a ‘fin plot’) is given to illustrate this issue and to guide filtering parameters. Conclusion We provide a method which gives unbiased estimates of relatedness, based on SNPs assayed by GBS, which accounts for the depth (including zero depth) of the genotype calls. This allows GBS to be applied at read depths which can be chosen to optimise the information obtained. SNPs with excess heterozygosity, often due to (partial) polyploidy or other duplications can be filtered based on a simple graphical method.
Population genomics of the Anthropocene: urbanization reduces the evolutionary potential of small mammal populations
Urbanization results in pervasive habitat fragmentation and reduces standing genetic variation through genetic drift. Loss of genome-wide variation may ultimately reduce the evolutionary potential of animal populations experiencing rapidly changing conditions. In this study, we examined genome-wide variation among 23 white-footed mouse (Peromyscus leucopus) populations sampled along an urbanization gradient in the New York City metropolitan area. Genome-wide variation was estimated as a proxy for evolutionary potential using more than 10,000 SNP markers generated by ddRAD-Seq. We found that genome-wide variation is inversely related to urbanization as measured by percent impervious surface cover, and to a lesser extent, human population density. We also report that urbanization results in enhanced genome-wide differentiation between populations in cities. There was no pattern of isolation by distance among these populations, but an isolation by resistance model based on impervious surface significantly explained patterns of genetic differentiation. Isolation by environment modeling also indicated that urban populations deviate much more strongly from global allele frequencies than suburban or rural populations. This study is the first to examine evolutionary potential along an urban-to-rural gradient and quantify urbanization as a driver of population genomics patterns.
The mysterious orphans of Mycoplasmataceae
The mysterious orphans of Mycoplasmataceae
Phylogeographic Inference Using Approximate Likelihoods
Phylogeographic Inference Using Approximate Likelihoods
The demographic history of most species is complex, with multiple evolutionary processes combining to shape the observed patterns of genetic diversity. To infer this history, the discipline of phylogeography has (to date) used models that simplify the historical demography of the focal organism, for example by assuming or ignoring ongoing gene flow between populations or by requiring a priori specification of divergence history. Since no single model incorporates every possible evolutionary process, researchers rely on intuition to choose the models that they use to analyze their data. Here, we develop an approach to circumvent this reliance on intuition. PHRAPL allows users to calculate the probability of a large number of demographic histories given their data, enabling them to identify the optimal model and produce accurate parameter estimates for a given system. Using PHRAPL, we reanalyze data from 19 recent phylogeographic investigations. Results indicate that the optimal models for most datasets parameterize both gene flow and population divergence, and suggest that species tree methods (which do not consider gene flow) are overly simplistic for most phylogeographic systems. These results highlight the importance of phylogeographic model selection, and reinforce the role of phylogeography as a bridge between population genetics and phylogenetics.
A simple approach for maximizing the overlap of phylogenetic and comparative data
A simple approach for maximizing the overlap of phylogenetic and comparative data
Biologists are increasingly using curated, public data sets to conduct phylogenetic comparative analyses. Unfortunately, there is often a mismatch between species for which there is phylogenetic data and those for which other data is available. As a result, researchers are commonly forced to either drop species from analyses entirely or else impute the missing data. Here we outline a simple solution to increase the overlap while avoiding potential the biases introduced by imputing data. If some external topological or taxonomic information is available, this can be used to maximize the overlap between the data and the phylogeny. We develop an algorithm that replaces a species lacking data with a species that has data. This swap can be made because for those two species, all phylogenetic relationships are exactly equivalent. We have implemented our method in a new R package phyndr, which will allow researchers to apply our algorithm to empirical data sets. It is relatively efficient such that taxon swaps can be quickly computed, even for large trees. To facilitate the use of taxonomic knowledge we created a separate data package taxonlookup; it contains a curated, versioned taxonomic lookup for land plants and is interoperable with phyndr. Emerging online databases and statistical advances are making it possible for researchers to investigate evolutionary questions at unprecedented scales. However, in this effort species mismatch among data sources will increasingly be a problem; evolutionary informatics tools, such as phyndr and taxonlookup, can help alleviate this issue.