Friday, February 28, 2020

Role of Hypoxia in breast cancer - Alternative Splicing and Methylation

Missing out on a great opportunity can be suffocating. Suffocation or asphyxiation is the deficiency of oxygen supply in the body. Such deprivation of oxygen is known as hypoxia. Hence, it is probably not surprising that looking for opportunities in hypoxia research can result in such outcomes. The recent Nobel Prize in Physiology was awarded to Semenza and colleagues for their pioneering work on hypoxia. No doubt these scientists have avoided suffocation for so long by incorporating some of the defense mechanisms that the body has developed to counteract the lack of oxygen. The role of hypoxia in tumour microenvironment makes it even more pertinent to understand hypoxia and how it is regulated.

One has to realize that while the body as a whole can experience hypoxia, it is the individual cells that respond to this condition. While some cells might die due to the acute lack of oxygen, certain adjoining cells might manage to survive as the oxygen deficiency was not as pronounced. Understanding this cell to cell heterogeneity in dealing with the lack of oxygen would be next step in unraveling the tumor micro environment. However, this needs sophisticated equipment like the 10x Chromium System that can generate single cell RNA-seq libraries from a cell suspension. Only when you have the instrument you can generate pilot data required by the grant agency. Not surprisingly, you need the said grant money to buy the instrument in the first place. This is probably what smart people call a catch 22. 

We decided to overcome this by categorizing entire tumor samples as hypoxic or normoxic as single cell resolution continues to evade our simple minds. In order to classify the tumor into hypoxic or normoxic we needed a signature that could act as a set of features. The molecular signature database is a great source for obtaining such lists of genes involved in a specific molecular function. These signatures have been meticulously assembled by combing through numerous other published datasets. The hypoxia signature on MSigDb is actually based on approximately 80 other lists taken from various studies (including those published by the great Semenza). In order to evaluate the biological meaningfulness of this signature, we used the ShinyGO tool (as it is very Shiny) to visualize the molecular functions that are prominent in this signature set consisting of 200 genes. 

Hypoxia hallmark signature from MSIGDB enrichment for molecular function (200 genes)
While functions involving sugar metabolism are prominent and make sense, they might not be the ideal set of genes to find hypoxic (single cell like) tumor patients. Hence, we decided to make our own new signature. Details of how this was done is described in the manuscript of Pant et. al., 2020. the new custom signature identified by Pant and colleagues is a leaner list that might help achieve the goal of studying hypoxia heterogeneity by looking at inter-individual variation until intra-individual variation becomes accessible. 

Hypoxia custom signature identified by Pant et.al., 2020
Using this super fantastic new signature that is identified by cleverly combining ChIP-seq and micro-array data, we stratified the public TCGA breast cancer data into hypoxia and normoxic patients. By comparing these two groups of tumors, we hoped to understand the differences that might exists between hypoxic and normoxic regions within the tumor. Since, we are specifically interested in Alternative Splicing (due to its ability to increase complexity without any increase in gene number) and its role in hypoxia, we identified differential spliced isoforms between these two groups. Exonic regions that were isoform specific among these isoforms were identified. Only for these exons, the role of DNA methylation was assessed by looking at correlations between the expression level of these exons and the methylation level of proximal DNA measured using arrays. Code required to replicate the results is available on the github repo here: Hypoxia splicing methylation correlation.

Wednesday, December 11, 2019

Correcting the nucleotide sequence of the tiger genome at the base-pair level

Tiger is the national animal of not just India but also South Korea, Malaysia and Bangladesh. Such importance accorded to this animal is a reflection of its true grandeur. Unfortunately, the historic range of tigers has diminished drastically in this century leading to tigers being classified as an endangered species. Being an endangered large cat, considerable efforts have been directed at conservation of the tiger. Recent conservation efforts have turned to using genomic tools to answer new questions (for example, see: "Conservation priorities for endangered Indian tigers through a genomic lens").

Most studies focusing on conservation using genetic tools have been dealing with magnitude of the diversity, demographic history and its interaction with geographic extent. These approaches have helped develop strategies to control illegal trade and associated poaching. However, the use of expensive genomic tools to aid conservation efforts is still not a mainstream topic. Despite discussions regarding conservation genomics and its utility, concrete examples of genomics making a difference on the ground are still rare and far between.

When the genomes of primates such as human and chimp were sequenced almost two decades ago, the promise of comparative genomics in identifying human specific traits was of great interest. A very compelling example of differences between human and chimp within the exonic region is that of exon2 in PRM1 gene. The alignment of human and chimp genomes for this region is given below:

Human      GGTGCTGCCGCCCCAGGTACAGACCGCGATGTAGAAGACACTAATTGCACAAAATAGCACATC
Chimpanzee GGTGCTGCCGCCGCAGGTCCAGAATGAGACGTAGAAGACACTAATTGCACAGAATAGCACATC 
Originally, this pattern of four amino-acid encoding differences within a single exon was reported by Sabeti et al 2006 (Positive Natural Selection in the Human Lineage).


Human                    CGCCCCAGGTACAGACCGCGATGTAGAAGACACTAATTGC
Bonobo                   CGCCGCAGGTCCAGACTGAGACGTAGAAGACACTAATTGC
Chimpanzee               CGCCGCAGGTCCAGAATGAGACGTAGAAGACACTAATTGC
Gorilla                  CGCCGCAGGAACAGACTGAGACGTAGAAAACACTAATTGC
Orangutan                CGCCGCAGGTACAGACTGAGATGTAGAAGACACTAATTGC
Gibbon                   CGCCCCAGGTACAGGCTGAGACGTAGAAGACACTAATTGC
Sooty mangabey           CGCCGCAGGTACAGGCTGAGGTGTAGAAGATACTAATTGC
Drill                    CGCCGCAGGTACAGGCTGAGGTGTAGAAGATACTAATTGC
Olive baboon             CGCCGCAGGTACAGGCTGAGGTGTAGAAGATACTAATTGC
Gelada                   CGCCGCAGGTACAGGCTGAGGTGTAGAAGATACTAATTGC
Crab-eating macaque      CGCCGCAGGTACAGGCTGAGGTGTAGAAGATACTAATTGC
Macaque                  CGCCGCAGGTACAGGCTGAGGTGTAGAAGATACTAATTGC
Pig-tailed macaque       CGCCGCAGGTACAGGCTGAGGTGTAGAAGATACTAATTGC
Vervet-AGM               CGCCGCAGGTACAGGCTGAGGTGTAGAAGATACTAATTGC
Angola colobus           CGCCGCAGGTACAGGCTGAGGTGTAGAAGATACTAATTGC
Ugandan red Colobus      CGCCGCAGGTACAGGCGGAGGTGTAGAAGATACTAATTGC
Black snub-nosed monkey  CGCCGCAGGTACAGGCTGAGGTGTAGAAGATACTAATTGC
Golden snub-nosed monkey CGCCGCAGGTACAGGCTGAGGTGTAGAAGATACTAATTGC
Ma's night monkey        CGCCGCAGGTATAAGCCGCGGTGTAGAAGACACTAATTGC
Marmoset                 CGCCGCAGGTACAAGCTGCCATGTAGAAGATACTAATTGC
Capuchin                 CGCCGCAGGTACAGACTGAGGTGTAGAAGATACTAATTGC
Bolivian squirrel monkey CGCCGCAGGTACAAGCTGAGGTGTAGAAGATACTAATTGC
Tarsier                  CGCCGCTCCTTCCGGCTGAGGTGTAGAAGATACTGA-CGC
Mouse Lemur              CGCCGCAGGTACAGGTGTAGAAGAAGAAGATACTAAATGC
Greater bamboo lemur     CGCCGCAGGTACAGG------TGTAGAAGATACTAAATGC
Coquerel's sifaka        CGCCGCAGGTACAG---GTGTAGAAGAAGATACTAAATGC
Bushbaby                 CGCCGCAGGTACAGGCTGAGGTGTAGAAGATACTAAACGC

Using a multiple sequence alignment that spans 27 primate species we are able to further delineate the changes that have occurred in the human lineage vs those that have happened in the chimp lineage. The PRM1 gene codes for a protamine protein that acts as a substitute for histones in the chromatin of sperm during the haploid phase of spermatogenesis. Striking patterns of positive selection and associated changes in the sperm morphology have been documented in various species. Identification of such amino-acid altering substitutions between species would contribute to a better understanding of the species and help define the entity that is the focus of conservation.

Given such interesting insights at the molecular level from genome sequencing, genome sequencing of any species has the potential to reveal interesting new information about a species. The genome of the tiger was first reported by Cho et al 2013 (The tiger genome and comparative analysis with lion and snow leopard genomes). Subsequent studies have used the tiger genome for comparative analysis in many high profile papers to identify patterns of protein evolution. 

Mittal et al 2019 (Comparative analysis of corrected tiger genome provides clues to its neuronal evolution) report corrections in the genome assembly sequence of the first tiger genome published by Cho et al 2013 and currently available as PanTig1.0 on ensemble as part of the release 98 (September 2019). In addition to the support from raw read data, the authors rely upon multiple sequence alignment based ancestral states and re-sequencing data from other individuals to ensure that the corrections that they are performing are correct. Having been on biorxiv for almost a year, this corrected tiger genome will hopefully motivate a speedy update of the tiger genome assembly on ensemble. The underlying program used for genome correction is called SeqBug. It is freely available for download on its own github page and is a better version of BCD.



Sunday, August 18, 2019

Wombats are herbivores with CDCA and 15-alpha-OH as the major bile salts

The wombat looks like a overgrown rat or even a cat with rat like features. However, it is neither a rodent nor a carnivore. Being a marsupial confined in its geographic distribution to Australia, many of us have probably never seen it. However, it does look similar to the koala bear in someways. Their claws and front teeth are used for burrowing as well as eating tough vegetation. These species feed on roots and bark. A very slow metabolism has been documented and is thought to help them survive in arid environments.
Vombatus ursinus -Maria Island National Park.jpg

The bile composition of the wombat (Vombatus ursinus) has been quantified using HPLC. It mainly consists of CDCA and 15-alpha OH bile acids. It is unclear whether the other two species of Northern and Southern hairy-nosed wombats (Lasiorhinus krefftii & latifrons) have a similar bile content. Given the frequent changes in bile composition of closely related species, it is possible that bile composition might be different in these other species.

Shinde et al., explores the signatures of relaxed selection in the CYP8B1 gene and finds strong patterns of relaxed selection in the wombat CYP8B1 gene. The time between biorxiving and acceptance of the paper is fairly short given the fast turnaround time of the journal of molecular evolution. All the code used for the analysis along with detailed instructions are posted on the github-CYP8B1 page that goes with the paper. In addition to the striking pattern of relaxation seen in the wombat CYP8B1 gene sequence the manuscript also explores few other aspects related to cetaceans, birds, afrotheria and technical challenges associated with detecting relaxed selection and gene loss. By investigating population level variability of the gene in chicken, we are able to identify the CYP8B1 gene that is not annotated in the latest build of the Gallus gallus genome. Located beside the ACKR2 gene seen in the picture below, it is conserved across a large number of chicken breeds despite having acquired a stop codon in the genome of the individual used for performing genome assembly. Future versions of the chicken genome will hopefully annotate this gene.

Figure 1: Lack of annotation for the CYP8B1 gene in the chicken genome.

Saturday, February 2, 2019

Incomplete Bhojeshwar Temple - a treasure trove of insights into ancient temple construction

India is without any doubt a country filled with temples. Every temple that i have been to has a large number of religious visitors and very few "cultural" only visitors. Fortunately, i visited a very interesting temple recently. This temple is incomplete and has been for almost a thousand years and tends to attract many non-religious visitors as it does not have traditional pooja (worship).

It is unclear why the temple construction was abandoned. However, anecdotal stories about why the construction was stopped range from war, natural calamities, superstition and even divine intervention. Nonetheless, the fact that temple construction is frozen in time has made it possible for experts to study temple construction methods of the 11th century. This is similar to freezing natural phenomenon in time to study them. Study of genomes to understand evolutionary processes seems very similar to studying an incomplete temple to understand construction methods. 

Saturday, October 6, 2018

Calculating instruments built by Rajput king Sawai Jai Singh II - The Jantar Mantar's

The most fascinating monuments in India for me are the Jantar Mantar's. Literally the word Jantar Mantar means "calculating instrument", but the term also has a mystical feeling of something very unique or even magical. Unlike other monuments that were built as forts, palaces, universities and religious sites, the Jantar Mantar's are arguably the only ones devoted to pursuit of science. Unlike a university that is largely involved in teaching, these are actual astronomical instruments. It is believe that the instruments were built to have a more accurate calendar that could be used to decide auspicious days for the royalty to undertake important tasks. 

The Jantar Mantar's (in five different cities) were built by the Rajput king Sawai Jai Singh II (3 November 1688 – 21 September 1743). Note the link to 3rd November and JA. The biggest of the Jantar Mantar's is located in Jaipur, the capital of Jai Singh II. As fate would have it, i visited this Jantar Mantar recently. It has 19 different instruments built to do fairly different tasks. The pictures are posted on the English wiki. While Jai Singh's contributions are well recorded, his chief astronomer Pandit Jagannatha Samrat is probably not as popular. It is conceivable that Samrat played a very important role in the design and construction of the Jantar Mantar's as it has to be noted that Jai Singh was also involved in conducting the royal duties of ruling and conducting war. Samrat is credited with writing Siddhānta-samrāṭ and Yantra-prakāra which provide useful details about the design, construction and use of astronomical instruments. A more recent general article that provides a comprehensive idea about the Jantar Mantar's is written by N Rathnasree

In this post we start with the Unnatamsa Yantra. Unnat meaning elevated and Amsa: division or degree of arc put together gives the word "Unnatamsa". The word Yantra meaning machine is added as a suffix to each of instruments. This instrument is used to measure the altitude - the angular height of an object in the sky. It consists of a large graduated brass circle that is hanging from a supporting beam and is pivoted to rotate freely around a vertical axis. A sighting tube is provided at the center of the circle at the intersection of two cross beams in the vertical & horizontal directions. The sighting tube can be moved in the vertical direction to align it towards celestial objects. The pivoting used in the Unnatamsa is analogous to the Alt-Azimuth mounting used in modern telescopes.

Unnatamsa Yantra

The altitude of celestial objects can be measured by looking at the graduations on the rim of the brass ring after locating the object using the sighting tube. The graduations on the rim of the circle are able to distinguish upto one tenth of a degree. Larger deviations of 1 degree and 6 degrees are marked with longer graduation marks. 

Sunday, September 9, 2018

Indian Peacock genome sequence, its comparative analysis and demographic history

It has become customary for me to put any manuscript on the biorxiv way before it gets published. My first manuscript to land on the pre-print server was the killer whale culture manuscript way back in Feb 2016. The date is actually very important as this  was when biorxiv really became main stream in biology with the number of monthly submission going from 60 to 200 per month. Subsequently, the bird comparative population genomics manuscript was on biorxiv before getting published.

Now, very recently the genome of the Indian Peacock was sequenced by Dr. Vineet Sharma from IISER Bhopal. The amount of press coverage was considerable. The Hindu, Mongabay and The times of India all had something to say about this genome. Being part of the team, pushing for posting of the peacock genome manuscript on biorxiv seemed the most obvious thing to do. The manuscript has now been accepted by the journal Frontiers in Genetics after two rounds of interactive review. Screenshot of the provisionally accepted abstract is given below. 

Interestingly, it seems that the Peacock genome is the first bird genome to come out of India. The International Chicken Genome Sequencing Consortium published the first bird genome in 2004. Almost one and half decades later, we have managed to claim a bird genome sequence, even if it is based on paired-end sequencing only. Great effort has been made on the data analysis front to make up for the lack of mate-pair, nanopore, pacbio, optical mapping, genetic map, BAC and other fancy forms of data. Yet, one has to remember that this goes to suggest that perhaps we lag behind rest of the world by almost 15 years atleast in this field.


Friday, July 27, 2018

SLR's whose function does not involve taking photographs

SLR's are a type of camera that let the person using it to see through the lens and actually see the image that is being captured based on a clever use of a mirror and prism system. However, we are not talking about these SLR's in this post. Rather, we come up with a new expansion for SLR's in the form of ssDNA binding protein-like receptors. After being on bioRxiv for a fairly short period, this manuscript has been published by the Immunobiology journal. It has to be noted that the wait times for editorial decisions were reasonable and reviewer comments were well thought. The initial submission was in Feb 2018 and the first revision was already resubmitted in June 2018. Hence, this journal is definitely a good avenue for publishing immunology research.

This paper is a moderately straightforward bioinformatic exercise.  However, the amount of domain knowledge that comes into play is extraordinary for an immunology novice. Interferon induction, innate immune sensors, host range breadth, Baltimore classification are just some of the concepts that all had to be weaved together to make this hypothesis come to life. Without a doubt, the bachelor course "Flow of Genetic Information" that I taught last semester was an important connecting link that made this paper possible. The fact that the title "A hypothetical new role for single-stranded DNA binding proteins in the immune system" itself has the word hypothetical in it makes this a very different paper than the large dataset based papers that have been my forte.  

Unfortunately, we don't go beyond some preliminary testing of our hypothesis. Hopefully, this paper will stimulate debate and fuel a more sophisticated search for much more single-stranded DNA sensors. At the very least, stronger and unequivocal evidence in favor of TLR9 and IFI16 would be a step in the correct direction. Fine-scale dissection of the abilities of these innate immune sensors to distinguish between very similar yet distinct molecules is without any doubt challenging to validate experimentally. An exceedingly clever experimental paper that does this would be the ideal outcome that we would like to see as a result of this brief bioinformatic venture.