Tuesday, February 12, 2013

Pleiotropy - mutations tend to affect more than 1 phenotypic character

As a continuation of the previous post on the Fundamental concepts in genetics series here is the next topic : "The pleiotropic structure of the genotype–phenotype map: the evolvability of complex organisms". The paper starts of with a description of the goals of genetics and how the Genotype-Phenotype map (GMP) is very useful in understanding genetics.

While pleiotropy has been sometimes defined as changes in one gene affecting traits that are seemingly unrelated, the part about "seemingly unrelated" is dropped probably as it is difficult to define what seems to be related or not.

The cost of complexity hypothesis postulates that complex organisms are fundamentally less evolvable compared to simpler organisms as complex organisms are more pleiotropic. Fisher's geometric model and the microscope analogy are described. While theoretical predictions are available, empirical data has only become available recently. Measurement of pleiotropy and the extent and patterns of pleiotropy are being studies with the datasets that have been generated. 

Measurement of pleiotropy is complicated by at least 2 different cases:
  1. Closely linked genes affecting 2 different traits tend to be co-inherited and can be appear to be due to pleiotropic effects of a single gene.
  2. Shared cis-regulatory element of 2 genes will affect the traits controlled by both genes. This has been called artefactual pleiotropy as pleiotropy can be defined as a character of a genes rather than that of a mutation.
 Apart from the technical difficulties involved is measuring pleiotropic effects, conceptual problems like the definition of a "phenotypic trait or character" make it difficult to have an objective measure of pleiotropy.
  • QTL data tend to overestimate pleiotropic effects due to biases caused due to linked genes
  • Gene knockout and knockdown experiments avoid the problem of closely linked genes but is only able it only measures mutations that lead to complete loss of gene function
  • Mesures of pleiotropy depend on the number and type of traits that are measured in an experiment
  • Traits that are beyond the detection limits of methods used to measure traits also tend to lead to a distorted picture of pleiotropy
  • Theoretical methods to estimate pleiotropy have used the relationship between the genetic load, effective population size and effective dimensionality of the phenotype (average pleiotropy) 
  • Distribution of fitness effects have also been used to predict phenotypic complexity based on predictions of FGM
Correlation among the traits also introduces many issues in estimation of pleiotropy.However, while universal pleiotropy was considered to be the norm, its being increasingly argued that the extent of pleiotropy is very minimal(variational modularity or restricted pleiotropy). With the datasets from Yeast, C.Elegans etc.. it is being seen that the degree of pleiotropy even with the upward biased estimates is rather low. 

While the molecular basis of Pleiotropy is not well studied, two types of pleiotropy were suggested by Gruneberg way back in 1938. Type I (initially called genuine) pleiotropy refers to multiple molecular functions of a single gene product. Type II (initially called spurious) pleiotropy refers to multiple morphological & physiological effects of a single molecular function. While, recent studies have shown type II pleiotropy to be most prevalent. Hence, the pleiotropy seen today could be a result of new biological functions being assigned to the same genes which still have the same molecular functions. However, the extent of pleiotropy is not something that is settled and will probably require a lot of work that would require solving the various issues involved in measuring and analyzing pleiotropy.


Unix for foreach loop with range and vector

To loop through values in unix or linux bash shell one can make use of the "for" loop.

Example: Specify range
 
for i in {1..24};
do
echo $i
done

One can also use C style for loops like

for ((i=1;i<=25;i++));
do
echo $i
done

It can be used like the perl's foreach loop by specifying an array. Note that range can also be descending. 


for i in {1..24} 25 50 {3..4} 100 150 {5..2} qw er st {a..z};
do
echo $i
done

You can also have strings and string range not just numbers.

Sunday, February 3, 2013

Epistasis - interaction between genes


Why are crows black? It seems the answer has two components, the first being the mechanistic and easier to answer while the second is evolutionary. While it might not be the ultimate answer (which is of course 42), it does go much further towards the "why" than just the "how".

Epistasis is simply defined as "interaction between genes". However, in his earlier paper "The Language of Gene Interaction" he explains the 2 slightly different ways the term has been used.

William Bateson is credited with "inventing" the term in 1908-09 to explain the disagreement between the segregation ratios expected based on the action of separate genes and the actual results of a dihybrid cross. The action of one locus(epistatic) masking the effects of alleles at another locus (hypostatic) gives the effect of the epistatic locus "standing upon" the hypostatic locus. However, the term has expanded to cover, 
  1. Functional relationship between genes--Functional epistasis--protein-protein interactions
  2. Genetic ordering of pathways--Compositional epistasis--"measures the effects of allele substitution against a particular fixed genetic background"
  3. The quantitative differences of allele-specific effects--Statistical epistasis--"measures the average effect of allele substitution against the population average genetic background"

Epistasis as a Tool- Flower color in sweet peas

Non-Mendelian  segregation ratio of 9:7 in the cross of two white flowers to produce violet flowers has been attributed to mutations in 2 different genes in the anthocyanin pathway. This framework has been extended to elucidate the order of genes in pathways by using knock-out mutations.

High-throughput approaches that test the effects of all possible combinations of genes in an organism are being done using comprehensive deletion and knockdown libraries along with high-throughput maintenance and screening methods. Both qualitative and quantitative experiments have tested different number of gene knockouts with various genetic backgrounds to understand the interactions between genes. However, the actual number of possible interactions is still a limiting factor. Moreover, interactions need not be a simple presence-absence effect, different expression levels, mutations at various positions in a gene etc can produce other possible combinations of interactions. While knockout libraries have been widely used in Yeast, it is not as easily applicable to many other systems. RNAi knockdown libraries have been used to overcome these limitations. 

Presence of epistatic interactions has been an obstacle to the mapping of many traits to their genes. Recently, few traits with epistatic effects have been identified in human disease genetics. Evolutionary basis of the advent of epistatic effects has been explained by atleast 3 different models. Further analysis of cases involving epistasis is required to understand the role played by evolution in generating or driving evolution.




Tuesday, January 8, 2013

Let It Be

While we have to "Imagine" , we also have to "Let It Be". Much more popular than other songs of the Beatles, this song says "Mother Mary comes to me Speaking words of wisdom". One could argue, that songs popularity comes from its religious connotation, but could also be due to its "wisdom" coming from a mother(Paul's). Another, interpretation i have heard is "Mary" being a herb!! This also makes sense, as "And in my hour of darkness she is standing right in front of me".

This song offers a solution, "And when the broken hearted people living in the world agree There will be an answer". So we can wait for the "broken hearted" to agree and "Let it be".


Sunday, January 6, 2013

Imagine there is no heaven

"Imagine" as its so often called refers to a Song that has been interpreted in various ways, criticized,  revered and even worshipped. The deceptively simple to understand lyrics and soothing music have probably contributed to its popularity in the 20th Century.

While many people have equated it to a communist recruitment song, it may indeed be a much bigger idea. In fact, you can imagine,whatever you want the song to mean. 

The very idea of a world with "Nothing to kill or die for" is so appealing to any human being irrespective of her religion, political ideology, country etc... its surprising that we still live in a world torn apart by war and strife. 

This probably is due to the human need to have something to kill or die for, a purpose for life even. Unless we have a different more attractive reason to "live for", we may just have to be content with dreams.

Only part of the song that baffles me is a "A brotherhood of man", why not a sisterhood of women? or better still a "Fellowship of Humanity"(this is actually the name of a church). One can only hope that, Lenon used the word brotherhood to mean "a group of people or all people"

A world with "no possessions" can be "Imagine"d today only as part of communism. Human pursuit of materialism can just be that "a pursuit". The irony of it was "probably" not lost on Thomas Jefferson who put "Life, liberty and the pursuit of happiness" in the US declaration of Independence. As the human race moves ahead it may "someday" come to realise that as long as we have greed we will have hunger "And the world will live as one".

Tuesday, October 16, 2012

Plant MicroRNA validation

Identification and validation of plant MicroRNA has been summarized by Meyers et.al., in the article Criteria for Annotation of Plant MicroRNAs. They build upon the previous article by Ambros et. al., which had expression and biogenesis criteria.

The expression criteria were:

  1. Hybridization of a "specific" RNA probe to a size selected RNA sample, using a method like Northern blotting.
  2. Identification of the "specific" RNA sequence in a size selected cDNA library. It is expected that this sequence matches the genomic sequence of the organism from which they were cloned.Various sequencing technologies have been used to generate size selected cDNA libraries.
 The Biogenesis criteria were:
  1. Structure prediction should support a fold-back precursor that contains the "specific" miRNA sequence within one arm of the hair-pin. Morever, this hairpin should have the lowest free energy as per a RNA-folding program like mfold "and must include at least 16 bp involving the first 22 nt of the miRNA and the other arm of the hairpin". Apart from meeting these criteria, the hairpin should also be free of any internal loops or large asymmetric bulges.
  2. The "specific" sequence and its predicted precursor fold-back secondary structure should be conserved.
  3. Increased accumulation of precursor in organisms with reduced Dicer function. 
Meeting any of these criteria individually is not sufficient to be considered a miRNA as even siRNA's meet the expression criteria and biogenesis criteria are not specific to miRNA's. Hence, both expression and biogenesis criteria are required for proper validation of miRNA as per Ambrose et.al., The more recent (2008) set of guidelines by Meyers et. al., tries to utilize the knowledge gained by studies in the 5 years since the initial criteria by Ambrose et.al in 2003.

The criteria set up by Meyers et. al., are grouped under primary and ancillary criteria with various precautions to be taken. It also has information about assigning miRNA's to families.

Primary criteria include presence of miRNA sequence in cDNA library and its validation by hybridization experiments and adherence to the expected stem-loop structure. Datasets with low coverage are to be treated with caution as they have chance of mistaking a siRNA as a miRNA.

While ancillary criteria can support a miRNA candidate that meets the primary criteria, it is considered neither necessary nor sufficient for a prediction. However, "clear" conservation is generally sufficient to annotate an miRNA, as long as it has satisfied the primary criteria in the organism in which it has an homolog.  Other criteria such as miRNA target prediction, DCL1 dependence, RDR and PolIV, PolV independence are useful for obtaining biological function information but not enough to validate a sequence as miRNA.

Recently, many new automated pipelines have been published for automating the process of miRNA identification, validation, annotation and target prediction.
  1. miRTour has a easy to use web interface to upload the EST/contigs but it has a 50Mb limit on the size of the dataset that can be uploaded.
  2. PIPmiR (Pipeline for the Identification of Plant miRNAs) provides an executable that can be downloaded and used.It has been used to predict both known and novel miRNAs in Arabidopsis.
  3. shortran: a pipeline for small RNA-seq data analysis is also available for download. 
  4. miRDeep-P is a version of miRDeep modified for use with plant transcriptomes.
  5. sRNA toolkit is a more general solution that identifies not only miRNA but other small RNA's.
The availability of numerous automated pipelines for plant miRNAs can be useful, but at the same time introduce its own set of artifactual problems.

Monday, October 15, 2012

Impossible genomes?

Few genomes are difficult (with current state of technology) to assemble due to their bizzare characteristics. The high cost of sanger sequencing, construction of FOSMID or BAC libraries, flow sorting of chromosomes and other wonderful methods makes the assembly of genomes with NGS methods difficult. However, even Sanger based methods find it difficult to sequence some genomes that are almost impossible to assemble to "completion". Genomes can be difficult due to reasons such as:

  1. Large size of genome: A genome that is very large (has many many bases) are difficult to sequence, mainly due to the higher costs involved in generating sufficient coverage. Largest known vertebrate genome is that of the Lungfish with a size of 133 Gb and canopy plant being the largest known plant genome with a size of 150 Gb. An amoeboid, Polychaos dubium might have the largest genome with a size of 670 Gb. Larger amounts of data are difficult to handle bioinformatically. Infact most assemblers would be unable to handle large amounts of data associated with these genomes. Moreover, these genomes are thought to be filled with repeats and genome duplication events making their assembly even more complicated.
  2. Repeat content of genome: Certain genomes are known to have very high transposon activity making them rich with repetitive content. These genomes need not be large, but can still be difficult to assemble due to the almost identical copies of DNA prevalent in the genome.
  3. Extremes in GC content: Certain genomes are known to have very high or very low GC content. This makes them difficult to sequence due to the bias involved in NGS methods. Although GC content extremes are constrained by the requirements imposed by the genetic code, Streptomyces coelicolor manages to have a GC-content of 72% while Plasmodium falciparum has just 20%. Apart from extremes in genome wide GC content, parts of the genome can have extremes in GC content making them difficult to sequence.
  4. Rarity of sample: Some organisms are so rare, that its almost as if they were extinct. Being able to find such species and obtaining enough DNA from them can be almost impossible. The situation is made more complicated by various legal, ethical and technical issues. Rarity of sample, could also be a result of the amazingly tiny amounts of DNA available in certain species. Cultivation of many microbial species in the lab is not yet possible and obtaining enough DNA from such species has driven research in the field of metagenomics and more recently single cell sequencing. DNA from extinct species is of lower quality and filled with many artifacts making correct assembly of genomes a daunting task. However, many of these problems have been overcome by novel methods and extinct species such as the mammoth, neanderthals, denisovan.... have been sequenced and assembled to a quality comparable to that of other NGS genome assemblies.
  5. Genome definition inconsistencies: To be able to assemble the genome of a certain species, it should be possible to define a species and what constitutes its genome. Symbiotic organisms can be difficult to delineate into distinct species, due to the high degree of inter-dependence of these species. The definition of species is a controversial subject and different interpretations of these definitions makes it controversial to claim sequencing of a particular "species".
  6. Dynamic nature of genomes: Genomes are generally though of as stable inherited genetic material which remain exactly the same over short periods of time. However, the genomes have many dynamic features. Telomeres change in length with age in almost all species. Similarly, small viral genomes with very high mutation rates can change drastically within the span of a few hours making them completely immune to a drug or conferring a new phenotype. Such changes will require re-sequencing of the genome to identify the changes to the genome.Species with different numbers of chromosomes along a cline are another type of dynamism.
Many other genome can be difficult to assemble due to other reasons?