Human genomes contain repeated segments of DNA through most of the genome. When the number of copies of the repeats vary between different human beings, we have a copy number variation. Few regions of the genomes are popular locations for such variations, these regions which have different number of copies are known as copy number variant regions.
Copy number variable regions(CNVR's) are of particular interest as they can be responsible for diseases and prototypic differences between individuals. Similar to single nucleotide polymorphisms (SNP's) which are disease markers, these regions have been associate with many conditions. These regions have also been associated with resistance from infection by HIV and Malaria.
The importance of these regions becomes apparent as copy number variation maps are being generated and updated into genomic variant databases. This type of data will be useful in understanding the relations between CNVR's and specific characters. Evolutinary impact of these regions could also be anlaysed to get an understanding of how evolution proceeded.
It will require a more clear understanding of the role of CNVR's to really appreciate how much they influence the different characters. May be they are root of all evil and good, but again they might just be a part of the bigger puzzle.
Sunday, April 18, 2010
Tuesday, April 13, 2010
Monday Morning - Swami and Friends
Our hate for Monday mornings may have something to do with the same things as Swaminathan's feelings of unpleasantness about this day. The joy of having enjoyed the freedom of Saturday and Sunday is overshadowed by the dreary nature of Monday mornings. Apart from the obviously long list of things to complete from the past two days, there is always this feeling of foreboding about Monday which is difficult to overcome.
Anything could happened on a Monday, may be Mr Ebenezar would take upon himself the duty of teaching us idiots the futility of idol worship. Worse still we might feel the insane urge to contradict and question him. Even if we do manage to get through all this ruckus, will we have the sense of not telling anybody at home during the meal about it and getting a stern letter written to the principal the next day?
Everything is not bad as we have few things to look forward to even on Mondays, there is the 12.30 mail to watch from the window. If we manage to solve the 5 arithmetic “puzzles” correctly life would not be that bad.
Anything could happened on a Monday, may be Mr Ebenezar would take upon himself the duty of teaching us idiots the futility of idol worship. Worse still we might feel the insane urge to contradict and question him. Even if we do manage to get through all this ruckus, will we have the sense of not telling anybody at home during the meal about it and getting a stern letter written to the principal the next day?
Everything is not bad as we have few things to look forward to even on Mondays, there is the 12.30 mail to watch from the window. If we manage to solve the 5 arithmetic “puzzles” correctly life would not be that bad.
Hydrogen from solar energy and water?
Industrialization has been largely driven by the continued discovery of oil reserves. However, the number of oil findings is decreasing. A future with no oil left to use is a reality we have to face. Apart from the obvious problem of scarcity the fuels such as oil, coal and gas have been known to contribute to the problem of global warming. The situation is further complicated by the growing need for energy from users who are yet to start using the energy resources.
Many alternative strategies such as solar energy, wind, nuclear, tidal, geothermal etc have been proposed to solve these problems. Although these alternative sources might be able to provide energy, it might not be possible to use them effectively as fuels for transportation systems. Transportation systems being the major consumer of fuels today may need a different approach. Loss of energy during the conversion process has made it necessary to have a direct product which can be used as a fuel.
Use of solar energy to produce fuels such as hydrogen has gained importance in this context. Hydrogen could be directly used as a fuel and lack of carbon in the fuel source makes it a rather clean source of energy. The problem of scarcity and global warming can be tackled with this interesting approach. Two main approaches are being pursued to achieve this goal of using water to produce hydrogen using solar energy. The first approach is the photo biological method which aims to create or alter a biological system to convert solar energy into hydrogen using water as raw material. The second approach is the chemical method, which uses photo systems or molecules that imitate photo systems coupled to other molecules to drive reaction that convert water into hydrogen.
Photosystem II uses solar energy to oxidize water releasing electrons. This reaction is rather efficient although the other steps happening in the biological systems are not as efficient. Hence, the aim is to mimic just this step of the process from nature. The chemical approach has used molecules such as ruthenium linked to the photo systems to act as electron acceptor from Manganese. This is used to drive the reaction to produce hydrogen from water. The enzyme hydrogenase which can catalyse the reaction to produce hydrogen is used in this second step of the reaction.
Biological systems such as Nostoc produce hydrogen in special cells from nitrogenase. Currently large and small scale reactors are being developed to produce hydrogen from such biological systems and make them as effective as possible.
The direct methods of producing fuels have been found to be much more effective than the indirect methods which require the energy to be converted to electricity which is then used to split the water molecules by electrolysis.
Many alternative strategies such as solar energy, wind, nuclear, tidal, geothermal etc have been proposed to solve these problems. Although these alternative sources might be able to provide energy, it might not be possible to use them effectively as fuels for transportation systems. Transportation systems being the major consumer of fuels today may need a different approach. Loss of energy during the conversion process has made it necessary to have a direct product which can be used as a fuel.
Use of solar energy to produce fuels such as hydrogen has gained importance in this context. Hydrogen could be directly used as a fuel and lack of carbon in the fuel source makes it a rather clean source of energy. The problem of scarcity and global warming can be tackled with this interesting approach. Two main approaches are being pursued to achieve this goal of using water to produce hydrogen using solar energy. The first approach is the photo biological method which aims to create or alter a biological system to convert solar energy into hydrogen using water as raw material. The second approach is the chemical method, which uses photo systems or molecules that imitate photo systems coupled to other molecules to drive reaction that convert water into hydrogen.
Photosystem II uses solar energy to oxidize water releasing electrons. This reaction is rather efficient although the other steps happening in the biological systems are not as efficient. Hence, the aim is to mimic just this step of the process from nature. The chemical approach has used molecules such as ruthenium linked to the photo systems to act as electron acceptor from Manganese. This is used to drive the reaction to produce hydrogen from water. The enzyme hydrogenase which can catalyse the reaction to produce hydrogen is used in this second step of the reaction.
Biological systems such as Nostoc produce hydrogen in special cells from nitrogenase. Currently large and small scale reactors are being developed to produce hydrogen from such biological systems and make them as effective as possible.
The direct methods of producing fuels have been found to be much more effective than the indirect methods which require the energy to be converted to electricity which is then used to split the water molecules by electrolysis.
Sunday, April 11, 2010
Brick BAT - iGEM brick biosaftey Assessment tool
With increasing popularity of genetic engineering and synthetic biology we are on the way to malicious biological content. We have seen programs like spyware and malware come out of the IT revolution. What horrors will come out of the advances in biology?
What if crops of entire nations are held to ransom by a pathogenic viral strain? worse still, the human population may be threatened. Organisms that copy our genetic material and take it for analysis without our permission are a distant possibility. Such spyteria (spy + bacteria) are a threat we must get ready to face.
A popular synthetic biology contest called iGEM ( International Genetically Engineered Machines) is encouraging an engineering based approach to synthetic biology. Although they are very serious about biosafety, a categorisation system for the various parts and components was not being used. I have come up with a simple categorisation scheme called Brick BAT. This is a set of questions to classify the components into different levels of biosafety.
What if crops of entire nations are held to ransom by a pathogenic viral strain? worse still, the human population may be threatened. Organisms that copy our genetic material and take it for analysis without our permission are a distant possibility. Such spyteria (spy + bacteria) are a threat we must get ready to face.
A popular synthetic biology contest called iGEM ( International Genetically Engineered Machines) is encouraging an engineering based approach to synthetic biology. Although they are very serious about biosafety, a categorisation system for the various parts and components was not being used. I have come up with a simple categorisation scheme called Brick BAT. This is a set of questions to classify the components into different levels of biosafety.
Tuesday, March 30, 2010
Blast - XML output parser
The below script parses the output from blast with XML as option and stores the values in an array.
#!/usr/bin/perl
#---blast output file
my $infile = "blastout";
open(IN, "$infile");
while((my $line =) && ($hitnumber[0] < $taketill)){ #reading the first $taketill numbers
chomp($line);
if ($line =~ /\/) {#reading hit number
push(@hitnums,split(/\<\/Hit_num\>/,(split(/\/, $line))[1]));
}
if ($line =~ /\/) {#reading definiton
push(@hitdefs,split(/\<\/Hit_def\>/,(split(/\/, $line))[1]));
}
if ($line =~ /\/) {#reading length
push(@hitlens,split(/\<\/Hit_len\>/,(split(/\/, $line))[1]));
}
if ($line =~ /\/) {#reading bitscore
push(@bitscores,split(/\<\/Hsp_bit-score\>/,(split(/\/, $line))[1]));
}
if ($line =~ /\/) {
push(@scores,split(/\<\/Hsp_score\>/,(split(/\/, $line))[1]));
}
if ($line =~ /\/) {#reading evalue
push(@evalues,split(/\<\/Hsp_evalue\>/,(split(/\/, $line))[1]));
}
if ($line =~ /\/) {#reading query match begin
push(@qfroms,split(/\<\/Hsp_query-from\>/,(split(/\/, $line))[1]));
}
if ($line =~ /\/) {#reading query match end
push(@qtos,split(/\<\/Hsp_query-to\>/,(split(/\/, $line))[1]));
}
if ($line =~ /\/) {#reading hit match begin
push(@hfroms,split(/\<\/Hsp_hit-from\>/,(split(/\/, $line))[1]));
}
if ($line =~ /\/) {#reading hit match end
push(@htos,split(/\<\/Hsp_hit-to\>/,(split(/\/, $line))[1]));
}
if ($line =~ /\/) {#reading query frame
push(@qframes,split(/\<\/Hsp_query-frame\>/,(split(/\/, $line))[1]));
}
if ($line =~ /\/) {#reading hit frame
push(@hframes,split(/\<\/Hsp_hit-frame\>/,(split(/\/, $line))[1]));
}
if ($line =~ /\/) {#reading identities
push(@identities,split(/\<\/Hsp_identity\>/,(split(/\/, $line))[1]));
}
if ($line =~ /\/) {#reading positives
push(@positives,split(/\<\/Hsp_positive\>/,(split(/\/, $line))[1]));
}
if ($line =~ /\/) {#reading alignment length
push(@algnlens,split(/\<\/Hsp_align-len\>/,(split(/\/, $line))[1]));
}
}
#!/usr/bin/perl
#---blast output file
my $infile = "blastout";
open(IN, "$infile");
while((my $line =
chomp($line);
if ($line =~ /\
push(@hitnums,split(/\<\/Hit_num\>/,(split(/\
}
if ($line =~ /\
push(@hitdefs,split(/\<\/Hit_def\>/,(split(/\
}
if ($line =~ /\
push(@hitlens,split(/\<\/Hit_len\>/,(split(/\
}
if ($line =~ /\
push(@bitscores,split(/\<\/Hsp_bit-score\>/,(split(/\
}
if ($line =~ /\
push(@scores,split(/\<\/Hsp_score\>/,(split(/\
}
if ($line =~ /\
push(@evalues,split(/\<\/Hsp_evalue\>/,(split(/\
}
if ($line =~ /\
push(@qfroms,split(/\<\/Hsp_query-from\>/,(split(/\
}
if ($line =~ /\
push(@qtos,split(/\<\/Hsp_query-to\>/,(split(/\
}
if ($line =~ /\
push(@hfroms,split(/\<\/Hsp_hit-from\>/,(split(/\
}
if ($line =~ /\
push(@htos,split(/\<\/Hsp_hit-to\>/,(split(/\
}
if ($line =~ /\
push(@qframes,split(/\<\/Hsp_query-frame\>/,(split(/\
}
if ($line =~ /\
push(@hframes,split(/\<\/Hsp_hit-frame\>/,(split(/\
}
if ($line =~ /\
push(@identities,split(/\<\/Hsp_identity\>/,(split(/\
}
if ($line =~ /\
push(@positives,split(/\<\/Hsp_positive\>/,(split(/\
}
if ($line =~ /\
push(@algnlens,split(/\<\/Hsp_align-len\>/,(split(/\
}
}
Friday, March 19, 2010
Strand specific translation of DNA to aminoacid sequence in PERL
PERL script to read the annotation file and update translated ORF's it into the mysql database and writing it into a multifasta file.
use DBI;
$dbh = DBI->connect('DBI:mysql:meta', 'root', 'password'
) || die "Could not connect to database: $DBI::errstr";
my ($infile) = @ARGV;
open(IN, "$infile");
my $outfile ="multifasta";
open(OUT,">>$outfile");
my $jcvi_read,$strand,$start,$stop,@seq,$ofset,$leng,$wrtseq;
while(my $line =){
chomp($line);
my @a = split(/ /, $line);
my @c = split(/\s+/, $line);
my @b = split(/\_/, $a[0]);
$jcvi_read=$b[2];
$start = $c[3];
$stop = $c[4];
$strand = $c[6];
$pep = $c[8];
$sth = $dbh->prepare("SELECT sequence FROM metadata WHERE jcvi_read='$jcvi_read'");
$sth->execute();
@seq = $sth->fetchrow_array();
$ofset=$start - 1;
$leng=$stop-$start;
$wrtseq=substr $seq[0], $ofset, $leng;
$sth->finish();
#function to translate DNA to Amino acid based on standard genetic code
sub codon2aa{
my($codon)=@_;
$codon= uc $codon;
my(%genetic_code) = (
'TCA'=>'S', #Serine
'TCC'=>'S', #Serine
'TCG'=>'S', #Serine
'TCT'=>'S', #Serine
'TTC'=>'F', #Phenylalanine
'TTT'=>'F', #Phenylalanine
'TTA'=>'L', #Leucine
'TTG'=>'L', #Leucine
'TAC'=>'Y', #Tyrosine
'TAT'=>'Y', #Tyrosine
'TAA'=>'_', #Stop
'TAG'=>'_', #Stop
'TGC'=>'C', #Cysteine
'TGT'=>'C', #Cysteine
'TGA'=>'_', #Stop
'TGG'=>'W', #Tryptophan
'CTA'=>'L', #Leucine
'CTC'=>'L', #Leucine
'CTG'=>'L', #Leucine
'CTT'=>'L', #Leucine
'CCA'=>'P', #Proline
'CAT'=>'H', #Histidine
'CAA'=>'Q', #Glutamine
'CAG'=>'Q', #Glutamine
'CGA'=>'R', #Arginine
'CGC'=>'R', #Arginine
'CGG'=>'R', #Arginine
'CGT'=>'R', #Arginine
'ATA'=>'I', #Isoleucine
'ATC'=>'I', #Isoleucine
'ATT'=>'I', #Isoleucine
'ATG'=>'M', #Methionine
'ACA'=>'T', #Threonine
'ACC'=>'T', #Threonine
'ACG'=>'T', #Threonine
'ACT'=>'T', #Threonine
'AAC'=>'N', #Asparagine
'AAT'=>'N', #Asparagine
'AAA'=>'K', #Lysine
'AAG'=>'K', #Lysine
'AGC'=>'S', #Serine#Valine
'AGT'=>'S', #Serine
'AGA'=>'R', #Arginine
'AGG'=>'R', #Arginine
'CCC'=>'P', #Proline
'CCG'=>'P', #Proline
'CCT'=>'P', #Proline
'CAC'=>'H', #Histidine
'GTA'=>'V', #Valine
'GTC'=>'V', #Valine
'GTG'=>'V', #Valine
'GTT'=>'V', #Valine
'GCA'=>'A', #Alanine
'GCC'=>'A', #Alanine
'GCG'=>'A', #Alanine
'GCT'=>'A', #Alanine
'GAC'=>'D', #Aspartic Acid
'GAT'=>'D', #Aspartic Acid
'GAA'=>'E', #Glutamic Acid
'GAG'=>'E', #Glutamic Acid
'GGA'=>'G', #Glycine
'GGC'=>'G', #Glycine
'GGG'=>'G', #Glycine
'GGT'=>'G', #Glycine
);
if(exists $genetic_code{$codon}){
return $genetic_code{$codon};
}
else{
return 'X';
}
}
if($strand == '-')
#reverse complementing -ve strands
{$wrtseq=reverse $wrtseq;
$wrtseq=~ tr/ATGC/TACG/;}
my $dna=$wrtseq;
#my $dna="TCATTCTCATTC";
my $protein='';
my $codon3;
for(my $i=0; $i<(length($dna)-2); $i+=3){ $codon3=substr($dna,$i,3); $protein.= codon2aa($codon3); } print $pep."\n"; $dbh->do("INSERT into iddata (jcvi_read,strand,start,stop,protein,orf_id) VALUES('$jcvi_read','$strand','$start','$stop','$protein','$pep')");
print OUT ">".$jcvi_read."|".$pep."\n";
print OUT $protein."\n";
}
$dbh->disconnect();
use DBI;
$dbh = DBI->connect('DBI:mysql:meta', 'root', 'password'
) || die "Could not connect to database: $DBI::errstr";
my ($infile) = @ARGV;
open(IN, "$infile");
my $outfile ="multifasta";
open(OUT,">>$outfile");
my $jcvi_read,$strand,$start,$stop,@seq,$ofset,$leng,$wrtseq;
while(my $line =
chomp($line);
my @a = split(/ /, $line);
my @c = split(/\s+/, $line);
my @b = split(/\_/, $a[0]);
$jcvi_read=$b[2];
$start = $c[3];
$stop = $c[4];
$strand = $c[6];
$pep = $c[8];
$sth = $dbh->prepare("SELECT sequence FROM metadata WHERE jcvi_read='$jcvi_read'");
$sth->execute();
@seq = $sth->fetchrow_array();
$ofset=$start - 1;
$leng=$stop-$start;
$wrtseq=substr $seq[0], $ofset, $leng;
$sth->finish();
#function to translate DNA to Amino acid based on standard genetic code
sub codon2aa{
my($codon)=@_;
$codon= uc $codon;
my(%genetic_code) = (
'TCA'=>'S', #Serine
'TCC'=>'S', #Serine
'TCG'=>'S', #Serine
'TCT'=>'S', #Serine
'TTC'=>'F', #Phenylalanine
'TTT'=>'F', #Phenylalanine
'TTA'=>'L', #Leucine
'TTG'=>'L', #Leucine
'TAC'=>'Y', #Tyrosine
'TAT'=>'Y', #Tyrosine
'TAA'=>'_', #Stop
'TAG'=>'_', #Stop
'TGC'=>'C', #Cysteine
'TGT'=>'C', #Cysteine
'TGA'=>'_', #Stop
'TGG'=>'W', #Tryptophan
'CTA'=>'L', #Leucine
'CTC'=>'L', #Leucine
'CTG'=>'L', #Leucine
'CTT'=>'L', #Leucine
'CCA'=>'P', #Proline
'CAT'=>'H', #Histidine
'CAA'=>'Q', #Glutamine
'CAG'=>'Q', #Glutamine
'CGA'=>'R', #Arginine
'CGC'=>'R', #Arginine
'CGG'=>'R', #Arginine
'CGT'=>'R', #Arginine
'ATA'=>'I', #Isoleucine
'ATC'=>'I', #Isoleucine
'ATT'=>'I', #Isoleucine
'ATG'=>'M', #Methionine
'ACA'=>'T', #Threonine
'ACC'=>'T', #Threonine
'ACG'=>'T', #Threonine
'ACT'=>'T', #Threonine
'AAC'=>'N', #Asparagine
'AAT'=>'N', #Asparagine
'AAA'=>'K', #Lysine
'AAG'=>'K', #Lysine
'AGC'=>'S', #Serine#Valine
'AGT'=>'S', #Serine
'AGA'=>'R', #Arginine
'AGG'=>'R', #Arginine
'CCC'=>'P', #Proline
'CCG'=>'P', #Proline
'CCT'=>'P', #Proline
'CAC'=>'H', #Histidine
'GTA'=>'V', #Valine
'GTC'=>'V', #Valine
'GTG'=>'V', #Valine
'GTT'=>'V', #Valine
'GCA'=>'A', #Alanine
'GCC'=>'A', #Alanine
'GCG'=>'A', #Alanine
'GCT'=>'A', #Alanine
'GAC'=>'D', #Aspartic Acid
'GAT'=>'D', #Aspartic Acid
'GAA'=>'E', #Glutamic Acid
'GAG'=>'E', #Glutamic Acid
'GGA'=>'G', #Glycine
'GGC'=>'G', #Glycine
'GGG'=>'G', #Glycine
'GGT'=>'G', #Glycine
);
if(exists $genetic_code{$codon}){
return $genetic_code{$codon};
}
else{
return 'X';
}
}
if($strand == '-')
#reverse complementing -ve strands
{$wrtseq=reverse $wrtseq;
$wrtseq=~ tr/ATGC/TACG/;}
my $dna=$wrtseq;
#my $dna="TCATTCTCATTC";
my $protein='';
my $codon3;
for(my $i=0; $i<(length($dna)-2); $i+=3){ $codon3=substr($dna,$i,3); $protein.= codon2aa($codon3); } print $pep."\n"; $dbh->do("INSERT into iddata (jcvi_read,strand,start,stop,protein,orf_id) VALUES('$jcvi_read','$strand','$start','$stop','$protein','$pep')");
print OUT ">".$jcvi_read."|".$pep."\n";
print OUT $protein."\n";
}
$dbh->disconnect();
Thursday, March 4, 2010
Updating mySql database from Perl
Perl script to read multifasta file calculate GC content,GC Skew and insert into mySql database.
#!/usr/bin/perl
use DBI;
$dbh = DBI->connect('DBI:mysql:meta', 'root', 'password'
) || die "Could not connect to database: $DBI::errstr";
my ($infile) = @ARGV;
open(IN, "$infile");
my $jcvi_read,$leng,$gccont=0,$gcskew=0,$seq="";
while(my $line =){
chomp($line);
if ($line =~ /^\>/) {#reading header line
if($jcvi_read !=0){
$gccont=($gccont/$leng)*100;
$dbh->do("INSERT into metadata (jcvi_read,full_length,gc_content,gc_skew_orient,sequence) VALUES('$jcvi_read','$leng','$gccont','$gcskew','$seq')");#inserting the data into mySql Database
}
$gccont=0;
$seq="";
my @a = split(/\//, $line);
my @b = split(/\_/, $a[0]);
$jcvi_read = $b[2];
my @c = split(/\=/, $a[1]);
$leng = $c[1];
# print $jcvi_read."\n";
}
else{#reading and analysing the sequence
$seq = $seq.$line;
my @char =split(//,$line);
foreach(@char){
if($_ eq "G"|| $_ eq "C"){$gccont++;}# Calculating GC content
if($_ eq "G"){$gcskew++;}#Calcualting GC skew
if($_ eq "C"){$gcskew--;}
}
}
#print $gcskew."\n";
}
$dbh->disconnect();
#!/usr/bin/perl
use DBI;
$dbh = DBI->connect('DBI:mysql:meta', 'root', 'password'
) || die "Could not connect to database: $DBI::errstr";
my ($infile) = @ARGV;
open(IN, "$infile");
my $jcvi_read,$leng,$gccont=0,$gcskew=0,$seq="";
while(my $line =
chomp($line);
if ($line =~ /^\>/) {#reading header line
if($jcvi_read !=0){
$gccont=($gccont/$leng)*100;
$dbh->do("INSERT into metadata (jcvi_read,full_length,gc_content,gc_skew_orient,sequence) VALUES('$jcvi_read','$leng','$gccont','$gcskew','$seq')");#inserting the data into mySql Database
}
$gccont=0;
$seq="";
my @a = split(/\//, $line);
my @b = split(/\_/, $a[0]);
$jcvi_read = $b[2];
my @c = split(/\=/, $a[1]);
$leng = $c[1];
# print $jcvi_read."\n";
}
else{#reading and analysing the sequence
$seq = $seq.$line;
my @char =split(//,$line);
foreach(@char){
if($_ eq "G"|| $_ eq "C"){$gccont++;}# Calculating GC content
if($_ eq "G"){$gcskew++;}#Calcualting GC skew
if($_ eq "C"){$gcskew--;}
}
}
#print $gcskew."\n";
}
$dbh->disconnect();
Subscribe to:
Posts (Atom)