Sometimes the mountain goes to Mohammed. This Tuesday, the 29th of May, scientists and representatives of BGI will present their services and scientific projects in a workshop entitled "Impact of breaking genomic techniques into disease diagnosis and mechanisms". The workshop will take place in the Aula Consigliare of the Faculty of Medicine, University of Brescia starting from 9.00. The meeting has been organized by the Department of Biomedical Sciences and Biotechnology. A more detailed program of the event is available at this link.
A blog with news and curiosity on genomics subjects with a particular interest for topics related to Next Generation Sequencing, Personal Genomics and Bioinformatics. We work at the University of Brescia (Italy) and are new in the field but with a lot of energy to share.
Sunday, 27 May 2012
BGI Workshop at the University of Brescia
Sometimes the mountain goes to Mohammed. This Tuesday, the 29th of May, scientists and representatives of BGI will present their services and scientific projects in a workshop entitled "Impact of breaking genomic techniques into disease diagnosis and mechanisms". The workshop will take place in the Aula Consigliare of the Faculty of Medicine, University of Brescia starting from 9.00. The meeting has been organized by the Department of Biomedical Sciences and Biotechnology. A more detailed program of the event is available at this link.
Monday, 14 May 2012
CGHub at UCSC: Cancer data at Petabyte scale
Recently there were several announcements about genomic data made available on the net, supporting the idea of a new era in genomics when data production greatly exceed data analysis capacity. So the solution is to make all these data avaiable to the community and let other research groups to continue the analysis. The recent CGHub initiative from UCSC and NCI goes exactly in this direction and stands out for the dimension of the dataset.
Indeed, University of California Santa Cruz, that everybody knows for its famous Genome Browser, is now building a petabyte-scale data repository for cancer genome projects, called CGHub (Cancer Genomics Hub).
The project is funded by the National Cancer Institute with over $10M and it started with 5 petabytes avaiable and the possibility to growth as needed by the three main NCI cancer genome projects: the Cancer Genome Atlas (TCGA); the Therapeutically Applicable Research to Generate Effective Treatments (TARGET) and the Cancer Genome Characterization Initiative (CGCI).
Therefore, CGHub will soon become the largest available dataset on cancer genomics ranging from adult cancer to childhood cancer to HIV-related tumors.
Currently, TCGA alone generates about 10 terabytes of data each month and its output is expected to increase rapidly getting to a total amount of 10 petabytes of DNA and RNA data from the 10,000 patients involved.
Besides providing the space, UCSC is also working on innovative software solutions to grant rapid and flexible access to this huge amount of data, as well as to compress them and simplfy the management of a petabyte scale storage.
"We would like to compress the data down to one tenth of its current size and that will not be possible without losing some information," told David Haussler, a professor of biomolecular engineering at UCSC in charge for the CGHub project. At present, the cancer genomics community is "working very hard to decide what information we can sacrifice in these very valuable data."
More information can be found on this post from GenomeWeb Bioinform and on the offical pages of TCGA, TARGET and CGCI.
Indeed, University of California Santa Cruz, that everybody knows for its famous Genome Browser, is now building a petabyte-scale data repository for cancer genome projects, called CGHub (Cancer Genomics Hub).
The project is funded by the National Cancer Institute with over $10M and it started with 5 petabytes avaiable and the possibility to growth as needed by the three main NCI cancer genome projects: the Cancer Genome Atlas (TCGA); the Therapeutically Applicable Research to Generate Effective Treatments (TARGET) and the Cancer Genome Characterization Initiative (CGCI).
Therefore, CGHub will soon become the largest available dataset on cancer genomics ranging from adult cancer to childhood cancer to HIV-related tumors.
Currently, TCGA alone generates about 10 terabytes of data each month and its output is expected to increase rapidly getting to a total amount of 10 petabytes of DNA and RNA data from the 10,000 patients involved.
Besides providing the space, UCSC is also working on innovative software solutions to grant rapid and flexible access to this huge amount of data, as well as to compress them and simplfy the management of a petabyte scale storage.
"We would like to compress the data down to one tenth of its current size and that will not be possible without losing some information," told David Haussler, a professor of biomolecular engineering at UCSC in charge for the CGHub project. At present, the cancer genomics community is "working very hard to decide what information we can sacrifice in these very valuable data."
More information can be found on this post from GenomeWeb Bioinform and on the offical pages of TCGA, TARGET and CGCI.
Friday, 11 May 2012
Flash Report: Life Technologies official reply to Loman's paper

Life Technologies has posted a reply to the "Performance comparison of benchtop high-throughput sequencing platforms" manuscript published in April on Nature Biotechnology.
Here is what they did:
How did we address some of the challenges?
• We obtained a sample of the EHEC E. coli isolate, the week after the paper was released
• We sequenced this sample using up-to-date Ion sequencing kits and software
• We directly compared the data from Loman et al July 2011 to Ion publicly released data from May 2012*
• We aligned the reads using the genome sequenced in Loman et al to measure accuracy
• We performed de novo assembly to measure how well the PGM performs for this application
Using the most recent chemistries, algorithms and the 318 Chip, the Ion Torrent PGM error rates are definitely lower (the majority of bases are Q30 or better), although in this area the MiSeq seems to have an edge on the competition, particularly in the homopolymer stretches.
It is however remarkable that in 9 months (July 2011 - April 2012) the indel and homopolymer accuracy has improved ~450% on the Ion Torrent platform.
Thursday, 10 May 2012
Genome vs Enviroment: the twin match.
"What fraction of the population would benefit from genome sequencing? “Benefit” in this context is defined as receiving information indicating that the risk of disease is increased or decreased to a degree that would alter an individual's lifestyle or medical management." This is the starting question that Roberts et al. wanted to answer to. Right now this seems an unsuccessfully effort, an impossible challange given the millions of variants that could contribute to the risk of a disease. However there is a field in which this challenge becomes easier: the genome of twins. Infact, as the authors wrote in their article published on Science Translational Medicine, in a pair of twins in which at least one is affected by a disease, the probability of the other twin developing a disease is dependent on the genome whenever that disease has some genetic component. The study was conducted on 53,666 twin pairs, analyzing 24 disease. The researchers, given the affected status of one of the two twins, used the prevalence of the disease in the second twin to calculate the genetic risk of a given disease for that genome. The results, elaborated without any type of sequencing effort, reveal that the majority of the twins would have tested negative for risk for 23 of the 24 diseases, substantially their risk was below the risk of the general population. This article received many critical comments, here you can find a series of them posted on the Nature blog, the main critics focused on the low novelty of the study. It is well known that twins could get sick of different diseases and also die for different causes that don't depend on their common genome.
What emerges from this work is that our genome, and its whole sequencing, is valuable and instrumental for the identification of the basis of many genetic disease, primarily for those with a family history. On the other hand, enviroment has the main role in many medical conditions, especially in complex diseases such as cancer. This is not a bad news because, indipendently of our genetic starting conditions, the match against complex diseases is open, and weapons such as diet, sports, and lifestyle, could make us the winners.
Wednesday, 25 April 2012
MiSeq vs PGM Ion Torrent vs 454 GS Junior. And the winner is...
By the way, this is the 100th post on this blog!
First Ion Proton systems installed and already operational at the Baylor College of Medicine
Life Technologies Corporation announced yesterday in a press release that multiple Ion Proton Sequencers have been installed (in early April) and are now operating at the Baylor College of Medicine Human Genome Sequencing Center, (BCM-HGSC) in Houston, Texas.
Dr. Richard Gibbs, the Director of the BCM-HGSC said: "We're pleased the Proton installation went so quickly and smoothly. We are now generating all the raw data needed for full exomes in just a few hours, and it's exhilarating to see what's to come."
Need a way to find useful data in the increasing stack of NGS studies? NGS catalog is a good starting point
With NGS technologies rapidly spreading and costs rapidly falling, we are now facing an amazing acceleration in the publication of new studies taking advantage of the increadible power of DNA sequencing. They range from ChIp-seq to whole genome resequencing, with exome sequencing and RNA-seq as the more widely used applications. All these studies produce a huge amount of data that are often publicly available to the scientific community.
But too much data could be as bad as no data, if researchers are unable to find the useful information within the huge amount of ACGT present in databases. The idea of a searchable database on NGS-relates studies is a really valuable one and such a resource could save you a lot of time. The idea just became a reality thanks to the NGS catalog web portal. This resource has been created by researchers at Vanderbilt University Medical Center, Nashville, TN and is described in the latest issue of Human Mutation. Give it a try!
Wednesday, 4 April 2012
BioBank is finally open: another gold mine of publicly available data
Yes it's true! After years of delay the BioBank is finally ready and open for public access.

As reported also on ScienceInsider, the ambitious project to build a repository for biological data from about 500,000 individuals is now completed. The £62 million project, funded mainly by the Medical Research Council and the Wellcome Trust, has tested one in every 50 people between the ages of 40 and 69 years, from across the UK country, between 2006 and 2010. The collection of samples and data will continue. Participants will be followed until they die, their details updated from National Health Service records, and some tests repeated. Each year brings more detail about each person and more cases of disease.
Using the web portal every researcher can now rely on this gold mine of biological data. Besides biological samples, a researcher can have access to a wide range of information on volunteers, including their height, weight, lung function, and blood pressure, as well as medical history and lifestyle data.
Tuesday, 3 April 2012
American College of Medical Genetics Meeting 2012

The 2012 annual ACMG meeting has just finished (March 27-31), and the topics that deserve attentions are many. One of the main point on which the meeting focused was the use of the next generation techniques in clinical and diagnostic field, as demonstrated by the high number of talks and posters related to this topic. Among them, one of the most interesting is the update given by the Undiagnosed Diseases Program director, Prof. Boerkoel, who presented an "extreme novel filtering" method (here is the link to the paper) to find disease-associated mutations behind conditions represented by just one or a few patients.
Another relevant event of the meeting was the release by the ACMG Board of Directors of the "Policy Statement" regarding Whole Genome and Exome Sequencing. This statement is divided in four section: Definitions, Indication for diagnostic testing, Clinical testing and result reporting, and Genetic screening. This is one of the first attempts to define and elucidate for clinicians and genetists what is next generation sequencing, how and when it can be used as a diagnostic tool, and how they should face and report the results.
Monday, 2 April 2012
Flash Report: the Amaz..ing 1000 Genomes in the cloud
Amazon and the U.S. National Institutes of Health (NIH) announced that the complete 1000 Genomes Project is being made available on Amazon Web Services (AWS) as free of charge public data set. The project has grown to 200 terabytes of genomic data (!!!) including DNA sequenced from more than 1,700 individuals. The 1000 Genomes Project aims to include the genomes of more than 2,662 individuals from 26 populations around the world, and the NIH will continue to add the remaining genome samples to the data collection this year.
The move to put the data up on AmazonWeb Services, aims to help speed up access to the research. Previously, researchers had to download data from government data centers (either NCBI or EMBL-EBI) on their own systems.
On the AWS site you can find more info on how to access the 1000 Genomes data.
Subscribe to:
Posts (Atom)





