Exabyte Scale of Genomics Data (and cat videos)

DNA, Image Source: http://www.publicdomainpictures.net/view-image.php?image=42718&picture=dna
DNA, Image Source

Its no surprise that genomics represents a terrific big data challenge, but noting that its data has doubled every seven months over the last ten years is remarkable given how the field is poised to really explode in the coming years.

This article points out the comparison with astronomy and social media:

The authors estimate that the genomics information so far, from sequencing different organisms and a number of humans, has produced data on the petabyte scale (a petabyte is a million gigabytes). However, over the last decade, genomic sequencing data doubled about every seven months, and will grow at an even faster rate as personal genome sequencing becomes more widespread. The researchers estimate that by 2025, genomics data will explode to the exabyte scale – billions of gigabytes. This surpasses even YouTube, the current title holder among the domains studied for most data stored.

Frankly, it is refreshing to see such a valuable area of study surpassing a repository of countless cat videos as a leading data management problem in our society.


