Showing posts with label proteomics. Show all posts
Showing posts with label proteomics. Show all posts

Thursday, June 14, 2007

What did ENCODE decode?


As recently as five days ago, I penned a post about what the Human Genome Project (HGP) had and had not accomplished. I wish I could say that I had written that with the full knowledge that it would be a great primer for a piece about the genome discoveries released today by ENCODE , the NIH follow-up effort to the HGP. I would be lying if I did.

Anyways, front and center at Nature.com is a pdf of the publication by the Encode Consortium outlining the highlights of their efforts to pass a fine toothed comb through approximately 1% of the human genome.

The publication is fascinating in both its breadth and detail. Before I expound on its virtues, let me first comment on my only suspicion about the project. From my own somewhat limited experience in biomedical research, I am not a big fan of large consortium efforts. While I love the concept of open source sharing of data and collaboration, I have usually found that huge efforts across many labs breed data inconsistencies as a result of methodological and analytical differences. Differences in variables as small as humidity in the lab can yield differences in datasets that can obscure the real story. All of that said, it would be very hard to argue with the key points that are coming out of this publication, because the key points make a lot more sense than the conventional wisdom that has been coming out of college biology text books for years (at least when I was in college).

Most of us have been taught at some point that DNA leads to RNA which leads to protein. Well, all of that is still true, but as time goes on, we continue to discover that there are more and more options for the RNA besides producing protein. Without further ado, here are the take home notes on the ENCODE project:

  • While it was once thought that a large proportion of DNA was "junk" which did nothing, it is becoming clearer that the vast marority of DNA does transcribe RNA. Many new non-protein coding RNA's have been discovered in the ENCODE effort.
  • Chromatin accessibility, basically how tightly the DNA is wound, has a huge effect on how readily it is transcribed to RNA. In turn, many RNA's can affect how tightly the DNA is wound.
  • We have evolved in a way that has rendered about 5% of our DNA inactive.
  • Some regions of our DNA are wildly variable from person to person, while other regions barely change (this isn't really news, but they've been able to pinpoint some of the specific variable regions).
  • RNA can do many things beside encode for protein. Some RNA's are used by the cell to suppress other RNA's...thereby regulating the genome. (this isn't really news either).
  • There is way too much RNA in cells for us to know what all of it does at this point in time
I do hope you take a look at the pdf file for the original article that I linked to up above. Science journals are tedious to read, especially if it is a new world for you, but it is worth tackling every now and then. There are usually pretty pictures.

Saturday, June 9, 2007

HUMAN GENOME PROJECT- Where is it now?

The human genome is the genome of Homo sapiens, which is composed of 24 distinct chromosomes (22 autosomal + X + Y) with a total of approximately 3 billion DNA base pairs. -Wikipedia



Let me start by acknowledging that the international public Human Genome Project (HGP) and the private Celera Genomics enterprise achieved a monumental task in sequencing the human genome in 2001. News media outlets trumpeted the accomplishment with verbal steamers and confetti, calling it, among other superlatives, “the Rosetta Stone of Life.” Press releases and magazine features foretold of new disease diagnostics, treatments, and even cures which would emerge from having “decoded” the blueprint for human life while warning about the ethical fallout that could emerge from DNA manipulation.


What They Forgot to Mention

The media outlets were mostly right. They, along with the scientists and the public relations professionals involved with the projects, told us that the sequencing of the human genome was the first step toward understanding and treating many human diseases. What they didn’t tell us was how many additional steps would still need climbing before we would reap those rewards. It seems that what the project yielded was far less “Rosetta Stone” and far more akin to the discovery of a tomb filled with unintelligible Egyptian hieroglyphs.

Many articles describing the HGP erroneously use the words “sequenced” and “decoded” interchangeably. In fact, Wikipedia describes the HGP as “a project to decode (i.e. sequence) more than three billion nucleotides contained in a haploid reference human genome and to identify all the genes presented in it.” The fact is, “decode” and “sequence” are quite different.


What Did the Human Genome Project Accomplish?

If by sequencing the human genome, the HGP and Celera didn’t actually decode it, then what exactly did they accomplish? To explain, we must first take a cursory look at the ingredients that make up the genome, deoxyribonucleic acid (DNA). DNA basically is made of chains of molecules called nucleotides. There are four different nucleotides which make up DNA. These are cytosine (C), thymine (T), adenine (A), and guanine (G). The HGP and Celera looked at the complete genome of one man and cataloged each of his 3 billion plus nucleotides. Basically, the hullabaloo in 2001 was simply the media fanfare that accompanied the inking of the correct order of a single human’s C’s, T’s, G’s, and A’s.

How will this correctly ordered catalog of letters eventually result in disease diagnostics and treatments? It will be a long and winding path along which the science community has only taken a few short steps. After having laid out the map of one man’s genome, researchers are now laying out the genomic maps of many others. In order to learn what each of the 30,000 genes does, researchers must compare the genomes of many humans, determine which nucleotides are ordered differently and how those differences in nucleotides translate into differences in how we each look, behave, grow, age, develop diseases, fight off diseases, and so forth. Objectively speaking, the cataloging of that first genome is no more valuable than any of the genomic catalogs which have followed; or will follow. The mapping of the first genome was a huge milestone, but any one of us could currently have our entire genomes mapped in the exact same way for a cost of approximately $200,000. The information obtained from your genome would have the same research value as the information gleaned from the entire multibillion dollar international effort of the HGP and Celera less than a decade ago.


Where will the Human Genome Project take us?

Research projects, like the international HapMap Project, are currently cataloging and comparing the genomes of humans across relatively tight clusters of human populations. We are learning what the predominant differences in the genetic codes of four distinct populations of African, Asian, and European ancestry. With each of the populations representing variable risk factors for specific diseases and conditions, these comparisons could provide indications of which genes can lead to diseases or protect against diseases (this effort is a minefield of unresolved ethics issues…a topic for another post).

Efforts to take advantage of the ability to map each of our genomes are further confounded by the fact that our genetic codes represent only a small portion of the complexities of our molecular biological systems. The HGP tells us that we have approximately 30,000 genes. Suppositions before the Genome project were that each single gene encodes a single protein . We are left now to wonder how humans are built of approximately 120,000 proteins. To explain this, it must be the case that some genes can be toggled to create more than one gene product. The toggling must be controlled by other genes which might also be “toggle-able”.

The manipulation of genetic circuitry required to treat diseases is almost infinitely complex and requires that major strides are made by researchers who are cataloging the structure and functions of the products of the genome; proteins. This is the study of proteomics (yet another topic for another post).

In the end, like the first moon landing in 1969, the accomplishments of the Human Genome Project represent a huge milestone for humanity. Also, like the first moon landing which only represented a tiny step out into the unfathomable depths of our galaxy, the mapping of the human genome is only a tiny step into the unfathomable complexities of biology and life itself.