The Principle

The original stays in the vault

Picture DNA as the archive of a library: it holds all the important original documents, but they may never be checked out. Anyone who needs to know something gets a copy — one they can take with them and use until it wears out.

That is exactly how the cell works. DNA sits in the nucleus and is never read directly for protein production. Instead, a temporary working copy of each gene that is needed gets made: messenger RNA, or mRNA for short. The enzyme RNA polymerase reads the DNA and synthesizes a complementary RNA strand as it goes — nucleotide by nucleotide. This process is called transcription.

RNA polymerase unwinds the double helix step by step: one strand serves as the template, complementary ribonucleotides attach — and are immediately linked into the growing mRNA chain. When the polymerase reaches the end of the gene, the finished copy is released and leaves the nucleus through specialized nuclear pores. In the cytoplasm, a ribosome clamps onto the mRNA and reads the codons — triplets of the bases A, U, G, C — one after another. For each codon, a transfer RNA (tRNA) delivers the matching amino acid and hands it off to the growing chain. Once the stop codon is reached, the chain drops off the ribosome — and immediately begins to fold. What happens next is pure chemistry: the sequence of amino acids determines which regions attract each other and which repel. The chain inevitably finds the one stable three-dimensional shape that this sequence can adopt — the finished protein.

The mRNA leaves the nucleus, travels into the cytoplasm — and is read there by a ribosome. The ribosome translates the letter sequence of the mRNA into a chain of amino acids. This second step is called translation. In the end, the chain folds into a finished protein.

NUCLEUS DNA Transcription mRNA CYTOPLASM Export Amino acids Translation Protein Function DNA → mRNA → Protein · Central dogma of molecular biology

Why the detour? Separating DNA and mRNA has a decisive advantage: DNA stays protected. If it served directly as the template for proteins, it would be exposed to mechanical stress every single time it was read. Instead, an mRNA can be read thousands of times and then simply broken down — the DNA remains untouched.

The Code

Three letters make one amino acid

mRNA is a strand made of four letters: A, U, G and C — the RNA bases. The ribosome reads this strand in groups of three: each triplet, called a codon, corresponds to exactly one amino acid. With four letters and three positions there are 4³ = 64 possible codons — but only 20 different amino acids. So the code is redundant: several codons can encode the same amino acid.

At first that sounds wasteful. But it is actually a protective mechanism: if a single nucleotide changes through mutation, the ribosome often still ends up with the correct building block — because the new codon still stands for the same amino acid. Such changes are called silent mutations. The protein gets built as if nothing had happened.

Nonpolar
Polar
Acidic
Basic
Aromatic
STOP

Inside → outside: 1st base · 2nd base · 3rd base · amino acid. mRNA uses U instead of T.

What makes this code so fascinating: it is nearly identical across all known living things. From gut bacteria to humans, from yeast to blue whales — the same codon codes for the same amino acid. This is one of the strongest molecular pieces of evidence that all life on this planet traces back to a single common origin. (A few exceptions exist, though — for instance in mitochondrial DNA or in certain ciliates.)

AUGMethionine (Start)
UUUPhenylalanine
GAAGlutamic acid
GGUGlycine
CCUProline
ACGThreonine
UGGTryptophan
UAAStop
The Surprise

20,000 genes — over 100,000 proteins

When the human genome was sequenced to over 99% completion in 2003 (the last gaps weren't closed until the Telomere-to-Telomere Consortium's 2022 assembly), researchers expected around 100,000 genes. The actual number was sobering: about 20,000 — barely more than a roundworm. And yet the human body produces well over 100,000 different proteins. How does that add up?

The answer lies in a process called alternative splicing. Before mRNA leaves the nucleus, it is cut apart and reassembled. Genes consist of coding segments (exons) and non-coding stretches in between (introns). The introns are cut out — but which exons are then joined together varies depending on cell type, developmental stage, or external signal.

The same genetic building blocks can therefore give rise to entirely different proteins. A gene is not a rigid recipe but more like a modular kit: the cell chooses which parts it currently needs. This makes gene regulation a fascinating subject in its own right — and explains why a liver cell and a nerve cell can be so different even though both carry exactly the same DNA.

The origin of life, written in the code: If the same genetic code works in bacteria and in human cells — and if a human gene can be inserted into a yeast cell and be read correctly there — then all living things today must once have shared a common ancestor. The genetic code is the oldest molecular fossil there is.

The Exception

When information flows backward

The central dogma says: information flows from DNA to RNA to protein — never the other way. But nature has exceptions. Retroviruses like HIV carry their genetic material as RNA and bring along an enzyme called reverse transcriptase, which rewrites the RNA back into DNA. This DNA is then inserted into the host cell's chromosome — permanently, with no way to remove it.

This is not just medically relevant: a substantial portion of the human genome — an estimated 8 percent — comes from ancient retroviruses that embedded themselves in our ancestors millions of years ago. We carry the traces of past infections in our genome. Some of it is silent baggage. Other parts have been repurposed by evolution and now serve real functions.

→ Back to topic overview DNA Visualization Gene Regulation