
The living cell · 64 min · 14,165 words
From gene to protein: how a cell actually reads itself
DNA is transcribed, spliced, exported, translated and folded. A research peptide is the ligand at the end of that pipeline, already made. This essay is every machine in between — polymerases, spliceosome, ribosome — with the actual rates attached.
What this essay actually tells you
- DNA does not make protein. Pol II elongates at ~20–40 nucleotides/s. A 20 kb gene is minutes. Dystrophin (2.3 Mb) is hours. The spliceosome is ~3 MDa and edits the RNA while it is still being written.
- Eukaryotic ribosomes add ~5–6 amino acids a second. Translation error ~10⁻⁴; replication after mismatch repair ~10⁻⁹–10⁻¹⁰. The genome is sacred. The proteome is a draft.
- A lyophilised research peptide skips the entire factory. Merrifield solid-phase synthesis builds the chain on a resin. Same amide bond. Different building. Epithalon’s literature sits on TERT transcription — a promoter, not a ribosome.
What this actually means
Francis Crick's central dogma is still the spine: DNA to RNA to protein. The part a century of biochemistry added is the machinery and the clocks. Finding a gene in 3.1 billion base pairs. Opening chromatin. Assembling a pre-initiation complex. Transcribing at 20-40 nucleotides a second. Splicing out introns on a 3-megadalton spliceosome. Exporting through a nuclear pore. Translating at 5-6 amino acids a second on a 4-megadalton ribosome. Folding, modifying, shipping. A lyophilised peptide in the catalogue skips all of that. It is the finished ligand. That is why it can occupy a receptor in a dish, and why it is not a gene, a therapy, or a treatment plan. The vial is Merrifield chemistry or a recombinant tank. The cell is a search problem with a clock measured in minutes to hours, not milliseconds. Occupancy is the catalogue. Gene therapy and CRISPR are other floors. Research use only is the label on every one of them.
Diagram
A lyophilised research peptide skips every step after “protein”. It is the ligand already. That is the entire point of the catalogue, and the reason it is not a gene therapy.
Crick’s flow is still right. The numbers are the part textbooks skip: a mammalian polymerase is slow, splicing is a machine the size of a ribosome, and translation errors run about one in 10⁴ amino acids.
Crick, 1958, and the 1970 restatement: sequence information flows from nucleic acid to protein, not the other way. Reverse transcription was the scandal that wasn't — RNA to DNA is still nucleic acid to nucleic acid. Prions are conformation, not sequence. The dogma survived because it was about information, not about which polymerase you like. What it did not tell you is that a human gene is, on average, a small island of exons drowned in introns, that the spliceosome is as big as the ribosome, and that most of the genome is not a gene. Those are the facts a century of biochemistry added, and they are the facts that make a lyophilised peptide interesting. The peptide on the shelf is the last box on the diagram, already filled. Everything between a promoter and that box is a machine the size of a virus, running on a clock measured in minutes to hours. Sit with that for a moment. DNA does not make protein. It is copied, spliced, shipped, printed and folded. We are going to walk every arrow, with the rates on.
In short. DNA does not make protein directly. Sequence information flows nucleic acid to protein, never backwards, and a gene is mostly intron.
The 1958 paper was a symposium talk, On protein synthesis, delivered to the Society for Experimental Biology, and it is still the cleanest sentence in the field. Crick distinguished sequence information from the energetic and stereochemical arguments that surround a peptide bond, and he put the information on a one-way street. The 1970 Nature piece, Central dogma of molecular biology, was the restatement after reverse transcriptase had been found in RNA tumour viruses by Temin and Baltimore. People had queued up to declare the dogma dead. Crick pointed out, with the patience of a man who had already been through this once, that RNA to DNA is a transfer between nucleic acids. The forbidden transfer was protein to nucleic acid, or protein to protein as sequence. That prohibition still holds. We have not found a ribosome that reads a polypeptide and writes a gene. We have found a great deal of machinery between the gene and the polypeptide, which is the subject of this essay, and none of it reverses the arrow. Reverse transcriptase was a plot twist. It was not a rewrite.
In short. Crick put sequence information on a one-way street. Finding that RNA can be copied into DNA did not reverse that rule.
Roger Kornberg took the eukaryotic polymerase apart and put a crystal structure on it; the Nobel was 2006. Robert Roeder had already shown, in the late 1960s, that animal cells run three nuclear RNA polymerases with different jobs and different toxin sensitivities: Pol I for most ribosomal RNA, Pol II for messenger RNA and a zoo of noncoding transcripts, Pol III for tRNA, 5S rRNA and a handful of small stable RNAs. Alpha-amanitin, the death-cap peptide, still does the teaching demonstration: Pol II dies first, Pol III later, Pol I barely notices. The catalogue is not a polymerase. It is, at most, a ligand someone has claimed might talk to a promoter that a polymerase will eventually read. Those are different objects, and the rest of this piece exists to keep them that way. Three polymerases, three jobs, one genome. A research peptide that 'activates transcription' in a caption has not yet said which polymerase, which gene, or which step of the assembly. The machines have names. We are going to use them.
In short. Cells run three separate RNA polymerases. A research peptide is not one of them, even if a paper claims it talks to a promoter.

The genome as a search problem
A haploid human genome is about 3.1 billion base pairs. GRCh38, the reference most of the literature still sits on, is a mosaic of clones with gaps at centromeres, acrocentric short arms and a handful of unrepeatable repeats. T2T-CHM13, the telomere-to-telomere assembly of a hydatidiform-mole genome published by Nurk and colleagues in Science in 2022, closed those gaps and put the complete haploid length near 3.05 billion bases of actual sequence. Diploid, in G1, you carry two of those, minus the sex-chromosome arithmetic. Each base pair is 0.34 nm along B-DNA, so the diploid nuclear content stretches to about two metres, packed into a nucleus six to ten micrometres across. Finding a promoter in that is not a lookup. It is a search problem with a clock, a geometry, and a chromatin state, and the search is performed by proteins that spend most of their time unbound, scanning. That is already a kind of wonder: a protein the size of a few nanometres finding a few hundred base pairs among three billion, in a packed room, in minutes.
In short. A human genome is about 3.1 billion base pairs, two metres of DNA in a nucleus a few micrometres across. Finding a gene in that is a search, not a lookup.
GENCODE puts protein-coding genes near 20,000, plus thousands of long noncoding RNAs, miRNAs, and a census of pseudogenes that used to be a punchline and is now a regulatory argument. The median protein-coding gene spans tens of kilobases to make a 1-2 kb mRNA. Dystrophin (DMD) runs 2.3 million base pairs — the largest human gene — to make a 14 kb message. Pol II at 30 nt/s would need about 21 hours of pure elongation to transcribe DMD, which is why muscle nuclei stagger the work and why truncated isoforms exist. Titin (TTN) is the other monster, not because the gene is the longest but because the coding sequence is: hundreds of exons, a polypeptide that can exceed three megadaltons depending on the isoform, the largest protein the human ribosome is asked to print. The median gene is nothing like either of those. It is 20-25 kb of genomic span, eight to ten exons, a coding sequence of one or two kilobases, and a protein of a few hundred residues that nobody has named a disease after. Most of the work of being a cell is that median gene, copied, spliced, printed, again and again.
In short. Humans have about 20,000 protein-coding genes. Most are modest; dystrophin is a 2.3-million-base marathon, and titin is the largest protein the ribosome is asked to print.
More than 98% of the haploid sequence does not code for protein. That sentence has been true since the draft genome and remains true after T2T. What it does not license is the claim that 80% of the genome is 'functional' in the sense a geneticist means. The ENCODE Project Consortium's 2012 Nature papers reported biochemical activity — transcription, binding, modification — across a large fraction of the genome, and the press release rounded that to function. Evolutionary biologists pointed out that a transposon remnant can bind a factor without having been selected to do so. Pervasive transcription is real. Selected function is a smaller set. Estimates of the fraction under purifying selection sit well below the biochemical-activity number and well above the protein-coding number; eight to fifteen percent is the range honest reviews quote. We will not overclaim 80%. We will also not pretend the noncoding 98% is junk in the 1970s sense. Regulatory sequence, structural RNA, parasitic DNA and measurement noise share the same alphabet. Distinguishing them is the field.
In short. Over 98% of the genome does not code for protein. Activity on a stretch of DNA is not proof it was selected to do a job.
- Haploid genome
- 3.1 Gbp
- Protein-coding genes
- ~20,000
- Median gene span
- 20-25 kb
- Longest gene
- DMD ~2.3 Mb
- Largest coding sequence
- TTN
- Noncoding fraction
- >98%
GRCh38 still gapped at the centromeres; T2T-CHM13 closed them. Same order of magnitude either way.
GENCODE has sat near this number for years. The explosion was noncoding annotation, not a new proteome.
Eight to ten exons. The mRNA is 1-2 kb. Introns are most of the gene.
Dystrophin. A 14 kb message. Hours of elongation, not minutes.
Titin. Hundreds of exons. A polypeptide that can exceed 3 MDa.
Not a synonym for junk, and not a synonym for function. ENCODE measured activity. Selection is stricter.
A transcription factor does not know the address. It binds short motifs that occur thousands of times, most of them irrelevant, and it finds the relevant ones because those motifs sit in accessible chromatin, next to other motifs, inside a loop that has already brought an enhancer into the neighbourhood. The search is facilitated diffusion: three-dimensional hopping through nucleoplasm, one-dimensional sliding along DNA when the factor is in contact, residence times of seconds at non-specific sites and longer at the real ones. Single-molecule imaging of transcription factors in living nuclei — the work that took the cartoon of a protein sitting on a promoter and replaced it with a protein that mostly isn't there — is why occupancy in the catalogue sense and occupancy in the ChIP-seq sense are cousins rather than twins. A peptide occupying a receptor pocket is a binding event with a defined ligand. A factor occupying a promoter is a time-average of visits. The cell does not mind. The distinction is why we are writing this: so that a blot that goes up is not automatically a polymerase that moved.
In short. A DNA-binding protein does not know the address. It samples thousands of short matches and only stays where the chromatin is open and help is already nearby.
Diagram
- 0.1 nmHydrogen atomA proton and an electron. Chemistry starts here.
- 0.3 nmWater molecule70% of a cell by mass. The solvent life is.
- 1 nmAmino acidTwenty kinds. Peptide bonds string them.
- 2–4 nmResearch peptideA named chain. BPC-157 is 1.4 kDa, 15 residues.
- 4–10 nmGlobular proteinHaemoglobin, a GPCR’s extracellular face.
- 25 nmRibosomeThe factory that reads mRNA into protein.
- 5 nmMembraneA lipid bilayer. Every compartment starts here.
- 0.5–1 µmMitochondrionA bacterium the cell swallowed and kept.
- 6–10 µmNucleusTwo metres of DNA folded into a sphere.
- 10–30 µmTypical cellA city. 10¹⁰ proteins. One genome.
- 1 mmTissue grainA thousand cells talking across ECM.
- 1.7 mYou~36 trillion human cells. Most of them are red blood cells.
Lengths are characteristic, not exact. A research peptide is closer in size to a water molecule than to the cell that assays it — which is why a 15-mer can occupy a receptor pocket a small-molecule drug also wants.
Chromatin is the on/off switch you can actually inherit
147 bp of DNA around a histone octamer (two each of H2A, H2B, H3, H4) is a nucleosome. H1 pins the entry/exit. Tails stick out and get methylated, acetylated, phosphorylated, ubiquitinated. H3K4me3 marks active promoters. H3K27me3 (PRC2) marks repression. H3K9me3 marks heterochromatin. DNA methylation at CpG, laid down by DNMTs and read by MeCP2, is the longer-term mute. Pioneer transcription factors (FOXA, GATA, PU.1) can open nucleosomal DNA that ordinary factors cannot see. That is how a genome becomes a liver rather than a neuron without changing a letter. The sequence was always there. The accessibility was not. Chromatin is the on/off switch you can actually inherit, at least for a while, and it is why two cells with almost the same DNA are not the same city. When we say a gene is off, we usually mean the nucleosomes have not been asked to move, the marks have not been rewritten, and the pioneers have not yet arrived. Turning a gene on, in a mammal, is a furniture problem before it is a polymerase problem.
In short. DNA wraps 147 base pairs around a histone spool to make a nucleosome. Chemical marks on the tails, and who can open that spool, decide whether a genome is a liver or a neuron.
Diagram
- 2 nmB-DNA0.34 nm/bp. Diploid G1 is ~2 metres of this.
- 11 nmNucleosome147 bp around a histone octamer. ~30 million per nucleus.
- loopsCTCF / cohesinEnhancers meet promoters by folding, not by sliding.
- µmA/B compartmentsHi-C: open A, closed B, territories at the lamina.
- 6–10 µmNucleusThe room. The search problem is the entire point of gene regulation.
Packing is not storage. It is the first regulatory decision: a promoter buried in H3K27me3 is not a promoter, it is furniture. Transcription starts when this origami opens the right 1,000 base pairs among 3.1 billion.
The nucleosome is not a spool that sits still. It slides, it is evicted, it is remodelled by SWI/SNF, ISWI, CHD and INO80-family complexes that hydrolyse ATP to move DNA relative to the histone core. A promoter that is occupied by a nucleosome on the TATA or the initiator is a promoter the polymerase cannot see. Chromatin remodellers and histone chaperones (NAP1, FACT, Asf1, the CAF-1 complex behind a replication fork) are the reason a genome packed at this density is still readable. FACT, in particular, travels with elongating Pol II and peels nucleosomes ahead of the polymerase so the template is naked enough to copy, then puts them back so the gene does not become an open wound. Transcription is therefore also a nucleosome-traffic problem. People who draw a polymerase on a naked line of DNA are drawing a bacterium, and even the bacterium has HU. A packed genome that can still be read, a thousand genes at a time, without unpacking the rest, is one of the quiet masterpieces of the cell. ATP pays for it, every minute, in every nucleated cell you have.
In short. Nucleosomes do not sit still. Remodellers slide or evict them so a polymerase can see the promoter, then put them back so the gene is not an open wound.
Marks that mean open, marks that mean closed
The histone language is larger than the four letters everyone quotes, but those four are the ones that earn their keep in a sentence. H3K4me3 is deposited at active promoters by COMPASS-family methyltransferases; a ChIP-seq track for it is a decent map of where Pol II has been licensed. H3K27ac, laid down by p300/CBP, marks active enhancers and the promoters they talk to; it is the acetylation you look at when you want to know whether a loop is doing business. H3K27me3 is PRC2 (EZH2 as the catalytic subunit) and it is the memory of repression that can survive a cell division. H3K9me3 is the heterochromatin mark, HP1's ligand, the thing you find at telomeres, at pericentromeres, at the lamina, at a locus the cell has decided to treat as furniture rather than as a gene. None of these marks is a master switch on its own. They are the local language in which pioneer factors, remodellers and the polymerase negotiate. Learn four letters and you can read a chromatin essay. The rest of the alphabet is there when you need it.
In short. Histone tails carry a local language of marks. Some mean a gene is open for business; others mean it is shut, or treated as furniture rather than as a gene.
- H3K4me3 — active promoters, COMPASS-family methyltransferases, a decent map of licensed start sites.
- H3K27ac — active enhancers and promoters, p300/CBP, the acetylation of a loop that is actually working.
- H3K27me3 — Polycomb repression, PRC2/EZH2, heritable through mitosis if the recruitment holds.
- H3K9me3 — heterochromatin, HP1, telomeres and lamina-associated domains, the genome as furniture.
- DNA methylation at CpG — DNMTs write, TETs oxidise toward removal, MeCP2 reads. Silent, and longer-term than a histone acetylation.
DNA methylation at CpG dinucleotides is the modification that survives a generation of textbooks. DNMT1 maintains the pattern after replication by recognising hemimethylated DNA; DNMT3A and DNMT3B de novo-methylate. TET enzymes oxidise 5-methylcytosine toward 5-hydroxymethylcytosine and further, which is one route to demethylation. MeCP2 binds methyl-CpG and recruits repression, and its loss is Rett syndrome, which is the reminder that a mute on the genome is also a neurological gene. CpG islands at promoters of housekeeping genes are usually unmethylated; methylation of those islands is a silencing event associated with development, imprinting, and, when it goes wrong, tumour-suppressor shutdown. It is slower than a histone acetylation and faster than a mutation. It is also not a peptide target in any honest catalogue sentence, which is why it is in this essay rather than on a product page. A correctly spelled genome can still be unread. That is already floor 1 leaking into floor 2 without a single variant you would report on a diagnostic panel.
In short. A methyl group on DNA is a longer-term mute than a histone mark. It is not a peptide target, which is why it sits in this essay rather than on a product page.
Pioneer factors are the proteins that can bind their motif on nucleosomal DNA, not just on naked DNA. FOXA, GATA, PU.1 are the canonical eukaryotic examples. They do not transcribe. They make a nucleosome discussable. Lineage-determining transcription factors are, to a first approximation, pioneers plus a network, which is how a haematopoietic stem cell becomes a macrophage without rewriting GRCh38. The sequence was always there. The accessibility was not. Talking about turning a gene on without talking about chromatin is talking about a plasmid in E. coli, and a human gene is not a plasmid. Pioneers crack the door; remodellers open it; the polymerase walks through. That order is why a small peptide claiming to activate transcription has a list of mechanistic questions waiting: did it recruit a pioneer, a remodeller, a pre-initiation complex, or a pause-release kinase? A blot that goes up does not tell you which. The furniture has to move before the reader arrives.
In short. A few proteins can bind their motif even on packed DNA. They do not transcribe; they make a closed gene discussable, which is how one genome becomes a macrophage.
Histone acetyltransferases (p300/CBP, GCN5/PCAF, MYST-family) put acetyl groups on lysine; histone deacetylases take them off. Acetylation neutralises a positive charge on the tail, weakens the tail-DNA conversation, and creates bromodomain-binding sites for the next complex. HDACs of the classical zinc-dependent families (HDAC1-11 in humans) reverse that. Sirtuins deacetylate histones using NAD+ as co-substrate. That sentence is why a redox cofactor is a chromatin cofactor, and why caloric restriction, NAMPT salvage and CD38-high inflammation all leak into gene regulation. NAD+ is nicotinamide adenine dinucleotide — a small molecule that shuttles electrons, and here also a coin chromatin enzymes spend. The 1000 mg NAD+ vial is not a histone deacetylase. It is the coin those enzymes spend. Spend it on PARP1 after DNA damage and the chromatin programme notices because the coin ran out. A cofactor vial is not a transcription factor. It is the budget those factors share with repair.
In short. Adding acetyl groups loosens histone tails; taking them off tightens them. Sirtuins spend NAD+ to do the taking-off, which is why a cofactor vial is not a transcription factor.
SIRT1 is nuclear and cytoplasmic and has transcription-factor substrates as well as histones. SIRT6 is chromatin-bound and has DNA-repair and ageing phenotypes in mouse genetics that are more persuasive than most of the supplement literature that borrowed its name. SIRT7 sits in the nucleolus. All of them consume NAD+ and produce nicotinamide plus O-acetyl-ADP-ribose; nicotinamide is itself a sirtuin inhibitor until NAMPT salvages it. PARP1, after a DNA break, can drain the nuclear NAD+ pool in minutes. The nuclear budget for deacetylation and the nuclear budget for repair are therefore the same coin. That is the non-brochure reason a redox cofactor keeps walking into a chromatin essay. It is also why this desk will not let a vial of NAD+ pretend to be a transcription factor. Ageing, in many tissues, correlates with a smaller NAD+ pool, which is a measurement. Repair and chromatin share that pool. A 1000 mg cake of β-NAD+ is the coin, lyophilised, for the bench. It is not a promoter, and it is not a years-added product. It is the budget.
In short. Different sirtuins sit in different rooms and all spend NAD+. After a DNA break, PARP1 can drain the nuclear pool in minutes, so repair and deacetylation share a coin.

The 3D genome: a promoter is not a flag on a linear string
Enhancers can sit 1 Mb away and still loop onto a promoter via cohesin and CTCF. The 3D genome is not a metaphor. Hi-C maps show A (open) and B (closed) compartments, topologically associating domains, and loops that bring a piece of DNA that does not code into physical contact with a transcription start site. Gene regulation is geometry plus chemistry. A promoter is not a flag on a linear string of letters. The flag is on a loop, the loop is held by a ring, the ring is loading and unloading on a timescale of minutes, and the enhancer that actually matters may not even be the nearest one on the map. That is why deleting a few bases of a CTCF site can fuse neighbourhoods a megabase apart, and why a congenital malformation can have a boring karyotype. Sequence change small; geometry change large. Once you have seen a Hi-C map, you will never again think of a gene as a line in a FASTA file. It is a point in a folded polymer, in a room six micrometres across, among two metres of neighbours.
In short. A control stretch of DNA can sit a million bases away and still loop onto the start site. Gene regulation is geometry plus chemistry, not a flag on a string.
CTCF is an eleven-zinc-finger protein that binds a motif and, when two motifs face each other, can stall cohesin. Cohesin is a ring that extrudes DNA loops until it hits that stall. The resulting loop is a topologically associating domain, a TAD, a neighbourhood in which enhancers and promoters meet each other more often than they meet the next neighbourhood. Delete a CTCF site and you can fuse TADs; fuse the wrong ones and you can put a limb-enhancer onto the wrong gene and get a developmental catastrophe that looks, in the clinic, like a congenital malformation with a boring karyotype. The sequence change is a few bases. The geometry change is a megabase. This is why a promoter is not a flag planted on a linear string of letters. The flag is on a loop, the loop is held by a ring, the ring is loading and unloading on a timescale of minutes. A research peptide does not move CTCF. The geometry is the furniture the polymerase walks into, and it was arranged before the ligand arrived.
In short. A ring protein pushes DNA into loops until it hits a stopper, making neighbourhoods where the right control DNA meets the right gene. Move the stopper and a limb-enhancer can land on the wrong gene.
A/B compartments are a coarser layer. A compartments are gene-rich, open, early-replicating, interior. B compartments are gene-poor, heterochromatic, later-replicating, often lamina-associated. A locus can switch compartments as a cell differentiates; the switch is visible in Hi-C as a change in whose DNA it now contacts. Laminopathies, the nuclear-envelope diseases, scramble this geography as well as the mechanics of the nucleus, which is one reason a structural protein of the lamina can present as a gene-regulation disease. The next essay in this desk, the nucleus piece, takes the envelope, the lamina and the telomere as its subject. This essay only needs you to believe that the promoter you are about to assemble a polymerase on is a point in a folded polymer, not a line in a FASTA file. Gene-rich DNA sits toward the interior; packed, quiet DNA often sits at the envelope. A cell type is, among other things, a folding of that polymer. The letters were always there. The folding is the difference between a hepatocyte and a neuron.
In short. Gene-rich DNA sits toward the interior; packed, quiet DNA often sits at the envelope. A promoter is a point in a folded polymer, not a line in a file.
Mediator is the 26-subunit, ~1.4 MDa complex that sits between enhancer-bound transcription factors and the polymerase at the promoter. It is not a nice-to-have. Mediator mutations are developmental diseases; Mediator biochemistry is why an enhancer loop is a transcriptional argument rather than two pieces of DNA that happen to be near each other. The Tail module of Mediator sees activators; the Head and Middle see Pol II and the general factors; the CDK8 kinase module is the optional repressor that, when bound, is not helping. Cryo-EM has put this thing on a map that would have looked like science fiction when Roeder was fractionating HeLa extract. The cartoon of a transcription factor touching a polymerase was always a fib about scale. Mediator is the fib's replacement: a 1.4-megadalton adaptor, as big as a small virus, whose job is to make a loop into a decision. Two stretches of DNA being near each other is not, by itself, a decision to transcribe. Mediator is how the nearness becomes a licence.
In short. Mediator is a huge adaptor sitting between distant control DNA and the polymerase. Two stretches of DNA being near each other is not, by itself, a decision to transcribe.
Diagram
Closed chromatin (H3K27me3, DNA methylation) hides the promoter. Pioneer factors and histone acetyltransferases open it.
PIC: TFIID, TFIIH, Mediator, Pol II. Ser5 phosphorylation of the CTD lets the polymerase leave the promoter.
Elongation ~20–40 nt/s. Capping, splicing, cleavage and polyadenylation happen on the still-growing RNA.
Human genes are islands in 3.1 billion base pairs of mostly noncoding sequence. Promoter, enhancers, chromatin state and the Mediator complex decide whether Pol II is allowed to fire. Epithalon’s literature sits on TERT and pineal clocks — two of the rare promoters anyone bothers to name in a peptide essay.
The pre-initiation complex is a machine that has to be rebuilt every time
TFIID (TBP + TAFs) finds the promoter. TFIIH helicase unwinds, kinase phosphorylates Ser5 of Pol II's C-terminal domain heptapeptide repeats (YSPTSPS, 52 of them in humans), and the polymerase is allowed to leave. Mediator — 26 subunits, ~1.4 MDa — sits between enhancers and Pol II and is not optional. Promoter-proximal pausing (DSIF, NELF) holds Pol II 20-60 nt downstream until P-TEFb (CDK9/cyclin T) phosphorylates Ser2 and NELF and the polymerase is actually allowed to work. A lot of transcriptional regulation is the decision to release the pause, not the decision to bind the promoter. Heat shock genes taught us that. Everything else borrowed it. The pre-initiation complex is a machine that has to be rebuilt every time. A gene that fires all day is a gene that has been re-licensed all day. Many genes are loaded and paused, not off. That is a more interesting off than a vacant promoter, and it is why a signal can produce mRNA in seconds rather than in the minutes it would take to assemble a PIC from scratch.
In short. Many machines must assemble at the start site before the polymerase is allowed to leave, and it often pauses just downstream. A lot of regulation is releasing that pause, not arriving at the promoter.
TBP, the TAFs, and a bend in the DNA
TATA-binding protein is the saddle. It sits on the minor groove, bends the DNA through about 90 degrees, and that bend is the nucleation point for the rest of the pre-initiation complex. Most human promoters do not have a textbook TATA box. They have CpG islands, or a combination of Initiator, downstream promoter element, motif ten, and a chromatin configuration that TFIID can still recognise through the TAF subunits. TFIID is TBP plus a dozen or so TAFs, a complex large enough to see nucleosomes and histone marks as well as sequence. TFIIA and TFIIB stabilise TBP on DNA. TFIIF comes with Pol II. TFIIE recruits TFIIH. TFIIH is the ten-subunit factory that contains XPB and XPD helicases — the same subunits that, in another life, are nucleotide-excision repair — and the CDK7 kinase that phosphorylates the CTD at Ser5. Open the duplex, mark the polymerase as initiating, and you have a transcription bubble and a CTD that can now recruit the capping enzyme. The PIC is assembled, used, and largely rebuilt. A gene that fires all day has been re-licensed all day.
In short. TATA-binding protein saddles the DNA and bends it, even at promoters with no TATA box. The start complex is assembled, used, and largely rebuilt; a gene that fires all day has been re-licensed all day.
The C-terminal domain of RPB1, the largest Pol II subunit, is a tandem array of the heptapeptide YSPTSPS. Yeast has about 26 repeats; humans have 52. Serine 5 phosphorylation is the initiation mark, laid down by CDK7 in TFIIH, and it is what the capping machinery and some of the early elongation factors read. Serine 2 phosphorylation is the elongation mark, laid down later by P-TEFb's CDK9, and it is what the splicing and 3'-processing machineries read. Serine 7, threonine 4, tyrosine 1 have their own kinases and their own readers; the CTD code that Buratowski named in 2003 is still being decoded, but the Ser5-then-Ser2 handoff is the part you can take to the bank. A polymerase with the wrong CTD phosphorylation is a polymerase that cannot recruit the next machine. Elongation is not just moving. It is moving while dressed correctly for the processing factors that travel with you. The polymerase wears a repeating tail whose phosphorylation tells the nucleus what stage it is at. Initiation and elongation are different outfits.
In short. The polymerase wears a repeating tail whose phosphorylation tells the nucleus what stage it is at. Initiation and elongation are different outfits; the wrong dress means the next machine will not board.
Promoter-proximal pausing: many genes are loaded and paused, not off
The textbook of 1990 had a polymerase that bound, initiated, and ran. The textbook of now has a polymerase that often transcribes 20 to 60 nucleotides past the start site and then sits there, held by DSIF (a heterodimer of SPT4 and SPT5) and NELF (four subunits, the negative elongation factor), until someone phosphorylates it out of the pause. P-TEFb is that someone: CDK9 plus cyclin T, the kinase that phosphorylates Ser2 of the CTD, phosphorylates DSIF into a positive elongation factor, and evicts NELF. Flavopiridol and other CDK9 inhibitors freeze the pause genome-wide; the experiment is as close to a clean genetic-pharmacological demonstration as transcription gets. John Lis's laboratory, first on Drosophila heat-shock genes and then genome-wide with GRO-seq, made the pause impossible to ignore. Karen Adelman's work made it the default for developmental genes, not a heat-shock curiosity. A large fraction of human genes have a paused polymerase. They are not off. They are loaded, and the decision is release.
In short. Pol II often copies 20 to 60 nucleotides and then waits, held until a kinase releases it. Many genes are loaded and paused, not off; the decision is release.
Heat-shock genes remain the teaching example because they are rude about it. HSF1 trimerises on heat, binds the Hsp70 promoter, and within seconds the paused polymerase is released and the gene is producing mRNA at a rate that looks like an emergency. The polymerase was already there. The chromatin was already open. The minutes you might have spent assembling a PIC from scratch are minutes a stressed cell does not have. Developmental genes use the same trick on a slower clock: pause, wait for the signal, release. HIV Tat hijacks P-TEFb to release a pause that the viral LTR depends on; the virus understood the CTD before half the textbooks did. When a paper claims that a small peptide activates transcription of a gene, the mechanistic questions write themselves. Did it open chromatin. Did it recruit a PIC. Did it release a pause. Did it stabilise an mRNA that was already being made. Those are different experiments. A blot that goes up does not tell you which one you did. Ask, and the paper gets more interesting.
In short. Heat-shock genes already have a paused polymerase waiting, so the response is seconds. A blot that goes up does not tell you whether chromatin opened, a pause released, or an mRNA was merely stabilised.
Elongation: twenty to forty nucleotides a second, with pauses
Mammalian RNA polymerase II typically elongates at about 20-40 nucleotides per second, which is about 1.2-2.4 kb per minute, commonly rounded to ~2 kb/min for a typical gene. Those numbers are averages over a lot of sitting still. Single-gene imaging and live-cell observations (Darzacq and colleagues on a reporter; Singh and Padgett on endogenous human genes; Jonkers, Kwak and Lis genome-wide) agree on the order of magnitude and disagree, productively, on the exact rate, because the rate is not a constant. The polymerase pauses at splice sites, at nucleosomes, at sequences that make the RNA:DNA hybrid misbehave. It backtracks: the 3' end of the RNA slips out of the active site and the polymerase has to be rescued by TFIIS, which stimulates a cleavage that realigns the 3' end. A 20 kb gene, at 2 kb/min of productive elongation, is ten minutes. Add pausing and it is longer. DMD, at 2.3 Mb, is not ten minutes. It is many hours of elongation even before you count the pauses, which is why a muscle fibre's nuclei are a factory with a night shift.
In short. Mammalian Pol II typically copies at about 20-40 nucleotides a second, with a lot of sitting still. A median gene is minutes; dystrophin is many hours.

Bacterial RNA polymerase is faster, as bacteria tend to be: 40-80 nt/s is a typical range for E. coli, and a 1 kb operon is a short errand. The eukaryotic tax is nucleosomes, long genes, and co-transcriptional processing that the polymerase is expected to wait for. Proofreading at the active site is modest. Transcription error rates sit around 10^-5 per nucleotide on the classical biochemical measurements, with in vivo transcriptome-wide estimates (Gout and colleagues in C. elegans, and others) in the 10^-6 to 10^-5 neighbourhood. The polymerase is not sacred and does not need to be. The message turns over. The gene does not. That gradient of care is the same one we will meet at the ribosome, one error in ten thousand amino acids, and at the replisome, one in a billion bases after mismatch repair. DNA is the archive. RNA is a draft. A 2 kb mRNA will often have a substitution, and the cell lives with it because the next message is already being written. Speed and care are a trade the cell made a long time ago, and it is a good trade.
In short. Bacteria copy RNA faster because they lack nucleosomes and long genes. Transcription error is about one in a hundred thousand; the message turns over, the gene does not.
Speed is also regulated. Pol II can be slow through a first exon and faster through a long intron; it can slow at the 3' end to give the cleavage machinery time to catch the polymerase. Phosphorylation of SPT5, the big subunit of DSIF, tracks with elongation rate. Histone marks along the gene body (H3K36me3, deposited by SETD2 travelling with Pol II) are both a consequence of elongation and a signal to the DNA-repair and splicing systems that this stretch is being read. A polymerase that is moving is a polymerase that is decorating the chromatin behind it. The gene after a round of transcription is not the gene before. That is one of the quieter facts in the field, and it is why chromatin state is a movie, not a photograph. Elongation is not a constant-speed motor. It slows for splicing and cleavage, and it leaves histone marks behind so the next round, and the repair systems, know where the reader has been. A gene is a place that remembers being read.
In short. The polymerase is not a constant-speed motor. It slows for splicing and cleavage, and it leaves histone marks behind, so chromatin state is a movie, not a photograph.
Co-transcriptional processing: the RNA is being finished while it is still being written
Capping happens as soon as the 5' end is free: 7-methylguanosine, 5'-5' triphosphate linkage. Without a cap the message is garbage and the ribosome will not start. Polyadenylation at the 3' end (CPSF, CstF, PAP, ~200 A's in mammals) is coupled to termination. Histone mRNAs are the famous exception — stem-loop, no polyA, replication-coupled. The rest of the proteome takes the polyA. Co-transcriptional processing means the RNA is being finished while it is still being written. That is not a post-hoc edit on a finished tape. The cap goes on when the RNA is only 20-30 nucleotides long, while the polymerase is still at the pause, or just leaving it. Most human introns are spliced while the polymerase is still writing the next exon. Cleavage and a poly(A) tail also stop the polymerase. Processing is the job, not a downstream afterthought, and a message that fails any of these steps should not see a nuclear pore. When it does, you get disease. When the filter is too tight, you get no protein.
In short. A cap is put on the 5' end as soon as it is free; without it the message is garbage. Most messages then get a poly(A) tail; histone mRNAs are the famous exception.
The cap is a three-enzyme argument in metazoans, condensed into a bifunctional capping enzyme plus a methyltransferase. RNA triphosphatase removes the gamma phosphate from the 5' end; guanylyltransferase hangs a GMP on through a 5'-5' triphosphate; guanine-N7 methyltransferase methylates the guanine. Banerjee's reviews from the 1980s are still the place the structure of that cap was made obvious to people who do not live inside RNA biochemistry. The cap is put on when the nascent RNA is only 20-30 nucleotides long, which is to say: while the polymerase is still at the pause, or just leaving it. CBC, the cap-binding complex (CBP80/CBP20), then rides the 5' end through splicing and export, and is swapped for eIF4E in the cytoplasm when it is time to translate. A message without a cap is a message the exonuclease XRN will eat, and a message the ribosome will not initiate on by the scanning route. Capping is not decoration. It is the licence to exist. The 5'-5' linkage is chemistry you will not find in DNA, and it is how a message proves it was born properly.
In short. The cap goes on when the RNA is only 20-30 nucleotides long, while the polymerase is still pausing. Without it the message is eaten, and the ribosome will not start.

Splicing, treated at length below, is co-transcriptional for most introns in most human genes. U1 snRNP can be at the 5' splice site of an intron before the 3' splice site has even been transcribed. That coupling is why elongation rate and splice-site choice talk to each other: a slow polymerase gives a weak splice site more time to win; a fast polymerase can skip it. The 3' end is a different machine. CPSF (cleavage and polyadenylation specificity factor) recognises the AAUAAA hexamer, CstF the downstream U- or GU-rich element, and together they recruit the endonuclease that cuts the RNA and the poly(A) polymerase that adds roughly 200 adenosines in mammals. Nick Proudfoot's laboratory spent decades showing that cleavage is also how Pol II terminates: once the RNA is cut, XRN2 chases down the still-transcribing polymerase and helps dismantle it, the torpedo model. Moore and Proudfoot's 2009 Cell review is the map of how processing reaches backward to transcription and forward to translation. Processing of an mRNA is a set of machines that have to fire in order, on a polymer that is still coming out of the polymerase, in a nucleus that is crowded.
In short. Most human introns are spliced while the polymerase is still writing. Cleavage and a poly(A) tail also stop the polymerase; processing is not a post-hoc edit on a finished tape.
Splicing is not editing. It is the gene.
The spliceosome is five snRNPs (U1, U2, U4, U5, U6) plus hundreds of proteins, assembled anew on every intron. Branch-point adenosine attacks the 5' splice site, then the exons are ligated and the intron is thrown away as a lariat. About 95% of human protein-coding genes are alternatively spliced. One gene, many proteins. DSCAM in Drosophila can theoretically make tens of thousands of isoforms; humans are less baroque and still routinely make a handful per gene. Antibody class switching and V(D)J are different chemistry (DNA recombination). Do not confuse them with splicing — people do, constantly, because both rearrange information, but one rearranges RNA and the other rearranges the archive. The gene on the chromosome is not the gene on the message. That 1977 observation still organises the field. Exons of a hundred or so nucleotides; introns that can be a hundred times longer; a 3-megadalton machine that has to find the right GU and the right AG among a sea of both. The genomic sequence is the recipe, with most of the lines crossed out.
In short. The spliceosome is rebuilt on every intron, cutting the intervening RNA out and joining the kept pieces. About 95% of human genes are alternatively spliced, so one gene is many proteins.
Diagram
U1 finds the 5′ splice site. U2 finds the branch point. U4/U5/U6 join. Two transesterifications. Lariat discarded.
Alternative splicing: cassette exons, mutually exclusive exons, intron retention, alt 5′/3′. One gene, many proteins.
Quality control: an exon-junction complex left downstream of a stop codon triggers NMD. Bad messages die in the cytosol.
~95% of human multi-exon genes are alternatively spliced. A cell type is, among other things, a splice-isoform programme. The spliceosome is comparable in mass to the ribosome and works co-transcriptionally — exons are joined before Pol II has finished the gene.
Sharp and Roberts, 1977, adenovirus late mRNAs, split genes. The Nobel was 1993, shared, and it is one of the few Nobels you can explain to a civilian in a sentence: the gene on the chromosome is not the gene on the message. Berget, Moore and Sharp, and Chow, Gelinas, Broker and Roberts, published the same year on the same virus. Everything about human gene architecture follows from that observation. Exons of a hundred or so nucleotides; introns that can be a hundred times longer; a machine that has to find the right GU and the right AG among a sea of both. The major spliceosome handles the GT-AG (GU-AG in RNA) majority. The minor, U12-type spliceosome handles a small class of AT-AC and variant introns with its own snRNPs. Miss a splice site and you have a disease: spinal muscular atrophy is a failure to include SMN2 exon 7 efficiently; a large fraction of disease mutations that look like missense or nonsense on a codon table are actually splice mutations once you look at the RNA. The gene is the spliced product.
In short. The gene on the chromosome is not the gene on the message: that 1977 observation still organises the field. Many disease mutations that look like codon errors are actually splice errors.
Alternative splicing is a cell-type programme
Exon skipping is the commonest alternative event in humans. Mutually exclusive exons, alternative 5' or 3' splice sites, intron retention, and alternative first or last exons make up the rest. The decisions are not random. They are made by RNA-binding proteins — the SR family, hnRNPs, Nova, PTB/nPTB, MBNL, RBFOX — whose concentrations and modifications differ by tissue, by developmental stage, by signalling state. A neuron is, among other things, a splice-isoform programme: thousands of exons that a hepatocyte skips, hundreds of mutually exclusive choices (the classic DSCAM comparison is fly, but mammalian neurexins and cadherin-related genes play a similar game at lower combinatorial explosion). Nilsen and Graveley's 2010 Nature review is the piece that made the exception into the rule for anyone still holding out. A proteome is larger than a gene count, and the extra is not only PTMs. It is exons. Which pieces are kept is a cell-type programme. The ribosome will print whatever the spliceosome joined. Identity, in a mammal, is partly an isoform.
In short. Which pieces are kept is a cell-type programme. A neuron includes thousands of exons a liver cell skips; the proteome is larger than the gene count, and the extra is exons.
Intron retention, long treated as a failure, is in some cases a regulatory choice: a retained intron can keep a message in the nucleus, or trigger NMD if it carries a premature stop, or add sequence to a protein. The tropomyosin genes, the FGFR family, and a long list of RNA-binding proteins themselves are spliced in ways that change function rather than just decorating it. Autoregulatory splicing of splicing factors is one of the quieter feedback loops in the nucleus: an SR protein that includes a poison exon and NMD-kills its own excess is a thermostat. The spliceosome is therefore not only a machine for making a message. It is a machine for deciding which message, in which cell, today. Keeping an intron is sometimes a choice, not a failure. Splicing factors even splice themselves as a thermostat. Once you have that picture, 'the gene' as a single protein becomes a convenience for a textbook, not a description of a tissue. A cell type is, among other things, which exons it kept this morning.
In short. Keeping an intron is sometimes a choice, not a failure: it can hold a message in the nucleus or kill it. Splicing factors even splice themselves as a thermostat.
Cech and Altman, Nobel 1989, catalytic RNA: the Tetrahymena intron that splices itself, and RNase P, which is a ribozyme that processes tRNA and happens to have a protein subunit. The spliceosome's active site, once the protein costume is stripped away by the structural work of the last decade, looks like those ancestors. U6 and U2 RNAs form a catalytic triplex that positions two magnesium ions; the chemistry is two transesterifications, the same chemistry the group II self-splicing introns use, which is why the textbooks now draw a straight line from a bacterial intron that needed no snRNPs to a 3-megadalton assembly that still uses RNA to do the cutting. Protein made it regulatable. RNA kept the scissors. If that disappoints you, look at a ribosome, which is the same joke told in rRNA: the peptidyl transferase centre is RNA, the proteins are furniture, and the joke is 4 billion years old. Catalysis by RNA is not a trivia fact. It is why RNA was probably first, and why a 3 MDa machine is still using an active site that looks like the Archaean.
In short. The spliceosome's cutting is done by RNA, dressed in protein. That is the same joke as the ribosome, and a straight line from self-splicing bacterial introns.
Quality control: a bad message should not become a protein
Nonsense-mediated decay is the best-named pathway in the nucleus-and-cytoplasm. If a stop codon sits more than about 50-55 nucleotides upstream of an exon-exon junction, the exon-junction complex left behind by the spliceosome is still there when the ribosome terminates, UPF1 is recruited with UPF2 and UPF3, and the message is degraded rather than being allowed to make a truncated protein. That is why a frameshift or a nonsense mutation often produces no protein rather than a short one, and why some recessive alleles are really NMD alleles. The first round of translation is, in this view, a quality-control lap. EJC-independent NMD also exists, reading long 3' UTRs and other insults. Nuclear retention of unspliced or incompletely processed RNA is the earlier filter: a message that still looks like a pre-mRNA should not see a nuclear pore. The cell would rather miss a protein than print a wrong one, which is the opposite of the error-rate philosophy it uses once translation has started — and the reconciliation is simple. A truncated protein can poison a complex. A substituted protein is usually just a wasted synthesis.
In short. If a stop codon arrives too early, the cell destroys the message rather than print a truncated protein. It would rather miss a protein than poison a complex.
There are other decay routes, and they matter because mRNA level on a blot is not a transcription rate. Deadenylation-dependent decay, the default, shortens the poly(A) tail until the cap is at risk; then DCP1/DCP2 decap and XRN1 eats 5' to 3', or the exosome eats 3' to 5'. AU-rich elements in 3' UTRs recruit TTP and related proteins and can dump a cytokine message in minutes. miRNAs, the Fire and Mello world (Nobel 2006, though the animal miRNA field is Lee, Feinbaum, Ambros and then Reinhart, Slack, Ruvkun), put RISC on a 3' UTR and repress translation and/or destabilise the RNA. None of this is transcription. All of it looks like transcription if you only measure steady-state mRNA. A peptide literature that reports a qPCR change and calls it gene activation has not earned the phrase until someone has asked whether the polymerase moved. Steady-state mRNA is not a transcription rate. Decay and microRNAs look like gene activation if you only measure the pile. Ask about the polymerase, and the pile becomes a rate.
In short. Steady-state mRNA is not a transcription rate. Decay and microRNAs look like gene activation if you only measure the pile, which is why a qPCR bump has not earned that phrase.
Nuclear export: mRNA is licensed to leave
The nuclear pore complex is ~110 MDa, one of the largest machines in the cell, ~3,000 per nucleus, a channel that passively admits ~40 kDa and actively escorts everything else via importins, exportins and the Ran-GTP gradient. An mRNP (message plus proteins, including the exon-junction complex) is too big to wander through. NXF1/NXT1 is the export receptor for bulk mRNA. If splicing failed, the message should not leave. When it does, you get disease. When it doesn't, you get no protein. mRNA is licensed to leave. It does not leak. Small proteins wander through; mRNA is walked through by an export receptor after it has been licensed as fully spliced. The honest cellular message is a passport-holding particle. Some viral RNAs have learned to cheat, which is how HIV Rev and the CTE of simpler retroviruses became textbooks. The cellular default is stricter, and it is why undergraduates who think translation happens in the nucleus have skipped a step that is, in mass, one of the largest machines they will ever meet.
In short. The nuclear pore is enormous and will not let an unfinished message wander out. If splicing failed, the RNA should stay put; when that filter fails, you get disease or no protein.
Diagram
Out → in
Proteins, TF, histones
Importins + Ran-GTP cycle. A transcription factor that cannot clear the pore is not a transcription factor. It is a cytosolic rumour.
The mesh
NPC · FG nups
Passive cutoff a few nanometres. A ribosomal subunit is assembled in the nucleolus and exported as cargo, not as a wanderer.
In → out
mRNA, assembled ribosomes
TREX, NXF1/NXT1. Unspliced RNA is retained on purpose. Export is a licence, not a leak.
~3,000 pores per nucleus. ~30 nucleoporins. FG-repeat mesh that lets small molecules through and makes macromolecules show a passport (NLS, NES, NXF1 for mRNA). The nucleus is not a bag. It is a gated compartment.
The pore is a ring of rings: concentric coats of nucleoporins, a nuclear basket on the nucleoplasmic face, cytoplasmic filaments on the other, and a central channel lined with FG-repeat proteins that form a selective phase. Small proteins wander through. Everything else needs a transport receptor. Importins and exportins for proteins use the Ran-GTP gradient — Ran-GTP nuclear, Ran-GDP cytoplasmic, RCC1 the nuclear GEF, RanGAP the cytoplasmic GAP — as the direction signal. mRNA bulk export is different. TREX (transcription-export complex) is recruited co-transcriptionally, Aly/REF and other adaptors land on the processed mRNP, and NXF1 (also called TAP) with its partner NXT1 (p15) is the receptor that actually walks the particle through the FG hydrogel. The exon-junction complexes are part of the passport: a fully spliced mRNA looks different to the pore than an unspliced one. Some viral RNAs have learned to cheat. The honest cellular message is licensed. It does not leak. Being the right shape, with the right escorts, is the hard part. Finding the door, among three thousand of them, is not.
In short. Small proteins wander through the pore; mRNA is walked through by an export receptor after it has been licensed as fully spliced. The honest cellular message does not leak.
Three thousand pores sounds like a lot until you count the traffic. Ribosomal subunits, assembled in the nucleolus, are a large fraction of it; tRNAs, snRNPs that have to mature in the cytoplasm and come back, importins recycling, histones in S phase, transcription factors on their futile cycles of entry and exit. mRNA is one cargo among many. A nucleus that has been damaged at the pore — a laminopathy, a nucleoporin mutation, a viral attack on the basket — is a nucleus that has a gene-expression problem even if every polymerase is healthy. Export is a step. Skipping it in a diagram is how you get undergraduates who think translation happens in the nucleus. The cytoplasm is the next room. It is not empty. The message has to leave the nucleus, find a ribosome, and, if it is a secreted or membrane protein, find the ER. Three thousand customs posts, a queue of ribosomes and histones and recycling importins, and a particle that has earned its passport. That is the door between the archive and the draft.
In short. Three thousand pores still have a queue: ribosomes, tRNAs, recycling importins, histones in S phase. Skip export in a diagram and undergraduates think translation happens in the nucleus.

Translation initiation: scanning to the first honest AUG
Translation in eukaryotes starts with the 40S subunit, eIF2-GTP-Met-tRNAi, and the eIF4F cap-binding complex. The 43S pre-initiation complex scans from the cap to the first good AUG in a Kozak context. 60S joining makes 80S. Elongation: eEF1A brings the charged tRNA, the peptidyl transferase centre (ribosomal RNA, not a protein) makes the peptide bond, eEF2 translocates. Three nucleotides per residue, 5-6 residues a second, GTP spent at both steps. Termination: a stop codon, eRF1 looking like a tRNA, eRF3, then recycling. Secreted and membrane proteins are born into the ER through the Sec61 translocon, signal peptide first, SRP as the usher. The cytosol never sees the inside of a GPCR's extracellular loops. Folding is PDI and calnexin and a quality-control argument that can end in ERAD and the proteasome. The ribosome scans from the cap to a decent start codon, then adds amino acids at five to six a second. A 400-residue protein is about a minute of elongation. Initiation is often the wait.
In short. The ribosome scans from the cap to a decent start codon, then adds amino acids at five to six a second. Secreted proteins are born into the ER, not the cytosol.
eIF4E binds the cap. eIF4G is the scaffold. eIF4A is the helicase that melts structure in the 5' UTR so the 40S can scan. Together they are eIF4F, and eIF4E is the subunit everyone argues about because it is a node: 4E-BPs, phosphorylated via mTORC1, release eIF4E when the cell is well-fed; a starved cell keeps eIF4E off the cap and global initiation drops. That is one of the actual mechanisms connecting amino-acid status to protein synthesis, and it is why mTORC1 inhibitors change the proteome without changing the genome. The 43S pre-initiation complex is the 40S subunit plus eIF1, eIF1A, eIF3, eIF5, and the eIF2-GTP-Met-tRNAi ternary complex. It lands at the cap via eIF4F and scans 5' to 3' until an AUG in a Kozak context (gccRccAUGG, the important letters being the purine at -3 and the G at +4) is a good enough match to stop it. eIF5 and eIF5B help the 60S join. IRES-driven initiation, used by some viruses and a minority of cellular messages, skips the cap. Most of your proteome does not.
In short. A cap-binding complex lands the small ribosomal subunit, which then scans to the first honest AUG. When the cell is starved, that landing is withheld, so protein synthesis drops without touching the genome.
Five to six amino acids a second, and initiation is still the wait
A mammalian ribosome adds residues at about 5-6 per second. Yeast is a little quicker. Bacteria run 12-21, which is why an E. coli protein of 300 residues can be finished in under half a minute of elongation and why bacterial polysomes look so busy on a classic electron micrograph. A 400-residue human protein is, on elongation time alone, a bit over a minute. Initiation is often slower than that minute. Getting the 43S onto the cap, scanning a long or structured 5' UTR, waiting for a ternary complex when eIF2 is phosphorylated (the integrated stress response: PERK, GCN2, PKR, HRI — four kinases, one serine on eIF2-alpha, a global initiation brake) — those are the waits. Polysomes are the compensation: many ribosomes on one mRNA, spaced roughly every 30-100 nucleotides depending on the message, so that a single successful initiation is not a single protein, it is a queue. Ribosome profiling (Ingolia, Weissman and colleagues) made the queue countable. A translation rate for a gene is occupancy times elongation, and occupancy is mostly initiation.
In short. A mammalian ribosome adds five to six amino acids a second; a 400-residue protein is about a minute of elongation. Initiation is often the wait, and polysomes put many ribosomes on one message.
Elongation, stalling, drop-off, release
eEF1A (eEF1-alpha) delivers aminoacyl-tRNA to the A site, GTP is hydrolysed if the codon-anticodon match is good, the tRNA accommodates, and the peptidyl transferase centre — a ribozyme, catalytic RNA, the same joke as the spliceosome — makes the peptide bond. eEF2 then translocates, another GTP, and the A site is empty again. A mismatch that escapes decoding is about one in ten thousand amino acids, the translation error rate. Stalling happens on rare codons, on polyproline stretches (eIF5A, the hypusine-containing factor, is the specialist rescue for those), on damaged mRNA, on a missing stop. Drop-off is the ribosome leaving before a stop; it is rare in a healthy elongation but not zero. Bacteria have tmRNA (SsrA) to rescue a ribosome stuck at a broken message: a chimeric tRNA-mRNA that tags the truncated protein for degradation and frees the subunit. Humans do not have tmRNA. They have Pelota/HBS1L and the SKI complex and a set of rescue factors that are not tmRNA and should not be described as if they were. The bacterial solution is elegant and it is not ours.
In short. The peptide bond is made by RNA, with about one error in ten thousand. Bacteria rescue stuck ribosomes with tmRNA; humans use other factors, and should not be described otherwise.
Termination is a stop codon in the A site, eRF1 shaped like a tRNA, eRF3 a GTPase, hydrolysis of the completed chain off the P-site tRNA, then recycling by ABCE1 and the initiation factors that will use the 40S again. Stop-codon readthrough happens; it is how some viruses extend a polyprotein and how some cellular genes make a minor longer isoform. It is also how aminoglycosides, in a research setting, can force a ribosome past a premature stop — an experiment, not a protocol, and not a catalogue claim. The finished chain, if it is a cytosolic protein, starts folding while it is still coming out of the tunnel. If it has a signal peptide, the SRP has already caught it, the ribosome has already been parked on Sec61, and the chain is being born into the ER lumen, not into the cytosol. Blobel and Dobberstein, 1975, the signal hypothesis; Blobel's Nobel was 1999. A GPCR's extracellular face has never been in the cytoplasm. Reconstituting a receptor in a dish from a lyophilised peptide is occupancy. It is not reconstituting that birth.
In short. A stop codon releases the finished chain. If it has a signal peptide, it was born into the ER, not the cytosol; a lyophilised peptide in a dish is not reconstituting that birth.

Folding: the chain is not the protein until it has a shape
Hsp70 family chaperones meet nascent chains. They bind exposed hydrophobic stretches, hydrolyse ATP, and give the chain another chance not to aggregate. Hsp90 takes on a more select clientele — steroid receptors, some kinases, a list that has kept cancer biologists in geldanamycin analogues for decades — and holds them in a near-native state until ligand or phosphorylation decides. TRiC/CCT is a double-ring chaperonin, the eukaryotic GroEL cousin, essential for actin and tubulin and a set of other clients that will not fold without a cage. These are not optional quality-of-life features. They are why a 300-residue protein in a crowded cytosol, at 200-300 mg/ml macromolecules, does not become a blob of wasted synthesis. Hartl, Horwich, Bukau, Lindquist: the chaperone field is not short of names, and the 2011 Hartl review in Nature is a door into it. The ribosome made a polymer. The chaperones made a protein. Sometimes they fail, and then ubiquitin and the proteasome (Ciechanover, Hershko, Rose, Nobel 2004) turn the polymer back into amino acids. Folding is not spontaneous in a crowded cell. Anfinsen's small proteins in dilute buffer were a different planet.
In short. Chaperones meet the growing chain so a crowded cell does not turn synthesis into a blob. The ribosome made a polymer; the chaperones made a protein, and the proteasome undoes the failures.
In the ER the problem is different. The chain is being N-glycosylated as it enters, disulphides are being made and isomerised by PDI-family enzymes, and calnexin and calreticulin are lectin chaperones that hold a glycoprotein until its glycan is trimmed to the form that means 'try again' or 'you are done'. BiP (the ER Hsp70) watches the rest. Failure is the unfolded protein response: IRE1 splices XBP1 mRNA, a rare cytoplasmic splicing event, to make a transcription factor; PERK phosphorylates eIF2-alpha and slows initiation so the ER is not fed more chains; ATF6 leaves for the Golgi, is cleaved, and becomes another transcription factor. The three of them upregulate chaperones, expand ER, and, if that fails, move the cell toward apoptosis. Walter and Ron's 2011 Science review is the map. A secreted peptide that you synthesise on a resin never sees this. A recombinant protein made in CHO cells sees nothing else. That is one of the quieter differences between a 15-mer and a 191-residue hormone, and it is not a value judgement. It is a factory tour.
In short. In the ER, sugars and disulphides are added while chaperones hold the chain. A peptide made on a resin never sees this; a recombinant hormone in CHO cells sees nothing else.
PTMs: a protein is a family of proteoforms
Phosphorylation is the one everyone can name. A proteome snapshot finds phosphate on the order of a few percent of residues at a time, which sounds small until you remember that the sites are concentrated on the proteins that are doing the signalling, and that occupancy at a given site can be anywhere from a trace to saturation. Kinases write (about 500 in the human genome, the kinome), phosphatases erase, SH2 and 14-3-3 and a hundred other domains read. A phosphoprotein is not a binary. It is a set of occupancy patterns. Glycosylation is the ER-and-Golgi decoration: N-linked on asparagine in the sequon Asn-X-Ser/Thr, O-linked on serines and threonines, a branching carbohydrate chemistry that a ribosome does not know about. Ubiquitin is a 76-residue tag, itself a protein, attached as a monomer or as chains of different linkages; K48-linked chains are the proteasome ticket, K63-linked chains are more often a signalling scaffold. Acetylation of lysines, on histones as we have seen and on thousands of other proteins as well, is the acetyl-CoA conversation. Lipidation is how a protein that has no transmembrane helix still lives on a membrane.
In short. Phosphate, sugar, ubiquitin, acetyl and lipid decorations make one gene product a family of molecules. The ribosome does not know about any of that chemistry.
Neil Kelleher and others have been pushing the word proteoform for this reason: the product of a gene is not a protein, it is a family of molecules that share a backbone and differ in splicing and modification. A Western blot with one band is a convenience. A top-down mass spectrum is an argument. The catalogue sells defined sequences, often unmodified, occasionally with a disulphide or an amidated C-terminus if the published structure has one. That is a proteoform of one. The cell's version of the same sequence, if the sequence is even endogenous, may be glycosylated, phosphorylated, cleaved, and present at a dozen masses. Occupying a receptor with the synthetic version is still occupancy. It is not identity with the endogenous proteoform. Honesty about that gap is the difference between a research ligand and a pretend hormone. The cell's version of a sequence may be a family. The vial is one member, characterised, for the bench. Both can bind. They are not the same molecule until someone has shown they are.
In short. The cell's version of a sequence may be glycosylated, phosphorylated, cleaved, and present at a dozen masses. Occupying a receptor with the synthetic version is still occupancy; it is not identity with the endogenous molecule.
- Phosphorylation — kinases write, phosphatases erase, a few percent of residues in a snapshot, concentrated on the signalling set.
- Glycosylation — N-linked in the ER, O-linked in the Golgi, a chemistry the ribosome does not speak.
- Ubiquitin — 76 residues, chains with different linkages, proteasome ticket or signalling scaffold depending on the geometry.
- Acetylation — acetyl-CoA as the donor, histones and a thousand other proteins, sirtuins and classical HDACs as the erasers.
- Lipidation — myristoyl, palmitoyl, prenyl, GPI. A membrane address without a transmembrane helix.
Error rates: DNA is sacred, protein is disposable
Replication, after polymerase proofreading and mismatch repair, runs at about 10^-9 to 10^-10 errors per base pair. Kunkel's reviews are the canonical numbers; a human diploid genome of 6 Gbp, copied, would be expected to pick up a handful of mutations per cell division once you do the arithmetic, which is both reassuringly small and how cancer still happens over a lifetime of divisions. Transcription error rates are about 10^-5 per nucleotide on the classical measurements, with some in vivo estimates a little lower. A 2 kb mRNA will often have a substitution. Translation error rates are about 10^-4 per amino acid. A 400-residue protein has a few percent chance of a wrong amino acid. The cell lives with that because proteins turn over, because most substitutions are boring, and because the alternative — making the ribosome as careful as the replisome — would be too slow and too expensive for a draft. DNA is the archive. RNA and protein are drafts. Ageing is partly what happens when the archive still drifts, and when mitochondria, copying 16.6 kb circles in a ROS-rich matrix with a thinner repair budget, drift faster.
In short. DNA is copied at about one error in a billion bases; RNA and protein are allowed to be sloppier because they are drafts. Ageing is partly what happens when the archive still drifts.
Diagram
- DNA replication + MMR10⁻⁹ to 10⁻¹⁰A genome of 6 Gbp (diploid) accumulates a handful of mutations per division.
- Transcription~10⁻⁵RNA is disposable. The cell can afford a wrong letter in a message that lasts hours.
- Translation~10⁻⁴One wrong amino acid per ten thousand. Proteins turn over. DNA does not.
- mtDNA10–100× nuclearNo histones, ROS next door, weaker repair. The second genome ages faster.
The genome is sacred, the message is cheap, the protein is cheaper. Ageing is partly what happens when the sacred copy still drifts — and when mitochondria, which never got the nuclear repair budget, drift faster.
This gradient of care is one of the deepest facts in molecular biology and one of the least quoted on product pages. A research peptide, synthesised chemically, has a different error structure: deletion sequences from missed couplings, incomplete deprotections, aspartimide, racemisation at some residues. HPLC is the filter; mass spec is the identity check. The error is not a ribosome's 10^-4. It is a synthetic chemistry problem, and it is why a chromatogram with a main peak at ≥98% is the sentence we actually sell. The cell and the chemist both make amide bonds. They fail differently. They are checked differently. Neither of them is a reason to confuse a vial with a nucleus. The replisome is careful because the copy is permanent. The polymerase and the ribosome are allowed to be sloppy because the product is not. A synthesiser is careful because the customer has a chromatogram. Three factories, three error budgets, one class of bond. Holding those three in the head at once is the literacy this essay is for.
In short. A research peptide fails by missed couplings and chemistry, not by a ribosome's one-in-ten-thousand. HPLC and mass spec are the filter; the cell and the chemist both make amide bonds, and fail differently.
Clocks on one gene: tens of minutes to hours, not milliseconds
Put one ordinary gene on a stopwatch. Chromatin opening at a primed locus can be minutes; at a silent heterochromatic locus it can be much longer and may need a cell division. PIC assembly is minutes. Pause-release is seconds to minutes after the signal. Elongation of a 20-25 kb median gene, at ~2 kb/min of productive polymerase, is on the order of ten to fifteen minutes, plus pausing. Co-transcriptional splicing is happening during those minutes, not after them. Cleavage, polyadenylation and export are more minutes; some messages sit in the nucleus longer than they sit in the cytoplasm. Initiation on a freshly exported mRNA is seconds to minutes depending on 5' UTR and eIF4E availability. Elongation of a 400-residue protein is about a minute. Folding of a simple domain can be co-translational; folding of a multi-domain ER protein, with glycans and disulphides, can be many minutes and may fail. The whole pipeline, chromatin-open to a folded protein you could put on a blot, is tens of minutes for a cooperative, already-open, median gene, and hours for something like DMD. It is not milliseconds. Receptor occupancy, once you have the ligand, can be milliseconds to seconds.
In short. From open chromatin to a folded protein is tens of minutes for a typical gene, hours for dystrophin. Occupancy, once you have the ligand, can be milliseconds: that gap is the commercial fact.
- Chromatin opening and enhancer looping — minutes at a primed locus, longer if you have to fight H3K27me3 or a TAD boundary.
- PIC assembly, TFIIH, Ser5 phosphorylation, the pause — minutes, and in many genes the polymerase is already waiting.
- Elongation at ~2 kb/min — a median gene is ten to fifteen minutes of productive polymerase; DMD is many hours.
- Capping, splicing, cleavage, poly(A) — co-transcriptional for most of it, not a separate shift.
- Export through ~3,000 pores — minutes, with TREX and NXF1/NXT1 as the licence.
- Initiation, a minute or so of elongation for a 400-residue protein, folding and PTMs — the cytoplasm's contribution, and the ER's if you are secreted.
The cell takes hours to find a gene, minutes to transcribe it, seconds to translate a domain, and milliseconds to occupy a receptor. We sell the last object. We do not sell the hours.
How a research peptide skips all of this
Somatropin — 191 residues, the ligand the GH receptor actually wants — is made industrially in E. coli or mammalian cells as a recombinant protein. That is the pipeline above, hijacked. IGF-1 LR3 is an 83-residue analogue, also a biosynthesis product. BPC-157 is 15 residues and is almost always solid-phase chemistry: Merrifield resin, Fmoc, TFA cleavage, HPLC. The amide bond is the same bond. The factory is not. Saying peptides are just small proteins skips the spliceosome; saying proteins are just big peptides skips the translocon. Both sentences are missing a machine. Somatropin is this whole pipeline hijacked in a tank; a fifteen-residue peptide is almost always chemistry on a bead. Same amide bond, different factory. Recombinant hormones see chaperones, glycosylation if the host is mammalian, a signal peptide if you left one on. A 15-mer on a resin sees a chemist and a chromatogram. Both can occupy a pocket. Only one of them was born into a lumen. That is not a value judgement. It is a factory tour, and it is why the two objects sit on different tills even when they share a catalogue.
In short. Somatropin is this whole pipeline hijacked in a tank; a fifteen-residue peptide is almost always chemistry on a bead. Same amide bond, different factory.
R. B. Merrifield, 1963, Journal of the American Chemical Society: Solid Phase Peptide Synthesis. I. The Synthesis of a Tetrapeptide. He anchored a C-terminal residue to an insoluble resin and added the chain, one protected amino acid at a time, washing away excess reagents instead of purifying a soluble intermediate at every step. The Nobel was 1984. The chemistry evolved from Boc/benzyl to Fmoc/tBu, which is what almost everyone uses now: base-labile Fmoc off, couple the next residue with an activating reagent, cap any unreacted amines so they do not become deletion sequences, repeat. A four-residue peptide is four cycles. A fifteen-residue peptide is fifteen. Each cycle is a chance to fail, which is why a 50-mer is a different proposition from a 10-mer and why a 191-residue hormone is not, in any serious factory, an SPPS product. At the end you cleave from the resin with acid (TFA for Fmoc chemistry), scavenge the protecting groups, and you have a crude peptide that is a mixture. HPLC is the argument that the main peak is the sequence you named. Mass spectrometry is the argument that the mass is the mass.
In short. Merrifield anchored a chain to a resin and added one residue per cycle. HPLC and mass spec then argue that the main peak is the named sequence; a 191-residue hormone is not a bead product.

The same amide bond. The different factory. In the cell, the peptide bond is made by an RNA active site in a 4-megadalton ribosome, at 5-6 residues a second, directed by a message that was itself transcribed and spliced, from a gene that was found in 3.1 billion base pairs of chromatin. On the bench, the peptide bond is made by a chemical activator on a bead, at one residue per cycle, directed by a synthesiser programme a chemist wrote, from a catalogue of protected amino acids. There is no intron. There is no nuclear pore. There is no chaperone, unless you count the chemist staring at a chromatogram. There is no PTM unless you put it in on purpose. The lyophilised cake in the vial is the finished ligand. Reconstitution is not translation. It is dissolving a solid. That gap is a superpower in a dish: you can occupy a pocket without asking a nucleus to transcribe anything. It is also a legal class on a vial. Research use only. Characterised sequence. Not a gene, not a therapy, not a treatment plan.
In short. In the cell the bond is made by a ribosome reading a spliced message; on the bench it is a chemist and a bead. Reconstitution is dissolving a solid, not translation.

This is the sentence the catalogue is built on. A research peptide is a defined sequence you can occupy a pocket with, without asking a nucleus to transcribe it. That is a superpower in a dish and a legal class on a vial. It is not gene therapy, it is not a recombinant hormone unless it actually is (HGH, IGF-1 LR3), and it is not boosting transcription unless you are specifically talking about a literature that claims that — Epithalon on TERT, which is the next essay — and even then you are talking about a paper, not a protocol. Occupancy is the catalogue. Gene therapy and CRISPR are other floors. We stock characterised sequences so a bench can ask a pocket question with a mass on the vial. We do not stock a promoter, a polymerase, or a spliceosome. The machines in this essay are the reason that sentence has to be long. Four residues claiming a promoter still have to go through every machine named above, or around them, and the blot will have to say which.
In short. A research peptide occupies a pocket without asking a nucleus to transcribe it. That is a superpower in a dish and a legal class on a vial, not gene therapy or a transcription protocol.
Four residues claiming a promoter
Epithalon is Ala-Glu-Asp-Gly. 390 daltons. Four residues. The Khavinson literature, a Soviet and then Russian programme that is large, internally consistent, and thinner in independent Western replication than a molecule this famous should be, reports effects on TERT expression and telomerase activity, and on pineal melatonin amplitude. TERT is a reverse transcriptase. Its job, with TERC, is to extend telomeres. It is off in most somatic cells on purpose — the tumour-suppression bargain — and on in the germline, in stem cells at a trickle, and in most cancers as a hijack. A claim that a tetrapeptide moves TERT is a claim about transcription, or about the stability of a message, or about a signalling pathway that ends at that promoter. It is not a claim about a ribosome printing TERT faster. Four residues do not elongate a polypeptide of that size. They might, if the papers are right, speak to the promoter that licenses the polymerase that makes the message that the ribosome would then print. That is a gene-regulation claim. It sits on this pipeline at the PIC-and-chromatin end, not at the elongation-and-folding end. The next essay is the nucleus, the telomere, and those papers.
In short. Epithalon is four residues with a literature that points at TERT transcription, not at a ribosome. A tetrapeptide on a till and a promoter in a 3.1-billion-base search problem are different objects.
HGH is the control thought-experiment. 191 residues. The pituitary somatotroph transcribes GH1, splices it (the 22 kDa isoform is the main product; a 20 kDa splice isoform exists), exports the mRNA, translates it into the ER, folds it with a disulphide geometry the GH receptor knows, and secretes it. The vial labelled somatropin is that protein, made in a tank, by a recombinant host that was given the coding sequence and skipped the 2.3-megabase dramas of a gene like DMD. The vial labelled Epithalon is four residues made on a resin. Both are research ligands in we. Only one of them is a protein the ribosome would recognise as a job. Only one of them has a literature that points at a promoter rather than at a receptor. Holding those two facts in the head at the same time is the whole literacy this essay is for. Growth hormone is 191 residues, made by a cell or a tank, aimed at a receptor. Epithalon is four residues on a resin with a promoter-level literature. Same catalogue. Different factories. Different floors. Different questions a blot can answer.
In short. Growth hormone is 191 residues, made by a cell or a tank, aimed at a receptor. Epithalon is four residues on a resin with a promoter-level literature; holding both facts is the point.
Occupancy, not gene therapy, not CRISPR
Gene therapy puts a coding sequence into a cell and asks the cell to run this entire pipeline on a cargo it did not have yesterday. CRISPR, in its editor forms, changes the archive and then asks the cell to run the pipeline on a slightly different gene. A research peptide puts a finished ligand next to a receptor and asks about occupancy. Those are three different acts, on three different floors of the building, with three different failure modes and three different regulatory categories. This catalogue is the third act. It is labelled research use only because that is what it is: characterised sequences for experiments that occupy pockets, stain blots, and test whether a literature still stands when the chromatogram is honest. It is not a protocol, a dose, a physiology, or a transcription factor, even when the literature under discussion (Epithalon, TERT, the pineal) is a transcription-level literature. The machines in this essay are the reason that sentence has to be long. Demand the machine, and the claim gets possible.
In short. Gene therapy adds a coding sequence; CRISPR edits the archive; a research peptide puts a finished ligand next to a receptor. This catalogue is the third act, and it is labelled as such.
If you have read this far you have watched a genome be searched, a nucleosome negotiate, a loop form, a polymerase pause, a spliceosome rebuild itself on every intron, a pore refuse an unfinished message, a ribosome scan, a chaperone argue, and a chemist skip the whole lot with a resin bead. The clocks are real. The masses are real. The error rates are a gradient of care that evolution settled on because archives and drafts have different jobs. The vial is the draft, already drafted. What you do with it in a dish is occupancy. What the cell does to make the endogenous version, when there is an endogenous version, is this essay. The nucleus is next door. TERT is off in most of your soma on purpose. Four residues claiming to talk to it will have to go through every machine named above, or around them, and the blot will have to say which. We stock the tetrapeptide so that question can be asked with a mass on the vial. We do not stock a promoter. Research use only is the label on every one of them.
In short. You have now watched every machine between a gene and a folded protein, and a chemist skip the lot with a bead. The vial is the draft, already drafted; occupancy is the dish experiment.
A promoter is a place in a folded genome. A peptide is a sequence in a vial. The literature that puts the second on the first is a claim about a machine this essay just walked through. Demand the machine.
Questions the essay actually answers
- Does DNA make protein directly?
- No. DNA is transcribed to pre-mRNA, spliced to mRNA, exported, and translated by ribosomes. The dogma is DNA to RNA to protein. Reverse transcriptases (including TERT) are the famous exception for DNA from RNA, not protein from DNA. Crick's 1970 restatement was about information flow, not about which polymerase you prefer.
- How does transcription work?
- Chromatin opens, a promoter is found, TFIID and Mediator and Pol II assemble a pre-initiation complex, TFIIH unwinds and phosphorylates the CTD at Ser5, the polymerase pauses 20-60 nucleotides downstream, P-TEFb releases the pause, and Pol II elongates at about 20-40 nucleotides per second while the RNA is capped, spliced and eventually cleaved and polyadenylated. That is transcription. Binding a promoter is the least of it.
- How fast is a ribosome?
- A mammalian ribosome adds about 5-6 amino acids per second. Bacterial ribosomes run 12-21. A 400-residue protein is roughly a minute of elongation on one ribosome. Initiation is often slower than that, and polysomes put many ribosomes on the same message, so the cell is not waiting on a single machine.
- How is a research peptide made, versus a protein in a cell?
- The cell transcribes, splices, exports, translates and folds. A research peptide is almost always solid-phase synthesis: Merrifield resin, one residue per cycle, cleavage, HPLC, mass spec. Same amide bond. Different factory. Recombinant hormones such as somatropin are the pipeline in a tank, not a resin bead.
- What is splicing?
- Most human genes are split into exons separated by introns. The spliceosome, a 3-megadalton assembly of five snRNPs, cuts the introns out of the still-growing RNA and ligates the exons. About 95% of human multi-exon genes are alternatively spliced. A cell type is, among other things, a splice-isoform programme.
- What is a promoter?
- A promoter is the DNA sequence, plus the chromatin state and the looping geometry, at which RNA polymerase II is licensed to start. TATA boxes and CpG islands are motifs, not the whole story. Enhancers can sit a megabase away and still loop onto it. A promoter is not a flag on a linear string.
- Where does Epithalon sit on this pipeline?
- The Khavinson literature sits on TERT transcription and pineal melatonin — gene-level claims, not a ribosome. Four residues. A promoter. Different floors of the building. The nucleus essay next door is that argument in full.
- How long does a cell take to go from gene to folded protein?
- Tens of minutes to hours, depending on the gene. Chromatin opening and PIC assembly are minutes. A 20 kb transcription unit is minutes of elongation. Dystrophin is many hours. Splicing is co-transcriptional. Export and initiation add minutes. A 400-residue elongation is about a minute. Folding can be co-translational or take longer in the ER. It is not milliseconds.
- What is promoter-proximal pausing?
- Pol II often transcribes 20-60 nucleotides and then waits, held by DSIF and NELF, until P-TEFb (CDK9/cyclin T) phosphorylates the CTD at Ser2 and the pause factors. Many genes are loaded and paused, not off. Heat-shock genes taught the field that. Release of the pause is a large fraction of what people mean by transcriptional regulation.
- Why is DNA copied more carefully than protein is made?
- Replication, after proofreading and mismatch repair, runs at about one error in a billion to ten billion base pairs. Transcription is about one in a hundred thousand. Translation is about one in ten thousand amino acids. DNA is the archive. RNA and protein are drafts. The cell treats them accordingly.
Hypothetical research reconstitution
How these vials are typically mixed
Hypothetical research reconstitution for the named catalogue vial. Not a protocol, not medical advice, not a use instruction. These amounts sit in published and commonly cited laboratory ranges. The vial is labelled for research use only — not for human or veterinary administration.
Epithalon
50mg
Mix with 5 ml bacteriostatic water → 10 mg/ml
- Hypothetical aliquot
- 5–10 mg
- 0.50–1.00 ml · 50–100 units on a U-100 syringe
- How often
- Once daily, evening, for 10–20 consecutive days
- 10–20 days, two cycles a year in the Khavinson-school notes
Bench steps
- Let the vial sit until it is no longer cold to the touch.
- Wipe the stopper with 70% isopropyl alcohol. Let it dry.
- Draw 5 ml bacteriostatic water (0.9% benzyl alcohol).
- Run the water slowly down the inside glass — do not blast the cake.
- Roll between finger and thumb until the cake is gone. Do not shake.
- Label the date. Store the solution at 2–8 °C. Do not freeze. Use within 30 days unless the note below says otherwise.
Tetrapeptide (AEDG). Short pulses, not a daily-forever molecule in that literature.
HGH
24 IU
Mix with 2 ml bacteriostatic water → 12 IU/ml
- Hypothetical aliquot
- 1–2 IU
- 0.08–0.17 ml · 8–17 units on a U-100 syringe
- How often
- Once daily, usually an evening aliquot in the somatropin notes
- 8–12 weeks, then a pause
Bench steps
- Let the vial sit until it is no longer cold to the touch.
- Wipe the stopper with 70% isopropyl alcohol. Let it dry.
- Draw 2 ml bacteriostatic water (0.9% benzyl alcohol).
- Run the water slowly down the inside glass — do not blast the cake.
- Roll between finger and thumb until the cake is gone. Do not shake.
- Label the date. Store the solution at 2–8 °C. Do not freeze. Use within 30 days unless the note below says otherwise.
24 IU in 2 ml. Two IU is about 17 units on the syringe. Gentle roll only — somatropin denatures if you beat it.
IGF-1 LR3
1000mcg
Mix with 1 ml bacteriostatic water → 1,000 mcg/ml
- Hypothetical aliquot
- 20–50 mcg
- 0.02–0.05 ml · 2–5 units on a U-100 syringe
- How often
- Once daily
- 4–6 weeks, then a pause
Bench steps
- Let the vial sit until it is no longer cold to the touch.
- Wipe the stopper with 70% isopropyl alcohol. Let it dry.
- Draw 1 ml bacteriostatic water (0.9% benzyl alcohol).
- Run the water slowly down the inside glass — do not blast the cake.
- Roll between finger and thumb until the cake is gone. Do not shake.
- Label the date. Store the solution at 2–8 °C. Do not freeze. Use within 30 days unless the note below says otherwise.
A thousand micrograms, not milligrams. 50 mcg is 5 units. Over-mixing the cake with a large water volume makes the marks unreadable — 1 ml is the point.
NAD+
1000mg
Mix with 10 ml bacteriostatic water → 100 mg/ml
- Hypothetical aliquot
- 50–100 mg
- 0.50–1.00 ml · 50–100 units on a U-100 syringe
- How often
- Two or three times per week in published infusion and assay notes
- 4–8 weeks, then a pause
Bench steps
- Let the vial sit until it is no longer cold to the touch.
- Wipe the stopper with 70% isopropyl alcohol. Let it dry.
- Draw 10 ml bacteriostatic water (0.9% benzyl alcohol).
- Run the water slowly down the inside glass — do not blast the cake.
- Roll between finger and thumb until the cake is gone. Do not shake.
- Label the date. Store the solution at 2–8 °C. Do not freeze. Use within 30 days unless the note below says otherwise.
A 1000mg cake wants 10 ml. Protect from light. Solution yellows as it oxidises — that is the cofactor dying, not a flavour. Use promptly.
Bacteriostatic water and sterile syringes ship with peptide orders over £75. Kit details · 10 ml bacteriostatic water
The vials this essay sits on
Named sequences the essay maps — Epithalon, HGH, IGF-1 LR3, NAD+. Hypothetical research neighbourhood, not a protocol, not a medicine. One press puts every in-stock vial in the bag.
Research use only. Not a combined-use instruction.
Read next

65 min · long read · The living cell
The living cell is a city, and you are 36 trillion of them
A 70 kg adult is on the order of 36 trillion cells, most of them red blood cells with no nucleus. A typical nucleated cell holds about ten billion proteins and two metres of DNA. The body recycles 40–60 kg of ATP a day. Those figures are published; this essay is what they mean.

54 min · long read · The living cell
The nucleus, telomeres, and the four residues of Epithalon
Two metres of DNA folded into a nucleus a few micrometres across. TERT is off in most somatic cells on purpose. Epithalon is four residues, Ala-Glu-Asp-Gly, with a TERT and pineal literature. The machines are real. A large Western trial of telomere length in adults is not.

82 min · long read · The living cell
Pathophysiology from the genome to a person who notices
Disease is a stack: genome, transcriptome, proteome, metabolome, organelle, cell fate, tissue, organism. A peptide binds one node. The rest of the stack keeps running. CFTR, type 2 diabetes and a tendon as worked examples.

48 min · long read · Peptide research
Somatropin: the 191-residue ligand
Recombinant human growth hormone is one of the most studied proteins in endocrinology. Pulsatility, GHR–JAK2–STAT5b, lipolysis, IGF-1 generation — the map is public.
More in this desk

70 min · long read · The living cell
How peptides talk to cells: occupancy, amplification, arrestin
A peptide is a ligand. Most of the catalogue binds a GPCR on the cell surface: one occupancy, then enzymes make thousands of second messengers. That amplification is real, and it is not magic. Desensitisation is why more ligand is not more signal forever.

64 min · long read · The living cell
Mitochondria: the bacterium you kept, the genome it kept, the peptides it writes
You turn over 40–60 kg of ATP a day using a 16,569-base genome that still uses a bacterial genetic code. NAD+ is the hydride carrier Complex I spends. MOTS-c is a 16-mer translated from mitochondrial 12S rRNA — Lee, Kim, Cohen, 2015. That last sentence is real, and it is surprising.

57 min · long read · The living cell
Proteostasis: the cell that eats its own mistakes
Ten billion proteins, a 76-residue tag, a 2.5-megadalton proteasome, and autophagy for whole organelles. Hershko, Ciechanover and Rose, 2004. Ohsumi, 2016. DSIP is a sleep-isolation nonapeptide, not an autophagy ligand.

51 min · long read · The living cell
Membranes: a five-nanometre wall
Every cell, every organelle, every synapse sits on a lipid bilayer about five nanometres thick. Singer and Nicolson, Hodgkin and Huxley, a billion lipids, and why a peptide ligand at a GPCR does not need to enter the cell.
Essays describe published research. They are not medical advice and they do not authorise human use of any catalogue item.



