Gene Structure and Function
A gene is not simply “a piece of DNA.” A typical human protein-coding gene is an organised sequence of regulatory and transcribed regions arranged along DNA. Its structure determines where transcription begins, which RNA is produced, how that RNA is processed, where translation begins and ends, and how much functional protein is made. Understanding this arrangement is the foundation for interpreting variants and explaining inherited disease.
The gene map: read it from 5′ to 3′
A typical eukaryotic protein-coding gene
| 5′ side of DNA | Region in order | What it controls or produces |
|---|---|---|
| Upstream | Enhancers and silencers | How strongly, where and when the gene is expressed |
| Upstream | Promoter and transcription-start site | Recruitment of transcription factors and RNA polymerase II; beginning of RNA synthesis |
| Transcribed | 5′ untranslated region | mRNA stability and efficiency of ribosome recruitment |
| Translated | Start codon | Defines the beginning of the open reading frame |
| Translated | Coding exons separated by introns in the primary transcript | Sequence that determines the amino-acid order after splicing |
| Translated | Stop codon | Signals termination of translation |
| Transcribed | 3′ untranslated region | mRNA stability, localisation and post-transcriptional regulation |
| Downstream | Polyadenylation and termination region | Cleavage, poly-A addition and termination of transcription |
Important: the DNA gene contains exons and introns, but the mature mRNA does not normally contain the introns. The “coding sequence” is only the part translated into amino acids; the whole gene is larger because it also includes regulatory and untranslated regions.
1. DNA is the physical information store
DNA is a double helix made from two antiparallel nucleotide strands. Each nucleotide contains phosphate, deoxyribose sugar and a nitrogenous base. Complementary pairing—A with T and C with G—allows information to be copied and repaired. The order of bases, not merely the amount of DNA, carries the instruction.
A chromosome contains one long DNA molecule packaged around histone proteins. A locus is the physical position of a gene. An allele is one version of a gene. A genome includes coding genes, regulatory sequences, non-coding RNA genes, repetitive DNA, centromeres and telomeres. Therefore, a gene must be understood within chromatin and chromosome organisation.
Clinical connection
When a laboratory reports a variant, the doctor must ask: Which gene? Which transcript? Which genomic position? Which allele? Is it coding, regulatory or splice-related? What is the zygosity? How does it fit the phenotype and family history?
2. Regulatory DNA: deciding when a gene is used
Enhancers
An enhancer is a regulatory DNA element that can increase transcription when activator proteins bind to it. Enhancers may lie thousands of bases away, upstream, downstream or within an intron. DNA looping brings the enhancer-bound proteins into contact with the promoter through mediator and chromatin-associated proteins.
Silencers and repressors
Silencers reduce transcription when repressor proteins bind. Repressors may block transcription-factor binding, recruit enzymes that compact chromatin, or interfere with mediator and RNA polymerase. A pathogenic regulatory variant can therefore lower gene expression even though every coding exon is normal.
Insulators and boundaries
Insulators help separate neighbouring regulatory domains. They prevent an enhancer from activating the wrong promoter and help maintain tissue-specific expression. Disruption of a boundary can cause a normal gene to be expressed in the wrong place or at the wrong time.
3. The promoter and transcription-start site
The promoter is the landing and assembly region for transcription. General transcription factors recognise promoter features and position RNA polymerase II. The transcription-start site is the nucleotide at which the initial RNA transcript begins. The promoter is not the same as the start codon: transcription can begin before translation begins.
Promoter activity is tissue-specific. A liver cell, neuron and erythroid cell may use different transcription factors on the same genome. Promoter methylation or a promoter sequence variant can prevent the correct amount of RNA from being produced.
4. The 5′ untranslated region and start codon
The first part of the transcript is the 5′ UTR. It is transcribed into RNA but not translated into protein. It can contain secondary structures and regulatory signals that affect ribosome scanning, translation initiation and mRNA stability. The start codon, usually AUG in mRNA, establishes the reading frame for the open reading frame.
A mutation before the start codon may reduce translation. A new upstream start codon may divert the ribosome. A change in the start codon may abolish normal initiation or cause use of an alternative start site, producing a shortened or abnormal protein.
5. Exons: what remains in mature mRNA
Exons are the segments retained after RNA splicing. Some exon sequence is untranslated, while coding exons contribute codons to the protein. Exons may encode catalytic domains, membrane-spanning regions, signal peptides, ligand-binding sites or protein-interaction domains.
Exon boundaries matter because a variant can remove an entire exon, change the reading frame or alter a domain. Alternative exon choice also allows one gene to make different proteins in different tissues.
6. Introns: removed sequences with important functions
Introns are transcribed into the initial pre-mRNA but removed during splicing. They are not simply useless gaps. Introns can contain enhancers, regulatory sequences, non-coding RNA genes and alternative exons. Their length and sequence can affect transcription and splicing efficiency.
Splicing depends on the 5′ donor site, branch point, polypyrimidine tract and 3′ acceptor site. A variant near any of these signals may cause exon skipping, intron retention or activation of a cryptic splice site. The mature mRNA may then encode an abnormal protein or be destroyed by nonsense-mediated decay.
How a splice variant causes disease
- DNA change weakens the normal splice signal.
- The spliceosome chooses an abnormal site.
- An exon is skipped or intronic sequence is retained.
- The reading frame may shift or a premature stop may appear.
- Abnormal RNA may be degraded or translated into a dysfunctional protein.
7. The coding sequence and open reading frame
The coding sequence is read in groups of three bases called codons. Each codon specifies an amino acid or a stop signal. The open reading frame begins at the start codon and continues in one fixed frame until a stop codon. Changing the frame changes every downstream codon.
| Change in DNA | What happens to the reading frame | Typical effect |
|---|---|---|
| One base substitution | Frame usually preserved | Synonymous, missense or nonsense change |
| Insertion/deletion of 1 or 2 bases | Frame shifted | Abnormal downstream amino acids and early termination |
| Insertion/deletion of 3 bases | Frame preserved | One amino acid added or removed |
| Splice disruption | Exon structure altered | Abnormal transcript and possible frameshift |
8. Stop codon, 3′ UTR and polyadenylation
The stop codon ends translation; it does not necessarily end transcription immediately. After the coding sequence, the 3′ UTR regulates mRNA stability, localisation and translation. MicroRNAs and RNA-binding proteins can bind the 3′ UTR and influence how long the message survives and how efficiently it is translated.
A polyadenylation signal directs cleavage of the transcript and addition of a poly-A tail. The tail protects mRNA from degradation and supports nuclear export and translation. A variant in the polyadenylation region can produce an unstable, mislocalised or improperly terminated transcript.
9. From gene to pre-mRNA
After promoter activation, RNA polymerase II moves along the DNA template and creates a complementary primary transcript. The template strand is read 3′ to 5′ while RNA is synthesised 5′ to 3′. The primary transcript contains the information from exons and introns and is called pre-mRNA.
The pre-mRNA receives a 5′ cap, undergoes intron removal and exon joining, and receives a poly-A tail. These processing steps are part of gene function; a gene can be transcribed normally but still fail because its RNA is processed incorrectly.
10. Mature mRNA and translation
Mature mRNA leaves the nucleus and binds a ribosome. The ribosome reads codons from 5′ to 3′. tRNA molecules bring amino acids using complementary anticodons. Translation proceeds through initiation, elongation and termination. The amino-acid sequence then folds into a functional structure.
The product may be a membrane receptor, enzyme, hormone, channel, structural protein, transcription factor or functional RNA. A gene does not always produce a protein; genes for rRNA, tRNA, microRNA and other non-coding RNAs act directly as RNA molecules.
11. Protein targeting and post-translational function
Signal peptides direct some proteins into the endoplasmic reticulum for secretion or membrane insertion. Other targeting sequences send proteins to mitochondria, lysosomes, peroxisomes or the nucleus. Proteins may require cleavage, phosphorylation, glycosylation, disulfide bonding or assembly with other subunits.
This explains why a variant outside the catalytic site may still cause disease. It may prevent folding, trafficking, cleavage, membrane insertion, interaction with a partner, or removal of a damaged protein.
12. Why the same gene behaves differently in different cells
Most nucleated cells carry the same genome but express different genes. Cell-specific transcription factors, chromatin accessibility, DNA methylation, histone modification and non-coding RNAs determine which genes are active. This is why a variant may affect one organ most severely even though the gene is present throughout the body.
13. Prokaryotic versus eukaryotic gene structure
Eukaryotic genes commonly contain promoters, enhancers, introns, exons, UTRs and separate transcription and translation compartments. Prokaryotic genes often lack introns, may be arranged in operons, and can couple transcription to translation because there is no nuclear membrane. The clinical genetics curriculum focuses mainly on human eukaryotic genes, but the comparison clarifies why bacterial gene regulation and antibiotic targets differ.
14. Gene variants mapped onto gene structure
Promoter variant
Changes transcription-factor binding and usually alters the amount of RNA.
Splice variant
Changes exon–intron processing and may create abnormal mRNA.
Missense variant
Changes one amino acid and may alter folding, binding or enzyme activity.
Frameshift variant
Changes the reading frame and commonly introduces premature termination.
3′ UTR variant
Changes microRNA or RNA-binding-protein regulation of transcript stability.
Copy-number variant
Changes the number of gene copies and can disturb dosage-sensitive pathways.
15. Clinical interpretation: what the doctor must ask
- Which region of the gene contains the variant?
- Is the region regulatory, untranslated, splice-related or coding?
- Does the variant change the reading frame or amino-acid sequence?
- Is the variant present in the correct transcript and genome build?
- Is the result germline or somatic?
- Does the patient’s phenotype fit the gene and mechanism?
- Is the classification pathogenic, likely pathogenic, uncertain, likely benign or benign?
- What action follows, and should relatives receive counselling or testing?
16. Worked example: a splice-site disease mechanism
Suppose a patient has a pathogenic change at the 5′ splice donor of a tumour-suppressor gene. The change may cause exon skipping. If the skipped exon is not a multiple of three bases, the reading frame shifts and a premature stop appears. The abnormal transcript may be degraded. The cell then loses tumour-suppressor protein, allowing accumulation of additional abnormalities.
The important reasoning chain is: DNA position → RNA processing → protein quantity or structure → cellular pathway → tissue phenotype → clinical disease.
Summary
The essence of gene structure is the ordered relationship between regulatory DNA, promoter, transcription start, 5′ UTR, start codon, exons, introns, stop codon, 3′ UTR and termination signals. Each region has a job. Each can be altered by disease-causing variation. A doctor who understands this map can explain how a laboratory result becomes a molecular mechanism and a patient’s phenotype.
References
- NCBI Bookshelf. Overview: Gene Structure.
- NCBI Bookshelf. From DNA to RNA.
- NCBI Bookshelf. Genetics 101.
- OpenStax Biology. Regulation of Gene Expression.
- SlideShare. Gene Structure and Function.
- SlideShare. Eukaryotic and Prokaryotic Gene Structure.
