Doctors Revision

Gene Structure and Function: From DNA Sequence to Clinical Phenotype

Gene Structure and Function

A gene is not simply “a piece of DNA.” A typical human protein-coding gene is an organised sequence of regulatory and transcribed regions arranged along DNA. Its structure determines where transcription begins, which RNA is produced, how that RNA is processed, where translation begins and ends, and how much functional protein is made. Understanding this arrangement is the foundation for interpreting variants and explaining inherited disease.

The central idea: DNA regulatory information controls transcription; the primary RNA transcript is processed into mature mRNA; the mRNA is translated into a polypeptide; and the polypeptide is folded, modified, transported and regulated to produce a phenotype.

The gene map: read it from 5′ to 3′

A typical eukaryotic protein-coding gene

5′ side of DNARegion in orderWhat it controls or produces
UpstreamEnhancers and silencersHow strongly, where and when the gene is expressed
UpstreamPromoter and transcription-start siteRecruitment of transcription factors and RNA polymerase II; beginning of RNA synthesis
Transcribed5′ untranslated regionmRNA stability and efficiency of ribosome recruitment
TranslatedStart codonDefines the beginning of the open reading frame
TranslatedCoding exons separated by introns in the primary transcriptSequence that determines the amino-acid order after splicing
TranslatedStop codonSignals termination of translation
Transcribed3′ untranslated regionmRNA stability, localisation and post-transcriptional regulation
DownstreamPolyadenylation and termination regionCleavage, poly-A addition and termination of transcription

Important: the DNA gene contains exons and introns, but the mature mRNA does not normally contain the introns. The “coding sequence” is only the part translated into amino acids; the whole gene is larger because it also includes regulatory and untranslated regions.

1. DNA is the physical information store

DNA is a double helix made from two antiparallel nucleotide strands. Each nucleotide contains phosphate, deoxyribose sugar and a nitrogenous base. Complementary pairing—A with T and C with G—allows information to be copied and repaired. The order of bases, not merely the amount of DNA, carries the instruction.

A chromosome contains one long DNA molecule packaged around histone proteins. A locus is the physical position of a gene. An allele is one version of a gene. A genome includes coding genes, regulatory sequences, non-coding RNA genes, repetitive DNA, centromeres and telomeres. Therefore, a gene must be understood within chromatin and chromosome organisation.

Clinical connection

When a laboratory reports a variant, the doctor must ask: Which gene? Which transcript? Which genomic position? Which allele? Is it coding, regulatory or splice-related? What is the zygosity? How does it fit the phenotype and family history?

2. Regulatory DNA: deciding when a gene is used

Enhancers

An enhancer is a regulatory DNA element that can increase transcription when activator proteins bind to it. Enhancers may lie thousands of bases away, upstream, downstream or within an intron. DNA looping brings the enhancer-bound proteins into contact with the promoter through mediator and chromatin-associated proteins.

Silencers and repressors

Silencers reduce transcription when repressor proteins bind. Repressors may block transcription-factor binding, recruit enzymes that compact chromatin, or interfere with mediator and RNA polymerase. A pathogenic regulatory variant can therefore lower gene expression even though every coding exon is normal.

Insulators and boundaries

Insulators help separate neighbouring regulatory domains. They prevent an enhancer from activating the wrong promoter and help maintain tissue-specific expression. Disruption of a boundary can cause a normal gene to be expressed in the wrong place or at the wrong time.

Do not call non-coding DNA “junk” automatically. A promoter, enhancer, silencer, splice element or untranslated region can be clinically important even when it does not encode amino acids.

3. The promoter and transcription-start site

The promoter is the landing and assembly region for transcription. General transcription factors recognise promoter features and position RNA polymerase II. The transcription-start site is the nucleotide at which the initial RNA transcript begins. The promoter is not the same as the start codon: transcription can begin before translation begins.

Promoter activity is tissue-specific. A liver cell, neuron and erythroid cell may use different transcription factors on the same genome. Promoter methylation or a promoter sequence variant can prevent the correct amount of RNA from being produced.

4. The 5′ untranslated region and start codon

The first part of the transcript is the 5′ UTR. It is transcribed into RNA but not translated into protein. It can contain secondary structures and regulatory signals that affect ribosome scanning, translation initiation and mRNA stability. The start codon, usually AUG in mRNA, establishes the reading frame for the open reading frame.

A mutation before the start codon may reduce translation. A new upstream start codon may divert the ribosome. A change in the start codon may abolish normal initiation or cause use of an alternative start site, producing a shortened or abnormal protein.

5. Exons: what remains in mature mRNA

Exons are the segments retained after RNA splicing. Some exon sequence is untranslated, while coding exons contribute codons to the protein. Exons may encode catalytic domains, membrane-spanning regions, signal peptides, ligand-binding sites or protein-interaction domains.

Exon boundaries matter because a variant can remove an entire exon, change the reading frame or alter a domain. Alternative exon choice also allows one gene to make different proteins in different tissues.

6. Introns: removed sequences with important functions

Introns are transcribed into the initial pre-mRNA but removed during splicing. They are not simply useless gaps. Introns can contain enhancers, regulatory sequences, non-coding RNA genes and alternative exons. Their length and sequence can affect transcription and splicing efficiency.

Splicing depends on the 5′ donor site, branch point, polypyrimidine tract and 3′ acceptor site. A variant near any of these signals may cause exon skipping, intron retention or activation of a cryptic splice site. The mature mRNA may then encode an abnormal protein or be destroyed by nonsense-mediated decay.

How a splice variant causes disease

  1. DNA change weakens the normal splice signal.
  2. The spliceosome chooses an abnormal site.
  3. An exon is skipped or intronic sequence is retained.
  4. The reading frame may shift or a premature stop may appear.
  5. Abnormal RNA may be degraded or translated into a dysfunctional protein.

7. The coding sequence and open reading frame

The coding sequence is read in groups of three bases called codons. Each codon specifies an amino acid or a stop signal. The open reading frame begins at the start codon and continues in one fixed frame until a stop codon. Changing the frame changes every downstream codon.

Change in DNAWhat happens to the reading frameTypical effect
One base substitutionFrame usually preservedSynonymous, missense or nonsense change
Insertion/deletion of 1 or 2 basesFrame shiftedAbnormal downstream amino acids and early termination
Insertion/deletion of 3 basesFrame preservedOne amino acid added or removed
Splice disruptionExon structure alteredAbnormal transcript and possible frameshift

8. Stop codon, 3′ UTR and polyadenylation

The stop codon ends translation; it does not necessarily end transcription immediately. After the coding sequence, the 3′ UTR regulates mRNA stability, localisation and translation. MicroRNAs and RNA-binding proteins can bind the 3′ UTR and influence how long the message survives and how efficiently it is translated.

A polyadenylation signal directs cleavage of the transcript and addition of a poly-A tail. The tail protects mRNA from degradation and supports nuclear export and translation. A variant in the polyadenylation region can produce an unstable, mislocalised or improperly terminated transcript.

9. From gene to pre-mRNA

After promoter activation, RNA polymerase II moves along the DNA template and creates a complementary primary transcript. The template strand is read 3′ to 5′ while RNA is synthesised 5′ to 3′. The primary transcript contains the information from exons and introns and is called pre-mRNA.

The pre-mRNA receives a 5′ cap, undergoes intron removal and exon joining, and receives a poly-A tail. These processing steps are part of gene function; a gene can be transcribed normally but still fail because its RNA is processed incorrectly.

10. Mature mRNA and translation

Mature mRNA leaves the nucleus and binds a ribosome. The ribosome reads codons from 5′ to 3′. tRNA molecules bring amino acids using complementary anticodons. Translation proceeds through initiation, elongation and termination. The amino-acid sequence then folds into a functional structure.

The product may be a membrane receptor, enzyme, hormone, channel, structural protein, transcription factor or functional RNA. A gene does not always produce a protein; genes for rRNA, tRNA, microRNA and other non-coding RNAs act directly as RNA molecules.

11. Protein targeting and post-translational function

Signal peptides direct some proteins into the endoplasmic reticulum for secretion or membrane insertion. Other targeting sequences send proteins to mitochondria, lysosomes, peroxisomes or the nucleus. Proteins may require cleavage, phosphorylation, glycosylation, disulfide bonding or assembly with other subunits.

This explains why a variant outside the catalytic site may still cause disease. It may prevent folding, trafficking, cleavage, membrane insertion, interaction with a partner, or removal of a damaged protein.

12. Why the same gene behaves differently in different cells

Most nucleated cells carry the same genome but express different genes. Cell-specific transcription factors, chromatin accessibility, DNA methylation, histone modification and non-coding RNAs determine which genes are active. This is why a variant may affect one organ most severely even though the gene is present throughout the body.

13. Prokaryotic versus eukaryotic gene structure

Eukaryotic genes commonly contain promoters, enhancers, introns, exons, UTRs and separate transcription and translation compartments. Prokaryotic genes often lack introns, may be arranged in operons, and can couple transcription to translation because there is no nuclear membrane. The clinical genetics curriculum focuses mainly on human eukaryotic genes, but the comparison clarifies why bacterial gene regulation and antibiotic targets differ.

14. Gene variants mapped onto gene structure

Promoter variant

Changes transcription-factor binding and usually alters the amount of RNA.

Splice variant

Changes exon–intron processing and may create abnormal mRNA.

Missense variant

Changes one amino acid and may alter folding, binding or enzyme activity.

Frameshift variant

Changes the reading frame and commonly introduces premature termination.

3′ UTR variant

Changes microRNA or RNA-binding-protein regulation of transcript stability.

Copy-number variant

Changes the number of gene copies and can disturb dosage-sensitive pathways.

15. Clinical interpretation: what the doctor must ask

  1. Which region of the gene contains the variant?
  2. Is the region regulatory, untranslated, splice-related or coding?
  3. Does the variant change the reading frame or amino-acid sequence?
  4. Is the variant present in the correct transcript and genome build?
  5. Is the result germline or somatic?
  6. Does the patient’s phenotype fit the gene and mechanism?
  7. Is the classification pathogenic, likely pathogenic, uncertain, likely benign or benign?
  8. What action follows, and should relatives receive counselling or testing?
Never diagnose from the word “variant” alone. A variant of uncertain significance should not be used to make irreversible treatment, reproductive or prophylactic decisions.

16. Worked example: a splice-site disease mechanism

Suppose a patient has a pathogenic change at the 5′ splice donor of a tumour-suppressor gene. The change may cause exon skipping. If the skipped exon is not a multiple of three bases, the reading frame shifts and a premature stop appears. The abnormal transcript may be degraded. The cell then loses tumour-suppressor protein, allowing accumulation of additional abnormalities.

The important reasoning chain is: DNA position → RNA processing → protein quantity or structure → cellular pathway → tissue phenotype → clinical disease.

Summary

The essence of gene structure is the ordered relationship between regulatory DNA, promoter, transcription start, 5′ UTR, start codon, exons, introns, stop codon, 3′ UTR and termination signals. Each region has a job. Each can be altered by disease-causing variation. A doctor who understands this map can explain how a laboratory result becomes a molecular mechanism and a patient’s phenotype.

References

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top