Chapter 18은 DNA sequencing을 기술로드맵의 case study로 다룬다. 출발점은 A·T·G·C라는 biological information이지만, 성능곡선을 실제로 바꾼 것은 chemistry 하나가 아니었다. Chain termination, fluorescent detection, capillary arrays, massively parallel reads, PCR/sample preparation, semiconductor sensing, precision optics, nanotechnology, computational assembly, storage와 workflow가 서로 연결되면서 read length·accuracy·throughput·run time·cost라는 Figures of Merit가 함께 이동했다. DNA sequencing의 핵심 사례는 단일 invention보다 complementary technologies의 network가 technology S-curve와 시장을 동시에 바꾼다는 것이다.
DNA as an Information Substrate — 생명정보는 네 종류의 nucleotide와 상보적 pairing 규칙 위에 놓인다
Sequencing이 가능한 이유는 DNA가 단순히 복잡한 분자가 아니라 복제 가능한 symbolic sequence이기 때문이다.
DNA의 구조
DNA는 deoxyribonucleic acid이며 known life의 building blocks와 operating procedures를 long double-helix molecules 형태로 encode하는 nucleic-acid family라고 chapter는 정의한다. Source는 Watson과 Crick의 1953 Nature paper를 double-helix geometry의 첫 설명으로 소개하고 1962 Nobel Prize를 언급한다.
DNA와 RNA는 nucleic acids이며 proteins, lipids, complex carbohydrates와 함께 known life에 필수적인 macromolecule classes로 제시된다. 각 DNA strand는 polynucleotide이고 nucleotide는 cytosine(C), guanine(G), adenine(A), thymine(T) 가운데 하나의 nucleobase, deoxyribose sugar, phosphate group으로 구성된다. Nucleotides는 sugar와 phosphate 사이 covalent bonds로 연결되어 alternating sugar-phosphate backbone을 만든다.
두 strand의 bases는 A–T, C–G pairing rule과 hydrogen bonds로 double-stranded DNA를 구성한다. 한 strand만 알면 complementary strand를 infer할 수 있다는 성질이 mitosis와 sequencing에 핵심이라는 것이 source의 설명이다.
Mendel: sequence 이전에 inheritance rule이 있었다
Gregor Mendel의 pea-plant hybridization experiments는 1856–1863에 수행되었고 dominant/recessive genes를 포함한 Mendelian inheritance의 규칙을 정립했다고 chapter는 설명한다. Fig. 18.2는 parental generation에서 시작해 다음 generations으로 trait가 전달되는 모습을 보여주며, phenotype에서 보이지 않던 recessive trait가 genetic inheritance를 통해 뒤에서 다시 나타날 수 있음을 시각화한다.
이 절의 역할은 sequencing 기술 이전에 “어떤 정보가 세대를 넘어 보존되는가”라는 biological problem이 있었다는 점을 연결하는 데 있다. Sequencing은 이 inheritance substrate를 직접 읽는 기술로 이어진다.
From Genes to Genomes — sequencing target은 짧은 gene에서 3-billion-base-pair human genome으로 확장된다
Sequencing이 읽는 범위
DNA sequencing은 individual genes, larger genetic regions, full chromosomes, whole genomes까지 읽을 수 있으며 open reading frames를 통해 RNA와 proteins를 간접적으로 sequence하는 데도 이용된다고 source는 설명한다. 응용분야로 biology, medicine, forensics, anthropology가 제시된다.
Milestones
Source는 bacteriophage φX174의 complete genome이 1977 처음 sequenced되었고, 1984 Medical Research Council scientists가 Epstein–Barr virus 172,282 nucleotides를 prior genetic profile 없이 decipher한 것을 turning point로 설명한다.
Human Genome Project는 main text에서 1990 launch, April 2003 accomplished로 서술된다. Human body는 약 10 trillion cells, 23 chromosome pairs, 약 3 billion DNA base pairs를 가지며 source는 이 information을 약 1000 large textbooks 또는 약 3 gigabits에 비유한다.
또한 source는 약 100 million base pairs, 약 3%,를 “active coding regions”로, 나머지를 “non-coding DNA” 또는 “junk DNA”라고 표현하면서 evolutionary origin과 functional significance가 active research라고 덧붙인다. 이 표현과 비율은 원서의 terminology와 framing을 그대로 요약한 것이며 이 글에서 현대 genomics consensus로 재검증하거나 수정하지 않았다.
Main text는 HGP를 1990 시작·2003 완료로 서술하지만, 뒤의 footnote는 “Human Genome Project launched by Craig Venter in the early 2000s”라고 적고, 또 다른 문장은 mass-production Sanger가 “first human genome in 2001”을 만들었다고 서술한다. 이 세 표현은 같은 역사적 사건을 서로 다른 방식으로 기술하므로 본문에서는 임의로 하나로 통합하지 않고 source 내부의 chronology 차이로 남겨 둔다.
The Sanger System — chemistry가 발명되고 fluorescence·capillary·arrays·software가 그 발명을 production technology로 바꾼다
Maxam–Gilbert sequencing
Allan Maxam과 Walter Gilbert는 1977 chemical modification와 base-specific cleavage를 사용하는 method를 발표했다. Purified double-stranded DNA를 cloning 없이 사용할 수 있었지만 radioactive labeling과 technical complexity가 extensive use를 제한했고 Sanger method가 refined되면서 주류에서 밀려났다고 source는 설명한다.
Process는 DNA 한쪽 end의 radioactive labeling과 purification, G / A+G / C / C+T 네 reactions에서 controlled chemical cleavage, denaturing acrylamide gel electrophoresis에 의한 size separation, X-ray-film autoradiography로 이어진다. Dark bands에 대응하는 fragment lengths를 분석해 sequence를 infer한다.
Sanger chain termination
Frederick Sanger와 coworkers가 1977 개발한 chain-termination method는 relative ease와 reliability 때문에 method of choice가 되었고, 초기부터 Maxam–Gilbert보다 toxic chemicals와 radioactivity 사용이 적었다고 source는 설명한다.
핵심은 ddNTPs(ddGTP, ddATP, ddTTP, ddCTP)이다. Dideoxynucleotide에는 3′-hydroxyl group이 없어 DNA polymerase가 다음 phosphodiester bond를 만들 수 없으므로 그 지점에서 chain elongation이 멈춘다. 서로 다른 종료 길이와 fluorescence를 읽으면 A·T·G·C sequence를 reconstruct할 수 있다.
Conceptual redraw based on Fig. 18.4 and the surrounding supported explanation; not a reproduction of the source figure.
Page 525의 Sanger chemistry paragraph에는 “ionization beam”, “thermolazer”, “CHRISPRE”라는 표현이 실제 인쇄본에 들어 있다. 바로 뒤 page와 Fig. 18.4는 ddNTP의 3′-OH 부재에 의한 chain termination, heat denaturation, fluorescence readout을 설명한다. 이 글은 anomalous terminology를 정상적인 sequencing step처럼 재해석하거나 조용히 고치지 않고, source anomaly로 명시한 뒤 주변에서 명확히 지지되는 mechanism만 기술한다.
Automation ladder
Sanger는 1980s부터 mid-2000s까지 prevailing method였으며 fluorescent labeling, capillary electrophoresis, automation이 efficiency와 cost를 지속적으로 개선했다. 그러나 later in the decade radically different architectures가 market에 들어오며 cost curve의 slope를 다시 바꾼다.
Next-Generation Sequencing — 한 개의 긴 sequence를 읽는 문제를 수많은 fragments를 병렬로 읽고 computationally assemble하는 문제로 바꾼다
Parallel sequencing and assembly
High-throughput 또는 next-generation sequencing은 genome sequencing/resequencing, RNA-Seq, ChIP-sequencing, epigenome characterization에 적용된다. 한 individual의 genome만으로 species의 variation을 대표할 수 없기 때문에 resequencing이 필요하다고 source는 설명한다.
Cost reduction을 이끈 architecture 변화는 thousands 또는 millions of sequences를 concurrently 생산하고 shorter reads를 overlap에 따라 assemble하는 parallelization이다. Fig. 18.5는 fragment의 beginning/end에 known subsequence를 두고 overlap으로 reference sequence를 조립하는 개념을 제시하며, 30× coverage의 single human genome data를 약 90–110 Gb로 표시한다. Source는 ultra-high-throughput에서 최대 500,000 sequencing operations를 parallel하게 실행할 수 있고 whole human genome을 as little as one day에 sequence할 수 있었다고 서술한다. 2006–2012 시기에는 2D surface parallelization이 중요한 architecture로 등장한다.
Table 18.1 — sequencing technology FOM landscape
Chapter가 제시하는 common Figures of Merit는 read length, single-read accuracy, reads per run, time per run, cost per one million bases이다. 아래 값은 source table의 snapshot을 그대로 옮긴 요약이다.
| METHOD | READ LENGTH | SINGLE-READ ACCURACY | READS / RUN | TIME / RUN | COST / 1M BASES |
|---|---|---|---|---|---|
| Single-molecule real-time · Pacific Biosciences | 30,000 bp N50; max >100,000 bases | 87% raw-read | 500,000 / Sequel SMRT cell; 10–20 Gb | 30 min–20 h | $0.05–0.08 |
| Ion semiconductor · Ion Torrent | Up to 600 bp | 99.6% | Up to 80 million | 2 h | $1 |
| Pyrosequencing · 454 | 700 bp | 99.9% | 1 million | 24 h | $10 |
| Sequencing by synthesis · Illumina | MiniSeq/NextSeq 75–300; MiSeq 50–600; HiSeq2500 50–500; HiSeq3/4000 50–300; HiSeq X 300 bp | 99.9% · Phred30 | MiniSeq/MiSeq 1–25M; NextSeq “130-00 Million” as printed; HiSeq2500 300M–2B; HiSeq3/4000 2.5B; HiSeq X 3B | 1–11 days, instrument/read-length dependent | $0.05–0.15 |
| cPAS · BGI/MGI | BGISEQ-50 35–50; MGISEQ200 50–200; BGISEQ500/MGISEQ2000 50–300 bp | 99.9% · Phred30 | BGISEQ50 160M; MGISEQ200 300M; BGISEQ500 1300M/flow cell; MGISEQ2000 375M FCS / 1500M FCL | 1–9 days | $0.035–0.12 |
| Sequencing by ligation · SOLiD | 50+35 or 50+50 bp | 99.9% | 1.2–1.4 billion | 1–2 weeks | $0.13 |
| Nanopore | Library-dependent; up to 2,272,580 bp reported | ~92–97% single read | User-selected read-length dependent | Real-time streaming; 1 min–48 h | $500–999 / flow cell; per-base cost experiment-dependent |
| Chain termination · Sanger | 400–900 bp | 99.9% | N/A | 20 min–3 h | $2400 |
Illumina NextSeq “Reads per run” cell은 인쇄된 Table 18.1 자체가 “130-00 Million”으로 보인다. 의미를 추정해 숫자를 보정하지 않고 source text를 그대로 표시했다. 또한 chapter는 Sanger가 sequencing-by-synthesis 등에 read length와 cost 측면에서 outpaced되었다고 서술하지만, 같은 table의 read length는 platform마다 정의와 범위가 달라 단일 scalar ranking으로 읽기 어렵다.
Single-read accuracy와 whole-genome accuracy를 구분한다
Footnote는 300–600 bp 정도의 single DNA fragment accuracy와 whole gene/chromosome/genome accuracy를 분리해야 한다고 강조한다. Repetition과 statistical analysis, 즉 coverage와 consensus를 이용하면 current technologies에서 whole-DNA reading accuracy가 >99.9%에 접근할 수 있다는 source의 설명이다. Sequencing system의 accuracy FOM은 read-level sensor 성능만으로 정의되지 않고 redundancy와 assembly까지 포함한다.
Cost Collapse — $100M genome에서 ~$1K genome까지의 곡선은 architecture switching과 complementary technology의 누적 효과를 보여준다
2001 → 2011 → late 2010s
Source는 Sanger mass-production era 이후 radically different approaches가 등장해 $100 million/genome in 2001 → $10,000 in 2011로 내려갔다고 설명한다. Fig. 18.6은 2001–2019 full-human-genome cost curve를 보여주며 late 2010s에는 약 $1,000 부근에서 flattening되는 모습을 담는다.
Figure caption은 progress를 “about five orders of magnitude”, 동시에 “a factor of 10,000”, “about 90% annual improvement”라고 서술한다. Five orders와 factor 10,000은 수학적으로 같은 표현이 아니므로 이 글은 하나로 정규화하지 않는다. 기술 trend를 사용할 때는 source가 제시한 slope와 endpoint뿐 아니라 metric definition과 internal consistency도 함께 검토해야 한다.
Why could sequencing improve so fast?
Chapter가 제시하는 답은 technology network다. Controlled polymerase chain reactions, solid-state semiconductors, precision optics, nanotechnology와 zero-mode optical waveguides 등이 precision·throughput·price의 FOM targets에 동시에 도달하면서 Ion Torrent, Pacific Biosciences 같은 new sequencing architectures가 가능해졌다.
As of 2020의 source landscape는 Illumina, Qiagen, Thermo Fisher Scientific 등을 high-throughput product leaders로 언급한다. Early sequencing이 nonprofit research labs 중심이고 standardization이 약했다면, later sequencing와 gene-editing technology는 deliberate roadmapping과 standardization의 영향을 더 많이 받게 되었다고 chapter는 설명한다.
Discussion & Exercise 18.1
New Markets Need New Workflows — sequencing machine의 가격이 내려가면 bottleneck은 sample flow·storage·retrieval·quality control과 해석으로 이동한다
Product families and commoditization
Source는 about 2016 이후 full human genome sequencing cost가 ~$1,000까지 내려가 individual DNA tests가 가능해졌다고 설명한다. Illumina는 single-purpose sequencer가 아니라 different research, medical, forensic needs에 대응하는 product family를 제공하는 example로 제시된다. Fig. 18.7에는 iSeq, MiniSeq, MiSeq, NextSeq, HiSeq, HiSeq X, NovaSeq가 나열된다.
Figure note는 used HiSeq machines가 source-era market에서 $65,000 or less에 available한 사례를 들며 capital cost 감소와 commoditization을 강조한다. 이 가격은 원서 시점의 example이며 현재 market price로 업데이트하지 않았다.
Sequencing value requires an information factory
저렴한 reads만으로 value가 생기지 않는다. Efficient workflows와 DNA sequences를 store/retrieve하는 IT infrastructure가 필요하며 Broad Institute에서는 industrial-engineering principles인 flow control, work-in-progress monitoring, quality control을 scale에 적용한 사례가 언급된다. Data-generation cost가 내려갈수록 operational excellence와 data infrastructure가 technology value의 더 큰 부분을 차지한다.
Application expansion
Source가 나열하는 applications는 cancer screening, immunology, gene therapy, cellular circuitry, epigenomics이다. Section title에 gene therapy가 포함되지만 chapter는 gene-therapy mechanism이나 clinical protocol을 자세히 전개하지 않는다. 따라서 이 글도 그 공백을 outside knowledge로 채우지 않는다.
Ancestry.com, 23andMe 같은 services는 source 시점에 under $100 consumer genetic testing을 제공하는 example로 제시된다. Primary market는 genealogy이며 specific disease biomarkers는 additional fee로 검사할 수 있다고 chapter는 서술한다. Ancestry application은 client DNA를 geographically tagged anchor populations와 비교해 fractional attribution을 추정하고, 보통 whole genome이 아니라 fraction만 sequence한다고 설명한다.
Population genomics and microbiome
Source는 2010 1000 Genomes Project와 diverse population characterization, microbiome characterization을 future demand drivers로 든다. Human body가 on the order of 10,000 other organisms를 host하고 이들의 DNA가 약 50–60 billion base pairs, human DNA의 약 20×라고 서술한다. 이를 위해 sequencing capability를 >1000× 확장하고 technology를 miniaturize해 field sequencing까지 가능하게 해야 한다고 chapter는 전망한다.
Beyond the Genome — capacity growth가 Moore’s Law보다 빠를 때 핵심 질문은 “얼마나 읽을 수 있는가”에서 “무엇을 읽고, 저장하고, 해석하고, 다시 쓸 것인가”로 이동한다
Fig. 18.8: sequencing capacity as a scaling system
Fig. 18.8은 Stephens et al. 2015의 growth graphic을 사용해 cumulative human genomes와 worldwide annual sequencing capacity를 함께 배치한다. Source figure에는 historical growth rate doubling every 7 months, Illumina estimate every 12 months, Moore’s Law reference every 18 months가 표시된다.
Milestones로 Human Genome Project(figure label 2001, ~3 billion bp), first personal genome, 1000 Genomes Project(2010, ~3,000 billion bp), Microbiome Project(2012, ~50 billion bp), environmental sequencing(>100,000 organisms, >>500 billion bp), ExAC, TCGA, early 454/Illumina/PacBio 등이 배치된다. Sequencing은 detector technology뿐 아니라 data-volume growth 자체가 roadmap driver가 되는 information infrastructure로 변화한다.
Sequencing is not gene editing
Chapter는 이 case study가 DNA sequencing만 다루며 CRISPR 같은 gene-editing technologies는 다루지 않는다고 경계를 명시한다. Fig. 18.6에 saturation signs가 보여도 source는 DNA technologies가 future decades에 계속 rapidly progress할 것으로 전망하고 biology를 technological evolution의 next frontier 후보로 본다.
Bioeconomy and DNA as storage
Source는 U.S. DNA/biology-related technologies와 industries가 2 million direct jobs, 8 million indirect jobs, 약 $2 trillion/year GDP output을 차지한다고 서술한다. 이는 원서 집필시점의 aggregate claim으로, 본 글은 현재 통계로 재검증하지 않았다.
또 다른 frontier는 DNA를 information medium으로 사용하는 것이다. Source는 “all of the 25 zettabytes of information created by humans on Earth today”가 one tube of DNA에 들어갈 수 있다고 예시하며, 이를 위해 DNA를 빠르고 정확하게 read할 뿐 아니라 write할 수 있어야 한다고 설명하고 Nicol et al. 2017 patent/application을 reference로 든다.
AI·Scientific Foundation Model 관점의 해석
Inference: Chapter 18을 AI4Science 관점에서 읽으면 sequencing은 “instrument → data → model” pipeline의 전형적인 사례가 된다. Data-generation FOM이 빠르게 개선될수록 bottleneck은 raw sequence acquisition에서 storage, provenance, quality control, phenotype/context linkage, scalable analysis와 hypothesis generation으로 이동한다.
이 구조는 genome foundation models나 multi-agent scientific systems를 설계할 때도 중요한 힌트를 준다. 모델이 더 커지는 것만으로는 sequencing technology value를 완성하지 못하며 sample provenance, coverage/accuracy semantics, cohort diversity, assay workflow, downstream validation을 하나의 evidence pipeline으로 묶어야 한다. 이는 Chapter 18의 technology-system logic을 현대 AI research에 확장한 inference이며 원서가 foundation models나 agentic AI를 실증한 결과는 아니다.
Source integrity checklist
- Chapter의 large fraction이 Wikipedia DNA sequencing article을 기반으로 한다고 source가 직접 밝힌다.
- HGP chronology 관련 표현은 main text와 footnote 사이에 차이가 있어 그대로 표시했다.
- Sanger chemistry paragraph의 “ionization beam / thermolazer / CHRISPRE”는 인쇄본에 존재하는 anomalous passage로 별도 표시했다.
- Fig. 18.6 caption의 “five orders”와 “factor 10,000”은 동치가 아니므로 임의로 통일하지 않았다.
- Table 18.1의 NextSeq “130-00 Million” 등 인쇄상 불완전해 보이는 값은 추정 보정하지 않았다.
- Vendor, price, jobs, GDP, storage-volume claims는 source-era로 표기했고 현재값으로 대체하지 않았다.