AI Research NotesChapter 9 · Effective Theory of the NTKCriticality · Universality · L/n · Gradient Stability
The Principles of Deep Learning Theory/Chapter 9/2026.08.22

깊어질수록 학습은
왜 흔들리는가

Criticality of the neural tangent kernel — universality, finite-width relevance, and stable gradient flow

Criticality and universal L over n scalingParallel and perpendicular susceptibilities are tuned to one. Scale-invariant and K-star-zero classes then exhibit different raw critical exponents but the same dimensionless L over n finite-width scaling. CRITICALITYχ∥ = 1χ⊥ = 1 SCALE-INVARIANTReLU · linearΘ∼ℓ · D,F∼ℓ² · A,B∼ℓ³ K* = 0tanh · sindifferent raw exponents DIMENSIONLESSL / nfinite-width cutoff TRAINstablescales different activations, different raw powers — the same dimensionless finite-width relevance
Chapter Thesis

8장은 NTK의 평균과 fluctuation, preactivation과의 cross-correlation이 층을 따라 어떻게 재귀되는지 계산했다. 9장은 그 긴 계산에 물리적 의미를 부여한다. 무엇이 깊어질수록 커지고, 무엇이 사라지며, 어떤 hyperparameter tuning이 그 흐름을 안정시킬 수 있는가.

결론은 두 갈래에서 한 점으로 모인다. ReLU와 linear가 속한 scale-invariant class와 tanh·sin이 속한 \(K^\star=0\) class는 raw critical exponent가 다르다. 그러나 NTK fluctuation과 NTK–preactivation cross-correlation을 적절히 무차원화하면 두 class 모두 같은 \(L/n\) scaling을 보인다. 깊이와 폭의 비가 finite-width 효과의 실질적인 cutoff가 되는 것이다.

Central Claim

criticality는 단순히 signal propagation을 안정화하는 초기화 규칙이 아니다. \(\chi_\parallel=1\)은 output scale과 parameter-local contribution의 폭주·소실을 막고, \(\chi_\perp=1\)은 chain-rule sensitivity와 layer-to-layer NTK contribution의 폭주·소실을 막는다. 9장은 이를 gradient dynamics의 언어로 다시 증명한다.

Universal finite-width scaling · Eq. 9.27
\[\boxed{p_A-2p_\Theta=p_B-2p_\Theta=p_D-p_\Theta-p_0=p_F-p_\Theta-p_0=-1}\quad\Longrightarrow\quad \text{normalized finite-width effects}\sim \frac{L}{n}.\]
\(\Theta\)frozen NTK. infinite-width limit의 deterministic NTK mean이며 leading training dynamics를 지배한다.
\(A,B\)NTK variance를 압축하는 finite-width tensors. normalized variance의 상대 크기를 결정한다.
\(D,F\)NTK와 preactivation 사이의 cross-correlation을 압축한다.
\(L/n\)서로 다른 universality class를 넘어 반복해서 나타나는 emergent finite-width scale이다.

이 글은 9장 전체인 §9.1–§9.4, 원서 pp. 227–246만을 대상으로 한다. 원문에는 이 장의 별도 Figure/Table이 없으므로 hero SVG와 비교표는 수식 관계를 설명하기 위한 새로운 개념 시각화다.

Part I · From §8 to §9

수식을 다 얻었으면,
이제 무엇이 중요한지 물어야 한다

Chapter 9 turns the RG recursions into statements about trainability.

8장에서 이미 kernel \(K\), four-point vertex \(V\), NTK mean, NTK–preactivation cross-correlation, NTK variance의 recursion은 얻었다. 첫 층 NTK는 deterministic했고, 두 번째 층부터 fluctuate했으며, 더 깊은 층에서는 fluctuation과 cross-correlation이 누적되었다. 9장의 목적은 식을 하나 더 만드는 데 있지 않다. 그 식들이 depth와 width의 함수로 어떤 hierarchy를 만드는지 밝히는 데 있다.

분석은 uniform hidden width \(n\), layer-independent initialization variances \(C_b,C_W\)를 기본으로 한다. 대신 training hyperparameter인 bias learning rate \(\lambda_b^{(\ell)}\)와 weight learning rate \(\lambda_W^{(\ell)}\)는 처음부터 layer dependence를 허용한다. 왜냐하면 activation universality class에 따라 균형 잡힌 learning-rate scaling이 달라지기 때문이다.

또 하나의 범위 제한이 있다. 9장은 leading \(1/n\) order의 single-input statistics에 집중한다. multi-input structure와 next-to-leading corrections는 이 장의 분석 대상이 아니다. 이 단순화 덕분에 깊이에 대한 scaling exponent를 선명하게 추출할 수 있다.

복잡한 이론에서 중요한 것은 변수의 개수가 아니다. 어떤 조합이 무차원이고, 깊이를 늘렸을 때 무엇이 커지는지 알아내는 일이다. 9장은 NTK를 그 기준으로 다시 정렬한다.
Part II · §9.1

frozen NTK와 agitated NTK를
같은 눈금으로 재야 한다

Mean, variance, and representation coupling become comparable only after dimensionless normalization.

frozen NTK: infinite width에서 얼어붙은 training geometry

finite-width NTK mean \(H^{(\ell)}\)를 \(1/n\)으로 전개하면 leading \(O(1)\) piece를 따로 떼어낼 수 있다. 책은 이를 \(\Theta^{(\ell)}\), 즉 frozen NTK라 부른다. 문헌에서 흔히 NTK라고 부르는 deterministic infinite-width kernel이 바로 이것이다.

Frozen NTK recursion · Eq. 9.6
\[\Theta^{(\ell+1)}=\lambda_b^{(\ell+1)}+\lambda_W^{(\ell+1)}g\!\left(K^{(\ell)}\right)+\chi_\perp\!\left(K^{(\ell)}\right)\Theta^{(\ell)}.\]

한 layer의 contribution은 bias term과 weight term으로 새로 더해지고, 이전 layer에서 올라온 NTK는 \(\chi_\perp\)만큼 곱해진다. 따라서 \(\chi_\perp\neq1\)이면 오래된 layer contribution은 깊이를 거치며 지수적으로 커지거나 사라질 수 있다.

9.1은 8장의 multi-input recursion을 leading single-input form으로 줄여 \(D,F,A,B\)의 recursion도 정리한다. 세부 계수는 길지만 구조는 명확하다.

TensorLeading recursion structure의미Initial condition
\(D\)\(\chi_\perp\chi_\parallel D\) + \(V,\Theta,\lambda_W\) termspreactivation–NTK cross-correlation의 same-pair channel0
\(F\)\(\chi_\parallel^2F\) + \(\Theta\) sourcecrossed neural-index channel의 cross-correlation0
\(B\)\(\chi_\perp^2B\) + \(\Theta^2\) sourcecrossed channel의 NTK variance0
\(A\)\(\chi_\perp^2A\) + \(V,D,\Theta,\lambda_W\) termssame-pair channel의 NTK variance0

첫 층 NTK가 deterministic하므로 \(A^{(1)}=B^{(1)}=D^{(1)}=F^{(1)}=0\)이다. 이후 layer에서 finite-width source가 이 tensors를 생성한다.

raw power가 아니라 무차원 비율을 본다

어떤 tensor가 \(\ell^2\) 또는 \(\ell^3\)로 자란다고 말하는 것만으로는 physics를 판단하기 어렵다. NTK variance는 frozen NTK의 제곱과 비교해야 하고, NTK–preactivation cross-correlation은 frozen NTK와 kernel의 곱과 비교해야 한다.

Dimensionless finite-width observables · Eqs. 9.23–9.24
\[\frac{A^{(\ell)}}{n\,[\Theta^{(\ell)}]^2},\quad \frac{B^{(\ell)}}{n\,[\Theta^{(\ell)}]^2},\quad \frac{D^{(\ell)}}{nK^{(\ell)}\Theta^{(\ell)}},\quad \frac{F^{(\ell)}}{nK^{(\ell)}\Theta^{(\ell)}}.\]

이 조합을 쓰면 activation별 raw exponent의 차이가 걷히고, 뒤에서 모두 \(\ell/n\)이라는 같은 scaling으로 모인다.

formal solution이 \(\chi_\perp\)의 역할을 드러낸다

Formal frozen-NTK solution · Eq. 9.28
\[\Theta^{(\ell)}=\sum_{m=1}^{\ell}\Bigl[\lambda_b^{(m)}+\lambda_W^{(m)}g(K^{(m-1)})\Bigr]\prod_{r=m}^{\ell-1}\chi_\perp^{(r)}.\]

한 층에서 만들어진 contribution이 output layer까지 가는 동안 perpendicular susceptibility가 계속 곱해진다. 이 곱은 두 input의 perpendicular perturbation이 깊이를 통과할 때 나타나는 바로 그 factor다. 5장에서 “가까운 입력의 차이를 보존하자”는 조건으로 등장했던 \(\chi_\perp=1\)이, 이제 “각 layer가 학습에 기여할 길을 보존하자”는 조건으로 재해석된다.

Part III · §9.2

ReLU 계열은 criticality에서
가장 단순한 power law를 만든다

Scale invariance turns susceptibilities into constants and makes every layer contribute uniformly.

scale-invariant universality class의 대표는 ReLU와 linear activation이다. piecewise-linear activation에서 \(A_2=(a_+^2+a_-^2)/2\), \(A_4=(a_+^4+a_-^4)/2\)를 두면 helper function과 susceptibilities가 단순해진다. \(g(K)=A_2K\), \(h(K)=0\), 그리고 \(\chi_\parallel=\chi_\perp=C_WA_2\)이다.

Critical initialization for the scale-invariant class · Eq. 9.40
\[C_b=0,\qquad C_W=\frac1{A_2}\qquad\Rightarrow\qquad \chi_\parallel=\chi_\perp=1.\]

criticality에서는 kernel이 깊이에 따라 고정되고 four-point vertex \(V^{(\ell)}\)는 선형으로 증가한다. layer-independent \(\lambda_b,\lambda_W\)를 두면 frozen NTK recursion은 단순 누적합이 된다.

Frozen NTK at criticality · Eqs. 9.43–9.44
\[\Theta^{(\ell+1)}=\Theta^{(\ell)}+\lambda_b+\lambda_WA_2K,\qquad \Theta^{(\ell)}=(\lambda_b+\lambda_WA_2K)\,\ell.\]

여기서 중요한 것은 선형 성장 자체보다 그 의미다. NTK가 이전 모든 layer contribution의 합이므로 linear growth는 각 층이 같은 order로 기여한다는 뜻이다. noncritical regime처럼 특정 층의 기여가 지수적으로 지배하거나 사라지지 않는다.

agitated NTK는 더 빠르게 자라지만, 상대 크기는 \(\ell/n\)이다

finite-width recursions을 풀면 \(D,F\)는 quadratic, \(A,B\)는 cubic으로 증가한다. critical exponents는

Scale-invariant critical exponents · Eq. 9.53
\[p_0=0,\quad p_\Theta=-1,\quad p_D=p_F=-2,\quad p_A=p_B=-3.\]

이다. 책의 convention \(O^{(\ell)}\sim \ell^{-p_O}\)를 기억하면 음수 exponent는 성장이다. raw하게 보면 fluctuation tensor가 mean보다 빠르게 커진다. 그러나 \(A/(n\Theta^2)\), \(D/(nK\Theta)\)처럼 normalize하면 모두 \(\ell/n\)으로 정리된다.

Part IV · §9.3

tanh와 sin은 같은 목적지로 가지만
learning-rate 지도가 다르다

The K-star-zero class reaches the same finite-width scaling through different depth powers.

두 번째 universality class는 nontrivial fixed point가 \(K^\star=0\)에 있는 smooth activation이다. 대표적으로 tanh와 sin이 포함된다. 조건은 activation이 origin에서 0이면서 slope는 0이 아닌 것이다.

Critical initialization for the K-star-zero class · Eqs. 9.55–9.56
\[\sigma(0)=0,\quad \sigma'(0)=\sigma_1\neq0,\qquad C_b=0,\quad C_W=\frac1{\sigma_1^2}.\]

이 class에서는 criticality가 fixed point 주변의 perturbative expansion으로 나타난다. \(\chi_\parallel\)과 \(\chi_\perp\)는 1에 접근하지만 정확히 layer-independent constant는 아니다. single-input kernel은 큰 깊이에서

Kernel asymptotics · Eq. 9.65
\[K^{(\ell)}\sim \frac{1}{(-a_1)\ell},\]

로 천천히 0에 접근하고, normalized four-point interaction은 역시 \(\ell/n\) scale을 만든다. perpendicular perturbation의 exponent는 \(p_\perp=b_1/a_1\)이며 tanh와 sin에서는 \(p_\perp=1\)이다.

naive한 layer-independent learning rate는 두 가지 불균형을 만든다

frozen NTK의 formal solution을 보면 bias source와 weight source가 서로 다른 \(\ell\)-dependence를 가진다. layer-independent learning rate를 그대로 쓰면 첫째, 깊어질수록 weight contribution이 bias contribution에 비해 polynomial하게 줄어든다. 둘째, \(p_\perp>0\)이면 output에 가까운 deeper layer contribution이 shallower layer를 지배한다.

이를 고치기 위해 9.3은 layer-wise scaling을 먼저 도입한다.

Layer balancing before global depth normalization · Eq. 9.70
\[\lambda_b^{(\ell)}=\frac{\widetilde\lambda_b}{\ell^{p_\perp}},\qquad \lambda_W^{(\ell)}=\frac{\widetilde\lambda_W}{\ell^{p_\perp-1}}.\]

이렇게 하면 bias와 weight contribution이 같은 parametric order를 갖고, frozen NTK의 exponent는

K-star-zero frozen-NTK exponent · Eq. 9.72
\[p_\Theta=p_\perp-1.\]

이 된다. tanh와 sin처럼 \(p_\perp=1\)이면 frozen NTK는 asymptotically constant order다.

agitated statistics의 raw exponent는 달라도 normalized physics는 같다

large-depth recursion을 풀면

K-star-zero critical exponents · §9.3 summary
\[p_D=p_F=p_\perp-1,\qquad p_A=p_B=2p_\perp-3,\qquad p_0=1.\]

을 얻는다. tanh·sin의 \(p_\perp=1\)을 넣으면 \(\Theta,D,F\)는 order-one이고 \(A,B\)는 linear하게 성장한다. 그러나 kernel 자체가 \(1/\ell\)로 작아진다는 점까지 함께 normalize하면 네 finite-width ratio는 다시 모두 \(\ell/n\)이다.

원문은 \(p_\perp=1\)인 tanh와 sin에서 single-input \(D\)와 \(A\)의 leading term이 weight learning rate에 독립이라는 흥미로운 특수성도 지적한다.

Part V · Scaling Law

activation이 달라도
finite width의 눈금은 \(L/n\)이다

Universality does not erase microscopic differences; it identifies the dimensionless ratio that survives them.

scale-invariant class와 \(K^\star=0\) class는 겉모습이 매우 다르다. 하나는 kernel이 고정되고 frozen NTK가 linear하게 자란다. 다른 하나는 kernel이 \(1/\ell\)로 줄고 learning rate도 layer별로 조정해야 한다. 그런데 무차원 finite-width observables를 계산하면 같은 법칙이 나온다.

Universality classKernelFrozen NTKRaw D,FRaw A,BNormalized effect
Scale-invariant\(K\sim \ell^0\)\(\Theta\sim\ell\)\(\sim\ell^2\)\(\sim\ell^3\)\(\ell/n\)
\(K^\star=0\)\(K\sim\ell^{-1}\)\(\Theta\sim\ell^{-(p_\perp-1)}\)\(\sim\ell^{-(p_\perp-1)}\)\(\sim\ell^{-(2p_\perp-3)}\)\(\ell/n\)

따라서 finite-width correction의 relevance는 activation의 미세한 모양보다 depth-to-width ratio에 의해 더 보편적으로 정리된다. \(L/n\ll1\)이면 infinite-width description이 좋은 leading approximation이 된다. 반대로 depth가 width에 비해 커질수록 NTK fluctuation과 representation coupling은 perturbatively 더 중요해진다.

무한폭은 기준점이고, 유한폭은 오차항이 아니다. \(L/n\)이 커지면 유한폭 correction이 학습의 구조 자체를 바꾸는 relevant interaction이 된다.이 해석은 Chapter 9의 leading-order large-width expansion 안에서 성립한다. \(L/n\)이 order one에 가까워지면 higher-order terms를 무시할 수 없으므로 이 장의 perturbative formula 자체가 충분하지 않을 수 있다.
Part VI · §9.4

exploding·vanishing gradient는
두 susceptibility의 문제로 다시 쓰인다

The backward instability and the forward kernel instability are two views of the same criticality problem.

전통적인 exploding/vanishing gradient 문제는 깊은 network에서 chain rule을 반복했을 때 matrix product가 지수적으로 커지거나 작아지는 현상으로 설명된다. 9.4는 gradient를 세 factor로 분해해 이 현상을 criticality와 직접 연결한다.

Gradient factorization · Eq. 9.87
\[\frac{d\mathcal L_A}{d\theta_\mu^{(\ell)}}=\sum_{\alpha,i_L,i_\ell}\underbrace{\varepsilon_{i_L;\alpha}}_{\text{error}}\underbrace{\frac{dz^{(L)}_{i_L;\alpha}}{dz^{(\ell)}_{i_\ell;\alpha}}}_{\text{chain rule}}\underbrace{\frac{dz^{(\ell)}_{i_\ell;\alpha}}{d\theta_\mu^{(\ell)}}}_{\text{local / trivial}}.\]

\(\chi_\parallel\): output scale과 local parameter signal을 동시에 지킨다

MSE에서는 error factor가 \(z^{(L)}-y\)다. kernel이 폭발하면 typical output과 error factor도 폭발한다. 이를 막으려면 \(\chi_\parallel\le1\)이 필요하다. 반대로 kernel이 지수적으로 사라지면 activation도 작아지고 weight에 대한 local derivative가 억제되어 deep-layer weights가 거의 update되지 않는다. 이를 피하려면 \(\chi_\parallel\ge1\)이어야 한다.

Parallel criticality
\[\chi_\parallel(K^\star)\le1\quad\text{and}\quad\chi_\parallel(K^\star)\ge1\quad\Longrightarrow\quad\boxed{\chi_\parallel(K^\star)=1}.\]

즉 forward kernel의 exploding/vanishing problem은 gradient problem의 일부로 그대로 나타난다.

\(\chi_\perp\): chain-rule path와 shallow-layer influence를 지킨다

chain-rule factor의 평균 제곱 contribution은 NTK forward recursion의 multiplicative factor와 동일한 \(\chi_\perp\)로 연결된다. \(\chi_\perp>1\)이면 deeper-layer NTK가 지수적으로 커져 training dynamics가 불안정해질 수 있다. \(\chi_\perp<1\)이면 shallow layer의 bias와 weight가 output-side NTK에 미치는 영향이 지수적으로 사라진다.

Perpendicular criticality · Eq. 9.93 interpretation
\[\mathbb E\left[\left(\frac{dz^{(\ell+1)}}{dz^{(\ell)}}\right)^2\right]=C_W\langle\sigma'(z)^2\rangle_{K^{(\ell)}}=\chi_\perp^{(\ell)}\quad\Longrightarrow\quad\boxed{\chi_\perp(K^\star)=1}.\]

5장에서 두 input의 차이를 보존하기 위해 얻었던 두 criticality condition이 9장에서는 gradient descent의 안정성과 layer contribution의 균형을 위해 다시 등장한다. 같은 수식이 forward signal propagation과 backward trainability를 함께 묶는다.

책은 이 effective-theory setting에서 criticality가 exploding/vanishing gradient 문제를 충분히 완화한다고 강하게 주장한다. 이를 arbitrary architecture와 optimizer에서 gradient clipping이 언제나 불필요하다는 일반 명제로 확대해서 읽을 수는 없다. 이 장의 직접 계산 대상은 critical initialization을 갖는 deep MLP의 leading large-width regime이다.

Part VII · Equivalence Principle

criticality가 지수적 불균형을 막는다면,
equivalence는 polynomial 불균형까지 막는다

Training hyperparameters should make parameter groups and layers contribute at the same parametric order.

\(\chi_\parallel=\chi_\perp=1\)은 서로 다른 층의 contribution이 지수적으로 달라지는 것을 막는다. 9.4는 여기서 한 단계 더 나간다. 특정 parameter type이나 특정 layer가 polynomial하게 training을 독점하는 것도 피하자는 것이다. 책은 이를 learning-rate의 equivalence principle이라고 부른다.

Scale-invariant class: layer-independent, depth-normalized

ReLU 계열에서는 layer별 bias·weight learning-rate coefficient를 동일하게 둘 수 있다. 다만 output NTK를 overall \(O(1)\)으로 유지하려면 전체 depth \(L\)로 normalize하고, weight term은 이전 층 width로도 normalize한다.

Scale-invariant equivalence scaling · Eq. 9.94
\[\eta\lambda_b^{(\ell)}=\frac{\eta\widetilde\lambda_b}{L},\qquad \frac{\eta\lambda_W^{(\ell)}}{n_{\ell-1}}=\frac{\eta\widetilde\lambda_W}{L\,n_{\ell-1}}.\]

\(K^\star=0\) class: layer, depth, width를 함께 보정한다

이 class에서는 \(\ell\)-dependent balancing까지 필요하다. 9.3의 layer scaling에 overall depth normalization을 결합하면

K-star-zero equivalence scaling · Eq. 9.95
\[\eta\lambda_b^{(\ell)}=\eta\widetilde\lambda_b\left(\frac1\ell\right)^{p_\perp}L^{p_\perp-1},\qquad \frac{\eta\lambda_W^{(\ell)}}{n_{\ell-1}}=\frac{\eta\widetilde\lambda_W}{n_{\ell-1}}\left(\frac{L}{\ell}\right)^{p_\perp-1}.\]

tanh와 sin처럼 \(p_\perp=1\)이면 식이 더 단순해진다.

Odd smooth activations with p-perp = 1 · Eq. 9.96
\[\eta\lambda_b^{(\ell)}=\frac{\eta\widetilde\lambda_b}{\ell},\qquad \frac{\eta\lambda_W^{(\ell)}}{n_{\ell-1}}=\frac{\eta\widetilde\lambda_W}{n_{\ell-1}}.\]

즉 tanh에서는 bias learning rate가 layer index에 따라 줄어들지만 weight coefficient는 이런 추가 \(\ell\)-rescaling이 필요하지 않는다. 책은 ReLU의 empirical popularity가 이런 간단한 layer-independent scaling과 부분적으로 관련될 가능성을 질문 형태로 제시한다. 이는 원문의 가설적 해석이지, 9장이 경험적으로 입증한 결론은 아니다.

9장이 남긴 설계 원리

  • 01Initialization과 optimization을 따로 튜닝하지 않는다. \(C_b,C_W\)가 susceptibility를 정하고, susceptibility가 NTK contribution의 depth scaling을 정하므로 두 hyperparameter 계층은 연결되어 있다.
  • 02Raw learning rate보다 scaled learning rate가 비교 단위다. width와 depth가 다른 architecture 사이에서 의미 있는 비교를 하려면 \(n\), \(L\), 필요하면 \(\ell\) dependence를 명시해야 한다.
  • 03Mean만으로 finite-width training을 설명할 수 없다. normalized NTK variance와 NTK–representation cross-correlation이 모두 \(L/n\)으로 relevant해진다.
  • 04criticality는 forward와 backward를 동시에 묶는다. \(\chi_\parallel\)은 kernel/output scale을, \(\chi_\perp\)은 chain-rule sensitivity와 interlayer influence를 안정화한다.

읽을 때 지켜야 할 경계

9장의 universality와 \(L/n\) 법칙은 leading large-width, single-input, deep-MLP effective theory에서 얻어진다. multi-input statistics, higher-order \(1/n\) corrections, 실제 finite-step training trajectory와 generalization은 이 장에서 직접 풀지 않는다. 또한 \(L/n\)이 충분히 작다는 perturbative 전제가 깨지면 higher-order finite-width terms가 중요해진다.

깊이는 parameter 수를 늘리는 장치가 아니다. 한 layer의 정보와 학습 신호가 다음 layer로 전달되는 규칙을 반복하는 장치다. 9장의 criticality는 그 반복이 지수적으로 무너지지 않게 하고, equivalence principle은 특정 층이 polynomial하게 독점하지 않게 한다.이 두 원리가 합쳐질 때 depth와 width가 달라져도 비교 가능한 training scale을 정의할 수 있다.

References & Source Notes

01
Daniel A. Roberts, Sho Yaida, Boris Hanin. The Principles of Deep Learning Theory.

Cambridge University Press. Chapter 9, “Effective Theory of the NTK at Initialization,” pp. 227–246. DOI: 10.1017/9781009023405.011.

02
Jacot, Gabriel, Hongler — Neural Tangent Kernel.

원저는 frozen NTK의 문헌적 맥락으로 NTK의 infinite-width formulation을 연결한다. Chapter 9의 핵심은 그 deterministic limit에 finite-width fluctuations와 cross-correlations를 결합해 effective theory로 확장하는 데 있다.

03
Exploding and vanishing gradients.

원저는 이 문제의 역사적 맥락으로 Hochreiter, Bengio–Frasconi–Simard, Pascanu–Mikolov–Bengio 등을 참고문헌 [58]–[60]에서 연결하며, Chapter 9에서는 이를 kernel criticality와 NTK forward equation의 언어로 다시 해석한다.

04
Visual note.

본문 hero SVG와 universality comparison table은 Chapter 9의 scaling equations을 설명하기 위해 새로 구성한 개념도다. 측정된 실험 그래프나 원저의 figure를 재현한 것이 아니다.

본문은 원문을 문장 단위로 번역하지 않고, 논리적 명료성·근거 중심의 전개·복잡한 수학의 단계적 해설을 중시하는 한국어 기술 에세이로 재구성했다. 책이 직접 계산한 결과와 본문의 해석·제한사항을 구분해 서술했다.