Temporal alignment
어떤 prefix feature와 새 target을 짝지을 것인가.
Falcon reframes recurrent sequence modeling as continual learning inside the forward pass: every token becomes a causal training example for a fixed-size fast-memory state.
이 논문의 핵심은 새로운 attention kernel 하나를 제안하는 데 있지 않다. recurrent state update를 architecture equation이 아니라 forward pass 안에서 수행되는 online continual learning rule로 다시 해석한다는 데 있다.
Transformer는 새 문맥을 확장되는 KV cache에 외부화한다. 반면 linear attention, fast-weight memory, selective SSM은 문맥을 고정 크기 recurrent state에 압축한다. 이때 state transition은 새로운 evidence를 state에 결합하는 local online learning rule이 된다.
read-after-write autoregressive semantics에서 step \(t\)에 새로 드러난 target \(v_t\)를 학습할 때 prefix-prediction objective에 맞는 training pair는 흔히 쓰이는 \((\phi(k_t),v_t)\)가 아니라 \((x_t,y_t)=(\phi(k_{t-1}),v_t)\)이다. same-step association도 causal이지만 다른 내부목적을 최적화한다.
KV cache, SSM, linear attention, fast weights를 문맥 저장과 online adaptation이라는 하나의 축에서 비교한다.
문맥이 늘어날수록 KV cache가 커진다. standard attention은 \(O(N^2)\) attention matrix와 증가하는 memory traffic을 요구한다.
확장되는 문맥을 fixed-size recurrent state \(S_t\)로 압축한다. inference state는 일정하지만 어떤 evidence를 어떻게 state에 쓰는지가 학습규칙 자체가 된다.
논문은 classical state-space/control, Kalman filtering, adaptive filtering에서 현대 linear attention, Mamba/Mamba-2, DeltaNet까지 연결한다. Mamba-2의 Structured State Space Duality는 recurrent SSM과 causal linear attention 사이 bridge를 제공하고, Falcon은 그 위에서 state update의 internal objective를 다시 설계한다.
어떤 prefix feature와 새 target을 짝지을 것인가.
새 evidence가 state를 얼마나 크게 수정하는가. \(\beta_t\)와 realized \(\eta_t\)가 담당.
ridge/shrinkage \(\lambda_t\)로 오래된 state를 얼마나 감쇠하는가.
sliding window로 최근 causal examples를 재사용해 noise와 local dependency를 다루는가.
논문은 read-after-write(RAW) convention을 사용한다. token \(t\)를 관측하고 state에 쓴 뒤 updated state \(S_t\)를 읽어 token \(t+1\)을 예측한다. 새로 드러난 \(v_t\)는 그 target을 예측할 때 사용 가능했던 prefix feature와 짝지어야 한다.
표준 DeltaNet/linear-attention-style update가 흔히 사용하는 \((\phi(k_t),v_t)\)도 미래정보를 쓰지는 않는다. 그러나 그것은 same-step cache association을 학습한다. Falcon은 “예측 당시 사용 가능했던 prefix representation이 새 target을 얼마나 잘 예측했는가?”를 local objective로 삼는다.
observe \(t\) → write \(S_t\) with \(\phi(k_{t-1})\to v_t\) → read \(S_t\) → predict \(t+1\).
\(\phi(k_t)\leftrightarrow v_t\)를 같은 step에서 결합. causal이지만 다른 fast-memory objective.
read \(S_{t-1}\)로 token \(t\)를 예측한 뒤 write하므로 update timing과 local objective가 달라진다.
\(S_0=0,\;x_1=0,\;\eta_1=0\). carried state의 첫 boundary step이 data 없이 decay하지 않게 한다.
pre-update prediction은 \(\hat y_t=S_{t-1}^\top x_t\), residual은 \(r_t=y_t-\hat y_t\)이다. 한 번의 normalized online gradient step은 다음과 같다.
\(\lambda_t=0,\varepsilon=0\)이면 classical NLMS recursion과 연결된다. normalization은 local smoothness에 step size를 맞춰 per-step descent를 제공하지만 cumulative online loss나 outer autoregressive training objective의 monotone decrease를 뜻하지 않는다.
모든 value channel이 scalar \(\eta_t\)를 공유. shared dynamics와 hardware efficiency.
\(\eta_t\in\mathbb{R}^{d_v}\). 각 value channel이 독립 plasticity trajectory를 가짐.
최근 \(B\) causal pairs의 window-average loss에 mini-batch gradient step. bounded rehearsal.
residual 대신 target을 직접 additive write. energy-normalized scalar gain.
direct target write를 유지하면서 value channel마다 별도 gain 적용.
window-average cross-covariance를 write하고 window energy로 magnitude 안정화.
이미 state가 예측한 부분을 residual로 제거한 뒤 error-driven edit를 수행한다. denominator는 local curvature/smoothness와 연결된다.
\(-\langle S^\top x_t,y_t\rangle+\frac{\lambda_t}{2}\|S\|_F^2\)로 target을 직접 write한다. energy normalization은 write-magnitude stabilizer다.
step size는 \(\mu_t^{(B)}=\lambda_{\max}(\bar C_t^{(B)})\)에 맞춘다. window average를 사용하므로 nominal \(B\)가 커진다고 injection이나 decay fraction이 선형증가하지 않는다. exact segment continuation에는 matrix state뿐 아니라 최근 \(B-1\) causal pair tail도 필요하다.
논문의 Figures 2, 4, 5, 7, 8, 9, 10은 recurrent form과 parallel form, SSD-style chunk-wise form의 대응을 반복해서 보여준다. 목표는 continual-learning update를 도입하면서도 현대 GPU 학습의 chunk parallelism을 유지하는 것이다.
Falcon-2는 channel마다 \(\eta_{t,j}\)가 다르지만 모든 channel이 같은 write-feature geometry를 공유한다. chunk key matrix \(K\)에서 shared Gram \(G=K^\top K\)를 한 번 만들고, channel별 rate는 작은 \(C\times C\) triangular system에만 반영한다. value path와 projected-history path가 같은 factor를 사용하므로 두 solve를 하나의 residual right-hand side로 합쳐 one TriSolve per chunk로 줄인다.
| Component / chunk | Falcon-2 full | Falcon-1 |
|---|---|---|
| Gram Matrix | O(dC²) | O(dC²) |
| Rate-dependent system build | O(dᵥC²) | O(C²) |
| Forward residual TriSolve | O(dᵥC²) | O(dᵥC²) |
| State update / output projection | O(ddᵥC) | O(ddᵥC) |
Falcon-1은 scalar \(\eta_t\)를 공유해 channel별 system construction/factorization을 제거하고 single multi-RHS solve로 바꾼다. asymptotic solve cost는 남지만 GPU 활용 관점의 practical savings가 커진다.
ridge는 \(\gamma_t=1-\eta_t\lambda_t\) carry를 만든다. long context에서 decay product를 직접 누적하면 reduced precision underflow/overflow가 생길 수 있다. 구현은 \(\alpha_t=\eta_t\lambda_t\)를 \(1-\varepsilon_\gamma\) 아래로 clamp하고 \(\log\gamma_t=\log1p(-\alpha_t)\)를 fp32로 계산한다. global cumulative product 대신 각 chunk에서 log-prefix를 0부터 다시 누적하고 state, step size, write target을 일관되게 재정규화한다.
124M–130M params
FineWeb-Edu
100k steps · seq 1,024 · batch 480 · ≈49.2B tokens
single 4-GPU H100/H200 node
Transformer baseline은 LLaMA-style RoPE/SwiGLU, recurrent baselines는 RetNet/LightningAttn, Mamba-2, DeltaNet, Gated DeltaNet이다. 학습은 bfloat16 AdamW, tied embeddings, Pre-Norm RMSNorm, no dropout, µP-style width scaling, base LR 1e−3 cosine decay, 2,000 warmup, weight decay 0.1, gradient clipping 1.0을 사용한다.
| Model | Wiki. | LMB. | FineEdu ↓ |
|---|---|---|---|
| Transformer (RoPE) | 33.25 | 47.43 | 17.38 |
| RetNet / LightningAttn | 36.86 | 65.16 | 18.79 |
| Mamba-2 | 34.53 | 48.74 | 17.70 |
| DeltaNet | 34.19 | 52.84 | 17.84 |
| Gated DeltaNet | 30.99 | 46.70 | 17.32 |
| Falcon-1A.1 | 34.41 | 47.93 | 17.70 |
| Falcon-1A.2 | 34.20 | 51.01 | 17.70 |
| Falcon-1A.3 | 34.02 | 49.84 | 17.40 |
| Falcon-1.3 | 33.00 | 48.70 | 17.10 |
FineWeb-Edu에서는 Falcon-1.3이 17.10으로 표 전체의 최저 perplexity다. baseline 최강은 Gated DeltaNet 17.32이며 inner-product Falcon 중에는 Falcon-1A.3 17.40이 가장 좋다. 저자들은 이를 uniform win이 아니라 competitive language-model quality로 해석한다.
| Model | PIQA | Hella. | Wino. | ARC-e | ARC-c | OBQA | Social IQA | SciQ | Avg. |
|---|---|---|---|---|---|---|---|---|---|
| Zero-shot | |||||||||
| Transformer | 65.67 | 37.54 | 51.70 | 52.36 | 27.65 | 31.60 | 38.84 | 79.90 | 48.16 |
| RetNet/LightningAttn | 64.91 | 35.36 | 49.64 | 57.62 | 26.28 | 32.20 | 37.97 | 80.40 | 48.05 |
| Mamba-2 | 66.32 | 36.89 | 50.75 | 58.16 | 26.62 | 32.60 | 38.38 | 80.70 | 48.80 |
| DeltaNet | 66.38 | 37.15 | 52.33 | 57.37 | 26.79 | 34.00 | 39.00 | 78.00 | 48.88 |
| Gated DeltaNet | 65.51 | 37.75 | 49.72 | 58.88 | 27.90 | 31.60 | 38.28 | 80.60 | 48.78 |
| Falcon-1A.1 | 66.05 | 37.27 | 50.67 | 57.62 | 27.65 | 31.60 | 38.95 | 81.10 | 48.86 |
| Falcon-1A.2 | 67.03 | 37.29 | 52.33 | 57.37 | 25.94 | 33.20 | 38.84 | 82.40 | 49.30 |
| Falcon-1A.3 | 66.10 | 37.55 | 50.12 | 59.01 | 26.79 | 32.00 | 37.82 | 82.20 | 48.95 |
| Falcon-3A.3 | 65.34 | 37.30 | 50.99 | 57.37 | 26.37 | 33.60 | 39.41 | 81.60 | 49.00 |
| Falcon-1.3 | 65.83 | 38.38 | 52.25 | 58.96 | 26.62 | 31.40 | 38.69 | 81.30 | 49.18 |
| One-shot | |||||||||
| Transformer | 66.43 | 37.55 | 50.28 | 59.64 | 29.01 | 30.00 | 39.82 | 84.60 | 49.67 |
| RetNet/LightningAttn | 65.13 | 35.19 | 49.64 | 56.78 | 25.85 | 28.80 | 36.44 | 81.50 | 47.42 |
| Mamba-2 | 66.70 | 36.61 | 51.07 | 58.63 | 26.96 | 32.40 | 37.97 | 82.90 | 49.16 |
| DeltaNet | 66.27 | 36.67 | 50.36 | 57.70 | 27.39 | 32.00 | 37.72 | 79.90 | 48.50 |
| Gated DeltaNet | 65.40 | 37.81 | 51.30 | 58.29 | 26.88 | 30.00 | 37.26 | 81.60 | 48.57 |
| Falcon-1A.1 | 66.81 | 37.41 | 50.75 | 57.53 | 27.99 | 32.40 | 37.92 | 81.80 | 49.08 |
| Falcon-1A.2 | 66.38 | 36.97 | 52.96 | 57.32 | 27.82 | 31.40 | 37.56 | 83.20 | 49.20 |
| Falcon-1A.3 | 65.78 | 37.22 | 49.72 | 58.88 | 27.73 | 31.40 | 37.97 | 82.40 | 48.89 |
| Falcon-3A.3 | 65.72 | 36.47 | 51.78 | 58.29 | 26.54 | 32.00 | 38.23 | 83.20 | 49.03 |
| Falcon-1.3 | 65.67 | 38.09 | 52.80 | 59.55 | 29.01 | 30.40 | 37.46 | 83.40 | 49.54 |
Falcon-1A.2가 zero-shot 평균 49.30으로 listed recurrent model 중 최고이고, Falcon-1.3은 one-shot 평균 49.54로 recurrent model 중 최고다. 다만 Transformer one-shot 평균 49.67보다 높지는 않다.
1–32 digit addition으로 학습하고 reversed sum을 생성하며 33–48 digit을 OOD length generalization으로 사용한다. storage와 carry propagation을 직접 스트레스하는 controlled diagnostic이다.
| Model | Best step | Val. acc. | Mean 33–48 ↑ | Acc@d33 / d48 |
|---|---|---|---|---|
| Transformer | 2000 | 100.0 | 65.8 | 97 / 49 |
| RetNet/LightningAttn | 2000 | 99.7 | 82.9 | 99 / 63 |
| Mamba-2 | 2000 | 100.0 | 75.2 | 100 / 51 |
| Falcon-1A.1 | 1900 | 100.0 | 80.6 | 100 / 59 |
| Falcon-1A.2 | 2000 | 100.0 | 85.2 | 100 / 63 |
| Falcon-1A.3 | 1900 | 99.8 | 85.9 | 100 / 69 |
| Falcon-3A.3 | 2000 | 99.9 | 87.2 | 100 / 69 |
| Falcon-1.3 | 2000 | 100.0 | 68.8 | 100 / 48 |
Falcon-3A.3이 mean 87.2로 가장 높고 Falcon-1A.3이 85.9다. 저자들은 이 결과를 primary result가 아니라 shifted/normalized update가 storage와 carry propagation이 지배적인 환경에서 extrapolation을 개선한다는 supporting evidence로 위치시킨다.
Linear Attention, Performer, Hyena, RetNet, RWKV, S4, Mamba, Mamba-2.
Hinton-Plaut, Schmidhuber, Fast Weight Programmers, DeltaNet, Gated DeltaNet.
LMS/NLMS, RLS. classical scale-robust online regression 원리를 fast memory로 이동한다.
implicit gradient descent, Test-Time Training, test-time regression, MesaNet, Titans, ATLAS.
Falcon의 차별점은 strict next-latent causal alignment, objective-matched normalization, per-column/sliding variants, SSD-compatible chunk-parallel implementation을 함께 묶는다는 점이다. RWKV-7/Kimi Linear처럼 richer gating을 쓰는 delta-style model에도 이 alignment/normalization을 local update replacement로 적용할 수 있다고 설명한다.
Falcon-1/2/3 scalar, per-channel, sliding update 비교.
denominator-free linear attention의 recurrent / masked / chunk-parallel view.
same-step association과 next-latent prediction의 objective alignment.
Falcon-1 rank-one recurrent ↔ WY ↔ per-chunk WY/Gram.
Falcon-2 shared Gram + per-channel batched TriSolve.
Falcon-1A/2A/3A direct inner-product writes.
Falcon-1A scalar decay-mask attention.
Falcon-2A per-channel decay masks.
Falcon-3 rank-B affine recurrence / ParallelFlow.
Falcon-3A window-induced decay mask와 chunk overlap.
RAW/RBW × shifted/unshifted timing convention 2×2 grid.
Falcon-2 chunk-parallel forward와 Falcon-3 recurrent reference.
Falcon-3 ParallelFlow와 Falcon-3A masked chunk attention.
batch-size-1 first-order online ridge reference.
DeltaNet WY forward/backward와 merged residual solve.
Falcon-1 forward/backward, NLMS map과 chunk-local log-decay chain rule.
RLS exact online ridge, first-order ridge, SSM discretization. RLS는 exact cumulative solution이지만 parallelization cost가 높다.
optimizer, precision, µP scaling, learning-rate schedule, H100/H200 hardware.
scaled recurrence, signed-feature denominator caveat, gain/gating, TTT/Titans relation, timing/boundary, log-space renormalization.
structured mask identity, chunk-local evaluation, backward pass.
affine WY representation과 chunk-parallel forward/backward.
per-channel dynamics, shared Gram geometry, batched one-TriSolve.
hardware-efficient Falcon-1, scalar/vector rates, stable positive-decay, complexity.
affine chunk map, matrix-valued CDE, low-rank drivers, tensorInv와 rank-B mapping.
Falcon-2/2A와 Falcon-3는 정의·구현되지만 main tables에서 별도 benchmark되지 않는다. Language modeling은 uniform win이 아니라 competitive quality다. per-step descent는 instantaneous local objective에만 해당한다. normalized read는 signed feature에서 denominator sign instability가 있을 수 있다. Falcon-3/3A exact continuation에는 \(B-1\) causal-pair tail이 필요하다.
Falcon은 hidden state를 passive summary가 아니라 forward pass 도중 갱신되는 parameter로 본다. 그러면 각 token은 fixed-size state를 위한 self-supervised online training example이 된다. architecture 설계는 kernel 선택에서 끝나지 않고 training-pair alignment, plasticity, forgetting, bounded rehearsal, parallel kernel의 공동설계 문제가 된다.
second-order/RLS, nonlinear predictor, task-conditioned local objective로 확장할 수 있다.
\(\beta,\lambda,B\)를 context와 uncertainty에 따라 동적으로 조절하는 연구공간이 열린다.
장기 실행 agent가 slow-weight update 없이 fixed-size state로 새로운 evidence를 온라인 학습하는 연결점이 된다.
Fast Weight Attention for Continual Learning은 recurrent attention을 압축된 attention으로만 보지 않는다. 매 순간 새 evidence로 스스로를 갱신하는 작은 온라인 학습기로 다시 정의한다.
본 게시물은 첨부된 54쪽 PDF의 Abstract, Introduction, Background, Autoregressive Next-Latent Prediction, Falcon derivation, Figures, Algorithms, Experiments, Related Work, Conclusion, References 및 Appendix A–H를 전체적으로 검토해 웹 읽기 흐름으로 재구성했다. 논문에 없는 외부 실험수치나 주장으로 결과를 보강하지 않았다.