01 · The modeling question
Temporal Knowledge Graph의 기본 단위는 event인데, 학습의 중심은 entity였다
TKG의 한 사건은 (s, r, o, t) quadruple이다. subject와 object entity, relation, timestamp가 하나의 사건을 이룬다. 그럼에도 기존 representation learning은 사건을 entity와 relation으로 분해하고, entity를 node, relation을 edge로 둔 “primary graph”에서 representation을 학습하는 데 집중했다.
GNN 계열은 entity의 pairwise structural dependency를 message passing으로 학습한다. 더 최근의 방법은 community, entity group, hypergraph, evolutionary cluster 같은 derived structure를 추가해 reachable하지 않은 entity 사이의 high-order correlation까지 모델링한다. 하지만 저자들의 문제 제기는 단순하다. event 자체 사이의 상관관계는 어디에 있는가?
논문은 event가 TKG의 core constituent임에도 기존 연구가 event-event correlation을 체계적으로 모델링하지 않았다고 주장한다. 사건은 shared entity를 통해 직접 연결될 수도 있고, entity를 공유하지 않아도 representation space에서 의미적으로 가까울 수 있으며, 여러 사건이 하나의 고차원적 pattern을 만들고 이 pattern이 시간에 따라 다른 pattern으로 이동할 수도 있다.
02 · Figure 1
러시아-우크라이나 전쟁 이후의 에너지 재편은 왜 entity graph만으로 충분하지 않은가
페이지 2의 Figure 1은 논문의 동기를 사건 수준에서 보여준다. 2022년 2월 러시아-우크라이나 전쟁 발발(Event A) 이후 유럽연합과 미국이 러시아 에너지 의존을 줄이는 사건(Events B, C)이 발생한다. 이어 러시아-인도 할인 원유 계약(Event D), 카타르-독일 가스 공급 계약(Event E), 러시아-중국 가스 공급 계약(Event F)이 이어진다.
여기에는 서로 다른 종류의 correlation이 겹친다. A-B와 A-C처럼 entity를 공유하거나 직접적으로 이어지는 pairwise relation이 있고, B-E처럼 사건의 당사자가 완전히 동일하지 않아도 의미적으로 가까운 correlation이 있다. B와 C는 “러시아 에너지 의존 축소”라는 high-order pattern을 만들고, D·E·F는 “대체 혹은 재편된 에너지 공급”이라는 또 다른 pattern을 형성한다. 첫 pattern의 변화가 두 번째 pattern으로 이어진다.
Co-entity correlation
두 event가 subject/object entity를 하나 이상 공유한다. 관측 가능한 구조적 연결이다.
Proximity correlation
entity를 공유하지 않더라도 learned event representation에서 가까운 kNN이면 edge를 만든다.
High-order correlation
여러 event가 하나의 soft cluster에 동시에 속하며, cluster 자체가 시간에 따라 align되고 진화한다.
04 · Preliminaries
Primary graph → Event graph → Event cluster graph의 세 층을 명시적으로 정의한다
시점 t의 사건 집합은 G^t로 표기한다.
node는 entity, edge는 relation이다. 기존 TKG representation learning의 기본 graph이다.
node가 event이고 edge가 event 사이의 heterogeneous pairwise correlation이다.
node는 event cluster, edge는 cluster 사이의 latent correlation이다. event cluster가 high-order event correlation을 표현한다.
event prediction task는 query (s, ?, o, t)에서 subject와 object, 과거 G^{1:T-1}이 주어졌을 때 가능한 relation의 조건부 확률 p(r̂ | s,o,G^{1:T-1})을 예측한다. 즉 link prediction의 relation slot을 temporal history를 사용해 채우는 문제이다.
05 · Architecture overview
HEART는 세 단계의 graph를 따라 event correlation을 점점 더 높은 차수로 올린다
entities + relations
events + co-entity/proximity edges
soft overlap + temporal alignment
latent edges + message passing
각 timestamp에서 relation-aware RGCN이 entity/relation representation을 갱신한다. relation-aware tensor event encoder가 (s,r,o)를 하나의 event vector로 만든다. event graph construction은 co-entity와 proximity edge를 붙인다. fuzzy c-means가 overlapping event cluster를 만들고, Hungarian matching이 과거 cluster와 현재 cluster를 정렬한다. self-supervised cluster optimization과 learnable sparsification이 cluster graph를 정제한다.
cluster-level 정보는 membership matrix를 통해 event로 되돌아가고, event에서 entity representation으로 역전파된다. time residual gate는 이전 timestamp representation과 현재 input을 섞고, attentive temporal encoder가 여러 timestamp의 entity/relation history를 position-enhanced self-attention으로 통합한다. 마지막에는 ConvTransE decoder가 relation probability를 예측한다.
06 · Event graph construction
Event를 단순히 s+r+o로 합치지 않고 relation-aware tensor interaction으로 encode한다
먼저 primary graph에서 RGCN이 structural dependency를 학습한다. object o의 다음 layer representation은 incoming (s,r) neighbor의 h_s + h_r를 degree-normalized aggregation하고 self-loop h_o를 더한 뒤 RReLU를 통과한다.
초기 entity representation은 in-degree가 같은 entity끼리 동일하게 배정해 update를 빠르게 하고, relation은 random initialization한다. relation representation은 해당 relation과 연결된 entity representation과 이전 시점 relation representation을 mean-pooling해 갱신한다.
event encoder의 핵심은 3-way tensor T이다. 같은 entity pair라도 relation이 다르면 사건 의미가 완전히 달라질 수 있으므로 subject-relation과 object-relation interaction을 별도로 계산한다.
이렇게 만들어진 x가 event graph의 node가 된다. event 사이에는 두 종류 edge를 만든다.
E_x^t = E_coe^t ∪ E_prox^t로 합쳐 heterogeneous event graph를 만들고, 다시 RGCN을 적용해 local structural/semantic event correlation을 학습한다.
07 · Event-aware multi-step evolutionary clustering
Event는 하나의 cluster에만 속하지 않는다. Soft membership을 유지한 채 시간축에서 cluster identity를 맞춘다
한 event는 여러 high-order pattern에 동시에 참여할 수 있으므로 hard clustering 대신 fuzzy c-means를 사용한다. membership U_{i,j}^{t,m}은 event x_i^t가 cluster centroid μ_j^t에 속하는 정도를 표현한다.
cluster representation은 membership-weighted event sum으로 만들고, membership은 cosine similarity와 fuzzy smoothing parameter m>1로 계산한다.
문제는 timestamp가 바뀌면 event에 persistent identifier가 없고 cluster index도 임의라는 점이다. HEART는 event cluster에 속한 entity distribution을 이용해 과거와 현재 cluster를 align한다. event-to-entity mapping matrix M^t와 entity occurrence normalization D^{-1}를 사용한다.
과거 cluster와 현재 cluster entity distribution의 cosine similarity로 affinity matrix를 만든다.
Hungarian algorithm이 one-to-one maximum-similarity matching을 다항 시간에 구한다. 이렇게 얻은 permutation π^τ로 현재 cluster와 모든 aligned historical cluster를 평균해 multi-step representation을 만든다.
마지막으로 aligned cluster가 시간에 따라 급격히 흔들리지 않도록 cosine distance 기반 temporal smoothness loss를 둔다.
08 · Event cluster graph message passing
Cluster를 만든 뒤에는 어떤 cluster끼리 실제로 상호작용할지 다시 학습한다
self-supervised optimization은 비슷한 cluster의 cohesion을 높이고 다른 cluster의 differentiation을 키우려는 장치이다. 먼저 Euclidean cluster distance L과 inverse-distance weight K를 계산한다.
cosine similarity F로 similarity adjustment S를 만든다.
distance와 similarity adjustment를 합쳐 cluster representation을 shift한다.
shift가 원 representation의 방향을 과도하게 바꾸지 않도록 self-supervised cosine loss를 둔다.
cluster graph edge는 implicit correlation encoder가 만든다. 두 cluster를 concatenate한 뒤 MLP와 ReLU로 latent relation z_{i,j}를 만들고, convolution+sigmoid로 edge intensity q_{i,j}를 0–1 사이에 둔다. learnable threshold κ보다 작은 edge는 0으로 만든다.
남은 edge intensity를 weight로 message를 aggregate한다.
그 뒤 fuzzy membership matrix를 이용해 cluster-level message를 event로 내리고 x_i'^t = Σ_j U_{i,j}^{t,m} ĉ̂_j^t로 갱신한다. event construction을 역으로 적용해 entity representation도 업데이트한다.
09 · Temporal integration and event prediction
Event-centric 구조가 최종적으로 entity와 relation representation을 다시 바꾼다
HEART가 event를 직접 예측하는 별도 decoder를 쓰는 것은 아니다. event graph와 cluster graph가 entity/relation representation을 개선하고, time residual gate와 attentive temporal encoder가 여러 시점의 representation을 통합한다. 저자들은 이 두 component의 상세식은 기존 연구 [7,39,40]를 참조하고, position-enhanced self-attention으로 final representation H를 얻는다고 설명한다.
decoder는 ConvTransE이다. entity pair를 convolution한 feature와 relation matrix를 곱해 가능한 relation의 확률을 예측한다.
TKG prediction loss는 multi-label binary cross entropy 형태이다.
전체 objective는 prediction, temporal smoothness, self-supervised cluster shift의 weighted sum이다.
10 · Appendix A
Event-centric modeling은 표현력을 늘리는 대신 event 수의 제곱항을 지불한다
| Module | Time complexity |
|---|---|
| RGCN | O((Ne+Nr)D²) |
| Event graph construction | O(Nx²D + NxD²) |
| Multi-step evolutionary clustering | O(T(NxNcD + NxNcNe + Nc³)) |
| Event cluster graph message passing | O(Nc²D + NxNcD) |
| Time residual gate | O(D²) |
| Attentive temporal encoder | O(T²D) |
| Event prediction | O(D) |
논문이 정리한 total complexity는 O((N_e+N_r+N_x)D² + N_x²D + T N_x N_c(D+N_e))이다. 특히 event graph construction의 N_x²와 clustering의 N_c³ Hungarian-related term은 event/cluster 규모가 커질 때 중요한 비용 요소이다.
11 · Datasets
정치 사건에서 15분 단위 GDELT, 연 단위 WIKI/YAGO까지 시간 granularity가 크게 다르다
| Dataset | #Entity | #Relation | Training | Validation | Test | Interval |
|---|---|---|---|---|---|---|
| ICEWS14 | 7,128 | 230 | 74,845 | 8,514 | 7,371 | 24 hours |
| ICEWS14C | 205 | 171 | 35,665 | 7,369 | 7,068 | 24 hours |
| ICEWS18 | 23,033 | 256 | 373,018 | 45,995 | 49,545 | 24 hours |
| ICEWS18C | 208 | 164 | 34,497 | 4,412 | 4,661 | 24 hours |
| GDELT | 7,691 | 240 | 1,734,399 | 238,765 | 305,241 | 15 mins |
| WIKI | 12,554 | 24 | 539,286 | 67,538 | 63,110 | 1 year |
| YAGO | 10,623 | 10 | 161,540 | 19,523 | 20,026 | 1 year |
ICEWS14/18은 Integrated Crisis Early Warning System의 정치 사건을 하루 단위로 기록한다. C 버전은 country-related event만 필터링한 dataset이다. GDELT는 human behavior/event를 15분 단위로 기록한다. WIKI와 YAGO는 연 단위 temporal aggregation이다. 이 granularity 차이는 뒤에서 HEART의 성공과 실패를 가르는 중요한 변수로 다시 등장한다.
12 · Experimental settings
PyTorch 2.4.1, RTX 3090/A800, NNI 25 trials로 key hyperparameter를 찾는다
HEART는 Python/PyTorch 2.4.1로 구현하고 NVIDIA RTX 3090 24GB와 A800 80GB GPU에서 학습한다. Neural Network Intelligence(NNI) toolkit과 Tree-structured Parzen Estimator를 사용해 최대 25 trial로 key hyperparameter를 search한다.
| Hyperparameter | Search space | ICEWS14 | ICEWS14C | ICEWS18 | ICEWS18C | GDELT | WIKI | YAGO |
|---|---|---|---|---|---|---|---|---|
| Nc | {2,4,6,8,10,12,14,16,18,20} | 16 | 12 | 16 | 10 | 6 | 4 | 4 |
| Nlayer | {1,2,3,4,5} | 1 | 1 | 2 | 2 | 2 | 3 | 1 |
| Nwindow | 1…14 | 12 | 11 | 5 | 5 | 3 | 3 | 11 |
| α (=β) | {0.1,0.2,0.3,0.4,0.5} | 0.3 | 0.2 | 0.1 | 0.5 | 0.4 | 0.1 | 0.2 |
| k in kNN | {1,2,3,4,5} | 2 | 5 | 3 | 1 | 1 | 1 | 4 |
λ=γ=0.1로 고정해 temporal/self-supervised regularization이 prediction objective를 지나치게 압도하지 않게 한다. Adam learning rate는 0.01, batch size 16, representation dimension 200이며 time residual gate와 attentive temporal encoder hidden size도 200이다. 결과는 3 independent run 평균이다.
metric은 MRR과 Hits@1/3/10이다. MRR은 ground-truth relation의 reciprocal rank 평균이고 Hits@k는 correct relation이 top-k 안에 들어간 비율이다.
13 · Baselines
13개 baseline은 shallow encoder → GNN → derived structure의 세 세대로 나뉜다
TTransE · HyTE
TTransE는 timestamp를 entity representation에 넣고, HyTE는 timestamp를 hyperplane으로 표현한다.
RE-NET · Glean · TeMP · RE-GCN · DACHA · TiRGN
GCN/RGCN으로 structural influence를, RNN/GRU/self-attention으로 temporal dependency를 모델링한다.
TITer · EvoExplore · GTRL · DHyper · DECRL
path, dynamic community, entity group, hypergraph, evolutionary entity cluster를 이용해 distant/high-order correlation을 학습한다.
TiRGN은 저자들이 사용하는 SOTA GNN baseline이고, DECRL은 가장 강한 derived-structure baseline이다. HEART와 DECRL의 차이는 부록에서 별도로 강조된다. DECRL은 entity-level evolutionary cluster이고 persistent entity ID가 있다. HEART는 event graph/event cluster graph를 직접 만들고 entity distribution 기반 multi-step alignment가 필요하다. 또 DECRL은 adjacent timestamp smoothness 중심인 반면 HEART는 multi-step temporal loss와 cluster-level self-supervised objective를 추가한다.
14 · Main benchmark results
HEART는 평균적으로 강하지만, “모든 dataset의 모든 metric에서 1위”는 아니다
논문은 ICEWS14, ICEWS18, GDELT를 main table로 제시한다. 별표는 pairwise t-test 95% confidence에서 HEART가 비교 방법을 유의하게 앞서는 항목을 뜻한다. 아래 표는 runner-up과 HEART 중심으로 먼저 읽을 수 있게 요약하고, 이어지는 접이식 표에 모든 baseline 값을 보존한다.
| Dataset | Method | MRR | Hits@1 | Hits@3 | Hits@10 |
|---|---|---|---|---|---|
| ICEWS14 | DECRL | 62.61 | 48.73 | 70.57 | 93.03 |
| HEART | 64.80 | 51.63 | 72.33 | 92.57 | |
| ICEWS18 | DECRL | 63.30 | 50.13 | 70.72 | 90.82 |
| HEART | 67.53 | 54.37 | 74.79 | 95.63 | |
| GDELT | Best prior by metric | 31.94 (DHyper) | 18.85 (DHyper) | 33.73 (DHyper) | 64.29 (DHyper) |
| HEART | 34.12 | 19.69 | 36.90 | 73.51 |
ICEWS14에서 HEART는 MRR +3.50%, Hits@1 +5.95%, Hits@3 +2.49%지만 Hits@10은 DECRL 93.03보다 낮은 92.57로 −0.49%이다. ICEWS18에서는 67.53/54.37/74.79/95.63으로 네 metric 모두 DECRL을 앞선다. GDELT는 34.12/19.69/36.90/73.51이고, 특히 Hits@10이 DHyper 64.29 대비 14.34% 개선된다.
저자들은 shallow < GNN < derived-structure의 전반적 순서를 관찰한다. GDELT의 절대 성능은 다른 dataset보다 낮다. 논문은 높은 false-positive rate와 POLICE, GOVERNMENT 같은 abstract conceptual entity가 많아 특정 국가 맥락을 식별하기 어렵다는 기존 분석을 이유로 든다.
전체 benchmark 표 — ICEWS14
| Approach | MRR | Hits@1 | Hits@3 | Hits@10 |
|---|---|---|---|---|
| TTransE | 23.79 | 14.24 | 29.17 | 34.56 |
| HyTE | 29.60 | 18.15 | 30.15 | 45.37 |
| RE-NET | 45.77 | 37.98 | 49.07 | 58.87 |
| Glean | 42.20 | 36.86 | 47.68 | 52.39 |
| TeMP | 46.04 | 39.07 | 49.84 | 59.74 |
| RE-GCN | 41.06 | 38.09 | 50.37 | 62.44 |
| DACHA | 45.44 | 37.88 | 49.47 | 58.69 |
| TiRGN | 47.71 | 39.83 | 52.17 | 63.95 |
| TITer | 46.12 | 39.08 | 50.76 | 60.39 |
| EvoExplore | 42.30 | 40.68 | 52.37 | 65.94 |
| GTRL | 46.25 | 40.11 | 51.09 | 65.79 |
| DHyper | 56.15 | 43.76 | 65.46 | 85.89 |
| DECRL | 62.61 | 48.73 | 70.57 | 93.03 |
| HEART | 64.80 | 51.63 | 72.33 | 92.57 |
전체 benchmark 표 — ICEWS18
| Approach | MRR | Hits@1 | Hits@3 | Hits@10 |
|---|---|---|---|---|
| TTransE | 11.96 | 13.97 | 12.79 | 24.33 |
| HyTE | 28.60 | 16.86 | 25.64 | 41.86 |
| RE-NET | 42.25 | 33.81 | 44.98 | 52.72 |
| Glean | 37.11 | 34.15 | 42.56 | 47.35 |
| TeMP | 43.24 | 38.77 | 45.04 | 55.94 |
| RE-GCN | 40.53 | 37.59 | 44.34 | 57.42 |
| DACHA | 43.87 | 37.11 | 47.47 | 57.69 |
| TiRGN | 46.64 | 38.13 | 50.66 | 62.90 |
| TITer | 45.44 | 39.78 | 48.77 | 58.73 |
| EvoExplore | 40.80 | 40.05 | 50.07 | 58.35 |
| GTRL | 46.35 | 40.95 | 51.19 | 60.18 |
| DHyper | 54.22 | 42.16 | 63.26 | 75.38 |
| DECRL | 63.30 | 50.13 | 70.72 | 90.82 |
| HEART | 67.53 | 54.37 | 74.79 | 95.63 |
전체 benchmark 표 — GDELT
| Approach | MRR | Hits@1 | Hits@3 | Hits@10 |
|---|---|---|---|---|
| TTransE | 8.62 | 7.73 | 11.03 | 23.34 |
| HyTE | 11.20 | 8.39 | 14.23 | 28.79 |
| RE-NET | 17.55 | 11.73 | 18.14 | 35.52 |
| Glean | 15.60 | 10.35 | 17.61 | 37.40 |
| TeMP | 19.19 | 11.07 | 19.84 | 40.52 |
| RE-GCN | 19.22 | 10.80 | 21.09 | 43.65 |
| DACHA | 21.91 | 11.27 | 17.49 | 47.13 |
| TiRGN | 24.91 | 13.78 | 25.66 | 49.02 |
| TITer | TLE | TLE | TLE | TLE |
| EvoExplore | 18.50 | 10.74 | 19.45 | 42.07 |
| GTRL | 22.44 | 12.48 | 18.03 | 50.82 |
| DHyper | 31.94 | 18.85 | 33.73 | 64.29 |
| DECRL | 27.58 | 15.74 | 29.16 | 59.54 |
| HEART | 34.12 | 19.69 | 36.90 | 73.51 |
TLE는 논문 표기 그대로 single epoch가 24시간을 초과했다는 뜻이다.
15 · Appendix B.3
Country-only ICEWS에서는 이득이 이어지지만, 연 단위 WIKI/YAGO에서는 오히려 뒤진다
| Dataset | Method | MRR | Hits@1 | Hits@3 | Hits@10 |
|---|---|---|---|---|---|
| ICEWS14C | DECRL | 58.55 | 44.62 | 66.52 | 82.06 |
| HEART | 61.71 | 47.66 | 69.41 | 91.31 | |
| ICEWS18C | DECRL | 61.37 | 46.28 | 67.01 | 86.79 |
| HEART | 65.38 | 53.22 | 72.28 | 90.43 |
ICEWS14C improvement는 MRR 5.40%, Hits@1 6.81%, Hits@3 4.34%, Hits@10 11.27%이다. ICEWS18C에서는 6.53%, 15.00%, 7.86%, 4.19%이다.
전체 appendix 표 — ICEWS14C
| Approach | MRR | Hits@1 | Hits@3 | Hits@10 |
|---|---|---|---|---|
| TTransE | 11.79 | 13.24 | 19.97 | 24.88 |
| HyTE | 22.17 | 18.15 | 27.28 | 35.37 |
| RE-NET | 43.27 | 36.97 | 47.08 | 55.19 |
| Glean | 40.24 | 34.62 | 45.48 | 50.09 |
| TeMP | 44.17 | 37.37 | 47.78 | 55.66 |
| RE-GCN | 41.76 | 36.67 | 45.37 | 51.74 |
| DACHA | 44.26 | 37.59 | 44.18 | 53.19 |
| TiRGN | 44.73 | 38.13 | 49.77 | 60.91 |
| TITer | 44.86 | 39.37 | 48.84 | 55.79 |
| EvoExplore | 49.77 | 40.12 | 54.37 | 65.83 |
| GTRL | 50.95 | 40.31 | 52.09 | 64.89 |
| DHyper | 54.16 | 41.45 | 62.03 | 75.35 |
| DECRL | 58.55 | 44.62 | 66.52 | 82.06 |
| HEART | 61.71 | 47.66 | 69.41 | 91.31 |
전체 appendix 표 — ICEWS18C
| Approach | MRR | Hits@1 | Hits@3 | Hits@10 |
|---|---|---|---|---|
| TTransE | 9.84 | 10.29 | 11.04 | 18.89 |
| HyTE | 22.23 | 16.27 | 25.68 | 33.39 |
| RE-NET | 41.05 | 32.87 | 42.78 | 50.43 |
| Glean | 35.58 | 32.26 | 40.44 | 46.49 |
| TeMP | 43.08 | 36.07 | 43.18 | 53.03 |
| RE-GCN | 40.27 | 36.35 | 41.75 | 49.25 |
| DACHA | 40.11 | 36.11 | 46.17 | 52.37 |
| TiRGN | 43.57 | 37.23 | 47.67 | 54.44 |
| TITer | 44.07 | 38.85 | 46.44 | 49.79 |
| EvoExplore | 47.33 | 38.96 | 49.37 | 56.15 |
| GTRL | 49.33 | 40.15 | 53.39 | 60.74 |
| DHyper | 52.11 | 41.04 | 60.03 | 73.22 |
| DECRL | 61.37 | 46.28 | 67.01 | 86.79 |
| HEART | 65.38 | 53.22 | 72.28 | 90.43 |
| Approach | WIKI MRR | YAGO MRR |
|---|---|---|
| RE-GCN | 97.92 | 97.74 |
| TiRGN | 99.04 | 99.30 |
| DHyper | 99.38 | 99.31 |
| DECRL | 99.67 | 99.56 |
| HEART | 99.14 | 98.41 |
| Improvement | −0.53% | −1.16% |
저자들은 WIKI/YAGO가 1년 단위 timestamp라 event correlation의 fine-grained evolution을 포착하기 어렵다고 설명한다. HEART는 각각 99.14, 98.41의 높은 MRR을 유지하지만 DECRL보다 낮다.
16 · Ablation on ICEWS18
가장 큰 손실은 temporal smoothness와 cluster self-supervision을 제거할 때 나타난다
| Module | Variant | MRR | Hits@1 | Hits@3 | Hits@10 |
|---|---|---|---|---|---|
| EGC | w/ A-MLP | 67.07 | 54.73 | 73.94 | 92.49 |
| w/ C-MLP | 65.32 | 53.32 | 72.72 | 91.66 | |
| w/ CNN | 62.02 | 48.79 | 68.21 | 89.98 | |
| w/ Att | 64.07 | 50.65 | 70.89 | 92.27 | |
| w/ BERT | 65.38 | 53.36 | 71.32 | 91.45 | |
| EMEC | w/o alignment | 63.90 | 52.28 | 68.72 | 89.73 |
| w/o fusion | 63.21 | 49.20 | 70.96 | 91.27 | |
| w/ Att-fusion | 67.23 | 51.79 | 74.46 | 93.85 | |
| w/o L_temporal | 62.63 | 49.11 | 69.47 | 91.60 | |
| ECGMP | w/o ICE | 65.76 | 53.15 | 72.49 | 92.08 |
| w/o SSO | 62.05 | 48.45 | 68.41 | 91.27 | |
| w/o LS | 63.87 | 50.18 | 71.53 | 90.76 | |
| HEART | 67.53 | 54.37 | 74.79 | 95.63 | |
EGC. sum/concat MLP, CNN, self-attention, BERT event encoder보다 tensor-based subject–relation/object–relation interaction이 전반적으로 우수하다. 다만 A-MLP의 Hits@1 54.73은 HEART 54.37보다 약간 높다. 저자의 결론은 단일 metric이 아니라 네 metric의 전체 경향에 기반한다.
EMEC. alignment 제거, fusion 제거, temporal loss 제거는 모두 큰 손실을 만든다. self-attention fusion도 MRR 67.23으로 근접하지만 HEART의 simple average multi-step fusion이 최종적으로 더 좋다.
ECGMP. implicit correlation encoder를 없애 fully-connected uniform edge로 만들거나, self-supervised optimization을 제거하거나, learnable threshold 대신 0.2 fixed threshold를 쓰면 모두 성능이 내려간다. 특히 SSO 제거의 MRR은 62.05로 큰 하락이다.
17 · Case studies
t-SNE cluster와 NATO-Ukraine/Russia relation ranking에서 event-centric 효과를 해석한다
페이지 8 Figure 3은 ICEWS14C entity representation을 t-SNE로 시각화한다. red point는 국가이고 blue label은 Belt and Road Initiative와 관련된 East Asian nation—China, Thailand, Vietnam, Laos, Malaysia, Philippines, Cambodia—이다. 저자들은 DECRL보다 HEART representation이 더 compact하고 cluster separation이 크다고 관찰한다. 이를 intra-cluster variance 감소와 inter-cluster discrimination 증가의 정성적 증거로 본다.
Table 6은 두 relation prediction 사례의 top-5를 비교한다.
| Query | DECRL top-5 | HEART top-5 |
|---|---|---|
| (NATO, ?, Ukraine, 2014/12/16) | ✓ Consult Discuss by telephone Praise or endorse ✓ Appeal for aid Make statement | ✓ Consult ✓ Appeal for aid Make statement ✓ Host a visit Discuss by telephone |
| (Russia, ?, NATO, 2014/12/24) | Discuss by telephone ✓ Express intent to meet or negotiate ✓ Make statement Engage in negotiation Threaten | ✓ Make statement ✓ Express intent to meet or negotiate ✓ Make an appeal or request Discuss by telephone Consult |
2014년 러시아-우크라이나 ceasefire 이후 Ukraine은 NATO 지원을 구하고 Russia도 NATO와 추가 대화를 요청했다. 저자들은 두 sample에서 HEART가 correct relation을 더 많이, 더 높은 rank에 놓는다고 해석한다. DECRL이 두 번째 sample에서 “Threaten”을 top-5에 넣은 반면 HEART는 dialogue-oriented relation을 더 정확히 올린다.
18 · Running time
학습은 느리지만 inference는 조금 빠르다
| Approach | Training time (s) | Inference time (s) |
|---|---|---|
| DECRL | 4066.57 | 42.48 |
| HEART | 5285.58 | 37.06 |
ICEWS14 전체 epoch 기준 HEART training은 DECRL보다 약 30% 길지만 inference는 더 빠르다. 저자들은 event graph memoization으로 중복 computation을 cache하고, stronger modeling으로 convergence에 필요한 epoch가 줄어 training overhead를 일부 상쇄한다고 설명한다. 같은 dataset에서 MRR +3.50%, Hits@1 +5.95%를 얻기 때문에 performance-efficiency trade-off가 유리하다고 평가한다.
19 · Hyperparameter sensitivity
가장 민감한 두 축은 cluster 수와 history window 길이이다
페이지 12 Figure 4는 ICEWS14에서 하나의 hyperparameter만 바꾸고 나머지는 Table 7의 값을 고정해 sensitivity를 측정한다. N_c는 16까지 증가할수록 좋아지다가 이후 급격히 떨어진다. cluster가 너무 적으면 high-order correlation을 충분히 분리하지 못하고, 너무 많으면 cluster graph message passing에 noise가 늘어난다는 설명이다.
N_window는 12에서 최적이다. window가 너무 짧으면 역사 정보가 부족하고 너무 길면 redundant data가 많아진다. 반면 HEART layer 수, α, kNN neighbor k는 넓은 범위에서 performance fluctuation이 작다.
N_c와 N_window라고 결론내리고, N_layer, α, k는 상대적으로 non-critical하다고 보고한다.20 · Limits and interpretation
HEART의 성공 조건은 event density와 temporal resolution에 묶여 있다
Coarse time granularity. 가장 명확한 source-supported limitation은 WIKI/YAGO이다. 연 단위 timestamp에서는 HEART가 DECRL보다 MRR −0.53%, −1.16% 낮다. 논문 스스로 fine-grained event correlation을 포착할 시간 정보가 부족하기 때문이라고 설명한다.
Quadratic event graph cost. event graph construction은 O(N_x²D) 항을 가진다. event가 한 timestamp에 매우 많을수록 event-centric pairwise structure의 비용이 커진다. 논문은 memoization을 쓰지만 scalability experiment를 별도로 제시하지 않는다.
Training overhead. ICEWS14에서 HEART training time은 5285.58초로 DECRL 4066.57초보다 길다. inference는 빠르지만 학습비용이 사라지는 것은 아니다.
Benchmark family. 7개 dataset은 TKG community에서 널리 쓰이지만 ICEWS 계열 정치 사건이 네 개이고, GDELT도 event/news 기반이다. WIKI/YAGO는 coarse-grained counterexample을 제공한다. 다른 산업/과학 event graph에서 동일한 improvement가 유지되는지는 이 논문이 직접 검증하지 않는다.
“First event-centric” claim. 저자들은 TKG에서 HEART를 최초 event-centric approach라고 명시한다. 다만 related work에는 event representation/event graph 연구가 존재한다. novelty의 정확한 범위는 “TKG event prediction 안에서 event graph + evolving event-cluster graph를 중심 representation mechanism으로 사용하는 것”으로 읽는 편이 안전하다.
21 · Key takeaways
핵심 정리
22 · References
논문 참고문헌 전체
Event representation, TKG reasoning, graph structure
- Bai et al. Integrating deep event-level and script-level information for script event prediction. EMNLP, 2021.
- Bai et al. FTMF: Few-shot temporal knowledge graph completion based on meta-optimization and fault-tolerant mechanism. World Wide Web, 2023.
- Chen et al. DACHA: A dual graph convolution based temporal knowledge graph representation learning method using historical relation. TKDD, 2021.
- Chen and Chen. DECRL: A deep evolutionary clustering jointed temporal knowledge graph representation learning approach. NeurIPS, 2024.
- Chen et al. Local-global history-aware contrastive learning for temporal knowledge graph reasoning. ICDE, 2024.
- Dasgupta et al. HyTE: Hyperplane-based temporally aware knowledge graph embedding. EMNLP, 2018.
- Deng et al. Dynamic knowledge graph based multi-event forecasting. KDD, 2020.
- Ding et al. Event representation learning enhanced with external commonsense knowledge. EMNLP-IJCNLP, 2019.
- Du et al. A graph enhanced BERT model for event prediction. ACL Findings, 2022.
- Gao et al. Improving event representation via simultaneous weakly supervised contrastive learning and clustering. ACL, 2022.
- Granroth-Wilding and Clark. What happens next? Event prediction using a compositional neural network model. AAAI, 2016.
- Jin et al. Recurrent event network: Autoregressive structure inference over temporal knowledge graphs. EMNLP, 2020.
- Lee et al. Weakly-supervised modeling of contextualized event embedding for discourse relations. EMNLP Findings, 2020.
- Li et al. EventKGE: Event knowledge graph embedding with event causal transfer. Knowledge-Based Systems, 2023.
- Li et al. TiRGN: Time-guided recurrent graph network with local-global historical patterns for temporal knowledge graph reasoning. IJCAI, 2022.
- Li et al. Constructing narrative event evolutionary graph for script event prediction. IJCAI, 2018.
- Li et al. Search from history and reason for future: Two-stage reasoning on temporal knowledge graphs. ACL, 2021.
- Li et al. Temporal knowledge graph reasoning based on evolutional representation learning. SIGIR, 2021.
- Liu et al. TLogic: Temporal logical rules for explainable link forecasting on temporal knowledge graphs. AAAI, 2022.
- Lv et al. Integrating external event knowledge for script learning. COLING, 2020.
- Niu and Li. Logic and commonsense-guided temporal knowledge graph completion. AAAI, 2023.
- Pichotta and Mooney. Statistical script learning with multi-argument events. EACL, 2014.
- Saxena et al. Question answering over temporal knowledge graphs. ACL, 2021.
- Schlichtkrull et al. Modeling relational data with graph convolutional networks. ESWC, 2018.
- Shang et al. End-to-end structure-aware convolutional networks for knowledge base completion. AAAI, 2019.
- Sun et al. TimeTraveler: Reinforcement learning for temporal knowledge graph forecasting. EMNLP, 2021.
- Tang and Chen. GTRL: An entity group-aware temporal knowledge graph representation learning method. TKDE, 2024.
- Tang et al. DHyper: A recurrent dual hypergraph neural network for event prediction in temporal knowledge graphs. TOIS, 2024.
- Trivedi et al. Know-Evolve: Deep temporal reasoning for dynamic knowledge graphs. ICML, 2017.
- Wu et al. TeMP: Temporal message passing for temporal knowledge graph completion. EMNLP, 2020.
- Xia et al. MetaTKG++: Learning evolving factor enhanced meta-knowledge for temporal knowledge graph reasoning. Pattern Recognition, 2024.
- Xu et al. Temporal knowledge graph reasoning with historical contrastive learning. AAAI, 2023.
- Zhang et al. Temporal knowledge graph reasoning with dynamic memory enhancement. TKDE, 2024.
- Zhang et al. Temporal knowledge graph representation learning with local and global evolutions. Knowledge-Based Systems, 2022.
Supporting models, clustering, datasets, optimization
- Bergstra et al. Hyperopt: A python library for model selection and hyperparameter optimization. Computational Science & Discovery, 2015.
- Boschee et al. ICEWS coded event data. Harvard Dataverse, 2015.
- Cover and Hart. Nearest neighbor pattern classification. IEEE Transactions on Information Theory, 1967.
- Devlin. BERT: Pre-training of deep bidirectional transformers for language understanding. 2018.
- Kingma and Ba. Adam: A method for stochastic optimization. 2014.
- Kuhn. The Hungarian method for the assignment problem. Naval Research Logistics Quarterly, 1955.
- Leblay and Chekol. Deriving validity time in knowledge graph. WWW, 2018.
- Leetaru and Schrodt. GDELT: Global data on events, location, and tone, 1979–2012. ISA Annual Convention, 2013.
- Mahdisoltani et al. YAGO3: A knowledge base from multilingual wikipedias. CIDR, 2013.
- Pu et al. EM-IFCM: Fuzzy c-means clustering algorithm based on edge modification for imbalanced data. Information Sciences, 2024.
- Shahverdy et al. Driver behavior detection and classification using deep convolutional neural networks. Expert Systems with Applications, 2020.
- van der Maaten and Hinton. Visualizing data using t-SNE. Journal of Machine Learning Research.
- Ward et al. Comparing GDELT and ICEWS event data. Analysis, 2013.
- Weber et al. Event representations with tensor-based compositions. AAAI, 2018.
- Chen et al. SHA-SCP: A UI Element Spatial Hierarchy Aware Smartphone User Click Behavior Prediction Method. IEEE Transactions on Human-Machine Systems, 2025.
Primary source: Qian Chen and Ling Chen. Rethinking Temporal Knowledge Graph Representation Learning: From Entities to Evolutionary Event-Centric Clusters. Proceedings of the 32nd ACM SIGKDD Conference on Knowledge Discovery and Data Mining V.1 (KDD ’26), August 9–13, 2026, Jeju Island, Republic of Korea. DOI: 10.1145/3770854.3780173.