스키마 매칭은 오래된 데이터 통합 문제다. 이름이 비슷한 두 column 가운데 어느 것이 같은 의미인지 고르는 일은 간단해 보인다. 그러나 실제 스키마에서는 column name과 description만으로는 뜻을 확정할 수 없는 경우가 적지 않다. 같은 table의 sibling field, 더 넓은 section의 주제, cross-field dependency, 비슷해 보이는 다른 후보와의 차이가 의미를 결정한다.
여기서 흔한 대응은 두 극단으로 갈린다. local prompting은 prompt를 작게 유지하지만 결정에 필요한 outside-column evidence를 잃는다. global prompting은 전체 스키마를 넣지만 수백·수천 attribute가 있는 현실적 schema에서는 budget을 넘기고, long-context model조차 묻힌 정보를 항상 잘 쓰는 것은 아니다. ConStruM의 관점은 다르다. 문제를 “더 많은 context”가 아니라 budgeted evidence packing으로 재정의한다.
column 하나만 보면 애매하고, schema 전체를 보면 너무 많다
ConStruM은 이 양극단 사이에서 “결정에 필요한 증거를 제한된 prompt 안에 어떻게 선택하고 조직할 것인가”를 핵심 문제로 둔다.
CHARTTIME과 STORETIME: sibling 하나가 의미를 가르는 경우
논문 1쪽의 Figure 1은 MIMIC-III의 CHARTEVENTS table을 예로 든다. CHARTTIME은 문서상 “observation이 만들어진 시간” 정도로 설명되어 있어 observation_time과 recorded_time 사이에서 애매할 수 있다. 그런데 sibling column인 STORETIME은 “observation이 수동으로 입력되거나 검증된 시간”이라는 더 분명한 description을 가진다. STORETIME이 별도로 존재한다는 사실을 보면 CHARTTIME은 “관측이 발생한 시간” 쪽으로 해석하는 것이 자연스러워진다.
이 예의 핵심은 column 자체보다 column 밖의 비교 evidence가 의미를 정한다는 것이다. 실제 schema에서 그런 evidence는 local neighborhood, table/section membership, simple relation, cross-reference, 더 넓은 module scope 등 여러 수준에 흩어져 있을 수 있다.
Figure 2가 보여주는 세 가지 prompting 전략
Local
source-target pair 또는 작은 shortlist만 LLM에 보여준다. 저렴하고 단순하지만 outside-column clue가 필요한 ambiguity를 풀지 못할 수 있다.
Global
source와 target schema 전체를 한 prompt에 넣는다. 전체 그림은 보이지만 현실적 대형 schema에서는 context budget과 attention 문제가 생긴다.
Structure-guided
schema-derived structure에서 query-specific evidence만 추려 compact context pack으로 제공한다. ConStruM이 택한 위치다.
논문은 long-context 자체를 부정하지 않는다. 오히려 “long context냐 RAG냐”의 진영 선택에서 한 걸음 물러난다. 중요한 것은 어떤 mechanism을 쓰든 query마다 가장 진단적인 context를 budget 안에서 구성하는 것이라는 입장이다.
forced-choice column selection
source schema를 \(S=\{s_1,\ldots,s_m\}\), target schema를 \(T=\{t_1,\ldots,t_n\}\)라 두고, 각 column은 name과 optional description 같은 intrinsic metadata를 가진다. table/section membership, key relation, meaningful presentation order에서 유도되는 neighborhood 같은 lightweight structure도 사용할 수 있다.
평가 task는 각 source column \(s\)가 정확히 하나의 target column \(t\)를 고르는 forced-choice selection이다. global one-to-one constraint는 downstream application이 강제하지 않는 한 요구하지 않는다. pairwise score \(f(s,t)\)를 먼저 계산해 argmax하는 식으로 볼 수도 있지만, 논문은 직접 selection problem을 정의하며 이 설정에서는 accuracy가 micro-F1과 동일하다.
ConStruM의 novelty는 matcher를 바꾸는 데 있지 않고 decision evidence를 재설계하는 데 있다
후보 생성은 upstream method에 맡기고, 마지막 한 번의 LLM selection에 들어갈 context만 구조적으로 증강한다.
후보 shortlist \(C_0\)가 이미 있다고 가정한다
ReMatch와 Matchmaker처럼 많은 LLM matcher는 upstream retrieval이나 reasoning 절차가 다르더라도 마지막에는 query/source column \(s\)와 작은 candidate set \(C_0\)를 놓고 한 번의 LLM call로 target을 고른다. 이 final LLM은 prompt에 담긴 evidence를 합쳐 결정을 내리는 aggregator로 볼 수 있다.
그렇다면 candidate generator를 다시 설계하지 않고도 decision-time evidence를 개선할 수 있다. ConStruM은 바로 이 modularity를 이용한다. upstream은 그대로 두고, 최종 prompt에 query-specific context를 추가한다.
schema context는 반복 query에 재사용되는 자산이다
이 분리는 단순한 성능 최적화가 아니다. “context를 언제 계산하는가”를 schema preprocessing과 decision-time instantiation으로 분리해, 반복되는 matching workload에서 비용을 재사용 가능한 구조로 바꾼다.
coarse-to-fine scope와 explicit contrast
논문은 naive context organization이 두 가지 방식으로 실패한다고 본다. 첫째, bounded prompt에서는 모든 evidence를 넣을 수 없으므로 broad scope는 compact하게 요약하고 local detail은 diagnostic한 부분에만 budget을 배분해야 한다. 둘째, 서로 매우 비슷한 후보가 있을 때 similarity reasoning만으로는 결정을 못 내린다. 후보를 나란히 놓고 어떤 축에서 다른지를 명시해야 한다.
가까운 세부와 먼 범위를 한 prompt 안에 동시에 넣는 방법
context tree는 column을 leaf로 두고, local group–table–module–database root로 올라가며 점점 넓은 scope를 짧은 summary로 제공한다.
schema dump 대신 column-to-root lineage
논문 5쪽 Figure 4의 핵심은 단순하다. query column에서 시작해 root까지 올라가는 lineage를 따라가면 local detail과 global scope를 같은 길이의 prompt 안에 배치할 수 있다. tree internal node는 점점 넓은 column group을 natural-language summary로 압축한다. useful local relation이 있으면 “A가 B의 applicability를 제한한다”, “C가 D의 unit/type을 제공한다” 같은 짧은 relation snippet도 저장한다.
이 구조는 RAPTOR의 multi-granularity retrieval과 닮았지만 대상이 document가 아니라 schema다. column 하나를 완전히 고립시키지도 않고, database 전체를 통째로 붙이지도 않는다.
wide table은 네 단계 coarse-to-fine construction으로 나눈다
결과 group/span이 lowest-level group budget \(B\)보다 크면 재귀적으로 반복한다. \(B\)가 tree height의 stopping condition을 사실상 결정한다. small table은 table root와 column leaf만으로 얕은 tree가 되고, wide table은 여러 level을 가진 deep subtree가 된다.
table들을 semantic module로 다시 묶는다
모든 table root를 하나의 super-root에 바로 붙이면 database tree가 거의 flat해져 intermediate context가 사라진다. ConStruM은 각 table root summary를 embedding하고 hierarchical clustering으로 related table들을 bottom-up merge한다. 새 parent를 만들 때 LLM이 해당 group을 짧게 요약한다.
distance threshold \(\delta\)가 module granularity를 제어한다. 작은 \(\delta\)는 conservative merge로 더 깊은 hierarchy를 만들고, 큰 \(\delta\)는 aggressive merge로 더 얕은 hierarchy를 만든다. 최종 tree는 column leaf → within-table summary → table root → table cluster module → database root의 well-defined lineage를 가진다.
query마다 꺼내는 것은 tree 전체가 아니라 lineage + relation snippet이다
Column: [query column] Description: [short description] Q_CONTEXT Path to root (summaries): - [1] Query column leaf: [local summary] - [2] Parent summary: [broader span summary] - [3] Higher-level summary: [table/module-level scope] - [...] Dataset summary: [dataset-level scope cues] Relation snippet (optional): - [Field A qualifies Field B] - [Field C provides the unit/type for Field D]
이 pack이 query와 각 candidate에 대해 생성되어 final prompt의 structural grounding이 된다.
비슷한 후보를 찾는 것과 비슷한 후보를 구별하는 것은 다른 문제다
global similarity hypergraph는 “서로 닮은 column 묶음”을 재사용 가능한 구조로 만들고, final prompt에서 직접 대비하도록 한다.
pairwise similarity보다 confusable set이 필요하다
실제 schema에는 event_time, ingest_time처럼 이름과 description이 서로 겹치는 tight cluster가 있다. embedding retriever는 이런 후보를 함께 잘 찾아낼 수 있다. 문제는 그 다음이다. “가장 비슷한 것”을 고르라고 하면 모두 비슷해 선택 축이 드러나지 않는다. 그래서 ConStruM은 query와 candidate 양쪽에서 confusable group을 만들고, group-wise differentiation cue를 생성한다.
논문 7쪽 Figure 5처럼 hyperedge 하나는 두 개가 아니라 여러 column을 동시에 묶는다. 서로 겹치는 hyperedge를 허용하므로 한 column이 여러 confusion group에 속할 수 있다. 이 표현의 목적은 graph theory 자체보다 한 번에 비교해야 할 대안 집합을 materialize하는 데 있다.
embedding → thresholded links → connected components
target 후보 확장과 source query clarification을 동시에 지원한다
Target side. initial shortlist \(C_0\)의 상위 seed에서 threshold \(\tau\)를 만족하는 neighbor를 제한적으로 추가해 working set \(C\)를 만든다. 이렇게 하면 hard-to-distinguish candidate가 초기 shortlist에서 빠졌을 위험을 줄일 수 있다. working set 안에서 connected component를 다시 materialize해 필요한 group만 사용한다.
Source side. query column \(s\) 주변의 작은 confusable set도 materialize한다. description이 generic할 때 sibling 또는 near-duplicate와 비교하면 query 자체의 intended meaning이 더 분명해진다. Figure 1의 CHARTTIME/STORETIME이 바로 이 상황이다.
논문 7쪽 Figure 6은 \(\{C_5,C_8\}\) shortlist에서 overlapping group을 따라 \(\{C_3,C_4,C_6,C_7\}\)을 추가하고, 관련 hyperedge 2·3·4의 contrast cue만 final prompt에 넣는 예를 보여준다.
commonality + difference axis + one cue per candidate
Differentiation among candidates (Group #1): Summary: [All measure the same concept; differ in scope / units / timeframe.] - cid 12: [applies to subset A; unit = ...] - cid 37: [applies to subset B; unit = ...] - cid 58: [same scope; different timeframe → ...]
LLM은 candidate metadata만 보고 contrast를 만드는 것이 아니라 각 candidate의 tree-derived context도 함께 본다. 따라서 cue가 surface token 차이가 아니라 scope와 semantics 차이를 반영하도록 설계한다.
offline structure를 만들고, online에서는 작은 working set만 조립한다
Figure 3과 Section 6을 합치면 ConStruM의 interface는 다섯 단계의 decision-time augmentation으로 정리된다.
upstream shortlist는 외부, evidence packing부터 ConStruM
upstream matcher를 monolithic하게 다시 쓰지 않는다
초기 \(C_0\)는 ConStruM의 일부가 아니다. 저자 구현에서는 embedding cosine top-\(k\)를 쓰지만 ReMatch-style retrieval이나 Matchmaker가 만든 shortlist도 사용할 수 있다. 이 add-on 설계가 중요한 이유는 기존 matcher의 retrieval/shortlisting logic을 유지한 채 final decision prompt만 바꿔 효과를 분리해서 평가할 수 있기 때문이다.
context-stress benchmark에서는 구조화된 evidence가 압도적이었고, 표준 benchmark에서는 경쟁력 있는 수준을 유지했다
HRS-B는 context가 반드시 필요한 상황을 의도적으로 어렵게 만든 benchmark이고, MIMIC-2-OMOP은 더 일반적인 schema matching setting이다. 두 결과는 같은 의미가 아니다.
실험 구성
LLM
GPT-5.4를 tree construction, grouped differentiation, match decision에 사용한다.
Embedding
embedding-based retrieval에는 text-embedding-3-small을 사용한다.
HRS-B shortlist
initial top-\(k=20\), similarity threshold \(\tau=0.95\), candidate expansion을 켜고 final set은 최대 24개로 제한한다.
Context tree
ordered documentation chunk window 250 columns, recursion stop budget \(B=50\).
MIMIC-2-OMOP에서는 Matchmaker-style shortlist-and-decide pipeline을 사용하되 self-improvement는 꺼서 ConStruM의 decision-time evidence 효과를 분리한다.
Health and Retirement Study의 Employment section을 context-stress benchmark로 만든다
HRS는 장기간 longitudinal survey로 wave마다 variable name, description, cross-reference가 포함된 긴 sequential documentation을 제공한다. HRS-B는 2006–2022년 Section J(Employment)를 사용한다. cross-wave true match는 original HRS question identifier를 이용해 label을 자동 유도한다.
하지만 identifier가 그대로 있으면 너무 쉽다. 저자들은 이를 sequential variable ID로 치환하고 text 안의 identifier reference도 다시 쓴다. 그리고 모든 variable을 넣지 않고 Jaccard similarity로 문서상 멀리 떨어져 있으면서 매우 비슷한 variable이 있는 confusable region을 골라 matched variable을 남긴다. local adjacency만으로 푸는 shortcut을 억제해 broader, nonlocal context가 필요하도록 만든 것이다.
190개 forced-choice query에서 100.00%
| Role | Method | Accuracy (%) | Wilson 95% CI |
|---|---|---|---|
| Basic embedding | Embed-1NN | 33.16 | [26.86, 40.13] |
| No-context LLM | LLM rerank | 54.21 | [47.11, 61.14] |
| Local | ReMatch | 38.95 | [32.30, 46.04] |
| Broader context | RAG | 38.95 | [32.30, 46.04] |
| Broader context | GraphRAG | 60.53 | [53.44, 67.21] |
| Broader context | LC-1/3 | 90.00 | [84.91, 93.50] |
| Ablation | Tree only | 97.37 | [93.99, 98.87] |
| Ablation | Diff only | 96.84 | [93.28, 98.54] |
| Full | ConStruM | 100.00 | [98.02, 100.00] |
표의 숫자는 context-stress benchmark의 성격과 함께 읽어야 한다. Embed-1NN 33.16%는 surface embedding만으로 문제가 풀리지 않는다는 것을 보여준다. no-context LLM rerank는 54.21%로 올라가지만 여전히 제한적이다. GraphRAG는 60.53%, LC-1/3은 90.00%까지 올라가므로 broad context 자체의 가치도 확인된다. 그러나 구조화된 tree/context와 contrast를 사용한 full ConStruM이 190/190을 맞힌다.
tree와 differentiation은 각각 강하고, 둘을 합치면 완전해진다
Tree only는 97.37%, Diff only는 96.84%다. no-context LLM 54.21%와의 paired exact McNemar test는 각각 \(p=9.67\times10^{-23}\), \(p=1.74\times10^{-23}\)로 매우 강한 차이를 보인다. full system은 Diff only보다 3.16 point 높고 이 차이는 \(p=0.03125\)다. Tree only보다 2.63 point 높은 차이는 \(p=0.0625\)로 conventional 0.05 threshold를 넘으므로 저자들은 통상적 의미의 유의차라고 주장하지 않는다.
비용은 offline construction과 online decision으로 나뉜다
HRS 2022의 651 columns에 local relation annotation까지 포함해 context tree를 구축할 때 61 LLM calls, 총 925,797 tokens(745,563 prompt + 180,234 completion), wall-clock 1178초(약 19.6분)가 들었다. 일부 lowest-level annotation call이 병렬이어서 개별 call latency 합은 3555초(약 59.2분)다.
HRS-B 190 query 전체에서 online stage는 평균 3.00 LLM calls/query, 약 42K total tokens/query, 약 239초 end-to-end latency/query를 기록한다. offline tree는 year별로 한 번 구축해 이후 query에 재사용한다.
표준 benchmark에서는 59.69%, matched control 대비 +15.46 point
| Method | Accuracy (%) | Note |
|---|---|---|
| ConStruM | 59.69 | same shortlist + structured evidence |
| MM-style w/o ConStruM | 44.23 | authors' simplified control, self-improvement disabled |
| Matchmaker* | 62.20±2.40 | reported from Matchmaker paper |
| ReMatch* | 42.50 | reported result |
| LLM-DP* | 29.59±2.00 | reported result |
| SMAT* (50-50) | 10.85±6.00 | reported result |
| Jellyfish* 13B | 15.36±5.00 | reported result |
가장 공정한 비교는 Matchmaker 62.20±2.40과 직접 경쟁한다고 보기보다, 같은 shortlist를 쓰는 MM-style control 44.23%와 ConStruM 59.69%를 보는 것이다. structured context pack과 differentiation을 final decision에 추가했을 때 15.46 point가 오른다. Matchmaker의 full self-improvement pipeline과는 실험 구성 자체가 다르므로 저자도 이를 reproduction이라고 부르지 않는다.
HRS-B 26개 year-pair 전체 결과
원문 Appendix Table 3의 per-pair forced-choice accuracy를 그대로 옮기면 다음과 같다. full ConStruM은 모든 year-pair에서 100%를 기록한다.
Table 3 전체 보기 · 26 source→target year pairs
| Year pair | n | Embed-1NN | LLM rerank | ReMatch | RAG | GraphRAG | LC-1/3 | Tree only | Diff only | ConStruM |
|---|---|---|---|---|---|---|---|---|---|---|
| 2006→2008 | 23 | 21.74 | 100.00 | 47.83 | 73.91 | 95.65 | 91.30 | 100.00 | 95.65 | 100.00 |
| 2006→2010 | 19 | 31.58 | 94.74 | 31.58 | 89.47 | 84.21 | 89.47 | 100.00 | 100.00 | 100.00 |
| 2006→2012 | 5 | 60.00 | 60.00 | 80.00 | 80.00 | 60.00 | 80.00 | 100.00 | 100.00 | 100.00 |
| 2006→2014 | 4 | 25.00 | 100.00 | 50.00 | 50.00 | 100.00 | 75.00 | 75.00 | 100.00 | 100.00 |
| 2006→2016 | 4 | 75.00 | 100.00 | 50.00 | 0.00 | 50.00 | 75.00 | 100.00 | 100.00 | 100.00 |
| 2006→2018 | 6 | 16.67 | 0.00 | 33.33 | 0.00 | 0.00 | 83.33 | 100.00 | 100.00 | 100.00 |
| 2006→2020 | 6 | 16.67 | 16.67 | 0.00 | 0.00 | 16.67 | 83.33 | 100.00 | 100.00 | 100.00 |
| 2006→2022 | 6 | 33.33 | 16.67 | 33.33 | 0.00 | 33.33 | 66.67 | 100.00 | 100.00 | 100.00 |
| 2008→2010 | 35 | 20.00 | 82.86 | 31.43 | 71.43 | 77.14 | 97.14 | 100.00 | 91.43 | 100.00 |
| 2008→2012 | 4 | 50.00 | 75.00 | 50.00 | 25.00 | 75.00 | 100.00 | 100.00 | 100.00 | 100.00 |
| 2008→2014 | 4 | 75.00 | 75.00 | 75.00 | 50.00 | 75.00 | 100.00 | 100.00 | 100.00 | 100.00 |
| 2008→2016 | 4 | 75.00 | 75.00 | 50.00 | 0.00 | 75.00 | 100.00 | 100.00 | 100.00 | 100.00 |
| 2008→2018 | 4 | 25.00 | 0.00 | 50.00 | 0.00 | 25.00 | 100.00 | 100.00 | 100.00 | 100.00 |
| 2008→2020 | 4 | 25.00 | 25.00 | 25.00 | 0.00 | 25.00 | 100.00 | 100.00 | 75.00 | 100.00 |
| 2008→2022 | 3 | 66.67 | 33.33 | 66.67 | 0.00 | 66.67 | 100.00 | 100.00 | 100.00 | 100.00 |
| 2012→2014 | 2 | 50.00 | 100.00 | 50.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 |
| 2014→2016 | 4 | 50.00 | 100.00 | 50.00 | 0.00 | 100.00 | 100.00 | 100.00 | 100.00 | 100.00 |
| 2014→2018 | 4 | 25.00 | 0.00 | 50.00 | 25.00 | 25.00 | 100.00 | 100.00 | 100.00 | 100.00 |
| 2014→2020 | 4 | 50.00 | 25.00 | 0.00 | 0.00 | 25.00 | 75.00 | 100.00 | 100.00 | 100.00 |
| 2014→2022 | 4 | 50.00 | 0.00 | 50.00 | 0.00 | 25.00 | 50.00 | 100.00 | 100.00 | 100.00 |
| 2016→2018 | 6 | 16.67 | 0.00 | 0.00 | 0.00 | 16.67 | 66.67 | 83.33 | 100.00 | 100.00 |
| 2016→2020 | 6 | 66.67 | 0.00 | 0.00 | 0.00 | 66.67 | 100.00 | 83.33 | 100.00 | 100.00 |
| 2016→2022 | 8 | 12.50 | 0.00 | 37.50 | 0.00 | 37.50 | 87.50 | 87.50 | 100.00 | 100.00 |
| 2018→2020 | 4 | 75.00 | 0.00 | 75.00 | 0.00 | 75.00 | 100.00 | 100.00 | 75.00 | 100.00 |
| 2018→2022 | 4 | 25.00 | 0.00 | 0.00 | 0.00 | 25.00 | 100.00 | 100.00 | 75.00 | 100.00 |
| 2020→2022 | 13 | 30.77 | 15.38 | 69.23 | 23.08 | 30.77 | 92.31 | 100.00 | 100.00 | 100.00 |
| Total | 190 | 33.16 | 54.21 | 38.95 | 38.95 | 60.53 | 90.00 | 97.37 | 96.84 | 100.00 |
ConStruM은 “schema matching을 LLM으로 한다”보다 더 데이터베이스적인 질문을 던진다
무엇을 prompt에 넣을지 구조적으로 고르는 retrieval/indexing problem이 최종 LLM reasoning quality를 좌우한다는 점이 핵심이다.
외부 knowledge를 넣기보다 schema 내부 구조를 evidence index로 만든다
classical schema matching은 COMA/COMA++처럼 linguistic, constraint, structural signal을 결합해 왔다. embedding/neural 계열은 scalable candidate generation과 task-specific matching을 강화했고, LSM, Unicorn, Jellyfish 등은 pretrained/local model을 data integration task에 맞춘다. LLM prompting 계열은 ReMatch, Matchmaker, Schemora, KCMF, GRAM, LLMATCH, Magneto 등으로 확장됐다.
ConStruM의 차이는 monolithic matcher가 아니라 fixed-budget final decision을 위한 structure-guided context module이라는 데 있다. KG-RAG4SM이나 SMoG처럼 external graph를 query하는 흐름과도 다르다. ConStruM은 schema 자체의 hierarchy와 similarity structure를 mining해 reusable evidence index를 만든다. GraphRAG/LightRAG의 graph organization과 RAPTOR의 hierarchical summaries를 schema matching에 맞게 가져온 셈이다.
selection condition이 아니라 prompt evidence pack
과거 “contextual schema matching” 연구에서는 correspondence가 어떤 predicate/selection condition 아래 valid한지를 context라고 부르기도 했다. ConStruM에서 context는 다른 의미다. multi-level neighborhood/scope summary, lightweight relation, contrast cue를 포함한 budgeted query-specific evidence pack이다.
TURL과 Starmie처럼 context-heavy data task도 관련 있지만, 이들은 table understanding이나 data-lake unionability를 대상으로 한다. ConStruM은 value-populated table 자체보다 schema metadata를 대상으로 cross-schema correspondence를 고르고, final LLM shortlist decision을 fixed prompt budget 아래 지원한다.
context가 필요 없는 문제에서는 얻는 것이 적을 수 있다
저자들이 명시한 첫 한계는 적용 범위다. column과 table description만으로 이미 충분히 discriminative한 schema에서는 extra structural context의 incremental benefit가 작을 수 있다. HRS-B는 의도적으로 context-critical case를 모아 만든 benchmark이므로 100%라는 숫자를 일반적인 모든 schema matching workload에 그대로 투영해서는 안 된다.
둘째, context tree와 similarity hypergraph는 structural context의 잠재력을 완전히 사용한 것이 아니다. tree는 hierarchy를 잘 압축하지만 richer relational pattern은 제한적으로만 relation snippet에 담는다. 저자들은 future work로 context module을 tree에서 richer relational representation, 예컨대 hypergraph로 확장하고, budgeted mechanism으로 decision-critical evidence를 선택하겠다고 제시한다.
셋째, schema graph 자체의 connectivity pattern과 hub-like versus isolated column 같은 node role도 추가 structural hint로 사용할 수 있다고 본다.
LLM data system의 병목은 reasoning model만이 아니라 evidence architecture다
이 논문을 schema matching 바깥으로 읽으면 한 가지 설계 원리가 드러난다. LLM의 context window가 커져도 “관련 evidence를 어느 granularity로, 어떤 대비 구조로, 어떤 순서로 보여줄 것인가”라는 data-system problem은 사라지지 않는다. 오히려 model이 강해질수록 prompt 안의 evidence organization이 final decision variance를 크게 만들 수 있다.
이 해석은 논문의 reported benchmark 결과를 넘어선 분석이다. 다만 context tree, similarity group, offline/online split, fixed-budget final decision이라는 설계 자체가 이 방향을 명확하게 뒷받침한다.