멀티모달 entity linking은 문장 속 mention과 이미지가 가리키는 실제 entity를 knowledge base에서 찾아 연결하는 문제다. 문제는 제품 세계의 entity가 서로 너무 닮았다는 데 있다. 같은 ASUS ROG laptop이라도 화면 크기, SSD 용량, 색상처럼 몇 개의 세부 속성만 다를 수 있다. 이름과 설명, 이미지의 전역적 유사성만으로는 마지막 한 걸음을 넘기 어렵다.
AMELI는 이 지점을 정면으로 다룬다. entity를 더 이상 하나의 atomic symbol로 보지 않고, title + description + images + structured attributes의 묶음으로 본다. mention 쪽에서도 review text와 image에서 fine-grained attribute를 추론해 candidate entity의 속성과 대조한다. 논문의 주장은 간결하다. 멀티모달 정보가 많아질수록 더 많은 feature를 넣는 것보다, 무엇이 서로 다른지 말해주는 속성이 더 중요해질 수 있다.
같은 종류의 물건을 구별하는 데는 category보다 차이를 만드는 속성이 필요하다
기존 multimodal entity linking이 text와 image를 함께 보면서도 entity의 structured meta-information을 충분히 활용하지 않았다는 것이 출발점이다.
멀티모달이라고 해서 자동으로 fine-grained한 것은 아니다
전통 entity linking은 text mention을 Wikipedia나 Wikidata 같은 KB entity에 연결하는 문제로 발전해 왔다. 최근에는 mention과 entity에 image를 더해 social media, news, video 같은 환경으로 확장됐다. 하지만 AMELI 저자들은 기존 연구가 KB entity를 대체로 atomic symbol처럼 다루며, 각 entity의 meta-information인 attribute를 의사결정의 중심에 놓지 않았다고 지적한다.
E-commerce에서는 이 생략이 특히 치명적이다. 사용자가 “ASUS laptop”이라고 쓰고 노트북 사진을 올렸다고 하자. candidate가 같은 모델 family의 여러 variant라면 title similarity와 overall image similarity가 모두 높다. 이때 Screen Size: 14 inches, SSD: 1000 GB, Color: White 같은 attribute가 variant-level linking을 결정한다.
Table 1: AMELI는 mention/entity 양쪽의 text·image에 structured attribute까지 더한다
| Dataset | Attribute | Mention Images | Mention Text | Entity Images | Entity Text |
|---|---|---|---|---|---|
| Zhou et al. (2021) | ✗ | ✓ | ✓ | ✓ | ✓ |
| WikiDiverse | ✗ | ✓ | ✓ | ✓ | ✓ |
| WIKIPerson | ✗ | ✓ | ✗ | ✓ | ✓ |
| OVEN-Wiki | ✗ | ✓ | ✗ | ✓ | ✓ |
| ZEMELD | ✗ | ✓ | ✓ | ✓ | ✓ |
| MEL_Tweets | ✗ | ✓ | ✓ | ✓ | ✓ |
| M3EL | ✗ | ✓ | ✓ | ✓ | ✓ |
| ✗ | ✓ | ✓ | ✓ | ✓ | |
| SnapCaptionsKB | ✗ | ✓ | ✓ | ✓ | ✓ |
| VTKEL | ✗ | ✓ | ✓ | ✗ | ✓ |
| Guo & Barbosa (2018) | ✗ | ✗ | ✓ | ✗ | ✓ |
| Zeshel | ✗ | ✗ | ✓ | ✗ | ✓ |
| AMELI | ✓ | ✓ | ✓ | ✓ | ✓ |
entity linking과 attribute extraction이 만나는 지점
textual entity linking은 dense retrieval, QA, autoregressive generation, knowledge-enhanced representation 등으로 발전했다. multimodal entity linking은 noisy social image를 걸러내거나, contrastive learning으로 text-image entity representation을 만들거나, entity name을 직접 생성하는 방향을 탐색했다.
한편 product attribute extraction은 title/description에서 sequence tagging 또는 QA로 attribute value를 뽑거나, product image의 visual clue까지 결합하는 별도 연구축을 만들었다. AMELI는 이 두 흐름을 연결해, noisy review에서 추론한 attribute를 entity linking의 disambiguation evidence로 쓴다.
benchmark를 만드는 일부터 fine-grained ambiguity를 보존해야 한다
Best Buy의 product page와 user review를 연결하되, variant-level gold link, 이미지 정제, informative-review filtering, mention verification을 거친다.
Best Buy product catalog + multimodal reviews
저자들은 product name, category, description, structured attribute category/value, multiple images가 정형적으로 존재하고 사용자 review가 text와 image를 함께 포함할 수 있다는 이유로 Best Buy를 선택한다. Requests로 page 정보를 수집하고, attribute가 button click 뒤에 노출되는 경우 Selenium을 사용한다.
초기 수집 규모는 38,329 product entities와 6,500,078 reviews다. 이 원시 데이터에서 실제 entity linking에 적합한 review만 남기는 것이 다음 단계다.
“review가 있다”와 “entity를 식별할 정보가 있다”는 다른 말이다
네 가지 signal로 “링킹 가능성”을 먼저 평가한다
각 review-product pair에서 (1) string match로 세는 mentioned attribute 수, (2) CLIP 기반 review-product image max similarity, (3) SBERT 기반 review text-product description similarity, (4) SBERT 기반 review text-product title similarity를 계산한다. 500쌍을 사람이 informative/uninformative로 annotation하고 threshold를 탐색한다.
논문은 이 threshold method가 informative review 예측에서 precision 85%, recall 82%를 얻었다고 보고한다. footnote에서는 SVM 등 classifier와 비교했다고 밝히며 threshold method가 가장 높은 accuracy를 냈다고 적는다.
product title/category에서 후보 명사를 만들고 review n-gram과 맞춘다
spaCy로 product title과 category의 root word를 분석해 product-name candidate를 만든다. review text의 1-gram부터 6-gram까지 순회하며 surface form 또는 lemma가 candidate와 맞으면 mention 후보로 삼는다. review에 여러 후보가 있으면 SBERT로 target product title과 가장 유사한 mention을 고른다.
200 review의 manual assessment에서 이 mention detector는 91.9% accuracy를 기록한다. 이후 전체 review에 적용하고 mention이 없는 review를 제거하며, Test set mention은 한 annotator가 모두 수동 검증·수정한다.
Table 2: mention benchmark와 multimodal KB
| Statistic | Train | Dev | Test |
|---|---|---|---|
| # Reviews / Mentions | 12,148 | 1,846 | 2,741 |
| # Review Images | 21,780 | 3,369 | 5,323 |
| Avg. # Images / Review | 1.79 | 1.83 | 1.94 |
| Avg. # Attributes / Review | 1.22 | 1.62 | 3.54 |
abstract와 Data Statement는 benchmark를 16,735 mentions + 30,472 review images + 34,690 entities + 177,873 entity images + 798,216 attributes로 정의한다. Table 2의 split 합도 12,148 + 1,846 + 2,741 = 16,735다.
강한 negative를 text와 image 두 방향에서 만든다
사람에게 34,690 entity 전체를 비교시키는 대신 gold product를 기준으로 SBERT title similarity top-10 negative와 CLIP image similarity top-10 negative를 합치고 gold entity를 포함한 candidate set을 만든다. 12 annotator가 target을 선택하며 overall accuracy는 79.73%다. 어떤 annotator라도 target을 맞히지 못한 review는 제거해 최종 Test 2,741건을 구성한다.
Test set에서는 한 annotator가 mention에 드러난 모든 Gold Attributes도 직접 label한다. 이는 뒤의 System Attribute와 Gold Attribute 비교를 가능하게 한다.
먼저 3만여 entity를 text와 image로 줄인다
two-step pipeline의 첫 단계는 gold entity가 포함될 가능성이 높은 top-K를 만드는 일이다. 이 단계의 miss는 이후 disambiguation이 고칠 수 없다.
review mention 하나를 multimodal KB의 유일한 entity로 연결한다
review \(r\)는 text \(t_r\), 여러 image \(\bar V_r\), entity mention \(m_r\)로 구성된다. KB entity \(e_j\)는 title, description \(d_{e_j}\), 여러 image \(\bar V_{e_j}\), attribute set \(\bar A_{e_j}\)를 가진다. title도 attribute의 하나로 취급한다.
전체 task는 Candidate Retrieval과 Entity Disambiguation으로 나뉜다. 첫 단계가 top-K 후보를 만들고, 두 번째 단계가 그 안에서 gold \(e^+\)를 선택한다. retrieval error 때문에 gold가 candidate set에 없을 수 있다는 사실을 모델 정의에 명시한다.
SBERT + InfoNCE
review text와 entity text를 SBERT로 encode해 cosine similarity를 계산한다. entity text에는 preliminary experiment 결과에 따라 title, description, attributes를 함께 붙인다.
SBERT는 gold entity를 positive로, 다른 product category의 standard negatives와 same-batch candidates를 negative로 두는 InfoNCE objective로 fine-tune한다.
CLIP + max pairwise similarity
review image와 entity image를 CLIP으로 encode하고 InfoNCE로 fine-tune한다. review와 entity가 multiple image를 가질 수 있으므로 모든 image pair의 cosine similarity 가운데 최대값을 retrieval visual score로 사용한다.
text와 image similarity를 weighted sum으로 합친다
\(\lambda\)는 Dev set에서 탐색한다. 이 score로 entity를 rank하고 top-K를 candidate로 넘긴다.
Table 3: multimodal fine-tuning이 Recall@100 95.11%까지 올린다
| Modality | Method | R@1 | R@10 | R@20 | R@50 | R@100 |
|---|---|---|---|---|---|---|
| T | Pre-trained SBERT | 19.52 | 46.63 | 57.06 | 71.18 | 82.52 |
| V | Pre-trained CLIP | 14.45 | 39.00 | 47.25 | 59.25 | 68.77 |
| T+V | Pre-trained CLIP/SBERT | 27.14 | 59.76 | 67.75 | 77.49 | 82.60 |
| T | Fine-tuned SBERT | 32.65 | 66.65 | 76.87 | 87.34 | 93.32 |
| V | Fine-tuned CLIP | 28.06 | 62.82 | 71.76 | 80.48 | 86.25 |
| T+V | Fine-tuned CLIP/SBERT | 48.12 | 85.84 | 90.26 | 93.69 | 95.11 |
저자들은 fine-tuning 뒤 Recall@10이 평균 25.3% 개선됐다고 보고한다. text와 image를 결합한 모델은 single modality보다 높으며, Recall@100 95.11%는 대부분의 gold entity를 100개 후보 안에 넣는다는 뜻이다. 그러나 실제 disambiguation 실험은 K=10을 사용하므로 Recall@10 85.84%가 error propagation의 현실적인 상한을 더 직접적으로 설명한다.
Table 7: title + description + attributes의 조합이 Recall@100 최고
| Entity text field | R@1 | R@10 | R@20 | R@50 | R@100 |
|---|---|---|---|---|---|
| Title | 12.29 | 38.12 | 48.92 | 63.81 | 75.88 |
| Desc | 14.85 | 40.42 | 50.31 | 63.81 | 74.17 |
| Attri | 12.81 | 36.88 | 47.46 | 63.77 | 77.67 |
| Title+Attri | 16.42 | 43.56 | 54.10 | 67.68 | 79.53 |
| Title+Desc | 19.08 | 47.21 | 57.10 | 69.21 | 79.82 |
| Attri+Desc | 17.62 | 45.06 | 55.45 | 69.06 | 80.48 |
| Title+Desc+Attri | 19.52 | 46.63 | 57.06 | 71.18 | 82.52 |
모든 cutoff에서 동일 조합이 최고인 것은 아니다. R@10과 R@20에서는 Title+Desc가 약간 높다. 저자들은 overall preliminary result를 바탕으로 candidate retrieval entity text에 title+description+attributes를 사용한다.
후보를 찾은 뒤에는 속성을 추출하고, 문장과 속성이 서로 함의하는지 묻는다
Figure 3의 오른쪽 절반은 NLI 기반 text module과 contrastive image module, retrieval score를 최종 weighted fusion으로 결합한다.
정확한 문자열, OCR, LLM QA, language completion을 함께 쓴다
OCR
review image 안의 brand/model text를 EasyOCR로 읽는다.
String Match
top-K candidate가 가진 attribute value가 review text에 그대로 나타나면 해당 attribute를 추출한다.
ChatGPT multiple-choice QA
“16 GB”처럼 표면형이 다를 때 attribute category를 질문, candidate value를 option으로 만들어 추출한다. 비용 때문에 11개 common category에만 적용한다.
GPT-2 zero-shot completion
나머지 attribute category는 Attribute Value Extraction: review / key: prompt로 completion한다.
어떤 extractor든 top-K candidate entity의 실제 attribute value와 맞는 결과만 유지한다. 이렇게 얻은 set이 System Attribute다. System Attribute와 맞지 않는 candidate는 filter한다. Train/Dev에는 manual review-attribute label이 없기 때문에 gold entity의 attribute와 맞지 않는 System Attribute를 제거해 training용 Gold Attribute를 구성한다.
“review가 이 속성을 함의하는가, 모순되는가”로 바꾼다
review가 “more of a gray-toned light pink”라고 말하면 pink product attribute는 함의하고 black attribute는 모순한다. 저자들은 이 intuition을 NLI로 formalize한다. mention의 System Attribute가 cover하는 category만 골라 candidate의 attribute value와 entity description을 review text와 pair로 만든다.
각 pair는 DeBERTa encoder를 통과한다. 여러 attribute representation과 description representation을 concatenate해 MLP가 text disambiguation score를 출력한다.
training은 top-K candidate 가운데 gold entity의 softmax probability를 높이는 cross-entropy objective를 사용한다.
global CLIP feature를 task-oriented space로 adapt한다
review image와 candidate entity image를 CLIP으로 encode한 뒤 feed-forward adapter와 residual connection을 적용한다. disambiguation에서는 preliminary experiment 결과에 따라 review/entity 각각 한 장씩, 서로 CLIP similarity가 가장 높은 image pair를 선택한다.
batch의 다른 entity를 negative로 쓰는 contrastive loss로 fine-tune해 fine-grained visual discrimination을 학습한다.
text NLI + image + retrieval score의 세 항을 합친다
\(\lambda_1,\lambda_2\)는 Dev set에서 tune한다. candidate retrieval을 완전히 버리지 않고 final disambiguation에도 prior score로 남긴다는 점이 특징이다.
Appendix E: attribute extraction을 QA와 completion으로 바꾼다
ChatGPT·Vicuna·LLaVA용 multiple-choice template는 review text/image를 제시하고 “이 mention의 product title 또는 attribute category가 무엇인가”라고 묻는다. top-10 candidate value를 A, B, C… option으로 둔다. GPT-2용 template는 few-shot demonstration 뒤 review text와 attribute category를 주고 value를 completion하게 한다.
attribute가 가장 큰 차이를 만들지만, machine-human gap은 여전히 크다
Disambiguation은 gold가 top-10에 있는 subset에서, End-to-End는 retrieval miss까지 포함한 전체 흐름에서 평가한다.
Attribute-aware multimodal model의 핵심 결과
| Modality | Attribute | Method | Disambig. F1 | End-to-End F1 |
|---|---|---|---|---|
| - | No | Random | 10.00 | 8.58 |
| V | No | V2VEL | 19.27 | 16.78 |
| T | No | V2TEL | 19.57 | 17.07 |
| T+V | No | V2VTEL | 31.37 | 30.22 |
| T+V | No | LLaVA | 23.33 | 20.03 |
| T+V | No | GHMFC | 12.52 | 12.11 |
| T+V | Filter | GHMFC* | 23.25 | 21.78 |
| T+V | No | WikiDiverse | 12.95 | 10.93 |
| T+V | Filter | WikiDiverse* | 24.57 | 20.48 |
| T+V | No | Our w/o Attribute | 52.53 | 44.85 |
| T | System | Our w/o Image | 44.40 | 38.52 |
| V | System | Our w/o Text | 42.61 | 36.64 |
| T+V | System | Our Approach | 60.30 | 51.54 |
| T+V | Gold | Our Approach | 73.08 | 62.87 |
| T+V | No | Human | 80.00 | 74.00 |
System Attribute를 쓰는 full model은 60.30% Disambiguation F1, 51.54% End-to-End F1로 표의 model baseline을 앞선다. attribute를 제거한 같은 architecture는 52.53/44.85다. 단순 post-filter만 추가해도 GHMFC와 WikiDiverse가 크게 오르는 점은 attribute가 architecture-specific trick보다 일반적인 discriminative evidence임을 시사한다.
Gold Attribute를 쓰면 73.08/62.87까지 상승한다. 이것은 attribute 활용 모듈의 잠재력과 동시에 attribute extraction이 현재 병목이라는 사실을 보여준다.
51.54%와 74.00% 사이
human performance는 50 review씩 random sample을 만들고 10 candidates를 제시해 두 annotator가 manual linking한 결과다. 두 사람이 모두 true label을 골라야 correct로 인정한다. Fleiss \(\kappa\)는 Disambiguation set에서 0.69, End-to-End set에서 0.71이다.
attribute label을 완벽하게 준다고 해도 human 74%에 아직 닿지 않는다. 따라서 gap은 extraction만의 문제가 아니라 reasoning, fine-grained vision, retrieval에도 남아 있다.
Table 8: 네 extractor를 합치면 F1 76.39%
| Extractor | Precision | Recall | F1 |
|---|---|---|---|
| String Match | 97.82 | 37.88 | 54.61 |
| Zero-shot GPT-2 | 92.38 | 37.57 | 53.41 |
| Zero-shot ChatGPT | 64.57 | 17.29 | 27.28 |
| OCR | 98.46 | 12.54 | 22.24 |
| Match+GPT2+ChatGPT+OCR | 94.33 | 64.18 | 76.39 |
| Few-shot GPT-2 | 90.04 | 44.39 | 59.47 |
| Few-shot Vicuna | 74.05 | 48.04 | 58.28 |
개별 extractor는 precision은 높지만 recall이 낮다. 서로 다른 실패 양상을 합치는 ensemble이 76.39 F1로 크게 상승한다. Vicuna는 계산 비용 때문에 max token length 64로 제한했고, 저자들은 이것이 성능을 낮췄을 수 있다고 적는다.
오답을 보면 attribute-aware라는 말의 실제 난도가 드러난다
50개 System-Attribute 오류를 분석하면 extraction, reasoning, fine-grained vision, retrieval이 서로 다른 비율로 실패한다.
10% + 18% + 32% + 26%
Extraction. review의 informal wording, idiom, typo, 3만개가 넘는 attribute value space 때문에 key attribute를 놓친다. “10 programmable buttons”처럼 숫자 하나가 gold를 가르는 경우가 대표적이다. logo recognition과 더 나은 OCR도 개선 여지다.
Reasoning. System Attribute에 필요한 signal이 있어도 모델이 abundant multimodal context 속에서 distinctive attribute를 놓친다. phone compatibility나 carafe capacity 같은 조건을 reasoning에 명시적으로 사용해야 한다.
Fine-grained vision. computer case의 미세 패턴, refrigerator dispenser의 internal/external location처럼 global image similarity로는 어려운 texture·part-level 차이가 있다. 저자들은 visual attribute로 attention을 guide하는 방향을 제안한다.
Retrieval. 26%의 오류는 gold가 top-10에 없어 disambiguation으로 복구할 수 없다. attribute-aware retrieval과 fine-grained image retrieval이 필요하며, review image 중 irrelevant image가 similarity 계산에 noise를 넣는 문제도 관찰한다.
숫자·플랫폼·용량·texture·부품 위치가 gold를 가른다
Appendix Figure 7은 여섯 사례를 나란히 보여준다. gaming mouse는 17 buttons 대 10 buttons, headset은 phone/Android compatibility, coffee maker는 12 cups 대 14 cups, PC case는 미세 외형 pattern, refrigerator는 water dispenser의 internal/external location이 핵심 clue다. 또 ROCCAT mouse 사례는 OCR로 image 속 “Burst”를 읽어 model difference를 포착할 수 있음을 보여준다.
Test-set candidate verification의 12명 결과
Table 5 전체 보기
| Annotator | #Correct | #Finished | Accuracy |
|---|---|---|---|
| 1 | 244 | 330 | 73.94 |
| 2 | 256 | 330 | 77.58 |
| 3 | 274 | 330 | 83.03 |
| 4 | 240 | 330 | 72.73 |
| 5 | 270 | 324 | 83.33 |
| 6 | 253 | 330 | 76.67 |
| 7 | 124 | 170 | 72.94 |
| 8 | 290 | 330 | 87.88 |
| 9 | 222 | 330 | 67.27 |
| 10 | 295 | 330 | 89.39 |
| 11 | 272 | 330 | 82.42 |
| 12 | 285 | 330 | 86.36 |
| Overall | 3025 | 3794 | 79.73 |
가장 많은 product category도 전체의 2.44%에 불과하다
Table 6 상위 10 category
| Category | # Products | Percentage |
|---|---|---|
| All Refrigerators | 847 | 2.44 |
| Action Figures (Toys) | 730 | 2.10 |
| Dash Installation Kits | 682 | 1.97 |
| Wall Mount Range Hoods | 680 | 1.96 |
| Nintendo Switch Games | 628 | 1.81 |
| Gas Ranges | 603 | 1.74 |
| Building Sets & Blocks (Toys) | 576 | 1.66 |
| Nintendo Switch Game Downloads | 574 | 1.65 |
| PC Laptops | 554 | 1.60 |
| Cooktops | 547 | 1.58 |
모든 baseline을 같은 top-K 후보 조건에서 비교한다
V2VEL은 image-image ResNet 기반 visual linker, V2TEL은 CLIP으로 entity text와 mention image를 encode하는 모델, V2VTEL은 두 모델을 retrieval-then-rerank로 결합한다. GHMFC는 text-guided visual attention과 visual-guided text attention, gated fusion, contrastive training을 사용한다. WikiDiverse baseline은 patch-level image와 token-level text를 self-attention transformer로 융합한다.
GHMFC*와 WikiDiverse*는 AMELI가 제안한 Attribute Filter를 단순 post-process로 붙인 버전이다. LLaVA는 동일한 top-10 retrieval 뒤 multiple-choice QA 형식으로 candidate title을 직접 고르게 한다. V2VEL/V2TEL/V2VTEL/GHMFC/WikiDiverse는 모두 AMELI에서 fine-tune해 같은 top-K setting에서 비교한다.
Table 9: mention context에는 다른 MEL dataset에서도 attribute가 나타난다
| Dataset | Extractor | # Attributes | # Mention Context | Attr / Mention |
|---|---|---|---|---|
| AMELI | System | 27,533 | 16,735 | 1.65 |
| AMELI Test | Human | 9,716 | 2,741 | 3.54 |
| MELBench-Richpedia | System | 36,705 | 17,800 | 2.06 |
| Sun (2017) | System* | 2,198 | 1,000 | 2.20 |
| Hu & Liu (2004) | System* | 348 | 314 | 1.11 |
AMELI의 자동 extractor는 review당 평균 1.65 attributes를 찾지만 human label은 3.54다. 저자들은 이 차이를 더 나은 extractor가 아직 회수하지 못한 attribute가 많다는 증거로 해석한다. Richpedia에서도 text만으로 평균 2.06 attributes가 추출돼, attribute-aware EL이 e-commerce 밖에도 적용될 여지가 있다고 주장한다.
실험 자원은 가볍지 않다
candidate retrieval model 한 번의 training은 1× NVIDIA A40에서 10시간, entity disambiguation model 한 번은 4× NVIDIA A40에서 7시간이 걸린다고 보고한다. disambiguation learning rate search space는 \(\{10^{-2},10^{-3},10^{-4},10^{-5},5\times10^{-3},5\times10^{-4}\}\), batch size는 \(\{12,16,20,24,32\}\)다.
attribute를 넣었다고 문제를 다 푼 것이 아니라, 어디가 아직 비어 있는지 더 선명해졌다
논문은 NLI 기반 활용 방식이 attribute의 잠재력을 충분히 쓰지 못한다고 인정하고 attribute-aware encoding, zero-shot learning, retrieval을 미래 방향으로 제시한다.
NLI는 attribute를 사용하는 첫 방식이지 마지막 방식이 아니다
저자들은 현재 NLI-based framework가 attribute의 잠재력을 완전히 활용하지 못한다고 명시한다. future work로 attribute-aware encoding, attribute-based zero-shot learning, attribute-aware retrieval을 제안한다. Remaining Challenges의 결과까지 합치면 extraction뿐 아니라 retrieval 단계부터 attribute를 넣고, vision에서도 part/texture를 attribute-guided하게 보는 방향이 자연스럽다.
PII와 offensive review에 대한 처리
저자들은 ACM Code of Ethics를 따랐으며 notable harmful effect, environmental impact, fairness, privacy, security risk를 발견하지 못했다고 보고한다. dataset에는 name, address, phone number 같은 sensitive personally identifiable information이 없다고 설명한다. user review에 offensive claim이 있을 수 있어 profanity-containing review를 preprocessing에서 제거했다.
Appendix K: CC BY 4.0 dataset, Apache 2.0 code
dataset은 CC BY 4.0, crawler와 baseline code는 Apache License 2.0으로 제공된다. intended use는 English e-commerce product/review 기반 attribute-aware multimodal entity linking이다. trained model은 user post를 product 또는 general entity에 연결하는 용도, 즉 user interest detection에도 사용할 수 있다고 적는다. image를 쓰지 않는 text-only entity linking에도 dataset을 활용할 수 있다.
KB와 review benchmark를 재사용 가능한 파일 구조로 제공한다
product_images/와 bestbuy_products.json. JSON에는 product_category, product_name, overview_section.description, image_path, image_url, Spec(attribute key/value), id, url이 들어간다.review_images/, cleaned_review_images/, bestbuy_reviews.json. JSON에는 header, body, mention, review image path/url, predicted_attribute, gold_attribute, review_id, fused_candidate_list(top-10 IDs), gold_entity_info(id/name/category)가 포함된다.멀티모달 linking을 representation matching에서 evidence comparison으로 바꾼다
이 논문의 기여는 “image와 text에 attribute를 하나 더 concatenation했다”는 식으로 축소하면 잘 보이지 않는다. candidate retrieval은 여전히 dense similarity에 의존하지만, final disambiguation은 review가 어떤 attribute를 말하고 있는지, candidate의 attribute가 그 말과 entail/contradict하는지를 묻는다.
이 구조는 product EL에 국한되지 않을 수 있다. 사람·장소·문서·바이오 entity처럼 서로 닮은 후보가 많은 domain에서도 global embedding이 candidate set을 만들고, fine-grained structured evidence가 마지막 decision을 책임지는 architecture를 생각할 수 있다. 다만 이 확장은 원문이 직접 평가한 결과가 아니라 연구적 inference다.
무엇이 입증됐고 무엇이 아직 남아 있는가
원문이 연결한 entity linking·multimodal learning·attribute extraction 연구
아래 목록은 첨부 논문의 bibliography를 주제 흐름이 보이도록 compact하게 정리한 것이다. 연도와 venue 표기는 원문을 따른다.