Building Neo4j-Powered Applications with LLMs · Part III — Neo4j·Spring AI·LangChain4j로 짓는 지능형 추천 시스템

CALL gds.louvain.write('communityGraph', {...})

지능형 추천 시스템
만들다

벡터 유사도는 한 사람의 이웃을 다섯 명까지 보여줄 뿐이다. KNN이 임베딩 사이에 SIMILAR 관계를 긋고 Louvain이 그 그물을 커뮤니티로 접으면, 1,242명이 한 취향으로 묶인다. 이 장은 그 커뮤니티의 힘으로 협업 필터링과 콘텐츠 기반 접근을 결합해 추천을 완성한다.

Ravindranatha Anthapu · Siddhant Agarwal Packt Publishing, 2025 Chapter 10 · pp. 205–227

데이터를 그래프에 적재하고, LangChain4j와 Spring AI로 그래프를 증강하며 추천을 생성하는 방법까지 보았으니, 이제 Graph Data Science(GDS) 알고리즘과 기계학습을 지렛대 삼아 추천을 개선하는 더 먼 길을 간다. Neo4j가 제공하는 GDS 알고리즘을 검토해 앞 장에서 만든 추천 시스템 너머로 나아가고, GDS 알고리즘으로 협업 필터링(collaborative filtering)콘텐츠 기반(content-based) 접근을 지어 추천을 제공하는 법을 배운다. 알고리즘 실행 후의 결과도 들여다보며, 우리의 접근이 잘 작동하는지, 더 나은 추천 시스템을 짓는 옳은 길 위에 있는지 검토한다. 이 알고리즘들이 앞 장에서 구현한 접근보다 왜 더 나은지도 이해하려 한다.

기술 요건 — 환경 설정

LangChain4j·Spring AI 프로젝트 작업에는 Java IDE 환경을 쓴다. 설치되어 있어야 하고 다루는 법도 알아야 한다. 우리는 앞 장에서 만든 그래프 데이터베이스에서 출발하며, 코드는 Neo4j 5.21.2 버전 데이터베이스에서 테스트되었다. 환경 설정에는 다음 플러그인이 설치된 Neo4j Desktop이 필요하다 — APOC 플러그인 5.21.2, Graph Data Science 라이브러리 2.9.0. 설치 방법은 이렇다(그림 10.1). Neo4j Desktop에서 DBMS를 선택하면 오른쪽에 상세 정보가 나타난다. Plugins 탭을 클릭해 플러그인을 선택하고, 펼쳐지면 Install and Restart 버튼을 클릭한다.

데이터베이스 준비

권장 — 덤프에서 시작하기

시작하기 전에 커뮤니티를 만들어 두어야 한다. 유사도·커뮤니티 탐지 알고리즘이 완료되기까지는 시간이 꽤 걸린다. 따라서 데이터베이스 덤프(hmreco_post_augment_with_summary_communities.dump)를 내려받아 Neo4j Desktop에 데이터베이스를 만드는 것을 권장한다. 덤프의 적재 방법은 공식 안내를 따른다. 이 덤프에는 모든 SUMMER_2019_SIMILAR 관계가 만들어져 있고 커뮤니티도 식별되어 있다.

(:Section {n: 1})-[:ENHANCES]->(:Graph)

GDS 알고리즘으로 추천을 개선하다

이 절에서는 더 나은 추천 시스템을 짓기 위해, 그래프에 대한 더 많은 통찰을 얻도록 그래프를 한층 더 강화하는 방법을 본다. 앞 장에서 만든 그래프 데이터베이스에서 출발한다(참고용으로 덤프를 내려받을 수 있다). Neo4j GDS 알고리즘이 그래프의 강화를 돕는다. 과정은 두 단계다.

  1. 고객 사이의 유사도를 계산한다. 우리가 만든 임베딩에 기초해 유사도를 계산하고, 그 고객들 사이에 similar 관계를 만든다. 이 목적에는 K-최근접 이웃(K-Nearest Neighbors, KNN) 알고리즘을 지렛대로 쓴다.
  2. 커뮤니티 탐지 알고리즘을 실행한다. similar 관계에 기초해 고객들을 그룹으로 묶는다. 이 목적에는 Louvain 커뮤니티 탐지 알고리즘을 지렛대로 쓴다.

KNN 알고리즘으로 유사도를 계산하다

K-최근접 이웃(KNN) 알고리즘은 노드 쌍을 모아, 노드와 그 이웃들 사이의 거리 값을 계산하고, 노드와 상위 K개의 이웃 사이에 관계를 만든다. 거리는 노드 프로퍼티에 기초해 계산된다. 이 알고리즘에는 동종 그래프(homogeneous graph)를 제공해야 한다. 모든 노드와 관계가 같을 때 동종 그래프라 부른다. KNN에 제공하는 노드 쌍에는 노드 레이블이나 관계 타입이 필요 없다. KNN 알고리즘은 연결된 노드 쌍과, 그 사이 관계의 문맥으로 쓸 수 있는 선택적 프로퍼티만 있으면 된다. 자세한 내용은 KNN 문서에서 읽을 수 있다.

이 알고리즘을 쓰려면 두 단계를 따른다. 첫째, 알고리즘을 적용할 관심 그래프를 투영(project)한다. 둘째, 적절한 구성으로 알고리즘을 호출한다. 알고리즘에는 세 가지 모드가 있다.

Stream

인메모리 그래프에 알고리즘을 적용하고 결과를 스트리밍한다. 결과를 검사해 우리가 원하는 것인지 확인할 때 쓴다.

Mutate

인메모리 그래프에 알고리즘을 적용하고 데이터를 인메모리 그래프에 되쓴다. 실제 데이터베이스는 변하지 않는다. 인메모리 그래프를 갱신해 두고 나중에 다른 목적으로 처리하고 싶을 때 쓴다.

Write

인메모리 그래프에 알고리즘을 적용하고 관계들을 실제 데이터베이스에 되쓴다. 과정을 확신하고 결과를 즉시 그래프에 기록하고 싶을 때 쓴다.

그래프 투영부터 시작한다. 임베딩을 SUMMER_2019 관계에 기록해 두었으므로 그것을 처리에 쓴다. 다음 Cypher가 알고리즘을 호출할 수 있도록 그래프를 메모리에 투영한다.

projection — myGraph 인메모리 투영cypher
MATCH (c:Customer)-[sr:SUMMER_2019]->()
WHERE sr.embedding is not null
RETURN gds.graph.project(
  'myGraph',
  c,
  null,
  {
    sourceNodeProperties: sr { .embedding },
    targetNodeProperties: {}
  }
)

보통은 노드 위의 프로퍼티로 투영을 만든다. 그런데 우리는 임베딩 값을 관계에 기록했다. 이 임베딩이 2019년 여름 시즌 구매를 표현하기 위해 만들어졌기 때문이다. 그 임베딩을 Customer 노드에 기록했다면, 다양한 시나리오의 고객 구매 행동을 이해하고 싶을 때 Customer 노드 하나에만 그것들을 기록하기 위해 영리한 이름 짓기가 필요했을 것이다. 임베딩을 관계에 기록함으로써 우리는 임베딩의 문맥을 그래프 안에 보존하고 있다. 위 Cypher에서 보듯 관계에서 임베딩을 꺼내 투영의 소스 노드 프로퍼티로 더하고 있다. 이제 알고리즘을 호출해 similar 관계를 그래프에 되쓴다.

knn — write 모드 호출cypher
CALL gds.knn.write('myGraph', {
    writeRelationshipType: 'SUMMER_2019_SIMILAR',
    writeProperty: 'score',
    topK: 5,
    nodeProperties: ['embedding'],
    similarityCutoff: 0.9
})
YIELD nodesCompared, relationshipsWritten

Cypher에서 보듯 알고리즘의 write 모드를 호출하고 있다. 알고리즘은 임베딩에 대한 코사인 유사도로 Customer 사이의 유사도를 계산하고, 컷오프 점수 0.9를 적용해, 유사도 점수 순으로 상위 5명의 이웃을 골라, 그 고객들 사이에 SUMMER_2019_SIMILAR라는 이름의 관계를 기록한다.

Note — 코사인 유사도

코사인 유사도는 두 벡터 사이의 각도를 계산한다. 벡터들이 서로 멀수록 유사도 값은 0에 가까워지고, 서로 가까울수록 1에 가까워진다. 더 읽고 싶다면 Wikipedia 항목을 참고한다.

유사 점수는 0과 1 사이일 수 있다. 두 개체 사이 점수가 0에 가까우면 서로 유사하지 않고, 1에 가까우면 더 유사하다. 우리가 유사도 컷오프로 0.9를 쓰는 이유는, 생성된 요약 텍스트에 기초해 임베딩을 만들고 있어 고객들의 유사도 점수가 높게 나올 수 있기 때문이다. 몇몇 키워드가 비슷하다는 이유만으로 고객 사이에 similar 관계가 생기는 것을 원하지 않는다. 이 가정은 이후 단계에서 검증한다. 더 가까운 추천을 얻기 위해 상위 다섯(k=5)의 유사 고객 행동으로 스스로를 제한한다.

Note — 비결정적 알고리즘

KNN 알고리즘은 기본적으로 비결정적(non-deterministic) 알고리즘이다. 실행할 때마다 다른 결과가 나올 수 있다는 뜻이다. 공식 문서에서 더 배울 수 있다. 결정적 결과를 원한다면 concurrency 파라미터를 1로 설정하고 randomSeed 파라미터를 명시적으로 설정해야 한다.

알고리즘을 호출했으면 그래프 투영을 제거해야 한다. 그렇지 않으면 데이터베이스 서버의 메모리를 계속 쓴다.

cleanup — 투영 제거cypher
CALL gds.graph.drop('myGraph')

이 Cypher가 그래프를 삭제해 투영이 쓰던 메모리를 정리한다. 다음은 SUMMER_2019_SIMILAR 관계에 기초한 커뮤니티 탐지다. 가장 대중적인 커뮤니티 탐지 알고리즘인 Louvain을 지렛대로 쓴다.

Louvain 알고리즘으로 커뮤니티를 탐지하다

Louvain 알고리즘은 개체 사이의 유사도 점수에 의지해 그들을 커뮤니티로 묶는다. 크고 그물처럼 얽힌 데이터를 받아, 이웃들과 그 관계들을 들여다봄으로써 더 작고 촘촘하게 짜인 커뮤니티들로 묶는 것이다. 이 계층적 군집화 알고리즘은 커뮤니티들을 재귀적으로 단일 노드로 병합하고, 응축된 그래프 위에서 모듈성(modularity) 군집화를 실행한다. 각 커뮤니티의 모듈성 점수 — 커뮤니티 안의 노드들이 무작위 네트워크에서 연결되었을 경우와 견주어 얼마나 더 촘촘히 연결되어 있는지의 평가 — 를 최대화한다. 우리는 더 폭넓은 추천을 제공하기 위해, 더 자동화된 방식으로 고객들을 더 촘촘한 그룹으로 묶고자 한다. 자세한 내용은 Louvain 문서에서 읽는다.

접근은 KNN 알고리즘 호출과 매우 비슷하다. 관심 그래프를 투영하고, 적절한 구성으로 알고리즘을 호출한다. 모드도 같은 셋 — Stream, Mutate, Write — 이다. 그래프 투영부터 시작하자. SUMMER_2019_SIMILAR 관계와 그 관계에 저장된 score 값으로 커뮤니티 탐지를 수행한다.

projection — communityGraph (무방향)cypher
MATCH (source:Customer)-[r:SUMMER_2019_SIMILAR]->(target)
RETURN gds.graph.project(
  'communityGraph',
  source,
  target,
  {
    relationshipProperties: r { .score }
  },
  { undirectedRelationshipTypes: ['*'] }
)

위 Cypher는 communityGraph라는 이름의 인메모리 투영을 만든다. 소스 노드, 타깃 노드, 그리고 SUMMER_2019_SIMILAR 관계의 score를 받아 투영을 짓는다. 투영이 지어지면 이 Cypher로 커뮤니티 탐지를 수행할 수 있다.

louvain — write 모드 호출cypher
CALL gds.louvain.write('communityGraph', { writeProperty: 'summer_2019_community' })
YIELD communityCount, modularity, modularities

이 Cypher가 커뮤니티 탐지를 수행하고, 커뮤니티 ID를 Customer 노드의 summer_2019_community라는 프로퍼티로 되쓴다. 커뮤니티 탐지가 끝나면 그래프 투영을 삭제해야 한다 — CALL gds.graph.drop('communityGraph'). 이 Cypher가 그래프를 삭제해 투영이 쓰던 메모리를 정리한다. 몇 개의 커뮤니티가 만들어졌는지는 이 Cypher로 검사할 수 있다.

inspect — 커뮤니티별 고객 수cypher
MATCH (c:Customer) WHERE c.summer_2019_community IS NOT NULL
RETURN c.summer_2019_community, COUNT(c) as count
ORDER BY count DESC

이 Cypher는 모든 커뮤니티를, 얼마나 많은 고객이 속했는지의 순서로 준다. 응답은 그림 10.2와 같은 모습이다.

c.summer_2019_communitycount
1729
2110
6010
1375
133
1242
2778
966
649
901
713
696
Started streaming 17 records after 8 ms and completed after 299 ms.

그림 10.2 — 고객 수와 함께 본 커뮤니티들

Note — 비결정적 알고리즘

Louvain 커뮤니티 탐지 알고리즘도 기본적으로 비결정적 알고리즘이다. 실행마다 다른 결과가 나올 수 있다. 공식 문서에서 더 배울 수 있다.

커뮤니티를 지었으니, 생성된 커뮤니티들을 들여다볼 차례다. 다음 절에서 이 커뮤니티 몇 개를 검사해, 구매 행동에 기초해 고객들이 묶였는지 관찰한다.

(:Section {n: 2})-[:OBSERVES]->(:Community)

커뮤니티의 힘을 이해하다

앞서 9장에서 우리는 벡터 유사도로 유사 고객을 찾는 법과 고객에게 추천을 제공하는 법을 보았다. 그림 9.10과 그림 9.11을 다시 떠올려 보자. 그림 9.10은 특정 고객과 유사한 고객들의 구매 이력을 보여주었고, 그림 9.11은 유사 고객들의 구매에 기초한 고객 추천을 보여주었다. 그 구매 이력과 고객 추천은 벡터 유사도의 활용을 이해하기 위해 미세조정 절에서 다뤘던 Cypher 질의들의 결과였다. 이 절에서는 커뮤니티를 더 깊이 들여다보며, 유사 고객을 찾는 데 커뮤니티가 단순한 벡터 유사도의 활용보다 왜 더 나을 수 있는지 본다.

Note

아래의 Cypher들은 기술 요건 절에서 공유한 데이터베이스와 관련된 것이다.

앞 절 Louvain 알고리즘으로 커뮤니티를 탐지하다에서 실행한 Cypher에서, 고객이 제법 많은 커뮤니티 하나를 고르자. 약 1,242명의 고객이 속한 커뮤니티 133을 들여다본다. 다음 Cypher는 처음 다섯 고객의 구매 요약을 품목 상세 없이 표시한다.

query — 커뮤니티 133의 구매 요약 5건cypher
MATCH (c:Customer)-[r:SUMMER_2019]->()
WHERE c.summer_2019_community=133
WITH split(r.summary, '\n') AS s
WITH CASE WHEN s[2] <> '' THEN s[2] ELSE s[3] END AS d
return d LIMIT 5
주의 — 커뮤니티 ID는 실행마다 다르다

커뮤니티 탐지를 실행할 때 생성되는 커뮤니티 ID는 실행마다 다를 수 있다. 자신의 Cypher 스크립트로 고객 커뮤니티를 직접 만들었다면, 커뮤니티들을 들여다보고 그 ID들을 써서 데이터를 검증해야 한다.

위 Cypher 스크립트를 실행하면 출력은 이렇다.

community 133 · customer 1The customer demonstrates a preference for stylish and modern pieces, particularly favoring dresses and lingerie that offer both comfort and elegance. The consistent choice of midi and short dresses paired with a variety of non-wired bras suggests a desire for chic yet relaxed fashion options. Additionally, the inclusion of tailored blouses and fashionable outerwear indicates an appreciation for versatile styles suitable for various occasions.
community 133 · customer 2The customer exhibits a preference for comfortable yet stylish clothing, favoring light and soft colors such as light pink and light blue. Their purchases reflect a blend of casual and lingerie items, indicating a focus on both everyday wear and intimate apparel. The selection features a mix of high-waisted denim and lace detailing, suggesting an appreciation for modern, flattering silhouettes.
community 133 · customer 3The customer demonstrates a strong preference for versatile and stylish pieces, favoring bold colors like pink and orange while incorporating comfortable fabrics such as jersey and cotton. Their purchases include a mix of casual wear, activewear, and lingerie, suggesting a balanced lifestyle that values both comfort and aesthetics. The frequent selection of shorts and dresses indicates a preference for easy-to-wear, fashionable items suitable for various occasions.
community 133 · customer 4The customer demonstrates a preference for comfortable and stylish lingerie, favoring soft materials with unique design details such as lace trims and laser-cut edges. Additionally, their choice of everyday wear leans towards light and airy fabrics, showcasing a blend of casual and chic styles suitable for various occasions. The color palette reflects a soft and neutral aesthetic, with light pinks, beiges, and whites dominating their selections.
community 133 · customer 5The customer exhibits a preference for comfortable and functional clothing, particularly in the realm of casual and lingerie wear. There is a clear inclination towards basic styles in neutral colors such as black and white, complemented by playful accents in perceived colors like orange and pink. The focus on versatile pieces suggests a desire for practicality combined with style.

요약 서술들에서, 이 커뮤니티의 고객들이 캐주얼과 란제리 의류를 함께 사기를 선호한다는 사실을 볼 수 있다.

이 고객들이 구매한 품목들을 보자.

query — 커뮤니티 133 고객 10명의 첫 품목 3건씩cypher
MATCH (c:Customer)
WHERE c.summer_2019_community=133
WITH c LIMIT 10
MATCH (c)-[:SUMMER_2019]->(start)
MATCH (c)-[:FALL_2019]->(end)
WITH c, start, end
CALL {
    WITH start, end
    MATCH p=(start)-[:NEXT*]->(end)
    WITH nodes(p) as nodes
    UNWIND nodes as n
    MATCH (n)-[:HAS_ARTICLE]->(a)
    WITH a LIMIT 3
    RETURN collect(a.desc) as articles
}
WITH c, articles
RETURN articles

이 Cypher는 다음의 출력을 준다.

커뮤니티 133 — 고객별 첫 품목 3건발췌
customer 1["Calf-length dress in a crinkled weave with a V-neck, wrapover front with ties at the waist and short sleeves with a slit and ties. Unlined.", "Calf-length dress in a crinkled weave with a V-neck, wrapover front with ties at the waist and short sleeves with a slit and ties. Unlined.", "Soft, non-wired bras in cotton jersey with moulded, padded triangular cups for a larger bust and fuller cleavage. Adjustable shoulder straps that cross at the back and lace at the hem. No fasteners."]
customer 2["High-waisted jeans in washed superstretch denim with hard-worn details, a zip fly and button, back pockets and skinny legs.", "Lace push-up bra with underwired, moulded, padded cups for a larger bust and fuller cleavage. Adjustable shoulder straps and a hook-and-eye fastening at the back.", "Lace push-up bra with underwired, moulded, padded cups for a larger bust and fuller cleavage. Adjustable shoulder straps and a hook-and-eye fastening at the back."]
customer 3["Vest top in cotton jersey with a print motif.", "Soft, non-wired bras in microfibre with padded cups that shape the bust and provide good support. Adjustable shoulder straps and a hook-and-eye fastening at the back.", "Chino shorts in washed cotton poplin with a zip fly, side pockets, welt back pockets with a button and legs with creases."]
customer 4["Microfibre Brazilian briefs with laser-cut edges, a low waist, lined gusset, wide sides and half-string back.", "Hipster briefs in microfibre with lace trims, a low waist, lined gusset and cutaway coverage at the back.", "Blouse in an airy weave with a V-neck, covered buttons down the front, short dolman sleeves and a tie detail at the hem."]
customer 5["Round-necked T-shirt in soft cotton jersey.", "Round-necked T-shirt in soft cotton jersey.", "Thong briefs in cotton jersey and lace with a low waist, lined gusset, wide sides and string back."]

많은 데이터를 들여다보지 않도록 첫 세 품목으로 스스로를 제한했다. 구매된 품목들에서, 요약이 고객의 구매 행동을 잘 요약하고 있음을 볼 수 있다. 이제 고객 연령 그룹과 커뮤니티의 상관이 어떻게 존재하는지 보자. 다음 Cypher는 각 커뮤니티에서 가장 빈번하게 나타나는 연령 그룹을 준다.

query — 커뮤니티별 최빈 연령 그룹과 비율cypher
MATCH (c:Customer) where c.summer_2019_community is not null
WITH  c.summer_2019_community as community, toInteger(c.age) as age,  c
WITH community,
    CASE WHEN age < 10 THEN "Young"
         WHEN 10 < age < 20 THEN "Teen"
         WHEN 20 < age < 30 THEN "Youth"
         WHEN 30 < age < 50 THEN "Adult"
         ELSE "Old"
    END as ageGroup,
    c
WITH community, ageGroup, count(*) as count
CALL {
    WITH community
    MATCH (c:Customer) where c.summer_2019_community=community
    RETURN count(*) as totalCommunity
}
WITH community, ageGroup, count, totalCommunity
WITH community, ageGroup, round(count*100.0/totalCommunity, 2) as ratio
WITH community, collect({ageGroup: ageGroup, ratio:ratio}) as data
CALL {
    WITH community, data
    UNWIND data as d
    WITH community, d
    ORDER BY d.ratio DESC
    RETURN community as c, d.ageGroup as a, d.ratio as r
    LIMIT 1
}
RETURN c as community, a as ageGroup, r as ratio
ORDER BY r DESC

결과는 그림 10.3과 같은 모습이다.

Community · Age GroupRatio (%)
1899 · "Youth"
60.87
5823 · "Adult"
56.68
770 · "Youth"
47.9
1729 · "Youth"
46.92
4602 · "Youth"
45.71
133 · "Youth"
44.61
921 · "Youth"
44.17
3444 · "Youth"
41.73
649 · "Youth"
41.62
1881 · "Youth"
41.26
1696 · "Old"
41.06
2381 · "Old"
40.94
6010 · "Old"
37.67
713 · "Youth"
37.64
760 · "Youth"
36.09
1875 · "Old"
35.09
2778 · "Youth"
34.47

그림 10.3 — 각 커뮤니티에서 가장 빈번하게 나타나는 연령 그룹과 그 비율

대부분의 커뮤니티가 20세에서 30세 사이의 연령 그룹 Youth에 지배되고 있음을 볼 수 있다. Youth 연령 그룹이 지배적이지 않은 커뮤니티 하나를 보자. 커뮤니티 5823이다. 다음 Cypher가 커뮤니티 5823의 처음 다섯 고객 구매 요약을 준다.

query — 커뮤니티 5823의 구매 요약 5건cypher
MATCH (c:Customer)-[r:SUMMER_2019]->()
WHERE c.summer_2019_community=5823
WITH c, split(r.summary, '\n') AS s
WITH c, CASE WHEN s[2] <> '' THEN s[2] ELSE s[3] END AS d
RETURN d LIMIT 5

특정 고객의 벡터 임베딩으로 유사도 검색을 수행하면, 결과는 주로 그 고객의 벡터와 가까운 벡터 표현을 지닌 다른 개별 고객들이다. 어느 정도의 이질성을 관찰하게 될 수도 있다. 중요한 단서 하나 — 벡터 거리에 기초해 타깃 고객과 유사한 고객을 찾는 데만 기댄다면, 잠재적으로 적절한 추천들을 놓칠 수 있다는 것이다. 다음 결과들을 보자.

community 5823 · customer 1The customer demonstrates a strong preference for versatile and stylish pieces, with a notable inclination towards swimwear and casual skirts, reflecting an active and chic lifestyle. The selection features a mix of practical and trendy items, highlighting an appreciation for both comfort and aesthetics. The color palette leans towards soft tones and earthy shades, suggesting a preference for understated elegance.
community 5823 · customer 2The customer exhibits a preference for stylish yet comfortable footwear and swimwear, favoring pieces that blend functionality with trendy elements. The consistent use of white and orange in swimwear suggests a bold and lively aesthetic, while the choice of soft organic cotton for kids' basics indicates an appreciation for quality and sustainability. Overall, there is a clear inclination towards versatile and fashionable pieces suitable for both leisure and casual settings.
community 5823 · customer 3The customer's fashion preferences indicate a strong inclination towards relaxed and comfortable styles, particularly in children's denim wear. The consistent choice of blue tones across multiple purchases suggests a preference for classic and versatile colors. Additionally, the inclusion of a dress with a structured yet casual design highlights an appreciation for both practicality and style in their wardrobe choices.
community 5823 · customer 4The customer exhibits a preference for versatile and comfortable clothing, with a notable inclination towards knitwear and soft fabrics. Their choices reflect a balance of casual and practical styles suitable for everyday wear, particularly in hues of black, dark orange, and grey, complemented by accents of pink. The selected items also indicate a focus on functionality, especially with the inclusion of nursing bras.
community 5823 · customer 5The customer demonstrates a preference for versatile and stylish pieces that blend comfort with contemporary design. They appreciate a mix of youthful and sophisticated styles, as seen in their selection of both kids' dresses and women's wear. The choice of colors suggests a fondness for neutral tones with pops of color, reflecting both playful and elegant aesthetics.

이 요약들은 이 커뮤니티가 아이가 있는 사람들 쪽으로 기울어 있음을 보여준다. 커뮤니티들과 그 안의 고객 요약 몇 건을 관찰하고 나면, 단순히 벡터 유사도만 따라가는 것보다 구매 행동을 더 잘 이해할 수 있게 된다.

다음 걸음은 더 나은 추천을 위해 협업 필터링과 콘텐츠 기반 접근을 결합하는 것이다.

(:Section {n: 3})-[:COMBINES]->(:Approach)

협업 필터링과 콘텐츠 기반 접근의 결합

협업 필터링(collaborative filtering)은 구매에 기초한 고객 유사도 — 우리는 그것으로 고객 커뮤니티를 지었다 — 나 특성에 기초한 품목 유사도를 이용해 추천을 제공하는 일이다. 콘텐츠 기반 필터링(content-based filtering)은 품목의 속성이나 특성에 기초해 추천을 제공하게 해 준다. 이 두 접근을 결합해 더 나은 추천을 제공하는 방법을 본다. 시도할 시나리오는 둘이다 — 시나리오 1: 다른 커뮤니티에 속한 품목의 필터링, 시나리오 2: 특성으로, 그리고 다른 커뮤니티 소속으로 품목을 필터링. 시나리오 1부터 논하자.

시나리오 1 — 다른 커뮤니티에 속한 품목을 걸러내다

이 시나리오에서는 먼저 같은 커뮤니티의 모든 고객이 구매한 품목 전부를 찾는다. 다음으로 다른 커뮤니티에 속한 고객들이 구매한 품목을 찾는다. 다른 커뮤니티에 속한 품목들은 제거된다. 이어서 이 품목들(다른 커뮤니티 소속)을 걸러내는(제거하는) 작업이 뒤따른다. 이 시나리오를 위해 커뮤니티 1696의 고객 000ae8a03447710b4de81d85698dfc0559258c93136650efc2429fcca80d699a를 고른다. 이 고객의 구매 요약부터 보자.

query — 고객 000a..699a의 구매 요약cypher
MATCH (c:Customer)-[r:SUMMER_2019]->()
WHERE c.id='000ae8a03447710b4de81d85698dfc0559258c93136650efc2429fcca80d699a'
WITH c, split(r.summary, '\n') AS s
WITH c, CASE WHEN s[2] <> '' THEN s[2] ELSE s[3] END AS d
RETURN d

이 Cypher는 다음의 출력을 준다.

community 1696 · customer 000a..699a"The customer's fashion preferences indicate a strong inclination towards comfortable yet stylish pieces, often favoring soft fabrics and relaxed fits. The choices reflect a taste for versatile items that can be dressed up or down, particularly in a palette that leans towards darker shades with hints of pink. Overall, there is a notable emphasis on casual wear that combines simplicity with a touch of elegance."

이제 이 고객을 위한 추천 — 이전에 구매하지 않은 품목들 — 을 다음 Cypher로 얻는다. 질의는 단계별로 짚는다.

scenario 1 · 단계 2~3 — 고객 본인의 구매 품목cypher
MATCH (c:Customer {id:'000ae8a03447710b4de81d85698dfc0559258c93136650efc2429fcca80d699a'})
WITH c
CALL {
    -- 이 고객이 구매한 품목들을 얻는다
    WITH c
    MATCH (c)-[:SUMMER_2019]->(start)
    MATCH (c)-[:FALL_2019]->(end)
    MATCH p=(start)-[:NEXT*]->(end)
    WITH nodes(p) as txns
    UNWIND txns as txn
    MATCH (txn)-[:HAS_ARTICLE]->(a)
    WITH DISTINCT a
    RETURN collect(a) as articles
}
WITH c, articles, c.summer_2019_community as community
scenario 1 · 단계 4 — 같은 커뮤니티 고객들의 품목cypher
CALL {
    WITH community
    MATCH (inc:Customer) WHERE inc.summer_2019_community = community
    MATCH (inc)-[:SUMMER_2019]->(start)
    MATCH (inc)-[:FALL_2019]->(end)
    MATCH p=(start)-[:NEXT*]->(end)
    WITH nodes(p) as txns
    UNWIND txns as txn
    MATCH (txn)-[:HAS_ARTICLE]->(a)
    WITH DISTINCT a
    RETURN collect(a) as inCommunityArticles
}
WITH c, articles, community, inCommunityArticles
scenario 1 · 단계 5 — 다른 커뮤니티 고객들의 품목cypher
CALL {
    WITH community
    MATCH (outc:Customer) WHERE outc.summer_2019_community is not null
        and outc.summer_2019_community <> community
    MATCH (outc)-[:SUMMER_2019]->(start)
    MATCH (outc)-[:FALL_2019]->(end)
    MATCH p=(start)-[:NEXT*]->(end)
    WITH nodes(p) as txns
    UNWIND txns as txn
    MATCH (txn)-[:HAS_ARTICLE]->(a)
    WITH DISTINCT a
    RETURN collect(a) as outCommunityArticles
}
WITH c, articles, community, inCommunityArticles, outCommunityArticles
scenario 1 · 단계 6~8 — 차집합과 10건 추천cypher
-- 원 고객이 속한 커뮤니티 밖 고객들이 구매한 품목을 제거한다
WITH c, articles, apoc.coll.subtract(inCommunityArticles, outCommunityArticles) as onlyInCommunity
-- onlyInCommunity에서 원 고객이 구매한 품목을 제거한다
WITH c, apoc.coll.subtract(onlyInCommunity, articles) as notPurchasedButInCommunity
-- 남은 목록에서 10개의 추천 품목을 제공한다. 단순화와 시연을 위해 10개로
-- 제한한다. 모든 품목을 보고 다른 측면으로 그룹화해 다른 추천을 제공할 수도 있다
UNWIND notPurchasedButInCommunity as article
RETURN article.id as id, article.desc as desc
LIMIT 10

여기서 우리는 고객이 구매한 품목을 먼저 얻는다. 그다음 고객이 속한 커뮤니티의 모든 품목을 가져오고, 그 뒤에 다른 커뮤니티 소속 고객들이 구매한 모든 품목을 얻는다. 커뮤니티 안의 고객들만 구매한 품목의 부분집합을 얻고, 그 집합에서 고객 본인이 구매한 품목을 제거해, 남은 품목들을 추천으로 제공한다. 이 질의의 출력은 그림 10.4와 같은 모습이다.

추천 — 고객 000a..699a (community 1696)10 articles
0708679001Slim-fit, ankle-length jeans in washed, superstretch denim with a high waist, zip fly, fake front pockets and real back pockets.
0834749001Oversized jumper in a soft rib knit containing some wool with a polo neck, low dropped shoulders, long, voluminous sleeves, and wide ribbing at the cuffs and hem. The polyester content of the jumper is recycled.
0513701002V-neck T-shirts in organic cotton jersey.
0522374003Jumper in a soft, fine knit with dropped shoulders, long sleeves and gently rounded hem.
0522374001Jumper in a soft, fine knit with dropped shoulders, long sleeves and gently rounded hem.
0687041002Long-sleeved, fitted top in soft, organic cotton jersey with a deep neckline, buttons at the top and a rounded hem.
0724567004Pyjamas with a strappy top and shorts in soft satin with lace details. Top with a V-neck and narrow adjustable shoulder straps. Shorts with narrow elastication at the waist.
0785086001Short satin nightslip with a V-neck, lace trims at the top and hem, and adjustable spaghetti shoulder straps.
0725353002Bell-shaped, knee-length skirt in woven fabric with a high waist and a concealed zip and hook-and-eye fastening in one side. Lined.
0604655007Pyjamas in printed cotton jersey. Short-sleeved top with a round neck. Bottoms with an elasticated waist and wide, gently tapered legs with ribbed hems.

그림 10.4 — 고객이 구매한 품목과 커뮤니티 밖 고객들이 구매한 품목을 걸러낸 추천

이 추천들은 고객의 구매 요약에 실제로 들어맞아 보인다.

이제 시나리오 2를 보자.

시나리오 2 — 특성으로, 그리고 다른 커뮤니티 소속으로 품목을 걸러내다

이 시나리오에서는 질의에 품목 특성을 더하고자 한다. 즉 한 고객에 대해, 먼저 특정 특성을 지닌, 커뮤니티가 구매한 모든 품목을 찾는다. 그런 다음 이 품목들이 속한 커뮤니티들을 찾아, 다른 커뮤니티에 속한 품목을 제거한다. 이 목적을 위해 커뮤니티 5823과 ID가 00281c683a8eb0942e22d88275ad756309895813e0648d4b97c7bc8178502b33인 고객을 고른다. 이 고객의 구매를 보자.

query — 고객 0028..2b33의 구매와 섹션cypher
MATCH (c:Customer) where c.id='00281c683a8eb0942e22d88275ad756309895813e0648d4b97c7bc8178502b33'
WITH c
CALL {
    WITH c
    MATCH (c)-[:SUMMER_2019]->(start)
    MATCH (c)-[:FALL_2019]->(end)
    MATCH p=(start)-[:NEXT*]->(end)
    WITH nodes(p) as txns
    UNWIND txns as txn
    MATCH (txn)-[:HAS_ARTICLE]->(a)-[:HAS_SECTION]->(s)
    WITH DISTINCT a,s
    RETURN collect({section:s.name, article:a.desc}) as articles
}
return articles

위 Cypher에 기초해 이런 출력을 얻는다.

고객 0028..2b33 — 구매 품목과 섹션4 items
Kids Boy"5-pocket jeans in washed stretch denim in a relaxed fit with an adjustable elasticated waist, zip fly and press-stud and tapered legs."
Kids Boy"5-pocket jeans in washed stretch denim with hard-worn details in a relaxed fit with an adjustable elasticated waist, zip fly and press-stud, and tapered legs."
Divided Collection"Short dress in woven fabric with a collar, buttons down the front and a yoke at the back. Narrow, detachable belt at the waist and long sleeves with buttoned cuffs. Unlined."
Kids Boy"Dungarees in washed stretch denim with a three-part chest pocket, adjustable straps with metal fasteners, and front and back pockets. Fake fly, press-studs at the sides, jersey-lined legs and a lining at the hems in a patterned weave."

이 고객이 "Kids Boy" 섹션의 의류를 사고 있으므로, 이 섹션에 속한 추천을 가져오자. 다음 Cypher는 앞의 질의에 이 섹션의 상세를 더해 추천을 준다. ID가 00281c..2b33인 Customer와 이름이 Kids Boy인 Section을 얻는 것부터 시작한다.

scenario 2 · 단계 2~3 — 섹션 필터 + 본인 구매cypher
MATCH (c:Customer {id:'00281c683a8eb0942e22d88275ad756309895813e0648d4b97c7bc8178502b33'})
MATCH (s:Section) WHERE s.name='Kids Boy'
WITH c,s
CALL {
    -- 관심 섹션에 속한, 이 고객이 구매한 품목들을 얻는다
    WITH c,s
    MATCH (c)-[:SUMMER_2019]->(start)
    MATCH (c)-[:FALL_2019]->(end)
    MATCH p=(start)-[:NEXT*]->(end)
    WITH nodes(p) as txns
    UNWIND txns as txn
    MATCH (txn)-[:HAS_ARTICLE]->(a)-[:HAS_SECTION]->(s)
    WITH DISTINCT a
    RETURN collect(a) as articles
}
WITH c, articles,s, c.summer_2019_community as community
scenario 2 · 단계 4 — 같은 커뮤니티 + 같은 섹션의 품목cypher
CALL {
    WITH community, s
    MATCH (inc:Customer) WHERE inc.summer_2019_community = community
    MATCH (inc)-[:SUMMER_2019]->(start)
    MATCH (inc)-[:FALL_2019]->(end)
    MATCH p=(start)-[:NEXT*]->(end)
    WITH nodes(p) as txns, s
    UNWIND txns as txn
    MATCH (txn)-[:HAS_ARTICLE]->(a)-[:HAS_SECTION]->(s)
    WITH DISTINCT a
    RETURN collect(a) as inCommunityArticles
}
WITH c, articles, community, inCommunityArticles, s
scenario 2 · 단계 5 — 다른 커뮤니티 + 같은 섹션의 품목cypher
CALL {
    WITH community, s
    MATCH (outc:Customer) WHERE outc.summer_2019_community is not null
        and outc.summer_2019_community <> community
    MATCH (outc)-[:SUMMER_2019]->(start)
    MATCH (outc)-[:FALL_2019]->(end)
    MATCH p=(start)-[:NEXT*]->(end)
    WITH nodes(p) as txns, s
    UNWIND txns as txn
    MATCH (txn)-[:HAS_ARTICLE]->(a)-[:HAS_SECTION]->(s)
    WITH DISTINCT a
    RETURN collect(a) as outCommunityArticles
}
WITH c, articles, community, inCommunityArticles, outCommunityArticles
scenario 2 · 단계 6~8 — 차집합과 10건 추천cypher
-- 커뮤니티 밖 고객들이 구매한 품목을, 원 고객의 커뮤니티 고객들이 구매한 품목에서 제거한다
WITH c, articles, apoc.coll.subtract(inCommunityArticles, outCommunityArticles) as onlyInCommunity
-- 앞 단계에서 얻은 품목 목록에서 원 고객이 구매한 품목을 제거한다
WITH c, apoc.coll.subtract(onlyInCommunity, articles) as notPurchasedButInCommunity
-- 이 품목들 중 10개를 추천으로 제공한다
UNWIND notPurchasedButInCommunity as article
RETURN article.id as id, article.desc as desc
LIMIT 10

이 질의를 실행하면 그림 10.5에 보인 출력을 보게 된다.

추천 — 고객 0028..2b33 · Kids Boy 섹션 (community 5823)10 articles
05055070035-pocket slim-fit jeans in washed stretch denim with an adjustable elasticated waist and zip fly.
0704150011Long-sleeved top in sweatshirt fabric with a motif on the front and ribbing around the neckline, cuffs and hem.
0701969005Shorts in soft, patterned cotton twill with an elasticated drawstring waist, fake fly and side pockets.
0595548001Shorts in soft, washed denim with an elasticated drawstring waist and a back pocket.
0704150006Long-sleeved top in sweatshirt fabric with a motif on the front and ribbing around the neckline, cuffs and hem.
0626380001Top in soft, patterned cotton jersey with long sleeves, an open chest pocket and slits at the hem. Slightly longer at the back.
0701972005Shorts in woven fabric with an adjustable elasticated waist and decorative drawstring. Zip fly and button, diagonal side pockets and welt back pockets.
0771489001T-shirts in airy cotton jersey with a chest pocket and short slits in the sides. Longer at the back.
0705911001Vest top in cotton jersey with a print motif and a ribbed trim around the neckline and armholes.
0666327011T-shirt in soft cotton jersey with a motif on the front.

그림 10.5 — 고객 본인과 커뮤니티 밖 고객들의 구매 품목을 걸러내, 구매와 품목 속성을 함께 고려한 추천

이 장의 시연들은 이 접근들을 씀으로써, 다른 고객들의 유사한 구매에 기초해 서로 다른 유형의 추천을 제공할 수 있음을 보여주었다.

요약

이 장에서 우리는 기본적인 추천 애플리케이션 너머로 나아가, 그래프 알고리즘을 지렛대 삼아 그래프를 강화하고 더 적절한 추천을 제공하는 방법을 보았다. KNN 유사도 알고리즘과 커뮤니티 탐지로 데이터의 숨은 통찰을 얻는 방법을 탐구했다.

다가오는 장들에서는 이 애플리케이션들을 클라우드에 배포하는 방법과, 배포를 위해 따를 수 있는 모범 사례를 들여다본다.

NEXT → Part 4 · Chapter 11 — Choosing the Right Cloud Platform for GenAI Applications