구조 모델, docking, MD, property prediction, sequence DB, cheminformatics library는 서로 다른 입력과 실행환경을 요구한다. ChemGraph가 다루는 핵심 문제도 setup·execution·validation을 둘러싼 이 수작업 비용이다.
Vecura의 학술적 의미는 “또 하나의 신약 예측 AI”가 아니라, 연구자의 자연어 목표를 근거·모델·도구·실험으로 번역하는 Agentic Life-Science Discovery Operating Layer에 있다.
무엇을 만드는가: Definition에서 Problem Definition으로
Vecura를 개별 모델이 아니라 장기 과학 워크플로를 구성하는 시스템으로 다시 정의한다.
Vecura-like Agentic Life-Science Discovery System
학술적으로 Vecura와 같은 시스템은 연구자의 고수준 과학적 의도(scientific intent)를 해석하고, 목표를 달성하기 위해 heterogeneous biomedical evidence와 specialized scientific models·tools를 동적으로 선택·조합·실행하며, 결과를 근거와 함께 평가해 다음 연구행동을 결정하는 tool-augmented autonomous scientific agent system으로 정의할 수 있다.
중요한 점은 LLM 그 자체가 시스템의 전부가 아니라는 사실이다. LLM reasoning + Scientific RAG/KG + model/tool orchestration + computation + verification + human feedback가 하나의 실행 계층으로 결합된다. Vecura의 현재 공개 페이지도 “실험을 기술하면 Agent가 단계를 계획하고 적절한 도구를 선택해 전체 workflow를 실행한다”고 설명하며, small molecule, peptide, repurposing, lead optimisation 등 서로 다른 discovery challenge를 고정 template 없이 연결하는 방향을 제시한다.
이 정의는 Biomni의 unified biomedical action space, retrieval-augmented planning, code execution과 특히 가깝다. Biomni는 수많은 biomedical tool, database, protocol을 하나의 action space로 구성하고 predefined template 없이 복합 연구 workflow를 동적으로 합성한다.
문제는 Molecule → Score가 아니라 Goal → Workflow → Evidence → Decision이다
신약개발 목표를 \(G\), 사용할 수 있는 데이터와 지식의 집합을 \(D\), 실행 가능한 도구와 모델을 \(T\), 비용·시간·안전성·선택성·합성가능성 같은 제약을 \(C\)라고 두자. Agentic discovery system이 찾아야 하는 것은 단일 예측값이 아니라 과학적 실행계획 \(W\)이다.
여기서 \(E\)는 새로 얻은 evidence, \(H\)는 갱신된 hypothesis, \(Y\)는 후보물질 또는 연구결론이다. 따라서 최적화 대상은 단순 accuracy가 아니다. 어떤 근거를 더 찾아야 하는지, 어떤 계산을 먼저 해야 하는지, 어떤 결과를 믿어야 하는지, 불확실성이 큰 경우 어떤 실험을 추가해야 하는지까지 포함하는 sequential scientific decision-making 문제이다.
2026년 Drug Discovery Today의 AI-agent review도 agentic drug discovery를 LLM이 perception, action, memory, scientific tool과 결합해 데이터 통합·계산·실험·가설 갱신을 반복하는 형태로 정리한다. 이 관점에서 Vecura의 핵심은 예측 정확도 하나가 아니라 연구 실행의 조합과 순서를 최적화하는 데 있다.
핵심은 지능이 아니라 조합, 검증, 추적 가능성이다
| Core concept | 의미 | 대표 연구 흐름 |
|---|---|---|
| Scientific Intent Understanding | 자연어 연구목표를 계산 가능한 목표·제약·평가조건으로 변환 | Co-Scientist, Biomni |
| Agentic Planning | 장기 목표를 여러 scientific task로 분해하고 의존성을 구성 | Co-Scientist, ChemGraph |
| Dynamic Tool Routing | 상황에 따라 structure, docking, ADMET, omics 도구를 선택 | Biomni, ChemGraph |
| Multi-Agent Collaboration | Planner, Retriever, Critic, Scientist 등의 역할 분담 | Virtual Lab, CLADD |
| Scientific RAG | 논문·DB·특허·내부 데이터를 검색해 reasoning을 grounding | CLADD, DeepEvidence |
| Evidence Graph | 주장과 관찰, 출처의 연결을 구조적으로 보존 | DeepEvidence |
| Self-Verification | 생성 claim을 curated database나 외부 도구로 재검증 | GeneAgent |
| Multimodal Foundation Models | sequence·3D structure·molecule·omics 표현을 specialized model로 처리 | Boltz-2, BoltzGen |
| Closed-loop Discovery | hypothesis→experiment→analysis→revision의 반복 | Robin, BioDiscoveryAgent |
| Human-in-the-Loop | 고위험 또는 고비용 결정에서 연구자가 steer·approve | Virtual Lab, Vecura |
| Provenance | 결과가 어떤 source/model/version/parameter에서 나왔는지 추적 | DeepEvidence, GeneAgent |
| Uncertainty-aware Decision | 단일 score가 아니라 모델별 uncertainty를 고려해 다음 행동 결정 | 차세대 핵심 연구문제 |
특히 중요한 원칙은 Agent의 자연어 추론과 과학 모델이 계산한 결과를 구분하는 것이다. 그럴듯한 문장과 실험·계산의 증거는 같은 것이 아니다. GeneAgent가 외부 biomedical database를 통한 self-verification으로 hallucination을 줄이고, DeepEvidence가 evidence graph로 출처 추적과 검증을 강조한 이유도 여기에 있다.
왜 지금 필요한가: Model Abundance가 만든 새로운 병목
개별 foundation model의 성능 경쟁이 orchestration, evidence integration, reproducibility의 문제로 이동한다.
질문이 바뀌었다: “무엇을 예측할까?”에서 “어떻게 끝까지 해결할까?”로
2022–2024년 AI 신약개발의 대표 질문은 구조를 얼마나 정확히 예측하는가, drug–target affinity를 얼마나 정확히 예측하는가에 가까웠다. 그러나 2025년 이후의 중심 질문은 여러 foundation model과 scientific tool을 어떻게 조합해 하나의 연구문제를 끝까지 해결하는가로 이동한다.
2025년 Drug Discovery Today의 foundation-model review는 2022년 이후 drug discovery 관련 foundation model이 빠르게 증가해 200개를 넘어섰다고 정리한다. target discovery, molecular optimization, preclinical research 등 모델이 필요한 대부분의 지점에 전문화된 모델이 들어오면서 오히려 새로운 병목이 생겼다.
Vecura는 이 지점에서 의미가 있다. 현재 공개 페이지에는 구조예측, docking, ADMET, protein design, molecular design, cheminformatics, dynamics, single-cell과 같은 다양한 도구가 하나의 Agent interface 아래 연결되어 있다. 따라서 학술적으로 Vecura는 개별 예측기의 경쟁보다 scientific model orchestration 문제를 전면에 놓는 사례라고 해석할 수 있다.
네 가지 마찰이 과학적 Agent를 요구한다
Affinity뿐 아니라 selectivity, solubility, permeability, toxicity, synthesizability, novelty, patentability를 동시에 고려해야 한다. 의사결정은 본질적으로 다목적 최적화다.
GeneAgent는 1,106개 gene set 평가에서 curated database를 이용한 self-verification이 일반 LLM보다 더 정확한 기능 설명을 만들 수 있음을 보였다. 근거 없는 fluent text는 과학적 증거가 아니다.
Robin은 literature search와 data-analysis agent를 연결해 hypothesis generation, experiment proposal, result interpretation, hypothesis update를 순환시켰다. 실제 과학은 이 반복 구조에 더 가깝다.
결국 Vecura류 플랫폼의 궁극적 목표는 workflow automation에 머무르지 않는다. 가설이 실험 결과에 의해 수정되고, 실패가 다음 탐색 공간을 바꾸는 closed scientific loop로 발전해야 한다.
2025–2026 연구 지형: 서로 다른 퍼즐 조각이 하나의 운영체제로 수렴한다
Co-Scientist, Virtual Lab, Robin, Biomni, GeneAgent, DeepEvidence, CLADD, ChemGraph의 기능을 Vecura 관점에서 재배치한다.
Vecura와 직접 경쟁하는 것은 하나의 논문이 아니라 여러 연구 흐름의 합이다
| System / Study | 핵심 메커니즘 | Vecura 관점의 의미 |
|---|---|---|
| Co-Scientist (Nature 2026) | Generation, Reflection, Ranking, Evolution 등 multi-agent와 tournament-style hypothesis evolution | scientific planning과 hypothesis quality를 test-time compute로 높이는 상위 reasoning 계층 |
| Virtual Lab (Nature 2025) | PI agent와 전문 scientist agent의 hierarchical collaboration | human-steered multi-agent 조직 설계와 실제 nanobody validation |
| Robin (Nature 2026) | literature search + data analysis + hypothesis revision | closed-loop discovery의 실험적 구현 |
| Biomni (2025 preprint) | unified action space, retrieval-augmented planning, code execution | 대규모 biomedical tool orchestration의 가장 가까운 공개 연구 비교군 |
| GeneAgent (Nature Methods 2025) | domain database를 통한 self-verification | claim-level factual grounding과 hallucination 억제 |
| DeepEvidence (Nature Machine Intelligence 2026) | breadth/depth search + incremental evidence graph | 근거 탐색, attribution, validation을 구조화하는 evidence layer |
| CLADD (AAAI 2026) | multi-agent RAG, biomedical KB retrieval, molecule contextualization | drug discovery에 특화된 retrieval collaboration |
| ChemGraph (Communications Chemistry 2026) | agentic computational chemistry workflow, decomposition, execution/validation | specialized scientific tool을 안정적으로 연결하는 execution layer |
| Boltz-2 (2025 preprint) | structure + binding affinity prediction | orchestrator가 호출할 수 있는 고성능 specialized foundation model |
| BoltzGen (2025/2026 preprint) | all-atom generative binder design across protein/peptide modalities | 생성·구조·조건 제약을 결합한 binder-design tool layer |
이들을 하나의 그림으로 보면 역할이 선명해진다. Co-Scientist는 “무엇을 생각할 것인가”를, Biomni와 ChemGraph는 “무엇을 실행할 것인가”를, GeneAgent와 DeepEvidence는 “무엇을 믿을 것인가”를, Robin은 “실험 뒤 무엇을 바꿀 것인가”를 각각 전진시킨다. Vecura의 흥미로운 위치는 이 네 질문을 하나의 산업형 연구 인터페이스에서 묶으려는 데 있다.
어디에서 실패하는가: Challenges와 Research Questions
Agent가 실행을 잘하는 것과 과학을 잘하는 것은 같은 문제가 아니다.
장기 계획, 이질적 근거, 불확실성, 재현성
1. Long-Horizon Scientific Planning
도구가 몇 개뿐이라면 LLM이 실행 순서를 고를 수 있다. 그러나 수백 개 model·database·tool이 있으면 workflow 생성 자체가 combinatorial search가 된다. ChemGraph의 결과도 복잡한 계산화학 작업에서는 task decomposition과 multi-agent 구성이 중요해짐을 보여준다.
2. Heterogeneous Biomedical Evidence
동일한 compound–target 관계도 assay type, species, dose, endpoint, protocol, time에 따라 의미가 달라진다. CLADD가 biochemical data의 heterogeneity, ambiguity, multi-source integration을 핵심 RAG 난제로 다룬 이유다.
3. Grounding and Hallucination
Scientific Agent가 그럴듯한 설명을 만드는 능력과 실제 계산·실험 증거를 연결하는 능력은 다르다. tool call 성공은 scientific validity가 아니다. GeneAgent와 DeepEvidence는 verification과 attribution을 시스템 내부 구조로 끌어들인다.
4. Uncertainty Propagation
Protein structure uncertainty, docking error, affinity prediction error, ADMET calibration error가 연쇄적으로 전달된다. 그러나 많은 agentic workflow는 서로 다른 confidence의 의미를 정규화하지 않은 채 최종 ranking으로 압축한다.
5. Benchmark Deficit
BixBench는 실제 bioinformatics 분석을 요구하는 장기 agent task에서 당시 frontier system의 성능이 충분하지 않음을 보여주었다. 원 논문 기준 open-answer accuracy 최고치가 17%에 머물렀다. 화려한 데모와 robust autonomous science 사이의 간극은 여전히 크다.
6. Reproducibility
동일한 질문에 Agent가 매번 다른 model version, parameter, database snapshot, tool order를 선택한다면 연구 결과의 재현성이 깨진다. 따라서 workflow trace 자체가 연구 산출물이어야 한다.
7. Privacy and IP
실패 assay, proprietary SAR, unpublished screen과 같은 내부 데이터는 가장 가치 있는 근거이면서 가장 민감하다. external evidence와 internal evidence를 함께 쓰려면 접근통제, data lineage, private deployment가 orchestration 설계의 일부가 된다.
다음 세대 Agentic Drug Discovery가 답해야 할 열 가지 질문
| RQ | Research Question |
|---|---|
| RQ1 | 자연어 연구목표로부터 고정 template 없이 최적 scientific workflow를 어떻게 생성할 것인가? |
| RQ2 | 수백 개 tool 중 accuracy·latency·cost·domain validity를 함께 고려해 어떤 tool을 선택할 것인가? |
| RQ3 | PubMed·ChEMBL·PubChem·patent·internal assay를 일관된 reasoning representation으로 어떻게 표현할 것인가? |
| RQ4 | Agent가 생성한 모든 scientific claim을 어떤 evidence와 연결해 자동 검증할 것인가? |
| RQ5 | structure·docking·affinity·ADMET의 이질적 uncertainty를 최종 candidate uncertainty로 어떻게 전파할 것인가? |
| RQ6 | 여러 agent/model이 상충할 때 다수결이 아닌 과학적으로 타당한 consensus를 어떻게 형성할 것인가? |
| RQ7 | 실패 후보와 negative experiment를 memory에 어떻게 축적하여 다음 hypothesis를 개선할 것인가? |
| RQ8 | workflow completion이 아니라 실제 wet-lab hit rate와 scientific novelty로 Agent를 어떻게 평가할 것인가? |
| RQ9 | Human-in-the-loop를 어느 단계에 배치해야 autonomy와 safety를 동시에 최적화할 수 있는가? |
| RQ10 | 성과가 Agent reasoning 때문인지 underlying model 때문인지 causal attribution을 어떻게 수행할 것인가? |
특히 RQ7은 중요하다. 성공한 evidence만 저장하는 시스템은 연구의 절반만 기억한다. negative-result-aware scientific memory와 explicit belief revision은 향후 Agentic AI Co-Scientist를 차별화할 가능성이 높다.
어떻게 설계할 것인가: Methods와 Key Applications
강한 LLM 하나보다 planning, evidence, specialized models, verification, feedback를 명시적으로 분리한다.
차세대 Vecura-like architecture
Co-Scientist는 Generation, Reflection, Ranking, Evolution, Proximity, Meta-review 같은 specialized agent와 tournament process를 사용해 hypothesis를 반복적으로 개선한다. 이것은 planner가 단일 pass로 계획을 내는 대신, 계획 자체를 토론·선택·진화시키는 구조가 가능함을 보여준다.
Virtual Lab은 PI Agent가 여러 전문 agent를 지휘하는 계층형 구조를 사용했다. 연구진은 이 시스템을 SARS-CoV-2 nanobody design에 적용해 92개 설계를 실험적으로 평가했다. 여기서 중요한 점은 multi-agent collaboration이 단순 역할극이 아니라 실제 wet-lab validation과 연결되었다는 사실이다.
Biomni는 biomedical action space를 구축하고 retrieval-augmented planning과 code-based execution을 결합한다. Vecura가 장기적으로 수백 개의 scientific tool을 운영한다면, 가장 가까운 학술 문제는 “tool registry를 얼마나 크게 만들 것인가”보다 “현재 질의와 증거 상태에 맞는 action sequence를 어떻게 합성할 것인가”이다.
CLADD는 molecule context를 biomedical knowledge base에서 동적으로 검색하고 여러 LLM agent가 이를 통합한다. DeepEvidence는 breadth-first와 depth-first 탐색을 결합하고 evidence graph를 점진적으로 구성한다. 둘을 결합하면 retrieval은 단순 top-k 문서 검색이 아니라 과학적 근거 공간을 탐색하는 search policy가 된다.
ChemGraph는 computational chemistry task를 분해하고 실행·검증하는 agentic framework를 제시한다. Boltz-2와 BoltzGen은 이 상위 orchestration이 호출할 수 있는 specialized model의 예다. 따라서 미래의 핵심 경쟁력은 더 큰 LLM 하나가 아니라 reasoning model, evidence engine, heterogeneous scientific tools, explicit verification의 조합에 있다.
Target에서 Hit, Lead, Repurposing, Omics까지
문헌, omics, genetics, clinical evidence를 통합해 target–disease 관계와 근거 강도를 평가한다. 핵심은 단순 entity retrieval이 아니라 evidence context와 contradiction을 함께 보존하는 것이다.
구조 retrieval/prediction → screening/docking → affinity → ADMET → candidate ranking을 연속 workflow로 구성한다. Vecura가 공개적으로 보여주는 대표적 사용 시나리오다.
Small molecule뿐 아니라 peptide, protein, nanobody와 같은 modality를 목표·제약 아래 생성하고 평가한다. BoltzGen은 이러한 universal binder design 흐름을 보여준다.
analog generation과 SAR reasoning을 potency, selectivity, ADMET, CNS penetration, synthesizability, IP constraint와 함께 다목적으로 최적화한다.
기존 약물의 새로운 disease/target context를 literature와 biological evidence로 재구성한다. Robin과 Co-Scientist의 biomedical validation이 이 방향을 뒷받침한다.
gene-set interpretation, perturbation experiment design, single-cell/spatial analysis를 agentic workflow로 묶는다. GeneAgent와 Biomni가 대표적인 학술 근거다.
결과적으로 Vecura 같은 시스템은 좁은 의미의 Drug AI라기보다 Computational Life-Science Research Agent로 확장될 가능성이 크다.
아직 풀리지 않은 것: Open Problems
실행 가능성, 과학적 타당성, 재현성, novelty를 같은 평가축에 올려야 한다.
“Agent가 할 수 있다”와 “과학적으로 옳다” 사이의 간극
첫째, workflow success와 scientific validity를 분리해야 한다. API가 정상적으로 호출되고 docking이 종료되었다고 후보물질이 타당한 것은 아니다. 실행 성공률은 과학적 성공률의 대리변수가 될 수 없다.
둘째, end-to-end benchmark가 부족하다. BixBench는 bioinformatics agent의 어려움을 드러내지만, target identification → molecule design → screening → ADMET → experimental validation 전체를 표준화해 평가하는 benchmark는 아직 충분하지 않다.
셋째, uncertainty가 compositional하지 않다. 서로 다른 모델의 confidence는 의미와 calibration이 다르다. 단순 평균이나 voting은 uncertainty propagation을 해결하지 못한다.
넷째, provenance가 더 세밀해야 한다. 단순 citation을 넘어 claim–source–assay–species–dose–time–protocol–model version–parameter–uncertainty가 함께 보존되어야 한다. 이 지점에서 일반 triple KG보다 qualifier를 가진 hyper-relational representation이 더 자연스러운 선택이 될 수 있다.
다섯째, scientific novelty와 hallucination의 경계가 어렵다. 이미 알려진 사실만 반복하는 Agent는 안전하지만 발견을 만들지 못한다. 반대로 evidence를 넘어서는 새로운 가설은 반드시 falsifiability와 counter-evidence search를 요구한다.
여섯째, negative science memory가 부족하다. inactive assay, toxicity failure, synthesis failure를 장기적으로 기억하지 않으면 Agent는 동일한 실패 공간을 반복 탐색한다. 성공사례 중심의 memory는 과학적 학습을 왜곡한다.
일곱째, Vecura-specific independent evidence가 제한적이다. 현재 공개 자료는 제품 기능과 활용 시나리오를 풍부하게 보여주지만, 전체 Agent architecture, tool-selection algorithm, controlled benchmark, prospective wet-lab performance를 공개적으로 다룬 peer-reviewed system paper는 확인되지 않았다. 따라서 제품 페이지의 속도·효율 메시지는 독립 benchmark와 구분해 읽어야 한다.
다음 단계: Agentic AI에서 Epistemic Agentic AI로
무엇을 실행할지 결정하는 시스템을 넘어 무엇을 알고, 무엇을 모르고, 왜 믿는지 설명하는 과학적 Agent로 간다.
일곱 가지 연구 방향
① Provenance-aware Scientific Agentic RAG
일반 vector RAG는 텍스트 chunk를 반환하지만 과학적 의사결정에는 qualifier가 필요하다. 예를 들어 “Drug A inhibits Protein B”가 아니라 assay, IC50, species, protocol, source, time, confidence까지 함께 저장해야 한다. DeepEvidence의 evidence graph를 확장해 hyper-relational evidence representation으로 만드는 방향이 유망하다.
② Uncertainty-Aware Tool Routing
Agent는 가장 유명한 모델이 아니라 expected information gain, cost, latency, uncertainty, domain validity를 함께 고려해 tool을 선택해야 한다. 초기 수만 개 후보에는 cheap predictor를, 상위 수백 개에는 구조·affinity model을, 최종 소수에는 FEP나 고비용 simulation을 배치하는 방식이다. 이는 일종의 Cost-Based Scientific Query Optimizer로 볼 수 있다.
③ Counter-Evidence Agent
여러 agent가 같은 방향으로 답을 늘어놓고 voting하는 것은 과학적 토론이 아니다. 하나의 Agent가 의도적으로 “이 후보가 왜 실패할 것인가?”를 찾도록 설계하면 confirmation bias를 줄일 수 있다. Candidate Agent와 Counter-Evidence Agent의 경쟁은 falsification-oriented multi-agent architecture로 확장될 수 있다.
④ Negative-Result Memory & Belief Revision
실패한 docking, inactive assay, toxicity, synthetic dead-end를 장기 memory에 기록하고 다음 탐색공간을 바꾸어야 한다. 이때 memory는 대화기록이 아니라 Scientific Belief State가 된다.
⑤ Closed-Loop Active Drug Discovery
궁극적 구조는 Design → Predict → Select → Experiment → Observe → Update이다. Robin과 Co-Scientist가 보여준 실험 연결을 drug discovery의 반복 optimization과 통합하면 Agent는 보고서를 쓰는 도구에서 실제 연구 cycle을 운영하는 시스템으로 이동한다.
⑥ Omni-Modal Scientific Reasoning
차세대 Agent는 text만 읽지 않는다. sequence, 3D structure, molecular graph, microscopy, omics matrix, assay curve, MD trajectory, patent structure, protocol을 함께 다뤄야 한다. 현실적인 설계는 하나의 거대한 multimodal LLM이 모든 것을 직접 처리하기보다 modality-specialized foundation model을 meta-reasoner가 orchestration하는 구조다.
⑦ Self-Driving Laboratory & Digital Twin
2026년 drug-discovery agent review가 제시하는 장기 방향은 autonomous laboratory와 digital twin이다. Computational Agent가 virtual experiment를 수행하고 robotic experiment 결과를 받아 hypothesis와 model preference를 다시 갱신하는 구조다.
따라서 가장 흥미로운 연구공간은 Scientific Agentic RAG + Evidence/Hyper-Relational KG + Uncertainty-Aware Tool Routing + Counter-Evidence Verification + Closed-Loop Experimentation의 결합이다. 이 조합은 Vecura가 보여주는 산업 플랫폼의 방향을 넘어, AI4Science와 AI4DrugDiscovery에서 독립적인 학술 문제로 확장할 수 있다.
References & Sources
Agentic workflow, scientific knowledge base, tool/model integration에 대한 현재 공개 설명. https://vecura.com/en
Drug discovery foundation model이 2022년 이후 200개 이상으로 증가한 연구 지형을 정리한다. ScienceDirect
Agentic drug discovery의 architecture, applications, privacy, benchmark, autonomous-lab 방향을 정리한다. DOI
Gemini 기반 multi-agent hypothesis generation, tournament evolution, test-time compute scaling. Nature
Hierarchical multi-agent scientific collaboration과 92개 nanobody의 experimental evaluation. Nature
Robin: literature search, data analysis, hypothesis update를 연결한 semi-autonomous scientific discovery. Nature
Unified biomedical action space, retrieval-augmented planning, code-based execution. bioRxiv
Domain database를 이용한 claim verification과 hallucination 억제. Nature Methods
Breadth-first/depth-first evidence exploration과 incremental evidence graph. Nature Machine Intelligence
CLADD: biochemical heterogeneity와 multi-source integration을 다루는 collaborative RAG. AAAI
계산화학 workflow를 계획·분해·실행·검증하는 agentic framework. Communications Chemistry
Complex structure와 binding affinity를 함께 다루는 structural biology foundation model. bioRxiv
Protein·peptide binder modality를 아우르는 all-atom generative design과 experimental campaigns. bioRxiv
실제 bioinformatics data analysis에서 long-horizon agent의 한계를 평가한다. arXiv