과학을 위한 언어모델의 다음 문제는 더 어려운 문제를 맞히는 데만 있지 않다. 서로 다른 형식의 증거를 읽고, 도구를 사용하고, 긴 작업을 끊기지 않게 이어가며, 실행 결과를 다시 다음 행동에 반영할 수 있는가가 더 중요해지고 있다.
Shanghai AI Laboratory의 Intern-S2-Preview는 이 문제를 scientific agentic foundation model이라는 이름으로 묶는다. 핵심 모델 Intern-S2-Preview-397B는 rendered scientific document, interleaved image-text, scientific corpus를 이용한 pre-training에서 출발한다. 이후 SFT, scalable multi-task RL, black/white-box agentic RL, on-policy distillation을 하나의 post-training pipeline으로 연결한다.
이 보고서의 진짜 흥미는 397B라는 숫자보다 과학적 지식, multimodal evidence, numerical signal, tool interaction, executable verification, long-horizon credit assignment을 하나의 학습 체계 안에 넣으려는 시도에 있다.
다만 논문 스스로도 이를 완성된 autonomous scientist로 규정하지 않는다. 결론에서 Intern-S2-Preview는 여전히 preview system이며, 더 긴 scientific workflow의 reliability, domain memory와 task environment, verifier, specialized tool integration이 향후 과제로 남는다고 명시한다.
과학 모델에서 과학 에이전트로
정적인 question answering이 아니라 heterogeneous evidence와 tools, environment를 엮는 장기 workflow가 목표다.
과학은 한 번의 답변으로 끝나지 않는다
기존 general-purpose LLM은 넓은 instruction-following과 reasoning 능력을 제공하지만 scientific modality, domain protocol, verifiable tool interaction에 특화되어 있지 않다. 반대로 scientific multimodal model은 microscopy, remote sensing, figure, time series 같은 전문 입력을 더 잘 다루지만 여전히 static QA로 평가되는 경우가 많다.
실제 연구는 다르다. 여러 modality의 evidence를 비교하고, 계획을 수정하고, code와 tool을 실행하고, 실패 결과까지 다음 판단에 반영한다. Intern-S2-Preview는 이 차이를 모델의 핵심 문제로 가져온다.
하나의 모델에 네 층을 쌓는다
Scientific multimodality
text, page image, figure, equation, table, microscopy, remote sensing, time series를 함께 다룬다.
Reasoning + generation
scientific reasoning뿐 아니라 molecule·material generation과 numerical forecasting까지 확장한다.
Agentic interaction
tool, files, terminal, coding environment에서 iterative action–observation loop를 수행한다.
Modular specialization
397B backbone을 고정한 채 external parametric memory로 새 scientific domain을 붙인다.
아키텍처: 기억은 붙이고, 숫자는 숫자로 예측한다
Memory Decoder와 time-series branch는 과학 모델이 마주치는 두 현실—지식의 빠른 변화와 수치 신호의 연속성—에 대응한다.
397B backbone을 다시 쓰지 않고 전문지식을 붙인다
Memory Decoder는 base Intern-S2-Preview-397B의 구성요소가 아니라 별도 extension이다. frozen backbone과 independently trained memory decoder가 같은 context를 병렬로 처리하고, lightweight token-level router가 두 next-token distribution을 동적으로 섞는다.
router는 두 모델의 hidden representation, confidence, entropy feature를 이용해 \(\lambda_t\in[0,1]\)을 예측한다. router training 동안 backbone과 memory는 frozen이며, domain example에는 memory usage를 장려하고 general example에는 억제하는 signed regularizer를 함께 쓴다.
논리적으로 중요한 점은 specialization을 backbone rewriting에서 memory attachment로 바꾼다는 데 있다. 과학지식은 long-tailed이고 계속 바뀐다. 매 domain마다 전체 모델을 fine-tune하면 general reasoning이나 agentic capability를 흔들 수 있다.
300,000 time steps를 압축해서 읽는다
새 encoder는 temporal chunking, normalization, CNN local feature extraction, Q-Former compression, channel-wise Transformer, global temporal Transformer를 연결한다. 입력 길이에 따라 patching을 동적으로 조절해 output token 수를 통제한다.
Intern-S1-Pro의 약 240K step에서 약 300K로 늘고, channel mean-pooling 대신 channel-wise Transformer로 inter-channel dependency를 학습한다. astronomy, geoscience, neuroscience, physiological signal, bioacoustics뿐 아니라 MHz radar signal까지 범위를 넓힌다.
연속값을 텍스트 토큰으로 흉내내지 않는다
forecasting은 별도 numerical branch로 수행한다. LLM의 semantic context와 time-series encoder의 temporal representation을 Q-Former가 추출하고, cross-attention으로 causal Transformer forecaster를 조건화한다. horizon predictor는 instruction에서 필요한 예측 길이를 읽는다.
이 설계는 “숫자를 말하는 언어모델”과 “수치 신호를 생성하는 모델”의 차이를 인정한다. 긴 forecast를 discrete token으로 출력하면 길이 제한과 numerical precision이 동시에 문제가 되기 때문이다.
Memory Decoder는 biology를 올리지만 모든 subtask를 올리지는 않는다
Intern-MemDec-4B를 붙이면 Biology-Instructions 평균이 56.92에서 60.32로 상승한다. 하지만 모든 task가 개선되는 것은 아니다. 예컨대 DNA-cpd는 63.11→72.57, promoter-enhancer interaction은 22.46→38.47, RNA-CRISPROnTarget은 6.61→17.18로 오르지만 antibody–antigen은 40.24→36.44, Protein-Thermostability은 58.44→53.97로 내려간다.
따라서 “plug-and-play memory가 domain을 안전하게 완벽히 향상시킨다”보다, 평균적인 target-domain gain을 만들면서 cross-domain profile을 크게 흔들지 않는 specialization path라고 읽는 편이 정확하다.
Pre-training: 논문은 텍스트 파일이 아니라 페이지다
figure, table, equation, layout과 주변 문맥의 관계를 잃지 않도록 rendered page와 interleaved document를 함께 학습한다.
OCR 이전의 페이지를 직접 학습한다
Visual Pre-training(VP)은 large-scale unlabeled scientific document를 page image로 render한 뒤 frozen visual encoder의 latent를 autoregressively 예측한다. foreground mask로 blank region을 제거하고 raster-scan sequence를 만든다.
temperature-scaled cosine similarity와 in-batch negative를 이용한 contrastive next-latent objective를 쓰고, continued pre-training에서는 text CE와 visual loss를 결합한다.
VP의 장점은 OCR/layout parser나 paired annotation 없이도 page structure와 visual pattern을 흡수할 수 있다는 점이다.
그림의 뜻은 caption 하나에만 있지 않다
일반 image-caption pair는 지역적인 semantic alignment에는 좋지만 scientific PDF의 핵심 문맥을 놓친다. figure 앞뒤 문장, equation, table, cross-page reference와 문서의 논리적 순서가 함께 중요하기 때문이다.
파이프라인은 MinerU2.5-Pro로 OCR/layout parsing을 수행하고 image, interline equation, table을 crop한다. text와 visual unit을 reading order대로 interleave한 뒤, visual input을 넣었을 때 page-text perplexity가 얼마나 줄어드는지를 visual gain으로 사용한다.
human review와 domain threshold를 결합해 informative page만 남기고 원래 page order로 concatenate한다. chunk는 최대 256K token, overlap은 512 token이며 life science, chemistry, materials science에 집중한다.
좋은 과학 이미지를 대규모로 다시 끌어온다
image corpus는 SHA256으로 deduplicate하고 8B embedding model로 1024-dimensional vector를 만든다. hundreds-of-millions scale을 처리하기 위해 Milvus collection을 shard로 구성한다. online stage에서는 text-to-image와 image-to-image retrieval을 모두 지원하고, image query는 image embedding과 caption-text embedding을 함께 써 visual/semantic recall을 결합한다. recall 뒤 duplicate filtering, reranker, quality score filtering을 수행한다.
RL 이전에 과학적 행동의 문법을 만든다
SFT mixture는 general conversation, instruction following, safety, coding, image-text understanding, spatial grounding, tool use, specialized science, long-horizon agentic trajectory를 포함한다. explicit reasoning task의 demonstration은 Intern-S1-Pro와 다른 open model을 이용한 rejection sampling으로 만들고, language model과 human domain expert가 factual correctness, reasoning quality, format consistency를 검증한다.
Reasoning RL: 길게 생각하는 모델을 빠르고 안정적으로 훈련하는 법
partial rollout, off-policy correction, length regularization, online speculative decoding, GEPO가 하나의 RL objective로 합쳐진다.
느린 한 개의 trajectory 때문에 GPU 전체가 기다리지 않게 한다
XTuner와 LMDeploy를 같은 GPU pool에 두고 rollout과 policy update를 번갈아 수행한다. 충분한 completed trajectory가 모이면 unfinished rollout을 버리지 않고 현재 prefix와 metadata를 저장한 채 pause한다. 같은 GPU에서 training을 수행한 뒤 새 weight를 inference engine에 sync하고 그 prefix에서 generation을 resume한다.
그러나 한 trajectory 안에 서로 다른 policy version이 만든 token이 섞인다. 따라서 모든 sampled token에 behavior-policy version과 generation-time log probability를 기록하고 importance ratio를 사용한다. 가장 오래된 retained segment가 현재 learner보다 세 policy update를 초과하면 trajectory를 버린다.
MoE에서는 LMDeploy의 expert routing을 XTuner에서 다시 재생하는 Rollout Routing Replay(R3)를 쓴다. residual numerical mismatch는 bidirectional binary KL mask로 outlier token을 제외한다.
어려운 문제는 탐색하게 두고, 이미 푸는 문제만 짧게 만든다
negative response에는 length penalty를 적용하지 않는다. 틀린 답이 왜 틀렸는지 알기 어렵기 때문이다. 또한 group의 positive response 비율이 threshold를 넘을 때만 성공 response 안에서 짧은 solution의 advantage를 더 크게 준다.
normalization으로 전체 positive advantage mass를 대략 보존하기 때문에 task reward를 새 auxiliary reward로 바꾸지 않고 성공한 reasoning 사이의 상대 선호만 조절한다. 35B 실험에서 reward curve는 유사한데 평균 output length가 크게 줄었다.
draft model도 계속 배우게 해야 한다
정책이 RL 중 계속 변하면 고정 draft model은 빠르게 stale해진다. Intern-S2-Preview는 최신 policy trajectory로 draft를 online update한다. hybrid LK loss는 초기에 forward KL로 policy를 안정적으로 따라잡고 acceptance가 올라가면 total-variation component의 비중을 높인다.
구현에서 draft는 \(K=4\) future position을 예측하고 \(\eta=3\)을 쓴다. 대규모 run에서 rollout generation은 약 2×, 전체 RL pipeline은 약 1.7× 빨라졌다고 보고한다.
서로 다른 task를 하나의 entropy 목표에 억지로 맞추지 않는다
heterogeneous task는 solution diversity와 uncertainty가 다르므로 같은 policy 아래에서도 entropy regime이 다르다. GEPO(Group-level Entropy-Controlled Policy Optimization)는 기존 grouped samples에서 group entropy를 추정하고, low-entropy group의 positive advantage와 high-entropy group의 negative advantage를 선택적으로 attenuate한다.
목표는 모든 task를 동일 entropy로 만드는 것이 아니라 task-dependent exploration regime을 유지하면서 update contribution의 bias를 줄이는 데 있다.
안정화 장치들은 reward가 아니라 optimization path를 고친다
기본은 leave-one-out REINFORCE다. reward가 모두 같은 group은 DAPO식 dynamic sampling으로 online filtering한다. 이후 GEPO를 적용하고 adaptive length regularization을 적용한다.
Muon optimizer를 learning rate \(10^{-6}\), weight decay 0.01로 사용한다. rollout batch는 8,192 completed responses, 8 mini-batch updates이며 maximum generation length는 65,536 token이다.
Agentic RL: harness와 task를 분리하고, 행동을 token까지 추적한다
과학 agent를 훈련하려면 environment-level outcome과 model-level token trace가 다시 연결되어야 한다.
agent runtime과 task distribution을 독립적으로 바꾼다
harness는 agent가 어떻게 instantiate되고 drive되고 observe되는지를 정의하고, task는 initial environment, natural-language objective, verifier-defined outcome을 정의한다. 두 축을 분리하면 white-box loop와 black-box CLI/API를 같은 RL protocol 아래에서 조합할 수 있다.
black-box harness로 OpenClaw, Claude Code, OpenCode, OpenHands, Mini-SWE 등이 언급된다. 각 harness의 native message/tool loop를 유지한 채 adapter가 session lifecycle과 model call을 공통 runtime에 연결한다.
semantic trajectory와 policy token을 따로 저장한 뒤 다시 합친다
Agent Rollout Runner와 Judger Adapter는 action–observation trajectory, reward, process annotation, session metadata를 Replay Buffer에 저장한다. LLM Serving은 exact token ID, label, behavior log probability, MoE router expert를 Rollout Trace Store에 기록한다.
Token-In–Token-Out(TITO) interface는 이미 저장된 tokenized prefix를 재사용하고 새 context만 tokenize한다. Rollout Trace Store는 session을 incremental PrefixTree로 관리해 branching interaction에서도 model call 경계를 보존한다. training 시 system/user/tool observation은 loss에서 mask하고 policy-generated span만 label을 유지한다.
coding·terminal task를 executable environment로 정규화한다
| Provider | Collection | # Tasks | # Environments |
|---|---|---|---|
| SWE-bench | SWE-smith | 59,136 | 222 |
| SWE-Gym | SWE-Gym | 2,438 | 2,401 |
| R2E-Gym | R2E-Gym-V1 | 7,480 | 8,101 |
| Nebius | SWE-rebench-V2 | 32,100 | 32,075 |
| AweAI-Team | Scale-SWE | 20,200 | 19,472 |
| NVIDIA | Nemotron-Terminal-Synthetic-Tasks | 80,000 | 8 |
| RUC-AIBOX | ClawGym-Task | 13,500 | 1 |
각 instance는 base repository/container, task objective, test/reward verifier를 포함하는 common contract로 materialize된다. 정적인 instruction-response pair가 아니라 실제 repository state와 program behavior에 reward가 묶인다.
community skill에서 다음 task distribution을 만든다
community-contributed skill을 seed로 사용해 unavailable authentication, external transaction, toxic content, infeasible/low-quality/redundant workflow를 제거하고 domain-balanced resampling을 한다. skill은 observable state와 state-transforming capability로 바꿔 skill-state graph를 만든다. compatible transition만 이어 variable-length path를 sampling한다.
그 path를 environment → instruction → verifier 순서로 stage-wise synthesis하고 각 단계에 rule-based 및 rubric-based validator를 붙인다. failure는 해당 stage에서 repair/regeneration한다.
online/offline rollout 뒤 각 step을 normal progress, tool-use error, repetitive failure, invalid recovery, premature termination, protocol violation, unsupported assumption, hallucinated observation 등으로 annotate한다. 오류 step은 context에는 남겨 causal history를 보존하면서 imitation loss에서는 skip할 수 있다. execution failure 통계를 다시 sampling weight, skill, environment template, prompt에 반영해 다음 task distribution을 만든다.
성공한 trajectory의 나쁜 행동까지 칭찬하지 않는다
한 session의 final outcome reward에서 group-relative advantage를 만들고 해당 session의 모든 eligible policy-generated segment에 기본 credit을 준다. 하지만 malformed output, invalid/repeated tool call, 불필요한 recovery가 포함된 성공 trajectory도 있을 수 있다. process annotator는 해당 assistant message에만 weight를 붙인다.
process weight는 positive credit만 suppress 또는 reverse하고 실패 trajectory의 negative signal은 그대로 둔다.
reward hacking 방지도 구체적이다. gold patch, held-out test, exact scoring identifier는 agent workspace에서 숨기고 repository history는 baseline commit 하나로 sanitize하며 remote reference를 제거한다. evaluation 때 canonical tests를 agent 종료 뒤 restore/overlay하고 gold test patch도 사후 적용한다. infrastructure failure와 genuine task failure를 분리한다.
On-policy distillation과 평가: 넓은 능력은 얻었지만 모든 영역의 1등은 아니다
reasoning과 agentic specialization을 두 expert로 나눠 충분히 학습한 뒤, 다시 하나의 student에 합친다.
두 expert의 장점을 unified model로 되돌린다
mixed RL 하나로 매우 heterogeneous한 task를 동시에 최적화하면 optimization conflict가 생긴다. 따라서 같은 SFT checkpoint에서 reasoning expert와 agentic expert를 각각 train하고, teacher trajectory로 student를 가볍게 SFT warmup한 뒤 OPD를 수행한다.
최대 context 256K에서 full vocabulary 또는 top-64 teacher logits를 매 token 전달하면 통신량이 크다. 두 teacher와 student가 같은 SFT origin이고 warmup으로 distribution mismatch를 줄였기 때문에 sampled token의 teacher log-probability만 전달한다. payload complexity는 \(O(HV)\) 또는 \(O(Hk)\)에서 \(O(H)\)로 줄어든다.
reasoning RL과 같은 clipped importance weighting, R3, BKL numerical mask를 재사용하며 advantage만 verifier-based sequence signal에서 teacher–student token-level signal로 바뀐다.
과학 benchmark의 결과는 강하지만 profile이 균일하지 않다
| Benchmark | Intern-S2 | Qwen3.5 | DeepSeek-V4-pro | Kimi-K2.7-Code | GLM-5.2 | GPT-5.5 | Gemini-3.1-Pro | Claude-Opus-4.8 |
|---|---|---|---|---|---|---|---|---|
| Biology-Instructions | 56.92 | 4.49 | 9.14 | 7.68 | 6.34 | 10.52 | 13.87 | 6.78 |
| Mol-Instructions | 52.37 | 11.65 | 12.06 | 24.56 | 19.58 | 40.49 | 38.84 | 38.35 |
| MolecularIQ | 61.49 | 41.48 | 44.43 | 52.81 | 60.91 | 76.41 | 38.94 | 66.78 |
| SciReasoner | 63.97 | 45.02 | 51.11 | 51.69 | 51.45 | 61.15 | 60.35 | 58.00 |
| TOMG-Bench | 65.66 | 54.06 | 57.63 | 58.28 | 57.89 | 69.89 | 62.67 | 61.38 |
| MP20 | 67.88 | 6.15 | 6.75 | 8.40 | 1.50 | 16.12 | 16.75 | 15.60 |
| ProteinBinder-9 | 4.36 | 1.64 | 1.88 | 1.92 | 2.01 | 2.13 | 2.21 | 2.40 |
| XLRS-Bench | 51.97 | 50.11 | – | 49.90 | – | 50.96 | 54.27 | 51.84 |
| MicroVQA | 68.81 | 68.71 | – | 61.04 | – | 63.63 | 71.02 | 61.80 |
| SFE | 61.67 | 62.97 | – | 50.76 | – | 52.09 | 59.57 | 59.08 |
| ObsCrisis | 26.07 | 19.22 | – | 32.63 | – | 28.33 | 25.71 | 24.24 |
| SciCode | 49.11 | 46.35 | 47.53 | 43.49 | 51.97 | 55.92 | 54.44 | 56.21 |
| SGI-Bench | 49.37 | 44.44 | 45.70 | 50.63 | 52.41 | 42.77 | 45.28 | 49.06 |
| ResearchClawBench | 18.44 | 15.86 | 13.69 | 15.40 | 23.35 | 17.00 | 14.54 | 21.74 |
저자들은 Biology-Instructions, Mol-Instructions, SciReasoner에서 strong open/closed models를 앞선다고 보고하고, MolecularIQ·TOMG·XLRS·MicroVQA에서는 open-source 중 최고라고 정리한다. 반면 MolecularIQ와 TOMG에서는 GPT-5.5가 더 높고, 여러 agentic benchmark에서는 GLM-5.2·Claude Opus 4.8 등 closed model이 앞선다.
MP20과 ProteinBinder-9는 보고서가 자체 평가 세트로 기술하는 부분이 있으므로 외부 공개 benchmark와 동일한 재현성 수준으로 읽지 않는 편이 안전하다.
general-purpose에서도 open model 상위권이지만 closed frontier를 일괄 추월하지 않는다
| Benchmark | Intern-S2 | Qwen3.5 | DeepSeek-V4-pro | Kimi-K2.7 | GLM-5.2 | GPT-5.5 | Gemini-3.1 | Claude-4.8 |
|---|---|---|---|---|---|---|---|---|
| MMLU-Pro | 89.75 | 87.80 | 86.86 | 87.10 | 87.22 | 88.20 | 91.00 | 90.12 |
| SimpleQA-Verified | 69.90 | 54.80 | 46.60 | 38.60 | 37.90 | 64.30 | 75.60 | 43.30 |
| AdvancedIF | 74.44 | 75.49 | 73.83 | 76.17 | 75.76 | 76.20 | 79.78 | 72.88 |
| HMMT-2026 | 91.57 | 87.88 | 91.76 | 90.34 | 92.50 | 97.06 | 94.70 | 95.36 |
| MMMU-Pro | 80.46 | 80.29 | – | 77.92 | – | 81.68 | 83.99 | 76.88 |
| ChartQAPro | 69.65 | 68.61 | – | 54.86 | – | 69.23 | 71.18 | 58.65 |
| SkillsBench | 50.03 | 35.58 | 49.53 | 55.63 | 53.19 | 49.59 | 37.20 | 54.40 |
| TerminalBench 2.1 | 67.42 | 51.30 | 64.00 | 66.29 | 77.90 | 79.40 | 73.80 | 84.60 |
| SWE-Bench-Pro | 61.56 | 43.55 | 55.40 | 57.59 | 62.10 | 58.60 | 54.20 | 69.20 |
| SWE-Multilingual | 81.67 | 65.00 | 72.44 | 78.56 | 82.00 | 73.33 | 44.00 | 77.00 |
| WildClawBench | 44.68 | 34.50 | 43.70 | 46.89 | 54.20 | 58.20 | 40.80 | 64.72 |
SciTS에서는 underlying signal을 직접 모델링하는 것이 큰 차이를 만든다
time-series understanding에서 Intern-S2-Preview는 Intern-S1-Pro와 공통인 9개 task 중 7개를 앞선다. PHU01 F1은 36.8에서 66.9로 상승한다. 새 radar task RAU01, RAU02도 각각 88.4와 60.2 F1을 기록한다.
forecasting에서는 dedicated numerical branch가 모든 보고 task에서 100% success rate를 기록한다. Intern-S2의 MAPE는 ENG02 60.2, ENG03 7.1, MEG03 32.8, NEG03 59.2, PHG02 72.2, URG01 138.9, URG05 60.6이다. horizon predictor accuracy는 99%, GIFT-Eval zero-shot MASE는 0.785다.
무엇을 증명했고, 무엇은 아직 증명하지 못했는가
강점은 단일 benchmark보다 multimodal evidence와 executable interaction을 하나의 training stack에 묶은 데 있다.
평가가 다루는 과학의 범위
scientific benchmarks는 Biology-Instructions의 multi-omics sequence understanding, Mol-Instructions의 molecular/protein/biomolecular instruction, MolecularIQ의 SMILES graph reasoning, SciReasoner의 9개 scientific domain·149 task, TOMG의 molecule editing/property optimization/custom generation, MP20 crystal generation, ProteinBinder-9의 9 target binder design을 포함한다.
multimodal scientific evaluation은 XLRS의 ultra-high-resolution remote sensing, 1,042문항 MicroVQA, 830 verified VQA와 66 task의 SFE, 4,202 sample·127 event·8 disaster category·61 country의 ObsCrisis를 다룬다. agentic science는 80 main problem·338 subproblem·16 subfield의 SciCode, 1,263 sample·10 domain·75 research direction의 SGI-Bench, 실제 논문 기반 40 task·10 domain의 ResearchClawBench를 포함한다.
general evaluation은 MMLU-Pro, 1,000 verified prompt의 SimpleQA, 1,645 human-written prompt의 AdvancedIF, HMMT-2026 33문제, MMMU-Pro, 1,341 chart·1,948 question의 ChartQAPro와 SkillsBench, Terminal-Bench 2.1, SWE-Bench Pro, SWE-bench Multilingual, WildClawBench까지 이어진다.
세 개의 통합이 특히 중요하다
Evidence integration
텍스트, 페이지 이미지, equation/table layout, scientific image, numerical signal을 같은 foundation에 넣는다.
Execution integration
white/black-box harness, sandbox, executable verifier, semantic trace와 exact policy token을 하나의 RL protocol에 연결한다.
Capability integration
reasoning expert와 agentic expert를 충분히 specialize한 뒤 on-policy distillation으로 unified model에 합친다.
Preview라는 단어를 제목 장식으로 읽으면 안 된다
논문 결론은 Intern-S2-Preview가 여전히 preview system이라고 명시한다. 향후 과제는 더 긴 scientific workflow의 reliability, domain-specific memory와 task environment의 확장, verifier 강화, specialized scientific tool과의 더 깊은 통합이다.
benchmark profile도 균일하지 않다. Biology-Instructions나 Mol-Instructions처럼 특화 pre-training의 이점이 큰 task에서는 매우 강하지만 SciCode, ResearchClawBench, TerminalBench, WildClawBench처럼 복잡한 agentic task에서 최고 closed model을 항상 넘지는 못한다. 이것은 “scientific + agentic”이라는 이름이 곧 autonomous scientific discovery의 완성을 뜻하지 않는다는 증거다.
또한 executable coding/terminal task의 clean verifier는 큰 장점이지만 실제 wet-lab science는 reward가 지연되고 noisy하며 부분 관측이다. 이 training recipe가 그러한 실험과학의 verifier 문제까지 해결했다고 논문은 주장하지 않는다.
과학 에이전트의 핵심은 ‘똑똑한 답’보다 추적 가능한 실행이다
그러나 scientific AI를 static benchmark responder에서 executable, verifiable, iterative workflow participant로 옮기려면 무엇을 학습하고 무엇을 기록해야 하는지를 매우 구체적인 시스템 언어로 보여준다.
PrefixTree와 TITO는 agent의 semantic action을 실제 policy token과 연결한다. 이 bookkeeping이 중요한 까닭은 long-horizon RL의 credit assignment가 결국 “어떤 행동이 결과를 만들었는가”를 정확히 복원하는 문제이기 때문이다. process-aware weight와 verifier integrity를 붙이면 agentic post-training은 단순 conversation fine-tuning보다 실행시스템에 가까워진다.
Memory Decoder도 같은 철학을 공유한다. 모든 전문지식을 backbone에 영구적으로 합치는 대신 교체 가능한 parametric memory로 domain expertise의 lifecycle을 관리한다. scientific knowledge가 빠르게 변한다는 현실을 모델 구조에 반영한 셈이다.
보고서가 연결하는 연구 계보
원 보고서는 28–35쪽에 총 112개 참고문헌을 수록한다. 주요 계보는 AlphaFold 3와 scientific multimodal models, Memory Decoder/MemSFT/MLP Memory, SciTS와 time-series foundation models, partial rollout·AReaL·CoPRIS·APRIL, speculative decoding·FastGRPO·LK losses, GEPO와 entropy-control RL, DAPO·Muon, Agent-FLAN·Lagent·T-Eval·MindSearch·SciExplore, SWE-Gym·SWE-smith·R2E-Gym·ClawGym·POLAR, process supervision, on-policy distillation, ODesign/RFdiffusion과 scientific-agent benchmarks로 이어진다.
- Intern-S2-Preview: Scientific Agentic Foundation Model · arXiv:2608.13505
- Intern-S2-Preview model · Hugging Face
- XTuner · agentic RL infrastructure reference in the report
이 글의 수치·방법·benchmark 설명은 첨부된 35쪽 보고서 전체와 그 안의 표·수식·도식을 기준으로 재구성했다. 원 보고서의 112개 bibliography entry 자체는 연구 계보로 요약했으며, 결과를 독립적으로 재검증하기 위한 외부 web research는 수행하지 않았다.