한 번에 더 똑똑하고 저렴해진다: 바이트 단위 정확 KV-캐시 접목이 고정된 소형 모델을 검증된 지식 플라이휠로 변환
Smarter and Cheaper at Once: Byte-Exact KV-Cache Grafting Turns a Frozen Small Model into a Verified-Knowledge Flywheel
July 15, 2026
저자: Sietse Schelpe
cs.AI
초록
본 논문에서는 가중치를 전혀 변경하지 않고도 고정된 소형 언어 모델의 성능을 향상시키면서 동시에 비용을 획기적으로 낮추는 방법을 보고한다. 검증된 지식은 바이트 단위로 정확한 키-값(KV) 상태 인공물로서 한 번 저장되었다가, 이후 새로운 추론 컨텍스트에 이식(graft)되어 복원된다. 복원 과정은 비트 단위로 정확하다: 고정된 결정적 구성 하에서, 이식된 로짓(logits)은 새로 계산한 결과와 바이트 단위로 완전히 동일하며(SHA-256 동등성 검증), KL 발산은 0이고 50개 샘플에 걸쳐 100% argmax 일치를 보인다. 부동소수점 회전 인코딩을 사용하는 모델에서 자기 위치 이식(own-position graft)이 유일하게 수치적으로 정확한 동작점임을 보이며, 두 모델 규모(12B, 31B)와 두 GPU 대상(그중 하나는 사전 등록된 재생을 통해)에서 바이트 단위 정확성을 검증한다. AIME 2025에서, 고정된 Gemma-4-12B는 검증된 솔루션 라이브러리가 이식되자 80.0%에서 93.3%로 성능이 상승했으며, 이는 자체의 발표된 기준점인 77.5% 및 31B 형제 모델의 89.2%를 상회한다. 반복 사례의 경우, 베이스 모델이 401,026 토큰 예산 내에서 결코 해결하지 못하는 여덟 개의 문제가 캐시된 검증 솔루션으로부터 총 61개의 디코드 토큰만으로 답변되었으며, 이는 6,574배 적은 토큰과 약 8,700배 적은 에너지 소비에 해당한다. 성능 향상 자체는 보류 전이(31B에서 7개 중 7개 성공)를 통해 입증된다. 동일한 바이트 단위 저장소는 사용 가능한 컨텍스트를 추가 가속기 메모리 없이 32,768에서 2,854,766 토큰으로 확장하며, 동일 아키텍처의 기기 간에 바이트 단위로 동일하게 전송된다. 시스템은 동작 수준에서 기술되며, 엔진은 독점적이지만 보고된 모든 수치는 확정된 입력 및 출력 해시로 뒷받침되므로 엔진 없이도 채점 결과를 재확인할 수 있다.
English
We report a way to make a frozen small language model both more capable and dramatically cheaper at once, without changing any weights. Verified knowledge is deposited once as a byte-exact key-value (KV) state artifact and later restored, by graft, into a fresh inference context. The restore is bit-exact: under a pinned deterministic configuration, the grafted logits are byte-for-byte identical to a fresh computation (SHA-256 equality), with zero KL divergence and 100% argmax agreement over fifty samples. We show that own-position graft is the unique numerically exact operating point on a model with floating-point rotary encoding, and we verify byte-exactness on two model scales (12B, 31B) and two GPU targets, one through a pre-registered replay. On AIME 2025, a frozen Gemma-4-12B moves from 80.0% to 93.3% once a verified solution library is grafted, above its own 77.5% and its 31B sibling's 89.2% published anchors. On the recurring case, eight problems the base model never solves within a 401,026-token budget are answered from cached verified solutions in 61 total decode tokens, a factor of 6,574 fewer tokens and about 8,700x less energy; the capability claim proper rests on held-out transfer (7 of 7 at 31B). The same byte-exact store widens usable context from 32,768 to 2,854,766 tokens at zero extra accelerator memory, and moves byte-identical between machines of the same architecture. We describe the system at the behavior level; the engine is proprietary, and every reported number is backed by committed input and output hashes so the scoring can be re-checked without it.