ChatPaper.aiChatPaper

AstroPT가 은하에 대해 알고 있는 것, 그리고 그것이 LLM에 대해 시사하는 바

What AstroPT knows about galaxies, and what that can teach us about LLMs

August 23, 2026
저자: UniverseTBD, Kshitij Duraphe, Aman Kumar, Michael J. Smith, Shashwat Sourav
cs.AI

초록

해석 가능성 연구는 점점 더 훈련 중 개념이 언제 출현하는지, 그리고 선형 프로브가 실제 구조를 복원하는지에 대해 질문하고 있지만, 언어 모델에서는 이러한 주장을 검증하기 어렵다. 언어는 개념 간의 실측 순서나 관계에 대한 정보를 거의 제공하지 않기 때문이다. 우리는 수백만 개의 은하 이미지로 훈련된 트랜스포머인 AstroPT를 통해 천문학적 실측 자료를 활용할 것을 제안하며, 이를 교정 테스트베드로 사용한다. AstroPT는 개념의 난이도 순서와 개념 간 관계가 사전에 알려진 도메인 내에서 훈련된 LLM 유사 모델이다. 체크포인트, 계층, 모델 크기, 목적 함수 선택에 걸쳐 동결된 표현을 프로빙한 결과, 은하 속성은 알려진 난이도를 따르는 고정된 순서로 출현함을 발견했다. 픽셀에 거의 직접적으로 기록된 양(밴드 등급)은 훈련 초기와 네트워크의 얕은 계층에서부터 해독 가능해지는 반면, 다중 밴드/스펙트럼 기반 및 추론된 양(적색편이, 특정 항성 형성률 등)은 더 늦게 그리고 더 깊은 계층에서 출현한다. 이러한 순서는 우리가 테스트한 훈련 목적 함수에 대해 불변하며, 용량에 따라 크기는 증가하지만 순서는 변하지 않는다. 또한 우리의 선형 프로브 방향은 은하 속성 간의 알려진 물리적 구조를 복원한다. 우리의 발견은 천문학이 우리가 그 외에는 맹목적으로 LLM에 적용하는 기계론적 해석 가능성 방법을 교정하기 위한 통제된 실험 환경을 제공함을 시사한다.
English
Interpretability research increasingly asks when concepts emerge during training and whether linear probes recover real structure, but in language models these claims are hard to validate because language offers little ground-truth ordering of concepts or relationships among them. We propose the use of astronomical ground truth through AstroPT, a transformer trained on millions of galaxy images, as a calibration testbed. AstroPT is an LLM-like model trained within a domain where the difficulty ordering of concepts and the relations among them are known in advance. Probing frozen representations across checkpoints, layers, model sizes, and objective choices, we find that galaxy properties emerge in a fixed order that tracks their known difficulty---quantities written almost directly into the pixels (band magnitude) become decodable early in training and shallow in the network, while multiband/spectra based and inferred quantities (such as redshift and specific star formation rate) emerge later and deeper. This order is invariant to our tested training objectives, and scales in magnitude but not in sequence with capacity. Our linear probe directions further recover the known physical structure among galaxy properties. Our findings suggest that astronomy offers a controlled sandbox for calibrating mechanistic interpretability methods we otherwise apply to LLMs blind.