ChatPaper.aiChatPaper

「AstroPT 對星系的認識,以及這對大型語言模型的啟示」

What AstroPT knows about galaxies, and what that can teach us about LLMs

August 23, 2026
作者: UniverseTBD, Kshitij Duraphe, Aman Kumar, Michael J. Smith, Shashwat Sourav
cs.AI

摘要

可解釋性研究日益關注概念在訓練過程中何時湧現,以及線性探針是否能夠恢復真實結構,但在語言模型中,這些主張難以驗證,因為語言本身缺乏概念的難易程度排序或其間關係的真實對照依據。我們提出透過 AstroPT 來利用天文學的真實標注——AstroPT 是一個在數百萬張星系影像上訓練的 Transformer,可作為校準測試平台。AstroPT 是類似大型語言模型的模型,但在一個概念的難易排序及其間關係皆可事先得知的領域中訓練。透過對跨檢查點、層數、模型規模與目標函數選擇的凍結表徵進行探測,我們發現星系性質以固定的順序湧現,且該順序與其已知的難易程度一致——幾乎直接寫入像素的量(如波段星等)在訓練早期即能被解碼,且在網路淺層即可解碼;而基於多波段、光譜的推定量(如紅移與比恆星形成率)則較晚出現,且出現於較深層。此順序不受我們所測試的各種訓練目標影響,且隨模型容量增加而規模放大,但順序不變。我們的線性探針方向進一步恢復了星系性質之間已知的物理結構。我們的研究結果表明,天文學提供了一個受控的沙盒環境,可用於校準我們在缺乏真實對照時應用於大型語言模型的機制性可解釋性方法。
English
Interpretability research increasingly asks when concepts emerge during training and whether linear probes recover real structure, but in language models these claims are hard to validate because language offers little ground-truth ordering of concepts or relationships among them. We propose the use of astronomical ground truth through AstroPT, a transformer trained on millions of galaxy images, as a calibration testbed. AstroPT is an LLM-like model trained within a domain where the difficulty ordering of concepts and the relations among them are known in advance. Probing frozen representations across checkpoints, layers, model sizes, and objective choices, we find that galaxy properties emerge in a fixed order that tracks their known difficulty---quantities written almost directly into the pixels (band magnitude) become decodable early in training and shallow in the network, while multiband/spectra based and inferred quantities (such as redshift and specific star formation rate) emerge later and deeper. This order is invariant to our tested training objectives, and scales in magnitude but not in sequence with capacity. Our linear probe directions further recover the known physical structure among galaxy properties. Our findings suggest that astronomy offers a controlled sandbox for calibrating mechanistic interpretability methods we otherwise apply to LLMs blind.