AstroPT对星系的认知,及其对我们理解大语言模型的启示
What AstroPT knows about galaxies, and what that can teach us about LLMs
August 23, 2026
作者: UniverseTBD, Kshitij Duraphe, Aman Kumar, Michael J. Smith, Shashwat Sourav
cs.AI
摘要
可解释性研究日益关注概念在训练中何时出现,以及线性探针能否恢复真实结构;但在语言模型中,这些主张很难验证,因为语言几乎没有提供概念之间顺序或关系的真值。我们提出借助AstroPT——一个在数百万星系图像上训练的Transformer——利用天文学真值作为校准测试平台。AstroPT是一个类似LLM的模型,其训练领域内概念的难度排序及概念之间的关系是预先已知的。通过跨检查点、层、模型规模和训练目标选择对冻结表征进行线性探针探测,我们发现星系属性以固定顺序涌现,这一顺序与已知难度一致:几乎直接编码在像素中的量(如波段星等)在训练早期即可解码且位于网络浅层,而基于多波段/光谱的量以及推断量(如红移和比恒星形成率)则在更晚阶段且更深层出现。这一顺序对测试的训练目标保持不变,并且随容量增大在幅度上增强、在顺序上不变。我们的线性探针方向还进一步恢复了星系属性之间已知的物理结构。我们的研究结果表明,天文学为校准机制可解释性方法提供了一个受控沙盒,而这些方法我们原本会盲目地应用于大语言模型。
English
Interpretability research increasingly asks when concepts emerge during training and whether linear probes recover real structure, but in language models these claims are hard to validate because language offers little ground-truth ordering of concepts or relationships among them. We propose the use of astronomical ground truth through AstroPT, a transformer trained on millions of galaxy images, as a calibration testbed. AstroPT is an LLM-like model trained within a domain where the difficulty ordering of concepts and the relations among them are known in advance. Probing frozen representations across checkpoints, layers, model sizes, and objective choices, we find that galaxy properties emerge in a fixed order that tracks their known difficulty---quantities written almost directly into the pixels (band magnitude) become decodable early in training and shallow in the network, while multiband/spectra based and inferred quantities (such as redshift and specific star formation rate) emerge later and deeper. This order is invariant to our tested training objectives, and scales in magnitude but not in sequence with capacity. Our linear probe directions further recover the known physical structure among galaxy properties. Our findings suggest that astronomy offers a controlled sandbox for calibrating mechanistic interpretability methods we otherwise apply to LLMs blind.