AstroPTは銀河について何を知っているのか、そしてそれが大規模言語モデルについて何を教えてくれるのか
What AstroPT knows about galaxies, and what that can teach us about LLMs
August 23, 2026
著者: UniverseTBD, Kshitij Duraphe, Aman Kumar, Michael J. Smith, Shashwat Sourav
cs.AI
要旨
解釈可能性研究は、訓練中に概念がいつ出現するのか、また線形プローブが実際の構造を回復するのかをますます問うようになっている。しかし言語モデルでは、言語が概念やその関係についてのグラウンドトゥルースの順序をほとんど提供しないため、これらの主張を検証することは困難である。我々は、数百万枚の銀河画像で訓練されたトランスフォーマーであるAstroPTを通じた天文学的グラウンドトゥルースの利用を、較正テストベッドとして提案する。AstroPTは、概念の難易度順序と概念間の関係が事前に既知である領域内で訓練された、LLM類似のモデルである。チェックポイント、層、モデルサイズ、目的関数の選択を横断して凍結表現をプロービングしたところ、銀河の特性が既知の難易度を追跡する固定された順序で出現することが分かった。すなわち、ピクセルにほぼ直接書き込まれる量(バンド等級)は訓練の初期かつネットワークの浅い層でデコード可能になる一方、多バンド・スペクトルに基づく量や推定される量(赤方偏移や比星形成率など)は後期かつ深い層で出現する。この順序は、我々が試験した訓練目的関数に対して不変であり、容量に応じてその大きさはスケールするが、順序は変化しない。さらに、我々の線形プローブ方向は、銀河特性間の既知の物理的構造を回復する。これらの発見は、天文学が、我々がそうでなければLLMに対して盲目的に適用している機構的解釈可能性手法を較正するための、制御されたサンドボックスを提供することを示唆している。
English
Interpretability research increasingly asks when concepts emerge during training and whether linear probes recover real structure, but in language models these claims are hard to validate because language offers little ground-truth ordering of concepts or relationships among them. We propose the use of astronomical ground truth through AstroPT, a transformer trained on millions of galaxy images, as a calibration testbed. AstroPT is an LLM-like model trained within a domain where the difficulty ordering of concepts and the relations among them are known in advance. Probing frozen representations across checkpoints, layers, model sizes, and objective choices, we find that galaxy properties emerge in a fixed order that tracks their known difficulty---quantities written almost directly into the pixels (band magnitude) become decodable early in training and shallow in the network, while multiband/spectra based and inferred quantities (such as redshift and specific star formation rate) emerge later and deeper. This order is invariant to our tested training objectives, and scales in magnitude but not in sequence with capacity. Our linear probe directions further recover the known physical structure among galaxy properties. Our findings suggest that astronomy offers a controlled sandbox for calibrating mechanistic interpretability methods we otherwise apply to LLMs blind.