ChronoLens: 시간, 언어, 언어 층위에 걸친 언어 변화 측정
ChronoLens: Measuring Language Change Across Time, Languages, and Linguistic Levels
August 4, 2026
저자: Gagan Bhatia, Julian Schlenker, Simone Paolo Ponzetto, Steffen Eger
cs.AI
초록
역사적 언어 변화는 형태론, 통사론, 의미론, 화용론에 영향을 미치지만, 계산적 연구는 일반적으로 이러한 수준들을 서로 호환되지 않는 표상으로 검토하므로 언어 전반에 걸쳐 이들이 함께 진화하는지 여부를 판별할 수 없다. 우리는 단일 분석 공간 내에서 변화의 크기와 방향이 언어적 수준, 언어, 역사적 시기에 따라 어떻게 달라지는지를 묻는 방식으로 이 문제를 다룬다. 우리는 동결된 다국어 언어 모델, 특징 정렬 교차 코더, 사후 언어적 개입을 결합한 프레임워크인 ChronoLens를 소개하고, 이를 1803~2026년에 걸친 다섯 개 의회 전통의 4,498만 개 문서와 약 172억 개 토큰에 적용한다. 그 결과 얻어진 희소 표상은 밀집 임베딩이나 통합 희소 오토인코더보다 언어 통계량과 훨씬 더 강하게 일치하며(ρ=0.72 대 0.29 및 0.28), 형태론, 통사론, 의미론, 화용론이 일반적으로 한 언어 내에서 비교 가능한 정도로 변화하는 반면, 언어들은 변화의 시기, 범위, 방향에 있어 현저한 차이를 보인다는 것을 밝혀낸다. 이러한 발견은 역사적 언어 변화가 구조화된 다차원적 과정임을 보여준다. 즉, 유사한 크기가 상이한 궤적을 은폐할 수 있으며, 의미 있는 언어 간 비교는 거리와 방향을 모두 측정해야 한다.
English
Historical language change affects morphology, syntax, semantics, and pragmatics, yet computational studies typically examine these levels with incompatible representations and therefore cannot determine whether they evolve together across languages. We address this problem by asking how the magnitude and direction of change vary across linguistic levels, languages, and historical periods within a single analytical space. We introduce ChronoLens, a framework that combines frozen multilingual language models, feature-aligned crosscoders, and post-hoc linguistic interventions, and apply it to 44.98 million documents and approximately 17.2 billion tokens from five parliamentary traditions spanning 1803--2026. The resulting sparse representations agree substantially more strongly with linguistic statistics than dense embeddings or a pooled sparse autoencoder (ρ=0.72 versus 0.29 and 0.28), and reveal that morphology, syntax, semantics, and pragmatics generally change by comparable amounts within a language, while languages differ markedly in when, how far, and in which direction they change. These findings show that historical language change is a structured, multidimensional process: similar magnitudes can conceal different trajectories, and meaningful cross-linguistic comparison requires measuring both distance and direction.