Puro-2B: $5090 예산으로 RTX 5090에서 학습된 Poor Lab의 Qwen2-1.5B
Puro-2B: Poor Lab's Qwen2-1.5B Trained on RTX 5090 within $5090
August 27, 2026
저자: Kairong Luo, Jiarui Cui, Yaorui Yin, Shengqi Chen, Yiming Yang, Linxiang Gao, Yanmohan Wang, Mingzhe Zhang, Kaiyue Wen, Kaifeng Lyu, Wenguang Chen
cs.AI
초록
언어 모델 사전 학습은 거의 감당하기 어려운 비용과 동의어가 되면서, 학계와 오픈소스 커뮤니티의 상당 부분이 접근하기 어려운 영역이 되었다. 오픈 가중치 모델과 오픈소스 훈련 레시피를 포함한 강력한 오픈소스 노력이 이미 존재하지만, 비용 효율적이고 하드웨어 접근성이 높은 오픈소스 사전 학습 레시피는 오랫동안 부재했다. 소규모에서도 Llama-3.2-3B를 학습하는 데는 150만 달러 이상이 들고, SmolLM3-3B를 재현하는 데는 70만 달러 이상이 필요하다. 본 보고서에서는 이러한 장벽을 낮추기 위해 설계된 오픈 사전 학습 레시피를 제시한다. 이 레시피를 사용하여 우리는 소비자용 RTX 5090 GPU에서 FP8 정밀도로 최대 1.4조 토큰에 대해 Puro-2B 모델 모음을 처음부터 학습한다. 모음의 모델들은 토큰 예산과 선택된 레시피 변형에 따라 다르다. 우리의 최고 모델은 6,900달러 미만의 컴퓨팅 비용으로 학습되었으며, 우리의 평가 프로토콜에서 Qwen2.5-1.5B 성능에 근접한다. 이러한 비용 효율성은 하드웨어 선택, 저정밀 훈련, 하이퍼볼 최적화, 커리큘럼 모델 평균화, 데이터 레시피를 포함한 여러 접근법의 결합으로 달성된다. 레시피 자체 외에도 우리는 두 가지 추가 결과를 제공한다. 첫째, Puro-2B 모음 전반에 걸쳐 훈련 비용과 평균 모델 성능을 연관 짓는 Puro 비용 스케일링 법칙(Puro Cost Scaling Law)을 도출한다. 적합된 법칙에 따르면 약 4.4K 달러, 즉 5,090달러 미만의 비용으로 Qwen2-1.5B의 성능에 도달하기에 충분하다. 둘째, 엔드투엔드 사례 연구로서 사전 학습 데이터 커리큘럼이 후속 훈련 이후의 다운스트림 성능을 어떻게 형성하는지 살펴본다. 이러한 통제된 연구는 모델 가중치만이 아니라 전체 사전 학습 파이프라인에 접근할 수 있기 때문에 가능하다. 우리는 데이터, 코드, 모델 가중치를 포함한 Puro-2B의 전체 훈련 레시피를 Apache 2.0 라이선스로 https://huggingface.co/collections/thu-pacman/puro-2b 에 공개한다.
English
Language model pretraining has become almost synonymous with prohibitive cost, placing it out of reach for much of the academic and open-source communities. Although strong open-source efforts already exist, including open-weight models and open-source training recipes, a cost-efficient, hardware-accessible, and open-source pretraining recipe has long been missing. Even at a small scale, training Llama-3.2-3B costs over \1.5M, and reproducing SmolLM3-3B needs over 700K. In this report, we present an open pretraining recipe designed to lower this barrier. Using this recipe, we train a collection of Puro-2B models from scratch on up to 1.4 trillion tokens with FP8 precision on consumer-grade RTX 5090 GPUs. The models in the collection differ in token budgets and selected recipe variants. Our best model is trained at a compute cost of less than \6.9K and approaches Qwen2.5-1.5B performance under our evaluation protocol. This cost efficiency is enabled by a combination of approaches, including hardware selection, low-precision training, hyperball optimization, curriculum model averaging, and the data recipe. Beyond the recipe itself, we provide two additional results. First, across the Puro-2B collection, we derive a Puro Cost Scaling Law that relates training cost to average model performance; the fitted law suggests that about 4.4K, less than \$5,090, is sufficient to reach the performance of Qwen2-1.5B. Second, as an end-to-end case study, we examine how pretraining data curricula shape downstream performance after post-training. Such controlled studies are enabled by having access to the full pretraining pipeline rather than model weights alone. We release the full training recipe for Puro-2B, including data, code, and model weights under Apache 2.0 at https://huggingface.co/collections/thu-pacman/puro-2b.