ChatPaper
AI 论文热榜
近 7 天全球 AI 研究者都在点赞的论文,人话解读
1
Repo-To-Skill:将 GitHub 代码仓库蒸馏为 AI4AI 技能
🔥 494
2026-09-02
深入阅读
arXiv 原文
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills
2
StudentSim:基于大语言模型的学生模拟器训练
🔥 458
2026-09-01
深入阅读
arXiv 原文
StudentSim: Training LLM-based Student Simulators
3
Qwen-Drive-1.0:迈向自动驾驶视觉语言基础模型的初步探索
🔥 337
2026-08-31
深入阅读
arXiv 原文
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving
4
HarnessDev:大语言模型能否创建并演进自身的智能体框架?
🔥 224
2026-09-01
深入阅读
arXiv 原文
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness?
5
Terminal-Universe:将智能体轨迹转化为可扩展的终端环境
🔥 213
2026-09-03
深入阅读
arXiv 原文
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments
6
LLaDA-Image:以全开放训练方案构建强大的图像生成器
🔥 196
2026-09-03
深入阅读
arXiv 原文
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes
7
Aspire:模型能否从模糊目标中自我进化?
🔥 173
2026-08-31
深入阅读
arXiv 原文
Aspire: Can Models Self-Evolve from Vague Goals?
8
懂得何时不重用:自主LLM后训练中的条件化经验迁移
🔥 139
2026-08-27
深入阅读
arXiv 原文
Knowing When Not to Reuse: Conditional Experience Transfer in Autonomous LLM Post-Training
9
SolarWM:面向长时程视频世界模型的开放数据与可扩展训练
🔥 133
2026-09-02
深入阅读
arXiv 原文
SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models
10
随机注意力:重新思考KV缓存清除以实现高效推理
🔥 123
2026-09-03
深入阅读
arXiv 原文
Random Attention: Rethinking KV Cache Eviction for Efficient Reasoning
11
EarlyEval:基于早期结果预测的低成本智能体评估
🔥 110
2026-09-02
深入阅读
arXiv 原文
EarlyEval: Cheaper Agent Evaluation via Early Outcome Prediction
12
LatentPress: Context Compression Beyond Text and Vision
🔥 102
2026-09-01
深入阅读
arXiv 原文
13
在策略蒸馏真的在蒸馏吗?从噪声教师到自我提升
🔥 98
2026-08-31
深入阅读
arXiv 原文
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement
14
LoopArena:模型作为循环工程运行时控制器的基准测试
🔥 98
2026-08-28
深入阅读
arXiv 原文
LoopArena: Benchmarking Models as Runtime Controllers for Loop Engineering
15
DreamX-Creator:普及2K分辨率的原生音视频生成
🔥 90
2026-08-31
深入阅读
arXiv 原文
DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution
16
DART-SD:面向多轮工具调用智能体自蒸馏的菱形拓扑感知检索与调优
🔥 88
2026-08-19
深入阅读
arXiv 原文
DART-SD: Diamond-topology Aware Retrieval and Tuning for Self-Distillation of Multi-Turn Tool-Calling Agents
17
超越数据扩展:面向视觉-语言-动作模型的表征中心持续预训练
🔥 84
2026-08-27
深入阅读
arXiv 原文
Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models
18
SMELT: 计算匹配的MoE循环Transformer的缩放定律
🔥 75
2026-09-01
深入阅读
arXiv 原文
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers
19
Lucida:面向可组合真实到仿真场景建模的解析、生成与放置
🔥 72
2026-08-31
深入阅读
arXiv 原文
Lucida: Parse, Generate, and Place for Composable Real-to-Sim Scene Modeling
20
为何门控DeltaNet能经受4比特量化考验:混合27B大语言模型中递归部分的NVFP4 W4A4量化方案
🔥 66
2026-09-03
深入阅读
arXiv 原文
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM
每天 5 分钟,跟上 AI 前沿
每日精选 5 篇热门论文,讲成人话,配一道小谜题。
去看今日刊 →