ChatPaper.ai
打開菜單
首頁
每日論文
arXiv
HuggingFace
定價
賬戶
工作台
🇭🇰
繁體中文
Loading...
•
•
•
•
•
•
•
•
•
•
AI研究論文每日精選
每日精選AI研究論文及翻譯
April 4th, 2025
場景中心的無監督全景分割
Scene-Centric Unsupervised Panoptic Segmentation
Oliver Hahn, Christoph Reich, Nikita Araslanov, Daniel Cremers, Christian Rupprecht, Stefan Roth
•
Apr 2, 2025
•
5
3
JavisDiT:基於層次化時空先驗同步的聯合音視頻擴散Transformer
JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization
Kai Liu, Wei Li, Lai Chen, Shengqiong Wu, Yanhao Zheng, Jiayi Ji, Fan Zhou, Rongxin Jiang, Jiebo Luo, Hao Fei, Tat-Seng Chua
•
Mar 30, 2025
•
54
4
SkyReels-A2:在视频扩散变换器中实现任意内容合成
SkyReels-A2: Compose Anything in Video Diffusion Transformers
Zhengcong Fei, Debang Li, Di Qiu, Jiahua Wang, Yikun Dou, Rui Wang, Jingtao Xu, Mingyuan Fan, Guibin Chen, Yang Li, Yahui Zhou
•
Apr 3, 2025
•
36
3
Whisper-LM:利用語言模型提升低資源語言的語音辨識模型效能
Whisper-LM: Improving ASR Models with Language Models for Low-Resource Languages
Xabier de Zuazo, Eva Navas, Ibon Saratxaga, Inma Hernáez Rioja
•
Mar 30, 2025
•
10
3
基礎代理的進展與挑戰:從類腦智能到演化、協作與安全系統
Advances and Challenges in Foundation Agents: From Brain-Inspired Intelligence to Evolutionary, Collaborative, and Safe Systems
Bang Liu, Xinfeng Li, Jiayi Zhang, Jinlin Wang, Tanjin He, Sirui Hong, Hongzhang Liu, Shaokun Zhang, Kaitao Song, Kunlun Zhu, Yuheng Cheng, Suyuchen Wang, Xiaoqiang Wang, Yuyu Luo, Haibo Jin, Peiyan Zhang, Ollie Liu, Jiaqi Chen, Huan Zhang, Zhaoyang Yu, Haochen Shi, Boyan Li, Dekun Wu, Fengwei Teng, Xiaojun Jia, Jiawei Xu, Jinyu Xiang, Yizhang Lin, Tianming Liu, Tongliang Liu, Yu Su, Huan Sun, Glen Berseth, Jianyun Nie, Ian Foster, Logan Ward, Qingyun Wu, Yu Gu, Mingchen Zhuge, Xiangru Tang, Haohan Wang, Jiaxuan You, Chi Wang, Jian Pei, Qiang Yang, Xiaoliang Qi, Chenglin Wu
•
Mar 31, 2025
•
270
7
OpenCodeReasoning:推進競技編程中的數據蒸餾技術
OpenCodeReasoning: Advancing Data Distillation for Competitive Coding
Wasi Uddin Ahmad, Sean Narenthiran, Somshubra Majumdar, Aleksander Ficek, Siddhartha Jain, Jocelyn Huang, Vahid Noroozi, Boris Ginsburg
•
Apr 2, 2025
•
15
3
音視頻控制下的視頻擴散與掩碼選擇性狀態空間建模 ——面向自然對話頭像生成的技術
Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation
Fa-Ting Hong, Zunnan Xu, Zixiang Zhou, Jun Zhou, Xiu Li, Qin Lin, Qinglin Lu, Dan Xu
•
Apr 3, 2025
•
44
7
解讀無模型強化學習中的湧現式規劃
Interpreting Emergent Planning in Model-Free Reinforcement Learning
Thomas Bush, Stephen Chung, Usman Anwar, Adrià Garriga-Alonso, David Krueger
•
Apr 2, 2025
•
12
2
NeuralGS:融合神經場與3D高斯潑濺技術,實現緊湊的3D表徵
NeuralGS: Bridging Neural Fields and 3D Gaussian Splatting for Compact 3D Representations
Zhenyu Tang, Chaoran Feng, Xinhua Cheng, Wangbo Yu, Junwu Zhang, Yuan Liu, Xiaoxiao Long, Wenping Wang, Li Yuan
•
Mar 29, 2025
•
11
2
WikiVideo:基於多部影片的文章生成
WikiVideo: Article Generation from Multiple Videos
Alexander Martin, Reno Kriz, William Gantt Walden, Kate Sanders, Hannah Recknor, Eugene Yang, Francis Ferraro, Benjamin Van Durme
•
Apr 1, 2025
•
36
3
指令引導的自回歸神經網絡參數生成
Instruction-Guided Autoregressive Neural Network Parameter Generation
Soro Bedionita, Bruno Andreis, Song Chong, Sung Ju Hwang
•
Apr 2, 2025
•
6
2
重新思考視覺語言模型的強化學習擴展:一個透明、從零開始的框架與全面評估方案
Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme
Yan Ma, Steffi Chern, Xuyang Shen, Yiran Zhong, Pengfei Liu
•
Apr 3, 2025
•
30
3
交錯式語音-文本語言模型的規模化分析
Scaling Analysis of Interleaved Speech-Text Language Models
Gallil Maimon, Michael Hassid, Amit Roth, Yossi Adi
•
Apr 3, 2025
•
28
2
基於大型語言模型的時間序列預測高效模型選擇
Efficient Model Selection for Time Series Forecasting via LLMs
Wang Wei, Tiankai Yang, Hongjie Chen, Ryan A. Rossi, Yue Zhao, Franck Dernoncourt, Hoda Eldardiry
•
Apr 2, 2025
•
16
2
ZClip:大型語言模型預訓練中的自適應尖峰抑制
ZClip: Adaptive Spike Mitigation for LLM Pre-Training
Abhay Kumar, Louis Owen, Nilabhra Roy Chowdhury, Fabian Güra
•
Apr 3, 2025
•
77
2
AI與機器人科學家在科學發現中的規模化定律
Scaling Laws in Scientific Discovery with AI and Robot Scientists
Pengsong Zhang, Heng Zhang, Huazhe Xu, Renjun Xu, Zhenting Wang, Cong Wang, Animesh Garg, Zhibin Li, Arash Ajoudani, Xinyu Liu
•
Mar 28, 2025
•
12
2
超越像素的想象:推理引导视觉编辑的基准测试
Envisioning Beyond the Pixels: Benchmarking Reasoning-Informed Visual Editing
Xiangyu Zhao, Peiyuan Zhang, Kexian Tang, Hao Li, Zicheng Zhang, Guangtao Zhai, Junchi Yan, Hua Yang, Xue Yang, Haodong Duan
•
Apr 3, 2025
•
67
2
稀疏自編碼器在視覺語言模型中學習單語義特徵
Sparse Autoencoders Learn Monosemantic Features in Vision-Language Models
Mateusz Pach, Shyamgopal Karthik, Quentin Bouniot, Serge Belongie, Zeynep Akata
•
Apr 3, 2025
•
10
2
GPT-ImgEval:全面診斷GPT4o圖像生成能力的基準測試
GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation
Zhiyuan Yan, Junyan Ye, Weijia Li, Zilong Huang, Shenghai Yuan, Xiangyang He, Kaiqing Lin, Jun He, Conghui He, Li Yuan
•
Apr 3, 2025
•
56
3
推理時期的通用獎勵模型縮放
Inference-Time Scaling for Generalist Reward Modeling
Zijun Liu, Peiyi Wang, Runxin Xu, Shirong Ma, Chong Ruan, Peng Li, Yang Liu, Yu Wu
•
Apr 3, 2025
•
54
6
GenPRM:通過生成式推理擴展過程獎勵模型的測試時計算能力
GenPRM: Scaling Test-Time Compute of Process Reward Models via Generative Reasoning
Jian Zhao, Runze Liu, Kaiyan Zhang, Zhimu Zhou, Junqi Gao, Dong Li, Jiafei Lyu, Zhouyi Qian, Biqing Qi, Xiu Li, Bowen Zhou
•
Apr 1, 2025
•
12
3
ShortV:通過凍結無效層中的視覺標記來實現高效的多模態大型語言模型
ShortV: Efficient Multimodal Large Language Models by Freezing Visual Tokens in Ineffective Layers
Qianhao Yuan, Qingyu Zhang, Yanjiang Liu, Jiawei Chen, Yaojie Lu, Hongyu Lin, Jia Zheng, Xianpei Han, Le Sun
•
Apr 1, 2025
•
21
2
FreSca:揭示擴散模型中的縮放空間
FreSca: Unveiling the Scaling Space in Diffusion Models
Chao Huang, Susan Liang, Yunlong Tang, Li Ma, Yapeng Tian, Chenliang Xu
•
Apr 2, 2025
•
19
2