ChatPaper.ai
打开菜单
首页
每日论文
arXiv
HuggingFace
定价
账户
工作台
🇨🇳
中文简体
Loading...
•
•
•
•
•
•
•
•
•
•
AI研究论文每日精选
每日精选AI研究论文及翻译
March 26th, 2025
基于下一帧预测的长上下文自回归视频建模
Long-Context Autoregressive Video Modeling with Next-Frame Prediction
Yuchao Gu, Weijia Mao, Mike Zheng Shou
•
Mar 25, 2025
•
72
2
将视觉预训练扩展至4K分辨率
Scaling Vision Pre-Training to 4K Resolution
Baifeng Shi, Boyi Li, Han Cai, Yao Lu, Sifei Liu, Marco Pavone, Jan Kautz, Song Han, Trevor Darrell, Pavlo Molchanov, Hongxu Yin
•
Mar 25, 2025
•
40
2
通过随机生成与滚动预算强制实现流模型的推理时缩放
Inference-Time Scaling for Flow Models via Stochastic Generation and Rollover Budget Forcing
Jaihoon Kim, Taehoon Yoon, Jisung Hwang, Minhyuk Sung
•
Mar 25, 2025
•
33
4
探索大型多模态模型在视频理解中的幻觉现象:基准、分析与缓解策略
Exploring Hallucination of Large Multimodal Models in Video Understanding: Benchmark, Analysis and Mitigation
Hongcheng Gao, Jiashu Qu, Jingyi Tang, Baolong Bi, Yue Liu, Hongyu Chen, Li Liang, Li Su, Qingming Huang
•
Mar 25, 2025
•
31
4
CoMP:面向视觉基础模型的持续多模态预训练
CoMP: Continual Multimodal Pre-training for Vision Foundation Models
Yitong Chen, Lingchen Meng, Wujian Peng, Zuxuan Wu, Yu-Gang Jiang
•
Mar 24, 2025
•
30
1
三思而后行:通过扩展多轮测试时思考提升大语言模型推理能力
Think Twice: Enhancing LLM Reasoning by Scaling Multi-round Test-time Thinking
Xiaoyu Tian, Sitong Zhao, Haotian Wang, Shuaiting Chen, Yunjie Ji, Yiping Peng, Han Zhao, Xiangang Li
•
Mar 25, 2025
•
26
5
识破伪造:基于大型多模态模型的合成图像检测与伪影解析
Spot the Fake: Large Multimodal Model-Based Synthetic Image Detection with Artifact Explanation
Siwei Wen, Junyan Ye, Peilin Feng, Hengrui Kang, Zichen Wen, Yize Chen, Jiang Wu, Wenjun Wu, Conghui He, Weijia Li
•
Mar 19, 2025
•
20
3
MDocAgent:面向文档理解的多模态多智能体框架
MDocAgent: A Multi-Modal Multi-Agent Framework for Document Understanding
Siwei Han, Peng Xia, Ruiyi Zhang, Tong Sun, Yun Li, Hongtu Zhu, Huaxiu Yao
•
Mar 18, 2025
•
19
2
ReSearch:通过强化学习让大语言模型掌握基于搜索的推理能力
ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
Mingyang Chen, Tianpeng Li, Haoze Sun, Yijie Zhou, Chenzheng Zhu, Fan Yang, Zenan Zhou, Weipeng Chen, Haofen Wang, Jeff Z. Pan, Wen Zhang, Huajun Chen
•
Mar 25, 2025
•
17
3
CoLLM:面向组合图像检索的大型语言模型
CoLLM: A Large Language Model for Composed Image Retrieval
Chuong Huynh, Jinyu Yang, Ashish Tawari, Mubarak Shah, Son Tran, Raffay Hamid, Trishul Chilimbi, Abhinav Shrivastava
•
Mar 25, 2025
•
14
2
WikiAutoGen:迈向多模态维基百科式文章生成
WikiAutoGen: Towards Multi-Modal Wikipedia-Style Article Generation
Zhongyu Yang, Jun Chen, Dannong Xu, Junjie Fei, Xiaoqian Shen, Liangbing Zhao, Chun-Mei Feng, Mohamed Elhoseiny
•
Mar 24, 2025
•
11
2
潜在空间超分辨率:基于扩散模型的高分辨率图像生成
Latent Space Super-Resolution for Higher-Resolution Image Generation with Diffusion Models
Jinho Jeong, Sangmin Han, Jinwoo Kim, Seon Joo Kim
•
Mar 24, 2025
•
10
1
FullDiT:具备全注意力机制的多任务视频生成基础模型
FullDiT: Multi-Task Video Generative Foundation Model with Full Attention
Xuan Ju, Weicai Ye, Quande Liu, Qiulin Wang, Xintao Wang, Pengfei Wan, Di Zhang, Kun Gai, Qiang Xu
•
Mar 25, 2025
•
8
2
DiffPortrait360:面向360度视角合成的连贯肖像扩散模型
DiffPortrait360: Consistent Portrait Diffusion for 360 View Synthesis
Yuming Gu, Phong Tran, Yujian Zheng, Hongyi Xu, Heyuan Li, Adilbek Karmanov, Hao Li
•
Mar 19, 2025
•
8
2
FirePlace:基于几何优化的LLM常识推理在3D物体摆放中的应用
FirePlace: Geometric Refinements of LLM Common Sense Reasoning for 3D Object Placement
Ian Huang, Yanan Bao, Karen Truong, Howard Zhou, Cordelia Schmid, Leonidas Guibas, Alireza Fathi
•
Mar 6, 2025
•
8
2
PhysTwin:基于物理约束的视频可变形物体重建与仿真
PhysTwin: Physics-Informed Reconstruction and Simulation of Deformable Objects from Videos
Hanxiao Jiang, Hao-Yu Hsu, Kaifeng Zhang, Hsin-Ni Yu, Shenlong Wang, Yunzhu Li
•
Mar 23, 2025
•
7
2
前瞻调优:通过部分答案预览打造更安全的语言模型
LookAhead Tuning: Safer Language Models via Partial Answer Previews
Kangwei Liu, Mengru Wang, Yujie Luo, Lin Yuan, Mengshu Sun, Ningyu Zhang, Lei Liang, Zhiqiang Zhang, Jun Zhou, Huajun Chen
•
Mar 24, 2025
•
5
3
通过微调迁移实现高效模型开发
Efficient Model Development through Fine-tuning Transfer
Pin-Jie Lin, Rishab Balasubramanian, Fengyuan Liu, Nikhil Kandpal, Tu Vu
•
Mar 25, 2025
•
4
2
FRESA:基于少量图像的前馈式个性化蒙皮虚拟角色重建
FRESA:Feedforward Reconstruction of Personalized Skinned Avatars from Few Images
Rong Wang, Fabian Prada, Ziyan Wang, Zhongshi Jiang, Chengxiang Yin, Junxuan Li, Shunsuke Saito, Igor Santesteban, Javier Romero, Rohan Joshi, Hongdong Li, Jason Saragih, Yaser Sheikh
•
Mar 24, 2025
•
4
2
xKV:面向KV缓存压缩的跨层奇异值分解
xKV: Cross-Layer SVD for KV-Cache Compression
Chi-Chih Chang, Chien-Yu Lin, Yash Akhauri, Wei-Cheng Lin, Kai-Chiang Wu, Luis Ceze, Mohamed S. Abdelfattah
•
Mar 24, 2025
•
4
1
基于直通式引导的Gumbel-Softmax流匹配技术用于可控生物序列生成
Gumbel-Softmax Flow Matching with Straight-Through Guidance for Controllable Biological Sequence Generation
Sophia Tang, Yinuo Zhang, Alexander Tong, Pranam Chatterjee
•
Mar 21, 2025
•
4
2
强基线:基于YOLOv12与BoT-SORT-ReID的多无人机目标跟踪
Strong Baseline: Multi-UAV Tracking via YOLOv12 with BoT-SORT-ReID
Yu-Hsi Chen
•
Mar 21, 2025
•
4
5
当文字超越视觉:视觉语言模型通过纯文本训练实现自我提升,助力以人为本的决策
When Words Outperform Vision: VLMs Can Self-Improve Via Text-Only Training For Human-Centered Decision Making
Zhe Hu, Jing Li, Yu Yin
•
Mar 21, 2025
•
4
2
迈向统一的哥白尼地球视觉基础模型
Towards a Unified Copernicus Foundation Model for Earth Vision
Yi Wang, Zhitong Xiong, Chenying Liu, Adam J. Stewart, Thomas Dujardin, Nikolaos Ioannis Bountos, Angelos Zavras, Franziska Gerken, Ioannis Papoutsis, Laura Leal-Taixé, Xiao Xiang Zhu
•
Mar 14, 2025
•
4
3
LLaVAction:面向动作识别的多模态大语言模型评估与训练
LLaVAction: evaluating and training multi-modal large language models for action recognition
Shaokai Ye, Haozhe Qi, Alexander Mathis, Mackenzie W. Mathis
•
Mar 24, 2025
•
3
2
Any6D:新型物体的无模型6D姿态估计
Any6D: Model-free 6D Pose Estimation of Novel Objects
Taeyeop Lee, Bowen Wen, Minjun Kang, Gyuree Kang, In So Kweon, Kuk-Jin Yoon
•
Mar 24, 2025
•
3
2
OpenCity3D:视觉-语言模型对城市环境了解多少?
OpenCity3D: What do Vision-Language Models know about Urban Environments?
Valentin Bieri, Marco Zamboni, Nicolas S. Blumer, Qingxuan Chen, Francis Engelmann
•
Mar 21, 2025
•
3
2
视觉语言模型能否在现实世界中应对面对面提问?
Can Vision-Language Models Answer Face to Face Questions in the Real-World?
Reza Pourreza, Rishit Dagli, Apratim Bhattacharyya, Sunny Panchal, Guillaume Berger, Roland Memisevic
•
Mar 25, 2025
•
2
2
克服词汇不匹配:词汇无关的教师引导语言建模
Overcoming Vocabulary Mismatch: Vocabulary-agnostic Teacher Guided Language Modeling
Haebin Shin, Lei Ji, Xiao Liu, Yeyun Gong
•
Mar 24, 2025
•
2
2
频率动态卷积用于密集图像预测
Frequency Dynamic Convolution for Dense Image Prediction
Linwei Chen, Lin Gu, Liang Li, Chenggang Yan, Ying Fu
•
Mar 24, 2025
•
2
2
LPOSS:基于图像块与像素的标签传播实现开放词汇语义分割
LPOSS: Label Propagation Over Patches and Pixels for Open-vocabulary Semantic Segmentation
Vladan Stojnić, Yannis Kalantidis, Jiří Matas, Giorgos Tolias
•
Mar 25, 2025
•
1
2
ST-VLM:面向视觉语言模型时空推理的运动学指令微调
ST-VLM: Kinematic Instruction Tuning for Spatio-Temporal Reasoning in Vision-Language Models
Dohwan Ko, Sihyeon Kim, Yumin Suh, Vijay Kumar B. G, Minseo Yoon, Manmohan Chandraker, Hyunwoo J. Kim
•
Mar 25, 2025
•
1
1
Co-SemDepth:航空影像的快速联合语义分割与深度估计
Co-SemDepth: Fast Joint Semantic Segmentation and Depth Estimation on Aerial Images
Yara AlaaEldin, Francesca Odone
•
Mar 23, 2025
•
0
2