ChatPaper.ai
打開菜單
首頁
每日論文
arXiv
HuggingFace
定價
賬戶
工作台
🇭🇰
繁體中文
Loading...
•
•
•
•
•
•
•
•
•
•
AI研究論文每日精選
每日精選AI研究論文及翻譯
March 25th, 2025
影片簡答QA:邁向大型影片語言模型的事實性評估
Video SimpleQA: Towards Factuality Evaluation in Large Video Language Models
Meng Cao, Pengfei Hu, Yingyao Wang, Jihao Gu, Haoran Tang, Haoze Zhao, Jiahua Dong, Wangbo Yu, Ge Zhang, Ian Reid, Xiaodan Liang
•
Mar 24, 2025
•
12
1
Aether:具備幾何感知的統一世界建模
Aether: Geometric-Aware Unified World Modeling
Aether Team, Haoyi Zhu, Yifan Wang, Jianjun Zhou, Wenzheng Chang, Yang Zhou, Zizun Li, Junyi Chen, Chunhua Shen, Jiangmiao Pang, Tong He
•
Mar 24, 2025
•
28
2
大型語言模型預訓練中的權重重新縮放方差控制
Variance Control via Weight Rescaling in LLM Pre-training
Louis Owen, Abhay Kumar, Nilabhra Roy Chowdhury, Fabian Güra
•
Mar 21, 2025
•
5
2
定位:交互式生成视频作为下一代游戏引擎
Position: Interactive Generative Video as Next-Generation Game Engine
Jiwen Yu, Yiran Qin, Haoxuan Che, Quande Liu, Xintao Wang, Pengfei Wan, Di Zhang, Xihui Liu
•
Mar 21, 2025
•
62
3
心靈之眼:從語言推理到多模態推理
Mind with Eyes: from Language Reasoning to Multimodal Reasoning
Zhiyu Lin, Yifei Gao, Xian Zhao, Yunfan Yang, Jitao Sang
•
Mar 23, 2025
•
3
2
DynamicVis:一種高效且通用的視覺基礎模型,用於遙感影像理解
DynamicVis: An Efficient and General Visual Foundation Model for Remote Sensing Image Understanding
Keyan Chen, Chenyang Liu, Bowen Chen, Wenyuan Li, Zhengxia Zou, Zhenwei Shi
•
Mar 20, 2025
•
0
2
FFN融合:重新思考大型語言模型中的序列計算
FFN Fusion: Rethinking Sequential Computation in Large Language Models
Akhiad Bercovich, Mohammad Dabbah, Omri Puny, Ido Galil, Amnon Geifman, Yonatan Geifman, Izhak Golan, Ehud Karpas, Itay Levy, Zach Moshe, Najeeb Nabwani, Tomer Ronen, Itamar Schen, Elad Segal, Ido Shahaf, Oren Tropp, Ran Zilberstein, Ran El-Yaniv
•
Mar 24, 2025
•
19
3
AgentRxiv:邁向協作式自主研究
AgentRxiv: Towards Collaborative Autonomous Research
Samuel Schmidgall, Michael Moor
•
Mar 23, 2025
•
22
2
等變圖像建模
Equivariant Image Modeling
Ruixiao Dong, Mengde Xu, Zigang Geng, Li Li, Han Hu, Shuyang Gu
•
Mar 24, 2025
•
15
1
OmnimatteZero:基於預訓練視頻擴散模型的免訓練實時Omnimatte生成
OmnimatteZero: Training-free Real-time Omnimatte with Pre-trained Video Diffusion Models
Dvir Samuel, Matan Levy, Nir Darshan, Gal Chechik, Rami Ben-Ari
•
Mar 23, 2025
•
25
2
全能判官:跨模态的多模态大语言模型评判系统
Judge Anything: MLLM as a Judge Across Any Modality
Shu Pu, Yaochen Wang, Dongping Chen, Yuhang Chen, Guohao Wang, Qi Qin, Zhongyi Zhang, Zhiyuan Zhang, Zetong Zhou, Shuang Gong, Yi Gui, Yao Wan, Philip S. Yu
•
Mar 21, 2025
•
20
2
RDTF:面向多帧動態貼圖生成的資源高效雙遮罩訓練框架
RDTF: Resource-efficient Dual-mask Training Framework for Multi-frame Animated Sticker Generation
Zhiqiang Yuan, Ting Zhang, Ying Deng, Jiapei Zhang, Yeshuang Zhu, Zexi Jia, Jie Zhou, Jinchao Zhang
•
Mar 22, 2025
•
3
2
無需訓練的瓶頸採樣擴散加速法
Training-free Diffusion Acceleration with Bottleneck Sampling
Ye Tian, Xin Xia, Yuxi Ren, Shanchuan Lin, Xing Wang, Xuefeng Xiao, Yunhai Tong, Ling Yang, Bin Cui
•
Mar 24, 2025
•
12
4
AMD-Hummingbird:邁向高效文本到視頻模型
AMD-Hummingbird: Towards an Efficient Text-to-Video Model
Takashi Isobe, He Cui, Dong Zhou, Mengmeng Ge, Dong Li, Emad Barsoum
•
Mar 24, 2025
•
5
2
Instruct-CLIP:利用對比學習進行自動數據精煉以提升指令引導的圖像編輯效果
Instruct-CLIP: Improving Instruction-Guided Image Editing with Automated Data Refinement Using Contrastive Learning
Sherry X. Chen, Misha Sra, Pradeep Sen
•
Mar 24, 2025
•
3
2
重探圖像融合技術於多光源白平衡校正之應用
Revisiting Image Fusion for Multi-Illuminant White-Balance Correction
David Serrano-Lozano, Aditya Arora, Luis Herranz, Konstantinos G. Derpanis, Michael S. Brown, Javier Vazquez-Corral
•
Mar 18, 2025
•
1
2
通過設計擊敗提示注入攻擊
Defeating Prompt Injections by Design
Edoardo Debenedetti, Ilia Shumailov, Tianqi Fan, Jamie Hayes, Nicholas Carlini, Daniel Fabian, Christoph Kern, Chongyang Shi, Andreas Terzis, Florian Tramèr
•
Mar 24, 2025
•
20
1
從潛在思維中推理學習
Reasoning to Learn from Latent Thoughts
Yangjun Ruan, Neil Band, Chris J. Maddison, Tatsunori Hashimoto
•
Mar 24, 2025
•
13
1
優化的最小化3D高斯噴濺
Optimized Minimal 3D Gaussian Splatting
Joo Chan Lee, Jong Hwan Ko, Eunbyung Park
•
Mar 21, 2025
•
13
2
Diffusion-4K:基於潛在擴散模型的超高解析度影像合成
Diffusion-4K: Ultra-High-Resolution Image Synthesis with Latent Diffusion Models
Jinjin Zhang, Qiuyu Huang, Junjie Liu, Xiefan Guo, Di Huang
•
Mar 24, 2025
•
6
2
Feather-SQL:一款面向小型语言模型的轻量级NL2SQL框架,采用双模型协作范式
Feather-SQL: A Lightweight NL2SQL Framework with Dual-Model Collaboration Paradigm for Small Language Models
Wenqi Pei, Hailing Xu, Hengyuan Zhao, Shizheng Hou, Han Chen, Zining Zhang, Pingyi Luo, Bingsheng He
•
Mar 22, 2025
•
13
2
Typed-RAG:面向非事實性問答的類型感知多維度分解
Typed-RAG: Type-aware Multi-Aspect Decomposition for Non-Factoid Question Answering
DongGeon Lee, Ahjeong Park, Hyeri Lee, Hyeonseo Nam, Yunho Maeng
•
Mar 20, 2025
•
6
2
我已在各方面做好準備:透過稀疏自編碼器解讀大型語言模型中的推理特徵
I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders
Andrey Galichin, Alexey Dontsov, Polina Druzhinina, Anton Razzhigaev, Oleg Y. Rogov, Elena Tutubalina, Ivan Oseledets
•
Mar 24, 2025
•
118
2
口語化過程監督引導出更優異的編程代理
Verbal Process Supervision Elicits Better Coding Agents
Hao-Yuan Chen, Cheng-Pong Huang, Jui-Ming Yao
•
Mar 24, 2025
•
2
2
QuartDepth:面向邊緣設備即時深度估計的訓練後量化技術
QuartDepth: Post-Training Quantization for Real-Time Depth Estimation on the Edge
Xuan Shen, Weize Ma, Jing Liu, Changdi Yang, Rui Ding, Quanyi Wang, Henghui Ding, Wei Niu, Yanzhi Wang, Pu Zhao, Jun Lin, Jiuxiang Gu
•
Mar 20, 2025
•
0
2
迷失在文化轉譯中:大型語言模型是否在跨文化情境下的數學表現上遇到困難?
Lost in Cultural Translation: Do LLMs Struggle with Math Across Cultural Contexts?
Aabid Karim, Abdul Karim, Bhoomika Lohana, Matt Keon, Jaswinder Singh, Abdul Sattar
•
Mar 23, 2025
•
6
2
CFG-Zero*:流匹配模型的改進版無分類器引導
CFG-Zero*: Improved Classifier-Free Guidance for Flow Matching Models
Weichen Fan, Amber Yijia Zheng, Raymond A. Yeh, Ziwei Liu
•
Mar 24, 2025
•
21
2
SimpleRL-Zoo:探索與馴化開放基礎模型在實際應用中的零樣本強化學習
SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild
Weihao Zeng, Yuzhen Huang, Qian Liu, Wei Liu, Keqing He, Zejun Ma, Junxian He
•
Mar 24, 2025
•
30
1
V-Seek:加速基於開放硬體伺服器級RISC-V平台的大型語言模型推理
V-Seek: Accelerating LLM Reasoning on Open-hardware Server-class RISC-V Platforms
Javier J. Poveda Rodrigo, Mohamed Amine Ahmdi, Alessio Burrello, Daniele Jahier Pagliari, Luca Benini
•
Mar 21, 2025
•
6
2
Video-T1:視頻生成的測試時間縮放
Video-T1: Test-Time Scaling for Video Generation
Fangfu Liu, Hanyang Wang, Yimo Cai, Kaiyan Zhang, Xiaohang Zhan, Yueqi Duan
•
Mar 24, 2025
•
88
1
引理:從錯誤中學習以促進大型語言模型的數學進步
LEMMA: Learning from Errors for MatheMatical Advancement in LLMs
Zhuoshi Pan, Yu Li, Honglin Lin, Qizhi Pei, Zinan Tang, Wei Wu, Chenlin Ming, H. Vicky Zhao, Conghui He, Lijun Wu
•
Mar 21, 2025
•
15
2
人類動作反學習
Human Motion Unlearning
Edoardo De Matteis, Matteo Migliarini, Alessio Sampieri, Indro Spinelli, Fabio Galasso
•
Mar 24, 2025
•
1
2
Vision-R1:通過視覺引導強化學習實現大型視覺語言模型的人類無監督對齊
Vision-R1: Evolving Human-Free Alignment in Large Vision-Language Models via Vision-Guided Reinforcement Learning
Yufei Zhan, Yousong Zhu, Shurong Zheng, Hongyin Zhao, Fan Yang, Ming Tang, Jinqiao Wang
•
Mar 23, 2025
•
19
2
AlphaSpace:透過語義標記化與符號推理實現機器人行動
AlphaSpace: Enabling Robotic Actions through Semantic Tokenization and Symbolic Reasoning
Alan Dao, Dinh Bach Vu, Bui Quang Huy
•
Mar 24, 2025
•
10
2
重新思考超分辨率中的图像评估
Rethinking Image Evaluation in Super-Resolution
Shaolin Su, Josep M. Rocafort, Danna Xue, David Serrano-Lozano, Lei Sun, Javier Vazquez-Corral
•
Mar 17, 2025
•
1
2
MagicComp:面向組合式視頻生成的無訓練雙階段精煉方法
MagicComp: Training-free Dual-Phase Refinement for Compositional Video Generation
Hongyu Zhang, Yufan Deng, Shenghai Yuan, Peng Jin, Zesen Cheng, Yian Zhao, Chang Liu, Jie Chen
•
Mar 18, 2025
•
8
2
CODA:重新利用連續變分自編碼器實現離散標記化
CODA: Repurposing Continuous VAEs for Discrete Tokenization
Zeyu Liu, Zanlin Ni, Yeguo Hua, Xin Deng, Xiao Ma, Cheng Zhong, Gao Huang
•
Mar 22, 2025
•
3
2
全域-局部樹狀搜索用於語言引導的三維場景生成
Global-Local Tree Search for Language Guided 3D Scene Generation
Wei Deng, Mengshi Qi, Huadong Ma
•
Mar 24, 2025
•
0
2
MetaSpatial:強化視覺語言模型在元宇宙中的三維空間推理能力
MetaSpatial: Reinforcing 3D Spatial Reasoning in VLMs for the Metaverse
Zhenyu Pan, Han Liu
•
Mar 24, 2025
•
3
2