ChatPaper.ai
メニューを開く
ホーム
今日の論文
arXiv
HuggingFace
料金プラン
アカウント
ワークスペース
🇯🇵
日本語
Loading...
•
•
•
•
•
•
•
•
•
•
AI研究論文デイリー
翻訳付きの日次キュレーションされたAI研究論文
October 18th, 2024
映画ジェン:メディア基盤モデルのキャスト
Movie Gen: A Cast of Media Foundation Models
Adam Polyak, Amit Zohar, Andrew Brown, Andros Tjandra, Animesh Sinha, Ann Lee, Apoorv Vyas, Bowen Shi, Chih-Yao Ma, Ching-Yao Chuang, David Yan, Dhruv Choudhary, Dingkang Wang, Geet Sethi, Guan Pang, Haoyu Ma, Ishan Misra, Ji Hou, Jialiang Wang, Kiran Jagadeesh, Kunpeng Li, Luxin Zhang, Mannat Singh, Mary Williamson, Matt Le, Matthew Yu, Mitesh Kumar Singh, Peizhao Zhang, Peter Vajda, Quentin Duval, Rohit Girdhar, Roshan Sumbaly, Sai Saketh Rambhatla, Sam Tsai, Samaneh Azadi, Samyak Datta, Sanyuan Chen, Sean Bell, Sharadh Ramaswamy, Shelly Sheynin, Siddharth Bhattacharya, Simran Motwani, Tao Xu, Tianhe Li, Tingbo Hou, Wei-Ning Hsu, Xi Yin, Xiaoliang Dai, Yaniv Taigman, Yaqiao Luo, Yen-Cheng Liu, Yi-Chiao Wu, Yue Zhao, Yuval Kirstain, Zecheng He, Zijian He, Albert Pumarola, Ali Thabet, Artsiom Sanakoyeu, Arun Mallya, Baishan Guo, Boris Araya, Breena Kerr, Carleigh Wood, Ce Liu, Cen Peng, Dimitry Vengertsev, Edgar Schonfeld, Elliot Blanchard, Felix Juefei-Xu, Fraylie Nord, Jeff Liang, John Hoffman, Jonas Kohler, Kaolin Fire, Karthik Sivakumar, Lawrence Chen, Licheng Yu, Luya Gao, Markos Georgopoulos, Rashel Moritz, Sara K. Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petrovic, Yuming Du
•
Oct 17, 2024
•
99
2
MixEval-X: 実世界データの混合からの任意対任意評価
MixEval-X: Any-to-Any Evaluations from Real-World Data Mixtures
Jinjie Ni, Yifan Song, Deepanway Ghosal, Bo Li, David Junhao Zhang, Xiang Yue, Fuzhao Xue, Zian Zheng, Kaichen Zhang, Mahir Shah, Kabir Jain, Yang You, Michael Shieh
•
Oct 17, 2024
•
76
2
JudgeBench: LLMベースの判定者を評価するためのベンチマーク
JudgeBench: A Benchmark for Evaluating LLM-based Judges
Sijun Tan, Siyuan Zhuang, Kyle Montgomery, William Y. Tang, Alejandro Cuadron, Chenguang Wang, Raluca Ada Popa, Ion Stoica
•
Oct 16, 2024
•
48
2
Fluid: 連続トークンを用いた自己回帰テキストから画像への生成モデルのスケーリング
Fluid: Scaling Autoregressive Text-to-image Generative Models with Continuous Tokens
Lijie Fan, Tianhong Li, Siyang Qin, Yuanzhen Li, Chen Sun, Michael Rubinstein, Deqing Sun, Kaiming He, Yonglong Tian
•
Oct 17, 2024
•
38
3
Janus: 統一されたマルチモーダル理解と生成のための視覚エンコーディングの切り離し
Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Chengyue Wu, Xiaokang Chen, Zhiyu Wu, Yiyang Ma, Xingchao Liu, Zizheng Pan, Wen Liu, Zhenda Xie, Xingkai Yu, Chong Ruan, Ping Luo
•
Oct 17, 2024
•
35
4
大規模言語モデルを用いた超人的な音声理解に向けたロードマップ
Roadmap towards Superhuman Speech Understanding using Large Language Models
Fan Bu, Yuhao Zhang, Xidong Wang, Benyou Wang, Qun Liu, Haizhou Li
•
Oct 17, 2024
•
35
2
MobA: 効率的なモバイルタスク自動化のための二層エージェントシステム
MobA: A Two-Level Agent System for Efficient Mobile Task Automation
Zichen Zhu, Hao Tang, Yansi Li, Kunyao Lan, Yixuan Jiang, Hao Zhou, Yixiao Wang, Situo Zhang, Liangtai Sun, Lu Chen, Kai Yu
•
Oct 17, 2024
•
33
3
WorldCuisines: グローバル料理に関する多言語および多文化のビジュアル質問応答のための大規模ベンチマーク
WorldCuisines: A Massive-Scale Benchmark for Multilingual and Multicultural Visual Question Answering on Global Cuisines
Genta Indra Winata, Frederikus Hudi, Patrick Amadeus Irawan, David Anugraha, Rifki Afina Putri, Yutong Wang, Adam Nohejl, Ubaidillah Ariq Prathama, Nedjma Ousidhoum, Afifa Amriani, Anar Rzayev, Anirban Das, Ashmari Pramodya, Aulia Adila, Bryan Wilie, Candy Olivia Mawalim, Ching Lam Cheng, Daud Abolade, Emmanuele Chersoni, Enrico Santus, Fariz Ikhwantri, Garry Kuwanto, Hanyang Zhao, Haryo Akbarianto Wibowo, Holy Lovenia, Jan Christian Blaise Cruz, Jan Wira Gotama Putra, Junho Myung, Lucky Susanto, Maria Angelica Riera Machin, Marina Zhukova, Michael Anugraha, Muhammad Farid Adilazuarda, Natasha Santosa, Peerat Limkonchotiwat, Raj Dabre, Rio Alexander Audino, Samuel Cahyawijaya, Shi-Xiong Zhang, Stephanie Yulia Salim, Yi Zhou, Yinxuan Gui, David Ifeoluwa Adelani, En-Shiun Annie Lee, Shogo Okada, Ayu Purwarianti, Alham Fikri Aji, Taro Watanabe, Derry Tanti Wijaya, Alice Oh, Chong-Wah Ngo
•
Oct 16, 2024
•
33
3
テキスト豊かなビジュアル理解のためのWebページUIの活用
Harnessing Webpage UIs for Text-Rich Visual Understanding
Junpeng Liu, Tianyue Ou, Yifan Song, Yuxiao Qu, Wai Lam, Chenyan Xiong, Wenhu Chen, Graham Neubig, Xiang Yue
•
Oct 17, 2024
•
32
2
DreamVideo-2: 主体によるゼロショット動画カスタマイズと正確なモーション制御
DreamVideo-2: Zero-Shot Subject-Driven Video Customization with Precise Motion Control
Yujie Wei, Shiwei Zhang, Hangjie Yuan, Xiang Wang, Haonan Qiu, Rui Zhao, Yutong Feng, Feng Liu, Zhizhong Huang, Jiaxin Ye, Yingya Zhang, Hongming Shan
•
Oct 17, 2024
•
25
2
MMed-RAG: 医療ビジョン言語モデル向けの多目的マルチモーダルRAGシステム
MMed-RAG: Versatile Multimodal RAG System for Medical Vision Language Models
Peng Xia, Kangyu Zhu, Haoran Li, Tianze Wang, Weijia Shi, Sheng Wang, Linjun Zhang, James Zou, Huaxiu Yao
•
Oct 16, 2024
•
23
3
MoH: マルチヘッド注意機構をヘッドの混合注意として
MoH: Multi-Head Attention as Mixture-of-Head Attention
Peng Jin, Bo Zhu, Li Yuan, Shuicheng Yan
•
Oct 15, 2024
•
22
2
BenTo: コンテキスト内転移を用いたベンチマークタスクの削減
BenTo: Benchmark Task Reduction with In-Context Transferability
Hongyu Zhao, Ming Li, Lichao Sun, Tianyi Zhou
•
Oct 17, 2024
•
20
3
PopAlign: より包括的なアライメントのための対照的なパターンの多様化
PopAlign: Diversifying Contrasting Patterns for a More Comprehensive Alignment
Zekun Moore Wang, Shawn Wang, Kang Zhu, Jiaheng Liu, Ke Xu, Jie Fu, Wangchunshu Zhou, Wenhao Huang
•
Oct 17, 2024
•
19
2
OpenAIのo1モデルの推論パターンに関する比較研究
A Comparative Study on Reasoning Patterns of OpenAI's o1 Model
Siwei Wu, Zhongyuan Peng, Xinrun Du, Tuney Zheng, Minghao Liu, Jialong Wu, Jiachen Ma, Yizhi Li, Jian Yang, Wangchunshu Zhou, Qunshu Lin, Junbo Zhao, Zhaoxiang Zhang, Wenhao Huang, Ge Zhang, Chenghua Lin, J. H. Liu
•
Oct 17, 2024
•
19
2
事後トレーニング済みの大規模モデルにおけるデルタパラメータ編集の統一された視点
A Unified View of Delta Parameter Editing in Post-Trained Large-Scale Models
Qiaoyu Tang, Le Yu, Bowen Yu, Hongyu Lin, Keming Lu, Yaojie Lu, Xianpei Han, Le Sun
•
Oct 17, 2024
•
17
2
FlatQuant: LLM 量子化においてフラットさが重要である
FlatQuant: Flatness Matters for LLM Quantization
Yuxuan Sun, Ruikang Liu, Haoli Bai, Han Bao, Kang Zhao, Yuening Li, Jiaxin Hu, Xianzhi Yu, Lu Hou, Chun Yuan, Xin Jiang, Wulong Liu, Jun Yao
•
Oct 12, 2024
•
15
2
VidPanos: カジュアルなパンニングビデオから生成されたパノラマビデオ
VidPanos: Generative Panoramic Videos from Casual Panning Videos
Jingwei Ma, Erika Lu, Roni Paiss, Shiran Zada, Aleksander Holynski, Tali Dekel, Brian Curless, Michael Rubinstein, Forrester Cole
•
Oct 17, 2024
•
13
2
LLMには政治的正確性がありますか?AIシステムにおける倫理的バイアスとジェイルブレイクの脆弱性を分析する
Do LLMs Have Political Correctness? Analyzing Ethical Biases and Jailbreak Vulnerabilities in AI Systems
Isack Lee, Haebin Seong
•
Oct 17, 2024
•
13
2
MLLM(Massive Language Models)は、中国語の画像の奥深い含意を理解できるか?
Can MLLMs Understand the Deep Implication Behind Chinese Images?
Chenhao Zhang, Xi Feng, Yuelin Bai, Xinrun Du, Jinchang Hou, Kaixin Deng, Guangzeng Han, Qinrui Li, Bingli Wang, Jiaheng Liu, Xingwei Qu, Yifei Zhang, Qixuan Zhao, Yiming Liang, Ziqiang Liu, Feiteng Fang, Min Yang, Wenhao Huang, Chenghua Lin, Ge Zhang, Shiwen Ni
•
Oct 17, 2024
•
11
2
前進する失敗:合成データと検索拡張を用いた音声認識のための生成誤り訂正の改善
Failing Forward: Improving Generative Error Correction for ASR with Synthetic Data and Retrieval Augmentation
Sreyan Ghosh, Mohammad Sadegh Rasooli, Michael Levit, Peidong Wang, Jian Xue, Dinesh Manocha, Jinyu Li
•
Oct 17, 2024
•
10
2
相互作用からの事後学習
Retrospective Learning from Interactions
Zizhao Chen, Mustafa Omer Gul, Yiwei Chen, Gloria Geng, Anne Wu, Yoav Artzi
•
Oct 17, 2024
•
9
2
記憶、検索、生成:無限のビジュアルコンセプトを理解する あなたのパーソナライズされたアシスタント
Remember, Retrieve and Generate: Understanding Infinite Visual Concepts as Your Personalized Assistant
Haoran Hao, Jiaming Han, Changsheng Li, Yu-Feng Li, Xiangyu Yue
•
Oct 17, 2024
•
9
2
MuVi: セマンティックアライメントとリズム同期を用いたビデオから音楽への生成
MuVi: Video-to-Music Generation with Semantic Alignment and Rhythmic Synchronization
Ruiqi Li, Siqi Zheng, Xize Cheng, Ziang Zhang, Shengpeng Ji, Zhou Zhao
•
Oct 16, 2024
•
9
2
MedMobile: 専門レベルの臨床能力を持つモバイルサイズの言語モデル
MedMobile: A mobile-sized language model with expert-level clinical capabilities
Krithik Vishwanath, Jaden Stryker, Anton Alaykin, Daniel Alexander Alber, Eric Karl Oermann
•
Oct 11, 2024
•
9
2
γ-MoD: 多モーダル大規模言語モデルのための深さ混合適応の探索
γ-MoD: Exploring Mixture-of-Depth Adaptation for Multimodal Large Language Models
Yaxin Luo, Gen Luo, Jiayi Ji, Yiyi Zhou, Xiaoshuai Sun, Zhiqiang Shen, Rongrong Ji
•
Oct 17, 2024
•
8
2
LoLDU: パラメータ効率のファインチューニングのための下三角-対角-上三角分解を用いた低ランク適応
LoLDU: Low-Rank Adaptation via Lower-Diag-Upper Decomposition for Parameter-Efficient Fine-Tuning
Yiming Shi, Jiwei Wei, Yujia Wu, Ran Ran, Chengwei Sun, Shiyuan He, Yang Yang
•
Oct 17, 2024
•
7
2
オープンマテリアルズ2024(OMat24) 無機材料データセットとモデル
Open Materials 2024 (OMat24) Inorganic Materials Dataset and Models
Luis Barroso-Luque, Muhammed Shuaibi, Xiang Fu, Brandon M. Wood, Misko Dzamba, Meng Gao, Ammar Rizvi, C. Lawrence Zitnick, Zachary W. Ulissi
•
Oct 16, 2024
•
7
1
長LRM:広範囲のガウススプラットのための長いシーケンス大再構築モデル
Long-LRM: Long-sequence Large Reconstruction Model for Wide-coverage Gaussian Splats
Chen Ziwen, Hao Tan, Kai Zhang, Sai Bi, Fujun Luan, Yicong Hong, Li Fuxin, Zexiang Xu
•
Oct 16, 2024
•
6
2
高品質データを鍵としてLLMから長い出力を解除するための最小チューニング
Minimum Tuning to Unlock Long Output from LLMs with High Quality Data as the Key
Yingda Chen, Xingjun Wang, Jintao Huang, Yunlin Mao, Daoze Zhang, Yuze Zhao
•
Oct 14, 2024
•
6
2
条件対照的整合を通じたガイダンス不要のARビジュアル生成に向けて
Toward Guidance-Free AR Visual Generation via Condition Contrastive Alignment
Huayu Chen, Hang Su, Peize Sun, Jun Zhu
•
Oct 12, 2024
•
5
2
AERO: 効率的なプライベート推論のためのSoftmax-Only LLMs
AERO: Softmax-Only LLMs for Efficient Private Inference
Nandan Kumar Jha, Brandon Reagen
•
Oct 16, 2024
•
4
2
TransAgent: 異種エージェントの協力によるビジョン言語基盤モデルの転移
TransAgent: Transfer Vision-Language Foundation Models with Heterogeneous Agent Collaboration
Yiwei Guo, Shaobin Zhuang, Kunchang Li, Yu Qiao, Yali Wang
•
Oct 16, 2024
•
4
2
SBI-RAG:スキーマベースの指導と検索拡張生成を通じた学生の数学ワード問題解決の向上
SBI-RAG: Enhancing Math Word Problem Solving for Students through Schema-Based Instruction and Retrieval-Augmented Generation
Prakhar Dixit, Tim Oates
•
Oct 17, 2024
•
3
2