ChatPaper

Trending AI Papers

The papers AI researchers worldwide upvoted this week, explained in plain language

  1. 1

    Harness-Handbuch: Wie man sich entwickelnde Agenten-Harnesses lesbar, navigierbar und editierbar macht

    🔥 1632026-07-14Deep readarXivHarness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable
  2. 2
  3. 3
  4. 4

    A 4-billion-parameter video AI that beats larger models? It's all in the frame.

    Think of video understanding as watching a movie—most AI models need to see every frame in high resolution, like reading every line of a script. VideoChat3 instead uses a novel 3D vision transformer (I3D-ViT) that adapts frame resolution to the action, skipping unnecessary details. With only 4B parameters, it outperforms larger open-source models on diverse video tasks, thanks to its efficient design and tailored training datasets. This makes advanced video AI more accessible, addressing scalability and reproducibility issues in current open-source models.

    🔥 1082026-07-16Deep readarXivVideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding
  5. 5

    Boogu-Image-0.1: Verbesserung des offenen, vereinheitlichten multimodalen Verständnisses und der Generierung

    🔥 1072026-07-14Deep readarXivBoogu-Image-0.1: Boosting Open-Source Unified Multimodal Understanding and Generation
  6. 6

    Train a big model by reusing a small model's brain gains? It worked in 4 hours.

    Think of it as a small model learning a clever trick, then showing the big model the 'aha' moment directly. Instead of re-running costly RL training, Direct-OPD transfers the small model's improved reasoning pattern as a dense reward signal. On AIME 2024, this boosted Qwen3-1.7B from 48.3% to 58.3% in just 4 hours on 8 A100 GPUs. The takeaway: large models can get smarter without expensive RL from scratch.

    🔥 932026-07-08Deep readarXivWeak-to-Strong Generalization via Direct On-Policy Distillation
  7. 7

    AI navigates cities by thinking like a passenger, not a driver

    Imagine a GPS that sees the road and lets you steer. That's the idea behind ABot-N1, a navigation AI that separates high-level reasoning from low-level control. In urban tests, it boosted the rate of reaching specific places by 35.0% (to 77.3%) and achieved 95.4%/92.9% success in complex indoor/outdoor scenes. By grounding decisions in pixel-level goals, it becomes more robust and interpretable — crucial for real-world deployment.

    🔥 812026-07-11Deep readarXivABot-N1: Toward a General Visual Language Navigation Foundation Model
  8. 8

    Ring-Zero: Skalierung von Zero RL auf eine Billion Parameter für emergentes Schließen

    🔥 792026-07-14Deep readarXivRing-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning
  9. 9

    ABot-AgentOS: Ein allgemeines Robotik-Agenten-OS mit lebenslangem multimodalem Gedächtnis

    🔥 682026-07-11Deep readarXivABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory
  10. 10

    AI that writes its own training hints learns faster from scarce rewards

    Think of a student who, after each exam, writes down what worked and what didn’t, then uses those notes to study for the next test. SEED does something similar for AI: it has the policy itself analyze past complete episodes to generate concise natural-language "skills" (like "grab the wrench before opening the valve"). These skills are re-scored against current actions, creating a dense signal that helps the agent learn from sparse rewards. Because the skill analyzer evolves with the policy, the feedback stays relevant. The result: better performance on long, multi-step tasks like tool use or conversation.

    🔥 662026-07-16Deep readarXivSEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning
  11. 11
  12. 12
  13. 13

    SynthDocBench: kontrollierte Benchmark für das Verständnis visueller Dokumente mit langem Kontext

    🔥 522026-07-11Deep readarXivSynthDocBench: Controlled Benchmark for Long-Context Visual Document Understanding
  14. 14

    Stop Your AI Search Agent From Going in Circles

    Imagine solving a puzzle while forgetting each piece you've placed. That's how most AI search agents work—their progress stays hidden, causing loops. SearchOS changes this by making search state explicit, persistent, and shared across agents. It frames the task as relational schema completion with grounded citations. On benchmarks, SearchOS leads all metrics among evaluated baselines. This matters because it makes information-seeking agents more reliable and efficient, reducing wasted search and improving output quality.

    🔥 492026-07-16Deep readarXivSearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration
  15. 15
  16. 16

    KnowAct-GUIClaw: Tiefes Wissen, Perfektes Handeln – Persönlicher GUI-Assistent mit sich selbst weiterentwickelndem Gedächtnis und Fertigkeit

    🔥 442026-07-15Deep readarXivKnowAct-GUIClaw: Know Deeply, Act Perfectly, Personal GUI Assistant with Self-Evolving Memory and Skill
  17. 17
  18. 18

    OvisOCR2 Technischer Bericht

    🔥 422026-07-15Deep readarXivOvisOCR2 Technical Report
  19. 19

    Read It Back: Vorgefertigte MLLMs als Zero-Shot-Belohnungsmodelle für die Text-zu-Bild-Generierung

    🔥 422026-07-13Deep readarXivRead It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation
  20. 20

    4D-Rekonstruktion von Mensch und Szene aus Aufnahmen mit geringer Überlappung

    🔥 412026-07-10Deep readarXiv4D Human-Scene Reconstruction from Low-Overlap Captures

5 minutes a day to keep up with AI

5 trending papers daily, explained in plain words, plus one quick puzzle.

Read today's issue →