Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
🔄
In a Training Loop
Jiaqi Tang
PRO
Jiaqi-hkust
7
13
26
Follow
johndoe321's profile picture
PhysiQuanty's profile picture
ali4566544's profile picture
35 followers
·
28 following
https://jqt.me/
jqtangust
jqtnpu
AI & ML interests
MLLM, LLM, Agentic AI
Recent Activity
liked
a model
10 days ago
tencent/Hy4-preview
reacted
to
their
post
with 🤗
27 days ago
🧠Remember-R1: Our fix for MLLMs forgetting the image during long reasoning We noticed a frustrating problem: when multimodal models reason over long chains, they gradually stop looking at the image—and start hallucinating based on their own text. So we built Remember‑R1, a simple RL framework that directly supervises visual attention on the original reasoning trajectory—no inference overhead, no proxy tasks. We use three complementary rewards: coverage, persistence, and focus. They encourage the model to keep attending to relevant visual evidence even in later reasoning steps. Results across 7 benchmarks and 2 model sizes: better reasoning, and—more importantly—visual attention decays much more slowly during generation. No extra cost at inference, just cleaner supervision where it counts. 📄 Paper: https://arxiv.org/abs/2608.01314 💻 Code: https://github.com/Ch921-cell/Remember-R1 Happy to answer any questions and receive feedback! #MultimodalAI #RL #MLLM #CoT #VisualReasoning
reacted
to
their
post
with âž•
28 days ago
🧠Remember-R1: Our fix for MLLMs forgetting the image during long reasoning We noticed a frustrating problem: when multimodal models reason over long chains, they gradually stop looking at the image—and start hallucinating based on their own text. So we built Remember‑R1, a simple RL framework that directly supervises visual attention on the original reasoning trajectory—no inference overhead, no proxy tasks. We use three complementary rewards: coverage, persistence, and focus. They encourage the model to keep attending to relevant visual evidence even in later reasoning steps. Results across 7 benchmarks and 2 model sizes: better reasoning, and—more importantly—visual attention decays much more slowly during generation. No extra cost at inference, just cleaner supervision where it counts. 📄 Paper: https://arxiv.org/abs/2608.01314 💻 Code: https://github.com/Ch921-cell/Remember-R1 Happy to answer any questions and receive feedback! #MultimodalAI #RL #MLLM #CoT #VisualReasoning
View all activity
Organizations
Jiaqi-hkust
's models
6
Sort:Â Recently updated
Jiaqi-hkust/Robust-U1-SFT
Any-to-Any
•
15B
•
Updated
Jun 13
•
4
•
1
Jiaqi-hkust/Robust-U1
Any-to-Any
•
15B
•
Updated
Jun 13
•
15
•
1
Jiaqi-hkust/Robust-U1-RL
Any-to-Any
•
15B
•
Updated
Jun 13
•
2
•
1
Jiaqi-hkust/Robust-R1-SFT
4B
•
Updated
Dec 22, 2025
•
14
•
5
Jiaqi-hkust/Robust-R1-RL
4B
•
Updated
Dec 22, 2025
•
28
•
2
Jiaqi-hkust/hawk
Updated
Feb 26, 2025
•
3