-
Unified Reward Model for Multimodal Understanding and Generation
Paper • 2503.05236 • Published • 124 -
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
Paper • 2505.03318 • Published • 94 -
CodeGoat24/UnifiedReward-Think-qwen35-9b
9B • Updated • 24 -
CodeGoat24/UnifiedReward-Think-qwen35-27b
3.05M • Updated • 40
SII-Yibin Wang
CodeGoat24
AI & ML interests
I'm part of Shanghai Innovation Institute, focusing on Multimodal RL and Generation.
Recent Activity
liked a model 26 days ago
CodeGoat24/UnifiedReward-2.0-qwen35-4b liked a model about 1 month ago
CodeGoat24/UnifiedReward-Flex-qwen3vl-32b upvoted a paper about 2 months ago
Light-WAM: Efficient World Action Models with State-Fusion Action DecodingOrganizations
UnifiedReward Flex
We updated the model weights and enhanced the training data to mitigate the position bias issue!!
-
Unified Personalized Reward Model for Vision Generation
Paper • 2602.02380 • Published • 20 -
CodeGoat24/FLUX.2-klein-base-9B-UnifiedReward-Flex-lora
Text-to-Image • Updated • 132 • 30 -
CodeGoat24/Wan2.2-T2V-A14B-UnifiedReward-Flex-lora
Text-to-Video • Updated • 30 • 13 -
CodeGoat24/Wan2.1-T2V-14B-UnifiedReward-Flex-lora
Text-to-Video • Updated • 5 • 6
UnifiedReward 2.0 Qwen3.5 Models
-
Unified Reward Model for Multimodal Understanding and Generation
Paper • 2503.05236 • Published • 124 -
Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-Tuning
Paper • 2505.03318 • Published • 94 -
CodeGoat24/UnifiedReward-Think-qwen35-9b
9B • Updated • 24 -
CodeGoat24/UnifiedReward-Think-qwen35-27b
3.05M • Updated • 40
UnifiedReward Flex
We updated the model weights and enhanced the training data to mitigate the position bias issue!!
-
Unified Personalized Reward Model for Vision Generation
Paper • 2602.02380 • Published • 20 -
CodeGoat24/FLUX.2-klein-base-9B-UnifiedReward-Flex-lora
Text-to-Image • Updated • 132 • 30 -
CodeGoat24/Wan2.2-T2V-A14B-UnifiedReward-Flex-lora
Text-to-Video • Updated • 30 • 13 -
CodeGoat24/Wan2.1-T2V-14B-UnifiedReward-Flex-lora
Text-to-Video • Updated • 5 • 6
spaces 4
pinned
Sleeping
Agents
3
UniGenBench Leaderboard (Chinese Long)
🏅
UniGenBench: a unified T2I generation benchmark.
pinned
Running
Agents
3
UniGenBench Leaderboard (Chinese)
🏅
UniGenBench: a unified T2I generation benchmark.
pinned
Sleeping
Agents
7
UniGenBench Leaderboard (English)
🏅
UniGenBench: a unified T2I generation benchmark.
pinned
Running
Agents
3
UniGenBench Leaderboard (English Long)
🏅
UniGenBench: a unified T2I generation benchmark.
models 55
CodeGoat24/UnifiedReward-Flex-qwen35-9b
9B • Updated • 14
CodeGoat24/UnifiedReward-Flex-qwen3vl-4b
5B • Updated • 4
CodeGoat24/UnifiedReward-Flex-qwen3vl-8b
9B • Updated • 582
CodeGoat24/UnifiedReward-Flex-qwen3vl-32b
33B • Updated • 14 • 2
CodeGoat24/UnifiedReward-Flex-qwen35-27b
27B • Updated • 14
CodeGoat24/UnifiedReward-Flex-qwen35-4b
5B • Updated • 7 • 2
CodeGoat24/UnifiedReward-Think-qwen35-27b
3.05M • Updated • 40
CodeGoat24/UnifiedReward-Think-qwen35-4b
5B • Updated • 8 • 2
CodeGoat24/UnifiedReward-Think-qwen3vl-2b
2B • Updated • 390
CodeGoat24/UnifiedReward-Think-qwen3vl-4b
4B • Updated • 146
datasets 14
CodeGoat24/UniGenBench-Eval-Images
Preview • Updated • 207 • 5
CodeGoat24/UnifiedReward-Flex-SFT-90K
Viewer • Updated • 1.39M • 167 • 3
CodeGoat24/UniGenBench
Updated • 105 • 4
CodeGoat24/UnifiedReward-2.0-T2X-score-data
Viewer • Updated • 337k • 278
CodeGoat24/VIDEOGEN
Viewer • Updated • 50.9k • 31
CodeGoat24/ShareGPTVideo-DPO
Viewer • Updated • 101k • 61
CodeGoat24/VideoFeedback
Viewer • Updated • 73.2k • 53
CodeGoat24/VideoDPO
Viewer • Updated • 29k • 67
CodeGoat24/OIP
Viewer • Updated • 21.4k • 40
CodeGoat24/LLaVA-Critic-113k
Preview • Updated • 31