arxiv:2602.02493
ZehongMa
zehongma
AI & ML interests
MLLMs, Image/Video Generation, Multi-modal Representation Learning
Recent Activity
upvoted a paper about 4 hours ago
MiniWorld: Democratizing the Training of Video World Models from Scratch authored a paper 17 days ago
PixelGen: Pixel Diffusion Beats Latent Diffusion with Perceptual Loss upvoted a paper about 2 months ago
UniDDT: Unifying Multimodal Understanding and Generation with Decoupled Diffusion TransformerOrganizations
None yet