Melvin Vivas
AI & ML interests
Recent Activity
Organizations
- Running on ZeroAgentsFeatured1.43k
Omni Video Factory
🏆1.43ktext to video, image to video, video extend
- Running on ZeroMCP3.03k
Wan2.2 14B Preview
🐌3.03kgenerate a video from an image with a text prompt
- Running on ZeroMCPFeatured3.41k
Wan2.2 14B Fast
🎥3.41kgenerate a video from an image with a text prompt
- Running on ZeroAgentsFeatured176
LTX 2.3 Sync
🕺176Portrait animation & lipsync with LTX 2.3
- Running1
Qwen-3-VL-8B OCR Receipts
🚀1structured data parser from receipt images
- RunningAgentsFeatured273
Qwen3 Omni Demo
⚡273Chat with AI using text, audio, images, or video
- Running on ZeroAgentsFeatured122
VLM Object Understanding
🦀122Explore object detection, visual grounding, keypoint Detecti
- SleepingAgents2
Dataset Card Drafter
😻2Create dataset descriptions and open PRs automatically
- Running on ZeroAgentsFeatured188
VibeVoice-Realtime-0.5B
🐨188Generate natural speech from text with selectable voices
-
microsoft/VibeVoice-1.5B
Text-to-Speech • 3B • Updated • 125k • 2.47k - RunningAgentsFeatured412
Qwen3 TTS Demo
🚀412Generate spoken audio from your text in many voices
-
mradermacher/Qwen3-1.7B-Multilingual-TTS-GGUF
2B • Updated • 1.07k • 12
- Running on ZeroAgents963
BRIA RMBG 2.0
🐢963remove background from any image
- Running on CPU UpgradeAgents2.46k
Omni Image Editor
🖼2.46kImage edit, text to image, image upscale, remove watermark
- Running on ZeroMCPFeatured2.67k
Qwen-Image-Edit-2511-LoRAs-Fast
🎃2.67kDemo of the Collection of Qwen Image Edit LoRAs
- Running on CPU UpgradeAgents1.03k
Open VLM Leaderboard
🌎1.03kVLMEvalKit Evaluation Results Collection
- Running on ZeroAgentsFeatured499
DeepSeek OCR 2 Demo
🚀499Try out DeepSeek-OCR-2 on your PDFs or images
- Running on ZeroMCP69
Multimodal OCR3
🌖69Chandra-OCR / Nanonets-OCR2 / olmOCR-2 / Dots.OCR
-
Qwen/Qwen3-VL-30B-A3B-Instruct
Image-Text-to-Text • 31B • Updated • 408k • • 593
-
openai/whisper-large-v3
Automatic Speech Recognition • 2B • Updated • 4.52M • • 6.19k -
openai/whisper-large-v3-turbo
Automatic Speech Recognition • 0.8B • Updated • 7.34M • • 3.27k - Running on ZeroMCPFeatured850
Whisper Large V3
🤫850Transcribe audio or YouTube videos to text
- Running on ZeroAgentsFeatured91
Kugel Audio
👀91Generate natural-sounding speech in European languages with voice cloning
- Running on ZeroAgents963
BRIA RMBG 2.0
🐢963remove background from any image
- Running on CPU UpgradeAgents2.46k
Omni Image Editor
🖼2.46kImage edit, text to image, image upscale, remove watermark
- Running on ZeroMCPFeatured2.67k
Qwen-Image-Edit-2511-LoRAs-Fast
🎃2.67kDemo of the Collection of Qwen Image Edit LoRAs
- Running on ZeroAgentsFeatured1.43k
Omni Video Factory
🏆1.43ktext to video, image to video, video extend
- Running on ZeroMCP3.03k
Wan2.2 14B Preview
🐌3.03kgenerate a video from an image with a text prompt
- Running on ZeroMCPFeatured3.41k
Wan2.2 14B Fast
🎥3.41kgenerate a video from an image with a text prompt
- Running on ZeroAgentsFeatured176
LTX 2.3 Sync
🕺176Portrait animation & lipsync with LTX 2.3
- Running1
Qwen-3-VL-8B OCR Receipts
🚀1structured data parser from receipt images
- RunningAgentsFeatured273
Qwen3 Omni Demo
⚡273Chat with AI using text, audio, images, or video
- Running on ZeroAgentsFeatured122
VLM Object Understanding
🦀122Explore object detection, visual grounding, keypoint Detecti
- SleepingAgents2
Dataset Card Drafter
😻2Create dataset descriptions and open PRs automatically
- Running on CPU UpgradeAgents1.03k
Open VLM Leaderboard
🌎1.03kVLMEvalKit Evaluation Results Collection
- Running on ZeroAgentsFeatured499
DeepSeek OCR 2 Demo
🚀499Try out DeepSeek-OCR-2 on your PDFs or images
- Running on ZeroMCP69
Multimodal OCR3
🌖69Chandra-OCR / Nanonets-OCR2 / olmOCR-2 / Dots.OCR
-
Qwen/Qwen3-VL-30B-A3B-Instruct
Image-Text-to-Text • 31B • Updated • 408k • • 593
- Running on ZeroAgentsFeatured188
VibeVoice-Realtime-0.5B
🐨188Generate natural speech from text with selectable voices
-
microsoft/VibeVoice-1.5B
Text-to-Speech • 3B • Updated • 125k • 2.47k - RunningAgentsFeatured412
Qwen3 TTS Demo
🚀412Generate spoken audio from your text in many voices
-
mradermacher/Qwen3-1.7B-Multilingual-TTS-GGUF
2B • Updated • 1.07k • 12
-
openai/whisper-large-v3
Automatic Speech Recognition • 2B • Updated • 4.52M • • 6.19k -
openai/whisper-large-v3-turbo
Automatic Speech Recognition • 0.8B • Updated • 7.34M • • 3.27k - Running on ZeroMCPFeatured850
Whisper Large V3
🤫850Transcribe audio or YouTube videos to text
- Running on ZeroAgentsFeatured91
Kugel Audio
👀91Generate natural-sounding speech in European languages with voice cloning