Ishant06/Qwen3.5-0.8B-Claude-4.6-Opus-Reasoning-Distilled Text Generation โข 0.8B โข Updated Mar 15 โข 21 โข 5
When Does Muon Help Agentic Reinforcement Learning? Paper โข 2607.16169 โข Published 6 days ago โข 13
PianoFlow: Music-Aware Streaming Piano Motion Generation with Bimanual Coordination Paper โข 2604.12856 โข Published Apr 15
When Does Muon Help Agentic Reinforcement Learning? Paper โข 2607.16169 โข Published 6 days ago โข 13
When Does Muon Help Agentic Reinforcement Learning? Paper โข 2607.16169 โข Published 6 days ago โข 13
Qwen2.5 Collection Qwen2.5 language models, including pretrained and instruction-tuned models of 7 sizes, including 0.5B, 1.5B, 3B, 7B, 14B, 32B, and 72B. โข 43 items โข Updated Mar 2 โข 732
bottlecapai/ThinkingCap-Qwen3.6-27B Image-Text-to-Text โข 27B โข Updated 12 days ago โข 12k โข โข 504
GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-Thinking-GGUF Text Generation โข 1B โข Updated 9 days ago โข 195k โข 297
view post Post 6012 DeepSeek-V4 can now run locally with Unsloth GGUFs! ๐ณRun lossless DeepSeek-V4-Flash on 168GB RAM or3-bit works on 110GB Mac, RAM, VRAM setups.Run via Unsloth Studio or llama.cpp.GGUF: unsloth/DeepSeek-V4-Flash-GGUFGuide: https://unsloth.ai/docs/models/deepseek-v4 See translation ๐ฅ 20 20 ๐ 5 5 ๐ 3 3 ๐ค 2 2 + Reply
view post Post 7665 Frontier models use distillation as a step of their post-training pipelines. In 2026 it has three jobs: compress a big model into a small one, merge RL experts into a single model, and let a model teach itself.I wrote up which frontier models use each one and how: https://huggingface.co/blog/sergiopaniego/distillation-2026It pairs with Class 2 of the Training an Agent series Ben and I are doing, where we teach these techniques hands-on with TRL! See translation 3 replies ยท ๐ 14 14 ๐ฅ 7 7 โค๏ธ 3 3 + Reply