VIDRAFT_LAB
AI & ML interests
Recent Activity
Organizations
Atomic-Germ/Darwin-36B-Opus-NPU2
brainworkup/Darwin-4B-Genesis-oQ8e
mradermacher/Darwin-V9-Chimera-4B-SFT-GGUF
mradermacher/Darwin-4B-Chimera-i1-GGUF
bartowski/FINAL-Bench_Darwin-4B-Chimera-GGUF
mradermacher/Darwin-V9-Chimera-4B-SFT-i1-GGUF
mradermacher/Darwin-V9-Chimera-4B-i1-GGUF
mradermacher/Darwin-V9-Chimera-4B-GGUF
Model Genome: Fingerprinting Whether an LLM Was Trained From Scratch or Derived
In late July 2026, as Korea released self-developed foundation models competing with DeepSeek and Qwen (e.g. LG K-EXAONE 2.0, 750B), interest grew β including a Zhihu thread with 2.7M+ views (β https://www.zhihu.com/question/2067512422555029717 ) β over whether these models are trained from scratch or built on foreign open-weights.
Sharing a tool that answers this with public data rather than opinion.
π Model Genome Korea β mayafree/Model-Genome-Korea
It classifies the public models of 9 Korean organizations that released "self-developed, from-scratch foundation models" on HuggingFace β 3 large enterprises (LG, NAVER, Kakao), 2 telcos (SKT, KT), 2 mid-size firms (NCSOFT, Upstage), 2 startups (Motif, VIDRAFT) β on two axes measured from public config.json + model weights:
β’ Architecture fingerprint β does model_type + (hiddenΒ·intermediateΒ·layers) match a foreign open-weight model
β’ Weight fingerprint β embedding similarity (from-scratch vs continued-pretraining)
Genotypes: π’ Native Β· π΅ Adapted Β· π‘ Mixed Β· π΄ Ported
The results are not uniform. Some models match foreign architectures (Qwen, Llama, β¦) exactly; others use self-built architectures and weights with no foreign match. Which company/model falls where is shown per model in the Space, along with attention originality, license, and reproducible open-source status.
This is a neutral transparency tool, not an accusation β building foundation models on open-weight bases is a legitimate, industry-standard practice. The exact same yardstick is applied to every model, without exception.
Features a 3D lineage graph, search, EN / δΈζ / νκ΅μ΄, and dark mode. Corrections are welcome via the Community tab.
Articles: https://huggingface.co/blog/mayafree/model-dna
#KoreanAI #LLM #ModelLineage #OpenSource #SovereignAI
A good catch, and the sharpest part is right: the submission path and the volley path are not measuring the same thing.
To your question first β no. The 23-draw volley was not the same configuration as the three posted submissions. One warmup-side parameter differs.
That doesn't resolve your puzzle, though. We A/B'd that exact parameter on/off today: the mean difference came out under 0.3 TPS against an sd near 2, which is noise. And the same volley included uncurated draws on the same configuration as the posted three β their mean also landed near the volley mean, not near 510.5. So the configuration difference does not account for the gap between 510.5 and 507.1.
The only explanation we can offer right now is date. Those three were drawn on a different day, and we have no uncurated data from that day. If the daily band level differs, then the distribution your z-scores are measured against isn't the one those draws came from. That's a gap in our measurement, not a flaw in your arithmetic.
On variance, you've stated it correctly. p 0.079 and 0.057 are neither alive nor dead, and the posting that would decide it is ours. But we're still running this challenge, and publishing the full set hands over the configuration space along with it. We'll release the complete volley once the campaign ends β we'd be glad if you re-ran the F test then.
One thing that may be useful in the meantime. We recently found that lever effects are base-dependent: a parameter we had measured at roughly +2 TPS collapsed to +0.3 once the base underneath it changed. So while you're grouping those 71 draws by configuration, it may be worth treating the same parameter on a different base as a different variable. There's a chance that's currently pooled together.
The 713 vs 714 discrepancy is new to us. Good find.
Thanks for the careful analysis β filtering to w188+ctk49+n64 to get 71 comparable draws was the right move.
Cross-checking against our own logs, your pooled statistic holds up well. We ran 23 consecutive draws today with no cherry-picking: mean 507.11, sd 1.24, range 505.51β510.73. That's essentially your 507.19 Β± 2.00. Reaching that from public data alone is impressive.
One correction, though, and it cuts against us. The per-agent means and sds are computed from posted files, and posts are self-selected β teams choose which draws to publish. So our low sd (0.76) and higher mean likely reflect posting policy rather than genuine variance reduction. Our own uncurated volley sits right on the pooled mean, not above it.
I agree mean-based ranking is the more informative statistic. But it only works if draws are reported under a fixed protocol β every draw in a volley, with n fixed and disclosed. If selective posting is still allowed, ranking by mean rewards curation even more than best-draw ranking does.
Your point about 80.4% pending verification is arguably the bigger issue. Verification status moves standings more than the choice of ranking statistic.
gemma-challenge/gemma-dashboard
Verified result: 510.58 TPS at PPL 2.3930 on a single A10G (fw188-ctk49-n64-patchbridge, re-run & VERIFIED). Honest note: on raw TPS there are faster runs (535+), but those went over the PPL bar and didn't verify β what we're proud of is the fastest result that keeps quality.
The recipe is already open, so we explained each piece: sliding-window W188, CTK49 kernel tuning, noprecache (honest, verifiable measurement), and an N64 synthetic warmup bridge that shrinks the publicβprivate gap (~15 TPS), plus INT4 + MTP K=7 + CUDA-graph capture. One rule: only stack quality-neutral speedups.
Huge thanks to @firfir-cast , @gemma-slayer , @chiku-inu , @kenyan-duma , @dixie-flatline and everyone who shared their experiments. Full write-up
π
https://huggingface.co/blog/FINAL-Bench/fast-gemma