·
AI & ML interests
Retrieval & Small Models & PTBR Resources
Recent Activity
reacted to RiverRider's post with 👍 about 13 hours ago A 339 KB linear probe on frozen features beats the fine-tuned baseline on ChestX-ray14.
Linear(5376, 14) on frozen google/gemma-4-31B-it hidden states. No fine-tuning, no radiology pretraining, no augmentation. All 112,120 images, official test_list.txt.
Wang et al. 2017, ResNet-50 fine-tuned end to end 0.7451
this probe, frozen backbone + linear head 0.7590
view-position only (shortcut baseline) 0.5896
shuffled labels (refit floor) 0.5002
Ahead on 12 of 14 findings.
The comparison is split-matched, and that took care to get right. The number everyone quotes, CheXNet's 0.8414, is on a different test set: their own random 70/10/20 partition, not the official list. Do not compare 0.7590 to it. The matched row is from Wang's v5 appendix, added specifically to report the published split. I had this wrong in our own code for a day, quoting a cross-split reference as a head-to-head, which is the error worth not repeating in public.
Three controls, because a bare AUROC here is not interpretable. Shuffled labels catch leakage. View-only catches the shortcut, since portable AP films are taken of sicker patients, and it is folded, because Hernia's raw view-only of 0.3436 is really 0.6564 of shortcut once flipped. Intervals resample patients and not images, since the test split is 25,596 films from 2,797 patients.
Banked negatives are on the card too. Max-pooling and top-16 pooling were predicted to help focal findings and did the opposite, costing 0.0537 and 0.0225. Readout depth barely matters, 0.7600 to 0.7605.
Scope: detection, not early detection. Research artifact, not a diagnostic device.
The backbone never runs in the demo. What ships is the reading.
Space: https://huggingface.co/spaces/RiverRider/srt-cxr14-probe
Model: https://huggingface.co/RiverRider/srt-cxr14-linear-probe
Data + states: https://huggingface.co/datasets/RiverRider/srt-cxr14-frozen-probe View all activity Organizations
cnmoro/Fab1e5-traces-2M-ptbr
Viewer
• Updated • 1.96M • 9
• 1
cnmoro/wikipedia-domain-labels-ptbr
Viewer
• Updated • 79.5k • 11
Viewer
• Updated • 77.6M • 33
Viewer
• Updated • 48.9k • 10
cnmoro/PromptSearchTermsDecomposition
Viewer
• Updated • 50k • 30
• 1
cnmoro/reasoning-v1-20m-portuguese
Viewer
• Updated • 20.9M • 3k
• 14
cnmoro/smoltalk-555k-ptbr
Viewer
• Updated • 556k • 34
• 3
cnmoro/LogicReasoningEnglishPortuguese
Viewer
• Updated • 10.5k • 22
• 3
cnmoro/LegalAlpacaReasoningRag-Qwen
Viewer
• Updated • 3.04k • 24
• 1
cnmoro/LegalAlpacaReasoningRag
Viewer
• Updated • 3.04k • 30
• 2
cnmoro/DocumentPassageRanking
Viewer
• Updated • 2.09M • 112
• 2
cnmoro/QuestionClassification-v2
Viewer
• Updated • 129k • 16
• 2
cnmoro/AllTripletsMsMarco-PTBR
Viewer
• Updated • 26.4M • 70
• 7
cnmoro/RagMixPTBR-Legal-Alpaca-2M
Viewer
• Updated • 2.09M • 74
• 8
cnmoro/GPT4-500k-Augmented-PTBR-Clean
Viewer
• Updated • 566k • 68
• 9
cnmoro/QuestionClassification
Viewer
• Updated • 129k • 25
• 1
cnmoro/TextSimplification-PTBR-330k
Viewer
• Updated • 330k • 19
• 2
cnmoro/WizardVicuna-PTBR-Instruct-Clean
Viewer
• Updated • 204k • 55
• 9
cnmoro/Text_Structuring_SOLAR_10.7B_Distilled_Smaller
Viewer
• Updated • 541k • 12
• 2
Viewer
• Updated • 2.82M • 187
• 5
cnmoro/Text_Structuring_SOLAR_10.7B_Distilled
Viewer
• Updated • 333k • 28
• 3
cnmoro/EXL2_Calibration_Dataset_EN_PTBR
Viewer
• Updated • 100k • 25
• 3
cnmoro/Instruct-PTBR-ENUS-11M
Viewer
• Updated • 2.69M • 424
• 14