AI & ML interests

None defined yet.

Recent Activity

jkminderĀ  updated a model about 21 hours ago
dlab-spp/vanilla-3b-instruct
jkminderĀ  updated a model about 21 hours ago
dlab-spp/vanilla-3b-base
jkminderĀ  updated a model about 21 hours ago
dlab-spp/vanilla-1.7b-instruct
View all activity

Organization Card

SPP: annotate pretraining data with normative reflections, inject at different pretraining stages, evaluate alignment and safety

Alignment — and the assistant identity itself — is normally introduced only after pretraining, once behavioral priors are already set. SPP installs the desired persona from token zero instead: we define it through normative values in a constitution, generate first-person moral reflections grounded in that constitution, and insert them throughout the pretraining corpus behind an <assistant> token. Post-training then binds the chat assistant identity to the installed persona. Pretraining up to 3B on 500B tokens, SPP improves constitution following and jailbreak robustness while preserving capabilities — and when the data arrives matters: models trained with reflections from token zero prioritize values differently and take fewer risky actions in out-of-distribution moral dilemmas than models given the exact same data only at the end of pretraining, an advantage that grows with scale.

Collections

šŸ“¦ Pretraining Datasets — the reflection data, the corpus selection manifest, safety scores, and verification files.

šŸ¤– Models — 3B Ā· Models — 1.7B — all five recipes, base and instruct, at both scales.

šŸ’¬ Post-training Dataset — SP-SFT, the mixture that performs persona binding.

šŸ“Š Evals — ConstitutionEval and an audited AIRiskDilemmas.

From EPFL DLAB.