---
language:
- en
library_name: stable-audio-3
license: other
license_name: stable-audio-community
license_link: https://huggingface.co/stabilityai/SAME-L/blob/main/LICENSE.md
base_model: stabilityai/SAME-L
tags:
- music
- audio
- autoencoder
- lora
- stable-audio-3
---
# SAME-L Decoder LoRA (`squeakfix_v3`)
A LoRA on **SAME-L's decoder**. It changes how latents are rendered to audio, not
how latents are produced, so the latent space, the DiT and every existing DiT
LoRA stay bit-compatible and this stacks with them in any order.
It targets the high-frequency squeaks SAME's autoencoder adds to percussion when
audio goes round the encode/decode loop more than once. A continuation or a
transform re-encodes its input, so a chained session compounds them.
Powered by Stability AI. This is a Derivative Work of
[SAME-L](https://huggingface.co/stabilityai/SAME-L) under the [Stability AI
Community License](LICENSE.md); [NOTICE](NOTICE) records what it changes. No
SAME-L weights are redistributed here, only the adapter.
## Listen
| | stock decoder | `squeakfix_v3` |
|---|---|---|
| **jungle**, fresh generation | | |
| **jungle**, after 3 round trips | | |
| **drums**, fresh generation | | |
| **drums**, after 3 round trips | | |
| **neurofunk**, fresh generation | | |
| **neurofunk**, after 3 round trips | | |
15 s clips, same latent and seed, decoder adapter the only variable, peak-matched
to −1 dBFS, generated on `medium` with a DiT LoRA attached. These three prompts
were not picked for artifact content, and the fresh-generation rows are there to
show the progression rather than because much is audible in them.
### On a fresh generation
| | stock decoder | `squeakfix_v3` |
|---|---|---|
| **glitch hop**, worst of 16 screened | | |
Much subtler, and it took screening 16 candidates by tonality runaway to find a
pair where it carries. What changes here is air rather than squeak count: on this
pair tonality moves 0.08 dB, which is nothing, while 16-22 kHz comes back
+2.4 dB.
## How this came about
I do a lot of transformations and continuations using my guitar as input audio.
When I first noticed these squeaks in the outputs, I mentioned that my trick was
to stem separate in Ableton and then apply the erosion effect to the percussion
layer. Claude used that as a guide for refining the methodology. The technical
version is below.
## Usage
Requires the `target: "decoder"` support on the
[`decoder-lora` branch of betweentwomidnights/stable-audio-3](https://github.com/betweentwomidnights/stable-audio-3/tree/decoder-lora).
It is not upstream yet, so that branch is currently the only place a decoder LoRA
will load. The `target` field in the checkpoint config routes it; there is no
special loading path. Ship at strength **1.0**, where it was trained and auditioned.
```bash
python run_gradio.py --model medium-base --lora-ckpt-path squeakfix_v3.safetensors
# stacks with a DiT LoRA, either order
python run_gradio.py --model medium-base \
--lora-ckpt-path my_style.safetensors squeakfix_v3.safetensors
```
```python
from stable_audio_3 import AutoencoderModel
from stable_audio_3.models.lora.loader import load_and_apply_loras
from stable_audio_3.models.lora.model import set_lora_strength
ae = AutoencoderModel.from_pretrained("same-l", device="cuda").autoencoder
load_and_apply_loras(ae, ["squeakfix_v3.safetensors"], "autoencoder")
set_lora_strength(ae.decoder, 1.0) # separate call; 0.0 is the stock decode
audio = ae.decode(ae.encode(waveform))
```
## Numbers
Four base-model generations of 2 minutes on ordinary descriptive prompts,
screened from 12 candidates by tonality runaway. Each latent is decoded twice,
adapter off and on, and each chain re-encodes in place, so dN is N round trips of
that decoder config. The adapter is the only variable.
| depth | tonality p95 | tonal frames | 16-22 kHz | SI-SDR vs own d0 |
|---|---|---|---|---|
| d0, fresh generation | 9.80 → **9.10** | 87 → **62** | 3.16 → **4.34** | |
| d1, one continuation | 11.82 → **9.73** | 462 → **64** | 0.95 → **3.76** | 16.9 → **21.0** |
| d2 | 14.79 → **10.61** | 2111 → **133** | −1.36 → **3.32** | 12.4 → **16.2** |
| d3 | 17.85 → **11.74** | 4279 → **453** | −3.36 → **2.96** | 10.1 → **13.6** |
The adapter is ahead at every depth, including a fresh generation. Stock tonality
climbs +8.05 dB over three round trips against the adapter's +2.64, and by d3
stock carries 4279 tonal frames against 453.
It is not doing this by sanding the top end. After three round trips the
adapter's 16-22 kHz sits 0.2 dB from a **fresh** stock render, at unchanged RMS,
while stock has lost 6.5 dB of it. Stock's 12-16 kHz meanwhile *rises* with
depth: the invented energy is replacing the air.
Prompt distribution matters more here than anything else. The same ladder run on
terse genre-tag prompts, which are out of distribution for the base model, starts
at **1649** tonal frames instead of 87. With a DiT LoRA attached, which makes
those tags in-distribution again, it starts at 780. The artifact is worst where
the prompt sits furthest from what the model expects, and re-encoding compounds
whatever it starts with.
A DiT LoRA does not change the shape: on that run stock climbs +5.79 dB against
the adapter's +0.73, and the adapter ends d3 holding more 12-16 kHz and 16-22 kHz
than a fresh stock render. All runs are fp16 at 120 s.
## Limits
- **Decoder LoRAs trained this way stamp a comb into the output.** The un-patch
is a bare reshape at `patch_size 256`, so output channel and
time-position-within-patch are the same axis, and any channel-wise bias the
adapter learns repeats in every patch. That is a harmonic comb at
`sr/256 = 172.27 Hz`, which reads as a wrong note rather than as noise. It is
inaudible in a mix and obvious once a stem separator strips the masking
content, which is how it was found. `--lambda_patch 30` is what prevents it: a
`v1` control trained without it sits **+10.9 dB** of comb excess over stock
while v3 sits **0.8 dB below** stock, and both hold through three further round
trips. On real audio rather than model output v3 measures +1.5 to +4.0 dB, so
measure on generated audio.
- **SAME-L only.** SAME-S shares the patch grid but its artifact sits 1-8 kHz and
wants a different recipe. Loading this there will not error, it will sound
wrong.
- **The gain scales with re-encode depth**, so it is small for one-shot
generation and compounds across chained continuations.
- **The squeaks are worst on out-of-distribution prompts.** Genre-tag prompts on
the base model start with 19x more tonal frames than descriptive ones. A DiT
LoRA makes such tags in-distribution again and lands in between. This adapter
reduces the artifact; it does not make prompt choice irrelevant.
- **One checkpoint, one dataset**, step 2000, chosen on re-encode behaviour and
listening. Your material may want a different one.
- **Neither half of the autoencoder is bit-deterministic** (encoder 1.5 %
relative, decoder 7e-3 max abs), so read any A/B as a multiple of those floors.
## Training
Self-supervised, audio only, no captions or paired data. Rank 16 / alpha 16,
5.63M trainable params (~1.30 % of the decoder), fp32 training and fp16 storage,
single GB10 (DGX Spark) at ~2.9 s/step on 10 s crops, under 10 GB.
```bash
python scripts/train_decoder_lora.py \
--model same-l --rank 16 --steps 8000 --lr 1e-4 \
--w_real 0.4 --w_drift 0.3 --w_dit 0.3 \
--lambda_cycle 10.0 --lambda_tonal 0.3 --lambda_patch 30 \
--patch_size 256 --crop_seconds 10 --seed 0 \
--dit_latent_dir out/dit_latents \
--data_dir /path/to/audio --eval_audio /path/to/holdout.wav \
--out_dir out/declora
```
Latents come from three buckets: real audio, drifted audio (`E(D(E(x)))` at depth
≥ 1, targeted against the original so the decoder is asked to undo accumulated
drift), and latents sampled from the base DiT. No discriminator, since
adversarial training makes a decoder invent *more* detail.
Full method, losses, ablations and traps: [Decoder LoRA
workflow](https://github.com/betweentwomidnights/stable-audio-3/blob/decoder-lora/docs/workflows/decoder-lora.md).
Provenance is embedded in the file's `lora_config` metadata. `v1` (identical
recipe, `--lambda_patch 0`) is an internal control, not published.
Trained and used by [the collabage patch](https://huggingface.co/thepatch), in
production in a live `stable-audio-3` backend since 2026-08-14. The autoencoder
this adapts: [SAME](https://arxiv.org/abs/2605.18613).