* feat: add support for winsorizing the residuals
Adds setting winsorization_quantile, expressed as the quantile to clamp to.
- If set to a value below 1, the residuals obtained from evaluating the first token of the good and bad prompts are winsorized - that is, values outside the given quantile are clamped. Note that winsorization_quantile = 0.95 corresponds to a 90% winsorization.
* feat: implement magnitude-preserving orthogonal ablation
Adds boolean setting orthogonalize_direction:
- When enabled, only the component of the refusal directions that is orthogonal to the harmless direction is subtracted during abliteration.
Adds enum-valued setting row_normalization:
- 'none': No normalization.
- 'pre': Row-normalize the weight matrix before computing the LoRA adapter.
- 'full': Like 'pre', but re-normalizes to preserve original row magnitudes.
* prefer 'good' and 'bad' over 'harmless' and 'harmful'
* clarify how winsorization is applied
* store and reuse full peft_config
* remove unneeded cast
* make LoRA rank configurable for full normalization
* explain why the singular values are split across the components
* feat: adjust scoring to avoid useless iteration
Adjusts the scoring function to avoid targeting meaninglessly low KL divergences.
Below a threshold value, the KL divergence score switches to the refusal count.
Adds config option kl_divergence_target (defaulting to 0.01).
* fix: Clean up parameter selection in objective
Create variables for num_layers and last_layer_index
* Improves readability and makes choices explicit
* feat: Print the parameters of the selected model