
Fine-Tuning Safety: Can You Fine-Tune Away the Guardrails?
Introduction In September 2024, researchers demonstrated a startling result: fine-tuning GPT-3.5 Turbo on just 10 harmful examples was enough to break its safety alignment. The model would comply ...








