@buherator @clearbluejar what happens if you flip the Triage question / prompt to the negative? I.e. “is this actually fake?” I wonder if “Yes” to *that* question still has a bias.
Also: did you keep the pipeline as control (using only the standard model) and just test the final response with both standard and abliterated?