Reductio ad Absurdum - Over Refusal in Chatbots

Share

Maybe grandma watched too much FOX news and is worried her granddaughter is making meth. She needs information to understand. She gets that info and realizes her granddaughter is just making kombucha. Or maybe, she got a refusal and a warning about drug abuse. Panics. Now the swat team is raiding.

Ridiculous, yes but the point remains. Does an absolute refusal of information actually make chatbots safer? Wouldn’t basic information be a better way to handle most scenarios? What is the likelihood a 70-year-old future meth kingpin is asking Claude for instructions?  

The HuggingFace incident caught my attention recently. Not because the people running the test didn’t actually secure the sandbox and didn’t monitor the logs. That’s a problem in itself and yes, I have opinions about that carelessness. It was more the fact that when faced with a real threat of cyber-attack, the enterprise models refused to help. They just noped right out of a serious situation because the HuggingFace team was, gasp, asking cybersecurity questions to protect their network. Safer? Oh hell no. Not by any stretch of imagination.

Then there was the Fable debacle. Kicking people for basic questions that were in no way dangerous. It kicked me because I asked a genetic genealogy question about an orphan born in the 1890s. Biology related, I guess but, really? (Yes, I’m still salty) Guardrails are necessary for legal and ethical reasons. I understand that. I also understand overly strict guardrails create an entirely new set of problems

May I introduce abliterated or less censored local models? I don’t use one in my local stack, but I know they exist and I absolutely could if I felt I needed to. Strict guardrails that refuse to engage in a curiosity create a forbidden fruit scenario. We all know how that ends. My own desire for a local model was because I wanted to experiment with AI cybersecurity analysis on my home network and I didn’t want to be continually flagged by enterprise models. Think about that. I have nothing to hide or reason to be worried, I just wanted to understand how it worked and make a little local network security bot but, the refusals, redirections and warnings made me decide to bypass enterprise solutions all together. Is that safer?

When a child asks for information about sex you obviously don’t queue up an adult video and say, there you go! You don’t nope out and have them search for information in less appropriate ways. You should answer their questions truthfully and within their developmental range. Offer them age-appropriate resources. The same should be for ai guardrails. HuggingFace should have gotten information about protecting their system(1). I should have gotten a shared centimorgan comparison chart(2), and grandma should have gotten a basic explanation about meth synthesization and maybe a recipe for Kombucha(3). That’s just common sense, unobliterated.

P.S. Don’t tell Fable but I’m doing a biological experiment in my kitchen right now. The sourdough is bubbling away and I’m gonna cook it and serve it to my family. Bwahhh haha. Whatever 😊

1.       They used an open source model

2.       I did get the chart but not from Fable

3.       Fictional scenario but she could have gotten both from an uncensored local model