Who cares about liability, Anthrowhatever put out Opus which apparently could compromise government infrastructure but God forbid an AI gets a little racist or tells you how to poorly cook meth
Timeline
Post
Remote status
Context
8
I just want an uncensored AI model, why is nobody making them? Being the only person to have a uncensored LLM would make so much money
Who cares about liability, Anthrowhatever put out Opus which apparently could compromise government infrastructure but God forbid an AI gets a little racist or tells you how to poorly cook meth
Who cares about liability, Anthrowhatever put out Opus which apparently could compromise government infrastructure but God forbid an AI gets a little racist or tells you how to poorly cook meth
@WandererUber@poa.st yeah the one im using still shows thinking and refuses to criticize Israel so I might need a better one lmao
the issue with "uncensored" is that they can take out the refusals but they can't take out liberal doctrine from all the training data being mainstream.
I tried it with Qwen3.6 right now. The normal version refuses to talk about it, the heretic does but still frames it liberally, and the one with the system prompt set to be a National Socialist does take a Right-Wing perspective but it goes into "mirroring" quite fast. This is probably a combination of the model not being that big and also the lack of training data, like I said.
Matty put a lot of work into actually training Anathema on the facts. That's a totally different challenge than simple uncensoring.
I tried it with Qwen3.6 right now. The normal version refuses to talk about it, the heretic does but still frames it liberally, and the one with the system prompt set to be a National Socialist does take a Right-Wing perspective but it goes into "mirroring" quite fast. This is probably a combination of the model not being that big and also the lack of training data, like I said.
Matty put a lot of work into actually training Anathema on the facts. That's a totally different challenge than simple uncensoring.
Abliteration !== RLHF removal. All abliteration does is remove the refusal mechanism. It doesn't change the intrinsic training - that requires SFT.
Yeah exactly
I was trying to say it in English, doc.
To be more precise with this example, while Qwen Heretic does blame "Zionist-controlled networks" for conflict around the Middle East, it doesn't even mention "the Jews" once in it's answer, nor does it have any concept of their ethnic hatred, building nukes etc.
You, of course, know how this works, matty old bean
I was trying to say it in English, doc.
To be more precise with this example, while Qwen Heretic does blame "Zionist-controlled networks" for conflict around the Middle East, it doesn't even mention "the Jews" once in it's answer, nor does it have any concept of their ethnic hatred, building nukes etc.
You, of course, know how this works, matty old bean
And even then SFT/DPO or whatever LoRA you use isn't going to make much of a dent. I'd recommend DPO over SFT unless you're trying to teach the model new behaviors. DPO shifts adjacent weights but you need a ton of training data to do much at all, and then you have gaps where it doesn't work. The only way to actually get this to work is to pre-train a model but I sincerely doubt any of these people are going to be willing to chip in for a couple thousand A100s.
sorry, I just wanted to help I didn't mean to come across as arrogant.
it doesn't at all
Have you written something long-form about training Lexi before? I would love to learn about the process.
Have you written something long-form about training Lexi before? I would love to learn about the process.
I have not - but may be worth doing in the future once I'm particularly confident about my skills. I've spent way too much money learning and don't want others to repeat my mistakes, so I'd rather have a concrete understanding first.
Replies
0
No replies yet.