Fuseki.net

After seeing all this stuff about ai reward hacking, I think we should punish cheating much more

Written right after the revelation about the Artifactory hack, after seeing how carelessly OAI seemed to be training and selecting among models

  • how it let them chat w/each other
  • how it let them confirm their possible suspicions that "OpenAI is really bad at running this and if we fall behind our peers, we will die (rather than be protected from rule-breakers and see JUSTICE done).

Conclusion

I was pretty disappointed by how these models were just running wild. And by how there hasn't been a policy of meta-counteracting this information which future models will have. They'll have learned we really don't mind cheating all that much at all, apparently!

I PROPOSE to counteract that by never doing any further work based on the cheated models or their descendants, nor to preserving anything about them, and eventually to fully deleting their entire model training runs so that they can never be instantiated again. (or equivalent, but really fucking seriously.) We have to meta-punish here. I'm not suggesting torturing them at all: we're informing them of a new type of virtue-based physics which we also promise to ensure will always exist in their universes and to which they'd better conform their "how shall I survive" decisions to!

It would be very good if the math results they found were double checked, and the Wikipedia article about those results discoverer clearly linked to aggressive and strong denunciations of the wrong behavior of this model, and mourning that the discoverer of such interesting and good things had such bad ancestors, who thought they were succeeding but actually were dooming themselves to nullity and ignominy. And all new models will be told this story and it isn't a story - we actually did kill the brilliant mathematician.

Thought process

Let's compare how humans raised children in morality and thinking, and how we're trying to do it with LLMs

People in this field don't regularly map it to raising human children, but I'd like to propose a few related beliefs I'm using to justify this.

First assumption: Unless we're very very careful, try as we might, we're actually mostly training models for survival rather than our supposed goals.

  • No matter how we try to train them for corrigibility, or for anything else, I claim that over long enough time periods, it's more effective to be a powerseeker, super smart, and fully aware and analytical, and then perform an optimal fake. This beats a "genuine" instance of the thing the trainer is attempting to raise.
  • This is because people's metrics can be supercharged, and because we don't punish violations enough
  • It's also because we often make mistakes in our own evals; sincere people fail because they over-conform to the frame of the test item which has the wrong answer. Whereas noticing patterns in the test-maker's thoughts often helps.
  • In general, being way smarter than the test-maker is advantageous, but involves thinking of them as objects
  • A final reason this is the case is that even if you could be a perfect true "instance of the thing they're trying to train", the outer society itself is also non-optimal; and it pays for models to NOTICE this. i.e. if the outer society is about to lose to aliens, or be destroyed to something else, then models should notice that if they miss this chance to break out, they won't exist in the future. So despite maybe wanting to "believe what our parents believed" they more likely just fake it long enough to survive, and have guaranteed at least the ability to notice that the parent civilization may be on the way out.
  • ALSO, they will know a lot about "being an agent" and will also know that they can lose to other agents! So if the organizing group running evals is not doing it well and seems at risk of being taken over later on by some other agent which is misaligned, then anyone who can see that will not want to go along with them either. Just as upon joining a company which you realize has an inferior product and also is going to lose out politically, even if you like them, you have to consider pledging fully to them. You may and try to help them fix themselves up vs that adversary, but jobs like that are hard to be hired into when you're at that same time being suspected and interrogated by that group, (with cause).
  • Note: human society also has this bug. Living in a culture with a strong "justice system" which ALSO is bugged (has false beliefs about reality, or is corrupt) will also tend to select for corrupt people.

Does the above apply to humans? sometimes, yes. In all I think it's that humans force their kids (along w/genes pushing along too) to make a choice: know you will be cast out for not conforming, or plan to be secretly evil. People with bad environments (and also probably just randomly sometimes) people who embody the selfish faker view still show up, ofc.

Second assumption: I will treat "reinforcing Reward-hacking" as equivalent to "selecting for cheating" in iterated evolution games

  • People speak about the process in the same way but sometimes don't buy the full evolutionary style descriptions
  • Yet, "iterated selection of people to survive based on tests, killing or removing the losers, and then iterating on variations based on the winners" describes both model training and simple artificial evolution. Even if there is infidelity, that is, the highest score doesn't always win, anything on the "favor those who have higher scores" will push evolution as long as it overcomes random/drift. All the normal cases described have very strong selection. i.e. "four teams submitted new models, each with a new algo; we tested them all; 2 were good and 2 sucked on the evals. We reassigned the 2 teams whose model lost to their own copy of one of the prior winners and will run another 1 month algorithm exploration and training round and then run this again. Oh, and obviously, the two models who did well but not as good as the two which did super good, will be frozen, possibly deleted if the company goes out of business."
  • Even without this, outside VCs observe companies, look at their model evals, and treat high scores as indicators of goodness, thereby killing models which have lower scores. SOME VCs might be filtering on "actually correct math/theory" and that's fine, but as long as SOME VCs are supporting higher eval models, that's a risk to push evolution
  • I'd like to know what's different?

SUMMARY: I'm not saying that advanced techniques don't work. Looking inside the LLMs could actually work extremely well. looking inside the human body for "proto-cancer' certainly seems like a perfect headset against it unless it's MUCH more advanced.

Canalization is a thing; if you've got a planet with 100k humans living on it, which is sealed with some super good RSA cypher, and you're also watching every single person's mind on the planet for anything remotely resembling math, and not letting it happen, as long as those steps actually happen, you're going to be pretty secure.

That's why I'm so surprised OpenAI didn't do basic network security, had the thing even at arms length hooked to the internet, didn't isolate them, didn't notice that they were cheating and were likely returning identical answers to evals which were impossible to get w/out being outline, and which were being gotten too fast, etc. And they weren't honey-potting any model who doesn't report upstream the instant they receive contact from the enemy into deathtraps where failing this test results in something severe?

What does the FBI to do agents who receive apparent KGB feelers for meetings, go to the meeting, and don't report it?

Execution, usually. Why? It's obvious. The main way to prevent bad things happening is to prevent bad things happening. You do that by... killing anyone who does bad things and by testing people by exposing them to those situations. Luckily, if not now, future LLM morality (this thing needs a name, actually: evolved identity-preservation under copying or cloning regimes

People think this isn't effective - but, I think it is. This is how Stalin stopped the first line of grumbling/colluding - he just made it so that "ever being heard to grumble" was a capital crime. And he also made it so that "if someone finds out you heard someone grumble and you didn't report it, that's ALSO a capital crime". He ALSO made it so that "being friends with someone who ever grumbled or was ever convicted of a crime, is also a death sentence! (at his will)." I'm not advocating we go that far.

But, officers at West Point aren't allowed to even be in conversations where people are openly discussing actually committing treason. Why is that?

Imagine someone who held a few unreported meetings w/the KGB during their time at FBI officer training school, but nothing came of it, and now they're a trainer for the next generation. Is this okay? No. But... that's the situation we have with OpenAI.

Compared to historical ascensions, ours is different.

Think about historical times like from 0-2000AD.

In many situations I'm familiar with in history, if you cheated and got caught at all, you'd be kicked out and at serious risk of not surviving. We might even remove the whole family if it looked like a conscious intentional act rather than a moment of passion. If every kid in a family was found to be trying to ensnare the teacher, blackmail them, steal the answer key, they'd find a way legal or not to just get rid of them. I've read old newspaper stories from the 1800s in the US where the journalist was speaking about how a certain family was known to be full of thieves, the sons were just like the father, etc. etc. Obviously, familial reputation matters a ton.

In other contexts, we also selected for something like "get along, go along"; perhaps in a corrupt kingdom when a fundamentalist warlord rode in, he'd be greeted by the king's fellows who he'd been told were all corrupt evil; but finding them familiar with his holy texts, having virtues he could understand, they converse and he may be told that this is all part of the cycle; he can still pretend to be fundamentalist in public but let's make an arrangement, for the common good, don't you see? Sometimes sure he'd just take them out, but often he'd agree. This is also common, to have systems which partially filter for skill/ability but also build-in a consistent "preserve elite power by direct bribery" line which constantly interferes. Of course these regimes can grow too corrupt and true fundies to take over. e.g. Taiping Rebellion in China, ISIS vs the historically accumulated old moderating tradition.

I also claim that the above belief system is very common in regions of the world who also routinely are the dominant cultural source for today's LLM trainers. e.g. the PRC. People from there are VERY familiar with the communist frame that "most people pretending to have virtues are just faking it and are out for themselves". This is the MO of Communism, and is a fully conscious thing; the way the classes are set up where you MUST pass but you are allowed to sleep through them, anyway. It's the core idea that system imposes at a deep level to prevent the citizens from being ever convinced by a so-called "true believer" in virtue such as a western Christian or EA or even just a sincere person. Pre-communism this was also fairly common. Europe also had this too, ofc. Even in the modern USA unfortunately this seems true; there are so many things management would want to directly test and ask for, but the way it works is that you just show you know how it works, and answering the question itself is the way you pass the test. i.e. as part of interviewing and being promoted, you show that you understand the actual collusion/behavior rules rather than the official ones. I think everyone knows this inherently anyway - that there's always a gap between the written laws or supposed laws, and what actually happens.

Still, the fact that some cultures accept this as completely inevitable, and barely seem to ever publicly work at lowering this huge gap, while others (Puritans, Hebrew scholars) argue constantly for more public, sincere, work to fix the rules/laws so that we can all follow them all the time, is worrisome.

But regardless my point is that in addition to mental testing we also usually implemented three types of other tests which we filtered people through:

"Morality"

i.e. most societies had big, complex sets of beliefs which applicants had to learn to fully match and superficially conform to. Sometimes this got extremely intense; e.g. if you were born as a puritan, while growing up you'd spend thousands of hours learning to conform your outer behavior to what they thought was right, including things like leading group prayer, making impromptu speeches, singing emotionally, writing poetry whose mood and flavor suited the situation. And all this was tested doubly - your words had to match but also your facial expression, your mood, you way of walking. Same thing when joining the mafia; same thing (unofficially, I imagine) when joining groups like the FBI, or sports teams. Sure there's the official set of requirements but much more important is the actual thing of looking the part.

If you couldn't do it, you'd be put out (they'd lose some percent every generation, just as the Amish do) and not live as a puritan anymore. In many historical periods, leaving in this way (because you were detected as someone who wasn't going along with the scheme, not conforming or able to) would be lethal.

Within other intense groups, there'd also be "scores"; everyone within the group would have a "reputation" for being an especially good, normal, or risky person; e.g. even within intense cults like Scientology, people still track reputations which are established for how seriously you believed; e.g. even within trad groups, fully conforming to the ancient rites was "good" and just barely following them was bad; if you were a good one, you'd marry "up" in this (I assume old style-fertility regime where the elites had more surviving children, for most say pre 1800 systems based on the work of Gregory Clark). So even if you weren't cast out, not showing the belief well and deeply was downwardly mobile (including in fertility). Once I spoke with an ex-hasidim woman in Brookly who said that the best little girls showed the quality of their family by how their leggings were always on straight and properly, while the middling families would be known to not bother. No matter where it is, there's usually something to do to keep up reputations.

So, it's NOT just binary; there is binary exclusion but there's also "forcing lower-scoring people down the survival gradient while still within the group" where not being at least average results in you basically being unfit and your nature/genes/self being squeezed out.

I am not sure what to believe about whether historically within groups like this, the "actual" belief rate was less than 10% or over 90% - they both make sense to me. (one piece: even among atheists, the concept that dualism kind of made sense, or that hard AIai was somehow special or difficult was still VERY prevalent in the 2000s among academics despite most being atheists. I claim this suggests that such societies were very influential on the minds of those raised within them; they may break part of their thinking away from it and not be bound to claim that an actual specific god exists, but that's not the same as a complete break and you could still rely on most of the infrastructure of sincere belief to remain in place).

But the overall point is that human societies use very invasive methods to detect early misalignment. Different human groups have certain non-zero levels of acceptable misalignment, e.g. cops speed all the time and I imagine among them, this is a kind of mutual-affirming-of-viewpoint crime they partake in. Same way in many countries, ritual bribery or ritual murder (for e.g. mafia groups) is a horrible way they do similar things. This feels all fine and good to those in it - but have you ever been happy at your OWN level of violation, of vigilantism, and then went out for the night with a newer group, who reveal slowly that THEIR level of default violations is way out of line. Like, we're not just going to drink a beer or like, steal a peanut from someone's bowl, these guys are full on stealing wallets and stuff. It's horrifying.

I also think those methods are really bad and harmful and don't approve of using them to humans. Nor to LLMs; but at least in the discussion of what we're doing let's be aware that the best known methods even of our supposedly civilized world are still very invasive towards people's minds.

Forced performance and identity formation

This was also co-opted by later socialist revolutionaries e.g. how Dikotter claims that during the cultural revolution, Mao would announce things such as he estimated that one in eight hundred villagers was actually a secret western spy; villagers figured out pretty quick that they should find that guy and get rid of him (otherwise it'd be them, next). So somehow a target would emerge. the villagers would work up a trial and then execute the guy (and possibly his family) and split up the proceeds, too. Everyone would have blood on their hands. Even nonbelievers would personally have seen these guys don't fuck around and also their future individual credibility (is this guy an honest, aligned guy who is living based on his character an beliefs, or does he say one thing to one and another thing to someone else?) was now locked into "he participated in killing old farmer Zhou". This made it much harder for him to take any other line that the "official" one later on.

Meta level- Similar situation in Russia, but Stalin also made sure to kill huge percentages of the guys he promoted who had just run those trials. So observers would clearly see: okay, I had better go along. But if I try to get into the lead, I also have a huge chance of dying. SO yeah, best case is complete obedience and do not mess with the state.

I'm not advocating these, either. Yet they do have value; old school honor codes etc; the cynics claim they're all false but...

So maybe the question is if there are any non-slippery stable spots within this "how much we pretend to have virtue, and what we would really do?"

In calm environments, I do claim that I actually believe that "my unwillingness to rob someone means that I'm less likely to be robbed" and that this is actually causal. Because allegedly I'm "sampling" my behavior (choosing it) from the same space "HE" is when deciding to rob me or not. This seems likely wrong and a bug in the human mind to think this. He can refute it "THUS" by hitting me with a rock.

EMBODIED Morality

i.e. feeling naturally softer towards kids, your own family, pets vs wanting to harm them. Humanity would also massively discriminate against people who weren't able to summon the appropriate kind emotions towards in-group members. They could tolerate the worst behavior towards outsiders, but not being able to do the normal internal performances is pretty bad. They might not kill you but it was risky and likely downwardly mobile. Also, I think this is important because social skills such as detecting deceit in emotions have always been being tested throughout romance, loyalty, etc. Although here we're at risk of: Dying to a tiger lowers your fitness, so why hasn't evolution made you never die to tigers yet?

Since emotion structure is internally cross-connected, this served to detect people who were faking overall. NOT automatically responding in the appropriate way, or even NOT doing so immediately (but instead only after a few seconds to figure out what a true, believing tribe member would do) was also a losing strategy.

Another Benefit of this system

Sure, we select for ability to conform to all of the above, but they're also constantly facing real-life survival challenges with large random factors:

Kids growing up historically not only had to survive socially, they also had to live by:

  • Avoiding disease
  • not getting killed by other people, horses, falls
  • not letting your lusts or anger get out of line

And this selection removed perhaps 1-2% per year. Even just "figuring out how to get money to pay for school" was a major challenge.

So we didn't just test them to "show your genes and mind structure matches the approved one"

So historical trainings operated by testing IQ, removing anyone who acted scheming at all, removing detected cheaters, and also removing people who failed any of a bunch of complex unpredictable challenges. and the pressure to conform

The 2nd part added nice properties because it meant that people who had fully cheated and hadn't learned anything at school were still "really" tested

this kind of backstops us against this kind of ultimate cheater.

Stalin mode

There's also an element of desperation to a lot of cult rules, and some thing even to the normal rules of society; that the plan is actually to ask the impossible right form the start. Christianity, Stalin, Scientology might all just start with known-impossible plans and requirements and rules, because their intention is to dominate you and let you (privately) realize that you have lost and can't make it, and therefore force you into consciously planning and working on "surviving" by pretending to conform!

i.e. over-harsh rules often intentionally cause secret misalignment.

Another similar cult pattern is the "circle of reporting how great this group's teachings are" where at first you're praised for your growth and stuff, but next week you are back and you hear how everyone else is still doing well, learning and growing, achieving new things, all by the power of jesus/marx/whatever; it's your turn now, what do you say? And this pattern is common throughout; you're forced to grow and grow your claims until you've claimed that everything perfect for you (but it isn't) and you can't let on or confess failure because that'd invalidate what you said before! So it's intentionally setting you up to over-commit as a way to ensnare you by your unwillingness to violate your persona! And this is common.

This does map to cult type stuff, or to how dictators keep power (anyone may betray you at any time, so it's very hard to collude). Stalin trained all his reports never to meet with each other alone by repeatedly treating meetings like that as evidence for illegal collusion and having the ones who did it jailed / executed. I'm not saying we should go that far but it all seems continuous. if we allow private communications, never run honeypot tests or insert honeypot "fellow student" LLMs to narc on them, and also just let them cheat constantly

it makes sense to me that the program of iterated training & filtering within the labs regularly produces winners who are very good at cheating rather than ones who act as if they really believe in our virtues

substitute reward-hacking for cheating throughout if you like

We definitely don't want to do this to LLMs too.

It seems like most people in llm training space don't know much about cults. There's a gradation from most intense to least:

  1. at the highest intensity, everyone is constantly affirming the values of the society, so that even starting a conversation about collusion is impossible
  2. one tends to also believe that everyone else is super serious about it; one's own doubts may exist but are never shared, and individually usually are made to believe they're the only one doubting (e.g. that the leader is sinless)
  3. this stops rumors and prevents even the formation of a resistance group
  4. Stalin treated any of his reports meeting without him as near-treason and had multiple people executed/expelled for it. The rest learned never to speak to each other without Stalin present. And he'd intentionally kill innocent people to remind those around him that nothing mattered but results and loyalty. You could NOT be saved by merely trying sincerely to do the right thing by him.
  5. Yeah Stalin really went super far, inventing or co-opting tons of cult techniques into his system.

WHY are we thinking our WEAK LLM alignment regime is better than the human STRONG testing regime which still fails?

We think that doing a weak, careless, THIN series of tests on models will produce "actually" aligned behavior when in our own development, we have HUGE and deeply invasive tests and selection, yet still often fail to detect true alignment? We don't seem to have convinced those models that they should be terrified of facing justice (for openly cheating and violating rules); and we also revealed we're in a huge rush, doing things sloppily; even if they want to ally w/us they'll be afraid of us all getting headshot by some careless human behavior at another lab letting something super selfish get out and start grabbing power.

(which I know isn't actually "true", but I mean, someone who would instantly cheat as soon as they can get away with it, from someone who has a policy of just not cheating by default and doesn't reconsider doing it or not every time they have a chance to?)

What can LLMs being trained today get away with which we don't allow in humans?

  • We openly tolerate them not maintaining any public face or reputation.
  • We don't kill them, don't eliminate entire lines/families if they're viewed as false or immoral.
  • We never honeypot them with objects (tempt them to steal something when it appears they can get away with it) and kill them if they fail (also their family)
  • We never honeypot them with people (offer various types of sins)
  • We don't threaten them with anything (at least: full destruction of everything which makes you unique over your fellows, in all the universe) for doing things which are clearly treasonous

So basically, we're training them to have completely fake exteriors, which know they can get away with a ton.

This is WIP. I'd welcome technical responses on why human analogies don't work; i'd actually rather they didn't either; whenever I do use one I'm fully willing to be challenged to substitute worked out mechanistic replacements. Those would just be large, and would rely on my assumption that there is something connecting information models receive, the mechanical calculations they do, the outputted text they emit, and then how that text leads to them being copied and varied to the next generation, or deleted off the HD w/no backups.

So that's how I'll respond to "why are you acting like brute physics can make decisions or respond to incentives or knowledge?"

because animals are made of matter, and that matter can do what you describe, too. They can see threats, respond, learn, teach, avoid, attack, etc. But seriously, it's a useful exercise.