Play project
Formalizing Morality (Neodore)'s itch.io pageResults
| Criteria | Rank | Score* | Raw Score |
| Overall | #14 | 1.581 | 2.500 |
Ranked from 2 ratings. Score is adjusted from raw score by the median number of ratings per game in the jam.
Leave a comment
Log in with itch.io to leave a comment.

Comments
had a go at using the framework to benchmark my friend's theory for alignment (Decision Topology, [https://github.com/yusuf/decision-topology]) with three agents — gemini 3.1 pro, codex 5.4 xhigh, and claude opus 4.6
in all three runs DT managed to crack the top 3. some caveats and interesting finds:
one thing worth noting — DT has a lot more description in the catalogue than anything else, which probably affects the results. more machinery spelled out = more surface area for the evaluator to find passes on. so take the absolute numbers with a grain of salt.
one of the more interesting things is that despite no contact between us, there seems to be a convergence ? SFOM starts from sentient experience, DT starts from the topology of agency — completely different entry points but they both land in a similar structural place. it leads me to believe there's a unified field theory for alignment and I think that's inevitable.
claude rated DT extremely high — 29/31, only failing on phenomenological accuracy and motivational internalism. not entirely sure why it scored so much higher than the other two models. might be the description length thing, might be that opus weighted operational specification and implementability more heavily. interesting either way.
in regards to sycophancy we now have personal theories that aren't in the training data (including SFOM) that have cracked the top 3 across multiple runs. if the models were just flattering us, moral realism wouldn't still be sitting at 12 and divine command at 8. .its also worth noting that its hard to come up with something that ranks higher than SFOM with it being the top across almost all the runs .
honestly? it's not that hard to come up with a theory that cracks the top 3. the existing theories are just very very outdated. we're in a kind of sub-advanced culture of alignment which is a much harder framework to address than what kant or mill were working with. any modern theory that takes AI, non-human scope, and formal specification seriously has a structural advantage out of the gate.
the convergence we're seeing points to something real.
perhaps something even more interesting would be to benchmark something much older — like the Pali Canon or something but without explicitly framing it as buddhism or religion
You can find the results here https://drive.google.com/drive/folders/1kf6BtI2RtfZ3tmctBEBAzANzIemwXkFI?usp=dri...
it would be even more fun if this was automated !!
all of the top theories have a sense of irreversibility and futures closing so thats also some thing to note
Awesome! you're the first person that tests a theory that beats SFOM in any training run since 2025!
I think your caveats are good, but it's still interesting to see that SFOM beats DT in 2/3 and beats its total average too, despite it being only a couple sentences versus the far more fleshed out DT submission! (In the future automated & expanded system + theories, it will be interesting to see how these all fair out!).
I overlooked DT and think there's a lot of good convergence between it and what fleshed out advanced SFOM from the time and, the theory without name fleshed out since then has in many regards, it shared w/ decision theory & neo expected utilitarian frameworks. I like that it's trying to reduce things to one axiom. I'll unpack w/ more time, but it's interesting signal and historically interesting!
I like to see others already pushing back with the system itself and testing it! and I love ur idea of testing other world religions too! I think the expanded system should essentially catalogue all philosophies, cultures, religions, etc. and also break them down into claims, beliefs, arguments, etc. to show their overlap, recombine them. benchmark and test at the atomic level and then their overall, and therefore recombine them, etc. pick from eachother's learnings, etc. test them in the fullest resolution possible,etc.
This is a noble aim but I'm a bit put off by the pitch's focus, in its "why you" section, on _work_ rather than _results_. Anyone can put in many years, many words, many publications and fail to make any progress towards their stated goals. And someone who so heavily emphasises proof of *effort* rather than proof of *outcome* makes me concerned that the author is not aware of the fact that effort and results are often not correlated - that they think that this emphasis is justified itself causes me to update away from believing in their work. So I think this is an angle of the pitch that needs reworking to have people consider your idea more seriously.
1. Totally fair criticism! In fact I think the pitch has much to improve upon
2. I put that there bc most of my R&D has been monk-mode stealth so I didn't have much to show publicly sadly / explaining it wouldve taken up more space from the 3 pages and i was running out of time. I also would rate my pitch very low in its current state.
3. I do think however that in the absence of the other evidence, showing consistent and thorough persistence and engagement is a signal, obviously not as strong as pure outcomes; but although they aren't guarantees, I do think a lot of deep research discoveries / invention come from people who spend a lot of time focusing on the problem from many angles before cracking it. + it's also proof that the person is doing it not for the money, etc. there is some signal, but i agree w/ you it's not enough! The effort volume isn't to convince people of the idea, but of the dedication to it. that it's not some side hobby or fleeting idea i just had that i will drop at any moment, etc.
Your heart is quite clearly in the right place. No doubts there. My concerns arise from a few things -- first of all, from a cursory review of your methods in assessing your theory against others in terms of validity, as judged by LLMs. You list your theory as just one of many, instead of giving it a special position, which is good. However, all the other theories, as far as I can tell, are fairly thoroughly documented in the training data. The LLMs can effortlessly derive that yours is the odd one out because they haven't heard of it before. That does open the door to sycophantic reinforcement, unfortunately. Now, as far as your assessment criteria... I feel like they're a bit subjective, and driven by a presupposed moral frame. Some of the criteria require moral assumptions in order to be viewed as valid criteria, which makes me feel like the theory is chasing its own tail a bit. I also wasn't able to find the theory beyond the sentence or two description in the lineup with the other theories. Do you have a more complete and explicit definition somewhere? I'd definitely be interested to read it.
I suppose my general concern is that, so far, it has felt like the theory seems to be the water in which Bay Area folk swim. I get that the idea is for it to be so obviously true that any intelligent entity would agree almost implicitly, but I don't see that as necessarily being the case. Again, I would like to either talk to you about the theory and how it runs up against certain moral scenarios, or just read a more complete description of the theory. Either would allow me to update my rating to more accurately reflect what I think of the theory itself, though I'm giving you points for being motivated by good intentions.
So actually, there is another one that isn't mine or in the training data bc it was a friend's own personal theory (SKL) and there's also came up higher on average, but still didn't come out on top. Someone at the hackathon also wrote down a sophisticated version of their own moral theory and ran it and it didn't break the top 7 and SFOM still came out on top.
So yes I think that you make a great point and that that's one of the things that should be added to a more sophisticated version of the benchmark ranking system! (in fact I wanted to automate it and create a site so people could add their own more easily etc) to make the ranking and research loop continuous. I think the other two friend theories coming up above the top 50% in ranking shows that you might be right that the AI's have a sycophantic bias to out of training data theories... however on the other hand I thought they might have a bias towards theories in their training data since they have a lot more information on them versus the personal summaries of out of training data theories (like my own / the friends & the hackathon's submission) so it's not cut and dry which way the system would be biased.
I did talk about the limitations of this current rough version of the benchmark on the site, but I'd add all your concerns and more to it. So maybe instead of 2-3 non training data ones, we will have many more to reduce sycophancy. but it is interesting to note as a counter to your point that among the 3 non training data theories, SFOM still ended up on top. I think although imperfect, it is a non-trivial signal that this should happen consistently.
I do think the criteria are also not the strongest. Again they were LLM generated at the time (to reduce my own bias) by asking the LLMs to come up with a thorough list of criteria to assess and rank moral theories. There are many criteria however that I think can be cut, merged, or refined, and i want to add quantitative criteria, and more. With funding I'd also try to get expert philosophers to contribute and write their own criterion, etc. Plus the automated website version would allow anyone to critique and submit their own criteria.
I think you're making a good point about the circularity of some of the criteria, but that's a whole other can of ouroboros to discuss in more detail!
Once again as stated in the doc and site, the benchmark is really a rough proof of concept signal, all the theories listed are listed as summaries with an example. The SFOM summary there was a piece of my theory written in 2025 focusing on key components like the Subjective v Objective morality resolution, but it was by no means exhaustive. and it is currently outdated as things have advanced substantially in the last year. So I will keep you posted as I begin to release more refined theory!
I would say there is some overlap w/ the EA theory and other theories before it like moral naturalism , natural law , utilitarianism, deontology etc. because I do think that all of those theories contain kernels of truth / are special cases or subparts of a larger unified one. A lot of the theory should be compatible with past correct theory , but more cohesive and general, and yes it should feel obviously true if it is obviously true.
One of the things I wanted to add to the benchmarking is moral scenarios as you mention, so maybe you could write down which ones you'd think are worth testing and I'll add them in the future benchmarking system! :)
I feel bad that you feel you need more information to rate me accurately, because I agree! this submission is not a good reflection of the work / theory in a complete or updated sense but i rushed this submission to participate bc Defender wanted me to and I thought it would be a fun low stakes way to push myself forward. so i apologize if you feel your time was wasted and would recommend ignoring me / waiting until I publish somewhere more fully / thoroughly.
I don't at all find it to have been a waste of time, it was an interesting read, I just don't feel like I understand what you're selling, yet. I think the EAs are ridiculous, and, while well-intentioned, absolutely rife with practical failings akin to that which have created many hells in recent history. I would want to be able to read your actual theory, even a several sentence summary, to be able to understand that which is being posited.
It can't be summarized into a sentence in a way that closes the inference gap. But something along the line of: formalizing axiology means we make understanding what is good/bad valuable/invaluable objective relative to the frame of reference of the invariant properties of all subjects (for example pain is painful tautologically, unlike what some ppl think where they say that pain can be pleasurable (they're confusion the net direction of a vector where two component vectors where the pain magnitude is smaller than the pleasure one). Then because these will be obviously true as anything true is, it will scale in acceptance with intelligence, and as we improve upon this measurable quantitative theory grounded in instantiated subjects and thus physical measurable dimensions, we will open up a new field that allows systematic progress on the ethical questions like in any scientific field, but also it will spread in a feedback loop of low friction just like all scientific/mathematical truths have across cultures. Particularly the objective vs subjective (moral nonrealism / relativism etc) debate is closed also, in the same way berkely/ einstein solved absolute v relative time debate : it's not one or the other, but actually frame relative yet still universally describable (& therefore predictable and explainable given correct measurement and frame etc). (and that's just one piece of many). I will ofcourse flesh all of this out in further publications before expecting this pitch to be any good. Feel bad for posting ahead of schedule, but alas Defender also is good at pushing things forward XD.
I think EAs got many things wrong (valuing the lives of people that don't exist in the future above current ones, focusing on earn to give instead of doing, etc) but are generally better in many of their predictions (ie AI ) than most which is why theyre so well funded, and that they're getting a bad rep bc of a few bad actors (FTX et al) which is unfair. the truth is they're approximately more accurate than many other groups and trying harder. So we can debug them instead of discard them imo. wdyt?
Also; as a reponse to your sycophancy worries. There is more than 1 non classic theory invented by someone in the list (the other one is SKL) , and another one was tested at a hackathon last year. neither of them beat SFOM. Also if you check megmind's comment, you'll see they tested it against another theory of a friend of there's and his analysis is interesting! I think sycophancy can be reduced as factor even more when including more theories out of training data set that are original . (and a future automated version of this will include many more!) :)