Fantastic update! And also an exceptional discovery to making open source LLMs work for game creation (Tested on a 12B model)
First of what I found (Worth testing and letting me know how you get on):
TL;DR (Use Llama.CPP (I used ROCM but the normal version should be okay))
My build: D:\ROCM\llama.cpp>.\build\bin\llama-server.exe -m "D:\ai\G125.gguf" -ngl 99 -c 256000 -b 512 -ub 512 -fa on -mg 0 --port 5001 -np 1 --jinja --reasoning off --temp 0.1
Ones to have: -c 256000 -fa on -mg 0 --jinja -np 1 --reasoning off --temp 0.1
Using Gemma 12B Q5 Heretic (Not certain if this is the one I used, but I do have the GGUF file if it's okay to upload a google drive here, but I'm pretty sure I used this: (https://huggingface.co/JMingo/gemma-4-12B-it-heretic-UD-GGUF)
Game settings: Max output tokens - 256000, reasoning effort - medium
Then create new game, wait up to a few minutes, could be up to 20K tokens in the terminal. After completion, use your original model (Recommended). If infinite loading, restart the game, model, backup & restore and try again.
Unsure on the quality yet. Theoretically doable on most models so long as you set the temperature extremely low - base level of 0.1
My theory is because the temperature is 0.1, it's being told to be AS predictable as possible, so that attention fade can't necessarily affect it during creation.
-----
I've been playing whack-a-mole with open source ai-models in hopes of finding one that will work to create a new game without reliance on external API, I have 24GBs of VRAM (RX 7900 XTX) and have repeatedly failed to create a new game with the only exception being the Gemini API.
I've tested: Qwen 3.6, 3.8, Gemma 4 31B, 26B A3B MoE, E4B, Cydonia 24B (& Magidonia 24B) Orion 26B A4B, Waifu Gemma A4B 26B , Deepseek R1 32B, Gemma 14b & 12b
Alongside many different quants to see if any differences would be made
However with all quants and models, classes fail to be created, the main idea behind it is that the logic on these smaller models are too inadequate.
It was further off-putting to see someone comment that GLM 5.1/5.2 didn't work (753B parameter model!)
However after some research, I found llama.cpp supports custom configuration inputs, so I had a play.
Long story short - I tried a stupid idea I had made it actually work on Gemma 12B Q5 model!
I thought it was a fluke, but it did it again successfully 4x, taking about 30-60 seconds each time and seems to be very detailed.
I ran into a weird problem where despite generating successfully, it'd infinitely generate whilst the quick start & gemini created ones would work just fine, but nothing seemed off in the logs.
Whilst testing to find a solution, it seems to have sorted itself when I changed model, restarted the game, backed up & restored. You might not receive this problem.
However I'd advise only using this method to create the new game, and then switch back to whatever you were using previously.
Setup:
D:\ROCM\llama.cpp>.\build\bin\llama-server.exe -m "D:\ai\G125.gguf" -ngl 99 -c 256000 -b 512 -ub 512 -fa on -mg 0 --port 5001 -np 1 --jinja --reasoning off --temp 0.1
-c 256000 -fa on -mg 0 -np 1 -jinja --reasoning off --temp 0.1
Game settings: Max output tokens - 256000, reasoning effort - medium
About the update and some small feedback:
Re-gen on text messages would be nice to see, but creating a save file and then loading from there works, or them being able to add a react emoji when you say bye would be a nice addition
Being able to give gifts when manually overriding the current conversation rather than waiting for a trigger
A symbol/indicator to let you know the scene is ending soon
Apart from that absolute brilliant, and hopefully that discovery helps someone out!
