Skip to main content

Indie game storeFree gamesFun gamesHorror games
Game developmentAssetsComics
SalesBundles
Jobs
TagsGame Engines

es4rtg

6
Posts
A member registered 13 days ago

Recent community posts

(1 edit)

Fantastic update! And also an exceptional discovery to making open source LLMs work for game creation (Tested on a 12B model)

First of what I found (Worth testing and letting me know how you get on):


TL;DR (Use Llama.CPP (I used ROCM but the normal version should be okay))

My build: D:\ROCM\llama.cpp>.\build\bin\llama-server.exe -m "D:\ai\G125.gguf" -ngl 99 -c 256000 -b 512 -ub 512 -fa on -mg 0 --port 5001 -np 1 --jinja --reasoning off --temp 0.1


Ones to have: -c 256000 -fa on -mg 0 --jinja -np 1 --reasoning off --temp 0.1


Using Gemma 12B Q5 Heretic (Not certain if this is the one I used, but I do have the GGUF file if it's okay to upload a google drive here, but I'm pretty sure I used this: (https://huggingface.co/JMingo/gemma-4-12B-it-heretic-UD-GGUF) 


Game settings: Max output tokens - 256000, reasoning effort - medium


Then create new game, wait up to a few minutes, could be up to 20K tokens in the terminal. After completion, use your original model (Recommended). If infinite loading, restart the game, model, backup & restore and try again.


Unsure on the quality yet. Theoretically doable on most models so long as you set the temperature extremely low - base level of 0.1


My theory is because the temperature is 0.1, it's being told to be AS predictable as possible, so that attention fade can't necessarily affect it during creation.

-----


I've been playing whack-a-mole with open source ai-models in hopes of finding one that will work to create a new game without reliance on external API, I have 24GBs of VRAM (RX 7900 XTX) and have repeatedly failed to create a new game with the only exception being the Gemini API.


I've tested: Qwen 3.6, 3.8, Gemma 4 31B, 26B A3B MoE, E4B, Cydonia 24B (& Magidonia 24B) Orion 26B A4B, Waifu Gemma A4B 26B , Deepseek R1 32B, Gemma 14b & 12b


Alongside many different quants to see if any differences would be made


However with all quants and models, classes fail to be created, the main idea behind it is that the logic on these smaller models are too inadequate.

It was further off-putting to see someone comment that GLM 5.1/5.2 didn't work (753B parameter model!)


However after some research, I found llama.cpp supports custom configuration inputs, so I had a play.


Long story short - I tried a stupid idea I had made it actually work on Gemma 12B Q5 model!


I thought it was a fluke, but it did it again successfully 4x, taking about 30-60 seconds each time and seems to be very detailed.


I ran into a weird problem where despite generating successfully, it'd infinitely generate whilst the quick start & gemini created ones would work just fine, but nothing seemed off in the logs.

Whilst testing to find a solution, it seems to have sorted itself when I changed model, restarted the game, backed up & restored. You might not receive this problem.


However I'd advise only using this method to create the new game, and then switch back to whatever you were using previously.


Setup:


D:\ROCM\llama.cpp>.\build\bin\llama-server.exe -m "D:\ai\G125.gguf" -ngl 99 -c 256000 -b 512 -ub 512 -fa on -mg 0 --port 5001 -np 1 --jinja --reasoning off --temp 0.1


-c 256000 -fa on -mg 0 -np 1 -jinja --reasoning off --temp 0.1


Game settings: Max output tokens - 256000, reasoning effort - medium


About the update and some small feedback:


Re-gen on text messages would be nice to see, but creating a save file and then loading from there works, or them being able to add a react emoji when you say bye would be a nice addition


Being able to give gifts when manually overriding the current conversation rather than waiting for a trigger


A symbol/indicator to let you know the scene is ending soon


Apart from that absolute brilliant, and hopefully that discovery helps someone out!

Definitely frequency & repetition penalty modifiers would be great to see, maybe a system prompt input to further help direct a local language model to help it write a certain way, or ensure it uses sprites, change backgrounds & actually end the scene

top P & k too, but that might be a stretch as that's going into ui bloat

Also the Gemini Key problem was something on my end so that's of no concern and that's my bad

Hi again, I've sent you an email


The issue seems to happen on KoboldCPP with a header error, but after doing that fix, it works seemlessly on KoboldCPP & Llama.CPP, if someone else has the problem then that fix is there readily available


Hopefully the new update is out soon as it sounds like really big improvements are coming, I'd love to see the ability to modify temperature and other parameters to adjust outputs as I'm now using Qwen 3.8 27B, it works exceptionally but the dialogue is very strange 


(1 edit)

That's really good and looking forward to the next update, it's a brilliant project so far!



A few other things - Is 127.0.0.1 hard coded? I can't seem to get it to work over my local network but it works fine on the same device (Got it working by opening a local proxy on my client device - netsh interface portproxy add v4tov4 listenaddress=127.0.0.1 listenport=5001 connectaddress=(HostIP) connectport=5001 - & then using http://127.0.0.1:5001 to connect (KoboldCPP) which worked

Hopefully you've changed your API key as there seems to be one for Gemini in the .json file, unless that's a placeholder

Definitely worth adding a re-swipe since I can see what you mean by local are lackluster compared to Gemini, I found Gemma 12B IT Heretic GGUF Q6 to be quite good and it gives me the generation speed, but it does mess up once in a while so that & a force end would be amazing

One other bit of feedback is allowing the option to keep memory of previous events or changing how pre-determined events are created, maybe having it done when the scene is created rather than being statically baked in the initial generation - For example, if you've been talking about something specific with characters the morning/night before, and then go to visit their dorm, one of mine was pre-written text about them never meeting me before (Gemini Flash 3.6 Lite)


Again really good project, keep it up

(2 edits)

What open source model did you use, and were you able to get it to make a new game? I have 24GBs of RAM, but I'm unsure what backend to use and with what parameters as I keep getting a connection lost mid way through or something about the generated class catalog was not usable. Failed: the class catalog

Would be really cool if in future you could add additional context when creating a game to specify certain elements, or modify the context of existing characters

Additionally to force end a scene as I'm stuck in an infinite loop essentially, retry a scene if the ai messes up too

(1 edit)

What GPU are you using? I'm using the exact same model and quant, It's taking forever for me to load it up via llama.cpp on an RX 7900 XTX and I keep getting an error that it lost connection to the endpoint was lost mid-reply, is there something specific you did, what context size are you using?