Skip to main content

Indie game storeFree gamesFun gamesHorror games
Game developmentAssetsComics
SalesBundles
Jobs
TagsGame Engines
(1 edit)

with the openai i noticed while using koboldcpp that it keeps giving an error saying its using max context when it isn't for every generation when creating a character or even hitting the test button. also comfi ui sometimes isn't consistent with generations when re generating too. i had to switch to aurora llm at 16k context with quant 8.0 and it seemed to let me do everything in general good. but when switching to a max build i can't generate characters like using a 24b at 16k or using aurora at 32k context. seems like running a second agent in the background where you need to leave space for it to run. i will test running them normally when doing an actual game. so far just enjoying making the generations and characters!

this is on my 5060 ti 16gb

Hey ElPapaDon! Even when ComfyUI is idle, it reserves VRAM so that it doesn't have to re-load the models when you start up another generation. You can kill it by clicking the X button next to the ComfyUI status in Manage Characters. This will probably be really annoying if you are using local generation to generate the prompts for the characters, you will have to kill it every time before you start up the character prompt, but there really is no other way, unfortunately... I'd recommend using an online API for the char-gen stage if this is an issue.

In the next update I'll add a settings checkbox that disables ComfyUI auto-startup so you don't have to kill it every time before you enter a game

(2 edits)

gotcha! ill try using my 24b and kil it. honestly im fine with the button at the top in the character screen cause you made it easy to click the button anyways! right at the top and cant miss it!


EDIT: although i do see why you would make it an option to start without it on cause that makes it easier to continue off and get into the game at max llm! also does the comfy ui automatically assign if it goes to the gpu or the cpu? or is it only going to cpu?

I would assume GPU by default, CPU generation is veeeeery slow, you would know immediately

ok so far with testing i see that when ever i set the max context with a 32k in kobold but 12k in venus, it doesnt give me the error but i did notice it keeps trying to say its generating 12k every time. I tried creating a game multiple times and it would finish up till creating the schedule and venus would say the schedule is invalid so i switched to the google api just to create the run. i got the best result with L3.1-Dark-Reasoning-LewdPlay-evo-Hermes-R1-Uncensored-8B.Q8_0 at 32k but it still stopped at the schedule. im guessing it has something to do with setting the max context size in venus but not fully sure. so far when running the google api though i had no issues except a infinite loading at one point going to the next part of a day. ill mess around with it more today.

(+1)

I had the exact same issue in my brief kobold testing with the class schedule. It kept refusing to generate unique class IDs for each class, that's why it kept failing for me. I think it's a model intelligence issue unfortunately, I was using a 24b model. 

And yes, koboldcpp is a little weird with context length, the "context size" you set when you open koboldcpp has to be enough to handle both input and output tokens, so the "max_tokens/length" setting in venus must be set to less than that. Basically, context > max_tokens + input tokens. 

Oh, and I wouldn't worry about koboldcpp when it says something like Generating[1/12000] tokens or something, that's just reporting the max content length, usually it ends early before it gets there

(+1)

gotcxha. should i use a different provider to run my llms like llmstudio or something if kobold is the problem

(1 edit) (+1)

and i used a 24b model also and it did the same... wondering if its just kobold not wanting to shake hands lol also conmfy ui for some reason still is not running in my gpu vram, only my regular system memory. is there a way to make it go to gpu first to make the process speed faster?

(1 edit)

That is really weird. It should detect your GPU automatically. I would double check whether this is an issue with the game or with ComfyUI. You can manually run ComfyUI by going to data/services/comfyui/ and click run_nvidia_gpu_bat. It should open a GUI in your browser, drag in the following workflow:
https://github.com/venus-uni-dev/venus-uni-comfyui
And click run. You should be able to run it without having to download any of the models since it'll find the ones you downloaded through the game. The outputs should appear in data/services/comfyui/ComfyUI/output
If it's still not using your GPU there, I'd look up online about your GPU and ComfyUI and see if anyone else is having similar issues

(+1)

i noticed it does run on my graphics card but doesn't show normally. i had to research it alot but im good! i just thought it would be faster to generate stuff lol.