Your ideas are ones that I considered too. Waiting 30 seconds in a text conversation is sort of unacceptable for me in terms of gameplay smoothness, but the loading idea is great actually: you can still continue the convo/enter a scene while it's loading I'm assuming? That's actually how the photos feature is implemented right now, so all I need is to instruct the text prompt to add an extra field for an image prompt when it feels like it. Thanks for the idea!
Oh, the current photos feature uses frontier image model for generation, since local image models aren't good enough for prompt adherence. So that means a larger concern for me is cost. I feel like allowing images to be generated during text/feed might cause prices to skyrocket very quickly. I suppose a setting helps with that but it feels like something that it would be expensive for me to even test lol
You're welcome to post the repo. I might take a look, but usually the details of the implementation are what is interesting to me, not the actual code, so what you've provided is already more than sufficient. Thanks for sharing
Viewing post in v0.2.0 - Known Bugs, feedback, suggestions
Yes, exactly. The photo shows as a loading bubble and you can keep texting, close the phone, or go into a scene while it draws. The reply box stays usable the whole time, and the picture just fills in when it's ready.
Makes sense about cost. My version is built entirely around local generation on the player's own ComfyUI, so cost never came up for me, and how to handle it with a frontier model is your call.
On local models and prompt adherence, I hit the same problem, so I tried not sending the local model prose. The LLM writes a short caption of the photo, like "lying on her bed in her swimsuit, looking at the camera." Then code turns that caption into Danbooru tags, which the checkpoint understands much better than sentences:
- Outfit: if the caption mentions her swimsuit or gym kit, it uses her own outfit tags for it.
- Coverage: it works out what's still covering her and what the camera can see. For example, a photo from behind never mentions her chest.
- Pose: it picks pose and hand tags from the caption's wording, like lying down, sitting, or holding something.
For consistency I also added optional body fields per character (build, chest, hips, backside), each picked from a short list of tags the checkpoint knows and checked against each other, so her body stays the same across sprites, CGs and photos.
It's still far from perfect and needs more tuning. Outfits only roughly match her sprite, and odd captions can drift. But since the game already renders CGs locally with ComfyUI from the same kind of tags, I figured why not reuse that for photos too. It's the same pipeline and the same checkpoint, just pointed at something new, and it costs nothing extra for players who already have it set up.
One more design note: what she's willing to send is decided from the save before the model is asked (nothing if she's annoyed or hostile, flirty after a kiss, more only with a reason like a crush or having slept together), and her caption is checked so the model can't go past that. That keeps it in character.
Here's the repo if you want to look: https://github.com/naudh1r/venus-university/tree/photo-feature
It's your mirror with only the photo feature added, and each commit message explains the reasoning behind that change.
Thanks for taking the time, and good luck with the update!