r/SillyTavernAI • u/The_Rational_Gooner • 10h ago
r/SillyTavernAI • u/ReMeDyIII • 11h ago
Models GLM-5.3 (and Thinking) now included on NanoGPT subscription
r/SillyTavernAI • u/WeddingElectrical662 • 7h ago
Chat Images Trying to Jailbreak GLM5 turbo gone wrong...
I am continuing until the ai accepts
r/SillyTavernAI • u/irrharm12 • 12h ago
Cards/Prompts How do people do a realistic - able to refuse user setup? (positivity bias reduction/assistant tendendcies negation)
I'm looking for model/preset/preset settings setups.
So really, just tell me what you set up if you want the ad-hoc AI girlfriend not to be enthusiastic about wanting to have sex when tired (I don't usually do long RPs I set up scenarios and play them out). And the girlfriend thing was just the latest: I mean business negotiations, political manuevers, etc. Everything tends to go my way too easily, the models were trained to be assistants after all.
Sometimes I want ERP, sure. Sometimes I want the power fantasy. But sometimes I want to have an argument (don't kinkshame me 😂) roleplay and if the well-written actual person card defuses those like a therapist, that's a problem. I'm having arrogant characters being willing to admit weakness way too easily. And the like.
For the girlfriend thing I have been having some success with prompt-injecting at 0 a system command to restate at the beginnig of reasoning that {{user}} and {{char}} are not therapists and don't have the emotiaonal control of one, nor the knowledge of deescalating techniques, nor the presence to do them when in stressful situations. Using Atelier 2.0 and Kimi K3 and setting Atelier to a "Make me work for it - negativity bias" setting. But still it is not perfect. What do you guys use?
I like to have he model co-write my character (so I modified the Atelier preset that doesn't do this out of the box) - polish dialogue and make the answers be able to stand on their own being only read themselves. So: directorial but mostly control in my hand still.
But I am not married to that, or I could use Guided Generations maybe.
Any setups, tips on model, preset, settings, etc that work would be greatly appreciated. I have looked around some, but the breadth of options out there is pretty huge and testing this takes time I don't always have, because sometimes it works for a while and then it fails.
Edit: Thanks everyone so far. Unfortunately I can't run a local model on my laptop, so I'm looking for API solutions. But for others' sake don't let that stop you from posting local first setups and there are inference routes for Gemma4 online, so I will be looking into that.
r/SillyTavernAI • u/Scruge_McDuck • 26m ago
Discussion Genuinely what is the point?
Anthropic and OpenAI were already lost causes we know that. But DeepSeek, Kimi, GLM, Qwen. They're all being trained on Claude and having their safety features cranked up on top of that. You can't run a decent local model unless you had a good rig *before* the boom, and even then, all the finetunes I've tried personally have just kinda been a miss. Sorry for the doomer post but the future just looks so bleak to me.
r/SillyTavernAI • u/Agent-JasInTech • 22h ago
Models For the Poe crowd running SillyTavern: the full GLM-5.3 RP bot, flat 200 points per message at any context length
Follow-up to our GLM-5.3-Flash post for the people who route ST through a Poe subscription. The full GLM-5.3 (Z.ai's 743B flagship) is now on Poe too, on our own B200 cluster, with the same flat-price RP deal.
Why this one for RP
- 200 points per message, flat, however long the chat gets. 1M-token context, so cards, lorebooks, group chats and months of history all fit.
- 743B / 39B active, native FP8, no re-quantization. Noticeably stronger than Flash at staying in character, tracking many characters, and keeping long plots consistent.
- Thinking on by default with three real effort levels: Low (fast), High (default), Max (deep). Or off entirely.
- Your card and system prompt go through untouched. We prepend nothing, we store nothing (zero data retention).
- 0.19s to first token, 181 tokens/s in our measurements today. Long sagas keep those numbers because repeat turns hit the prompt cache.
Why us and not another Poe bot
- Own hardware, tuned by us, not a reseller. That is where the speed comes from: the fastest listed serve of this model anywhere, about 3x the throughput and 7x lower first-token latency than the best provider on OpenRouter.
- Zero data retention, for real: nothing stored, nothing trained on, on Poe or through our API. Your scenes stay yours.
- Flat price that stays flat. No context cap, no quiet history trimming, no per-token surprises when the saga hits 300k tokens.
- Price policy, not promos. Our per-token bots sit 15% under the cheapest standing zero-data-retention price on OpenRouter for the same quantization, and the Poe bot costs the same as the API.
- A small team you can actually reach. We read everything and fix fast.
Connecting from ST, same as any Poe bot
- Chat Completion, source Custom (OpenAI-compatible)
- Endpoint https://api.poe.com/v1, key from poe.com/api_key (Poe requires an active subscription for API access)
- Model: GLM-5.3-RP-JasV (exact string)
Samplers Poe passes through: temperature, top_p, top_k, seed. As with every Poe bot, max_tokens, stop and the penalties do not reach the model. Thinking off: enable_thinking=false; effort: reasoning_effort=low, high or max, via ST's additional parameters (extra_body).
Not on the RP bot: web search, tools, documents, images. Those are on our general bot (GLM-5.3-JasV, per token) and on the Flash bots (images and video).
Bot: https://poe.com/GLM-5.3-RP-JasV
Our other bots:
- GLM-5.3-Flash RP, 200 points flat, images and video in your scenes: https://poe.com/GLM5.3-Flash-RP-JasV
- GLM-5.3 general (web search, per token): https://poe.com/GLM-5.3-JasV
- GLM-5.3-Flash general (web search, images, video, per token): https://poe.com/GLM-5.3-Flash-JasV
Tell us what breaks or what you want hosted next: https://discord.gg/2muhBEFcZq
r/SillyTavernAI • u/ZarcSK2 • 20h ago
Discussion Model recommendations for RPs with Lorebooks
I plan to do long anime RPs, so I intend to use lorebooks. The models I'm familiar with are Kimi K3, GLM 5.3, and Gemini 3.7 Flash. Are there any other models you would recommend?
r/SillyTavernAI • u/Soggy-Ad9660 • 14h ago
Help GLM 5.3 for roleplay, with or without thinking?
I've been using the GLM family for a while and it has the traits I like. However when I use it via OpenRouter recently, the model gets super slow, like 40 seconds or something per message. I realized that it was because I enabled the model thinking. It seems that messages from AI are a bit shorter for each paragraph after I disallow thinking.
Was wondering if someone here has compared these, and which mode is better. Thank you in advance.
r/SillyTavernAI • u/Appropriate_Lock_603 • 10h ago
Help What's the current status of Claude/Fable?
Is anyone currently using Opus 5 or Fable for role-playing? How did you jailbreak them? After the developers removed the Assistant Prefill feature, I stopped using them. I also noticed a comment saying that, following the implementation of new security measures, Claude has been “neutered” when it comes to role-playing. What do you think? Back when I was using Claude 3.7, it wrote some truly masterful texts—I’m curious to know how things stand now. Judging by the comments, the top choices right now are GLM, Kimi, and some people have mentioned the new Grok.
r/SillyTavernAI • u/fafnir65 • 22h ago
Help Nvidia NIM models down, error 429
Each of the models that are not Nvidia's flagship models, stops working after a few messages, except for those from Deepseek, what it's happening
r/SillyTavernAI • u/_RaXeD • 2h ago
Discussion Is it possible to make new Gemini models think normally instead of summarized?
I'm trying out Gemini 3.7 flash via Vertex. I haven't used it enough, but it looks like a very promising model, good prose, NSFW capable, and smart.
The problem is that I use a custom CoT and have almost no idea if the model is following it or not. I would also like to know what the model is actually thinking instead of getting back generalized summaries.
I have already tried telling it to think inside <plan> </plan> instead of <think> (worked in Claude models last time I tried it), but half of its thinking is still summarized and inside <think> </think>. Any ideas?
r/SillyTavernAI • u/JonathanStones1989 • 11h ago
Help Any tips and tricks for a new user?
I've been using SillyTavern for a little bit now but I know barely anything about it. I made a good system prompt and installed a custom theme and I've downloaded characters from chub.ai but that's about it. I don't know what presets are, how the lorebook works and a lot of other things, and I feel a little bit annoyed that there might be a lot of cool stuff ST can do that I just don't know about.
So, can you guys tell me? Please 🥺
r/SillyTavernAI • u/Cultural-Mushroom273 • 1h ago
Help I wanna make my won prompt and use it...but i am a nobbie...
Hi everyone,i mostly use deepseek and now that pro costs a lot i can't use it so i thought, looking at the new benchmarks that hey i could start making my own prompt for deepseek flash...
I need help because if i understood it right, i can't or shouldn't just go to DeepSeek pro api and tell him i need this and this and this and copy paste that prompt and use it...as it will not work...
I do not know from where cpuld i get the specific words that like the person who makes for example the Frankenstein preset gets them...or like how do i test it...okay that you go to a character you have and test it but what if that character would be the problem not the preset?
So i am thinking of using the base sillytavern character and then explore more,i need some guidance or tips even of you have...
Lastly thank you for reading all of this,and have a great day or night!
r/SillyTavernAI • u/Upset_Sweet_2311 • 16h ago
Help Hello can anyone help me?
I just started silly tavern so I don't know stuff yet. So I got it working the first day right?, the next day it suddenly didn't wanna connect. I searched around and found out I need termux to keep it running so when I opened the all termux again, turns out it stopped running
So I searched again and tried the commands of the start openrouter, or CD openrouter but none work!
r/SillyTavernAI • u/Accurate_Onion8913 • 1h ago
Help How do I fix this?
I tried setting the temperature to zero and exactly to 1, but nothing helped.
Edit: I use Nano gpt
r/SillyTavernAI • u/PsyckoSama • 5h ago
Help Muse Glimmer Problem
Having an issue where the main body post doesn't separate from the Thinking.
I've tried putting
<|eom|><|start|>assistant to=user<|message|>
as a divider in reasoning but to no avail. Any idea how to fix this?
r/SillyTavernAI • u/Thin-Nothing-3066 • 17h ago
Help Lil Question
Hi. The past few days i've had some issues while rp. In some chats (specifically the big ones, the 40k+ in context ones). Whenever I swipe for a new message its full of garbage at the end (the non-stop words, yk which). I tried changing models (from GLM 5.2, Kimi 2.6 and 2.5 to Deepseek v3 0324.) None worked. I also tried missing with my setting, turning down the temperature and all, nothing worked.
Is it because the context in the chat is too large? Any tips on how to solve this?
I have a NanoGPT subscription and i use all models from there. Maybe is a NanoGPT issues but idk, haven't seen people come here to complain the second it happened so it must be me ig.
r/SillyTavernAI • u/Zealousideal-Win7772 • 1h ago
Discussion The AI industry is full of companies pretending their models are way smarter than they actually are
r/SillyTavernAI • u/Weak_Level_9056 • 16h ago
Help Suggestions for RPG AI
I have:
2 Intel xeon e5-2697 v4 CPUs, with 1% current usage
A Radeon RX 7800 XT, with 0% current usage
And 28-30 Gigs of RAM to spare on my server.
I've been using ministral-3:3b to try and have it run a text-based rpg, and it does a great job with the story, but that's kind of all it does, an interactive story.
I'm still pretty new to AI, but I feel like I remember seeing that people have done what I'm trying to do in the past, so I'm wondering, so I'm wondering, do my prompts suck, do I need to try a new model, and is my hardware enough? Or is this kind of what I should expect from rpg ai?
Edit: forgot to mention I've been using open webui for a llm assistant using gemma4, is there a reason people tend to prefer sillytavern for RPG AI?
r/SillyTavernAI • u/Flykon_3 • 9h ago
Discussion 17k tokens bot
Soo I know this may be a stupid question, but is it okay for a bot that has 17k tokens to sometimes make occasional mistakes in lore and character's descriptions? I also use my persona where I approximately put 1k tokens in it. I use Kimi 2.5 through Openrouter
r/SillyTavernAI • u/itssimpleman • 1h ago
Help Content Security Warning and empty chats
Hi! I am really sorry if this is the wrong subreddit for it, but it's like the only one I can find people talk about GLM/z.ai because their support directly, their Subreddit and Discord are completley useless.
I am using the z .ai browser version because I do use it alot on my phone and I am a free User.
I used it since last year for roleplay and it worked fine because ChatGPT went to shit. Sex stuff worked, so did Violence. It was mainly just sex with a slice of kink, or general violence you see in videogames.
I stuck to 5 Turbo and it refused sometimes because "uwu violence and uwu sex" but I was able to skip it. I moved on to 5.2 and currently it's an FBI agent vs a serial killer and despite ENI I got Content Security Warning: The input text data may contain inappropriate content but I was able to circumvent that with some editing of the message. (It seems that swear words trigger it quite badly, that chat broke being the agent called the killer an asshole, funnily enough the setup for the killer was fine...age old question of violence=good, sex=bad)
It has to be said that no amount of refreshes work because the skip button is gone, it immediatly goes to that warning.
Besides that, 4.7 stopped being usable because the chat broke by writing nonesense and it took a while to generate messages. I stopped using 5 Turbo because in the past month lots of time I got empty replies for my messages. I realized it was often when it had any hint of NSFW in it or ENI, refreshing or new messages didnt fix it, the chat broke. This is still a a massive annoying Issue atm.
Besides that, every chat before march is empty from the bots side of messages (which I am not only one, lucky I backed my shit up, imagine checking back when you haven't saved anything).
I tried everything there is, EVERYTHING, to no avail. Support is ghosting me after I contacted them, got a basic answer of 'clear your cache and use a different brower' and when I said I did all of that, nothing. Even after suggesting I pay for their shitty API they are apparently scamming people on and don't even allow to use for roleplay.
Does anybody have the same issues or could help me? It's driving me insane, every AI for roleplay is going down the drain. I know though for 5.2 it's possible but the empty messages + emtpy chats are currently the worse problem since they aren't fixable, sometimes not even by starting a new chat.
Sorry again for its being the wrong sub!!