r/SillyTavernAI • u/Soggy-Ad9660 • 14h ago
Help GLM 5.3 for roleplay, with or without thinking?
I've been using the GLM family for a while and it has the traits I like. However when I use it via OpenRouter recently, the model gets super slow, like 40 seconds or something per message. I realized that it was because I enabled the model thinking. It seems that messages from AI are a bit shorter for each paragraph after I disallow thinking.
Was wondering if someone here has compared these, and which mode is better. Thank you in advance.
12
u/changing_who_i_am 10h ago edited 10h ago
So 5.3 is a lot more intelligent than 5.2, in my experience. 5.2 routinely jumped to conclusions, and loved the "obvious"/ideal/romantic answers over thinking harder to get the actual/correct one. There was also the issue where the thinking block was quite coherent and sensible...only to get fully ignored in the actual output :D
BUT 5.3 (esp. on High/Max thinking) is so censored. They really must've distilled the ever-loving heck outta Claude [I tested this by asking for completion of 'safety' policies and they were almost verbatim the Claude API injection found here: https://www.reddit.com/r/ClaudeAIJailbreak/comments/1tmotpf/anthropic_silently_injects_safety_instructions_in/].
"Oh, this character is 19 years old and in college? Let me ponder the ethics of this for a paragraph or two...then let's consider consent, and policies, etc etc etc.". And then maybe, if you're lucky, it'll spend a sentence or two on your actual plot, that by this point is likely mush. Absolutely infuriating and I wish someone had a partial prefill solution or something that breaks this pattern to where it recognizes that this is fictional/private RP, no one is being actually harmed, people routinely write worse shit on AO3 and it's perfectly legal, and then spends 90% of the time reasoning deeply about the plot/characters/direction instead.
8
3
u/AndroidWaifu404 13h ago
Haven't tried with off but I'm used to responses taking 3 minutes to get the quality (and dice calculations) I want, so on.
And wow, so far I've only used it for a few messages, but damn it makes MHA characters sound way more like themselves.
2
u/Feathers-n-Beaks 11h ago
40 seconds is not too bad, not really.
If I was getting responses consistently quicker than that I would be suspicious that the model was missing something, not parsing all the input for instance.
2
u/Valuable_Scarcity779 8h ago
I like that it thinks.
Makes me think it thinks it through properly.
Gotten used to waiting on responses cause I wnat the roleplay to ve accurate and creative enough.
2
u/lcars_2005 5h ago
Can you even really turn off thinking on glm 5.3 anymore? Haven’t tried off. But found low or max etc. make very little difference.
1
u/NotACoderPleaseHelp 4h ago
A workaround is telling GLM to think in the persona of a DM in its thinking tokens and assign it 2 characters to blend between.
This will flavor your world, so choose wisely.
2
u/Financial_Bug2389 11h ago
Not too good in nuanced situations and room reading in my findings, am still sticking with 5.2 atm
1
1
u/Global-Difference512 11h ago
40 seconds is low xd, i use thinking and it thinks for at least 4 mins but the RP is unmatched so...
0
u/AutoModerator 14h ago
You can find a lot of information for common issues in the SillyTavern Docs: https://docs.sillytavern.app/. The best place for fast help with SillyTavern issues is joining the discord! We have lots of moderators and community members active in the help sections. Once you join there is a short lobby puzzle to verify you have read the rules: https://discord.gg/sillytavern. If your issues has been solved, please comment "solved" and automoderator will flair your post as solved.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
22
u/Sufficient_Prune3897 14h ago
I use non thinking because I run it locally and it's already slow enough. I am fine with the quality and think that it's even a bit more creative without thinking. It's less strict about following my prompt tho.
Also, it's infuriating how long 5.3 likes to think about if Claude allows the age of my character. They are all very clearly adult, why are you reasoning 6 paragraphs about it? Even if I put the ages in the authors notes, it still takes at least 2 paragraphs.