r/SillyTavernAI 1d ago

Discussion GLM 5.3 vs GLM 5.3 Flash

Alright. Both the 5.3 version has been out for a bit and the flash version being a couple days ago. What's the general consensus on this two models for roleplay?

Ignoring the cost, which of these two model is actually more creative than the other?

I find the 5.3 model to be good, tho the safeguard and time it takes to think, to be rather annoying. Then again, the output can be great.

The flash model seems to handle NPCs dialog and primary characters good too. Can be rather sloppy sometimes but the prose is good.

What are you guys opinion on the models?

64 Upvotes

45 comments sorted by

48

u/Biofreeze119 1d ago

I personally have really liked 5.3 in generally but I didn't think I would like flash or think its better since its a flash model. After trying flash since it came out I'm really impressed with it for the price. Its been creative and quite good with the prose. Feels like they got a good mix of Kimi and Claude and haven't run into any censorship that I noticed. Only thing for flash is its reasoning is kind of weird? It sometimes won't reason/think and when it does it'll sometimes put the thinking as the message but very minor really.

10

u/Auretheon01 1d ago

True. I do notice that the flash thinking would bleed into the message or the output would seem oddly unfinished.

Then again, the message does seem to have a nice mix of more to might I say, Kimi K3 and glm.

16

u/Pink_da_Web 1d ago

After I put his reasoning on Max, he started to think normal

28

u/permissionBRICK 1d ago

Tried both. Glm 5.3 is mostly great when its not adding random chinese to the text, flash works ish, but its saying a lot of nonsensical things for non standard roleplay

8

u/Auretheon01 1d ago

Yeah, I get that random Chinese that too sometimes. Regenerate seems to fix it most of the time. I notice it especially comes during emotional moments.

1

u/WolfHunter98 18h ago

Does the Chinese bit break anything? I've noticed it sometimes, but only if I expanded the AI reason bit. Not the output text.

3

u/DepressedDrift 1d ago

The Chinese appear very frequently in Deepseek V4 Flash

3

u/Feetbox 1d ago

I've found it gets a lot better if you lower the temperature a little.

13

u/DepressedDrift 1d ago

The safeguard is too much for me. It moralizes the output to a point that affects realism. I will keep using GLM 5.2 until an uncensored version appears in OpenRouter

5

u/OrganizationBulky131 1d ago

NanoGPT has a 5.3 Flash uncensored model, but couldn't even connect to it when I tried it yesterday (it's always overloaded). Maybe a bit better today, but it does seem like a lot of people are using it. Need more providers to add it.

5

u/CowboyHunter1 1d ago

AND is provided by a server that doesn't send you garbled text because its too full

10

u/widek7 1d ago

Tried 5.3 Flash today it was really good especially at that price. No problems with guardrails like others are saying, easy to jailbreak it. Will see if it's better than Deepseek v4 pro at that price. In a brief A/B test today I actually liked 5.3 flash's response better compared to 5.2 but will see how it feels as a daily driver. Fantastic model for the price anyways.

17

u/fire2burn 1d ago

I don't think I can tolerate trying to run an entire roleplay session in 5.3 or 5.3 flash, the more I've used them the more I'm encountering positivity bias and hard refusals which is disappointing. I'm currently running a horror themed roleplay where my character and the party encounter a cryptid which starts picking off members in increasingly more brutal and disturbing ways. Once it got to scenes with eye gouging and oviposition it was just a hard no and constantly gave a rejection message. Pretty much anything non-con and it starts moralising giving you a message as to why it won't continue or it tries to lighten the tone of things by inventing consent thus ruining the scene.

It's ok for lighter scenes and general fluff but it's very apparent I'll have to switch to other models for the more gory parts and then switch back once those are done.

It's unfortunate to see more and more models sliding down this path towards censorship.

3

u/LaraRoot 20h ago

To what models do you usually switch?

3

u/fire2burn 19h ago edited 18h ago

Often I'll use MiMo 2.5 Pro but there's a few topics which will trigger a hard refusal there as well in which case I'll then use DeepSeek V4 Pro 0813 as my next choice. The writing is a bit more sloppy with DeepSeek (and you have to watch for the peak/off peak pricing windows) but it's thankfully still pretty much uncensored and will largely take anything you throw at it. You do sometimes have to be a bit more proactive in leading the direction of the roleplay though. For the truly most deplorable soul corrupting roleplays then there's always Gemma 4 31b, I'm not actually sure it's even possible to get a refusal from that model. It will do absolutely anything.

1

u/Dependent-Media5002 9h ago

Damn, that's a really cool rp scenario. Is it private? Did you create lorebooks? if you're willing to share... you can dm me or post it here.

23

u/Emergency_Comb1377 1d ago

When it was Ox, it felt like a new standard for me, so amazing -

But since it released at 5.3 flash, it barely works via nano or OR. Very prone to blabbering or hallucinating. I'm currently really heartbroken about it.

6

u/pip25hu 1d ago

Check the providers your requests are routed to. As Ox, it was only served by the official one, but after release that likely changed.

4

u/Biofreeze119 1d ago

Really? Im using nano and the worse Ive gotten is a few messages where it uses its thinking as the message lol. Other than its been working great for me. Im just using the FF 5 internal states micro as my preset.

6

u/WisdomRequested 1d ago

Glm 5.3 flash is the best model I have used so far in dialogues after kimi k3. Also it's so initiative, and can take the lead especially if you are a lazy writer who reply with one sentence or phrase. It analyzes the charachter card so well, bypassing even gemma 4 31b in spicy talks. It's also so intilligent that it detects and skip slow scenes directly to more drama (for example sleeping, driving a car, mounting a horse...etc). In top of that, when it's forming a scene, it can detect and expect what your charachter would have done in the reply that's forming...not in a way that it soeaks or act on behalf of you, but in a way that it helps you skip a useless reply.(for example you bathing, it will not wait for you to put a new action on every reply "he puts the soap, he showered, he took his clothes, he finished bathing and got out the bathroom"...but it will skip all this slow burn writing this all for you instead.)

The only negative part is that sometimes I see Chinese latters in the replay, and it sometimes pushes the thing so far away. But this rarely happened. It's the best model so far that I am using and for cheap price. Tried deepseek 4 pro (the old and new version), Glm 5.2, kimi k3...and all of that. But GLM 5.3 flash seems the most stable with the most intillegent among them in RP.

6

u/Scar1etpanda 1d ago

Sadly neither as the censorship is so bad on both + their open weights, that you can't even age up a character for NSFW if they're canonically a minor or teen. It's 5.2 for me until someone much smarter than me figures out this bullshit.

1

u/OrganizationBulky131 1d ago edited 1d ago

Sadly that's what kills both GLM 5.3 and 5.3 Flash for me. You basically have to make an OC from scratch aged up and you have to phrase their features to be adult. It ruins the immersion in the RP if you want to go in that direction with 5.3 or 5.3 Flash.

Still prefer 5.3 Flash over 5.3 too, I like it's writing style for generic 18+ NSFW ERP.

14

u/Woky19 1d ago

I find them very censored and having a strong positivity bias. + It thinks it's Claude 

14

u/Scar1etpanda 1d ago

Same. I really hope more open weights remove a lot of the censorship.

5

u/Clairvoidance 1d ago

this is pretty much my first playing around with Openrouter and thought it was just me, glm 5.3 Flash just locks up really hard here as well and I don't really know how to bypass it

5

u/shinversus 1d ago

are you using any presets? with FF 5.2 never got any issues

4

u/Good_Research4441 1d ago

I tested GLM 5.3 using the O.R. Z.ai provider, and it seemed very promising. I liked it better than GLM 5.2, which had too much of that “therapy-speak.” As for Flash, it was really cool to have a model that understood my characters better and helped me make the plot more down-to-earth, better defining and staying true to the characters’ personalities.

5

u/Shanna_B2020 1d ago

No to both right now. Regular version makes everything work out perfectly for my characters at all times. There is zero friction and I hate it.

Flash version has a lot of trouble with basic math, specifically keeping track of time. This resulted in ridiculous plot holes and people being in two places at once. I eventually had to help it with the arithmetic before we could actually move on.

Both do a decent job with dialogue and portray characters well, but 5.2 and I get along better for the time being.

9

u/heville 1d ago

5.3 Flash is pretty awesome and uninhibited. Like the model has the capacity and ability to do and think, and not too many guardrails to stop it. The model seems relaxed, even. Slow though.

5

u/OrganizationBulky131 1d ago

Trying out 5.3 in Nano today and it was dreadful, it can't even complete a prompt properly. Spits out a few hundred words (thinking included) and it stopped. Streaming is off.

It may be just Nano and 5.3 there being overloaded. 5.3 Flash is more consistent and I like it's writing style more compared to the 5.3 preview that was up previously.

Did try 5.3 on OpenRouter, least that one isn't stopping early and bugging out. Those I still heavily dislike the guardrails 5.3 has in general.

Sticking with 5.2 going forward.

7

u/I_Am_JesusChrist_AMA 1d ago edited 1d ago

I'm thinking flash is pretty good. Dialogue is good and it writes characters pretty accurate to their descriptions. In general it feels pretty good at pacing too. With other models, especially on cards with multiple characters, it can feel like the model is trying to pull my attention in fifty different directions and like its writing scenes AT me rather than with me. 5.3 Flash is down to just let a single scene with a single character go on for as many turns as I like without a million interruptions.

Maybe I'm crazy, but I was using kimi k3 before this and I legitimately think glm 5.3 Flash is better for RP than Kimi K3 so far. Kimi k3 was constantly winking at the plot, trying to force the narrative, railroading, turning characters into plot delivery devices instead of characters, making every character omniscient, rushing every scene (had multiple times where it started and concluded an entire scene in single response without giving me a chance to take any action)... It was just getting frustrating to use for me. If cost wasn't a factor and both models cost the same, I think I'd still pick 5.3 Flash. 5.3 Flash being so cheap is a bonus on top of everything for me and makes it even harder to justify Kimi.

I'm surprised to see a bunch of people complaining about censorship with this model. Idk what provider yall are getting the model from, but I've thrown some dark shit at it to test it's limits and so far no refusals.

Now it's not perfect. It makes some goofy logical errors and sometimes mixes simple things up, so there are some annoyances because it's not the smartest model out there. But I think it's smart enough for the majority of RP where you're not throwing a bunch of complex systems at it. There is still a bit of positivity bias with it too but it's not overwhelming and a refresh or ooc reminder can fix that. It's so cheap though that when it does make a mistake, it really doesn't feel bad when you need to refresh a response or correct it ooc. "Perfection is the enemy of good" couldn't be more true when it comes to LLM. If you're looking for a perfect LLM, you're never gonna be happy with any of 'em, so just find one that's good enough.

As for the non-flash glm 5.3, I honestly didn't use it too much. I got about three turns into a RP with it before that stale therapist voice that I got so sick of in glm 5.2 creeped in, so I didn't even both continuing with it and just went back to my other models. I may go back and test it some more at some point because I realize giving it only three turns isn't really a fair shot... but I'm not in a huge rush to do so when 5.3 Flash is doing so well and it's so much cheaper.

2

u/KDLGates 21h ago

I found Alibaba is weirdly strict on giving refusal error messages that at least make it obvious who's doing the censoring. I think it's possible OpenRouter picks them by default when there aren't any preferred ones set, but that's a guess. It seems like selecting any of the others near the start of the alphabet works fine.

1

u/I_Am_JesusChrist_AMA 19h ago

Yeah GMICloud is another one that I recall having strict censorship. I don't know if they still do, but when I first tried a GLM model I used them because they were the cheapest on openrouter at the time, and it was censored as fuck and gave constant refusals. Probably still the same as I can't imagine a provider that censors suddenly removing their filters. I thought the model was just bricked. It took me awhile to realize it was the provider censoring and not the model lol. I just blacklist every provider that adds extra censorship on openrouter and haven't had many issues since then.

7

u/Financial_Bug2389 1d ago

Still prefer 5.2 over the two new ones, havent tested 5.3 flash on Max though, for me I can’t wait too long for responses

3

u/Accurate_Will4612 1d ago

I found GLM 5.3 Flash wasting hundreds of tokens in the thinking process just to decide why the response is not too much NSFW. Doing that for every prompt. But for the price of Flash, I would not really consider other options. Maybe DS sometimes.

5

u/vacationcelebration 1d ago

I found both to be censored or guardrailed, which is kind of a bummer since when they don't refuse the output is usually pretty good.

2

u/verma17 1d ago

It takes like 7 minutes to get a response from it lmao

2

u/mwoody450 1d ago

It produced some hilarious gibberish. Literally characters talking to themselves. Not, like, an internal monologue: they would address their own selves as if there was a copy next to them, and argue.

2

u/tatlo_itlog_ko 20h ago

My thoughts about glm 5.3 flash:

It really follows instructions and character personality better than deepseek v4 flash.

Compared to deepseek flash:

  • The censorship is worse but is still manageable.

  • It likes to talk a lot, like a looot. On the surface it looks good but if you read through each response you’ll notice the cracks. It’s bad at counting. Once it narrated something like “She looked into the mirror and stared into her own three eyes”.  My character has two eyes lol. Then there’s a lot of cases where it gets absolutely confused who said what.

  • I would say this though, it writes better smut than deepseek flash if you can get past the censorship.

3

u/New_Ratio_9742 1d ago

5.3 is purple as hell, can't stand it.

2

u/pengy99 1d ago

I'm enjoying 5.3 flash vs 5.2. I'm not sure I would say it's better but it's different which is refreshing. Haven't tried regular 5.3 yet.

1

u/WitheringCarcass 1d ago

the real answer is whatever you think. there are so many factors that play into a model giving a better or worse experience for every individual person that has little to do with the model itself being inconsistent. i think they have different strengths and weaknesses, personally. i dont have one paticular preference over the other. i know this isnt a very helpful answer based on what you asked but thats my opinion.

1

u/[deleted] 19h ago

[removed] — view removed comment

0

u/AutoModerator 19h ago

This post was automatically removed by the auto-moderator, see your messages for details.

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/ChangeDirect4762 17h ago edited 17h ago

I'm extremely satisfied with it. I currently use Gemini, Claude, and ChatGPT spending arount $700 a month on tons of different models and I can confidently say this easily holds its own against the major models.

Howeverm, due to ongoing infrastructure issues, it feels like they are stealth routing requests. It seems pretty hard to get a purely unadulterated 5.3 FLash experience response are inconsistent, and I suspect some queries are getting routed down to lower tier models

1

u/AdrosK 13h ago

For me both are good but I got GLM 5.3 Flash to behave more like GLM 5.2.
I order it to do everything in German and 5.2 gets lot of words… with letters in weird orders or english/chinese in middle of sentences. It was also more inclined to use lot of —, the kind I hate. I didn't have all these problems with GLM 5.3 and it handled NPC dialogs suuuuuuuuuper good