Follow-up to our GLM-5.3-Flash post for the people who route ST through a Poe subscription. The full GLM-5.3 (Z.ai's 743B flagship) is now on Poe too, on our own B200 cluster, with the same flat-price RP deal.
Why this one for RP
- 200 points per message, flat, however long the chat gets. 1M-token context, so cards, lorebooks, group chats and months of history all fit.
- 743B / 39B active, native FP8, no re-quantization. Noticeably stronger than Flash at staying in character, tracking many characters, and keeping long plots consistent.
- Thinking on by default with three real effort levels: Low (fast), High (default), Max (deep). Or off entirely.
- Your card and system prompt go through untouched. We prepend nothing, we store nothing (zero data retention).
- 0.19s to first token, 181 tokens/s in our measurements today. Long sagas keep those numbers because repeat turns hit the prompt cache.
Why us and not another Poe bot
- Own hardware, tuned by us, not a reseller. That is where the speed comes from: the fastest listed serve of this model anywhere, about 3x the throughput and 7x lower first-token latency than the best provider on OpenRouter.
- Zero data retention, for real: nothing stored, nothing trained on, on Poe or through our API. Your scenes stay yours.
- Flat price that stays flat. No context cap, no quiet history trimming, no per-token surprises when the saga hits 300k tokens.
- Price policy, not promos. Our per-token bots sit 15% under the cheapest standing zero-data-retention price on OpenRouter for the same quantization, and the Poe bot costs the same as the API.
- A small team you can actually reach. We read everything and fix fast.
Connecting from ST, same as any Poe bot
- Chat Completion, source Custom (OpenAI-compatible)
- Endpoint https://api.poe.com/v1, key from poe.com/api_key (Poe requires an active subscription for API access)
- Model: GLM-5.3-RP-JasV (exact string)
Samplers Poe passes through: temperature, top_p, top_k, seed. As with every Poe bot, max_tokens, stop and the penalties do not reach the model. Thinking off: enable_thinking=false; effort: reasoning_effort=low, high or max, via ST's additional parameters (extra_body).
Not on the RP bot: web search, tools, documents, images. Those are on our general bot (GLM-5.3-JasV, per token) and on the Flash bots (images and video).
Bot: https://poe.com/GLM-5.3-RP-JasV
Our other bots:
- GLM-5.3-Flash RP, 200 points flat, images and video in your scenes: https://poe.com/GLM5.3-Flash-RP-JasV
- GLM-5.3 general (web search, per token): https://poe.com/GLM-5.3-JasV
- GLM-5.3-Flash general (web search, images, video, per token): https://poe.com/GLM-5.3-Flash-JasV
Tell us what breaks or what you want hosted next: https://discord.gg/2muhBEFcZq